Live Evaluation

Benchmark Dashboard

Real-time scientific quality scores updated after every analysis run. All metrics are computed from live data — no cherry-picking.

Loading scores…
Live Scores

Scientific Quality Metrics

Four core quality signals measured on every run and reported here from recent saved analysis records — a limited sample, since not every run persists a durable snapshot, so treat these as directional. Green = meets bar, amber = marginal, red = below threshold.

Citation Coverage
Snapshot coverage — % saved insights with ≥1 PMID
Grounded Ratio
% claims citation-verified (PMID)
Cite-F1
Citation precision × recall
PMID Hallucination
% fabricated citations (lower = better)
Historical Trend

Quality Over Time

Last 30 analysis runs. Each point represents one analysis session.

Score Trend

Citation coverage, grounded ratio, and confidence index across recent runs
Citation Coverage Grounded Ratio Confidence Index
Analysis Volume

Platform Activity

The first three cards report distinct platform counts: analyses processed by the learning engine, durable saved snapshots, and configured data sources. The fourth card is a separate canonical engineering-proof state — not a biological benchmark result.

Analyses Learned From
Saved Snapshots
76+
Data Sources
External benchmarking

GaiaWorld: auditability under a frozen evaluation contract

GaiaWorld tests a separate biological-prediction workflow by freezing the evaluation boundary, lanes, primary metric, and prediction hashes before one-time scoring. It publishes negative and positive results. This is evidence about an evaluation process—not a clinical, causal, safety, efficacy, or treatment claim.

Reproducibility

Verifiable, Not Cherry-Picked

GaiaLab's platform code is proprietary, but its scientific claims are held to a public standard: every metric on this page is computed on live data after each analysis, and forward-looking predictions are pre-registered and hashed before outcomes are known — so results can be independently checked without access to the engine.

# Prospective validation — pre-registered, tamper-evident
# Drug-repurposing study sealed 2026-07-17, readout 2028-01-17
#   357 Tier II+ predictions (no trial at lock) vs 148 matched controls
#   cohort hash: daaed957d2d38171b6a626b11d2a6df48a6262165fc55c3e685b1c3a64bfb96b
# The hash is a one-way fingerprint of the frozen cohort. At readout, anyone
# can confirm the prediction set was fixed in advance — no backdating possible.

We publish hits and misses. A refuted prediction is a valid, reported result. Independent timestamped pre-registration records are posted publicly (Zenodo); the outcome readout will be published on the evaluation date regardless of what it shows.