Live Validation Data — Updated Weekly via CT.gov

Prediction Accountability Dashboard

Every drug repurposing prediction GaiaLab generates is timestamped, assigned a confidence score, and automatically cross-referenced against ClinicalTrials.gov. This page is the unfiltered record.

Claims without calibration data are marketing. Here's ours. · Full calibration methodology →

🔬

Prospective validation evidence — building in real time

Each analysis run adds to this dataset. As predictions accumulate and ClinicalTrials.gov checks complete, calibration curves populate here automatically. No other biological intelligence platform publishes this data continuously — that is the standard we hold ourselves to.

Pre-registered Prospective Validation Protocol
We pre-register how forward predictions are judged — before outcomes are known — so the claim is falsifiable and cannot be tuned to a result. The registration is publicly timestamped with a DOI, so a third party can confirm it predates any outcome.
Public pre-registration (DOI): 10.5281/zenodo.21420447 — sealed 2026-07-17, indexed in OpenAIRE.
Public record: DOI 10.5281/zenodo.21420447 · protocol: docs/PREREGISTRATION_drug-repurposing-2026-07-17.md · power & blinded-adjudication amendment (2026-07-25): docs/PROSPECTIVE_ADJUDICATION_AND_POWER.md (to be published as a new DOI version).
Full Prediction Ledger
Every therapeutic candidate GaiaLab has scored — timestamped, unfiltered, and continuously cross-referenced against ClinicalTrials.gov. All entries in this name-matched ledger are retrospective benchmarks. Semantic matching has separately surfaced 15 prospective matches (2 off-label repurposing hypotheses + 13 on-label concordance) — see /new-trials.
Drug Disease Context Confidence Outcome Trial NCT IDs Recorded

Benchmark Summary · Last updated 2026-07-03

Every headline metric, with its scope and known limitation. Retrospective figures test correspondence with already-existing evidence; prospective figures — the only fair test of forward predictive ability — are pending. We publish these whether or not they flatter us.

Benchmark Dataset Type Metric Known limitation Readout
Drug repurposing — retrospective AUROC N=529, 22 disease areas Retrospective 0.545 AUROC Modest signal just above the 0.50 random baseline; not a clinical predictor. Complete
Drug repurposing — temporal holdout AUROC n=22 curated approvals, 8/8 neg. controls Retrospective 0.90 AUROC Small curated set; a different, easier test than the retrospective AUROC above. Not a clinical predictor and not a prospective result. Complete
Drug repurposing — prospective forward skill Sealed cohort; live at /new-trials Prospective Pending The genuine test. Pre-registered; 0 confirmed forward hits so far — reported honestly. 2028 (drug lockbox)
Calibration — Brier score N=608 resolved predictions Retrospective 0.08 (vs 0.25 no-skill) Below 0.25 partly reflects class imbalance, not proven calibration — read WITH ECE (~0.40). Live
LassaAI — national early-warning AUROC 313 weeks / 6,456 cases (2020–2025), 2024 hold-out Retrospective 0.880 (vs 0.849 naive) Beats a naive baseline by only ~0.03 — not statistically significant (DeLong p=0.61). Seasonality-driven. Preprint: Zenodo. Prospective 2026–27 season

Trial correspondence is not efficacy. A registered or completed ClinicalTrials.gov trial for a drug–disease direction indicates that direction is being pursued in research — it does not establish that the drug is an effective treatment, nor that GaiaLab predicted it before the trial existed.

Methodology & Definitions

How predictions are recorded: At the end of every analysis, up to 10 drug candidates are saved with their confidence score, disease context, target genes, and a timestamp. Records are immutable — no retroactive changes.

Trial-correspondence check: Each prediction is queried against ClinicalTrials.gov API v2 (clinicaltrials.gov/api/v2/studies) using the drug name + disease context. We use precise labels: completed-trial correspondence (a completed disease-matched trial exists — the internal status code is validated, but this means a matching trial was found, not that efficacy is proven); active-trial correspondence (an active/recruiting trial exists — direction pursued, outcome pending); and no trial correspondence identified (none found — counts against the correspondence rate). None of these establish treatment efficacy or forward-prediction skill.

What this is NOT: This does not measure whether GaiaLab identified the drug before the trial started (we don't have that date information). It measures whether the drug+disease direction is being/was pursued in a clinical setting — a proxy for research relevance, not therapeutic efficacy.

Calibration curve: A well-calibrated system shows higher-confidence predictions matching trials at higher rates than lower-confidence ones. This is research direction correspondence, not therapeutic outcome accuracy. A true efficacy calibration curve requires prospective trial completion data — that data will be added as it matures. We publish what we have, not what looks best.

Data source: GET /api/predictions · GET /api/predictions/calibration — public, no auth required.