GaiaWorld
A machine willing to publish when it loses.
GaiaWorld freezes biological prediction experiments before outcome reveal, scores them against declared baselines, and keeps negative results visible. Experiment #001 is scored and published as a negative primary result. Experiment #002 and Experiment #003 are scored and published as narrow positive primary results on distinct retrospective held-out boundaries.
Arc H1 scientific flight recorder
Retrospective held-out evaluation on a pre-frozen 100-gene H1 hESC CRISPRi panel. These are GaiaWorld panel metrics, not official Arc leaderboard metrics.
Arc H1 hESC CRISPRi
Verification record
Target-ridge performance
Negative on the primary metric. Mixed only in a secondary diagnostic.
The baseline beat the model on magnitude accuracy.
Lower panel MAE is better. The target-ridge model scored 0.13257 versus 0.07616 for no-change and 0.08195 for empirical mean. It won only 11% of cases against no-change and 8% against empirical mean.
The target-ridge lane did show a panel discrimination diagnostic of 0.5332 versus about 0.505 for both baselines, a +0.0282 difference. That measure was secondary, adapted to the frozen 100-gene panel, and is not official Arc PDS.
Abstention was preserved instead of converted into fake precision.
The separately audited direct-evidence route admitted zero Arc-compatible transition effects. It therefore remained at 100/100 abstentions and zero coverage rather than fabricating expression-delta magnitudes from generic literature confidence or association signals.
Experiment #001 is closed to tuning.
The revealed holdout can now inform diagnosis of failure, but it cannot be reused as untouched evidence for a revised model. Experiment #002 must use a genuinely new evaluation boundary and freeze its architecture, hyperparameters, metrics, baselines, prediction hashes and scorer before reveal.
Positive on the primary metric. By a narrow, fully-disclosed margin.
The redesigned model beat both baselines — narrowly.
On a genuinely new boundary (Replogle K562 CRISPRi, Gate 0b hidden 200-case partition committed before download), the frozen baseline-residual-ridge model scored panel MAE 0.02498 versus 0.02523 for empirical-mean and 0.04705 for no-change — strictly better than both, meeting the pre-written success criterion. Against the stronger baseline it won 35 cases, lost 9, and tied 156; the top 5 winning cases contribute about 47% of the gain.
The margin is ~1% relative over the stronger baseline, concentrated in the cases whose perturbation target is present in the measurement matrix. The panel discrimination diagnostic was 0.5123 versus 0.5025 for both baselines (+0.0098) — secondary, non-overriding, and not official Arc PDS. The evaluation partition's outcomes are public since 2022: "untouched" here means process-hidden under a pre-committed hash rule, and the record says so.
Positive on the primary metric. A separate RPE1 holdout.
The frozen ridge lane beat both declared baselines.
On the RPE1 CRISPRi 200-case declared hidden partition, the frozen baseline-residual-ridge lane scored panel MAE 0.23221 versus 0.23818 for empirical mean and 0.29546 for no-change. The relative margin over the stronger baseline is 2.5%.
The scorer reports 33 abstentions. This is a one-time, retrospective, held-out 100-gene panel evaluation with a pre-committed split and frozen predictions. It does not establish general biological prediction ability, causal validity, patient benefit, safety, efficacy, or clinical utility.
Scientific claims should increasingly arrive with receipts.
A result is more useful when the path to it can be inspected. GaiaWorld records the parts of a benchmark that can otherwise be quietly rewritten after outcomes are known:
What was frozen, when it was frozen, which code version ran, which data were available, and which baselines and primary metric were declared.
What changed, why it changed, what was excluded, and—where applicable—who reviewed the eligibility boundary.
What the system predicted before outcome reveal, what happened afterward, and the raw artifacts needed to check the account.
These receipts do not make a scientific result true, clinically useful, or independently validated. They make the process more inspectable and make it harder to quietly rewrite history after seeing the outcome.
Inspect the record. Challenge the method.
GaiaWorld's published experiments are self-published retrospective records, not externally validated findings. We invite independent reproduction, hash verification, and critique—including findings that identify limitations or discrepancies.