Methodology & corpus selection
An adversarial test we run on ourselves.
The audit exists to answer one question honestly: does Genomarker find real failures in real published science — and pass the work that holds up?
Study selection
Nine published gene-expression studies, 2004–2022, each with public data and a computational methods section, and none involving Genomarker. Round 1 deliberately mixes calibration cases — three papers later retracted, a signature that failed in phase III, a canonical confounding case — with highly cited studies and a clean counterexample.
Spec reconstruction
Each paper's methods are transcribed into a canonical analysis request exactly as published — ambiguities included. Where the publication under-specifies, the ambiguity itself becomes a finding, not a silent guess.
Re-execution
The reconstructed analysis runs on our production system, through the same checks every user's analysis gets. Nothing bespoke, nothing softened.
Classification
Findings are classified by severity — hard refusal, blocking, warning, confirmation — and by subject: the study, or Genomarker itself. A blocking finding stops the analysis until someone signs a justification; a hard refusal stops it outright.
Disclosure ethics
Aggregate results are public. Per-paper findings are shared with the original authors first, anonymized in public reporting, and never framed as accusations — the audit tests Genomarker, not the scientists.
Versioning
This is round 1 of a living audit, run on production in 2026. Round 2 re-runs the corpus through the replay engine when it ships, and every revision is archived.
SELECTION FIXED BEFORE RE-EXECUTION · REVISIONS ARCHIVED