Skip to content

Independent verification for AI-generated science

Science is about to be produced faster than it can be trusted

AI agents now run entire analyses unattended, and their results are rarely checked by anyone independent. Genomarker checks them: it refuses what the data can't support, and signs the rest as evidence anyone can verify.

Catch Seal Replay

Genomarker Assurance StatementGM-2026-08-14-8F29E1Example

Analysis
Differential expression
Assurance profile
GM-BIOMARKER-DE-1.3
Input integrity
PASS
Study design
PASS
Statistical methodology
PASS · DISCLOSED FINDING
Confounder control & leakage
PASS
Plan alignment
PRE-SPECIFIED
Execution & provenance
VERIFIED · COMPLETE
Independent replay
REPRODUCED_EQUIVALENT
Long-term reproducibility
GRADE A
SIGNED · VERIFIABLEgenomarker verify …8F29E1.gmk
Fig. 1 — An example Genomarker statement. It says exactly what was checked, what passed, and what it does not establish.

The problem

Results are trusted long before anyone re-runs them.

Of 27,271 code notebooks linked to biomedical papers, just 879 re-ran to their originally reported results (GigaScience, 2024). AI agents now multiply the volume — and every failure mode with it:

  1. Methods that were never defensible

    The wrong statistical test, an unsupported design, a violated assumption — chosen silently and reported confidently.

  2. Confounding and leakage no one caught

    Batch tracks treatment; test data leaks into training. The result looks strong precisely because it is wrong.

  3. Analyses that drifted from the plan

    What was pre-registered and what was run quietly diverge. Post-hoc becomes indistinguishable from pre-specified.

  4. Provenance that says what, not how

    Parameters unrecorded, software versions lost, the exact input data no longer identifiable.

  5. AI decisions nobody can inspect

    An agent made an inferential choice mid-analysis. The agent is gone. The justification never existed.

  6. Results that expire with their software

    Rerunnable today, unreproducible in five years. The environment decays; the claim stays in the literature.

COMPUTATION

Scientists, pipelines and, increasingly, AI agents — producing more analyses, faster and cheaper than ever.

GENOMARKER

The independent check: catch what the data can't support, seal what it can, replay it on demand.

SCIENTIFIC CLAIM

Papers, biomarker programs, investment decisions and regulatory filings that depend on the computation being right.

The proof

We audited published science we had no part in.

Nine published gene-expression studies, re-run end to end on our production system with data and methods as published. We report what we found in the studies — and in our own platform.

9

Published studies audited

re-run on our production system

7

Re-executed end to end

1 refused with a stated reason · 1 not auditable from its deposit

57

Structured findings

1 hard refusal · 11 blocking · 39 warnings · 6 confirmations

1

Hidden confound caught — analysis refused

recorded in no metadata field; read from the raw instrument files

THE CATCH

Hidden in the raw files.

In a large public patient study, none of the 48 days on which samples were scanned held both patients and controls. No metadata field recorded it. Genomarker read the scan dates from the raw instrument files and refused the comparison: the data cannot separate the biology from the scan day.

A statement about what the public deposit supports — not about the published conclusions.

See the full audit

How it works

Catch. Seal. Replay.

One engine runs every check, whether a scientist, a pipeline or an AI agent submitted the analysis.

CATCH

Refuses what the data can't support.

Before anything runs, Genomarker checks whether the data can answer the question: batch confounding, pseudoreplication, the wrong method for the data type, deviations from a pre-registered plan. It refuses or flags — and says why.

Live

SEAL

Signs what it can.

Inputs, code versions, parameters, rules and results are recorded and signed against published keys, so anyone can check the record without trusting us.

LivePortable packages — Q1 2027

REPLAY

Re-runs the work independently.

We re-executed seven published analyses end to end on our production system, and our core statistics match independent R reference implementations.

Proven on 7 studiesScheduled replay — mid-2027

Read how it works

Who it's for

Built for the people who carry the risk.

For investors & BD teams

Genomarker Diligence

Independent review of the computational claims behind a financing, licensing deal or trial decision — a written verdict, claim by claim, before you commit.

  • Life-science investors
  • Pharma BD & licensing
  • Translational reviewers

For research teams

Genomarker Runtime

Run supported analyses with the checks built in. Every run produces a signed record, so your results are ready for a reviewer, a partner or a data room.

  • Biotech R&D teams
  • Translational research groups
  • Core facilities

For AI-science platforms

Genomarker Assurance API

The checks your agents call before they claim a result — the same rules and the same evidence, whoever the actor is. Developer preview, mid-2027.

  • AI-scientist platforms
  • Pharma AI teams
  • Workflow systems

What a statement says

Precise claims, not “certified correct.”

Scientific truth can't be read off a workflow, so Genomarker never issues a vague verdict. Each statement says what was checked, what passed, what was flagged — and what it does not establish.

Assurance dimensions
Assurance dimensionThe question it answers
Input integrityDo we know exactly what data entered the analysis?
Study designDoes the experimental design actually support the contrast being tested?
Statistical methodologyWere appropriate methodological safeguards applied?
Confounding & leakageWere detectable confounders and information leakage caught?
Plan alignmentDid the analysis follow the pre-registered plan?
Execution & provenanceDid the requested computation run as specified — and can we reconstruct what happened?
Independent replayCan it be re-executed independently of the original run?
Long-term reproducibilityIs enough state preserved to replay it years from now?

Independent replay and long-term reproducibility are assessed as their engines ship (mid-2027); until then a statement marks them not assessed.

Read how it works

The north star

Results should not expire when their software does

Where this is going: a result sealed in 2026 that an independent researcher can still verify, reconstruct and replay in 2036 — without the original scientist, agent or software.

2026

Original analysis

  1. Analysis requested — by a scientist or an AI agent
  2. Design problem detected; execution refused
  3. Reason returned; plan corrected
  4. Analysis run; signed record issued
  5. Environment and dependencies preserved
  6. Evidence published as GM-2026-08-14-8F29E1

2036

Ten years later — no original researcher required

  1. Independent researcher retrieves GM-2026-08-14-8F29E1
  2. Verifies the signatures offline
  3. Inspects the original scientific decisions
  4. Reconstructs the computational environment
  5. Replays the computation
  6. Reproduced — equivalent result

Team

Who's building it.

Salar Sayyad

Founder & CEO

Built the Genomarker platform end to end, solo and full-time since Feb 2026 — five releases in under six months and a nine-study public audit run on production. 10+ years as an AI engineer, data scientist and product manager; BSc Computer Science, University of Toronto.

Dr. Saed Sayad

Scientific advisor · AI & bioinformatics

PhD in biochemistry & bioinformatics; adjunct professor at the University of Toronto and former professor of data science at Rutgers University, with 25+ years in machine learning and predictive modeling, focused on biomarker discovery and precision medicine. Inventor of the Real Time Learning Machine and creator of the widely used online text An Introduction to Data Science. Advises on statistical methodology and the rule library.

Dr. Mark Hiatt, MD, MBA, MS

Scientific advisor · clinical & regulatory

Stanford fellowship-trained physician executive in precision medicine — former VP of Medical Affairs at Guardant Health, SVP of Market Access & Strategy at BostonGene, and CMO at RadSite; 75+ articles and book chapters, 150+ conference presentations. Advises on clinical, regulatory and market-access context for evidence standards.

Proof before trust.

We're onboarding research teams, AI-science platforms and review clients in small cohorts. Bring us an analysis — or a claim you're about to bet on.