We built the
AI judgement engine
for therapeutic development.

Your data, proprietary or public, reasoned over the way a scientist weighs thin evidence, down to the low-magnitude signals that brute-force averaging destroys.

The problem

A drug programme is a chain of commitments, none of them scored.

The commitments are target, modality, delivery vehicle and patient population. Potency, exposure and cost carry numbers. The evidence underneath them does not. Each commitment forecloses the next, and the trial that settles it arrives years later saying the programme failed, not which commitment was wrong.

A better assistant cannot close this

The literature records what worked, rarely what a good scientist thought before they knew. No corpus teaches that judgement and no index retrieves it. A strong scientist gets faster, a weak chain of evidence stays invisible.

A person's judgement does not accrue

It leaves when they do and never reaches the programme next door. The judgement we encode sits in one schema, so a link found in cancer is live for an autoimmune programme on the same pathway. Upskilling scales with headcount, a written layer compounds.

Generating hypotheses is cheap now. Judging them still costs a clinical trial.

The thesis

Scale won everywhere the score was cheap and repeatable.

Here it is neither. The industry spends about a billion dollars in R&D for every new drug that reaches the market, and what would teach a model is not one trial but a representative sample across every target, modality and patient population, drawn from a supply of patients that does not regenerate. What is left is a ladder of proxies.

How far from a living humanHow much of it existsFaithfulness
Human outcomeThe one place biology answers straight
In vivo, animalA mouse is not a human
~10×
In vitroA cell measured outside its niche
~100×
In silicoNot measured at all
~1000×

Where a single layer stops

AlphaFold learns one transformation, and the representative data for it already existed. Its structures feed our graph, and a perfect structure still does not tell you whether inhibiting that protein helps a patient.

What we do about it

Elman scores each link on the strength of the evidence behind it, and judges how far it carries toward the human outcome.

Nobody can train their way to that, and distillation and synthetic data start from the same missing measurements. Factoring in every variable, representative data for one clinical question would take billions of patients.

What faithfulness is, and the benchmark that would test it.

Read the full thesis →

How it works

What it does today.

01

Unifies evidence

Single-cell RNA-seq, genomics, clinical and animal-model data, proprietary or published, resolved into one graph instead of analysed one layer at a time.

02

Scores every link

Mechanistic hypotheses built as cause-and-effect chains, each one held, disproven or sent back for more.

03

Plans the experiment

Commissions the wet-lab test that would settle the question, and reads the result back into the graph.

It runs on whichever model suits the subtask, a frontier LLM, a protein-folding model, a binding-affinity model.

Every edge resolves to what produced it, an experimental arm in a paper with magnitude, method, p-value and sample size, a bioinformatics run on your data or a wet-lab result, and never a model's memory of the literature.

What we are building toward is a number for how far evidence carries, tested against real trial outcomes.

Evidence

It named a mechanism where the field expected resilience.

The Allen Institute set the engine on one population of dying visual-cortex neurons. Nobody knew why. It converged on hyperexcitability.

225
hypotheses, three runs
10 minutes
and a few dollars of compute each
32 / 32
held and disproven, every verdict public
The value of the workflow is not that the first-pass output is correct, but that every claim is traceable to the experiments supporting it, and that the full reasoning is exposed for others to revise, extend, or refute.
From the preprint we co-authored with the Allen Institute, who funded the work under a research contract. Submitted to Cell
170,000
experimental facts, each tied to the experiment behind it, all in one schema. Evidence gathered to answer one question is already there for the next.

Our mission

Make therapeutic development predictable.

The clinical trial should confirm the answer, not reveal it.

If you are developing therapeutics, generating evidence, or want to help make this possible, we should talk.

hello@elman.ai →