Most models tell you a trial will fail. Ours tells you what to change.
Talaria predicts clinical-trial outcomes before the trial runs, across every therapeutic area — and returns the specific design levers that move the result, each on its own causal interval. counterfactual panel — the design changes that move the outcome, unranked, each with its own interval
The Instrument. A hand-authored synthetic fixture — not a Talaria prediction and not a measured success rate. Before the Day-0 datum the corridor is neutral: you cannot tell which future you are in until the decision is made. Band thickness is uncertainty, not decoration.
A score, and a ranking against other assets.
A calibrated interval plus the five design levers that move it — each on its own causal range.
After the protocol is locked, when nothing can change.
While the protocol is still editable — the only moment an answer can still alter the outcome.
One number, with no statement of its own limits.
Its uncertainty, subgroup disagreement and drift state — and, in writing, what it cannot yet do.
Two ways to start.
Join the founding cohort. Tell us who you are and what you are working on, and we will bring the right person to the call — with the methodology and the current limitations, in writing, before you rely on anything.
Got it — now pick a time.
Your details are with us. Choose a slot that suits you and we will come prepared.
This did not reach us. Your details are saved in this browser only. Send them directly so nothing is lost:
What a founding conversation holds
- Your levers, on your programme. We run one of your own trial drafts through the counterfactual panel, not a demo fixture.
- First access at deployment. Founding members are contacted before general availability.
- A say in the roster. Which predictions we build after the core is decided on customer pull — founding members are who we ask.
- The methodology, in full. The held-out design, the leakage audit and the calibration evidence — not a summary slide.
- The frontier, unhidden. You get told what the model cannot yet do, in writing, before you rely on it.
We use what you share to route you to the right surface and to prepare for the call — nothing else. No card is collected and no charge is made at any point on this page.
Stay up to date with our news. Occasional notes on what we are building, what we have validated and what we have had to retract. A name and an email are all we need — and we will stay in touch.
You're on the list.
We will be in touch with our news — and nothing else.
This did not reach us. Send it directly so you are not missed:
What we actually send
- What shipped. New prediction surfaces as they become available, with their scope stated.
- What we validated. Results that cleared the standards above — with corpus, cutoff and interval attached.
- What we retracted. When a number does not survive audit, you hear that too. It is the part most companies leave out.
Infrequent by design. If a month passes with nothing worth reporting, you will not hear from us that month.
A trial does not simply fail. It fails of something.
Talaria decomposes a predicted failure across five categories in one fixed order — warm to cool, patient-proximal to system-level. The order is the information; it is never reordered. This is the taxonomy the 5px band at the top of every page stands for.
Illustrative portfolio-level shape — not any single trial's composition, and not a Talaria measurement.
The failure is rarely the molecule. It's the translation.
Drugs that succeed in Phase II routinely fail in Phase III. The failure is rarely the compound — it is the design: the endpoint too insensitive to catch the effect, the population that diluted it, the trial powered for a difference larger than the drug delivers, the enrichment cutoff set one notch too wide.
The molecule was not the difference. The design was.
Every one of those was a decision made at design time, and every one was editable — free to change, right up until the trial locked and it wasn't. Talaria prices those decisions while they are still decisions, and hands back the ones worth changing.
A score tells you the odds. We tell you the edit.
Talaria is not a leaderboard model. It is a shared representation of a trial, anchored to a causal map, so every prediction answers not just “how likely?” but “which change would move it, and by how much?” That second answer is the product — and it is the one a score cannot give you at any level of accuracy.
| The question a programme actually asks | A leaderboard model | Talaria |
|---|---|---|
| “Will this design survive Phase III?” | a pointA probability, and nothing travels with it. | an intervalA calibrated probability with its own uncertainty range, and the state of its calibration today. |
| “Then what do we change?” | not answerableNothing. The model has no representation of an intervention, so it cannot be asked. | answered by contractFive design levers — endpoint, population, enrichment, powering, line of therapy — each on its own causal range. Attached to every prediction. |
| “Does it hold for our patients?” | pooledOne number averaged across every subgroup, which hides the strata that disagree. | stratifiedSubgroup estimates that are permitted to contradict the pooled figure — and are shown when they do. |
| “How sure is it — and sure of what?” | undifferentiatedA single confidence, if any, that cannot separate ignorance from noise. | decomposedEpistemic uncertainty, which shrinks with evidence, split from aleatoric, which does not — plus a flag when the trial is unlike anything learned from. |
| “Is it still trustworthy this quarter?” | staticFixed at training time. It reports a stale number without saying so. | self-withdrawingA live drift and coverage state that withdraws its own calibration when the ground moves, rather than quietly reporting on. |
| “Can we show how it decided?” | opaqueA score with no reproducible provenance behind it. | signedAn Ed25519-signed audit row naming model, features and substrate version for every prediction issued. |
This compares what the output carries, which is a property of how each system is built — not a scoreboard. Talaria captures a larger share of the achievable headroom than published Phase III baselines, but that is a contextual comparison across different corpora, not a paired head-to-head result — and we do not publish one until an external-corpus paired evaluation exists. No performance figure on this site is a measured Talaria result.
Anatomy of one prediction.
Everyone can return a number. The question is what travels with it. Below is one synthetic prediction, opened up — each register is something Talaria attaches that a leaderboard model does not. The vertical line is the same prediction throughout.
Every value here is a hand-authored synthetic fixture — illustrative of the structure a prediction carries, not a Talaria measurement and not a success rate. No leaderboard model produces this. A probability you cannot act on is a number; a probability with a lever is a decision.
Register 1, the bare number: 0.41, a point with nothing attached. Register 2, uncertainty envelope: an interval split into epistemic uncertainty, which shrinks with evidence, and aleatoric uncertainty, which does not, plus an out-of-distribution flag asking whether this trial resembles the ones the model learned from. Register 3, subgroup heterogeneity or CATE: biomarker-high 0.62 to 0.78 schematic, biomarker-low 0.21 to 0.37 schematic, elderly 0.30 to 0.66 schematic and wide because there are few events. The pooled number hid every one of these. Register 4, disagreement monitor: the same prediction computed along two causal paths that disagree, adjudicated as genuine epistemic uncertainty. Register 5, calibration trust: state ACTIVE, effective coverage 95.0 percent heuristic, drift GREEN. Register 6: the five levers in fixed order shown unranked and unvalued per Lock 5, and an Ed25519-signed audit row carrying model, features and substrate versions per Lock 11.
Asked at three moments in a trial's life.
The surfaces are the same engine addressed at different points — and when you ask is what changes the answer's worth. Each opens an interactive demo that runs on a synthetic composite.
Intelligence
Ask before you commit: will this design survive Phase III — and what would have to change for it to? The answer arrives as a calibrated interval with the five design levers that move it, each on its own causal range. It is built for the moment a protocol is still editable, because that is the only moment the answer is worth anything.
Open Intelligence · M01Revive
This asset failed. Was it the molecule, or the trial — and could a redesign bring it back? Revive separates a compound that cannot work from a compound that was asked the wrong question, then prices what a corrected trial would have to look like. Most shelved assets were never cleanly falsified — they were under-powered, mis-populated, or read out on the wrong endpoint.
Open Revive · M02Sentinel
The trial is running. Is the ground it was designed on still the ground — and is the model still trustworthy for it? Sentinel watches the standard of care, the competitive field and the enrolling population for movement under a protocol that can no longer be changed. When the substrate shifts far enough, it says so and withdraws its own calibration rather than quietly reporting a stale number.
Open Sentinel · M03Discovery to designations — and where an answer still changes something.
Read left to right as time. The band narrows because attrition is the base rate, and the two dashed rails are drawn as rails on purpose: trial design and designations recur throughout the arc rather than happening once. Durations are approximate industry-typical figures for orientation — not Talaria measurements.
One shared substrate; 57 output-adapter heads; the seven-learner cap never moves. These bars are stage counts — per-head status lives in the Atlas, where all 57 can be filtered by stage, status and decision value. Open the interactive Portfolio Atlas
What we measure before we ask you to believe it.
A prediction engine is only worth what its evaluation is worth. These are the standards a Talaria number has to clear before it leaves the building — and the reason no accuracy figure appears anywhere on this site yet.
Trained on the past. Judged on the future.
Every evaluation holds out trials by date, never at random. A random split lets a model learn from trials that had not happened yet — it flatters the score and predicts nothing. The cutoff is a frozen constant, and changing it requires a formal design review.
Every feature is presumed guilty.
Features are computed with the label inaccessible by construction, and each one faces an adversarial audit that tries to prove it encodes the outcome. Families that failed — including some of our own — were removed and the headline number restated downward.
A 30% must mean thirty in a hundred.
Ranking trials correctly is not enough; a probability you act on has to be true at its face value. We track calibration error alongside discrimination, and the engine withdraws its own calibration when the ground shifts rather than reporting a stale number.
A number without its population is noise.
Results are labelled with the corpus that produced them. A figure measured on oncology trials is not an all-therapeutics figure, and we will not present it as one. Where the broader corpus is not yet ready, we say so instead of borrowing the narrower number.
We estimate what is even achievable.
Some of a trial's outcome is irreducibly unknowable at design time. We estimate that ceiling first, so performance is read as share of achievable headroom captured — which is the honest denominator — rather than against a perfect score no method can reach.
You get told what it cannot do.
Every engagement begins with a written statement of the model's current limits — which questions it answers, which it does not, and where the evidence is still thin. You receive that before you rely on it, not after.
Where the evidence stands today
Talaria is pre-deployment. Our internal results are measured on a rich oncology corpus; the all-therapeutics substrate is built but its outcome label is not yet clean enough to certify, and we will not publish a number that rests on it.
So this site publishes no accuracy figure at all. Every value shown in the figures above is a hand-authored synthetic fixture, labelled as such, and no screenshot on this site places a probability beside a real asset's known outcome. When a figure clears the standards above, it will appear here with its corpus, its cutoff date and its interval attached.
Under evaluation and want the detail? The full methodology, the held-out design and the current limitations go to founding-cohort conversations directly. Ask us for it