TALARIA INTELLIGENCE
Predictive Analytics · Clinical Trial Outcomes
The counterfactual thesis

Most models tell you a trial will fail. Ours tells you what to change.

Talaria predicts clinical-trial outcomes before the trial runs, across every therapeutic area — and returns the specific design levers that move the result, each on its own causal interval. counterfactual panel — the design changes that move the outcome, unranked, each with its own interval

The Instrument · move the leversynthetic composite · illustrative
broadenriched

as designed — misses the endpoint as revised — one lever moved, it clears

The Instrument. A hand-authored synthetic fixture — not a Talaria prediction and not a measured success rate. Before the Day-0 datum the corridor is neutral: you cannot tell which future you are in until the decision is made. Band thickness is uncertainty, not decoration.

PREDICTION · synthetic composite DRAFT-0417interval
0.36 – 0.46 (schematic)
01Enrichment cutoffinterval · unranked
02Primary endpointinterval · unranked
03Populationinterval · unranked
04Powering & enrollmentinterval · unranked
05Line of therapyinterval · unranked
+ 2 more levers — powering & enrollment, line of therapy
counterfactual panel attached · Lock 6 — levers shown unranked and unvalued, per Lock 5
1
shared prediction engine
57
predictions across the R&D lifecycle — the designed surface
5
failure modes, in one fixed causal order
100%
of predictions carry their levers, by contract
What you get back

A score, and a ranking against other assets.

A calibrated interval plus the five design levers that move it — each on its own causal range.

When it is useful

After the protocol is locked, when nothing can change.

While the protocol is still editable — the only moment an answer can still alter the outcome.

What it admits

One number, with no statement of its own limits.

Its uncertainty, subgroup disagreement and drift state — and, in writing, what it cannot yet do.

Get in touch

Two ways to start.

Join the founding cohort. Tell us who you are and what you are working on, and we will bring the right person to the call — with the methodology and the current limitations, in writing, before you rely on anything.

No card, no obligation. Submitting opens our calendar so you can pick a time — and sends us your details so the call starts with context instead of a discovery questionnaire.

What a founding conversation holds

  • Your levers, on your programme. We run one of your own trial drafts through the counterfactual panel, not a demo fixture.
  • First access at deployment. Founding members are contacted before general availability.
  • A say in the roster. Which predictions we build after the core is decided on customer pull — founding members are who we ask.
  • The methodology, in full. The held-out design, the leakage audit and the calibration evidence — not a summary slide.
  • The frontier, unhidden. You get told what the model cannot yet do, in writing, before you rely on it.

We use what you share to route you to the right surface and to prepare for the call — nothing else. No card is collected and no charge is made at any point on this page.

The signature

A trial does not simply fail. It fails of something.

Talaria decomposes a predicted failure across five categories in one fixed order — warm to cool, patient-proximal to system-level. The order is the information; it is never reordered. This is the taxonomy the 5px band at the top of every page stands for.

patient-proximalsystem-level

Illustrative portfolio-level shape — not any single trial's composition, and not a Talaria measurement.

The translation gap

The failure is rarely the molecule. It's the translation.

Drugs that succeed in Phase II routinely fail in Phase III. The failure is rarely the compound — it is the design: the endpoint too insensitive to catch the effect, the population that diluted it, the trial powered for a difference larger than the drug delivers, the enrichment cutoff set one notch too wide.

The molecule was not the difference. The design was.

Every one of those was a decision made at design time, and every one was editable — free to change, right up until the trial locked and it wasn't. Talaria prices those decisions while they are still decisions, and hands back the ones worth changing.

The engine

A score tells you the odds. We tell you the edit.

Talaria is not a leaderboard model. It is a shared representation of a trial, anchored to a causal map, so every prediction answers not just “how likely?” but “which change would move it, and by how much?” That second answer is the product — and it is the one a score cannot give you at any level of accuracy.

What the answer carries · a leaderboard model vs. Talaria
The question a programme actually asks A leaderboard model Talaria
“Will this design survive Phase III?” a pointA probability, and nothing travels with it. an intervalA calibrated probability with its own uncertainty range, and the state of its calibration today.
“Then what do we change?” not answerableNothing. The model has no representation of an intervention, so it cannot be asked. answered by contractFive design levers — endpoint, population, enrichment, powering, line of therapy — each on its own causal range. Attached to every prediction.
“Does it hold for our patients?” pooledOne number averaged across every subgroup, which hides the strata that disagree. stratifiedSubgroup estimates that are permitted to contradict the pooled figure — and are shown when they do.
“How sure is it — and sure of what?” undifferentiatedA single confidence, if any, that cannot separate ignorance from noise. decomposedEpistemic uncertainty, which shrinks with evidence, split from aleatoric, which does not — plus a flag when the trial is unlike anything learned from.
“Is it still trustworthy this quarter?” staticFixed at training time. It reports a stale number without saying so. self-withdrawingA live drift and coverage state that withdraws its own calibration when the ground moves, rather than quietly reporting on.
“Can we show how it decided?” opaqueA score with no reproducible provenance behind it. signedAn Ed25519-signed audit row naming model, features and substrate version for every prediction issued.

This compares what the output carries, which is a property of how each system is built — not a scoreboard. Talaria captures a larger share of the achievable headroom than published Phase III baselines, but that is a contextual comparison across different corpora, not a paired head-to-head result — and we do not publish one until an external-corpus paired evaluation exists. No performance figure on this site is a measured Talaria result.

What a prediction carries

Anatomy of one prediction.

Everyone can return a number. The question is what travels with it. Below is one synthetic prediction, opened up — each register is something Talaria attaches that a leaderboard model does not. The vertical line is the same prediction throughout.

ONE PREDICTION · 0.41 probability (schematic) 0.2 0.4 0.6 0.8 01 · WHAT A LEADERBOARD MODEL RETURNS 0.41 a point. nothing travels with it. 02 · + UNCERTAINTY ENVELOPE epistemic — model ignorance. shrinks with evidence. aleatoric — irreducible. does not. out-of-distribution flag is this trial even like the ones we learned from? 03 · + SUBGROUP HETEROGENEITY (CATE) the one interval fans into strata that disagree — the pooled number hid every one of them biomarker-high 0.62 – 0.78 (schematic) biomarker-low 0.21 – 0.37 (schematic) elderly 0.30 – 0.66 (schematic) — wide: few events 04 · + DISAGREEMENT MONITOR causal path A causal path B — disagrees adjudicated: genuine epistemic uncertainty when two routes to the same number disagree, something has to say why — correlated error, leakage, or real uncertainty 05 · + CALIBRATION TRUST ACTIVE effective coverage 95.0% (heuristic) · drift GREEN is the model still trustworthy for this trial, today? 06 · + COUNTERFACTUAL PANEL & SIGNED AUDIT five levers, fixed order, unranked and unvalued (Lock 5) ed25519 signed · model · features · substrate — Lock 11

Every value here is a hand-authored synthetic fixture — illustrative of the structure a prediction carries, not a Talaria measurement and not a success rate. No leaderboard model produces this. A probability you cannot act on is a number; a probability with a lever is a decision.

Register 1, the bare number: 0.41, a point with nothing attached. Register 2, uncertainty envelope: an interval split into epistemic uncertainty, which shrinks with evidence, and aleatoric uncertainty, which does not, plus an out-of-distribution flag asking whether this trial resembles the ones the model learned from. Register 3, subgroup heterogeneity or CATE: biomarker-high 0.62 to 0.78 schematic, biomarker-low 0.21 to 0.37 schematic, elderly 0.30 to 0.66 schematic and wide because there are few events. The pooled number hid every one of these. Register 4, disagreement monitor: the same prediction computed along two causal paths that disagree, adjudicated as genuine epistemic uncertainty. Register 5, calibration trust: state ACTIVE, effective coverage 95.0 percent heuristic, drift GREEN. Register 6: the five levers in fixed order shown unranked and unvalued per Lock 5, and an Ed25519-signed audit row carrying model, features and substrate versions per Lock 11.

One engine, three surfaces

Asked at three moments in a trial's life.

The surfaces are the same engine addressed at different points — and when you ask is what changes the answer's worth. Each opens an interactive demo that runs on a synthetic composite.

Discovery to designations — and where an answer still changes something.

Read left to right as time. The band narrows because attrition is the base rate, and the two dashed rails are drawn as rails on purpose: trial design and designations recur throughout the arc rather than happening once. Durations are approximate industry-typical figures for orientation — not Talaria measurements.

THE R&D ARC · WHERE EACH SURFACE CAN STILL ACTIntelligencebefore a phase commits — the protocol is still editableSentinelwhile it runs — the protocol is frozenReviveafter a readout failsattritionDiscovery~4 yrIND · Ph 0–1~1.5 yrPhase 1–2~2 yrPhase 2–3~2 yrPhase 3~3 yrSubmission · FDA~1.5 yrPhase 4 · lifecycleopen-endedthe band narrows because most programmes stop before filing (schematic)Trial designre-decided before every phaseDesignations · int’lsought at several points
hue = position along the lifecycle, discovery through designations bar height = heads at that stage this ramp encodes stage order, not the failure taxonomy above

One shared substrate; 57 output-adapter heads; the seven-learner cap never moves. These bars are stage counts — per-head status lives in the Atlas, where all 57 can be filtered by stage, status and decision value. Open the interactive Portfolio Atlas

Leading indicators

What we measure before we ask you to believe it.

A prediction engine is only worth what its evaluation is worth. These are the standards a Talaria number has to clear before it leaves the building — and the reason no accuracy figure appears anywhere on this site yet.

01 · Temporal honesty

Trained on the past. Judged on the future.

Every evaluation holds out trials by date, never at random. A random split lets a model learn from trials that had not happened yet — it flatters the score and predicts nothing. The cutoff is a frozen constant, and changing it requires a formal design review.

02 · Leakage audit

Every feature is presumed guilty.

Features are computed with the label inaccessible by construction, and each one faces an adversarial audit that tries to prove it encodes the outcome. Families that failed — including some of our own — were removed and the headline number restated downward.

03 · Calibration, not ranking

A 30% must mean thirty in a hundred.

Ranking trials correctly is not enough; a probability you act on has to be true at its face value. We track calibration error alongside discrimination, and the engine withdraws its own calibration when the ground shifts rather than reporting a stale number.

04 · Scope, stated

A number without its population is noise.

Results are labelled with the corpus that produced them. A figure measured on oncology trials is not an all-therapeutics figure, and we will not present it as one. Where the broader corpus is not yet ready, we say so instead of borrowing the narrower number.

05 · The ceiling

We estimate what is even achievable.

Some of a trial's outcome is irreducibly unknowable at design time. We estimate that ceiling first, so performance is read as share of achievable headroom captured — which is the honest denominator — rather than against a perfect score no method can reach.

06 · The frontier, in writing

You get told what it cannot do.

Every engagement begins with a written statement of the model's current limits — which questions it answers, which it does not, and where the evidence is still thin. You receive that before you rely on it, not after.

Where the evidence stands today

Talaria is pre-deployment. Our internal results are measured on a rich oncology corpus; the all-therapeutics substrate is built but its outcome label is not yet clean enough to certify, and we will not publish a number that rests on it.

So this site publishes no accuracy figure at all. Every value shown in the figures above is a hand-authored synthetic fixture, labelled as such, and no screenshot on this site places a probability beside a real asset's known outcome. When a figure clears the standards above, it will appear here with its corpus, its cutoff date and its interval attached.

Under evaluation and want the detail? The full methodology, the held-out design and the current limitations go to founding-cohort conversations directly. Ask us for it