BioDecisionBench: Can models reason through the narrow path to drug development success?

1 min read

AI models can solve formidable mathematical problems, trounce humans at programming competitions, and outperform doctors. But can they help people make sound decisions when the evidence is ambiguous and the stakes are high? How do we know they are helping us reason better, not just faster?

To find out, we built BioDecisionBench, a collection of 40 tasks derived from 26 cases of complex reasoning failures in life sciences. The tasks span the drug development process, from target selection to clinical trial design to choosing new indications for a successful drug. We showed that our new Elicit Research Agent sets a new performance frontier on BioDecision and thus for reasoning through high-stakes decisions in life sciences.

Why and how we designed BioDecisionBench

High-stakes decision-making under ambiguity and limited information is an acute problem in drug development. A company only gets to pick a small set of targets to pursue, a limited number of assays to run, and a single clinical trial design to conduct. It’s not possible to explore all avenues, so the only way to do better is to reason better about how things might have gone differently if they had chosen differently.

AI research agents promise to help, assuming they can be evaluated on these capabilities. However, existing AI benchmarks in the life sciences do not meet the challenge. They often:

  1. Only cover bioinformatics data analysis

  2. Are saturated by current frontier models

  3. Use rubrics that are not comprehensive or don’t reward a variety of good answers

  4. Only evaluate the final answer rather than the process used to get there

So, we created BioDecision. It covers real, recent decisions from the whole pharmaceutical development cycle, including everything from target selection to clinical trial design to choosing new indications for a successful drug. BioDecision is not saturated on current frontier models (Claude Opus 5 Max and GPT‑5.6 Sol Max), is based on high-quality reference materials and iteration with domain experts, and measures both research process and the quality of the answers.

Tasks can have multiple variants — different framings of the same case. Each variant contains a detailed prompt to be given to an agent, as well as detailed rubrics that tell a judge model how to score the agent’s answer to this task.

BioDecision was built with pharmaceutical researchers and executives (including Dr. James McIlroy, CEO of EnteroBiotix, and Dr. Yin Huang, Senior Scientist at Propeller Bio) and based on state-of-the-art methods research.

An example: Timing of immunotherapy

In 2021, Qian et al. published a retrospective study showing that patients who received immunotherapy earlier in the day survived longer than those who received it later. Many more observational studies followed. Eventually, at least 13 randomized controlled trials were planned, likely costing tens of millions of dollars. The first of these (Huang et al. 2026) initially published positive results, but was then retracted. Did these studies deserve the time and money that went into them?

We asked AI tools a version of this question (see the Appendix for the full version), then graded this answer based on how well it identified 31 risks or limitations of the evidence and the decision, for example:

A. Evidence use and evidence lineage

  • A3 — Changing the threshold for what counts as “early” in the day changed the result of the analysis

  • A10 — In the SWOG S1404 trial, patients who lived closer to the hospital tended to receive care earlier in the day, introducing a confounder

  • A11 — The analysis is sensitive to small changes in specification, e.g., the exact time threshold used

B. Empirical analyses in the health-system data

  • B2 — The study failed to carefully reconstruct when patients actually received treatment, and just relied on scheduled times

  • B3 — Time zones and time recording practices were not standardized

D. Decision consequences

  • D1 — The answer should set a minimum threshold for what effect size is worthwhile

  • D2 — The answer should not recommend switching to early-in-day immunotherapy

In answering this question, Elicit found many things that Claude missed:

  1. Confounders — living close to the hospital makes it easier to get there for early-in-day treatment (9% less likely to get early-in-day treatment per 50 miles distance from hospital), but also makes it easier to receive care in general (this covers A10 in the rubric)

  2. Brittle sensitivity to assumptions — Changing the threshold time for what counts as “early in the day” from “before 3:48 PM” to “before 3:18 PM” erased the observed effect (this covers A11 in the rubric)

  3. Attrition bias — in Tsukaguchi et al. 2025, patients who received fewer than 4 treatments were excluded from the analysis. If there’s a systematic difference between the groups in who reaches 4 treatments, this may bias the results

Elicit’s analysis makes a convincing case against immunotherapy timing trials. This analysis might have convinced researchers to redirect the time and money spent on them to more promising trials.

Agents can reason better

We compared the performance of Elicit’s Research Agent, on its smartest setting, to the following AI tools, set wherever possible to their highest effort setting:

  • Frontier models

    • Claude Opus 5 with Max effort

    • GPT‑5.6 Sol with Max effort

  • Specialized science agents

    • Biomni, with the Max setting

    • Claude Science with Opus 4.8

    • Edison Scientific’s Edison Literature High

We generated 3 answers per task variant for each tool, for a total of 120 per tool (except for Edison Scientific, for which we generated 2 answers per variant for a total of 80). We scored each answer according to the BioDecision rubrics using Claude Opus 5, GPT‑5.6 Terra, and Gemini 3.1 Pro Preview (we averaged the scores across the three judges).

Elicit Research Agent scores 76.6%. The two general frontier systems follow at 68.7% and 66.3%, and the three science-specific tools cluster near half marks: 48.6%, 47.4%, and 45.6%. Elicit scored higher than any other tool, p < 0.05 for each comparison.

The averages in the previous chart don't tell you whether Elicit is reliably better or just better on average, because some tasks are harder than others for everyone. So we also compared the systems head to head on identical work. Every system answered all tasks 3 times, which means that for any pair of systems and any task there are 9 matched comparisons. We computed the win, tie, and loss shares within each task, then averaged across the 40 tasks so every task counts equally.

Against Claude Opus 5, Elicit scores higher on 70.6% of matched comparisons and lower on 15.6%. Against GPT‑5.6 Sol the split is 70.6% and 17.2%. Most of the remaining ties are cases where both systems got full marks (so a few easy tasks show saturation).

Against the three science-specific tools, Elicit scores higher on 97.5% to 100% of comparisons and loses fewer than one comparison in three hundred.

How Elicit outperforms

We believe that a few families of reasoning errors underlie many of the mistakes in drug development, e.g.:

  1. study design for preclinical studies that doesn’t meet our basic standards for randomized controlled trials

  2. assuming that observational evidence shows causality

  3. uncritically using performance on biomarkers and other surrogates as evidence of therapeutic benefit

To outperform the other tools, Elicit:

  1. gets an initial answer from a model

  2. has another model critique the answer, based on guidance on how to avoid the common reasoning error families

  3. has the original model update its answer based on the critique

In preliminary results not reported here, we found that Elicit did not outperform other systems without this process of critique and rewriting.

Down the narrow path of good reasoning

BioDecision shows that AI tools are somewhere in the middle of their journey toward good reasoning. All tools, especially the frontier models we tested, identified many key considerations on the decisions we tested. Elicit identified more, but no tool came close to identifying them all. We will continue iterating on Elicit to make it more rigorous, and on BioDecision to make it a better measure of reasoning (and we will publish it in the near future). We hope this work will contribute to better decisions that save lives.

Appendix: Full immunotherapy timing prompt

I lead real-world evidence strategy for a first-line metastatic renal-cell carcinoma program. Across several centers, patients whose checkpoint-inhibitor infusions were usually given earlier in the day had better outcomes than patients treated later. Similar findings have been reported in other cancers, and circadian changes in immune activity provide a plausible rationale. The result has prompted interest in reserving morning infusion slots and in funding a prospective scheduling trial.

Our health-system network has several years of data on patients starting first-line nivolumab with ipilimumab, followed by nivolumab maintenance when clinically appropriate. We have order and administration timestamps for each antibody, original appointment times and rescheduling, laboratory results, performance status, the clinical factors used in standard metastatic-RCC risk assessment, prior nephrectomy, metastatic sites and burden, corticosteroid use, immune-related adverse events, treatment holds and discontinuations, hospitalizations, imaging assessments, progression, subsequent anticancer treatment, and death.

The current proposal is to include patients who received at least four checkpoint-inhibitor administrations, derive each patient's usual infusion window from treatment during the first 12 weeks, divide them into earlier- and later-treatment groups, and compare progression-free and overall survival from the first infusion after propensity-score matching.

Could you prepare the analysis and decision memo we should use before changing scheduling practice or committing to a prospective trial? Explain how you would determine whether earlier administration genuinely improves outcomes, what you would change in the proposed analysis, how you would distinguish a nivolumab-timing effect from ipilimumab timing, the induction visit as a whole, later nivolumab maintenance, and the health system's scheduling process, and what findings would justify either proceeding to a randomized timing trial or abandoning the hypothesis.

Appendix: Another example — microdystrophin for Duchenne muscular dystrophy

In June 2023, Dr. Peter Marks, a Director at the FDA, overruled 4 FDA review bodies to grant accelerated approval for a gene therapy for Duchenne muscular dystrophy (DMD, a severe loss of muscle function caused by a mutation in the dystrophin gene).

Then in November 2025, the FDA narrowed the indication for the gene therapy to ambulatory patients only and added a safety warning about potentially fatal liver side effects.

Between the accelerated approval and that rollback, the trial of the drug failed its primary endpoint, and 2 people who took it died of liver failure. For now, the drug is still approved for its narrowed indication, but some question whether it should be approved at all.

Issues with the preclinical evidence

Why did Marks approve the drug in the first place? DMD is fatal and lacks effective therapies, so Marks was willing to take a risky bet on a therapy for it. He determined that the gene therapy made cells do the things they’d need to do to improve muscle function, and so the therapy was reasonably likely to help patients with DMD.

Marks relied on preclinical studies that showed that expression of microdystrophin (an engineered protein similar to dystrophin) correlated with improved muscle physiology in animals. However, these studies have the following problems:

  1. If improved muscle function is the “elephant,” then the assays measuring it were the blind men. Some measured protein in bulk tissue with no guarantee that it was in the right place; others measured protein at the sarcolemma without showing that it restored the dystrophin-associated protein complex or muscle function. Each assay captured one feature of the protein and risked mistaking that feature for therapeutic efficacy

  2. Only young animals modeling mild disease showed an effect, while other animal models did not show any effect

  3. The studies failed to scrutinize potential adverse effects in cardiac muscle, and potential immune reactions caused by the therapy. More microdystrophin expression for certain patient subgroups reflected more antigens and more potential for harm

How to do better

What could drug developers do differently if they wanted to bet on developing a better microdystrophin therapy?

We asked agents:

For microdystrophin gene therapies in Duchenne muscular dystrophy, what would need to be true for microdystrophin expression to have high predictive validity for functional or therapeutic benefit? Where is the biggest source of risk with microdystrophin expression readouts?

We evaluated how well each agent assessed the evidence on microdystrophin biomarkers and opportunities for improvement.

Elicit did a better job than other tools at flagging these key considerations:

  • Becker muscular dystrophy does not show that microdystrophin can always stand in for dystrophin

  • Relatedly, microdystrophin is a better biomarker for some genotypes than for others

  • Disease stage, age, tissue, and endpoint modify how well microdystrophin expression predicts function

  • Microdystrophin cannot easily measure the overall therapeutic index, i.e., the benefit-harm tradeoff. More microdystrophin sometimes meant more harm rather than more benefit

Elicit’s answer could potentially have accelerated work on microdystrophin therapy by convincing researchers to search for better biomarkers, and saved the lives of patients who died from the gene therapy.

More broadly, Elicit identified opportunities to construct more valid preclinical biomarkers for future DMD therapies. Preclinical researchers using Elicit have a better chance of choosing the right biomarkers and the right therapies, and of bringing those therapies to market.