Here’s the uncomfortable part up front: the first verdict didn’t feel uncertain. It felt finished. Confident, clean, done. It was just wrong by a full step — and the only thing separating it from the better verdict was how completely I’d assembled the evidence before I started judging.
That’s the whole point of this piece. The drug is just the example.
The example, fast
Late 2012. Eli Lilly has two failed Phase 3 trials of solanezumab — both missed their co-primary endpoints in mild-to-moderate Alzheimer’s. But a pre-specified pooled analysis of the milder patients shows a real cognitive signal (~34% slowing, p≈0.001). Lilly narrows the program to mild patients and launches a third pivotal trial.
The easy read: failed twice, cherry-picked a subgroup, kept going. Sunk cost.
The easy read is lazy, and reading the ending backward is how most of us review our own history.
Pass one (thin evidence): “this needed restructuring”
Working from topline results and trial design alone, the decision looked weak. Two failed primaries, a verdict resting on a subgroup, a mechanism that looked refuted by the field. Verdict: poor warrant, restructure.
Pass two (full evidence): “this was defensible”
Then I added three things that were knowable in 2012 but missing from the thin pass:
• The mechanism signal was differentiated, not blunt. Bapineuzumab failed that same year — but it cleared plaque and still didn’t work. Solanezumab did the opposite: didn’t clear plaque, but engaged soluble amyloid and carried the cognitive signal in milder patients. So 2012 wasn’t “anti-amyloid is dead.” It was a specific, reasoned signal — and a real argument to intervene earlier.
• The endpoint change went through FDA, pre-lock. Not a quiet post-hoc dredge into unblinded data. A regulator-aware adjustment before database lock. Much weaker objection.
• The safety profile was clean. ARIA-E under 1%, essentially placebo-level, against bapineuzumab’s 16–21% in some groups. For a multi-year drug in a high-need disease, that’s a legitimate reason to commit.
Same decision. Same method. Same reasoning. The verdict moved from “restructure” to “strengthen” — one full step — because the evidence base got more complete. Nothing about Lilly’s actual choice changed between my two passes.
Where the decision genuinely doesn’t hold
To be clear: defensible to run a confirmatory trial is not the same as the efficacy being real. The signal lived in a subgroup, in a secondary analysis, never in a met primary. Not once across two trials.
That’s exactly the right basis to test a hypothesis. It is not a basis to call the effect established — and some of the 2012 language leaned that way. The weakness was never continuing the program. It was treating a hypothesis-generating subgroup like a confirmed result.
(Yes, EXPEDITION3 later failed. No, that’s not needed to reach any of this. If your read of a decision only works once you know the ending, you haven’t judged the decision — you’ve described the outcome.)
The thing worth keeping
An under-grounded judgment doesn’t announce itself. It doesn’t feel thin from the inside. It feels done — and lands with the same confidence as a sound one, just off by a step.
Which means how completely you built the evidence isn’t a warm-up before the real analysis. When the stakes are a multi-year program, it’s the part that decides whether the analysis is worth anything at all.
Case #1 in a series on public drug-development decisions and the gap between a decision’s quality and its outcome. I do this structured retrospective work through HypotheX. More to come.


