The Refusal Engine
A rehearsal. Freeze a ruleset on old refusals, test it on ones it never saw, no peeking.
The point was to prove the machinery, and to see what the exhaust says about how a real decider refuses.
What was built
- 448 Complete Response Letters, each read once and tagged into six reason categories.
- Split at 1 January 2025. 357 refusals before, as the build set. 91 after, locked.
- 25,293 approvals and 336 rejections merged into one decision table. Generics benched: zero refusals, which makes it a checklist, not a judgment.
- Pre-decision signals for 2,994 drugs. Trial evidence and structural flags.
141,819 decisions and reviews across the three subjects.
1. Judgment lives in the no, and it is a stack
| Warning flags stacked | 0 | 1 | 2 | 4 | 5 |
|---|---|---|---|---|---|
| Reject rate | 4% | 9.7% | 12% | 28% | 43% |
A single flag barely moves the call. Five stacked is ten times the base rate. In the letters, two or three reasons co-fire in most rejections. Single-reason refusals are 2% of nos.
2. The refusal pattern is stable across time
Reasons ranked nearly identically before and after the wall, on letters the file never saw.
| Reason | Before 2025 | After 2025 |
|---|---|---|
| Incomplete application | 78% | 68% |
| Manufacturing or facility | 70% | 57% |
| Safety | 60% | 51% |
| Trial design | 34% | 33% |
| Efficacy not shown | 25% | 33% |
The drugs change every year. The reasons they get told no do not. A judgment file frozen on old refusals still describes new ones. That was the claim the rehearsal had to prove, and it held.
3. The real bar is operational, not scientific
Three independent tests agreed. Drugs are refused for incomplete filings at 77% and factory problems at 69%, far more than for failing to work, at 25%.
The common assumption is that the FDA mostly rejects drugs that do not work. It is backwards. The bar is complete data and a clean facility.
4. The visible signals separate, but weakly
| Signal | In approvals | In rejections |
|---|---|---|
| First-time filing | 13% | 55% |
| Biologic | 7.5% | 26% |
| No completed phase 3 | — | 2× base |
| Zero registered trials | — | 2.7× base |
Real, and modest. Bigger trials do not save a drug. Rejected drugs had higher average enrollment. Evidence size is not the bar.
5. Public data alone cannot predict FDA rejections
Of 193 rejections, 144 had zero or one visible warning flag. Three-quarters looked clean on everything public and were rejected anyway.
They were killed by manufacturing and completeness data that is auth-walled or FOIA-locked. The ceiling of a public-data FDA predictor is low, and the reason is known exactly. The decisive data is hidden.
6. ICLR: taste shows up where the score stops deciding
Unlike the FDA, a conference ranks. It picks a favoured 30% from a field where most submissions are competent. A threshold frozen on 2020 to 2023 agreed with the real decision 87.4% of the time on 14,395 held-out papers.
Hold the score constant at the borderline, where the number does not decide, and the topic moves acceptance by more than 20 points.
| Topic, at identical borderline scores | Acceptance |
|---|---|
| Pre-training | ~73% |
| Time series, planning | ~72% |
| Diffusion | ~70% |
| Imitation, theory | ~68% |
| Adversarial training | ~52% |
| Graph neural networks | ~51% |
| Data augmentation | ~50% |
| Active learning, deep RL | ~49% |
Base acceptance is about 30%. Two papers with the same review scores are not judged the same.
7. SCOTUS returned nothing
The minority class is the affirm, about 25%. Using only strictly pre-decision structured features, a rule frozen on 2005 to 2020 scored 67.5% on 2021 to 2025. The always-reverse baseline is 74%.
An earlier 96% result was thrown out. It leaned on a feature coded from the outcome.
Affirm or reverse cannot be predicted from public metadata, because the deciding reasoning lives in the opinion text, which is written after the fact. A validation layer that only reports its wins is not a validation layer.
What this proved
- The machinery works. Date wall, freeze, held-out test, no peeking, all clean.
- Judgment is a stackable structure that can be frozen and stays stable over time.
- The FDA’s judgment is operational, and that falls out of the exhaust from three directions.
- Where the signal was in the data it won. Where the signal was hidden or written afterwards, it returned nothing. It was never fooled into a false positive.
What this did not prove
- Live reject-versus-approve prediction from public data. Blocked by the hidden manufacturing data.
- That the method captures taste. The FDA approves. It never champions. This tested the veto half, not the selection half. Taste needs a decider that ranks, not one that only refuses.
- Transfer to a private operator. All three subjects are public strangers.
Verdict
The right rehearsal, and it succeeded as one. The plumbing is proven and the core finding held on unseen data: a frozen judgment file stays true over time.
The predictor’s ceiling is set by data access, not by the method. The taste half had to be proven somewhere else, on a decider that selects.