Predicting Judgment from Behavioral Exhaust
Approval is the default. The refusal is the signal. A decider approves most of what it sees and shows its judgment only in the rare no.
So the test is simple. Take the reasons behind past refusals, weight them by how much they actually drive the decision, and run them against decisions the model has never seen.
What was built
| Source | Rows | What it holds |
|---|---|---|
| ICLR papers, 2020–2025 | 33,043 | Scores, accept or reject |
| ICLR reviews | 69,300 | Reviewer prose and sub-scores |
| Michelin Guide, five yearly snapshots | ~19,000/yr | Award tier and the published review |
| FDA decisions | 26,691 | Approve or reject |
| SCOTUS cases | 9,341 | Affirm or reverse |
Every headline number carries a 95% confidence interval from 1,000 bootstrap resamples. Cutoffs and hyperparameters were chosen on training data only.
1. A refusal is a pile-up, not a single cause
Two or three reasons co-fire in most FDA rejection letters. Single-reason refusals are 2% of nos.
| Reasons stacked | 0 | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|---|
| Rejection rate | 4% | 10% | 12% | 9% | 28% | 43% |
Base rate 5.5%. One reason barely moves the odds. Three or more push rejection from unlikely to likely. There is no threshold count. The decision is a weighted sum.
2. The loudest complaint is not the deadliest
Volume is how often a reason appears. Weight is how much it raises the chance of rejection, measured against held-out outcomes. The two rankings disagree.
| Reason | Volume | Weight | Verdict |
|---|---|---|---|
| Lacks novelty | 13.2% | +10.9 | True veto |
| Unconvincing motivation | 4.6% | +10.2 | True veto |
| Missing baselines | 6.5% | +6.5 | Moderate |
| Unclear writing | 14.8% | +6.2 | Loud but mild |
| Weak or no theory | 8.3% | +4.8 | Moderate |
| Insufficient experiments | 8.7% | +2.9 | Weak |
| Poor reproducibility | 3.6% | −1.4 | Reflex, not a veto |
Unclear writing is raised more than anything else and sits mid-pack on decisiveness. Poor reproducibility is raised often and slightly predicts acceptance.
Ranking by volume echoes what a decider says. Ranking by weight predicts what it does. The decisive-but-rare reasons are the moat. Nobody reading the published rulebook can find them. They only surface by testing reasons against real outcomes.
3. A new signal helps only if it is independent
| Model | Blind accuracy |
|---|---|
| Always reject | 60.7% |
| Average reviewer score | 86.7% |
| Shape of the individual scores | 87.2% |
| Plus sub-scores: soundness, presentation, contribution, confidence | 87.8% |
| Plus the written complaints | 86.3% |
| Plus reviewer disagreement | 86.8% |
The complaints and the disagreement are both real. Both are already inside the score the reviewer gave. They restate, they do not add. Five attempts to beat 87% failed the same way.
4. Freshness beats volume
Every training year run against every test year, forward and backward, plus random pooled splits.
| Training slice | Rejections caught |
|---|---|
| Stale and distant, 2020 to 2025 | 82.0% |
| Adjacent, 2024 to 2025 | 87.2% |
| Same period, 2025 to 2025 | 93.1% |
| Best fresh-adjacent, 2021 to 2024 | 95.3% |
More data is not better. Fresher, matched data is better. A random pool of all years reaches about 88% by sheer volume and loses to a matched slice. Training on 2025 to predict 2020 still reaches 88.5%, so the structure is real and not a one-year artifact.
A judgment file goes stale as judgment moves. Freshness of the training window is as powerful a lever as adding an independent signal.
5. The method transfers to a decider with no score
ICLR hands the model a reviewer score that already sits next to the decision. Michelin publishes no score. The reasons have to stand alone. 1,863 recurring qualities were pulled from the review text, every word and two-word phrase appearing in at least 0.5% of reviews.
| Move | AUC | Stars caught | False alarms |
|---|---|---|---|
| Text only, default cutoff | 0.902 | 67% | 9% |
| Cutoff tuned on the build year | 0.902 | 83% | 19% |
| Plus price tier, cuisine, country | 0.946 | 90% | 14.5% |
| Plus pooled stable years | 0.973 | 93.4% | 10.4% |
| Same recipe, backward to 2022 | 0.972 | 93.0% | 10.1% |
0.973 sits in a 95% interval of 0.970 to 0.975. Stars caught, 92.6 to 94.2. The cutoff was free: the default was the wrong operating point. Metadata added the most because it is genuinely independent of the words. The words say how good the cooking is. The metadata says what kind of restaurant it is.
6. Stated criteria are not revealed weights
Michelin publishes five criteria, including consistency and creativity. Here is what the exhaust says.
| Quality | Weight | |
|---|---|---|
| Ceremony. Tasting menu, sommelier, theatre | +0.54 | Rarest positive quality, strongest driver |
| Execution | +0.40 | 0.48 to 0.71 across the ladder |
| Refinement | +0.35 | |
| Ambition and creativity | +0.00 | Equal in starred and unstarred reviews |
| Locality | −0.22 | 0.88 down to 0.59 |
| Comfort | −0.36 | 0.34 down to 0.21 |
| Modesty | −0.42 | Top penalty, every year |
Creativity is weightless. Ceremony is the hidden driver. Phrases that read like virtues are near-guaranteed no-stars: daily specials 3%, market fresh 2%, noodle soup 0%, against a 21% base. Those are the bistro signal, not the star signal.
7. Execution gets you in the door. Becoming a monument gets you loved.
Every starred restaurant already passed the veto, so one star to three is pure selection. Mentions per 100 starred reviews.
| Quality | 1★ | 2★ | 3★ |
|---|---|---|---|
| Exceptional | 1.6 | 5.9 | 11.6 |
| World-class, renowned | 4.9 | 10.4 | 16.4 |
| Caviar | 3.8 | 6.9 | 12.3 |
| Signature | 5.5 | 9.1 | 13.0 |
| Remains, passing, decades | 0.6–1.3 | 0.6–1.6 | 3.4–5.5 |
The star is earned by execution. Precision, technique, finesse, craftsmanship. The third star is earned by something else. Temple, destination, flagship, decades, a world reputation. What keeps a restaurant at one star is being a type rather than a place. A hotel dining room. An outpost of a group. A modern cuisine kitchen. Competent, starred, and never championed.
That gap, measured blind on a decider that publishes no score, is the taste delta.
8. Drift is a rate, and it differs by decider
| Decider | Cost of stale training | 95% CI |
|---|---|---|
| ICLR | −5.2 points of rejection-catch | 4.6 to 5.8 |
| Michelin | −0.039 AUC over two years | 0.036 to 0.043 |
At ICLR, pooling years hurts. Taste moved and old years flatten the median. At Michelin, pooling helps. The top ten star drivers learned in 2022 and in 2026 overlap nine out of ten. Same lever, opposite direction, because the deciders are different.
The re-slicing cadence of a judgment file is not a universal setting. Measure the decider first. Fresh slices for one that drifts, pooled history for one that does not.
What this proved
- A refusal is a stack of separable reasons, and the stack can be frozen and re-run on decisions it never saw.
- Volume and weight are different numbers, and only weight predicts.
- The method transfers. Same seven steps, unrelated domain, no score to lean on, AUC 0.973 blind.
- Both halves came out of the same exhaust: the veto at 0.97, the selection at 0.90 to 0.95.
What this did not prove
At ICLR the strongest predictor is the reviewer score, and that score sits one step upstream of the chairs’ decision. Predicting acceptance from it is closer to reading the scoreboard than extracting judgment. ICLR validates the plumbing and the volume-versus-weight finding, both of which stand on their own. The clean extraction result, published text to decision with no score, is Michelin.
Nothing here tests transfer to one private operator inside their own company.
What was got wrong
- An early test scored papers by counting complaints. It lost to the baseline. Volume is not weight.
- A SCOTUS test scored 96% and was thrown out as circular. It leaned on a feature coded from the outcome.
- Five attempts to beat 87% by adding written reasons and disagreement all failed.
- Early runs pooled 2020 to 2023 to predict 2024 and flattened the signal.
- The first Michelin test predicted which restaurants would lose a star. AUC 0.52. Every published review is praise written after the verdict, so there is no negative signal to find.
- Seven hand-written buckets scored 0.72 and were nearly mistaken for the ceiling. Data-driven extraction from the same text reached 0.90, then 0.97.
Bootstrap intervals, added last, caught two overstatements made earlier in this study. The ICLR stale-training loss was first reported as a collapse to 71%. The rigorous figure is 82% against 87% fresh. A real five-point drift, not a twenty-four point collapse. And Michelin was described as not drifting at all. It drifts by 0.039 AUC over two years. Small, and significant. The findings survive both. The magnitudes were wrong, and rigour is what caught them.
Verdict
A decider’s judgment can be pulled out of its exhaust and used to predict its future refusals, blind, at 87 to 90% accuracy and 90 to 95% recall on the nos. The path to 95% is not a cleverer algorithm. It is three disciplines. Rank by weight, not volume. Add only independent signals. Keep the window current.
The last one matters most. Judgment moves, so the file that models it has to move with it.