Read the validation studies here
Product
Edin CloudEmulator
What is taste engineering?
Taste Models
Your Taste ModelOur Taste HarnessBuilding Taste ModelsIntegrations & Security
Company
The TeamOur MissionResearchEdin LabsEdin’s ThesisContact
Request a demo
← Research
Peer review & restaurant awards

Predicting Judgment from Behavioral Exhaust

ICLRMichelin Guide31 August 2026

Approval is the default. The refusal is the signal. A decider approves most of what it sees and shows its judgment only in the rare no.

So the test is simple. Take the reasons behind past refusals, weight them by how much they actually drive the decision, and run them against decisions the model has never seen.

What was built

SourceRowsWhat it holds
ICLR papers, 2020–202533,043Scores, accept or reject
ICLR reviews69,300Reviewer prose and sub-scores
Michelin Guide, five yearly snapshots~19,000/yrAward tier and the published review
FDA decisions26,691Approve or reject
SCOTUS cases9,341Affirm or reverse

Every headline number carries a 95% confidence interval from 1,000 bootstrap resamples. Cutoffs and hyperparameters were chosen on training data only.

1. A refusal is a pile-up, not a single cause

Two or three reasons co-fire in most FDA rejection letters. Single-reason refusals are 2% of nos.

Reasons stacked012345
Rejection rate4%10%12%9%28%43%

Base rate 5.5%. One reason barely moves the odds. Three or more push rejection from unlikely to likely. There is no threshold count. The decision is a weighted sum.

2. The loudest complaint is not the deadliest

Volume is how often a reason appears. Weight is how much it raises the chance of rejection, measured against held-out outcomes. The two rankings disagree.

ReasonVolumeWeightVerdict
Lacks novelty13.2%+10.9True veto
Unconvincing motivation4.6%+10.2True veto
Missing baselines6.5%+6.5Moderate
Unclear writing14.8%+6.2Loud but mild
Weak or no theory8.3%+4.8Moderate
Insufficient experiments8.7%+2.9Weak
Poor reproducibility3.6%−1.4Reflex, not a veto

Unclear writing is raised more than anything else and sits mid-pack on decisiveness. Poor reproducibility is raised often and slightly predicts acceptance.

What this means

Ranking by volume echoes what a decider says. Ranking by weight predicts what it does. The decisive-but-rare reasons are the moat. Nobody reading the published rulebook can find them. They only surface by testing reasons against real outcomes.

3. A new signal helps only if it is independent

ModelBlind accuracy
Always reject60.7%
Average reviewer score86.7%
Shape of the individual scores87.2%
Plus sub-scores: soundness, presentation, contribution, confidence87.8%
Plus the written complaints86.3%
Plus reviewer disagreement86.8%

The complaints and the disagreement are both real. Both are already inside the score the reviewer gave. They restate, they do not add. Five attempts to beat 87% failed the same way.

4. Freshness beats volume

Every training year run against every test year, forward and backward, plus random pooled splits.

Training sliceRejections caught
Stale and distant, 2020 to 202582.0%
Adjacent, 2024 to 202587.2%
Same period, 2025 to 202593.1%
Best fresh-adjacent, 2021 to 202495.3%

More data is not better. Fresher, matched data is better. A random pool of all years reaches about 88% by sheer volume and loses to a matched slice. Training on 2025 to predict 2020 still reaches 88.5%, so the structure is real and not a one-year artifact.

What this means

A judgment file goes stale as judgment moves. Freshness of the training window is as powerful a lever as adding an independent signal.

5. The method transfers to a decider with no score

ICLR hands the model a reviewer score that already sits next to the decision. Michelin publishes no score. The reasons have to stand alone. 1,863 recurring qualities were pulled from the review text, every word and two-word phrase appearing in at least 0.5% of reviews.

MoveAUCStars caughtFalse alarms
Text only, default cutoff0.90267%9%
Cutoff tuned on the build year0.90283%19%
Plus price tier, cuisine, country0.94690%14.5%
Plus pooled stable years0.97393.4%10.4%
Same recipe, backward to 20220.97293.0%10.1%

0.973 sits in a 95% interval of 0.970 to 0.975. Stars caught, 92.6 to 94.2. The cutoff was free: the default was the wrong operating point. Metadata added the most because it is genuinely independent of the words. The words say how good the cooking is. The metadata says what kind of restaurant it is.

6. Stated criteria are not revealed weights

Michelin publishes five criteria, including consistency and creativity. Here is what the exhaust says.

QualityWeight
Ceremony. Tasting menu, sommelier, theatre+0.54Rarest positive quality, strongest driver
Execution+0.400.48 to 0.71 across the ladder
Refinement+0.35
Ambition and creativity+0.00Equal in starred and unstarred reviews
Locality−0.220.88 down to 0.59
Comfort−0.360.34 down to 0.21
Modesty−0.42Top penalty, every year

Creativity is weightless. Ceremony is the hidden driver. Phrases that read like virtues are near-guaranteed no-stars: daily specials 3%, market fresh 2%, noodle soup 0%, against a 21% base. Those are the bistro signal, not the star signal.

7. Execution gets you in the door. Becoming a monument gets you loved.

Every starred restaurant already passed the veto, so one star to three is pure selection. Mentions per 100 starred reviews.

Quality1★2★3★
Exceptional1.65.911.6
World-class, renowned4.910.416.4
Caviar3.86.912.3
Signature5.59.113.0
Remains, passing, decades0.6–1.30.6–1.63.4–5.5

The star is earned by execution. Precision, technique, finesse, craftsmanship. The third star is earned by something else. Temple, destination, flagship, decades, a world reputation. What keeps a restaurant at one star is being a type rather than a place. A hotel dining room. An outpost of a group. A modern cuisine kitchen. Competent, starred, and never championed.

That gap, measured blind on a decider that publishes no score, is the taste delta.

8. Drift is a rate, and it differs by decider

DeciderCost of stale training95% CI
ICLR−5.2 points of rejection-catch4.6 to 5.8
Michelin−0.039 AUC over two years0.036 to 0.043

At ICLR, pooling years hurts. Taste moved and old years flatten the median. At Michelin, pooling helps. The top ten star drivers learned in 2022 and in 2026 overlap nine out of ten. Same lever, opposite direction, because the deciders are different.

What this means

The re-slicing cadence of a judgment file is not a universal setting. Measure the decider first. Fresh slices for one that drifts, pooled history for one that does not.

What this proved

What this did not prove

At ICLR the strongest predictor is the reviewer score, and that score sits one step upstream of the chairs’ decision. Predicting acceptance from it is closer to reading the scoreboard than extracting judgment. ICLR validates the plumbing and the volume-versus-weight finding, both of which stand on their own. The clean extraction result, published text to decision with no score, is Michelin.

Nothing here tests transfer to one private operator inside their own company.

What was got wrong

Two numbers corrected

Bootstrap intervals, added last, caught two overstatements made earlier in this study. The ICLR stale-training loss was first reported as a collapse to 71%. The rigorous figure is 82% against 87% fresh. A real five-point drift, not a twenty-four point collapse. And Michelin was described as not drifting at all. It drifts by 0.039 AUC over two years. Small, and significant. The findings survive both. The magnitudes were wrong, and rigour is what caught them.

Verdict

A decider’s judgment can be pulled out of its exhaust and used to predict its future refusals, blind, at 87 to 90% accuracy and 90 to 95% recall on the nos. The path to 95% is not a cleverer algorithm. It is three disciplines. Rank by weight, not volume. Add only independent signals. Keep the window current.

The last one matters most. Judgment moves, so the file that models it has to move with it.