Read the validation studies here
Product
Edin CloudEmulator
What is taste engineering?
Taste Models
Your Taste ModelOur Taste HarnessBuilding Taste ModelsIntegrations & Security
Company
The TeamOur MissionResearchEdin LabsEdin’s ThesisContact
Request a demo
← Research
Ball-and-strike calls

Judgment Without Text

MLB umpires1 September 2026

No text. 124 umpires, one rulebook strike zone, 2.2 million called pitches.

Every other subject revealed itself in writing. This one has none. And the target is not ambiguous taste, it is a rectangle in the rulebook that every decider is bound to. If a personal signature survives here, it is a floor, not a ceiling. Softer targets should show more.

What was built

2,205,584Called pitches
15,107Games
124Umpires
2020–2026Seasons

Pitch location, the batter-specific zone, the call, the umpire. Ground truth from Statcast, independent of the umpire. No language of any kind.

1. The refusal is the unit

The strike is the default. The ball is the refusal. Most refusals prove nothing. A pitch a foot outside called a ball is not a judgment, the same way an approval nobody contested is not a judgment.

What carries signal is the considered no. A pitch the rulebook calls a strike, refused anyway. And its mirror, an out-of-zone pitch called a strike. Those are the wobblers, defined by geometry, not by a fitted model. The model does not get to pick its own hard cases.

2. Across the whole surface, the fingerprint disappears

A first pass modelled every called pitch. An umpire’s own history did not beat a pooled league model at predicting their own future calls. 49% win rate across 74 umpires. Mean log-loss lift about zero, interval spanning zero.

Read alone, that says personal judgment does not exist. It does. The full surface is mostly pitches every umpire calls identically. They carry no personal information and they vote in the average anyway. Diluting the signal with non-decisions manufactures a false null.

3. Isolated to the contested calls, it more than doubles

Scored onTop-1 identificationMean self-percentile
Full edge surface8–9%27–31%
Wobblers only20%top ~20th
What this means

Give it 500 wobbler calls from one blind season and an umpire’s yes-no pattern names them out of about 80 peers one time in five. Chance is 1.2%. Sixteen times chance, using nothing but where they said yes and where they said no.

4. It is one person, visible from either edge

The refusal side and the grant side belong to the same umpire. Correlation of self-identifiability across the two is r=0.31 across 74 umpires. Ten land in the top decile on both sides at once, against fewer than one expected by chance. Not two unrelated quirks. One personal zone.

5. A frozen fingerprint ages

Trained on 2020 and 2021 only, then used to re-identify the same umpires each year after.

Test yearGapTop-1 self-IDMean self-percentile
20221 year6%22%
20232 years9%25%
20243 years8%29%
20254 years3%34%
20265 years2%41%

Roughly 3 to 7 percentile points a year. A frozen profile is worth about three years before it is materially stale. The 2025 to 2026 drop is steeper than the trend. 2026 is the first season with challengeable calls. That is consistent with observability compressing individual variation toward the norm, and it does not prove it.

6. Neither pure model predicts the next call

Identification is not prediction. Naming an umpire from a batch of calls is a different question from calling one pitch. Pooled across 18,064 blind 2025 wobblers.

ModelAccuracyLog-loss
Always guess the majority60.3%
League, all umpires pooled66.7%0.6303
Personal, one umpire’s own year66.1%0.6375

Personal beat league on 40% of umpires. That is the honest ceiling on identity alone. A fingerprint visible across many decisions is not enough to beat a well-trained crowd on any one decision.

7. The blend beats both

Those two findings are not in conflict. They are the two halves of one architecture. A shared prior, corrected by individual signal in proportion to the evidence for it.

blended = w · personal + (1−w) · league,   w = n / (n + K)

n is the umpire’s own training volume. K is a shrinkage constant tuned on 2024 alone, by per-umpire cross-validation restricted to wobblers, never touching the blind year. K = 12,000. Typical personal weight, 0.25 to 0.30.

Model, blind on 2025AccuracyLog-loss
Always guess the majority60.3%
League only66.7%0.6303
Personal only66.1%0.6375
Blended66.6%0.6248
The load-bearing result

Best log-loss of the three, which is the metric that grades confidence rather than the coin flip. Beats pure league on 76% of umpires and pure personal on 78%. Mean improvement over league is +0.0054, 95% CI 0.0037 to 0.0074.

What this proved

What this did not prove

Verdict

A personal signature survives inside a legislated rule, and it concentrates where the rule runs out. Used alone it loses to the crowd. Used as a shrinking correction on top of the crowd, it wins.

That is the shape a judgment file should take. Not one model. A pair, with a weight between them that gets measured rather than assumed.