Judgment Without Text
No text. 124 umpires, one rulebook strike zone, 2.2 million called pitches.
Every other subject revealed itself in writing. This one has none. And the target is not ambiguous taste, it is a rectangle in the rulebook that every decider is bound to. If a personal signature survives here, it is a floor, not a ceiling. Softer targets should show more.
What was built
Pitch location, the batter-specific zone, the call, the umpire. Ground truth from Statcast, independent of the umpire. No language of any kind.
1. The refusal is the unit
The strike is the default. The ball is the refusal. Most refusals prove nothing. A pitch a foot outside called a ball is not a judgment, the same way an approval nobody contested is not a judgment.
What carries signal is the considered no. A pitch the rulebook calls a strike, refused anyway. And its mirror, an out-of-zone pitch called a strike. Those are the wobblers, defined by geometry, not by a fitted model. The model does not get to pick its own hard cases.
2. Across the whole surface, the fingerprint disappears
A first pass modelled every called pitch. An umpire’s own history did not beat a pooled league model at predicting their own future calls. 49% win rate across 74 umpires. Mean log-loss lift about zero, interval spanning zero.
Read alone, that says personal judgment does not exist. It does. The full surface is mostly pitches every umpire calls identically. They carry no personal information and they vote in the average anyway. Diluting the signal with non-decisions manufactures a false null.
3. Isolated to the contested calls, it more than doubles
| Scored on | Top-1 identification | Mean self-percentile |
|---|---|---|
| Full edge surface | 8–9% | 27–31% |
| Wobblers only | 20% | top ~20th |
Give it 500 wobbler calls from one blind season and an umpire’s yes-no pattern names them out of about 80 peers one time in five. Chance is 1.2%. Sixteen times chance, using nothing but where they said yes and where they said no.
4. It is one person, visible from either edge
The refusal side and the grant side belong to the same umpire. Correlation of self-identifiability across the two is r=0.31 across 74 umpires. Ten land in the top decile on both sides at once, against fewer than one expected by chance. Not two unrelated quirks. One personal zone.
5. A frozen fingerprint ages
Trained on 2020 and 2021 only, then used to re-identify the same umpires each year after.
| Test year | Gap | Top-1 self-ID | Mean self-percentile |
|---|---|---|---|
| 2022 | 1 year | 6% | 22% |
| 2023 | 2 years | 9% | 25% |
| 2024 | 3 years | 8% | 29% |
| 2025 | 4 years | 3% | 34% |
| 2026 | 5 years | 2% | 41% |
Roughly 3 to 7 percentile points a year. A frozen profile is worth about three years before it is materially stale. The 2025 to 2026 drop is steeper than the trend. 2026 is the first season with challengeable calls. That is consistent with observability compressing individual variation toward the norm, and it does not prove it.
6. Neither pure model predicts the next call
Identification is not prediction. Naming an umpire from a batch of calls is a different question from calling one pitch. Pooled across 18,064 blind 2025 wobblers.
| Model | Accuracy | Log-loss |
|---|---|---|
| Always guess the majority | 60.3% | — |
| League, all umpires pooled | 66.7% | 0.6303 |
| Personal, one umpire’s own year | 66.1% | 0.6375 |
Personal beat league on 40% of umpires. That is the honest ceiling on identity alone. A fingerprint visible across many decisions is not enough to beat a well-trained crowd on any one decision.
7. The blend beats both
Those two findings are not in conflict. They are the two halves of one architecture. A shared prior, corrected by individual signal in proportion to the evidence for it.
blended = w · personal + (1−w) · league, w = n / (n + K)
n is the umpire’s own training volume. K is a shrinkage constant tuned on 2024 alone, by per-umpire cross-validation restricted to wobblers, never touching the blind year. K = 12,000. Typical personal weight, 0.25 to 0.30.
| Model, blind on 2025 | Accuracy | Log-loss |
|---|---|---|
| Always guess the majority | 60.3% | — |
| League only | 66.7% | 0.6303 |
| Personal only | 66.1% | 0.6375 |
| Blended | 66.6% | 0.6248 |
Best log-loss of the three, which is the metric that grades confidence rather than the coin flip. Beats pure league on 76% of umpires and pure personal on 78%. Mean improvement over league is +0.0054, 95% CI 0.0037 to 0.0074.
What this proved
- Individual judgment is extractable from pure behaviour, with no language, inside a rule every decider enforces identically.
- It lives in the contested calls almost exclusively. Forcing a thin personal model onto easy cases adds noise, not signal.
- The blend ratio is a measurable number. For a typical umpire, w is about 0.25. A quarter of a close call is personal. Three-quarters is inherited convention.
- K is a trust threshold you can calibrate. It answers how much history a decider needs before their own signal outweighs the default.
- Drift is measurable on a person, not just an institution. Half-life around three years.
What this did not prove
- Nothing about text-to-decision extraction. There is no text here.
- The blend used location, count-adjacent context and batter handedness only. No catcher framing, no velocity or movement, no sequencing. The real ceiling is above 66.6%.
- w ≈ 0.25 belongs to a domain where the target is public and heavily conventionalised. Softer targets should run higher. What generalises is the way to measure it, not the number.
- The 2026 challenge data would score each umpire’s deviations against machine-adjudicated truth, separating judgment that is personal from judgment that is personal and right. Not pulled yet.
Verdict
A personal signature survives inside a legislated rule, and it concentrates where the rule runs out. Used alone it loses to the crowd. Used as a shrinking correction on top of the crowd, it wins.
That is the shape a judgment file should take. Not one model. A pair, with a weight between them that gets measured rather than assumed.