Read the validation studies here
Product
Edin CloudEmulator
What is taste engineering?
Taste Models
Your Taste ModelOur Taste HarnessBuilding Taste ModelsIntegrations & Security
Company
The TeamOur MissionResearchEdin LabsEdin’s ThesisContact
Request a demo
← Research
Editorial taste

Case Study Zero

Finding a difficult taste test in YouTube editorial decisions12 September 2026

Most explainer videos are a talking head that barely changes. The host delivers the same scripted line four or five times, the takes come out nearly identical, and one of them has to be chosen. That is editorial decision making at its most particular, and it makes one of the most unusual tests of taste and judgment we could find.

Abstract. We used footage from an art history YouTube explainer series, recovering 2,974 talking head editorial choices from five episodes. We did it by aligning transcripts of the raw footage against the finished cuts, with no project files and no manual labelling. In the 183 cases where a line was shot more than once and exactly one take survived, nothing inside a take predicted the choice: not the words, not pitch, energy or pace, not eyeline or expression. The taste sits only in how a take stands against its rivals. When we wrote this into a file and handed it to a general-purpose model, it moved that model from chance to the host’s own reliability.

+0.0 pts

A model with a taste emulator beat a blind model at editorial decisions by 16 points.

Identical model, identical 183 choices, identical prompt. The only difference was whether the host’s taste came with it. Blind it scored 37.7%, barely above the 32.0% you get by guessing. With the file, 54.1%.

95% CI [+6.0, +26.8] · interval excludes zero · qwen3:14b, temperature 0

What was built

0Decisions recovered
0Episodes
0Contested choices
0Bytes in the file

Every take that did not make the cut was rejected by one person for reasons that were never recorded. No project files, no edit decision lists, no notes. The only surviving trace of the decision is the difference between what was shot and what went out.

So we transcribed every raw camera take and every finished cut, then aligned the two word streams. Anything that survived into the final is kept, everything else is cut. The check that the labels are sound is how much of each finished episode traces back to an identified moment in the raw.

EpisodeRaw segmentsKeptFinal traced to raw
Episode 132926%91%
Episode 256220%96%
Episode 443030%88%
Episode 553721%87%
Episode 71,11626%92%

It holds on all five independently, not only on the episode the method was built against. No human labelled anything.

1. Where the taste actually sits

Because the piece is scripted, what belongs in the episode is settled before the edit begins. The 75% cut rate is arithmetic. A hundred minutes of coverage has to fit a twenty-minute script, so most of it goes regardless of quality.

What is left, once the script holds the content constant, is the one decision he alone makes. The script demands this line, he recorded it four times, and he keeps one. That choice is the taste, and it is the only place to look for it.

A controlled comparison, by accident

Content, speaker, script, camera, lighting and day are all held constant by the way he works. The only thing that varies between the attempts is the performance and the choice. Every earlier study had to construct its contested set. Here he built it himself, by shooting the line again.

2. Nothing inside a take predicts the choice

Three modalities, each scored blind, each holding out whole takes or whole episodes.

SignalBlind AUCVerdict
The words of the line0.508null
Delivery: pitch, energy, rate, contour0.526null
Vision: eyeline, expression, visible flubno variancenull
Relational: position among its rivals0.613clears chance
Relational and delivery together0.624clears chance
Three honest nulls

Pitch height, range, contour, final fall, loudness, dynamics and speaking rate were measured with voiced-filtered tracking and normalised inside each take. Across 138 contested groups not one separated on both a sign test and a Wilcoxon test. A vision model returned looking into the lens on 97% of 560 frames and never once flagged a flub, which says more about a locked-off camera than about the model. Fifteen text markers mined from the transcripts all landed between 49% and 54%.

Every intuitive explanation fails. He keeps the better-sounding take, the more expressive read, the one where he looked right: none of it survives measurement. The things a person would name as their reasons turn out to carry no weight, which is what every earlier study in this program found in its own decider.

3. One axis carries everything

His taste operates on the comparison between takes, not on any take by itself. Turned into a decision rule and scored the way an agent would use it:

RulePicks his takevs chance
First attempt21.9%below
Longest take30.1%below
Cleanest take, fewest stalls35.1%below
Closest to the written script40.4%+8.4
A 17-feature learned ranker42.0%+10.0
Last attempt48.1%+16.1

The 17-feature model scores worse than the one-line rule. 183 decisions cannot support 17 parameters. Every minimal model that beats the rule contains a position term, and every model without one falls to chance or below. Adding a second or third independent signal buys about one point, with an interval that swallows it.

One axis, not a stack

Across every other subject in this program a decision broke into several weighted reasons. This taste does not. It runs on a single axis, and adding anything to it earns nothing.

4. The decisive test

One local model, qwen3:14b, run twice over the identical 183 choices at temperature zero. Each prompt lists the competing takes in the order they were recorded and asks which one he kept. The first arm gets only that. The second gets the identical prompt with a 1,455-byte file prepended. Nothing else differs.

Chance
0.0%
Blind agent
no file
0.0%
The rule alone
executed mechanically
0.0%
Agent with the Emulator
0.0%
Held out by episode, with the file derived without ever seeing the test episode: 50.3%, a +12.6 point lift, 95% CI [+2.7, +22.4].
The result

A general model handed the raw material and no knowledge of the decider performs at 37.7%, barely above chance. His judgment is not inferable from the content. Handed a 1.5 KB text file it reaches 54.1%, a lift of +16.4 points with a 95% confidence interval of [+6.0, +26.8] that excludes zero.

5. The lift is judgment, not leaked statistics

The full file states its own accuracy and the chance baseline. Supplying those risks telling the model roughly how often the last take wins, which is information about the answer rather than about the decider. So the file was rebuilt with every number stripped out, leaving only the decision unit, the named checks, the rule, the exception and the list of what was tested and found not to matter.

File given to the agentAccuracy95% CI
None37.7%[31, 45]
Full, including its own accuracy figures53.0%[45, 60]
Stripped, judgment only54.1%[47, 61]

Removing every number changes nothing, a difference of 1.1 points with a 95% CI of [−4.4, +2.2]. The numbers contribute nothing and the judgment carries all of it.

The structure is worth ten points

A thinner file, with the named checks stripped out and only the conclusion left in, dropped to 39.9%. Same holdout, same rule, same model. Handing an agent the answer on its own is much weaker than handing it the answer surrounded by the considerations, their limits, and the list of what was ruled out.

6. Meeting the ceiling

The subject estimates his own self-consistency at roughly 50%. Shown the same competing takes twice, he would choose the same one about half the time. That is his estimate and it has not been measured, which is the largest open item in this study.

If it is close, 54.1% against a 50% ceiling means the file has captured most of what can be captured. No model exceeds a decider’s own reliability, so a 90% target was never available. He cannot hit 90% against himself.

What that reframes

When roughly half the rejected takes were acceptable anyway, asking whether the agent picked his exact take sets the bar in the wrong place. The question that matters is whether it picked a take he would accept. That number is necessarily higher, and an assistant who selects an acceptable take has done the job.

What this proved

What this did not prove

Verdict

One person’s taste, recovered from work he had already done, reduced to a single measured axis, written into a file smaller than this page, and handed to a model that had never seen him. The model went from guessing to choosing the way he chooses.

The file holds no weights and no training data. It is a document, readable by the person it describes, and honest about what it cannot do.