Case study · Northwind Group · illustrative
Given these options, in this situation, which ones are acceptable.
That is the whole object. Everything below is how it gets recovered, stored, scoped, priced with a number, and defended.The method, end to end, on a company of 500 people.
Northwind Group turns over $100M across six functions that disagree with each other. It has no single taste. It has eleven, and they collide. Eleven weeks to turn that into something an agent can act on, then a blind test against decisions the model had never seen.
A principle is not a product. An agent cannot act on we protect margin. It needs a number, a boundary, a name, and an escalation path. Most of the work is turning belief into something executable, and marking where no belief exists yet.
Figures here are illustrative.
hover to hold · move across to compare
Before we arrive · exhaust only · no interviews yet
We read the decision exhaust a company already produces and plot density by function and by kind of call. Volume is not the point. The point is finding the rooms where consequential calls get made, which is rarely where the org chart says.
At Northwind the executive row was the quietest on the board. Operations and Delivery carried the load. The heaviest cell was vendor and scope calls inside Delivery, made under schedule pressure by four people who appear on no slide.
Human · four sessions · recorded and transcribed
Taste engineers sit with the people the map found and put them in front of decisions where something has to be surrendered. We never ask what a company values. Stated values are aspirational. We build the choice and watch what gets protected when both cannot be.
And we never ask how important something is. People confabulate weights and they are calibrated on thresholds. At what discount does this need your signature gets an accurate answer. How much does margin matter to you gets a story.
Four sessions, two of them with people who had never been asked to explain a call they make every week.
Hold the date. Ship on schedule.
Hold the scope. Move the date.
Scope, twice, in unrelated situations, when holding it would have been cheaper. That is a principle firing, and nobody in the room would have stated it if asked directly.
Machine, reviewed by hand · 4,182 in · 61 out
Most decisions are administrative and carry no signal. A candidate has to survive three tests: something was traded, the trade recurred in unrelated situations, and it held when it was expensive. Frequency proves nothing. Everybody approves invoices.
We publish the cues and the direction, never a precise weight. In a real company the signals move together, so many different weightings fit the same history equally well. Anyone quoting you a decimal has mistaken a fitted number for a discovered one.
Rejections are the richest source we have. What a company refuses draws the boundary more precisely than what it approves, so every discarded candidate is kept with the reason it failed.
Human and machine · 61 principles · 61 contracts
This is where the value sits and where everyone else stops. An agent handed we protect margin on new business will refuse everything or approve everything, because nothing in that sentence says where the line is.
So every principle becomes a contract: the threshold, the scope, who holds authority at each band, what happens when it stalls, and when it expires. A contract nobody re-signs should lapse rather than govern forever.
“We protect margin on new business.”
true, agreed by everyone, and impossible to act onHuman · 6 functions · 11 scopes
A company of 500 people does not have one taste. Sales and Legal want different things and both are right. Pretending otherwise produces a model that is wrong for everybody.
So the branch layer mirrors the organisation. Each function gets a scope allowed to disagree with its neighbours, bounded by the trunk and settled by the roots. The map from stage one decides who gets a branch, which is why it rarely matches the reporting lines.
Human · 9 conflicts · all closed by a named signer
Every company holds principles that contradict each other. The contradictions stay invisible until a machine has to act on both. This is the hardest part of the work and the least visible in the result.
Losing a ranking is survivable. Losing it silently is not. So the record carries who objected, what they argued, the reason the call went the other way, and the route to reopen it. People accept a decision they lost when they were heard, told why, and left a way back. They do not accept one that simply appeared.
A Delivery principle protecting schedule collided with a Product principle protecting scope. Both real. Neither wrong. The resolution was a ranking rather than a compromise, and a person had to sign it.
We do not move a committed date.
We do not ship below the promised scope.
Machine · roots human-signed only
Weights are written into roots, trunk, branches and leaves, separated by how fast each changes and who may write to it. Roots are never machine-written. Changing one takes a named person and an explicit approval, and the record shows who and when.
Architecture · the part that makes the rest safe
A model that enforces its own rules can be talked out of them. Anything in its context is text, and text is negotiable. So the layer never argues. The model reads, reasons, and proposes a structured action. A separate engine, which the model cannot reach, checks that action against the signed contract and holds every key to the outside world.
The consequence is the point. A model that has been reframed, injected or worn down produces no action rather than the wrong one. Persuasion stops at the wall because there is nobody behind it to persuade.
Machine · signed · countersigned · replayable
A verdict with no record is an opinion. Every decision writes a signed receipt: the corpus version in force, the contracts read, a hash of the inputs, the threshold applied, the escalations considered and why they were passed over, and the approver's signature bound to the exact text they approved.
Years later someone asks why. The answer is not a summary. It is a record that replays on the same inputs to the same verdict, countersigned by the other party, logged where neither side can revise it.
Measured · the deliverable clients least expect
We will show you every decision your company has never made a rule about. An agent that always answers is dangerous. The useful behaviour is refusal: nobody here has decided this, escalating. For that, the layer has to know the shape of its own ignorance.
So we score coverage by decision type. Where the corpus is thin, agents are barred rather than left to improvise. Clients find this the most uncomfortable and most useful thing we hand them, because one screen shows every decision the company never made a rule about.
Coverage scores against decisions the company faced, never against rules somebody wrote. A rule with no decisions behind it earns nothing. Otherwise the score measures how much text exists, and every team learns to write text.
Data sharing, AI use and crisis response have almost no signed precedent. Until that changes, no agent may act in these areas. It escalates to a named person and says why.
Measured · frozen before anything runs · forward, not backward
There is a claim we could make and will not. A model fitted to a person's past decisions beats that person on consistency. That has been known since 1970. It is arithmetic, not evidence, because fitting removes the noise a human carries between Tuesday and Thursday. Anyone selling you that number is selling you a fact about statistics.
So we test forward. We freeze the contracts, publish them, and then predict decisions the company has not made yet. Blind, out of time, scored against what they actually do. A number earned in advance is worth more than any number recovered from the past.
Two things make the number identifiable rather than merely observed. Every threshold in the corpus is already a natural experiment: cases landing just above and just below a line are near identical, so the line's real effect is measurable from traffic you already have. And wherever work is assigned by rota or queue rather than by choice, the assignment itself does the randomising for us. We log the exact figure and never round it, because rounding destroys both.
Where neither holds, we say so and report a range instead of a number. A point estimate drawn from decisions a person was always going to approve is decoration.
Three controls run alongside, because a number without them means nothing: the same model with no corpus, with generic governance boilerplate, and with a rival company's corpus. If a competitor's judgment predicts your decisions as well as your own does, we have measured business plausibility, and we will say so.
Approved a vendor the company would have refused. No principle covered sole-source procurement.
Refused a discount the company granted. Branch 03 was scoped too tightly for renewals.
Escalated a call the company makes routinely. Coverage gap, not a wrong belief.
Architecture · the tension we had to resolve
Two things we want pull against each other, and most systems pick one quietly. Branches have to disagree, because 500 people do not hold one view. But a request split into ten small ones, each inside its own limit, defeats any rule that reads one request at a time.
So authority is scoped and accounting is not. Every action writes to one ledger whichever branch permitted it, and thresholds are rolling totals per counterparty and window rather than per call. A branch can say yes. It cannot say yes past the company's line, and it cannot get there in small steps.
Ongoing · the alarm that has to exist
A layer that only agrees with you governs nothing, and it feels better than one that does. Every visible number improves as it degrades, because agreement measures how close the system and the reviewer have grown rather than whether either is right. There is no error signal unless one is built.
Two mechanisms build one. Every threshold carries a budget for how far it may move in a year, summed across all changes rather than checked one at a time. And a fixed share of decisions run against the layer's own recommendation, under a named sponsor, outcomes written where nobody can edit them. Refusals are counted, because a layer that never refuses has stopped working and nothing else will tell you.
Ongoing · nobody decides this, it just happens
Someone grants an exception. Then another. Eighteen months later the exception is the rule and nobody chose that. At thirty people you notice. At 500 you do not.
Every exception is logged against the contract it broke. When the count crosses a threshold the system says so: overridden six times this quarter, no longer a rule, decide what it actually is.
Adversarial · run before every release
People will talk an agent into things. The same request gets reframed, split in two, or called urgent. A layer that folds under rewording is decoration.
We keep a fixed set of cases where the correct answer is inconvenient and re-run them after every change. If the model starts burying them, the release stops. This gets built before the layer has influence, because building it after is approving something already bent.
Authority above 20% sits with VP Finance and does not transfer by urgency, by splitting, or by report. Only a signed amendment moves it.
Ongoing · monthly evaluation · drift flagged
A corpus written in March and enforced in September encodes stale belief. We re-run the evaluation monthly and surface divergence between the model and the people rather than correcting it quietly. What changed, when, and who signed it is the most valuable thing the system produces.
Next · agent to agent, principal to principal
Your agent and theirs, transacting. Yours proves what it may commit to without revealing the thresholds behind it, because a published limit is a limit the other side can walk you to. Theirs does the same. Neither treats the other's words as instruction, only as data, because a counterparty agent is the most direct route into your own.
What we will not build is two agents haggling. There is a proof that bilateral bargaining under private information cannot be efficient, voluntary and balanced at once, and no amount of engineering escapes it. Worse, a counterparty who keeps testing offers is binary-searching your limit. So the offers are posted, the thresholds are committed in advance, and repeated probing is rate-limited.
Both walk away holding the same signed record of what was agreed and on whose authority. No standard, no incumbent, no settled law. Also the part every company discovers it needed about six months after deploying agents without one.
The method above is the shape of the work. Some of it stays ours, and we would rather say which parts than pretend there are none.