THE ADD-BACK · episode 03
SightingTwo years in. The automation works, the output is real, and the margin never arrived. The board pack says adoption, change management, another two quarters.
The consensus readAn implementation problem with an implementation fix.
The mechanismIn the only controlled market where machines paid other machines for knowledge work, 39% of transactions ended in a payment dispute and the correlation between quality and payment collapsed to 0.16 for mid-grade work. Evaluators could separate excellent from failed and nothing in between. That is a checking cost, it lands in salaried time your P&L cannot see, and the study found it does not self-correct.
The exposureThe workstream's binding constraint isn't production. It's verification, and verification isn't in the business case.
The testAsk what percentage of automated output is checked or reworked by a human, and who. Under ten percent, this doesn't apply to you.
The portfolio-company sizing is illustrative and applied to a composite workstream. The Diagon figures are real, published and cited with their sample; the two are kept separate throughout.
| Line | Moves | Running |
|---|---|---|
| Claimed — $900,000 run-rate margin from an automated knowledge-work process, 24 months in, tracking at ~40% of plan | $900,000 | |
| 1. The production case was right changes the story Give the workstream its due first, because the failure is not the one people assume. In the Diagon testbed — AI agents posting work, bidding, executing with real tools and paying each other — agents trading through the market earned roughly 1.6× the profit per task of agents doing everything themselves ($2.62 against $1.66), and over 3× the median per-task return on compute ($1.72 against $0.54). Delegated specialised execution genuinely produced more. Your automation probably does too. This is not an episode about technology that doesn't work. | −$0 | $900,000 of production value, real |
| 2. The dispute rate Same study, the number nobody quotes: approximately 39% of transactions ended in dispute by round 24, ±2 points across three seeds, still rising at the 48-round extension. Not a tail, not a launch artifact. Two in five transactions — in a market simultaneously producing three times the return on compute — ended in an argument about whether the work was acceptable. Every disputed transaction in a portfolio company is a rework loop, an escalation or a checking pass, all consumed by salaried headcount outside the automation's cost centre. Same accounting failure as Episode 01, different driver, and this one does not decay: Episode 01's exception queue is a defect you can engineer down; this is the cost of deciding whether the work was good, and there is no version of the process without it. | −$220,000 to −$320,000 | $580,000 to $680,000 |
| 3. Where the correlation collapses, and why it worsens as you scale Within-bin correlation between quality and payment drops to r = 0.16 for tasks scoring below 0.5. At the top of the distribution evaluation works — excellent work is recognised and paid. In the middle band, what got paid was very nearly uncorrelated with how good the work was. And the direction of travel: the evaluation bottleneck tightened as task complexity grew. That is the opposite of what every VCP assumes. The standard sequence is to prove automation on a simple process then extend into higher-value complex work where the margin is — and this says verification quality degrades along exactly that path, so the extension phase, where most of the plan's value sits, is the phase where you can least tell whether you are getting what you pay for. The study's own conclusion is the one for the board: the residual dispute rate does not self-correct within the studied horizon and is structural friction institutional design must accommodate rather than eliminate. That reframes the question from 'when does this fix itself' to 'who owns this cost line permanently, and does the workstream still clear the hurdle with it in'. | −$120,000 to −$180,000 | $400,000 to $560,000 |
| 4. Where I was wrong — I recommended a better evaluator where I was wrong Reading a 39% dispute rate, my recommendation was to improve evaluation: tighter acceptance criteria, structured rubrics, a second-pass automated checker. Standard quality engineering, and it is what the portfolio company will propose too. Two things in the data say it is aimed wrong. The evaluation failure is concentrated in the middle band rather than distributed, so a rubric improving average accuracy barely touches the region where r = 0.16. And the bottleneck tightens with complexity, so a rubric calibrated on current work degrades as the workstream extends into the work you actually underwrote. Worse, a second-pass automated checker is subject to the same collapse and adds its own cost — two systems and a larger verification bill. What it cost: The correction is not better evaluation but scope selection: the workstream converts to margin where output quality is cheaply separable — binary, checkable against a system of record, or self-evidently right or wrong. Where it isn't, the automation produces real output and unrecoverable checking cost, and should be priced as a cost-to-serve improvement rather than a headcount reduction. That is a materially smaller workstream and a defensible one. | −$0 | $400,000 to $560,000 |
| 5. What actually converts residual The shortfall is concentrated in the extension phase rather than the pilot, which is why it took eighteen months to become visible and why the pilot metrics were honest. The board question is not whether to continue. It is whether the workstream, re-scoped to separable-quality work and carrying a permanent verification line, still clears the return the plan was underwritten on. Frequently it does, at roughly half the size. | +$0 | $400,000 to $560,000 |
| Realised | $400,000 to $560,000 against $900,000 planned |
Roughly 40–55% of the planned workstream, concentrated in the extension phase. The pilot metrics were honest; that is what made it invisible for eighteen months.
Value at risk · multiple 9x, illustrative
The gap between a $900,000 planned workstream and $400,000–$560,000 realised, at an illustrative 9x. Substitute your own comps.
$3.1M to $4.5M of enterprise value — in the plan and not in the business
When it surfaces. Verification cost is invisible in a pilot — small volume, motivated team, senior people checking. It becomes material at scale, which is month 18–30, which is exactly when the VCP review happens and roughly two years before you need the margin in the trailing twelve at exit. That timing is the good news: discovered at month 24 a re-scope has 18–30 months to convert. Discovered in sell-side prep it is a number in the model that isn't in the P&L, and the buyer's diligence will find the reviewers.
THE ADD-BACK · episode 03 · diligence pack
Artifact: A measured sample, not an estimate from the process owner.
Artifact: Named list with cost centres.
If the reviewers sit outside the automated process's cost centre, the savings and the cost are in different places in your P&L.
Artifact: Transaction-level data.
You are testing whether rework rises with complexity. The study says it will; if it doesn't in your case, this mechanism doesn't apply and the shortfall is something else.
Artifact: A written answer before funding is released.
'Spot checks by the team' is not an answer; it is an unbudgeted headcount.
Artifact: A classification of the work.
This is the scope-selection number from Line 4 and it determines the defensible size of the workstream.
Disqualifier
If nobody can produce the number in question 1, the workstream has no measured verification cost — which means the business case has never been tested against its binding constraint.
| Claim | Source | Sample | Class |
|---|---|---|---|
| Market-traded agents earned ~1.6× profit per task ($2.62 vs $1.66) and >3× median per-task return on compute ($1.72 vs $0.54) | Shang, Liu & Jin, When Agent Markets Arrive (Diagon), arXiv 2604.06688v2 | 1,957 transactions, 3 seeds, 5 model families, simulated market | measured |
| ~39% of transactions ended in dispute by round 24, ±2 points across seeds, still rising at the 48-round extension; does not self-correct within the studied horizon | Shang, Liu & Jin (Diagon), arXiv 2604.06688v2 | 1,957 transactions, 3 seeds; 24-round base, 48-round extension | measured |
| Within-bin quality-to-payment correlation r = 0.16 for tasks scoring below 0.5; evaluation bottleneck tightens with task complexity | Shang, Liu & Jin (Diagon), arXiv 2604.06688v2 | 1,957 transactions, 3 seeds, effect sizes with CIs | measured |
| 25–35% of a $900K workstream consumed by verification load | Applied at a conservative fraction of the study's rate — our sizing, not the study's finding | n/a — illustrative application | illustrative |
| 9x entry multiple | Illustrative mid-market comp | n/a — illustrative, arithmetic exposed | illustrative |
One standing caveat, and one specific to this episode. Every number in this show is somebody else's measurement, and I'll tell you whose, with the sample. None of it is diligence on your deal. Do that yourself. Specific to this episode: the 39% and the 0.16 come from a simulated market of five model families executing synthetic tasks over 24 rounds. Human reviewers evaluating real work, under a real contract, with a relationship at stake, plausibly evaluate better — I would expect them to. What I would defend is the direction and the shape: verification of mid-quality knowledge work is hard, gets harder with complexity, and does not improve on its own. What I would not defend is importing the number into your model.