THE ADD-BACK · episode 10
SightingAsk a fund how AI is performing across its portfolio and the answer is almost always one word: uneven. Some companies show real EBITDA impact, most don't, and nobody says why.
The consensus readVariance in execution. Better vendors, better change management, better operating partners.
The mechanismUneven isn't a distribution of outcomes. It's a distribution of measurability. Attribution needs three things — a pre-change baseline, a defensible counterfactual, and cheaply separable output quality — and only small, low-value use cases have all three. So the initiatives that show attributable impact are the ones where attribution was possible, not the ones that worked best.
The exposureYour view of what works in AI is being formed by a biased sample, and the bias runs toward the least valuable work.
The testList the initiatives with proven impact and check whether they're the largest ones. If they are, ignore me.
Attribution survivorship
The initiatives visible in your results are the ones that could be measured, not the ones that worked.
The test. A correlation you can run in an afternoon. Plot each initiative's demonstrated impact against its underwritten value. If the proven wins cluster at the low-value end, you're reading a measurability distribution and calling it a performance distribution.
The portfolio pattern is a composite. Flyvbjerg's database figures and the Diagon correlation are real and cited with their samples; the market-context claim is internal research and flagged as such.
| Line | Moves | Running |
|---|---|---|
| Claimed — A fund-level read informing the next allocation of value-creation capital | AI delivers real but uneven EBITDA impact | |
| 1. The three conditions for attribution changes the story Each has been an episode in this set, which is why this one closes the arc. A pre-change baseline (Episode 02): perishable, capturable only before implementation, routinely not captured because in the first hundred days nobody is thinking about a diligence process four years out. A defensible counterfactual (Episode 04): any saving characterised as 'enabled by' or 'contributed to' has no counterfactual, cannot be tied to a document, and will not be credited by a hostile reader. Cheaply separable output quality (Episode 03): where correctness is checkable against a system of record, impact is demonstrable; where it requires judgment, quality-to-payment correlation collapses to r = 0.16 in the mid band and the evaluation bottleneck tightens as complexity grows. Which initiatives satisfy all three? Invoice matching, document classification, routine reconciliation, standard-form drafting — high volume, binary correctness, short cycles, small unit value. Which fail at least one? Anything touching pricing judgment, complex quoting, customer resolution, underwriting support — the work with real margin in it. | −$0 | attribution available on the cheap end, unavailable on the expensive end |
| 2. The bias that produces changes the story Follow the consequence through a portfolio review. Initiatives showing clean attributable impact are disproportionately the ones with all three conditions, so they are disproportionately small. Large initiatives either show nothing or show contested numbers a sceptic can dismantle, which in a review reads the same as showing nothing. The fund observes: small automation works, large automation is uneven. And it allocates accordingly — more of what's proven, more caution on the ambitious. But the observation is an artifact of the measurement, not of the performance. Some large initiatives are working and cannot demonstrate it; some are failing and cannot be caught. Both are invisible in the same way, which is the property that makes this dangerous rather than merely inconvenient: the failure mode is symmetric, so you cannot correct for it by assuming pessimism or optimism. The word 'uneven' is doing the concealing — it describes a spread in results, when what is actually spread is your ability to see them. | −$0 | allocation decisions being made on a biased sample |
| 3. The falsification test, borrowed changes the story You don't have to take this on argument, and the method comes from the best-evidenced work on large-project failure. Flyvbjerg's megaproject database — more than 16,000 projects, with only about 8.5% meeting cost and schedule and roughly 0.5% also delivering promised benefits, and cost overruns roughly constant across ninety years, 104 countries and six continents — includes a genuinely elegant argument that overruns are not technical error: if inaccurate estimates were caused by technical causes, errors in overestimating costs would have been of the same size and frequency as errors in underestimating. They aren't. The distribution is one-sided, which implicates something structural rather than random. Run the same test on your own portfolio. If AI impact were genuinely uneven in the performance sense you would expect roughly symmetric surprise — about as many initiatives materially beating their case as missing it, because random execution variance is two-tailed. If instead almost everything either lands near its case or disappears into 'hard to attribute', with very few upside surprises, that asymmetry is not performance variance. It is a measurement floor, and the missing tail is the work you can't see. That is a check you can run on a spreadsheet you already have. | −$0 | allocation decisions being made on a biased sample, now testable |
| 4. Where I was wrong — I read it as a capability distribution where I was wrong My first read of 'uneven impact' was the natural one: variance in vendor quality, operating-partner attention and portfolio-company readiness. The recommendation followed — tighten vendor selection, standardise readiness assessment, concentrate operating-partner time on the weaker companies. Every element of that is aimed at a performance distribution. What it cost: If the spread is in measurability, none of it moves the number — and worse, it reinforces the bias. Standardising vendor selection on demonstrated impact selects vendors who do measurable work, which is the cheap end. You would systematically optimise the portfolio toward low-value automation while believing you were raising quality. That is the part that changed my view of the priority: this is not an analysis error with a modest cost, the recommendation actively compounds the problem, and it is the recommendation a competent person reaches first. The correction sits in Episode 02 and belongs at fund level: baseline capture becomes a condition of value-creation funding. No baseline, no release of capital. A few days of analyst time per initiative, and the only intervention that widens what you can see rather than optimising within it. | −$0 | unchanged, priority raised |
| 5. What the bias costs residual The fund isn't losing money on the initiatives it can measure — those are real and they work. It is losing the ability to allocate. With attribution available only at the cheap end, the value-creation budget drifts toward small, safe, demonstrable work, which is defensible in every individual decision and, aggregated over a hold period, means the operational alpha the fund is now judged on gets pursued where it is easiest to prove rather than where it is largest. | +$0 | a portfolio that optimises for provable impact rather than for impact |
| Realised | A portfolio that optimises for provable impact rather than for impact, one reasonable decision at a time |
No dollar figure is offered, and that is deliberate — see the downside.
Value at risk · multiple None stated — and the omission is the point. Every other episode in this set uses an illustrative 9x; this one does not, because the exposure is unobservable by construction.
There isn't a clean dollar figure and manufacturing one would breach this show's own evidence rule. The exposure is the difference between the operational alpha you pursued and the one you'd have pursued with unbiased information, which cannot be measured from inside the biased sample. The illustrative multiple used elsewhere in this set would add false precision here.
Unquantifiable by construction — stated as such rather than estimated
When it surfaces. It doesn't. This is the only episode in the set with no discovery event, because the failure mode is invisibility itself. There is no retrade, no write-down, no buyer's diligence finding — just a fund that spent a hold period getting steadily better at proving small things.
Why there is no number here
Every other episode in this set prices its exposure at an illustrative 9x. This one does not, and the absence is the finding rather than a gap in the work. Attribution survivorship is a bias in what you can see, so its cost is the difference between the operational alpha you pursued and the one you would have pursued with unbiased information - a quantity that cannot be measured from inside the biased sample. Any figure offered here would have to be derived from the very initiatives the mechanism says are unrepresentative. Publishing a number would make this artifact more persuasive and less true, and this show's standing caveat is that every number comes with its sample. This one has none, so it gets no number.
THE ADD-BACK · episode 10 · diligence pack
Artifact: A two-column table.
If demonstrated impact clusters at the low-value end, the term applies.
Artifact: Two counts.
Per Line 3 you are testing for symmetry. A one-sided distribution means a measurement floor, not execution variance.
Artifact: One word per initiative.
Converts a vague category into a fixable list, and the answer is usually 'baseline'.
Artifact: A line in the funding process.
Cheapest item in this pack and the only one that changes what's visible next year.
Artifact: An honest inventory.
Some is recoverable from system data; the informal layer is not, and knowing which is which prevents a false sense of coverage later.
Disqualifier
If item 1 shows demonstrated impact spread evenly across the value range, attribution is working in this portfolio and this episode doesn't apply. That would be an unusual and genuinely good result.
| Claim | Source | Sample | Class |
|---|---|---|---|
| Only ~8.5% of projects met cost and schedule and ~0.5% also delivered promised benefits; cost overruns roughly constant across ninety years; one-sided error distribution implicates structural rather than technical causes | Flyvbjerg, Oxford megaproject database | More than 16,000 projects, 104 countries, six continents, ~90 years | measured |
| Quality-to-payment correlation r = 0.16 in the mid band; evaluation bottleneck tightens with complexity | Shang, Liu & Jin, When Agent Markets Arrive (Diagon), arXiv 2604.06688v2 | 1,957 transactions, 3 seeds, 5 model families, simulated market | measured |
| Returns now depend far more on operational value creation than financial engineering; real attributable EBITDA impact from AI inside portfolio companies remains uneven across the industry | Internal market research | n/a — not a published measurement; specifics held loosely, direction only | reported |
| A nine-portfolio-company fund observing 'uneven' AI impact | Composite | n/a — composite | composite |
One standing caveat. Every number in this show is somebody else's measurement, and I'll tell you whose, with the sample. None of it is diligence on your deal. Do that yourself.