Cost per Completed Outcome is the fully loaded cost of one resolved unit of a business's actual output — compute, Steward intervention time, and residual Coordination Tax — divided by the number of outcomes actually completed, not merely executed, reported by Task Tier rather than as a single blended figure. It is not a proxy. It is not a capability score. It is the number a business needs to know before it can claim autonomy pays for itself, and it is, as far as the published record shows, a number nobody in the current wave of agent launches has produced.
What the industry measures instead
Every frontier lab currently reports capability in one of three currencies: tokens consumed, benchmark scores, or completion claims. None of the three is a business metric. Tokens measure computation, not value delivered — a task solved efficiently and a task solved wastefully can consume the same tokens and produce entirely different economics. Benchmark scores measure performance on a fixed evaluation set, not on the business's actual task distribution, which is why a system can top a leaderboard and still fail the specific exception pattern that determines whether a real deployment holds up. Completion claims measure that a task finished, not what finishing it cost relative to what the business earned from it.
Memo #10: The Stewardship Model established the two metrics Arco already tracks that come closer: MTTI, the average time between required human interventions, with a target greater than 72 hours, and Escalation Rate, the proportion of task executions that require escalation to a human Steward. Both are necessary. Neither is sufficient. MTTI and Escalation Rate measure whether a system is running autonomously — how often the architecture holds without a human decision. They say nothing about whether it is running economically. A system can post an excellent MTTI and still be unprofitable per completed outcome, if the interventions it does require are expensive enough, or if the definition of "completed" has quietly narrowed to exclude the work that made the number look good. Process health and unit economics are different questions, and the current metric set only answers the first one.
The formula only works if both halves are defined precisely, because both halves are where the number gets gamed.
Completed means the outcome the business exists to produce — the resolved customer request, the closed transaction, the delivered report — not the step that got handed off to the next stage of the pipeline. A task that a T1 agent processes correctly but that later requires rework, reversal, or a customer complaint was not completed. It was executed. The distinction matters because a system optimising for throughput at the task level can post excellent numbers while producing a business result that never actually lands.
Cost is the full stack, not the visible slice of it. It is compute at the point of execution, plus the fully loaded cost of every Steward intervention the outcome required, plus the residual Coordination Tax — the meetings, approvals, and manual handoffs — that survives at whatever tier of the task the human is still involved in. A business that reports only its compute spend and excludes Steward time from the numerator is not measuring Cost per Completed Outcome. It is measuring the cost of the part of the business it has already automated, which is a different and much friendlier number.
Where the number comes from
The cost stack is not flat across a business — it is set by the Task Tier the outcome sits in. At T1, fully automatable work, Escalation Rate runs approximately 1:100 and Workforce Arbitrage — the measurable cost delta between a human workforce and an equivalent agentic stack performing the same work — reaches 37 to 50 times. Cost per Completed Outcome at T1 is close to pure compute, because escalation is rare and the Steward's fully loaded cost is amortised across a hundred outcomes for every one that reaches them. At T2, where Escalation Rate runs 1:5 to 1:10, the Steward's cost is amortised across far fewer outcomes, and Cost per Completed Outcome rises accordingly — not because the agent got worse, but because the denominator shrank relative to the fixed human cost sitting inside it. At T3, where human judgement is mandatory, the metric converges toward the cost structure of the business it replaced, which is the correct result: Operational Arbitrage is not capturable where human judgement is required by the nature of the task, and a metric that pretended otherwise would be measuring something other than reality.
This is why Cost per Completed Outcome has to be reported by tier, not as a single blended figure for the business. A blended average lets a business with a small number of expensive T3 outcomes hide behind a large number of cheap T1 ones — technically accurate, structurally misleading, and exactly the kind of average that makes a struggling process look healthy on a dashboard.
Nominal Cost per Completed Outcome
MTTI has a known failure mode: Nominal MTTI, the condition in which a long measured interval between interventions reflects not genuine autonomous operation but a Steward who has quietly stopped reviewing the audit surface. The number looks like architectural certainty. It is actually unmonitored drift, and Memo #15: Auditable Autonomy already made the case that a metric an auditor cannot trace is not evidence of anything.
Cost per Completed Outcome has the same failure mode, and it will be gamed the same way if the definition is left loose. A business can report an excellent figure by narrowing what counts as "completed" until only the cheap, uncontested outcomes qualify — reclassifying a reversed transaction as a new task rather than a failed one, or counting a T2 escalation as resolved the moment it leaves the agent's queue rather than when the customer's problem is actually closed. Call this Nominal Cost per Completed Outcome: a favourable number produced not by genuine cost efficiency but by a definition that has been quietly narrowed until the expensive cases no longer count. The defence is the same one Arco applies to MTTI — the definition of "completed" has to be fixed before the number is reported, tiered rather than blended, and auditable against the same Proof of Action trail that keeps MTTI honest. A cost metric that cannot be traced back to which specific outcomes it counted is not a cost metric. It is a number chosen to look good.
The Operator's Verdict
A business that cannot state its Cost per Completed Outcome, by tier, with a fixed definition of "completed," does not know whether its autonomy is profitable. It knows that it is fast. Those are not the same claim, and the industry's current metrics — tokens, benchmarks, completion counts — are built to let the first one stand in for the second.
The number is uncomfortable because it is honest. Blended, it flatters you. Reported by tier, with "completed" defined tightly enough to survive an audit, it shows exactly where the architecture is still expensive — the only version worth having.
Technology changes what is possible. Cost per Completed Outcome determines what is worth building.
KEY TAKEAWAY
What is Cost per Completed Outcome, and why does it matter for autonomous businesses?
Cost per Completed Outcome is the fully loaded cost of one resolved unit of business output — compute, Steward time, and residual Coordination Tax — divided by outcomes actually completed, not merely executed. It replaces tokens, benchmarks, and completion claims as the metric that determines whether a system is economically superior to the process it replaced, not just faster. It is reported by Task Tier, since a few expensive T3 outcomes can hide inside a large volume of cheap T1 ones. Like MTTI, it has a gaming failure mode — Nominal Cost per Completed Outcome — produced by narrowing 'completed' rather than by genuine cost efficiency. Key metric: At T1, Escalation Rate runs approximately 1:100 and the figure approaches pure compute cost; at T2, Escalation Rate rises to 1:5–1:10 and Steward time dominates.
