The Delegation Ladder
A task-by-task rule for what AI may draft, what it may decide, and what a person must own. The rule is settled by one figure: the cost of the task's worst unsupervised mistake.
Category: Governance. Written by Jaime Garcia, Founder, SnowRock. Published . 16 min read.
In short
- Delegation should be priced on the worst case rather than the average one. Score each task by how reversible its worst unsupervised mistake is and how far that mistake reaches, and the sum sets a ceiling on how much independence the task may ever hold.
- Five rungs reduce to three actions. A task may be allowed to draft, to decide something reversible, or to own an action outright, and only work that can fail without serious harm should ever reach the last one.
- Independence is earned on evidence and lost by standing rule. A task starts one rung below its ceiling, is promoted on a schedule written in advance, and is demoted the moment it makes the mistake the ceiling exists to prevent.
Every serious AI conversation with an operator arrives, sooner or later, at the same question, and it is never about the model. It is some version of this: fine, but what is the thing actually allowed to touch? May it answer the customer, or only draft the answer? May it post the invoice, or only suggest the coding? May it move money? The answer to that last one is no. The interesting part is saying precisely why, and what else belongs in the same category as moving money.
Today almost every company answers these questions by feel, and the result is failure in both directions at once. Over-delegate and you get the support agent that quietly refunded its way through a week of tickets, or the automation that emailed the wrong price list to the entire book of business. Under-delegate and you get the drafting tool whose every output is re-read, forever, by the very person it was bought to relieve.
BCG’s 2026 survey of nearly twelve thousand workers measured that second failure. Roughly half now spend more time reviewing, correcting, or directing AI than doing the underlying work, four in ten say the technology has lengthened their decision-making, and about half report that their employer has set no clear rules for how mixed human-and-machine work is governed. That is not safety. That is paying for an employee and then doing the job yourself.
Executives already sense the shape of the answer. When PwC asked which tasks leaders would trust an agent to handle, the responses fell into a clean gradient: most would trust it with data analysis, fewer with routine collaboration, fewer still with autonomous interactions, and fewest with financial transactions. That gradient is an intuitive ranking of tasks by how much a mistake would cost, and the intuition is right. It has simply never been written down as a rule anyone can run.
AI delegation: what the existing ladders leave out
The consulting literature already offers ladders, and each of them grades the machine rather than the decision. Deloitte’s six-level autonomy ladder, the best of the genre, sorts what an agent can do, from manual execution up to a self-governing mesh, and recasts the human at each level as operator, guide, monitor, supervisor, strategist, orchestrator. It is a maturity escalator. What it cannot tell you is whether this particular task, the wire transfer or the quote or the job posting, belongs at level two or level four, because its unit of analysis is the enterprise and nothing in it prices an error.
The others share the gap. Bain grades technical architecture and compresses the oversight question into a single honest line about building Iron Man suits rather than autonomous Iron Man robots. McKinsey's agentic work names governed autonomy, behavior controlled through embedded policies and escalation, as a requirement, then leaves the writing of those policies to the reader. NIST's AI Risk Management Framework is admirably explicit about the omission, scaffolding the process of managing risk while declining, on purpose, to set anyone's risk tolerance. Every one of these tells the reader to set the boundary, and none of them provides one.
Two precedents point at the right axis. The SAE's levels of driving automation contributed the idea of an operational design domain, the principle that autonomy is granted only within stated conditions and never in general. OpenAI's 2023 guidance on agentic systems proposed calibrating human approval to harm severity and reversibility, the correct two ingredients, offered as principle and left as principle. The Delegation Ladder takes that unfinished thought and turns it into something a team can actually run: a score, five rungs, a promotion rule, and a demotion trigger.
Price the worst unsupervised mistake
Ignore, for a moment, everything the system does well. Delegation is not priced on the average case; it is priced on the worst one. For any task, ask a single question: if this ran unsupervised and produced its worst plausible error, what would that cost? Two smaller questions settle the answer.
The first is reversibility. Score it zero if undoing the mistake takes one step, such as editing a draft or reverting a page. Score it one if recovery is possible but takes real effort, a correction and an apology and an hour of cleanup. Score it two if recovery costs money or trust: a clawback, a make-good, a discount nobody planned. Score it three when there is no undo at all, because the money is gone, the data has leaked, the filing has been made, or the relationship has ended.
The second is blast radius. Score it zero if the mistake reaches only the person using the tool, one if it reaches the team, two if it reaches customers or vendors or the public, and three if it reaches the balance sheet, a regulator, or the law. Add the two numbers. That sum, the worst-case score, sets the highest rung the task may ever occupy: a zero clears everything, a one or two caps the task at the exception rung, a three at review, a four at draft, and a five or six keeps it off the ladder entirely.
The ceiling is a maximum, not a placement. Tasks begin below it and earn their way up. Two modifiers can lower a ceiling and never raise it: an error that would not surface on its own drops the ceiling one rung, and so does novel volume, because a thousand unsupervised executions a month will compound a small error into a large one.
Five rungs, and the three actions behind them
The score names a ceiling, and the ceiling is one of five rungs. The five rungs reduce to three actions that a room of non-engineers can argue about.
| Rung | What it means | A human still owns | Typical tasks |
|---|---|---|---|
| R0 · Off the ladder | AI may inform, flag, and reconcile, but never act. The worst mistake is permanent and reaches the balance sheet or the law. This is a category, not a phase you graduate from. | Everything | Wire transfers, vendor bank-detail changes |
| R1 · Draft | AI composes; a person commits. Every output crosses a human before it crosses the company’s boundary. | Every send | Contract redlines, hiring screens, key-account email |
| R2 · Review | AI acts in production; humans audit a defined sample, say a tenth of outputs plus every escalation. | The audit | Routine support replies, catalog updates |
| R3 · Exception | AI decides inside written numeric bounds; anything outside them lands in a human queue with context attached. | The envelope and the exceptions | Invoice coding, reorders within set bands |
| R4 · Own | AI runs the task and no one reviews transactions; a human owns the number the system answers to and the authority to demote it. | The metric | Meeting notes, data hygiene, tagging |
The framework reduces to one rule. AI may draft anything, may decide what is reversible, and may own only what it can break without breaking the business. A person owns every send at R1, every audit at R2, every envelope at R3, every metric at R4, and at R0 the action itself.
How a task earns promotion
The score sets the ceiling and deliberately says nothing about the starting rung, because trust in a system, like trust in a new hire, is built from evidence rather than from an org chart. Three rules govern that movement, and without them the ladder is only a diagram.
Start one rung below the ceiling. A task cleared for the exception rung begins at review; a task cleared for review begins at draft. The first weeks in production are where the failure modes nobody imagined introduce themselves, and it is far cheaper to meet them while a person is still on the line.
Promote on evidence, against criteria written before the pilot begins. Our standard is four consecutive weeks or two hundred executions, whichever is longer, at or above the human baseline for quality, with zero incidents of the class the ceiling worries about. Only then does the task move up one rung. Writing the promotion rule in advance is worth the effort. In our own engagement data, a definition of done set before the build tripled the odds of ever reaching production.
Demote without a meeting. Any realized mistake of the ceiling class drops the task one rung immediately, by standing rule, and any mistake that reveals the score itself was wrong, the reversible error that turned out not to be, sends the task back for rescoring while the postmortem runs. Demotion by rule rather than by debate is precisely what makes promotion safe to grant.
Two failure modes that hollow out a rung from inside
The first is the rubber stamp. Reviews decay. The academic work on autonomy levels calls it the slide into meaningless rubber-stamping, and every operator has watched it set in by the second month. So measure the review itself and count the reviewer's intervention rate. If a human approves more than ninety-eight percent of outputs for a full quarter, one of two things is true: the task has earned a promotion, or the review has become theater and needs a smaller sample, a sharper checklist, or a different reviewer. A review that never intervenes has stopped being oversight, and it adds delay and liability at payroll cost.
The second is the missing envelope. Borrow the best idea in the driving-automation levels and scope every rung to written conditions: these order types, these dollar bands, these customer tiers, this system of record. Outside the envelope, a task does not improvise at its rung; it drops to draft automatically and asks. A quoting agent cleared for the exception rung on standard products under twenty-five thousand dollars becomes, by that same document, a drafter the moment a custom order or a new currency appears. The envelope is what lets you say yes with confidence, because you are never delegating “quoting.” You are delegating one defined slice of it.
Ten ordinary tasks, scored
The scores below are typical rather than universal, because reversibility depends on your contracts, your systems, and your regulators. That dependence is the point. The grid is filled in by someone who knows the business, in about an hour, one task at a time, which is the hour most AI strategies skip.
| Task | Reversibility | Blast | Score | Ceiling |
|---|---|---|---|---|
| Meeting notes and CRM hygiene | 0 | 0 | 0 | R4 · Own |
| Inventory reorders within bands | 1 | 1 | 2 | R3 · Exception |
| Invoice coding | 1 | 1 | 2 | R3 · Exception |
| Routine support replies | 1 | 2 | 3 | R2 · Review |
| Website price corrections | 1 | 2 | 3 | R2 · Review |
| Quote drafting, standard SKUs | 2 | 2 | 4 | R1, earns R2 |
| Contract redlines | 2 | 2 | 4 | R1 · Draft |
| Hiring screens | 2 | 3 | 5 | R1, capped |
| Vendor bank-detail changes | 3 | 3 | 6 | R0 |
| Wire transfers | 3 | 3 | 6 | R0 |
The Ladder does not make the machine more trustworthy. It makes the trust you place in it explicit, and explicit trust is the only kind you can safely extend.
Where each task belongs
Every maturity model smuggles in the assumption that up is good. The Delegation Ladder refuses it. A healthy company runs tasks at every rung permanently: wire transfers at R0 in 2026 and in 2030 and for as long as wire fraud exists, contract redlines at R1 for as long as a signature carries liability, note-taking at R4 because nothing about it can hurt you. Capability forecasts describe what the ceiling could someday allow. The score describes what it may allow now. The companies that get this right will not be the ones highest on anyone’s ladder. They will be the ones whose every task sits exactly where its worst mistake says it should.
Two closing cautions. The Ladder governs delegation rather than selection. It decides what the machine may touch, not which machine to use, and that second question has its own rule, the Half-Life of a Problem. And the Ladder is cheap to adopt and expensive to skip. It costs one scoring hour per workflow, a two-line promotion rule, and a demotion trigger written into the runbook. The alternative, which is the current default almost everywhere, is delegation by feel, and delegation by feel holds up until the day it does not.