The Delegation Ladder

A task-by-task rule for what AI may draft, what it may decide, and what a person must own. The rule is settled by one figure: the cost of the task's worst unsupervised mistake.

Category: Governance. Written by Jaime Garcia, Founder, SnowRock. Published . 16 min read.

In short

Every serious AI conversation with an operator arrives, sooner or later, at the same question, and it is never about the model. It is some version of this: fine, but what is the thing actually allowed to touch? May it answer the customer, or only draft the answer? May it post the invoice, or only suggest the coding? May it move money? The answer to that last one is no. The interesting part is saying precisely why, and what else belongs in the same category as moving money.

Today almost every company answers these questions by feel, and the result is failure in both directions at once. Over-delegate and you get the support agent that quietly refunded its way through a week of tickets, or the automation that emailed the wrong price list to the entire book of business. Under-delegate and you get the drafting tool whose every output is re-read, forever, by the very person it was bought to relieve.

BCG’s 2026 survey of nearly twelve thousand workers measured that second failure. Roughly half now spend more time reviewing, correcting, or directing AI than doing the underlying work, four in ten say the technology has lengthened their decision-making, and about half report that their employer has set no clear rules for how mixed human-and-machine work is governed. That is not safety. That is paying for an employee and then doing the job yourself.

Executives already sense the shape of the answer. When PwC asked which tasks leaders would trust an agent to handle, the responses fell into a clean gradient: most would trust it with data analysis, fewer with routine collaboration, fewer still with autonomous interactions, and fewest with financial transactions. That gradient is an intuitive ranking of tasks by how much a mistake would cost, and the intuition is right. It has simply never been written down as a rule anyone can run.

AI delegation: what the existing ladders leave out

The consulting literature already offers ladders, and each of them grades the machine rather than the decision. Deloitte’s six-level autonomy ladder, the best of the genre, sorts what an agent can do, from manual execution up to a self-governing mesh, and recasts the human at each level as operator, guide, monitor, supervisor, strategist, orchestrator. It is a maturity escalator. What it cannot tell you is whether this particular task, the wire transfer or the quote or the job posting, belongs at level two or level four, because its unit of analysis is the enterprise and nothing in it prices an error.

The others share the gap. Bain grades technical architecture and compresses the oversight question into a single honest line about building Iron Man suits rather than autonomous Iron Man robots. McKinsey's agentic work names governed autonomy, behavior controlled through embedded policies and escalation, as a requirement, then leaves the writing of those policies to the reader. NIST's AI Risk Management Framework is admirably explicit about the omission, scaffolding the process of managing risk while declining, on purpose, to set anyone's risk tolerance. Every one of these tells the reader to set the boundary, and none of them provides one.

Two precedents point at the right axis. The SAE's levels of driving automation contributed the idea of an operational design domain, the principle that autonomy is granted only within stated conditions and never in general. OpenAI's 2023 guidance on agentic systems proposed calibrating human approval to harm severity and reversibility, the correct two ingredients, offered as principle and left as principle. The Delegation Ladder takes that unfinished thought and turns it into something a team can actually run: a score, five rungs, a promotion rule, and a demotion trigger.

Price the worst unsupervised mistake

Ignore, for a moment, everything the system does well. Delegation is not priced on the average case; it is priced on the worst one. For any task, ask a single question: if this ran unsupervised and produced its worst plausible error, what would that cost? Two smaller questions settle the answer.

The first is reversibility. Score it zero if undoing the mistake takes one step, such as editing a draft or reverting a page. Score it one if recovery is possible but takes real effort, a correction and an apology and an hour of cleanup. Score it two if recovery costs money or trust: a clawback, a make-good, a discount nobody planned. Score it three when there is no undo at all, because the money is gone, the data has leaked, the filing has been made, or the relationship has ended.

The second is blast radius. Score it zero if the mistake reaches only the person using the tool, one if it reaches the team, two if it reaches customers or vendors or the public, and three if it reaches the balance sheet, a regulator, or the law. Add the two numbers. That sum, the worst-case score, sets the highest rung the task may ever occupy: a zero clears everything, a one or two caps the task at the exception rung, a three at review, a four at draft, and a five or six keeps it off the ladder entirely.

The ceiling is a maximum, not a placement. Tasks begin below it and earn their way up. Two modifiers can lower a ceiling and never raise it: an error that would not surface on its own drops the ceiling one rung, and so does novel volume, because a thousand unsupervised executions a month will compound a small error into a large one.

Five rungs, and the three actions behind them

The score names a ceiling, and the ceiling is one of five rungs. The five rungs reduce to three actions that a room of non-engineers can argue about.

RungWhat it meansA human still ownsTypical tasks
R0 · Off the ladderAI may inform, flag, and reconcile, but never act. The worst mistake is permanent and reaches the balance sheet or the law. This is a category, not a phase you graduate from.EverythingWire transfers, vendor bank-detail changes
R1 · DraftAI composes; a person commits. Every output crosses a human before it crosses the company’s boundary.Every sendContract redlines, hiring screens, key-account email
R2 · ReviewAI acts in production; humans audit a defined sample, say a tenth of outputs plus every escalation.The auditRoutine support replies, catalog updates
R3 · ExceptionAI decides inside written numeric bounds; anything outside them lands in a human queue with context attached.The envelope and the exceptionsInvoice coding, reorders within set bands
R4 · OwnAI runs the task and no one reviews transactions; a human owns the number the system answers to and the authority to demote it.The metricMeeting notes, data hygiene, tagging
The five rungs. Read the rungs by their verb. AI may draft anything, decide the reversible, and own only the trivially recoverable, while a human owns every send, every audit, every envelope, every metric, and everything at R0. Source: The Delegation Ladder, SnowRock..

The framework reduces to one rule. AI may draft anything, may decide what is reversible, and may own only what it can break without breaking the business. A person owns every send at R1, every audit at R2, every envelope at R3, every metric at R4, and at R0 the action itself.

How a task earns promotion

The score sets the ceiling and deliberately says nothing about the starting rung, because trust in a system, like trust in a new hire, is built from evidence rather than from an org chart. Three rules govern that movement, and without them the ladder is only a diagram.

Start one rung below the ceiling. A task cleared for the exception rung begins at review; a task cleared for review begins at draft. The first weeks in production are where the failure modes nobody imagined introduce themselves, and it is far cheaper to meet them while a person is still on the line.

Promote on evidence, against criteria written before the pilot begins. Our standard is four consecutive weeks or two hundred executions, whichever is longer, at or above the human baseline for quality, with zero incidents of the class the ceiling worries about. Only then does the task move up one rung. Writing the promotion rule in advance is worth the effort. In our own engagement data, a definition of done set before the build tripled the odds of ever reaching production.

Demote without a meeting. Any realized mistake of the ceiling class drops the task one rung immediately, by standing rule, and any mistake that reveals the score itself was wrong, the reversible error that turned out not to be, sends the task back for rescoring while the postmortem runs. Demotion by rule rather than by debate is precisely what makes promotion safe to grant.

Two failure modes that hollow out a rung from inside

The first is the rubber stamp. Reviews decay. The academic work on autonomy levels calls it the slide into meaningless rubber-stamping, and every operator has watched it set in by the second month. So measure the review itself and count the reviewer's intervention rate. If a human approves more than ninety-eight percent of outputs for a full quarter, one of two things is true: the task has earned a promotion, or the review has become theater and needs a smaller sample, a sharper checklist, or a different reviewer. A review that never intervenes has stopped being oversight, and it adds delay and liability at payroll cost.

The second is the missing envelope. Borrow the best idea in the driving-automation levels and scope every rung to written conditions: these order types, these dollar bands, these customer tiers, this system of record. Outside the envelope, a task does not improvise at its rung; it drops to draft automatically and asks. A quoting agent cleared for the exception rung on standard products under twenty-five thousand dollars becomes, by that same document, a drafter the moment a custom order or a new currency appears. The envelope is what lets you say yes with confidence, because you are never delegating “quoting.” You are delegating one defined slice of it.

Ten ordinary tasks, scored

The scores below are typical rather than universal, because reversibility depends on your contracts, your systems, and your regulators. That dependence is the point. The grid is filled in by someone who knows the business, in about an hour, one task at a time, which is the hour most AI strategies skip.

TaskReversibilityBlastScoreCeiling
Meeting notes and CRM hygiene000R4 · Own
Inventory reorders within bands112R3 · Exception
Invoice coding112R3 · Exception
Routine support replies123R2 · Review
Website price corrections123R2 · Review
Quote drafting, standard SKUs224R1, earns R2
Contract redlines224R1 · Draft
Hiring screens235R1, capped
Vendor bank-detail changes336R0
Wire transfers336R0
Typical scores for common mid-market tasks. Notice what the grid does to the argument. The unanswerable “do we trust AI?” becomes “is a mis-coded invoice a one or a two on reversibility?”, a question a controller can settle before lunch. Source: The Delegation Ladder, SnowRock. Illustrative defaults; rescore against your own contracts, systems, and regulators..

The Ladder does not make the machine more trustworthy. It makes the trust you place in it explicit, and explicit trust is the only kind you can safely extend.

Where each task belongs

Every maturity model smuggles in the assumption that up is good. The Delegation Ladder refuses it. A healthy company runs tasks at every rung permanently: wire transfers at R0 in 2026 and in 2030 and for as long as wire fraud exists, contract redlines at R1 for as long as a signature carries liability, note-taking at R4 because nothing about it can hurt you. Capability forecasts describe what the ceiling could someday allow. The score describes what it may allow now. The companies that get this right will not be the ones highest on anyone’s ladder. They will be the ones whose every task sits exactly where its worst mistake says it should.

Two closing cautions. The Ladder governs delegation rather than selection. It decides what the machine may touch, not which machine to use, and that second question has its own rule, the Half-Life of a Problem. And the Ladder is cheap to adopt and expensive to skip. It costs one scoring hour per workflow, a two-line promotion rule, and a demotion trigger written into the runbook. The alternative, which is the current default almost everywhere, is delegation by feel, and delegation by feel holds up until the day it does not.