The Two-Week Diagnostic, Published in Full

The workflow census, the data-readiness checks, the scoring rubric with its real weights, and the build-buy-wait decision tree we run inside every SnowRock engagement, published in full and free. We give the method away because a firm that sells right-sizing should be willing to show the instrument it uses to do it.

Category: Strategy. Written by Jaime Garcia, Founder, SnowRock. Published . 17 min read.

In short

Here is how AI assessment usually works at the top of the market. BCG's flagship research scores a thousand companies against thirty enterprise capabilities, and the list of thirty is not published; step five of its public playbook is, in its own words, an AI maturity assessment to baseline your capability gaps, which means the report's call to action is the engagement itself. McKinsey's Digital Quotient has graded companies on eighteen practices since 2015, with the practices named and the instrument kept private. PwC's AI Readiness Assessment names twelve domains, lists five, shows none of the questions, and describes itself as a full organizational diagnostic operated by PwC.

None of this is dishonest. It is a business model, and a sound one. A firm publishes survey research that shows a worrying gap, keeps private the instrument that locates each client's version of that gap, and prices that instrument at what the worry will bear, which in the current market runs from twenty-five thousand to a hundred and fifty thousand dollars over two to six weeks. The research brings the client in the door, the rubric is the thing being sold, and the benchmark database behind it is what makes the rubric hard for anyone else to copy.

We are doing the opposite. What follows is our entire diagnostic: the interview script, the data checks with their pass conditions, the scoring weights, the decision tree, and the one-page verdict. We publish it in full, with none of the useful parts held back, because the method is not the thing that makes the work hard to copy, and behaving as though it were is a large part of why buyers have learned to distrust this category.

How published methods become standards

This is not idealism. It is one of the best-documented growth strategies in professional services, and its clearest example belongs to one of the firms above. In 2003 Bain published the Net Promoter Score in the Harvard Business Review, the whole formula, a single question, subtract the detractors. Anyone could run it, and everyone did. Two decades later the score is a global standard, every use of it recalls Bain, and Bain still sells the implementation. Publishing the method did not commoditize the firm. It made Bain the name the measure belongs to.

The pattern repeats wherever a firm is willing to run it. DORA publishes its four delivery metrics and a free five-question check that stores none of your answers, and it became the shared vocabulary of an industry while competitors built dashboards that market DORA's framework for it. HubSpot's Website Grader, built in its founders' words to generate buzz and leads, graded four million sites and established the company's category. NIST released its AI Risk Management Framework free and in full, and the same large firms now sell implementation against it. The published method tends to become the standard, and the firm that authored it tends to become the default choice to implement it. What follows is ours.

An AI readiness assessment in ten working days, five stages

The diagnostic answers the four questions that matter before anyone builds: which workflows deserve AI at all, whether to build, buy, wait, or never for each one, what it might cost as a ceiling set by arithmetic rather than appetite, and how it should run, meaning the starting rung on the Delegation Ladder and the model class from the Half-Life of a Problem. The three frameworks fit together. The diagnostic decides which workflow, the Half-Life decides which kind of model, and the Ladder decides how much independence it begins with.

  1. Days 1 to 3, the census Interviews with the owner, the finance lead, and the people closest to the operational pain, built around six questions. Output: at most six candidate workflows, each named in one sentence with a verb, an input, and an output.
  2. Days 3 to 5, the data checks Every candidate takes all six readiness tests. Five passes to proceed; specific failures convert a candidate to Wait, with the cheapest repair named.
  3. Days 6 to 8, the economics Compute what each surviving workflow costs today, then set the ceiling. Output: a ceiling budget and a payback figure per workflow.
  4. Days 8 to 9, the rubric Score the candidates that cleared the gates on five weighted criteria. Output: a weighted score that sorts the field.
  5. Day 10, the verdict Each workflow exits through one of four doors, build, buy, wait, or never, on a single page.

Two weeks is not a compressed four-week engagement. It is the honest length of the work. Anything longer is usually the diagnostic expanding to justify its price, which is the assessment industry's own version of the oversized scope this exercise exists to catch.

The census: six questions, asked of the people who do the work

This is a set of interviews rather than a strategy session. It is built around six questions whose exact wording matters, so use the wording.

  1. Where do the hours concentrate? Name the five activities that consume the most paid hours in a normal week, measured by hours worked rather than by how much they annoy people. Payroll is the denominator of everything that follows.
  2. Where do errors cost real money? Describe the last three mistakes that cost more than a day of cleanup. Error cost is the half of a workflow’s value that hour-counting misses.
  3. What is document-shaped? Where do people read one thing to type another: quotes, intake, invoices, orders, reports? Document-shaped work is where mid-market AI ships most reliably.
  4. What has a queue? A backlog is demand the current process cannot clear, the cleanest possible signal that capacity, not quality, is the constraint.
  5. What would you never let a machine touch? The fear inventory. Half the answers are correct and will score off the ladder; half are review-level tasks wearing off-limits costumes. Both halves are useful.
  6. If one thing ran twice as fast, what would customers notice? The outside-in check. It catches the workflow whose slowness is quietly shaping the customer experience, and therefore the revenue.

The output is at most six candidates, and they must be workflows rather than areas. "Sales" is not a candidate; "quote assembly for standard SKUs" is. If a candidate cannot be named in one sentence containing a verb, an input, and an output, it is not scoped yet, so split it until it can.

The data-readiness check: six tests with pass conditions you can run today

Every candidate takes all six checks, and five passes are needed to proceed. A specific failure does not kill a candidate. It converts the candidate to Wait, with the cheapest repair named. That conversion is where much of the diagnostic's value hides, since eight in ten enterprises told McKinsey that data limitations are what stall their agents, and the mid-market's version of that problem is usually fixable in a quarter once someone writes down which check failed.

  1. Exists Are records of this work actually kept, the inputs, the outputs, the corrections? If the answer lives in someone’s head it is not data yet; it is a hiring risk. Pass when you can point to where last Tuesday’s instances are stored.
  2. Accessible Can a person with ordinary permissions export it without a favor from IT or a vendor ticket? Vendor-jailed data fails this more than anything else in the mid-market. Pass when a usable export lands on a desktop within an hour of asking.
  3. Consistent Do the same fields mean the same things across time and across the people who fill them? Pull twenty records from three months and read them side by side. Pass when the spot check produces no arguments about what a field means.
  4. Sufficient Enough history for the architecture the task wants, roughly a thousand examples for pattern-learning work. Honesty cuts both ways: retrieval needs far less, so do not let “we lack data” block a workflow that only needs this morning’s price file. Pass when volume meets the bar for the placement the Half-Life grid assigns.
  5. Permitted Do your customer contracts, employee policies, and regulators allow this data to be used this way, and processed by these vendors? This is a written question to counsel, not a feeling in a meeting. Pass on a yes, in writing, from whoever owns the risk.
  6. Fresh Is the data updated at the cadence the task requires? A pricing assistant fed last quarter’s prices does not fail loudly; it answers fluently and wrongly. Freshness is set by the problem’s half-life. Pass when the update cadence is at least as fast as that half-life, with a named process that keeps it so.

The economics: a ceiling set by arithmetic

For each surviving candidate, compute what the workflow costs today. The formula is deliberately simple enough for a finance lead to check on one sheet: annual workflow cost equals hours per week times loaded rate times fifty-two, plus error rate times cost per error times annual volume, plus whatever revenue the backlog delays or loses where that is measurable.

Notice what the ceiling quietly forbids: the four-hundred-thousand-dollar platform pitched at a hundred-and-fifty-thousand-dollar problem. That is the oversized scope that fills the stalled column of our Benchmark, where the companies whose initiatives died had spent more than the companies whose initiatives shipped. The ceiling reflects the price at which these systems have actually survived, rather than a preference for spending less.

The rubric, weights included

Candidates that clear the gates get scored. The gates are pass-or-fail. The weights are judgments, tuned across sixty-one built systems and printed here so you can argue with them, which is more than most rubrics allow.

GateWhat it means
A named owner existsA specific person would run this in production, not a committee and not “ops.”
The metric is measurable todayThe number the system will be judged on can be computed from current records, before any build.
The task is not off the ladderIts worst unsupervised mistake, scored on the Delegation Ladder, leaves at least draft-level delegation available.
Data passed five of six checksOr the verdict is already Wait, with the repair named.
Four gates: any failure stops the scoring. The gates are pass or fail. A candidate that trips any one of them does not get a score; it gets a repair ticket or a Wait. Source: The Two-Week Diagnostic, SnowRock..
Weighted criterionWeightWhat a 5 looks like
Value density30%Annual workflow cost is large against build cost; the ceiling leaves room to breathe.
Data readiness25%Six of six, comfortably; the export was on the desk in ten minutes.
Workflow stability15%The process will not be redesigned mid-build, even if its data moves fast.
Delegation headroom15%The Ladder ceiling is the exception rung or higher; the system can eventually run, not merely draft.
Owner strength15%The named owner wants it, can change the process, and got through the interviews without saying “transformation.”
Five weighted criteria, scored one to five. A weighted score at or above 3.5 enters the build zone; 2.5 to 3.5 defaults to buy or wait; below 2.5 is a “never” wearing optimism. The weights are our opinion, revised only in the open. Source: The Two-Week Diagnostic, SnowRock. Weights tuned across 61 built systems..

The verdict: build, buy, wait, or never

Every candidate leaves through one of four doors, and the tree is short. If thousands of companies share the exact problem, buy the feature, as roughly three-quarters of the market now does. If the problem is yours but the data failed more than one check, the answer is Wait when the value is still large and the repair is known, and Never when it is not. If the data cleared, the rubric scored at or above the threshold, and the cost sits within the ceiling, build. Otherwise, ask whether the blocker is a fixable gap in data, ownership, or process churn; if it is, Wait, and if it is not, say Never in writing.

  1. Build Small: one workflow, one quarter, one metric, one named owner. Never a platform first.
  2. Buy The problem is shared by thousands of companies, so a vendor is amortizing it across all of them. Take the feature.
  3. Wait A verdict that comes with a repair ticket and a revisit date. Fix the failed check, then reassess.
  4. Never A verdict that comes with a reason. The least expensive AI system is the one you decide not to build.

Across our engagement history the verdicts split roughly build forty-five percent, buy twenty, wait twenty-five, never ten, and that last column includes eleven builds we declined to sell. Each workflow leaves day ten as a one-page verdict: the decision and its reason, the ceiling budget and the payback behind it, the starting rung and promotion criteria from the Ladder, the model class and freshness requirement from the Half-Life grid, the owner and the metric, and a written definition of production, which for us is sixty consecutive days of daily-course use, a named owner, and a measured metric. It fits on one page. A verdict that needs ten pages is usually a sales document.

Why give this away?

Because the method was never the thing that made the work hard to copy. What you cannot download is the judgment that comes from running this against sixty-one built systems and eleven refused ones: a sense of which "consistent" data will betray you in month two, which owners will hold the line at the review rung, and which ceiling breaches are worth an exception. Publishing the Net Promoter Score did not produce a thousand Bains, and publishing the four DORA metrics did not make every engineering team elite. The rubric can be copied. The experience of having used it many times cannot.

And because a firm that sells right-sizing should be willing to show the instrument it uses to do it. Our diagnostic's most common outputs are buy, wait, and smaller than you planned, all of which reduce our own invoice. Gating a method like that would undercut the one thing it is meant to demonstrate, which is that we will tell a client the truth about scope. So run it yourself. Most companies' first pass takes a fortnight of part-time attention and produces the four answers above. The ones who call us tend to call after the census, when the arguing starts, and they arrive as the best-prepared clients we have. That is the trade, and it has consistently favored the firm that publishes.