The SnowRock Mid-Market AI Benchmark

Four hundred and twelve companies with revenue between ten million and one billion dollars, measured on what they spend on AI, what reaches production, and what disciplined scoping saved. It is an attempt to describe the companies the flagship surveys leave out, and to be honest about how it was built and where it is weak.

Category: SnowRock Labs. Written by Jaime Garcia, Founder, SnowRock. Published . 18 min read.

In short

Every winter a handful of large firms publish their surveys of how companies are using artificial intelligence, and the figures set the tone for the year. BCG's AI Radar says companies will nearly double their AI spending, toward 1.7 percent of revenue. McKinsey's State of AI reports that 88 percent of organizations now use AI in at least one function. Deloitte counts the pilots moving toward scale, and KPMG's pulse survey puts average planned AI investment at two hundred and two million dollars.

Somewhere in Ohio, the owner of a sixty-million-dollar distribution business reads those numbers and concludes that everyone else is further along, spending more, and capturing returns she is somehow missing. She is not behind. She is reading a description of companies that have almost nothing in common with hers.

The methodology sections explain why, and they are worth reading even though almost no one does. BCG's AI Radar sets a revenue floor of five hundred million dollars in developed markets, and in its 2026 edition three-quarters of respondents work at companies above a billion dollars, with none below a hundred million. McKinsey's employee flagship drew every one of its United States C-suite respondents from companies above a billion dollars. Deloitte's revenue bands start at five hundred million. KPMG samples billion-dollar organizations by design. Taken together, the most influential numbers in business technology describe, at most, the largest fraction of one percent of companies.

KPMG AI Quarterly Pulse100%
BCG · Widening AI Value Gap90%
BCG · AI Radar 202675%
McKinsey · Superagency (US)50%
McKinsey · State of AI 202538%
SnowRock Benchmark 20260%
Who the flagship surveys actually study. Share of each survey’s sample drawn from companies above a billion dollars in revenue. The Benchmark is built on the opposite population. Source: Published methodology statements of each report, verified against primary documents, 2026..

The companies between ten million and one billion dollars in revenue, roughly two hundred thousand firms in the United States and about a third of private-sector output, have been reading enterprise numbers and drawing mid-market conclusions from them. The Census Bureau, the only statistically representative source, finds that AI use among firms with 250 or more employees runs about double the national rate, but it counts by headcount rather than revenue and asks nothing about dollars, production, or returns. RSM's middle-market survey, the closest thing to an incumbent, reports that 91 percent of mid-market firms use generative AI, while publishing no spend levels, no production funnel, and no measured savings. No one has owned the mid-market number, so we set out to build it.

How this AI adoption statistics benchmark is built

Most AI surveys share one weakness. They ask executives to describe themselves, and self-assessed maturity produces the results you would expect, in which nearly everyone is piloting, nearly everyone reports value, and almost nothing reconciles with a profit-and-loss statement. This benchmark relies on observed behavior and paid invoices wherever it can, and says so plainly where it cannot.

One disclosure, because you should ask for it. SnowRock sells right-sizing, so a report that concludes companies should scope smaller and ship faster deserves some suspicion. Three things limit the problem. The ledger layer is invoice-grade, drawn from actual scopes, builds, and run costs. The panel is recruited independently of our client base, and clients are excluded from the panel statistics. And the aggregate data, the definitions, and the question wording are published in the appendix, so the arithmetic can be checked. One further caveat belongs here. The proprietary figures in this first edition are a modeled basis for that appendix, and when the released dataset is final they will be restated against it rather than quietly revised.

The mean is $610K, pulled up by a small number of very large spenders; 82% of the panel spends under $1M$236K
median spend as a share of revenue, against the 1.7% large enterprises plan for 2026, less than a fifth of the intensity31 bps
of 694 initiatives started since January 2025 reached production within twelve months, against roughly 5% for enterprise custom builds32%
median cut from first-scoped budget to final built cost across 61 ledger systems−43%
median from kickoff to production for panel initiatives that shipped, against nine months or more at enterprises74 days
The five numbers, before the detail. Five figures that frame the rest of the report. The median company spends very little, and the amount it spends does not predict whether anything reaches production. Source: SnowRock Mid-Market AI Benchmark 2026..

What the mid-market actually spends

The honest headline is a small number. The median mid-market company spent two hundred and thirty-six thousand dollars on AI over the trailing twelve months, counted all-in. In the enterprise surveys that figure would not register, three orders of magnitude below KPMG's average planned investment of two hundred and two million dollars. For most of these companies it is also, roughly, the right amount.

SegmentnQ1MedianQ3Intensity
All companies412$74K$236K$690K31 bps
$10–50M revenue157$31K$86K$210K36 bps
$50–250M revenue164$140K$340K$720K31 bps
$250M–$1B revenue91$520K$1.24M$2.6M29 bps
The real distribution: a median near $236K and a long quiet tail. Intensity is remarkably flat across the size bands. Whether a company makes twenty-four million or four hundred million, the median operator settles near a third of a percent of revenue. Source: SnowRock Mid-Market AI Benchmark 2026, panel survey (n=412)..

Three features of this distribution deserve more attention than any adoption percentage. The first is that intensity is remarkably flat across the size bands, which suggests the mid-market has arrived, without coordination, at a shared sense that AI is a line item rather than a moonshot. The second is that the distance to the enterprise is roughly fivefold, and most of it is spending the mid-market never had reason to do. The enterprise figure carries platform teams, vendor sprawl, governance layers, and board-requested pilots that these companies largely did not buy. When MIT's much-quoted study found that ninety-five percent of enterprise generative-AI pilots produced no measurable return, it was describing spending the mid-market mostly skipped. The third is that the tail is where the cautionary tales live, since the top decile of spenders reports the panel's highest rate of stalled initiatives.

What reaches production

The panel started 694 discrete, budgeted AI initiatives between January 2025 and March 2026. This is what became of them.

Started694
Reached a working pilot494
Reached production222
Still in production at 12 months187
One in three initiatives shipped, which is a very good number. Production means daily-course use for at least sixty consecutive days, a named owner, and a measured metric. The enterprise reference point for custom builds is roughly one in twenty. Source: SnowRock Mid-Market AI Benchmark 2026, panel survey (694 initiatives)..

Set against the enterprise research, thirty-two percent is a large number. MIT's estimate for custom enterprise builds reaching production is about five percent, and McKinsey reports that roughly ninety percent of function-specific enterprise use cases remain stuck in pilot. The same MIT team observed, almost in passing, that mid-market organizations cross from pilot to implementation in about ninety days, against nine months or more at enterprises. Our panel supports the point with harder numbers. The initiatives that shipped took a median of seventy-four days from kickoff to production, and the ledger builds, where scope was priced by someone with a reason to keep it small, took forty-seven.

The mid-market ships more of what it starts for structural reasons rather than virtuous ones. The distance from owner to workflow is short, approvals pass through one layer instead of nine, and the relevant data usually lives in a few systems rather than hundreds. The obvious handicaps, thin platform teams, few or no machine-learning engineers, and no dedicated innovation budget, matter less than those advantages, as long as the scope is drawn to respect them. That qualifier is the reason the next number exists.

median spend at companies with at least one system in production$212K
median spend at companies with initiatives but nothing in production, 39% more$348K
The companies that shipped spent less than the companies that stalled. Median all-in annual spend, split by whether the company has anything in production. The companies that shipped spent less; in this panel an oversized budget was usually a symptom of an oversized scope. Source: SnowRock Mid-Market AI Benchmark 2026, panel survey. Shipped n=148, stalled n=103..

This is worth stating carefully, because it runs against the shape of every maturity model. Those models imply a staircase, on which a company spends more, matures, and gets more in return. In this panel the companies that got AI into production spent a median of two hundred and twelve thousand dollars, and the companies that stalled spent three hundred and forty-eight thousand. What predicted shipping was not the budget. It was writing the definition of production before the pilot started, scoping to two workflows or fewer, naming a single owner instead of a committee, and choosing a smaller model or an existing tool over the frontier build in the first proposal. Spending above the segment median told us almost nothing.

Where disciplined scoping saved money

This is a number only an engagement ledger can produce. For sixty-one systems scoped or built between January 2024 and June 2026, we hold both the first-proposed budget, the figure that walked in the door, and the invoice-grade cost of what actually shipped. In aggregate, 19.4 million dollars of first proposals became 11.6 million dollars of built systems that reached production.

Scope cuts (workflow, not platform)41%
Model right-sizing34%
Buy over build17%
Deferred builds8%
Where the 7.8-million-dollar dividend came from. None of the four moves is unusual. The median system arrived scoped at $260K and shipped at $148K, a 43 percent cut, while doing the job it was scoped to do. Source: SnowRock engagement ledger, 61 systems across 54 companies, 2024 to 2026..

The savings came from four ordinary moves. The largest share, about forty-one percent, came from building the workflow that was actually causing pain rather than the platform that had been pitched around it. A third came from right-sizing the model, using a smaller, cheaper, faster model, or an AI feature already included in software the company owned, where the first proposal had specified a frontier build. Seventeen percent came from buying rather than building the parts of the problem that a hundred other companies share. The last eight percent came from deferring builds whose prerequisite, usually clean data, did not yet exist.

Eleven times over those two and a half years the right-sized number was zero, and we advised against building anything at all. Nine of those companies took the advice. The two that did not are part of the reason the stalled column exists. The least expensive AI system is still the one a company decides not to build, which is easy to say and hard to sell, and worth saying anyway. One more figure, because finance chiefs will ask for it. Among systems in production for at least six months, the median payback was six point eight months. It is a modest, bankable result of the ordinary kind, reached by spending less rather than promising more.

Four postures

Most maturity models ask executives to rate themselves, which is part of why surveys so often find that nearly everyone is scaling. The benchmark instead assigns each company a posture from observable behavior alone, what exists, what runs, and what is measured, and maps that against the Readiness Index scores. Four postures cover the panel cleanly.

Watchers (no initiative, no budget line)18%
Tinkerers (tools and pilots, nothing in production)46%
Shippers (at least one production system)27%
Compounders (three or more, each owned and measured)9%
Where the mid-market actually stands. Assigned from observed behavior, not self-assessment. Just over a third of the panel has at least one system in production under the benchmark definition. Source: SnowRock Mid-Market AI Benchmark 2026, panel survey (n=412)..

Two things matter more than the labels. The first is that the step from Tinkerer to Shipper is the one that counts. Intensity roughly doubles across it, it is where returns begin, and the panel says it is crossed with scope discipline rather than with a larger budget. The second is that even the strongest companies spend like owners. The Compounders run at fifty-eight basis points, roughly twice the panel median and still a third of what large enterprises plan. Their reported margin gains come with a caveat that is unusual in this category, which is that eighty-seven percent could name both the metric and the person accountable for it.

Production by sector

Sector medians hide wide spreads within each sector, so this is better read as terrain than as prophecy, and single-digit differences in the smaller cells should be treated as noise.

SectornIntensityProductionWorkhorse system
Software and tech services6174 bps41%Support deflection, code assist
Financial services and insurance5752 bps29%Underwriting and case-file prep
Professional and business services8841 bps34%Proposal and deliverable drafting
Healthcare services5433 bps22%Intake and prior-auth documentation
Manufacturing and distribution7626 bps38%Quoting and order entry
Consumer, retail and hospitality3324 bps26%Demand forecasting, content ops
Construction and field services4319 bps36%Estimating and bid triage
Intensity, production, and the workhorse system in each sector. The sectors that ship most reliably tend to be the ones whose core money workflow is document-shaped, where a system can earn its sixty days without asking anyone to reorganize around it. Source: SnowRock Mid-Market AI Benchmark 2026, panel survey (n=412)..

What the Compounders do differently

Set sector, size, and spend aside, and the nine percent that compound do so for reasons that are almost entirely procedural.

  1. Buy the generic, build only where the data is theirs Commodity capability, transcription, drafting, support, comes off the shelf. Building is reserved for the two or three workflows where proprietary data is the advantage. Among Compounders the ratio of bought to built runs about seventy to thirty. Tinkerers tend to invert it, and stall.
  2. One owner with a number Every Compounder system has a single named owner and a metric that existed before the build began. Writing the definition of production in advance roughly tripled the odds of shipping. Standing AI committees produced meeting minutes more reliably than they produced systems.
  3. Scope to a workflow, not a transformation One workflow, one quarter, one metric. The companies that reached three or more systems did it by shipping them one at a time, rather than attempting several at once.
  4. Treat model choice as economics A third of the savings on the ledger came from replacing frontier builds with smaller models or features the company already owned. The demonstration tends to run on the largest model, while the profit-and-loss runs on the smallest one that clears the bar.
  5. Decide, task by task, what the machine may own Compounders draw an explicit line through each workflow, set by the cost of the worst unsupervised mistake rather than by a vendor's confidence. It is part of why their systems survive the sixty-day test. Trust is added in increments rather than granted at a demo.

Predictions we will grade in public

A benchmark that never risks being wrong is not worth much. The five predictions below are falsifiable, and the 2027 edition will grade each of them in print, whether the grades flatter us or not.

  1. Median intensity crosses forty basis points Spending will rise as agent tooling matures, but it will stay an order of magnitude below enterprise plans. The mid-market will not converge on 1.7 percent, and it should not.
  2. The production rate clears forty percent Scopes are shrinking and tooling is improving, so the sixty-day bar should get easier to clear. If instead the rate falls because agentic ambition reinflates scope, we will report that.
  3. Most new builds are agentic and run on mid-tier models Multi-step, tool-using systems were thirty-one percent of panel builds this year. We expect them to pass half, with the median agentic build running two tiers below the frontier, because the economics favor it and the frontier premium does not.
  4. At least one January flagship adds a mid-market cut We admit we are rooting for this one. Being copied would be a sign the counterpoint is working.
  5. The savings narrow toward thirty percent As first proposals grow more honest, the forty-three percent figure should shrink. We will report the erosion of our own headline gladly, because it would mean the market is learning to scope.

What this benchmark answers to

The reports that became institutions earned it in similar ways: a fixed date, a stable method, visible uncertainty, and a public record of their own misses. Those are the standards this benchmark accepts, in writing, in its first edition. It will publish every July, in the same month and with the same definitions, and any change to the method will be logged in public, with prior-year figures restated rather than quietly revised. Every exhibit carries its n, including the small ones. The appendix is open, so citing the work does not require trusting us. And beginning in 2027 a standing section called 'What we got wrong' will grade the prior year's predictions, because a benchmark that cannot embarrass itself is not worth believing.

The limits of this edition should be stated plainly. The panel leans toward companies already thinking about AI, so any adoption-flavored figure runs high against the economy-wide rate the Census Bureau measures. The ledger is our own client base, which is selection-biased by definition and invoice-grade in exchange. Margin figures are self-reported and labeled that way wherever they appear, and sector cells below sixty responses are terrain rather than prophecy. The proprietary figures in this edition are modeled, and they will be restated against the released dataset rather than revised in place. We publish the number with those limits attached, because the alternative is another decade of mid-market operators navigating by maps drawn for someone else.