8 min read
Anthropic’s October 7 release puts another sharp reduction in AI prices on the table. For businesses, the larger question is how much it costs to produce work someone can actually use.

Key takeaways
- Haiku 5.5 makes routine AI processing cheaper, but a lower model bill does not automatically mean a cheaper business process.
- The useful comparison includes failed attempts, employee review and the number of outcomes that meet the same standard.
- Companies that can measure completed work will be better placed to benefit from the next price cut.
Artificial intelligence became cheaper again this week. Whether that makes a company’s work cheaper is a separate calculation.
On October 7, Anthropic introduced Haiku 5.5, a model aimed at high-volume work. For prompts of up to 100,000 tokens, its announced prices are $0.10 per million input tokens and $0.50 per million output tokens. Above that threshold, the rates are $0.50 and $2.50. Tokens are the units into which a model divides the text it reads and writes. Anthropic’s release and pricing set out the limits.
The company says Haiku 5.5 costs about 75 percent less to run on average than Haiku 4.5, allowing for different request lengths and changes in token use. That is a vendor estimate, not a promise about an individual customer’s bill. Anthropic also halved Sonnet 5.5’s cache-read price to $0.10 per million tokens. A cache lets a system reuse material it has already processed.
The reductions matter. A service handling thousands of documents or customer requests can reconsider work that was previously too expensive to automate. But a model’s processing charge is only one part of the cost. Somebody still has to define the job, provide the right information, handle exceptions and decide whether the result is good enough to leave the building.
This is where the public conversation about AI prices and the private economics of a business begin to diverge. The model company sells computing. The customer needs a resolved problem.
The Unit That Matters
A price per million tokens resembles the price of a raw material. It tells a buyer something important about one input. It says much less about the finished product. A low price for lumber does not tell a furniture maker what it will cost to deliver a table. Waste, labor, design, shipping and defects still matter. The same distinction applies when software assembles an answer instead of a physical object.
For a business, the relevant unit might be an invoice correctly reconciled, an insurance file ready for a reviewer, or a customer issue resolved without another call. Each has a different definition of completion. A draft summary can be excellent while the underlying account remains wrong. Ten thousand generated replies can coexist with a growing queue of customers waiting for help.
The FinOps Foundation, which develops practices for managing technology spending, makes a related distinction between resource measures and business measures. Its unit-economics guidance encourages organizations to connect AI consumption to outcomes such as assists or completed cases. The implication is practical: count the work the system helps finish, alongside what it consumes.
A company can begin with a straightforward measure: the total cost of a workflow over a defined period, divided by the number of results that meet an agreed standard. The numerator includes model charges, other software, employee review and correction. Failed attempts belong in the cost even when they produce nothing useful. Otherwise, the system improves on paper whenever its hardest cases disappear from the report.
How a Cheaper Model Can Cost More
Consider a hypothetical team processing 1,000 routine documents. Two approaches perform the same job, using the same acceptance standard. The first spends ten cents per document on AI and requires an average minute of employee review. The second spends forty cents on AI but requires eighteen seconds of review. Assume employee time costs $60 an hour, including the employer’s associated costs.
In the first approach, model processing costs $100 and review costs $1,000. If 900 documents meet the required standard, the direct cost is approximately $1.22 for each accepted document. The second approach spends $400 on the model and $300 on review. If 950 documents are accepted, the comparable cost is about 74 cents. The more expensive model produces the less expensive accepted work in this illustration.
A Hypothetical Cost Comparison
| Measure | Approach A | Approach B |
|---|---|---|
| Model processing | $100 | $400 |
| Employee review | $1,000 | $300 |
| Direct cost | $1,100 | $700 |
| Accepted documents | 900 | 950 |
| Direct cost per accepted document | $1.22 | $0.74 |
Both approaches process 1,000 documents. Review time is valued at $60 per hour. Integration, other software and later remediation would need to be added in a real deployment.
SnowRock illustration. These are assumptions, not measured results for any AI model.
This example does not establish that larger models are always better value. A simpler model might achieve the same quality with the same review time at a fraction of the price. It shows why a purchase decision cannot stop at the rate card. Small differences in human effort can outweigh large percentage changes in a computing charge.
The calculation also leaves something important unresolved: what happens to the rejected documents? If employees must finish them manually, that work needs to appear in the complete process budget. If the organization can defer them, there may still be a cost in delayed payment or service. A useful financial model follows the job until the business has actually dealt with it.
A Model for Each Part of the Job
The latest releases also make a single-model purchasing strategy harder to justify. In its September 28 Sonnet 5.5 announcement, Anthropic positioned Sonnet for defined everyday tasks and Opus for more complex work requiring sustained judgment. It reported lower task costs for Sonnet despite unchanged input and output rates at launch. The amount of processing needed matters alongside the price of each unit.
A reasonable business response is to separate a workflow into decisions that require different levels of judgment. Sorting an incoming file by document type is different from deciding whether its financial explanation is credible. Extracting a customer number is different from interpreting an unusual contract clause. Those activities should not automatically receive the same model, the same amount of processing or the same authority.
Imagine an order-management team. A smaller model could identify missing fields and prepare a short account summary. A more capable model could investigate conflicting delivery instructions. A person could approve an unusual concession to a major customer. The point is to spend carefully on the difficult part while keeping predictable work inexpensive. This is a proposed design, not a claim about a particular deployed system.
Routing has costs of its own. The software that sends a job to the right model can misjudge its difficulty, and a mistaken first attempt can make a second attempt more expensive. The sensible comparison is between complete workflows. A sophisticated combination of models should earn its complexity through better results, rather than receive credit simply for containing more moving parts.
The Queue Behind the Automation
The most revealing failure may appear after an AI system succeeds at producing more work. If employees can generate proposals ten times faster but the same two people must approve every proposal, the organization may move its bottleneck rather than remove it. More drafts arrive. Review becomes hurried. Important work competes with material that nobody needed in the first place.
This changes the meaning of productivity. A faster individual can still be part of a slower process. One team may report hours saved while another absorbs the checking, reconciliation and customer follow-up. Managers who look only at the department purchasing the AI service will miss that transfer. The operating result should include the people downstream who inherit its output.
There is reason to be careful about headline productivity figures as well. In February, the research organization METR said it was redesigning its developer-productivity experiment because participation and task selection had changed. Some developers did not want to work without AI, and time measurement became harder when they ran agents alongside other work. METR said its data was weak evidence for the size of the current benefit, rather than a reliable universal estimate.
That finding does not settle the value of AI in an accounting office or service company. It demonstrates how easily the measurement can change while the technology changes. A company should distinguish between reduced effort, faster delivery, higher quality and additional output. All can be valuable. They are not interchangeable, and none should be assumed merely because the software was used.
What to Measure Before Switching
A useful comparison starts with a sample of actual work, including ordinary cases and the exceptions employees find difficult. The organization should decide what an acceptable result looks like before seeing which system produced it. Otherwise, a persuasive answer can quietly change the standard by which it is judged.
Anthropic’s engineering guidance on agent evaluations distinguishes the record of an agent’s actions from the resulting state of the world. An agent can report success without having completed the underlying task. The guidance also describes combining automated checks with expert judgment and testing across repeated attempts. A benchmark should measure the promised job, including its failure cases.
For a business trial, the working record should connect each request to its computing cost, time waiting, review effort and final disposition. Later corrections need to attach to the original job, not vanish into a separate support budget. If two approaches produce different quality, the report should show that difference rather than compress it into a single savings claim.
There is no need to estimate every future expense with false precision. Implementation can be reported separately from ongoing operation. Uncertain costs can be shown as a range. A small trial may establish that a proposal deserves a larger trial without proving that it should run across the company. Honest boundaries make a result more useful to the person approving the next investment.
What Lower Prices Actually Make Possible
Lower processing costs can change more than an existing budget. They may make it worthwhile to search a neglected archive, classify older customer questions or review a larger set of product feedback. These jobs often compete poorly for employee time because no single item justifies the effort. Cheap processing can make the collection worth examining, provided somebody can turn the result into action.
That creates an important choice. A company can use the savings to do the same work for less, or it can spend more in total to do substantially more useful work. A growing AI bill is not automatically a failure, just as a shrinking bill is not automatically a success. The question is what the spending buys and whether the added result is worth its cost.
The October price cuts create a good moment to repeat that calculation. Some jobs will move to smaller models. Some will benefit more from clearer instructions or better source material. Others will remain expensive because the difficult part belongs to a human reviewer, a customer or a process the software cannot change.
The companies best positioned to benefit will know the difference. They will be able to say how many jobs were completed, what standard those jobs met, how much employee effort remained and what the full process cost. The next model release will then be a purchasing opportunity they can evaluate, rather than another impressive number they cannot put to work.