OpenAI Releases Its Most Capable Model Yet
GPT-5.6 Sol arrives alongside a workplace agent built to operate software on a user’s behalf, intensifying the contest with Anthropic and reopening the question of whether frontier models should be restricted or widely deployed.
Category: Strategy. Written by Jaime Garcia, Founder, SnowRock. Published . 14 min read.
In short
- For an operator, the agent matters more than the model. GPT-5.6 Sol is an incremental gain in intelligence, and ChatGPT Work is what turns that intelligence into action, moving AI from a system that advises to one that operates.
- The frontier contest is decided less by benchmarks than by return on inference: how much commercially valuable work a system completes per dollar, under real rules, at acceptable risk.
- For agents, the design problem is permission more than intelligence. The strongest enterprise products pair capability with granular permissions, audit trails, and reversibility, because an autonomous agent is only as safe as the access it is granted.
OpenAI has released GPT-5.6 Sol, its most advanced model to date, together with a system built to perform everyday professional work across applications, websites, and business software. The pairing puts OpenAI back at the leading edge of the frontier-model market after a stretch in which regulatory uncertainty delayed the technology’s public deployment.
Sol is aimed at complex reasoning, software development, financial analysis, legal work, and cybersecurity. Alongside it, OpenAI introduced ChatGPT Work, an agentic product powered by the model that can act on spreadsheets, calendars, email systems, and other digital tools on a user’s behalf. That combination is more consequential than a conventional model upgrade, because it changes what the technology is for.
Earlier generations of AI were used mainly to generate things: answers, documents, images, code. A workplace agent is built to execute sequences of actions. A user can ask it to analyze a spreadsheet, flag the anomalies, draft a summary, book a meeting, and send the findings to colleagues, and rather than explaining how each step should be done, the system navigates the applications and does the work. That is the shift beneath the announcement. The technology has been an advisory layer that tells you things, and it is becoming an operational one that carries them out.
It also sharpens the rivalry with Anthropic, whose Fable 5 sits in a similar position at the commercial frontier. Independent evaluations put the two systems close across many standardized tests, with Sol looking particularly strong in applied legal and financial work. Increasingly the difference between them turns less on raw intelligence than on price, safety restrictions, software integrations, and how much autonomy each company will grant. The real question has moved past which model is most intelligent to which model becomes the operating layer for professional work.
What an agent actually does
A chatbot waits for a question and returns a response. An agent receives an objective, works out which steps are required, selects the tools it needs, and attempts to finish the assignment. That demands more than language generation: the system has to read software interfaces, hold context across several actions, interpret information as it changes, and recognize when a step needs human approval.
The appeal is plain, because office work is scattered across applications. People spend a great deal of paid time moving information between email, spreadsheets, calendars, databases, and chat, and much of it follows predictable patterns while still demanding continuous attention. An effective agent absorbs that coordination burden: assembling a weekly operating report from several systems, triaging incoming messages, catching a scheduling conflict, updating a customer record after a meeting. The user supplies the goal and the system manages the workflow. It is the same market Anthropic is chasing through Claude Cowork, and that Microsoft, Google, and a widening field of enterprise-software companies are building toward, and the commercial prize is large precisely because an agent’s value is tied to labor. A model that writes a paragraph saves minutes. A model that completes a recurring process can replace hours, or a whole administrative function.
Why Sol matters, and why benchmarks no longer settle it
Frontier models are getting harder to tell apart in ordinary use. Most of the leaders summarize documents, produce competent prose, answer questions, and write software; the daylight between them shows up in work that requires sustained reasoning, technical depth, or interaction with external tools, which is exactly where Sol is meant to improve, across financial modeling, legal research, advanced software development, scientific reasoning, and cybersecurity investigation.
The significance is economic rather than technical. Companies will not reorganize because a chatbot got moderately better at conversation; they may reorganize when a system can complete work that used to require expensive professionals. So the contest is moving away from novelty and toward reliability under commercial conditions. It is not enough to produce a good answer once. A system has to perform consistently, use the correct data, follow the organization's rules, and avoid creating new legal, financial, or security exposure. That standard, rather than a leaderboard, is where Sol and Fable will be judged.
The cost of frontier intelligence
The most capable models are expensive to run. Sol and Fable consume far more computation than smaller systems, and their prices reflect advanced processors, data-center infrastructure, and the extra work of complex reasoning. For businesses that produces a tiered market: route routine work to inexpensive models and reserve the frontier for the hard or high-value assignment. A company might classify messages or draft simple summaries with a small system while keeping Sol for legal analysis, financial decisions, or a difficult software problem, which is roughly how human organizations already allocate their most experienced people.
OpenAI and Anthropic therefore have to show that the extra performance of a flagship earns its cost, and benchmarks will not close that argument. A model that scores slightly higher while consuming twice the resources can be the worse buy at scale, while a pricier system wins when a mistake is expensive or when superior work eliminates significant labor. The metric that matters is becoming return on inference: how much commercially valuable output each dollar of operation produces.
A delayed release, and a new precedent
Sol arrived after an unusual bout of government involvement. Washington had raised concerns that the newest models from OpenAI and Anthropic could materially increase cyber risk, sought voluntary access to advanced systems before release, and at first encouraged the companies to limit distribution. OpenAI prepared to make Sol available only to a restricted set of approved organizations; after further discussion, the government allowed both companies to release more broadly. The episode sets a precedent that will outlast this release: frontier launches are no longer treated purely as product decisions but as events with possible national-security weight, and that framing will harden as models grow more capable in software exploitation, biological research, autonomous operation, and strategic analysis.
It leaves an unresolved policy question about where oversight should end. A model that can find software vulnerabilities helps an attacker compromise a system and helps a defender find and fix the weakness first. Restricting access reduces the misuse and reduces the defensive benefit in the same motion.
The cybersecurity dilemma
Sol and Fable are unusually good at examining software. They can spot programming errors, explain how systems interact, and hunt for weaknesses across large codebases, which makes them valuable to security teams and, in the same breath, to attackers. The immediate danger is less that a model launches an attack on its own and more that it accelerates one, compressing the time and expertise once required for work only skilled professionals could do. The same request can be defensive or offensive depending on who is asking, and a model cannot always tell which context is legitimate, which is why developers have added guardrails around malware, exploitation, biological threats, and other sensitive areas. The disagreement is about how tight those guardrails should be.
OpenAI has taken the more permissive line, keeping broader access to advanced cybersecurity capability on the argument that defenders need the same tools attackers will seek. The reasoning rests on asymmetry: large enterprises and agencies run enormous software estates and need systems that can find vulnerabilities quickly, prioritize patches, and investigate suspicious activity, and over-broad restrictions can make a model safer in theory while stopping legitimate users from deploying it against real threats. The bet carries risk, since a permissive system is easier to manipulate through misleading context or fragmented requests, and can offer enough help across several small interactions to speed harmful work even while refusing a complete attack plan. OpenAI is wagering that monitoring, account controls, and targeted limits can hold that line.
Anthropic has gone the other way. It held back its most advanced system, Claude Mythos, over concern that its vulnerability-discovery and exploit-generation abilities were too dangerous to release widely, then shipped Fable 5, a more limited system that blocks sensitive cybersecurity and biological requests and routes them to a weaker model. That reduces access to the most dangerous capabilities and introduces a practical cost, because a security team trying to analyze malicious software or reproduce an attack may be refused precisely the assignment for which advanced reasoning is most useful. OpenAI is prioritizing utility and Anthropic is prioritizing containment, and neither choice resolves the underlying tension. A model permissive enough for serious defense is also useful for attack, and a model restrictive enough to stop attackers will frustrate the people trying to stop them.
Two theories of security
Underneath the dispute are two theories of technological security. One holds that dangerous capability should stay scarce, confined to trusted organizations, approved researchers, and government partners, on the logic that fewer users means less abuse. The other holds that defensive capability has to be distributed, because attackers will eventually reach comparable tools through other models, open-source systems, or their own development, and restricting responsible organizations mostly leaves defenders weaker. Both are partly right: scarcity slows proliferation without permanently preventing it, and distribution improves collective defense while enlarging the pool of capable bad actors. The workable balance probably differs by capability. A model that flags common software defects may be safe for broad release; a system that can autonomously discover and exploit unknown vulnerabilities may need tight control. The hard part is measuring where one category ends and the other begins.
The expanding role of government
The federal posture toward AI has been comparatively light, and that is getting harder to sustain as frontier models overlap with cyber operations, biotechnology, military systems, critical infrastructure, and geopolitical competition. Agencies will want earlier visibility into what these systems can do, and developers will resist oversight that delays releases, exposes proprietary information, or lets officials shape commercial technology. The likely settling point resembles regulation in other strategic industries: independent evaluations, documented risk controls, and capability reporting before deployment, with the most advanced systems held to different standards than ordinary commercial models.
The thorniest version is foreign access. Anthropic has already faced an order restricting certain foreign nationals from using Fable, a requirement that contributed to the system being taken offline for a time. Nationality-based rules are difficult to administer in a global software market where companies employ international staff, serve multinational customers, and run cloud infrastructure across jurisdictions, and a control designed to protect national security can collide quickly with ordinary business operations. The more defensible restrictions will focus on behavior, institutional risk, and sensitive applications rather than nationality alone, and even those will be hard to enforce.
Meta’s cheaper alternative and the three-tier market
Meta released its own model the same day as Sol. It is less capable than the flagships from OpenAI and Anthropic and considerably less expensive, and that gap may matter more than the capability gap, because the most powerful model does not automatically become the most used. Many commercial tasks do not need frontier reasoning, and a business will often prefer a cheaper system that performs adequately at high volume. Meta’s broader strategy leans on scale, low cost, and open distribution, which pressures the premium providers even when its benchmark numbers trail.
- Frontier The most capable and most expensive systems, reserved for the hardest and highest-stakes work. OpenAI and Anthropic occupy this end.
- Midrange Capable general-purpose models for the bulk of enterprise work, where reliability and price matter more than a few benchmark points.
- Compact Small systems that run cheaply, at high volume, or directly on devices. This is the layer Meta is positioning to commoditize.
As capabilities converge, the pricing pressure on the top tier only intensifies, and the interesting strategic question stops being who has the best model and becomes who captures each layer.
Permission becomes the design problem
An agent is useful only when it has access. To schedule a meeting it needs the calendar; to build a financial report it needs the files and the data; to answer email it needs the history; to update a record it needs write access to the database. Every additional permission raises utility and raises risk in equal measure, because a compromised or mistaken agent can expose confidential material, send an unauthorized message, alter records, or complete a transaction incorrectly.
So organizations will need fine-grained control over what an agent can see and do. It might be allowed to draft an email but not send it, to analyze an account but not move money, to propose spreadsheet changes but require approval before saving. That turns autonomy into a spectrum rather than a switch, and the strongest enterprise products will compete on granular permissions, complete audit trails, and reliable ways to reverse an action. Intelligence is only one part of such a product. Governance is the other, and it is the part most buyers have not yet learned to price.
The liability question
As agents grow more autonomous, responsibility gets harder to assign. Suppose a system misreads an email and cancels a critical meeting, or edits a financial model incorrectly and a decision is made on the result, or sends confidential information to the wrong customer. Who is accountable, the employee who issued the instruction, the company that deployed the agent, the software provider, the model developer, or the vendor whose application permitted the action? Contracts will try to allocate the risk, but the legal structure is underdeveloped, and until organizations can predict not only what these systems can do but what happens when they fail, most will confine agents to low-risk work.
What the new OpenAI model changes about the competitive map
Sol clarifies the shape of the field. OpenAI is pursuing broad distribution, enterprise agents, and comparatively permissive access to advanced capability. Anthropic is emphasizing professional performance and more restrictive safety controls. Meta is using price and availability to challenge the economics of premium models. Google and Microsoft hold distribution advantages through operating systems, cloud platforms, and productivity software. None of it will be decided by a single benchmark; it will be decided across capability, operating cost, safety policy, enterprise integration, developer adoption, distribution, data access, and trust, and a company can lead on one axis and lose the market on another. The most capable model may be too expensive, the safest too restrictive, the cheapest insufficient for high-value work, and the winner will be whoever combines adequate intelligence with the best economic and operational fit.
What this means for the mid-market
For a company between ten million and a billion dollars in revenue, almost none of the frontier drama changes the next decision, though one part of it changes a great deal. The contest between Sol and Fable is, for you, a spectator sport: you will rent whichever is winning that quarter, by the call, and the price will keep falling on its own. What is genuinely new is that software can now operate your software, and that is the moment our three rules stop being counsel and become a control system.
Score every workflow on the Delegation Ladder before an agent touches it, because "draft the email" and "send the email" are now different products with different blast radii. Match the model to the half-life of the problem, and let the cheaper tier do the routine work a frontier model is wasted on. And treat permissions and audit trails as the actual purchase, since an agent holding your calendar, your inbox, and your ledger is only as safe as the access you grant it. The firms that do well in the agent era will be the ones that decided, in writing, exactly what their software is allowed to do without asking. By then the decision that matters is a governance decision: how much of the work a company is willing to let its software do on its own.