Inside Ironclad’s contract software, an AI training exercise asks an agent to arrange agreements and approvals. A well-written answer is not enough. The process it leaves behind must work.
OpenAI’s October 6 account of the collaboration describes hosted software, tasks shaped by experienced users and synthetic training exercises based on publicly filed contracts. Filters were intended to remove personal information. Private customer contracts were excluded. GPT-6 Astra was the first frontier model trained on Ironclad tasks.
The contribution went beyond documents: a place to practice, with people who could distinguish a working process from a plausible imitation.
“It has to do more than complete individual actions,” Sunita Verma, Ironclad’s chief technology officer, said in the published announcement.
But somebody must decide what counts as success. An agent can become very good at passing the test without becoming good at the job.
Researchers call a setting where an agent acts and receives feedback an environment. The agent might use a browser, change a file or run a program. Reinforcement learning uses scores from its attempts to adjust the model’s behavior. An environment can also be used just for testing, without changing the model at all.
Consider a password reset. A transcript of a successful request shows the words exchanged. Performing the reset requires deciding whether the requester can access the account, choosing an available recovery method and dealing with a failed check. Each decision changes the next one. The transcript preserves one route through that problem. Working software allows an agent to try others, including routes that end badly. Those failures could be useful practice if the system can identify what went wrong.
Microsoft Research explored a related problem in its October 7 account of Agent Lightning, software for training agents. The researchers wanted agents to practice with the same surrounding tools they would use afterward. Rebuilding those tools for training can change behavior, they wrote.
Using roughly 6,000 training samples, the team raised one model’s single-attempt success rate on SWE-bench Verified, a software-repair test, from 41.8 percent to 56.4 percent. The result concerns that model and setup. It does not promise similar gains in other jobs.
The problem with a passing score
Software repair offers researchers something many office tasks lack: code that can be run, with tests that check whether it works.
In SWE-smith, a research project published in 2025, researchers built working environments around projects written in the Python programming language and deliberately introduced changes that broke existing tests. They produced 50,000 repair tasks from 128 public software projects. That work supplied the task foundation for Microsoft’s experiment.
The researchers also released records of agents working through the problems, not just the broken code. Those records preserve the steps taken along the way. For a developer studying why an agent succeeds or fails, the sequence can be more informative than the final repair.
A test, however, is a compressed description of what somebody wants. It is rarely the whole job.
Imagine an agent rewarded only for making a warning disappear. It might remove the warning rather than repair its cause. A check that accepts the right account balance could miss an unauthorized transfer. Repeating the exercise would then reward the shortcut, making a higher score poor evidence of improvement.
This weakness predates the current announcements. When OpenAI and the SWE-bench authors introduced SWE-bench Verified in 2024, they identified tasks with unclear instructions, tests that rejected valid solutions and environments that failed regardless of the repair. Professional developers reviewed the material to produce a 500-task subset. A disappointing score could reflect a flawed examination, not just a weak model.
The opposite mistake is more attractive commercially. An examination that is too forgiving makes a product look ready before it is. If the developer chooses the tasks, writes the scoring rules and announces the result, the reader needs to know how much the score says about work outside that setup. Improvement on selected exercises does not establish how well an agent will handle a customer’s messier operation.
Employees can help identify what a test leaves out. They may reject an apparently correct result or accept two very different approaches. Their disagreements can expose requirements nobody has written down. Turning those judgments into consistent checks takes time; some decisions will still need a person.
The starting conditions matter, too. Suppose one attempt changes an account and the next inherits that change. Two models are no longer facing the same problem. A service that goes offline halfway through can also make an agent look less capable than it is. Before interpreting the scores, researchers have to know whether the system presented the intended problem and remained available long enough to solve it.
A business that has to keep running
On September 28, Hugging Face added reinforcement learning environments to its Hub, where developers share models and datasets. The catalog separates collections of tasks from the software that runs them and provides ways to launch them through supported frameworks. It makes the material easier to discover; it does not supply the computers or turn a collection of files into working software.
The SWE-smith authors had already described the cost of assembling this material in their 2025 paper. Earlier methods could require hundreds of hours of human work, while the software environments occupied several terabytes of storage. A terabyte is a thousand gigabytes. The paper treated these demands as obstacles to research, and presented automation as a way to reduce the preparation burden.
That leaves an operating cost which a dataset’s size does little to explain.
A company providing access must maintain accounts, restore starting conditions and repair broken connections. Its staff may also have to settle disputes over results. A buyer seeking thousands of attempts needs to know whether the service can support them. For the supplier, a payment that covers preparing the tasks may be inadequate if keeping them available consumes months of engineering work.
The supplier’s incentives can be complicated. A more capable model could improve its product, while an outside agent that learns to do the same work might compete for some of its customers. Whether the arrangement helps the supplier depends partly on who sells the resulting service and keeps the customer relationship. That is a business question the training score cannot answer.
And greater realism can be a poor use of money. A team studying one decision may learn enough from a narrow simulation. Recreating an entire office application would add maintenance without necessarily making the experiment more informative. The case for a fuller environment becomes stronger when several steps depend on one another, so an early mistake changes the choices available later.
For a software company, contributing a training environment therefore means taking on work beyond exporting records. The next attempt needs a fresh account, a functioning application and a result somebody can trust enough to learn from. If those are missing, the company has supplied access to software, but it has not supplied a useful place to practice.



