🔍 Read the full analysis: A Question For Software Users: What Does Ironclad Say About OpenAI Training? on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
OpenAI described training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task rubric criteria, while estimated task times were simulations—not measured customer productivity gains. OpenAI says human oversight remains necessary and is inviting other software companies to propose research partnerships.
OpenAI said it trained its GPT-6 Astra frontier model in hosted copies of Ironclad’s contract-management software, testing it on 11 legal, commercial and procurement workflows. Astra met an average 55% of the criteria in the task rubrics, a result OpenAI presents as progress on complex software work—not evidence that the workflows are ready to run without human review.
OpenAI’s October 6 post, titled “Advancing computer use with Ironclad,” describes a collaboration in which Ironclad and OpenAI staff selected 11 tasks involving work such as creating nondisclosure agreements, building procurement approval processes and revising reusable contract clauses based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.
The tasks were graded against rubrics containing 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at high effort. An internal model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria; that result is for the showcase task, not the overall average.
OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered them to remove personal information. It said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. The post also reports estimated task times of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol. OpenAI’s footnote says those are simulated estimates based on assumed processing and generation speeds, not observed time savings for customers.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Workflow Accuracy Matters
The results matter because contract and procurement systems encode business rules and approval controls, not just steps to complete on a screen. If an agent creates a process but misses a spending threshold, a required Security review or a Legal check for nonstandard terms, the workflow may be unusable even if it satisfies many other rubric items. A 55% average is a measure of criteria met, not a 55% task-completion rate.
OpenAI’s own post acknowledges that losing track of a business rule limits the work companies can safely delegate, and says human oversight remains important. The performance figures therefore show a measured research result, not proof that businesses can hand contract workflows to an agent unattended. The simulated time estimates also cannot establish net productivity gains: they do not measure real customer use or the time required to verify and correct the output.
For software vendors, the collaboration points to a possible new role: providing the specialist workflows and secure environments in which models can practise. That may help agents work more effectively inside a product. It also makes the vendor’s underlying rules, records, audit trails and controls more central if customers increasingly interact through an agent rather than directly through the product interface. These are implications of the arrangement, not reported commercial outcomes.
contract management software with AI integration
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
The Ironclad work is framed by OpenAI as training models to understand an organization’s rules, carry out multi-step work in specialized software and check completed work against the original requirements. The tasks were chosen by Ironclad staff and OpenAI employees familiar with the product. The evaluation used a criterion-based rubric, so the reported scores represent the share of requirements met, rather than a simple count of tasks passed.
OpenAI says it used hosted copies of Ironclad’s product for model practice and synthetic tasks based on public SEC filings. It distinguishes those materials from non-public customer information, which it says was not used. The post also invites a small number of software companies to propose research partnerships. It asks prospective partners to bring a specific task that current agents struggle with, people who know the work, a secure test environment and data that can safely be used for research.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Show
The post does not establish how Astra would perform across Ironclad’s broader set of workflows or in day-to-day customer operations. The 11 research tasks are a limited sample, and the average criterion score does not reveal from the figures alone which requirements were missed on each task. The approximately 94% result applies to one showcase task and should not be generalized to the full set.
OpenAI’s simulated time figures are not measured customer outcomes, and the source does not report the time needed for a person to inspect, repair or approve an agent’s work. The post also does not provide details here about a commercial rollout, customer deployment results, or whether the model will be made available in a particular product. OpenAI says it did not use non-public Ironclad customer data, but further details about the secure testing setup and data-handling arrangements are not specified in the supplied material.
As an affiliate, we earn on qualifying purchases.
What OpenAI Is Seeking Next
OpenAI says it plans to work with a small number of software companies on tasks that current agents cannot reliably complete. Interested vendors are asked to provide a concrete failure case, subject-matter experts, a secure test environment and research-appropriate data. The post does not name additional partners or give a schedule for future collaborations.
For businesses evaluating agents, the immediate next step is to seek task-level evidence rather than rely on a single average score. Buyers can ask which rubric requirements failed, how approval controls are preserved, who reviews the result, and whether performance has been tested in their own workflows. Further information on deployment, customer validation and future partner projects would clarify whether the research translates into dependable operational use.
AI legal document drafting software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad test?
They tested a frontier model on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.
Does Astra’s 55% score mean it completed 55% of the tasks?
No. OpenAI says 55% is the average share of rubric criteria met across the tasks. It is not a task completion rate or a pass rate.
Were the reported time savings measured with customers?
No. OpenAI describes the figures as simulated estimates based on assumed processing and generation speeds, not measured customer time savings.
Did OpenAI use private Ironclad customer contracts?
OpenAI says it used synthetic tasks based on publicly filed SEC contracts and did not use non-public Ironclad customer data, OpenAI customer data or OpenAI internal contracts.
Can companies use these agents for contract work without review?
The reported results do not establish that. OpenAI says human oversight remains important, and the average score indicates that many rubric criteria were not met in the evaluation.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
