AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Question For Software Users: What Does Ironclad Say About OpenAI Training? on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task rubric criteria, while estimated task times were simulations—not measured customer productivity gains. OpenAI says human oversight remains necessary and is inviting other software companies to propose research partnerships.

OpenAI said it trained its GPT-6 Astra frontier model in hosted copies of Ironclad’s contract-management software, testing it on 11 legal, commercial and procurement workflows. Astra met an average 55% of the criteria in the task rubrics, a result OpenAI presents as progress on complex software work—not evidence that the workflows are ready to run without human review.

OpenAI’s October 6 post, titled “Advancing computer use with Ironclad,” describes a collaboration in which Ironclad and OpenAI staff selected 11 tasks involving work such as creating nondisclosure agreements, building procurement approval processes and revising reusable contract clauses based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.

The tasks were graded against rubrics containing 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at high effort. An internal model used during Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria; that result is for the showcase task, not the overall average.

OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered them to remove personal information. It said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. The post also reports estimated task times of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol. OpenAI’s footnote says those are simulated estimates based on assumed processing and generation speeds, not observed time savings for customers.

At a glance
reportWhen: Published October 6; results and partne…
The developmentOpenAI published details of a collaboration with Ironclad to train and evaluate a frontier model on workflows inside the contract-management product.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Workflow Accuracy Matters

The results matter because contract and procurement systems encode business rules and approval controls, not just steps to complete on a screen. If an agent creates a process but misses a spending threshold, a required Security review or a Legal check for nonstandard terms, the workflow may be unusable even if it satisfies many other rubric items. A 55% average is a measure of criteria met, not a 55% task-completion rate.

OpenAI’s own post acknowledges that losing track of a business rule limits the work companies can safely delegate, and says human oversight remains important. The performance figures therefore show a measured research result, not proof that businesses can hand contract workflows to an agent unattended. The simulated time estimates also cannot establish net productivity gains: they do not measure real customer use or the time required to verify and correct the output.

For software vendors, the collaboration points to a possible new role: providing the specialist workflows and secure environments in which models can practise. That may help agents work more effectively inside a product. It also makes the vendor’s underlying rules, records, audit trails and controls more central if customers increasingly interact through an agent rather than directly through the product interface. These are implications of the arrangement, not reported commercial outcomes.

Amazon

contract management software with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

The Ironclad work is framed by OpenAI as training models to understand an organization’s rules, carry out multi-step work in specialized software and check completed work against the original requirements. The tasks were chosen by Ironclad staff and OpenAI employees familiar with the product. The evaluation used a criterion-based rubric, so the reported scores represent the share of requirements met, rather than a simple count of tasks passed.

OpenAI says it used hosted copies of Ironclad’s product for model practice and synthetic tasks based on public SEC filings. It distinguishes those materials from non-public customer information, which it says was not used. The post also invites a small number of software companies to propose research partnerships. It asks prospective partners to bring a specific task that current agents struggle with, people who know the work, a secure test environment and data that can safely be used for research.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Do Not Show

The post does not establish how Astra would perform across Ironclad’s broader set of workflows or in day-to-day customer operations. The 11 research tasks are a limited sample, and the average criterion score does not reveal from the figures alone which requirements were missed on each task. The approximately 94% result applies to one showcase task and should not be generalized to the full set.

OpenAI’s simulated time figures are not measured customer outcomes, and the source does not report the time needed for a person to inspect, repair or approve an agent’s work. The post also does not provide details here about a commercial rollout, customer deployment results, or whether the model will be made available in a particular product. OpenAI says it did not use non-public Ironclad customer data, but further details about the secure testing setup and data-handling arrangements are not specified in the supplied material.

Amazon

contract review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What OpenAI Is Seeking Next

OpenAI says it plans to work with a small number of software companies on tasks that current agents cannot reliably complete. Interested vendors are asked to provide a concrete failure case, subject-matter experts, a secure test environment and research-appropriate data. The post does not name additional partners or give a schedule for future collaborations.

For businesses evaluating agents, the immediate next step is to seek task-level evidence rather than rely on a single average score. Buyers can ask which rubric requirements failed, how approval controls are preserved, who reviews the result, and whether performance has been tested in their own workflows. Further information on deployment, customer validation and future partner projects would clarify whether the research translates into dependable operational use.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI and Ironclad test?

They tested a frontier model on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software.

Does Astra’s 55% score mean it completed 55% of the tasks?

No. OpenAI says 55% is the average share of rubric criteria met across the tasks. It is not a task completion rate or a pass rate.

Were the reported time savings measured with customers?

No. OpenAI describes the figures as simulated estimates based on assumed processing and generation speeds, not measured customer time savings.

Did OpenAI use private Ironclad customer contracts?

OpenAI says it used synthetic tasks based on publicly filed SEC contracts and did not use non-public Ironclad customer data, OpenAI customer data or OpenAI internal contracts.

Can companies use these agents for contract work without review?

The reported results do not establish that. OpenAI says human oversight remains important, and the average score indicates that many rubric criteria were not met in the evaluation.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Raw-feed licensing. The contract that doesn’t exist yet.

A new licensing category for raw-feed downstream rewriting lacks an industry-standard contract, risking legal and economic issues in AI content use.

Georgia Republicans decline to redraw congressional map in defiance of Trump

Georgia Republicans have decided not to redraw congressional districts during a special session, opposing Trump’s push amid a Supreme Court ruling affecting voting rights.

Georgia Republican legislative leaders reject governor’s call for 2028 redistricting

Georgia Republican leaders refuse to consider redistricting during a special session, citing legal and strategic concerns amid civil rights opposition.

Estate And Inheritance Facilitator Marketplace

A new marketplace aims to simplify estate settlement by guiding executors through steps and connecting them with vetted facilitators, starting with a pilot program.