explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people will ask
  • What OpenAI actually built
  • The numbers, and how to read them
  • What the model was not trusted to do
  • Why this matters beyond contracts
  • What this means if you build or buy
  • What is still unknown
  • Related reading
← Back to blog

explainx / blog

OpenAI and Ironclad: How GPT-6 Astra Was Trained on Contracting Workflows

OpenAI, GPT-6 Astra, AI Agents, Computer Use, Legal Tech

OpenAI trained GPT-6 Astra on 11 Ironclad contracting tasks: 55.0% vs 41.6% for GPT-5.6 Sol, in 19.2 vs 37.0 minutes. What it means for agent builders.

Oct 7, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
OpenAI and Ironclad: How GPT-6 Astra Was Trained on Contracting Workflows

OpenAI says the fastest way to make agents good at specialized business software is to train them inside that software with the people who use it. On October 6, 2026, it published "Advancing computer use with Ironclad", describing a research collaboration with the contracting platform Ironclad. The headline number: on 11 research tasks, GPT-6 Astra scored 55.0% against 41.6% for GPT-5.6 Sol, in an estimated 19.2 minutes per attempt versus 37.0. Both are OpenAI's own figures on an internal evaluation.

This post walks through what was built, what the numbers do and do not show, and what builders can reuse. For the model itself, see our GPT-6 Astra launch numbers and the Astra vs Sol comparison.

TL;DR: the questions people will ask

table · 2 cols
QuestionShort answer
Who is involved?OpenAI research and Ironclad, a contract lifecycle management company
What is new?Contracting workflows turned into RL training and eval tasks; Astra is the first frontier model trained on them
How many tasks?11 tasks across legal, commercial, and procurement work, each scored on 8 to 50 criteria
Headline resultAstra (Max reasoning) 55.0% vs Sol (High reasoning) 41.6% mean rubric score
Speed19.2 minutes vs 37.0 minutes estimated average time per attempt
Best internal model63.7% on the same tasks, per OpenAI
Independent verification?None. Numbers come from OpenAI's post
Can I join?OpenAI invites a small number of software companies to apply
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What OpenAI actually built

OpenAI frames the project as a follow-on to what it showed with Astra on professional computer work, such as preparing documents and testing websites. The stated goal is to make agents better at "specialized software to solve complex business problems": understanding a company's business rules, executing multi-step workflows, and verifying the result against the original requirements.

Rather than collecting generic screen recordings, OpenAI partners with software companies that know the workflows best. Ironclad is the first partner. The setup has four parts:

  1. Tasks chosen with experts. Ironclad employees, plus people who use Ironclad at OpenAI, helped researchers identify 11 tasks. Examples include setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable legal clause so it reflects the jurisdiction a requester selects.
  2. Realistic length. OpenAI estimates an experienced user would need about 30 to 40 minutes per task on average.
  3. Detailed rubrics. Each task was graded against 8 to 50 criteria depending on complexity, so researchers can see which parts a model got right and where it failed.
  4. A practice environment. Ironclad supplied hosted copies of its product. OpenAI researchers built synthetic training tasks around representative workflows and used reinforcement learning so the model improves through practice and feedback.

The example OpenAI gives is a legal operations team setting up a software-purchasing process. Finance approves purchases above a threshold, Security reviews certain requests, Legal reviews nonstandard terms. The agent has to turn that short list into an intake form, document templates, approval rules, and a record of the final agreement. Getting each click right is not enough. In OpenAI's words, the finished process must work across the situations it was designed to handle, including requests above and below the spending threshold.

The numbers, and how to read them

table · 3 cols
MetricGPT-5.6 Sol (High reasoning)GPT-6 Astra (Max reasoning)
Mean rubric score41.6%55.0%
Estimated time per attempt37.0 minutes19.2 minutes

OpenAI compared each model at the reasoning setting where it scored highest. It also says an internal model used in developing Astra reached 63.7%, and that it aims to bring those gains to future models.

A few cautions before repeating the "32% higher" figure:

  • It is a relative change. 55.0 versus 41.6 is 13.4 points. Expressed as a ratio of average scores that is about 32%, which matches OpenAI's wording, but a reader skimming "32%" might assume 32 points.
  • The sample is 11 tasks. That is small. A few tasks moving can shift the mean noticeably.
  • Time is estimated and simulated. OpenAI says "estimated time per attempt," and elsewhere "simulated time," so this is not wall-clock time in a production environment.
  • Training overlap. Astra is the first frontier model trained on Ironclad tasks, and Sol was not. Part of the gap is plausibly the model having practiced this exact product. OpenAI presents that as the point, but it means the result says less about how Astra handles a contracting tool it has never seen.
  • Self-reported. There is no third-party replication. For comparison, see how our coverage of post-launch benchmark changes showed that headline numbers can move after publication.

What the model was not trusted to do

The most useful sentences in the post are the cautious ones. OpenAI writes that the result "underscores why human oversight still matters as agents get better at complex contracting tasks, and why a full contracting platform remains essential." For context on what the company sells, see the Ironclad site, which describes AI contract lifecycle management software.

The takeaway: a 55% mean rubric score means the agent missed a meaningful share of requirements on average. In a legal or finance workflow, a missed approval rule is not a cosmetic miss. Treat this as evidence the training recipe works, not as a green light to remove review.

Why this matters beyond contracts

The recipe is the story

Strip out the legal domain and the structure is generic:

  1. Find a task people still do by hand in specialist software.
  2. Have experts define what "done correctly" means as a rubric with many small criteria.
  3. Build a sandbox copy of the software.
  4. Train with reinforcement learning against the rubric.
  5. Re-measure against the previous model.

This is the same shape as the agent skills and evals practices that builders already use at smaller scale; our agent skills guide covers how to encode procedures for an agent, and a rubric-based eval is the natural way to check whether a skill works. If you build agents for one vertical, writing 10 to 15 tasks with 8 to 50 checks each is a realistic starting point.

Vendors are becoming training partners

OpenAI explicitly says it is partnering "with a small number of software companies." That is a shift in how enterprise software and model labs relate. Instead of only exposing an API or a plugin, a vendor can supply the environment and the definition of success, and get a model that is better at its product. It echoes the forward-deployed engineering trend where labs embed with customers, now applied to model training rather than only deployment.

For software companies, the open question is competitive: a model that learns your workflows is also learning workflows a competitor's agent could exploit. OpenAI's call for partners asks for "data that can be safely used for research," which hints at where the negotiation will happen.

Computer use keeps improving, unevenly

Computer use is the capability underneath all of this: seeing a screen and operating it. We covered OpenAI's expansion of it in Codex computer use on Windows and mobile, and the browser-use benchmark comparison shows how differently models score depending on the task. The Ironclad result adds a category: long, multi-step configuration work in business software, where the final state has to satisfy many rules at once.

What this means if you build or buy

  • If you buy legal or procurement software: expect vendors to announce agent features built on this kind of training. Ask what task set and rubric they measured against, and what percentage of criteria are met, not just that the agent "completes" workflows.
  • If you build agents: copy the rubric approach. Score partial credit per requirement, test above and below thresholds, and track time per attempt alongside accuracy.
  • If you are a software vendor: read the partner criteria in OpenAI's post. You need a concrete failing task, evidence of failure, a success definition, domain experts, a secure environment, and research-safe data.
  • If you work in legal ops: the near-term effect is faster first drafts of workflow configuration, with a human checking every approval path. OpenAI's separate legal product push is covered in our Astra for law post, and our AI and legal help guide covers where the limits lie.

What is still unknown

  • Which of the 11 tasks Astra failed, and which criteria it missed most often.
  • How the score changes on contracting tools the model never trained on.
  • Whether the gains hold at lower reasoning settings, since OpenAI compared each model at its best setting.
  • How much of the gap comes from the Ironclad-specific training versus general Astra improvements. OpenAI does not publish an Astra-without-Ironclad baseline.
  • Pricing or availability of any Ironclad product built on this work. The post describes research, not a launch.

If a follow-up appears, we will update this post.

Related reading

  • GPT-6 Astra launch: benchmarks and pricing
  • GPT-6 Astra vs GPT-6 Sol
  • OpenAI Astra for law
  • AI and law: legal help and contracts guide
  • GPT-6 Astra browser-use benchmark v2
  • Codex computer use on Windows and mobile
  • What are agent skills? Complete guide
  • Forward-deployed roles and the future of work

Figures and quotes are from OpenAI's October 6, 2026 post and are accurate as of publication; vendor research results can be revised.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 8, 2026

GPT-6 Astra Clears MazeBench and Every "I'm Not a Robot" Level

Two GPT-6 Astra capability demos went viral in the same 24 hours: a reported 7x lead over Claude Fable 5.1 on MazeBench, and a full clear of all 48 levels of the "I'm Not a Robot" browser game using computer-use tools. Here's what MazeBench measures, what the CAPTCHA clear actually shows about browser control, and why "beat a human test" isn't the same claim as AGI.

Sep 18, 2026

GPT-6 Astra Cracked a 1941 Enigma Message and a 1918 WWI Cipher

Two separate builders reported GPT-6 Astra decoding historical ciphers that had sat unsolved for decades — an 82-character 1941 German Army Enigma message (MVUEH) and a 1918 WWI German naval radio transmission from a public list of 50 unsolved ciphers. One result got direct sign-off from a working Enigma historian; the other has an honest, unresolved question about why a message using an already-known key sat unsolved for so long. Here's what actually happened, and what's still unverified.

Sep 18, 2026

Top 10 Things to Build With GPT-6 Astra (2026)

GPT-6 Astra's launch-week coverage produced dozens of demos, but most builders don't need a maze-solving CAPTCHA gauntlet — they need to know what's actually worth building with it today. These are ten concrete, buildable project ideas GPT-6 Astra is well-suited for, each grounded in a real demo or benchmark explainx.ai has already covered.