SDSignal Desk

Advancing computer use with Ironclad

Oct 6, 2026, 3:00 AM · OpenAI

Image: OpenAI

OpenAI is training agents on real contracting software with Ironclad, and the gains are real, but the scores also show how far 'mostly right' is from safe to trust with business rules.

Why it matters

OpenAI says it is partnering with a small number of software companies to turn hard, high-value workflows into training and evaluation tasks for its agents, and Ironclad, a contracting platform, is the first. Together they built 11 tasks across legal, commercial and procurement work, such as setting up nondisclosure agreements, building procurement approval processes, and updating a reusable clause to reflect a chosen jurisdiction. OpenAI estimates each would take an experienced user about 30 to 40 minutes.

GPT-6 Astra is the first OpenAI frontier model trained on these tasks. Across them, Astra averaged a 55.0 percent rubric score against 41.6 percent for GPT-5.6 Sol, with estimated time per attempt falling from 37.0 to 19.2 minutes. An internal model reached 63.7 percent.

From the desk

We like the method more than the marketing. Grading each task against 8 to 50 specific criteria, practicing in hosted copies of the real software, and defining success with people who actually do the work is how agent evaluation should look. It's a welcome step away from vague demos. OpenAI also says training tasks were built from public SEC filings with personal information filtered out, not from customer or internal contracts, which is the right line to draw.

Now the sober part. A 55 percent average means the agent misses a large share of what these tasks require. Even the showcase clip, where Astra met about 94 percent of criteria, leaves a few requirements unmet. In contracting, the missing few percent can be the approval rule that lets a large purchase skip Finance, or the clause that cites the wrong jurisdiction. OpenAI and Ironclad say as much: an agent that loses track of one rule halfway through limits what anyone can confidently ask it to do, which is why human oversight still matters.

We'd also flag what the numbers aren't. OpenAI's footnotes say the times are simulated estimates based on assumed model speeds, not measured customer savings, and the results cover only these 11 tasks, not Ironclad's workflows broadly. That's honest, and it means nobody should read this as proof that contract ops can be automated today.

Our read: this is the right kind of progress, and the partner model could make agents genuinely useful in specialized software. If it scales, expect frontier labs to shape their models around the workflows of whichever vendors join first, and expect pressure on legal and procurement teams to trust a system before its error rate earns that trust.

Context

OpenAI is inviting other software companies to apply, asking for concrete failing tasks, people who know the work deeply, secure test environments and data safe for research. Ironclad's CTO, Sunita Verma, framed the goal as agents that understand the full contracting lifecycle while preserving the controls teams rely on.

Who feels it

Legal and procurement teams
Agents are getting better at configuring contracting workflows, but error rates still demand careful human review.
Enterprise software vendors
A new path to have frontier models trained on their product's workflows, with influence over how success is defined.
Ironclad customers
Could see more capable AI features over time, built on models that have practiced on Ironclad-style tasks.
AI evaluators
Criteria-based, task-specific rubrics offer a more concrete benchmark model than broad leaderboards.

What to watch

  1. Which software companies join as the next research partners
  2. Whether OpenAI publishes measured, not simulated, time savings from real deployments
  3. When the internal model's 63.7 percent performance reaches a shipping model
  4. How Ironclad builds review and control steps around agent-configured workflows

Read the original

Continue at the source.

OpenAI

Companies: OpenAI