SDSignal Desk

GPT-6 Astra: The next generation in intelligence for work

Sep 9, 2026, 4:00 AM · OpenAI

Image: OpenAI

OpenAI's work-focused Astra pitch sells computer use inside everyday apps, leaner cost per task, and tighter enterprise controls—while confirming Critical cyber capability and off-by-default admin gates.

Why it matters

OpenAI is positioning GPT-6 Astra as its most capable model for business work, now in ChatGPT Work, Codex, and the API. The company claims state-of-the-art results on computer use, browsing, professional work, software engineering, cybersecurity, and science—and emphasizes that Astra can operate the same applications people already use, even when those apps lack an API.

Early customer quotes in the post span Cognition, Databricks, Hebbia, Box, Figma, Thomson Reuters Labs, Datacurve, Basis, CodeRabbit, and XTX Markets: better decks, fewer unsupported assertions, stronger coding and long-horizon agents. Pricing starts at $10 per million input tokens and $50 per million output. Enterprise access is off by default until administrators enable it.

Astra is also described as OpenAI's most aligned model yet and the first to reach the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger training against unauthorized actions and new admin controls over sites, desktop apps, uploads, and downloads.

From the desk

We're treating this post as the enterprise packaging of Astra, not a second origin story. The work claim that matters is computer use without a months-long integration project: write code, drive GUIs, follow brand templates, finish jobs inside tools that were never built for agents. If that holds outside launch quotes, it compresses the distance between "pilot" and "in the workflow." Internal anecdotes—multicamera footage cut into a developer video with hundreds of thousands of views in days, a memory-allocator fix that cut Codex turn latency dramatically—are colorful; customer evals are the sturdier signal.

Efficiency is the second pillar. OpenAI says Astra completes tasks in fewer tokens with fewer retries and sits on much of the cost-efficiency frontier for professional work and coding evals, citing Terminal Bench 4.0 at 57.9% versus 37.3% for GPT-5.6 Sol 2 and 55.8% for Claude Fable 5.1, at lower estimated API cost per task. That is a procurement argument as much as a research one: more useful work per dollar, not just a higher ceiling.

Then the control story. On an internal computer-use safety benchmark for hard business mistakes—leaking confidential data, oversharing dashboards, deleting data—OpenAI says Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1, with confirmation policies and automated review improving further. Pair that with Critical cyber capability and you get the modern frontier bind: the same model class that can help secure systems can also abuse them. OpenAI's answer is refusals, monitoring, enterprise allowlists, and plugins into Oracle Analytics, Power BI, Navan, and Avalara via desktop browser use.

Our read: useful AI for work earns the benefit of the doubt when admins start narrow and expand access. The harm path if this scales carelessly is obvious—agents with GUI reach and Critical cyber skill operating on production systems with rubber-stamp approvals. Benchmarks and vendor quotes are not audits. We're watching whether off-by-default Enterprise stays meaningful and whether independent teams reproduce the computer-use safety deltas once access is wide.

Context

Zero Data Retention remains available for eligible API customers on supported endpoints, subject to approval. The post frames Astra as succeeding prior GPT-5.6 Sol comparisons on several public and partner evaluations.

Who feels it

Enterprise IT and security
Off-by-default access, allowlisted apps/sites, confirmation policies, and Critical cyber designation should drive staged rollouts—not blanket enablement on day one.
Knowledge workers and analysts
Stronger template-following and document-grounded decks can cut rework; Box's point on fewer confidently wrong assertions is the feature to verify in real queues.
Software teams
Codex plus computer-use gains and partner coding evals suggest fewer steps on long-horizon engineering—still dependent on harness quality and review discipline.

What to watch

  1. Share of Enterprise tenants that enable Astra and how narrowly they configure allowlists.
  2. Independent reproductions of Terminal Bench and computer-use safety comparisons versus Sol and Claude Fable.
  3. Whether desktop enterprise plugins expand beyond the first Oracle Analytics, Power BI, Navan, and Avalara set.

Read the original

Continue at the source.

OpenAI

Companies: OpenAI