SDSignal Desk

Basis completes a tax workbook 2x faster with GPT-6 Astra

Sep 27, 2026, 5:00 PM · OpenAI

Image: OpenAI

Accounting-agent startup Basis says GPT-6 Astra finished a 50-tab tax workbook in half the time of GPT-5.6 Sol and lifted internal eval scores about 20%.

Why it matters

OpenAI’s customer story says Basis, which builds AI agents for accountants, compared GPT-6 Astra with GPT-5.6 Sol on a complicated 50-tab tax workbook. Co-founder Mitch Troyanovsky says Astra completed the workbook in half the time and better understood user intent, taking a more direct path with fewer corrections and more efficient token use.

Basis also uses Astra to dial reasoning effort up or down as a task progresses while keeping cache intact—aimed at cost and latency on long jobs. Internal evaluation scores improved about 20%, including when to ask questions, flag assumptions, and follow templates and primary sources.

From the desk

We’re filing this as a vendor case study with a useful data point: long-horizon accounting work is where agent models earn or lose trust.

Useful AI for accountants should shrink workbook grind so humans can do judgment. A 2x wall-clock win on a 50-tab book is the kind of evidence that moves buyers—if it replicates outside Basis’s harness. Adaptive reasoning effort is the quieter innovation; spending compute only on hard steps is how agent economics stop being silly.

I’m watching independence. This is OpenAI’s site quoting a customer on OpenAI’s newest model. Treat the percentages as directional until peers publish similar bakeoffs. Tax work also has a wrong-answer cost that chat demos never show; evals that check primary-source consultation matter more than speed alone.

If Astra-class models keep cutting long-task time without raising silent error rates, professional services AI gets real. If speed comes from skipping checks, firms will learn that the expensive way.

Context

OpenAI customer story, September 28, 2026. Basis is a North America technology startup using the OpenAI API.

Who feels it

Accounting and tax firms
Worth a controlled pilot on non-client workbooks; keep human review on filings.
Agent builders
Adaptive reasoning + cache-aware long tasks is a pattern to copy carefully.
OpenAI rivals
Expect similar vertical case studies; demand comparable task definitions.

What to watch

  1. Third-party replications of the 50-tab workbook comparison
  2. Error-rate disclosures alongside speed claims
  3. Whether Basis’s ~20% internal eval lift shows up in customer outcomes

Read the original

Continue at the source.

OpenAI

Companies: OpenAI