SDSignal Desk

Asana cuts model costs 76x in browser tests with GPT-6.1 Sol

Oct 9, 2026, 12:00 AM · OpenAI

Image: OpenAI

Asana says an AI coding agent found why its browser agent was burning money, then fixed it. The 76x headline is narrow, but the method behind it is the useful part.

Why it matters

OpenAI published a customer story in which Asana used GPT-6 Astra in Codex to tune the browser agent inside StackAI, the automation platform Asana acquired. On GPT-6.1 Sol, the optimized workflow averaged $0.47 in estimated model cost and about four minutes per run, which OpenAI and Asana describe as 76x cheaper and 5x faster than the original production setup on a rival model.

The fix was not exotic. Astra found that the agent cached its fixed instructions but resent its growing browsing history of page text and screenshots at full price on every request. It also kept trimming that history, which broke caching and sometimes forced it to revisit pages. Asana's StackAI CTO, Frank Hidalgo, says work he estimated at one to two months by hand took about a week.

From the desk

We think this is a genuinely useful story, as long as it is read for what it is. It is a vendor case study, published by OpenAI, about OpenAI's models. The competing models are anonymized as A, B and C. The test was a single task, collecting six fields for 32 books from a public demo catalog, run three times per configuration across 144 runs. That is a careful experiment, but it is not a broad benchmark, and the 76x figure compares an optimized workflow on one model against an unoptimized one on another.

The more honest number may be buried in the middle. Applying the same fixes to the original production model cut its cost from at least $36.21 to $1.24 per run, about 29x. Much of the win came from engineering hygiene, caching history properly and pruning screenshots in batches, not from switching models. That is good news for every team running agents: plenty of cost is self-inflicted.

The reliability result matters too. Giving GPT-6.1 Sol a larger history budget took it from three of 18 runs producing an answer to all 18, each correct. For browser agents, memory management is not just a cost lever. It decides whether the job gets done.

And the workflow is the real headline. An engineer set direction, an agent mapped the code, refactored it for parallel experiments, ran the study, and logged everything for review before changes went to production. Hidalgo's line that human attention, not shipping speed, is now the bottleneck rings true. The risk is the flip side: as agents run experiments overnight, reviewing their work properly becomes the scarce skill, and teams that skip review will ship confident mistakes faster.

Context

StackAI lets customers build no-code workflows that navigate websites, fill out forms and gather information. Prompt caching lets repeated input be billed at a fraction of the normal price; in this study 89 percent of input came from cache at 5 percent of the uncached price. Asana and StackAI published the complete study on their blogs.

Who feels it

Developers building agents
Caching growing context and avoiding constant history edits can cut costs dramatically, regardless of which model is used.
Enterprises buying automation
Lower per-run costs could make stronger models affordable for routine browser tasks, but vendor claims deserve independent testing on real workloads.
Engineering leaders
Agent-run experiments can compress weeks of work into days, shifting the bottleneck to human review.

What to watch

  1. Whether Asana's planned cost, runtime and quality evaluations in StackAI show similar gains on real customer workflows
  2. Independent replications using the published study
  3. How Asana's use of agents for pre-release QA testing holds up

Read the original

Continue at the source.

OpenAI

Companies: OpenAI