SDSignal Desk

Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

Sep 19, 2026, 6:00 AM · TechCrunch

Image: TechCrunch

Vals raised a $40 million Series A led by Andreessen Horowitz to sell private, domain-specific AI benchmarks as public leaderboards keep getting gamed.

Why it matters

Lucas Ropek’s TechCrunch report profiles Vals, a 2024-founded evaluation startup that closed a $40 million Series A led by Andreessen Horowitz last month after a prior seed led by 8VC and Bloomberg Beta. Co-founder Rayan Krishnan argues academic benchmarks lag frontier models and that public tests invite training-to-the-test.

Vals keeps test materials private and scores models on complex work in law, finance, and coding—asking whether output matches human-quality products in a domain, including negative outcomes if models “ran wild.” The company also says it is building benchmarks for recursive self-improvement, mental health, cybersecurity, biosecurity, and law of armed conflict / Geneva Convention application.

Revenue is described as eight times last year’s; headcount grew from eight at the start of the year to 25, with plans to hire 10–15 more. Vals recently launched a federal-agency evaluation program. Krishnan frames private evals as central to how AI companies will talk to investors and, eventually, public markets.

From the desk

We’re glad someone is building evaluations that look like real work instead of another trivia leaderboard—and we’re uneasy when the grade-keeper is paid by the student.

Krishnan’s diagnosis is fair: public benchmarks age fast, and models get optimized against them. Private task suites in law and finance are how enterprises should buy models. Checking for downside behavior, not only pass rates, is the right instinct for biosecurity and cyber.

The SAT analogy only goes so far. The College Board does not also sell consulting to the schools being ranked. Vals’ customers pay to be tested; that can fund better suites, or it can create pressure to soft-pedal harsh results. Independence is the product, and independence is hard when the invoice comes from the lab.

Useful AI needs trustworthy measurement. We’re for domain evals that survive marketing decks. I’m watching whether Vals publishes enough methodology for outsiders to trust scores without publishing the answer key—and whether federal use includes conflict-of-interest rules when the same vendor grades vendors.

Context

TechCrunch published the profile on September 19, 2026. Krishnan previously interned at Palantir and, as a Stanford undergrad, worked at Microsoft and Stanford’s AI lab; Vals’ San Francisco office is on Folsom Street in a former brewery building.

Who feels it

Enterprises
Private domain scores may beat public leaderboards for procurement—if methodology and failure modes are disclosed.
Frontier labs
Paying for evals only helps if tough scores are shared internally and with customers, not buried.
Federal buyers
Agency eval programs need clear independence rules when vendors grade the models they sell to.

What to watch

  1. Whether Vals publishes aggregate methodology and known failure modes without releasing test keys.
  2. How federal agencies structure contracts and conflict disclosures with Vals.
  3. Whether harsh private scores leak into public filings as AI companies approach IPOs.

Read the original

Continue at the source.

TechCrunch

Also covering this