2026-09-19 18:23 UTC
DANGMUAAI & Developer Tools, Decoded
BackIndustry

Vals Raised $40M for a Benchmark Labs Can't Train On

Vals closed a $40M Series A led by a16z, selling private test sets and domain-task evals. Revenue is eight times last year; headcount went from 8 to 25.

DangMua EditorialSep 19, 20263 min read
Vals Raised $40M for a Benchmark Labs Can't Train On

Vals, a benchmarking startup founded in 2024, raised a $40 million Series A led by Andreessen Horowitz last month. Its pitch is a benchmark that AI labs cannot study for.

The contamination problem it is selling against

Many benchmarking systems publish their test materials. That means a lab can train a model against those tests, which, as TechCrunch puts it, is "arguably cheating on their exam." Many of the widely cited benchmarks are also older, and were not built to measure what current models do.

Vals does not publicly disclose its specific test materials. It also skips general-knowledge testing in favor of complex tasks tied to particular industries — law, finance, and coding.

"Historically, I think evaluation has been done to evaluate intelligence in a very abstract way," co-founder Rayan Krishnan told TechCrunch. "Like, do models know enough information to be able to take a bar exam type test?" His stated goal is narrower and more testable: "Can they do work that produces a product of the same quality as a human within every domain?"

Who pays, and why that is odd

Labs pay Vals to test their own models — an arrangement that invites the obvious question of why a company would pay to learn its model underperforms. Krishnan's answer is the SAT: he compares the revenue model to a student paying the College Board to sit the exam. The score is the product, and a credible score needs an examiner the test-taker does not control.

The growth numbers Krishnan gave TechCrunch: revenue is currently eight times what it was last year, and headcount went from eight people at the start of the year to 25, with plans to add another 10 to 15. Before the a16z round, the company raised a seed led by 8VC and Bloomberg Beta. It recently launched a program providing model evaluations to federal agencies.

The test list is the interesting part

Vals says it is pushing past conventional industry tasks. "We have a benchmark on recursive self improvement," Krishnan said. "We're doing some work in mental health, cybersecurity, biosecurity, and even law of armed conflict to models to understand how to apply the Geneva Convention." He frames the work as checking for negative outcomes too — analyzing, in his words, what the implications would be "if these models ran wild in the world."

That is a list of exactly the capabilities that decide whether a model ships with restrictions. Krishnan's bet is that those evaluations become financial infrastructure: "AI companies are starting to go public," he said, arguing that benchmarks will become "a central part of how these companies submit public filings or talk about the prospective investments they're going to make in AI."

What to watch

The structural tension in this model is that private test sets solve contamination by removing external verification. Nobody outside Vals can audit a score they cannot see, and the customer paying for the evaluation is the one being evaluated. That is survivable while buyers treat the results as procurement input. It gets harder if Krishnan's prediction lands and the same private scores start appearing in regulatory filings.

More from DangMua