Benchmarks have been broken for years and everyone knows it. The moment a test is public you can train against it, and from then on it measures who read it. Vals, out of San Francisco, is trying the opposite: the test material stays locked away. TechCrunch went to see the founder.
Who is doing this
Rayan Krishnan is 25. He worked at Palantir and Microsoft, and at Stanford’s AI lab while studying. Vals has existed since 2024. The seed round came from 8VC and Bloomberg Beta; last month a $40 million Series A led by Andreessen Horowitz followed.
His starting point: “We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance.”
What makes it different
Two things. First, the actual test material is never published, so a model cannot memorise it. Second, Vals does not measure general intelligence but work: tasks from law, finance and coding, the way they turn up on the job. Krishnan’s yardstick is not whether a model passes an exam but whether models “can do work that produces a product of the same quality as a human within every domain”.
Then come the subjects you would not expect on a benchmark list: recursive self-improvement, mental health, cybersecurity, biosecurity — and the law of armed conflict, meaning whether a model applies the Geneva Conventions correctly. Negative outcomes are explicitly part of the brief: what it would mean “if these models ran wild in the world”.
The examinee writes the cheque
Model providers pay to be tested. Krishnan compares it to the College Board: a student pays for the right to sit the SAT. Revenue is eight times what it was last year, the team has grown from eight people to 25, and another ten to fifteen are planned. There is now a dedicated programme offering evaluations to federal agencies.
Krishnan expects demand to rise as the IPOs arrive. SpaceX is already public, Anthropic is aiming for November, and OpenAI will probably follow.
This is exactly the model currently under dispute
Vals sells examinations to the examined and keeps the questions secret. Both are defensible and both are attackable — in the same week the industry is arguing about embedded evaluators. Anthropic has brought Accenture in-house, and a hundred specialists responded with five conditions for independence. One of them, roughly stated: whoever pays must not decide what gets published.
Keeping the tasks secret solves a real problem and creates a new one. A test nobody can recompute is a promise to the reader, not evidence. If Vals genuinely wants to be the gold standard, it will eventually have to say who gets to publish the results when they come out badly.