Skip to content

Product1 publisher2 min readPublished

Vals sells a private AI exam to the labs it grades

Vals keeps its test materials private and charges the model developers it scores, and Andreessen Horowitz has now put $40 million behind that arrangement. Buyers reading the numbers cannot run the test themselves.

The Product Desk · Product desk

Photograph accompanying Vals sells a private AI exam to the labs it grades
Photo: techcrunch.com

What happened

  • Vals, a benchmarking startup formed in 2024, raised $40 million in a Series A led by Andreessen Horowitz last month, after a seed round led by 8VC and Bloomberg Beta.
  • Unlike benchmarks that publish their tests, which a developer can train against, Vals withholds its test materials and scores models on industry tasks in law, finance and coding.
  • The model developers being scored are the ones paying for the tests, an arrangement co-founder Rayan Krishnan compares to a student paying the College Board to sit the SAT.
  • Revenue is currently eight times what it was last year, according to the company, which is also looking for a significantly bigger office than its two floors on Folsom Street.
  • Vals recently launched a program providing model evaluations to federal agencies, and Krishnan casts benchmarking as how AI companies will establish public trust as they go public.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • exposure Anyone quoting a Vals score in a procurement memo is relying on a report the graded lab commissioned. The conflict lands on whoever reads the number.
  • constraint Keeping the items private stops labs training on the test and also stops a buying team from rerunning it, so the score stays unchecked against the tasks that team actually cares about.
  • decision Once an agency can cite a third-party evaluation in a bid, a model vendor has to choose between sitting the exam and explaining to the contracting officer why it did not.
  • precedent Eight-times revenue growth shows the graded party will pay for its own exam. The next benchmark firms will pitch labs first, and buyer-funded evaluation stays the harder business to fund.

A team choosing between two models this quarter is reading a chart the vendor drew. TechCrunch, in its piece on Vals, states the incentive plainly: good benchmarks pretty much always mean good PR [5]. The same piece reports that companies have figured out how to outwit legacy benchmarking systems, many of them older and not built to measure modern models [6]. One route is mechanical. A benchmark that publishes its questions can be trained against [7].

Krishnan's answer is to keep the questions in-house and to score domain work instead of general knowledge, in law, finance and coding [7][8]. "Historically, I think evaluation has been done to evaluate intelligence in a very abstract way," he said, describing bar-exam-style tests [11]. The question he says he wants answered instead: "Can they do work that produces a product of the same quality as a human within every domain?" [10]

The money runs the opposite way from what a buying team might assume. Model developers pay Vals to test them, and Krishnan likens the arrangement to a student paying the College Board to sit the SAT [12]. TechCrunch reports those evaluations are becoming key decision-making factors for companies looking to acquire new AI models [13]. So the developer pays for the test, and the tasks inside it stay private from the buyer leaning on the score [7][12]. Whether a client can keep a poor result unpublished goes unanswered [20].

Labs are buying. Vals started this year with eight people and now has 25, with another 10 to 15 planned [15]. That lands at 35 to 40, between four and five times the January headcount, inside twelve months [19]. The subject matter is widening into places where an outside grade may be the only grade: Krishnan lists a benchmark on recursive self improvement, plus work in mental health, cybersecurity, biosecurity and the law of armed conflict, testing whether models can apply the Geneva Convention [17]. He said the hope is to analyze how, "if these models ran wild in the world, what the negative implications would be" [18].

How much weight a benchmark number deserves in a vendor decision comes down to who paid for the test, whether your team can inspect the items, and whether the tasks resemble the work your users actually bring. Vals is paid by the developer, and the items stay in-house [12][7]. The match between its tasks and your users' work is the one a deployment team can settle without anyone's permission, by running the last fifty real tickets through both candidates and reading the output.

What to watch

  • Whether Vals ever publishes a score for a model whose developer did not commission the test.
  • Whether the federal evaluation program names the agencies buying it and the contract vehicle they buy it on.
  • Whether any lab cites a Vals result that is worse than the benchmark numbers it publishes itself.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories