Skip to content

Product2 publishers3 min readPublished

Snorkel books a $375m run rate a year after switching from tooling to finished datasets

The $350m Series E values Snorkel AI at $3.5b, roughly nine times the annualized revenue it reports. What a customer buys now is a written evaluation rubric, a graded dataset and a sandbox to train in.

The Product Desk · Product desk

Illustration accompanying Snorkel books a $375m run rate a year after switching from tooling to finished datasets

What happened

  • Snorkel AI disclosed a $350 million Series E at a $3.5 billion valuation, led by Insight Partners and S32, with GV, Addition, Lightspeed, Greylock and Wells Fargo among the participants.
  • The price is nearly triple the $1.3 billion valuation the company carried 17 months ago, when it raised $100 million in a Series D.
  • TechCrunch places peer data suppliers on a similar curve, with Mercor at $2 billion gross annualized revenue, Handshake past $1 billion this year and Micro1 at $500 million.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost A customer buying graded data is paying for expert time it no longer manages, and because that time is booked as cost of goods sold, the cost per verified task sits outside the reported run rate.
  • decision Teams that bought labeling automation to keep data work in-house now have to decide whether to keep running that pipeline or take delivery of finished datasets and rubrics written elsewhere.
  • exposure When the vendor owns the rubric and resolves reviewer disagreements, the customer's working definition of a correct answer lives outside the customer's own building.
  • contradiction TechCrunch treats 60-70% specialist payouts as the reason peer revenue figures overstate the business; Snorkel's answer moves the same payments to a different line of the income statement without changing who gets paid.

Two human reviewers score the same model answer differently. In Snorkel AI's account of its reinforcement learning work, that gap usually means the written evaluation criteria are inconsistent, so the criteria get rewritten [8]. The rubrics can run several pages. For a programming task, the guidance has to cover every cybersecurity and performance requirement the generated code must meet [7].

That document is now part of what the customer buys. Snorkel's first product was Snorkel Flow, software that cut the labeling work in supervised learning projects using statistical methods its founders developed at the Stanford AI Lab [4]. Last year the company began selling ready-to-use datasets instead, alongside the rubrics and the virtual environments a model trains in, such as a simulated developer workstation for a code generation model [5][9]. Tens of thousands of human experts generate the reinforcement learning tasks [6]. TechCrunch reports the approach is hybrid, with Snorkel's own software and models producing data synthetically next to the subject matter experts [15].

Investors put $3.5b on a $375m annualized run rate, about 9.3 times revenue [1][1]. "Since launching our new data-as-a-service offering nearly a year ago, we've grown over 18 times, and this week crossed an annualized revenue run rate of $375 million," co-founder and chief executive Alex Ratner said in a blog post [10]. Seventeen months ago, a $100m Series D valued the company at $1.3b [11].

Anyone pricing this category has to ask where the labor shows up. Mercor's gross annualized revenue has reached $2b, Handshake passed $1b earlier this year, and TechCrunch reported Micro1 at $500m [12]. Those companies hand roughly 60% to 70% of top-line income straight to the domain specialists. Net revenue lands well below the headline [13]. Run that range against $375m and $225m to $263m would go out to experts, leaving $113m to $150m [2]. Snorkel told TechCrunch that it sells environments and completed datasets, and that it does not sell human labor, so expert payments sit in cost of goods sold instead of coming off the reported run rate [14]. TechCrunch's report does not include the size of those payments [17].

Teams buying this tell themselves they are buying data, and data is inputs. They are also moving the definition of a correct answer to a supplier. Does the evaluation rubric leave with the vendor when the contract ends, and can your own staff reproduce a score on a sample the vendor already graded? If the answer to both is no, the rubric is the product you rented, and next year's renewal is priced against a standard you cannot rebuild in-house.

Snorkel says the new money goes to hiring engineers, AI safety initiatives and support for open-source model evaluation benchmarks [16].

What to watch

  • Whether Snorkel publishes a gross margin now that expert payments sit in cost of goods sold.
  • Whether the open-source evaluation benchmarks Snorkel says it will fund overlap with the rubrics it charges customers for.
  • Whether Mercor, Handshake or Micro1 begin reporting net rather than gross annualized revenue.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories