Skip to content

Product1 publisher3 min readPublished

TypeSafe leaves stealth with a model whose confidence scores developers have to calibrate themselves

TypeSafe's first model, Jev, answers structured questions with typed output and a confidence measure attached. The account of its launch says developers have to check those scores against real outcomes on their own data first.

The Product Desk · Product desk

Photograph accompanying TypeSafe leaves stealth with a model whose confidence scores developers have to calibrate themselves
Photo: dcvc.com

What happened

  • TypeSafe AI came out of stealth with $40 million in seed funding led by DCVC, and Forbes reported the firm valued the San Francisco company at $200 million, citing a person familiar with the transaction.
  • Chief executive Diogo Almeida founded the company in 2024 after working on reinforcement learning from human feedback, InstructGPT, ChatGPT and GPT-4 at OpenAI, with co-founders Erik Gafni and Sasha Sheng.
  • Its first model, Jev, answers structured questions with typed output such as a yes/no probability, a pick from a defined list or a score on a set scale, and TypeSafe calls it a System One Model.
  • TypeSafe's website lists Jev at 39 cents per 1,000 workflows against $3.31 for OpenAI's Gpt-5.6 Luna and a figure printed as R19.49 for Anthropic's Claude Haiku 4.5.
  • Jev is available only through an early-access waitlist.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Before pricing anything, a buyer has to settle whether the queue's answers can be enumerated up front. A typed model only holds a question whose answer set someone has written down.
  • cost The cheap per-workflow rate becomes a saving only after someone assembles a labelled sample to score the confidence numbers against. That work falls on the team that owns the queue.
  • contradiction TypeSafe's own account carries two speed multiples: up to 100 times faster in one line, nearly 194 times in another. A business case built on either needs the buyer's own workload test to stand up.
  • constraint Access is gated behind a waitlist, so evaluation happens on TypeSafe's schedule. Teams comparing it to an incumbent model run the bake-off when TypeSafe grants access, whatever their own planning cycle demands.

The developer wiring Jev into a claims queue gets branches instead of paragraphs. Each decision comes back with probabilities and a confidence measure. The developer sets the thresholds that decide whether the workflow proceeds on its own, goes back for more evidence, or lands on a person's desk, according to SiliconANGLE [5]. TypeSafe's own worked example is property underwriting: Jev estimates the likelihood a building catches fire, confident cases run through, ambiguous ones go to an underwriter [6].

That threshold is only useful if the confidence number tracks real accuracy. SiliconANGLE reports that developers need to test how well Jev's scores correspond to actual accuracy on their own data [14]. In practice that means a pile of invoices, tickets or claims where the right answer is already known and written down. A team with three years of adjudicated outcomes can run that check in a sprint. A team whose past decisions live in one senior person's judgement cannot run it at all, and will set the cutoff at 0.85 because 0.85 looks like a sensible number.

The price list is the easiest claim to check. At the listed rates, a million workflows costs $390 on Jev and $3,310 on Gpt-5.6 Luna, a difference of $2,920 [2]. That is a factor of 8.5 on the vendor's own page [1]. TypeSafe also says its internal tests found Jev about 445 times cheaper than the language models it compared against, roughly 52 times the ratio those two listed prices imply [11][3]. SiliconANGLE says the test figures have not been independently verified and will vary by workload, network location and comparison method [12].

Almeida, who worked on reinforcement learning from human feedback and ChatGPT before founding the company, put the design argument this way to Forbes: "We've been optimizing for humans, and we're superhuman at pleasing humans" [16]. The company's stated thesis is that the same qualities work against a model inside production software. There it can generate plausible but incorrect output, vary its method between requests, and state uncertain answers confidently [19]. DCVC general partner James Hardiman said TypeSafe is addressing "one of the biggest remaining challenges in AI" by making models reliable enough to embed in products at scale [17].

Where TypeSafe says the money is: high-volume plumbing such as classifying service requests, evaluating invoices, triaging security alerts and reviewing what AI agents produced. In that setup a specialised decision model acts as a control layer while a language model drafts text [15].

Start with whether the answer set can be written down before the answer arrives, as a yes/no, a list or a scale [4]. Then whether you already hold outcomes you can score a confidence number against. With both, this is a procurement conversation, and the figure to argue over is cost per workflow against reviewer hours saved. With the answer set but no outcomes, you are buying a confidence score you have to take on trust, and the first project is the label set. Without an answer set, a prose model is still the right component, because a typed answer cannot carry the question.

What to watch

  • Whether TypeSafe opens Jev beyond the early-access waitlist and publishes calibration data buyers can reproduce.
  • Any independent timing or cost benchmark of Jev against the frontier models TypeSafe compared itself to.
  • Whether TypeSafe or DCVC confirms the $200 million valuation Forbes attributed to a person familiar with the transaction.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories