Invest1 publisher2 min readPublished
Nearly half of one newsletter's 25 product openings ask for eval experience
Lenny's newsletter counted eval experience in nearly half of 25 product manager postings it listed, and the two practitioners it commissioned say most teams skip error discovery and write metrics for failures they never found.
The Investor · Invest desk

What happened
- Hamel and Shreya, writing from work with over 50 AI companies, say most teams jump straight to writing metrics and end up measuring the wrong things.
- Their process begins with discovering and analyzing errors, the stage they say most teams skip before moving to customized metrics and a continuous improvement loop.
- Among the post's examples, Ramp raised the accuracy of its automatic receipt collection product from 35% to 83% after investing in evals.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- contradiction One post holds both that writing evals is the most important thing to teach product people and that a free plugin plus a coding agent does the critical stage in about half an hour, so the buyer of the skill and the buyer of the tool are being quoted different prices for the same work.
- cost The priced item here is the teaching: the plugin costs nothing and the course carries a 25% code, so a team's real outlay is the staff hours spent looking at session records.
- exposure Anyone sizing a return on eval spending is working from gains self-reported by companies inside the authors' own client base, with no control group for the model upgrades that landed over the same period.
- decision Hiring managers drafting the next product req now have to decide whether eval experience screens candidates out or is a half-day of in-house training they can cover after the offer.
Ramp's automatic receipt collection went from 35% accuracy to 83% after the company invested in evals [6], so the miss rate fell from 65 in every 100 receipts to 17, about three-quarters fewer failures [14]. Shopify reported an AI workflow builder 2.2 times faster and 68% cheaper than the frontier-model system it replaced [7], and a 68% cut leaves the replacement running at 32 cents on the dollar [15]. Cursor reported 41% lower costs on its Auto Balance routing alongside higher user satisfaction [8], and Harvey nearly doubled an internal quality score on a rebuilt contract reviewer [9].
Those are company self-reports, collected in a post whose authors sell the training. The sample of more than 50 AI companies is Hamel and Shreya's own client work [2], the workflow ships with a free plugin [12], and the post closes with a 25% discount code for their course [11]. None of those companies said what the eval work cost to do.
The hiring evidence is lighter than the product numbers. Lenny wrote that nearly half of the 25 openings he shared on socials last week ask for experience writing evals [1], which is about 12 postings [13], drawn from a list curated by a newsletter that markets an evals course [11].
Mike Krieger, Anthropic's former CPO and now head of Labs, said that "if there's one thing we can teach product people, it's that writing evals is now probably the most important thing" [4]. Garry Tan, the CEO of Y Combinator, said that "evals are emerging as the real moat for AI startups" [5].
Set that against the workflow in the same post: three steps, a coding agent such as Codex or Claude, about 30 minutes once you know the basics, run with more than 50 companies and, the authors say, turning up major product flaws every time [10]. Evals exist because a prompt, model, or code change can improve one behavior while breaking another [18]. If the discovery work compresses to half an hour of agent time, a req asking for eval experience is not paying for the labour of reading session records. It is paying for a judgement about which failures count, and in my view that judgement is what the product job already was.
The way to know I have that wrong is compensation. If eval work splits out into its own titles with their own bands, it is a headcount category, and the skill is as scarce as the quotes claim.
What to watch
- A posting sample larger than one newsletter's 25, showing whether eval experience is asked for outside curated lists.
- Whether Shopify, Cursor, Ramp or Harvey publish the baseline and the cost of the eval work behind the reported gains.
- Whether the next round of model price cuts absorbs Cursor's 41% routing saving.