Skip to content

Product1 publisher3 min readPublished

Harvey post-trained a 27B open-weight model into the frontier band on its own legal benchmark

The band it joined is one where the best closed models still finish under 10% of tasks end to end at roughly $50 and 20 minutes each, so what post-training buys a law firm is price and hosting rather than capability.

The Product Desk · Product desk

Illustration accompanying Harvey post-trained a 27B open-weight model into the frontier band on its own legal benchmark

What happened

  • Harvey says even the strongest closed frontier models finish fewer than 10% of Legal Agent Benchmark tasks end to end, on a benchmark where a task counts only if every rubric criterion passes.
  • Reaching the top of LAB's closed-source leaderboard costs roughly $50 per task and over 20 minutes of latency, and Harvey argues quality only counts if it fits customers' cost and latency budgets.
  • With Baseten Research, Harvey post-trained a 27B open-weight model using LAB signal and its own legal-agent harness in the loop, and reports the model landing in the closed-source frontier band.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost At roughly $50 an attempt, the expensive part of a legal-agent pilot is the retries nobody budgeted for, and that bill lands on the partner who approved the pilot rather than on the vendor.
  • decision All-pass grading pushes a firm to scope agents to work a human finishes anyway, because LAB scores a nearly complete work product at zero while a reviewing lawyer would keep most of it.
  • capability Hosting the weights inside the firm turns the audit surface into reasoning traces, tool calls and intermediate decisions, which is the governance argument Harvey is making for the open-weight route.
  • exposure Anyone quoting 'frontier band' in a procurement memo is quoting the company that wrote and grades the benchmark, with no third-party run of the same tasks to lean on.

A LAB task is shaped like the week an associate dreads: a partner-style instruction, a closed universe of matter documents, and a work product someone will read line by line against a rubric an expert wrote [5]. Grading is all-pass, so a task counts only when every criterion passes [5]. That is why the under-10% end-to-end figure [2] and the 63.0% criterion pass rate from the training run [10] are not in tension. The first measures whether the work would have shipped unedited; the second is closer to how much of it a reviewing lawyer would keep [17].

The cost line is what should shape a pilot. Roughly $50 a task and over 20 minutes of latency at the top of the closed-source leaderboard [3] means running LAB's own taxonomy of more than 1,200 tasks across 24 practice areas [4] once would cost about $60,000 [14] and take 400 hours if the tasks ran one after another [15]. No firm runs the benchmark, but the unit economics carry over to any workflow where the agent needs several attempts before a human keeps one.

Before any training, the hold-out baselines already complicate the buy decision [18]. Sonnet 4.6, Opus 4.7 and GPT 5.5 lead, with GLM 5.1 and DeepSeek v4 mixed in among them and the other popular open-weight models below [8]. Two open-weight models sitting inside the leading group before Harvey touched anything is the strongest part of its case, and it is a case about the price floor rather than the 10% ceiling [2].

Harvey's account of where frontier quality comes from is about reading habits: Opus, Sonnet and GPT-5.5 opened most of the data room and read documents in full, often near 90% coverage, while weaker open-weight models used grep to surface passages and rarely read anything end to end [9]. The GRPO run on Qwen3.5-9B tested whether that habit can be learned from outcome reward alone, with reward coming only from the rubric and no retrieval-specific signal [10]. Criterion pass rate moved 20.5 points [16], though Harvey does not say in this post whether document coverage moved with it [10].

The conclusion Harvey draws is that post-training and harness optimization have to be developed together [11]. For a buyer that is the fine print: the gain belongs to a model and a harness as a pair, not to a 27B checkpoint anyone can lift into their own stack. The next-step list asks for richer, less lossy compaction and for compaction artifacts to anchor partial credit in long-horizon RL [12], which is a plain way of saying the long runs currently drop information they need and the reward signal is too coarse to see where.

Two questions sort the candidate workflows: whether a human reviews the output before it leaves the firm, and whether the cost of one attempt fits inside what the matter can absorb. Reviewed and cheap is where legal agents already earn their keep. Reviewed and expensive is the $50 quadrant [3], and it is the only one where post-training changes the answer, since price and hosting the weights in your own environment are what it is meant to buy [7]. Unreviewed and cheap runs into the under-10% number [2], and unreviewed and expensive has no case.

Harvey has not published the post-trained model's cost per task or its latency on LAB, so the frontier-band result is doing the work a cost number should do, and the score comes from the company that wrote the benchmark [19]. Until that price appears, the quadrant boundaries sit where they already sat.

What to watch

  • Whether the hold-out set and grading scripts are released so someone outside Harvey and Baseten can score the same tasks.
  • Whether the post-trained 27B model is actually shipped into customer-hosted environments rather than staying an internal result.
  • Whether the promised compaction work lifts end-to-end all-pass rates, or only criterion pass rates.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories