Skip to content

Science1 publisher2 min readPublished

A model trained on pathologists' pan-and-zoom logs caught every cancer-positive slide

A July Nature paper tested whether a general-purpose vision-language model reads a billion-pixel slide better when it searches the way a pathologist does, panning at low magnification before zooming in. The test set was lymph node tissue.

The Scientist · Science desk

Illustration accompanying A model trained on pathologists' pan-and-zoom logs caught every cancer-positive slide

What happened

  • A University of Pennsylvania team logged where eight pathologists panned and zoomed on digital slides, then filtered the raw traces down to the moments that looked like deliberate attention.
  • On human-labeled lymph node slides from colorectal cancer cases, Pathology-o3 flagged every cancer-positive slide, and 15.5% of the slides it called positive turned out to be negative.
  • OpenAI's o3, given the same slides, identified 87.5% of the positive ones, and 53.3% of the slides it called positive were negative.
  • Huang said the study was not built to beat specialized models trained disease by disease, but to test whether a general-purpose model can be taught to navigate a slide.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • decision No missed positives and roughly one wrong call in six positives fits a triage step ahead of a pathologist. A lab that adopts it re-reads about one flagged slide in six for nothing.
  • constraint The evidence covers one question, metastasis in lymph node tissue from colorectal cases, so it does not establish that behavior training helps where the diagnostic signal is spread across the whole slide.
  • capability Screen logs of expert navigation become a supervision source. That puts a training signal within reach of any imaging specialty whose experts already search before they judge.

Expect the 100 percent to travel furthest, and it is the figure that needs a denominator. The Live Science account does not state how many slides were in the test set [22]. Sensitivity is a fraction of the positive slides only, so on a set holding twenty positives, one miss would have printed as 95 percent [23]. The head-to-head with OpenAI's o3 carries more information, because both models worked the same labeled slides [14].

Read the false-alarm shares as precision instead. Of the slides Pathology-o3 called positive, 84.5 percent really were positive [18]. For o3, 46.7 percent were [19]. The share of positive calls that were wrong ran about 3.4 times higher for o3 [21], and o3 also missed 12.5 percent of the cancer-positive slides outright [20].

Strip it back and the design question is where to spend magnification. A whole slide can hold billions of pixels while the evidence of cancer occupies a tiny patch [3]. Splitting that slide into fixed-size patches treats every square of tissue as equally worth reading [1]. Pathology-o3 makes a low-resolution pass first, lets the behavior-trained model pick regions, and only then reads those at higher magnification [12]. Huang compared it to a search-and-rescue helicopter. "You don't start by inspecting one square meter of ground," he told Live Science [6].

I would most want to see the training data replicated. Most pathology systems learn from what a pathologist leaves behind at the end of a search, a labeled image or a diagnosis [24]. Huang's team recorded the search itself, and since raw logs include drift, overshoot and fiddling with magnification [8], the filtered output needed a check: they compared the regions their filter kept against eye-tracking data to confirm the software was capturing where the pathologists actually looked [9].

Pathologists stay in the loop on the language side as well. For each inspected region the model drafted a short rationale for why it was worth examining, and pathologists accepted, edited or rejected each draft; those judgements became further training data [10]. In one case the model called a region potentially metastatic and proposed zooming in to look for atypical cells [11].

The test also carries a structural limit. The slides were labeled by human pathologists before the test [14], so what was measured is agreement with a diagnosis already completed, not a change to one in progress.

The false positive share came out of a tuning choice. The researchers built the system to err toward flagging a region for another look instead of risking a miss, and Huang said that might explain the rate [17].

What to watch

  • Per-slide inference cost: whether a low-resolution scan plus selected high-magnification crops runs cheaper than exhaustive fixed-size patching, and at what magnification budget.
  • A prospective test on tissue other than lymph nodes, reporting the slide count and the number of positives.
  • Whether the 15.5% false positive share holds on slides from other labs and scanners, and with reviewers whose behavior did not train the model.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories