Skip to content

Science1 publisher2 min readPublished

An LLM pipeline pulled nearly 9,000 concrete records out of the literature in under an hour

A pipeline described in npj Computational Materials screened more than 27,000 concrete papers and reports an extraction F1 as high as 0.98 across several models. Rice University has a provisional patent pending on the method.

The Scientist · Science desk

Illustration accompanying An LLM pipeline pulled nearly 9,000 concrete records out of the literature in under an hour

What happened

  • A team writing in npj Computational Materials describes a large language model pipeline that extracts and structures materials data from unstructured papers, tested on concrete.
  • Within one hour the pipeline pulled nearly 9,000 records covering more than 100 attributes from a corpus screened out of more than 27,000 publications.
  • Reported extraction accuracy reaches an F1 of 0.98 for composition, process and property attributes, and the authors say performance held across a broad range of models.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • capability The reported performance holds across several models, so a group can run this on whichever model it already has access to. The recurring expense is inference tokens on a fixed corpus, with no training run to fund.
  • cost Model access for the work came through the OpenAI Researcher Access Program. The one hour of wall-clock time is not a quoted price for a team repeating a 27,000-paper screen on its own account.
  • constraint The paper is free to read under a licence that bars sharing adapted material, and the pipeline itself sits under a pending Rice University application. Reuse outside academia is a legal negotiation before it is an engineering task.
  • decision If the pipeline ports to other materials domains as the authors claim, an R&D group's hard question is which hundred attributes are worth defining.

Nearly 9,000 records out of a starting pool of more than 27,000 publications comes to roughly one record for every three papers screened [1]. The screen is a filter on what counts, since the schema is built around papers that report a composition, a process and a measured property together [3]. The authors used concrete because, in their words, it is "a representative and particularly challenging example" [2].

The throughput is larger than the record count suggests. More than 100 attributes across nearly 9,000 records is on the order of 900,000 individual field values. An hour holds 3,600 seconds, so the pipeline is placing roughly 250 values a second [2]. The abstract gives 0.98 as an upper bound across models and attributes; the low end and the size of the human-labelled set it was scored against go unreported [3].

An extraction F1 scores agreement with a human reading of the same passage. It tells you the pipeline copied the paper faithfully. It leaves open whether the paper's own number was sound, and whether two labs reporting the same attribute measured it the same way.

The authors report that machine learning analyses "underscore the importance of large, diverse, and information-rich datasets for enhancing both in-distribution accuracy and out-of-distribution generalization to unseen materials" [6]. More rows help, and rows unlike the ones already in hand help more, on their analysis [6]. The result is described as the largest open laboratory database for blended cement concrete [5].

The paper is open access under a CC BY-NC-ND 4.0 licence, which permits non-commercial use and sharing with credit and bars sharing adapted material [9]. All authors are inventors on U.S. provisional application No. 64/041,318, filed by Rice University on April 16, 2026, covering the extraction pipeline [8]. A company that wants to point the same method at its own internal reports is starting with a licensing question.

The portable part is the model-agnosticism. Performance held "across a broad range of LLMs" [3]. For a small group, that portability matters more than any single score, and the authors say the pipeline is readily adaptable to other materials domains [7]. The paper's own framing is that data-driven materials discovery remains constrained by the scarcity of large, high-quality, accessible experimental datasets [13]. The work was supported by Rice's civil and environmental engineering department, a Rice Academy Postdoctoral Fellowship and a Gulf Research Program Early-Career Research Fellowship [12].

What to watch

  • Whether another group can reproduce the 0.98 F1 on a held-out corpus it labelled itself.
  • What Rice University does with provisional application No. 64/041,318, since licence terms decide whether the method is reusable commercially.
  • Whether the same pipeline reaches comparable F1 in a domain whose papers report attributes less consistently than blended cement work does.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories