Science1 distinct publisher3 min readUpdated
A team led by Ho Won Jang used language models to pull 1,202 property records out of published work, screened about 150 million candidate compositions down to 37, and synthesized two.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Seoul National University's College of Engineering says a research team led by professor Ho Won Jang, of the Department of Materials Science and Engineering, has designed lead-free dielectric materials by combining data extracted from scientific literature with physics-informed machine learning, then synthesized two of the resulting compositions and measured them [1][2]. The work, published in Nature Communications, is interesting less for the two samples than for the ordering: the expensive step came last, after 448 already-published papers had been read by machine [3][4].
The target class matters. Dielectrics are insulators that block current while storing charge, and they are the core material in the multilayer ceramic capacitors used in smartphones, electric vehicles and other electronics [5]. A higher dielectric constant means more energy stored in the same volume [6], but a part is only useful if that performance holds when the device gets hot [7]. Relaxor ferroelectrics, whose electrical response changes relatively gradually with temperature, are the promising family because they can pair a high dielectric constant with stability across a broad range [8]. Even restricted to lead-free chemistries, the candidate count is effectively unbounded once you allow arbitrary element combinations and mixing ratios [9].
The obstacle is not that the data does not exist but that it is unusable as published. Relevant numbers sit in text, tables and graphs across different papers, and measurement conditions such as temperature, frequency and sample characteristics vary between studies, so nothing can be fed directly to a model [10]. The team used large language models to organize composition and processing conditions out of text and tables, and converted plotted curves into numerical data to recover temperature-dependent properties [11]. That produced 1,202 records covering composition, processing and dielectric properties from 448 papers [4], which works out to roughly 2.7 usable records per paper [1]. That is a thin yield per document, and still far cheaper than 1,202 syntheses.
To make the records comparable, the researchers added 22 physical descriptors including elemental composition and microstructure [12], then combined 30 independently trained models to predict three indicators tied to dielectric constant and temperature stability, using agreement between the models as a confidence ranking [13][14]. Applied to a virtual space of about 150 million compositions, the funnel returned 37 candidates [15], a reduction of roughly four million to one [2]. Two of those, about 5 percent of the shortlist [3], were made in the lab, and the team reports that both showed high dielectric constants and good high-temperature stability [16].
Read the validation narrowly. Two syntheses out of 37 shortlisted compositions confirm that the pipeline can surface something real; they do not establish a hit rate, and the university's account describes the measured performance only qualitatively, without reported values for dielectric constant or the temperature window [16][17]. The argument for literature mining as a first step does not depend on those numbers. It depends on the cost asymmetry: 448 papers of prior work, restructured once, replaced the front end of a search that would otherwise have been trial and error [18].
What to watch: whether the remaining 35 candidates get made and how many survive; whether the extraction and physics-screening stack is published in reusable form or stays in-house; and whether the same 22-descriptor unification holds for property classes where papers report even less about processing conditions.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Seoul National University College of Engineering announced that a research team led by professor Ho Won Jang of the Department of Materials Science and Engineering developed a technology for designing lead-free dielectric materials by combining data extracted from scientific literature with physics-informed machine learning.
The researchers synthesized two of the shortlisted compositions and experimentally confirmed both high dielectric constants and excellent high-temperature stability.
The findings are published in the journal Nature Communications.
The process yielded a dataset of 1,202 records covering composition, processing conditions and dielectric properties from 448 papers.
Dielectrics are insulating materials that prevent electricity from flowing directly while storing electric charge, and are key materials in multilayer ceramic capacitors (MLCCs) used in smartphones, electric vehicles and other electronic devices.
The higher the dielectric constant, the more electrical energy a component of the same size can store.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed and quantified, but single-sourced and unreproducible from what is supplied
The account is unusually specific for a press relay: exact dataset counts (1,202 records, 448 papers), pipeline parameters (22 descriptors, 30 models, three targets), a stated screening funnel (~150 million to 37), measured permittivities of 3,422 and 3,307, and compliance with named industry temperature classes benchmarked against barium titanate. It also sits on a Nature Communications publication. Against that: one publisher, a university announcement as the origin, no DOI or paper metadata, no model accuracy or extraction error rates, no data or code release statement, no independent replication, and a body that is truncated mid-sentence.
No adoption evidence beyond two laboratory samples
Nothing in the supplied material shows anyone using this outside the originating lab: no manufacturer trial, no licensing, no partnership, no dataset or model release, no third-party replication, and 35 of the 37 shortlisted compositions were never synthesized. Meeting X5R/X6R/X7R on two lab pellets is technical validation, not uptake, and inferring commercial traction from it would be guesswork.
Mildly overstated framing over a genuinely quantified lab result
The headline and lede sell speed and a transformation of materials discovery — 'could transform materials discovery from a trial-and-error process into a data-driven one' — on the strength of two synthesized samples out of a 37-item shortlist, with no time-to-discovery comparison against a conventional baseline to support the speed claim. That is a modest overstatement rather than a large one, because the underlying numbers are concrete and measured against external MLCC standards and an incumbent material. Working in the other direction, the cluster's own derived reading understates the result by treating the outcome as unquantified when the body reports 3,422 and 3,307 and the X5R/X6R/X7R temperature windows; that internal understatement partly offsets the promotional framing.
Institutional announcement relayed largely intact
The chain is a Seoul National University College of Engineering announcement, naming the lead professor, the first author and co-workers, carried by an aggregator whose research-news items typically restate institutional releases. That gives a clear promotional interest in emphasizing novelty, the 150-million-to-37 funnel and the 'transform discovery' framing, and in omitting failure modes, model error rates and manufacturability. Two things restrain the score rather than raise it: the underlying work is peer-reviewed in Nature Communications, and the relay still carries hard numbers and a comparison against an incumbent material, which promotional copy often drops. No commercial sponsor, vendor or funding conflict is disclosed in the supplied material.
Moderate: specific and peer-reviewed, but one publisher and one origin
The core factual claims — dataset size, pipeline parameters, screening funnel, measured permittivities, standards compliance, publication venue — are internally consistent and stated precisely, and the derived arithmetic follows directly from them. Confidence is capped by structural limits rather than internal contradictions: a single publisher relaying a single institutional origin, no primary-paper identifiers, no model validation metrics, no independent corroboration, a truncated body, and one derived cluster claim that the source text itself contradicts.
science
Complete rye centromeres turn a wheat breeder's theoretical gene pool into a working one1 distinct publisher
science
The safety catch on DNA replication now has a structure, and a mutation that breaks it1 distinct publisher
science
What you expect from your own old age shows up a decade later in who you still see1 distinct publisher
science
Narwhal tusks hide two spirals twisting against each other, and the mismatch is the point3 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 20, 2026