Skip to content

Science1 publisher3 min readPublished

Seoul National University mined 448 papers, then made two heat-stable lead-free dielectrics

A team led by Ho Won Jang used language models to pull 1,202 property records out of published work, screened about 150 million candidate compositions down to 37, and synthesized two.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Seoul National University mined 448 papers, then made two heat-stable lead-free dielectrics
Photo: nature.com

What happened

  • Seoul National University College of Engineering announced that a research team led by professor Ho Won Jang of the Department of Materials Science and Engineering developed a technology for designing lead-free dielectric materials by combining data extracted from scientific literature with physics-informed machine learning.
  • The researchers synthesized two of the shortlisted compositions and experimentally confirmed both high dielectric constants and excellent high-temperature stability.
  • The findings are published in the journal Nature Communications.
  • The process yielded a dataset of 1,202 records covering composition, processing conditions and dielectric properties from 448 papers.
  • Dielectrics are insulating materials that prevent electricity from flowing directly while storing electric charge, and are key materials in multilayer ceramic capacitors (MLCCs) used in smartphones, electric vehicles and other electronic devices.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

Seoul National University's College of Engineering says a research team led by professor Ho Won Jang, of the Department of Materials Science and Engineering, has designed lead-free dielectric materials by combining data extracted from scientific literature with physics-informed machine learning, then synthesized two of the resulting compositions and measured them [1][2]. The work, published in Nature Communications, is interesting less for the two samples than for the ordering: the expensive step came last, after 448 already-published papers had been read by machine [3][4].

The target class matters. Dielectrics are insulators that block current while storing charge, and they are the core material in the multilayer ceramic capacitors used in smartphones, electric vehicles and other electronics [5]. A higher dielectric constant means more energy stored in the same volume [6], but a part is only useful if that performance holds when the device gets hot [7]. Relaxor ferroelectrics, whose electrical response changes relatively gradually with temperature, are the promising family because they can pair a high dielectric constant with stability across a broad range [8]. Even restricted to lead-free chemistries, the candidate count is effectively unbounded once you allow arbitrary element combinations and mixing ratios [9].

The obstacle is not that the data does not exist but that it is unusable as published. Relevant numbers sit in text, tables and graphs across different papers, and measurement conditions such as temperature, frequency and sample characteristics vary between studies, so nothing can be fed directly to a model [10]. The team used large language models to organize composition and processing conditions out of text and tables, and converted plotted curves into numerical data to recover temperature-dependent properties [11]. That produced 1,202 records covering composition, processing and dielectric properties from 448 papers [4], which works out to roughly 2.7 usable records per paper [1]. That is a thin yield per document, and still far cheaper than 1,202 syntheses.

To make the records comparable, the researchers added 22 physical descriptors including elemental composition and microstructure [12], then combined 30 independently trained models to predict three indicators tied to dielectric constant and temperature stability, using agreement between the models as a confidence ranking [13][14]. Applied to a virtual space of about 150 million compositions, the funnel returned 37 candidates [15], a reduction of roughly four million to one [2]. Two of those, about 5 percent of the shortlist [3], were made in the lab, and the team reports that both showed high dielectric constants and good high-temperature stability [16].

Read the validation narrowly. Two syntheses out of 37 shortlisted compositions confirm that the pipeline can surface something real; they do not establish a hit rate, and the university's account describes the measured performance only qualitatively, without reported values for dielectric constant or the temperature window [16][17]. The argument for literature mining as a first step does not depend on those numbers. It depends on the cost asymmetry: 448 papers of prior work, restructured once, replaced the front end of a search that would otherwise have been trial and error [18].

What to watch: whether the remaining 35 candidates get made and how many survive; whether the extraction and physics-screening stack is published in reusable form or stays in-house; and whether the same 22-descriptor unification holds for property classes where papers report even less about processing conditions.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories