Science1 distinct publisher3 min readUpdated
Microsoft's learned exchange-correlation functional is now in one production code with four more integrations underway. The accuracy claim is the easy part; workflow policy is not.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
Microsoft's learned exchange-correlation functional is now in one production code with four more integrations underway. The accuracy claim is the easy part; workflow policy is not.
Microsoft Research says Skala, its deep-learning exchange-correlation functional, is now available in CP2K and is being integrated into Psi4, FHI-aims, ORCA and VASP [1][2]. That is the consequential part of the announcement: a learned functional stops being a paper with a checkpoint attached and becomes a keyword someone can type into an input file on a Tuesday.
The accuracy figures are the part everyone will quote. Microsoft reports that Skala-1.1, trained on 2.5 times more data than the first public release, reaches a weighted average error of 2.8 kcal/mol on GMTKN55 [3][4], the 55-category suite covering thermochemistry, reaction barriers and noncovalent interactions [5]. The company says that surpasses today's leading global and range-separated hybrid functionals while keeping the cost profile of a semi-local functional [6], and that the model also produces accurate electron densities, dipole moments and geometries [10]. The training data comes from expansions of Microsoft's own reference collection, MSR-ACC, built with expensive wavefunction methods, with electron affinities among the newly added categories [11].
Take those numbers at Microsoft's word for now and the interesting question is still operational. Skala is explicitly designed so that each release supersedes the last, in deliberate contrast to the accumulating "functional zoo" of conventional DFT [7]. That is defensible engineering and a real break with how the field has worked. It also means the functional name in a methods section no longer identifies a calculation. B3LYP in 2013 and B3LYP in 2026 are the same object; Skala-1.0 and Skala-1.1 are not, since the second was trained on a different and larger dataset [3][7]. Any group adopting this needs a version pin, a stored checkpoint, and a policy on whether a paper under review gets rerun when the model updates. Groups that regress against published internal baselines will find those baselines moving underneath them.
The second operational wrinkle is cost, and Microsoft has been unusually candid about it by shipping a living benchmark that tracks the computational performance of successive, increasingly optimised Skala releases across software packages and hardware platforms [9]. Publishing that reference is an admission that the same functional will not cost the same in CP2K as in ORCA, or on one accelerator versus another. The claim that improvements arrive at unchanged practical cost [8] is a claim about the model, not about any particular implementation. Until the benchmark has entries, "efficiency of a semi-local functional" [6] is an architectural statement, not a wall-clock promise for your cluster.
One scope note worth keeping in mind: the aggregate accuracy figure in this announcement is a molecular one, on GMTKN55, and no equivalent aggregate number is given for periodic or solid-state benchmarks [13], even though the integration list includes codes used well outside main-group molecular chemistry [1].
What to watch: whether the Psi4, FHI-aims, ORCA and VASP integrations land in shipped releases rather than branches [1]; whether the living benchmark publishes per-code, per-hardware timings that a group can compare against its current hybrid workflow [9]; and how the first Skala update after 1.1 is handled by anyone who has already published with it [7].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Microsoft Research announced that Skala is available in CP2K and is being integrated into Psi4, FHI-aims, ORCA and VASP, bringing it closer to the communities that use those codes.
Skala is Microsoft Research's deep-learning exchange-correlation functional for density functional theory.
Skala-1.1 was trained on 2.5 times more data than the first public version of Skala.
Skala-1.1 achieves a weighted average error of 2.8 kcal/mol on the GMTKN55 benchmark suite.
GMTKN55 is a widely used benchmark suite comprising 55 categories of chemistry, including thermochemistry, reaction barriers and noncovalent interactions.
Microsoft states that Skala-1.1's accuracy surpasses today's leading global (range-separated) hybrid functionals while retaining the efficiency of a semi-local functional.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quantified but first-party only
The cluster rests on one first-party announcement. It is unusually specific for a vendor post — a named public benchmark suite, a single aggregate number (2.8 kcal/mol on GMTKN55), a stated training-data multiple, and named integration targets — which makes the claims falsifiable. But there is no independent replication, no per-category data, no named hybrid-functional comparison set, and no periodic/solid-state metric, so evidentiary strength stays well below mid-range.
One production code live, four pending
Verifiable adoption is narrow: availability in CP2K plus an earlier (GPU4)PySCF/ASE community release, with Psi4, FHI-aims, ORCA and VASP only described as in progress. No user counts, deployment disclosures, downloads, or third-party papers using Skala appear in the material, so adoption is credited for real code-level distribution but not for demonstrated field use.
Overstated, moderately
Framing runs ahead of the demonstrated result. The post generalizes to 'predictive DFT' across chemistry, materials, catalysis, energy and drug discovery and claims superiority over leading hybrid functionals, while the only aggregate metric is one molecular benchmark, self-reported, with four of five integrations unfinished and no periodic accuracy figure for the materials-oriented codes named. The gap is moderate rather than severe because the central factual claims are specific and testable, and the availability claim is bounded honestly.
Vendor announcing its own model and its own benchmark
Every claim originates from the organization that built Skala, generated the MSR-ACC training data, negotiated the code integrations, and is now also establishing the living benchmark that will measure successive releases. Owning both the artifact and the yardstick, and publishing on its own research blog with no independent voice in the cluster, is a strong incentive alignment toward favorable framing.
Facts of announcement solid, substance unverified
High confidence that these statements were made and that CP2K availability is real, since a primary-source announcement is direct evidence of its own contents. Low confidence in the comparative accuracy and ecosystem-impact substance, which depends on unpublished detail and unreplicated benchmarks from a single interested party.
science
AI weather models cannot forecast what they never saw. A hybrid method aims at that gap.1 distinct publisher
science
Pasqal's prompt-to-circuit agent still needs a physicist in the loop1 distinct publisher
invest
AMD's Korea test lab is a bet that inference stops being a GPU-only purchase1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 20, 2026