Science1 publisher3 min readPublished
Skala reaches CP2K, and computational chemists inherit a versioning problem
Microsoft's learned exchange-correlation functional is now in one production code with four more integrations underway. The accuracy claim is the easy part; workflow policy is not.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Microsoft Research announced that Skala is available in CP2K and is being integrated into Psi4, FHI-aims, ORCA and VASP, bringing it closer to the communities that use those codes.
- Skala is Microsoft Research's deep-learning exchange-correlation functional for density functional theory.
- Skala-1.1 was trained on 2.5 times more data than the first public version of Skala.
- Skala-1.1 achieves a weighted average error of 2.8 kcal/mol on the GMTKN55 benchmark suite.
- GMTKN55 is a widely used benchmark suite comprising 55 categories of chemistry, including thermochemistry, reaction barriers and noncovalent interactions.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
Microsoft Research says Skala, its deep-learning exchange-correlation functional, is now available in CP2K and is being integrated into Psi4, FHI-aims, ORCA and VASP [1][2]. That is the consequential part of the announcement: a learned functional stops being a paper with a checkpoint attached and becomes a keyword someone can type into an input file on a Tuesday.
The accuracy figures are the part everyone will quote. Microsoft reports that Skala-1.1, trained on 2.5 times more data than the first public release, reaches a weighted average error of 2.8 kcal/mol on GMTKN55 [3][4], the 55-category suite covering thermochemistry, reaction barriers and noncovalent interactions [5]. The company says that surpasses today's leading global and range-separated hybrid functionals while keeping the cost profile of a semi-local functional [6], and that the model also produces accurate electron densities, dipole moments and geometries [10]. The training data comes from expansions of Microsoft's own reference collection, MSR-ACC, built with expensive wavefunction methods, with electron affinities among the newly added categories [11].
Take those numbers at Microsoft's word for now and the interesting question is still operational. Skala is explicitly designed so that each release supersedes the last, in deliberate contrast to the accumulating "functional zoo" of conventional DFT [7]. That is defensible engineering and a real break with how the field has worked. It also means the functional name in a methods section no longer identifies a calculation. B3LYP in 2013 and B3LYP in 2026 are the same object; Skala-1.0 and Skala-1.1 are not, since the second was trained on a different and larger dataset [3][7]. Any group adopting this needs a version pin, a stored checkpoint, and a policy on whether a paper under review gets rerun when the model updates. Groups that regress against published internal baselines will find those baselines moving underneath them.
The second operational wrinkle is cost, and Microsoft has been unusually candid about it by shipping a living benchmark that tracks the computational performance of successive, increasingly optimised Skala releases across software packages and hardware platforms [9]. Publishing that reference is an admission that the same functional will not cost the same in CP2K as in ORCA, or on one accelerator versus another. The claim that improvements arrive at unchanged practical cost [8] is a claim about the model, not about any particular implementation. Until the benchmark has entries, "efficiency of a semi-local functional" [6] is an architectural statement, not a wall-clock promise for your cluster.
One scope note worth keeping in mind: the aggregate accuracy figure in this announcement is a molecular one, on GMTKN55, and no equivalent aggregate number is given for periodic or solid-state benchmarks [13], even though the integration list includes codes used well outside main-group molecular chemistry [1].
What to watch: whether the Psi4, FHI-aims, ORCA and VASP integrations land in shipped releases rather than branches [1]; whether the living benchmark publishes per-code, per-hardware timings that a group can compare against its current hybrid workflow [9]; and how the first Skala update after 1.1 is handled by anyone who has already published with it [7].