These materials span different chemistries, crystal structures, and local Li-ion environments, enabling a comprehensive assessment of the generality of our observations.
For LLZO, we utilized the PBEsol dataset from our previous work18, which contains 1,978 LLZO configurations with 56 Li-ions per structure, totaling 110,768 local Li-ion environments.
This robust performance demonstrates that fewer than 1000 local Li-ion environments are sufficient to capture the essential interatomic interactions of LLZO.
The force RMSEs for LGPS show a slight upward trend as the number of Li-ion environments approaches this threshold; however, the impact on the predicted transport properties remains within acceptable ranges.
Across all three chemically diverse SSE families (oxide, halide, and sulfide), ~1000 Li-ion local environments appear sufficient to train a well-performing force field under the NEP framework.
Conventionally, SSEs are regarded as multi-component materials with relatively complex configurations due to the mobile ion sublattice. Therefore, to ensure adequate sampling of the PES, it is widely assumed that a large training dataset, typically consisting of at least a few thousand configurations, is required to develop an accurate MLFF for an SSE material. For instance, studies on the \(P\bar{3}m1\) phase of Li 3 YCl 6 utilized about 1700–1850 structures1922, while similar benchmarks across various ternary halides and sulfides consistently employed 1800 configurations23. The required dataset size increases dramatically when studying complex phenomena. For example, investigations into the crystallization of Li 3 PS 4 glass and Li 2 ZrCl 6 transport mechanisms have relied on much larger datasets, containing 34,510 and 60,000 structures, respectively2425. It appears that datasets containing thousands to tens of thousands of structures have become the “standard” for simulating complex ion transport and phase behavior in SSEs. However, several tests presented in this Perspective suggest a different picture: the PES of crystalline SSEs is relatively easy to sample, and therefore can be effectively captured with a small amount of training data.
In fact, the inherent structural characteristics of SSEs determine this ease of sampling. While the mobile ion sublattice renders SSEs more complex than typical rigid solids, their PESs remain significantly simpler than those of liquid, molecular, or polymer systems. In liquid or polymer systems, atoms can explore vast, unconstrained configurational spaces. In contrast, crystalline SSEs possess a rigid framework that provides a stable structural backbone. Within this framework, host atoms only vibrate around equilibrium positions with small displacements, while the mobile ions, despite highly diffusive, are restricted to predefined pathways and specific crystallographic sites. From a sampling perspective, this means SSEs behave more like traditional solids than unconstrained liquids, allowing a relatively small number of representative structures to provide sufficient coverage of the relevant local atomic environments.
To quantify this data efficiency, we conducted systematic dataset-size sensitivity tests on three representative SSE materials: the oxide Li 7 La 3 Zr 2 O 12 (LLZO), the halide Li 3 YCl 6 (LYC), and the sulfide Li 10 GeP 2 S 12 (LGPS). These materials span different chemistries, crystal structures, and local Li-ion environments, enabling a comprehensive assessment of the generality of our observations. For LLZO, we utilized the PBEsol dataset from our previous work18, which contains 1,978 LLZO configurations with 56 Li-ions per structure, totaling 110,768 local Li-ion environments. Regarding the MLFF framework, we adopted the fourth-generation NEP model proposed by Fan et al.10, which offers a balance between simulation accuracy and speed18 and helps minimize testing costs. The NEP model achieved sub-meV/atom precision when trained on the full dataset of LLZO: the root mean square errors (RMSEs) for energy and force are 0.77 meV/atom and 83.67 meV/ Å, respectively. We then progressively reduced the dataset by factors of 1/2n (n = 1, 2, . . . ) using the farthest-point sampling algorithm implemented in our GPUMDkit package26. Each reduced dataset was used to train a new NEP model, while the full 1,978-structure dataset served as a test set to evaluate how prediction accuracy evolves as the training set size decreases.
Remarkably, even aggressive dataset reduction has a minimal impact on model accuracy (Fig. 1a, b). For the LLZO system, the RMSE for energy and force remains nearly constant as the dataset is reduced to 1/32 of its original size (i.e., n = 5, corresponding to only 62 configurations containing 3472 Li-ions). A slight upward trend in error emerges when the reduction factor reaches n = 6 and n = 7. Even at an aggressive reduction factor of 1/128 (n = 7), where only 15 LLZO structures (corresponding to 840 Li-ions) are retained, the test errors increase only modestly to 0.88 meV/atom and 101.17 meV/ Å. This robust performance demonstrates that fewer than 1000 local Li-ion environments are sufficient to capture the essential interatomic interactions of LLZO.
Fig. 1: Dataset-size sensitivity tests. Full size image Dataset-size sensitivity tests on MLFF performance in Li 7 La 3 Zr 2 O 12 (LLZO), Li 3 YCl 6 (LYC), and Li 10 GeP 2 S 12 (LGPS). The validation root mean square errors of a energy and b force. c Li-ion diffusivities (1000 K for LLZO, 500 K for LYC and LGPS). d Li-ion activation energies.
More importantly, we evaluated whether these data-efficient models could maintain an accurate description of ion transport properties, which is a fundamental requirement for MLFF in SSE research. As shown in Fig. 1c, the calculated Li-ion diffusivities at 1000 K remain consistent across all models with different training set sizes, with values around 0.88–0.98 × 10−5 cm2/s. Similarly, the activation energies (E a ) calculated from the Arrhenius plots at four temperatures (950, 1000, 1100, 1200 K) cluster closely between 0.281 and 0.299 eV (Fig. 1d), which is well within an acceptable range. The complete Arrhenius plots are provided in Supplementary Fig. S1 of Supporting Information (SI). Additional dimer dissociation tests further confirm the microscopic robustness of NEP models trained on reduced datasets (Supplementary Fig. S2 in SI). To provide an external validation of the trained force fields, we also compared NEP-based MD with AIMD under identical small-supercell conditions, finding good agreement across the three materials (Supplementary Fig. S3 in SI).
Furthermore, we performed similar dataset-size sensitivity analyses on two chemically distinct SSE materials: the halide LYC and the sulfide LGPS. For LYC, we employed the PBE+optB88-vdW dataset for the P\(\bar{3}\) m1 phase from Wang et al.19, containing 1698 structures with 18 Li-ions each, totaling 30,564 local Li-ion environments. For LGPS, we utilized part of the PBEsol dataset from Huang et al.27, containing 1799 unit-cell structures with 20 Li-ions each, totaling 35,980 local Li-ion environments. As shown in Fig. 1, both halide and sulfide electrolytes exhibit trends similar to those observed in LLZO. Specifically, LYC exhibits insensitivity to dataset reduction, with its energy and force RMSEs remaining almost unchanged even when the number of Li-ions drops below 1000. The force RMSEs for LGPS show a slight upward trend as the number of Li-ion environments approaches this threshold; however, the impact on the predicted transport properties remains within acceptable ranges. Across all three chemically diverse SSE families (oxide, halide, and sulfide), ~1000 Li-ion local environments appear sufficient to train a well-performing force field under the NEP framework. Interestingly, our tests using the MACE framework show that MACE errors are more sensitive to dataset size, with RMSEs exhibiting a more gradual increase as the training set decreases (see Supplementary Fig. S4 in SI).
We attribute this disparity to the underlying optimization algorithms. While NEP employs the separable natural evolution strategy (SNES)28, MACE relies on gradient descent-based training, which typically requires larger datasets to effectively explore the complex parameter spaces. To isolate the impact of the optimization algorithm, we performed tests using the GNEP framework29, which employs the same architecture as NEP but uses a gradient descent-based optimizer similar to MACE. As shown in Supplementary Fig. S5, GNEP also displays a more pronounced sensitivity to data sparsity compared to the standard NEP. This confirms that the superior data efficiency of the standard NEP is also attributed to its SNES algorithm.
We also tested whether similar data-efficiency trends appear in foundation-model fine-tuning scenarios. Fine-tuned NEP89 and MACE-OMAT-0-MEDIUM models show consistent diffusivity and activation-energy predictions across dataset reduction levels, even when the number of additional system-specific configurations is very small (Supplementary Figs. S6, S7 in SI).
These results challenge the conventional wisdom that accurate simulation of ion diffusion in bulk SSEs requires massive training datasets. In fact, the rigid framework and constrained diffusion paths of Li ions make the PES relatively easy to learn. Furthermore, once Li ions begin to diffuse within these channels, they naturally explore most of the relevant local environments and can effectively self-sample the configuration space even in relatively short trajectories. This observation shifts the emphasis of MLFF development for SSEs: choosing structures that can effectively cover the relevant configuration space is far more important than accumulating a large dataset. It should be noted, however, that these conclusions are drawn from defect-free bulk crystals, and systems with higher configurational complexity, such as amorphous SSEs, heavily substituted or high-entropy compositions, lithium-doped structures, or materials undergoing lithium extraction, may require further investigation.