8.4 Export first-principles labels for ML workflows
8.4.1 A training set is a scientific measurement record
An ABACUS trajectory is not a trustworthy machine-learning dataset merely because a converter accepts it. Each frame needs a compatible energy, structure, atom ordering and any force or virial labels used in training. Failed SCF steps, different electronic states, mixed pseudopotentials and correlated train/test samples can create misleading model accuracy.
This unexecuted data-audit template starts from the ABACUS 3.9.0 trajectory-analysis case. It uses the current official dpdata ABACUS plugin API, whose version must be recorded independently from ABACUS. The plugin supports abacus/scf and abacus/md calculation directories, but its evolving parser is not assumed to interpret every older or newer output identically. Pin the installed dpdata version and validate one frame manually before bulk export.
No trajectory has been simulated or model trained here. The Python example is a proposed local conversion and audit, not an executed benchmark. Real source frames must exist, satisfy your reference-quality policy and include full calculation logs. A directory with STRU alone is unlabeled structure data; a System is not automatically a LabeledSystem.
8.4.2 Unit and identity contracts
The standard DeePMD data contract uses coordinates and cells in Å, energies and virials in eV, and forces in eV/Å. The virial is a nine-component energy-valued tensor, not a pressure printed in kbar. Converting a stress requires the chosen sign convention, cell volume and unit conversion; do not assume that copying nine numbers preserves its meaning.
For conservative reference forces,
This finite-difference check should be small only after the displacement and electronic thresholds are themselves converged. It detects sign or labeling mistakes but does not prove the reference functional is physically exact. Use small representative frame sets first; do not pay for a massive collection before checking the contract.
The arrows and split diagram are conceptual; there are no calculated forces or measured model errors in the illustration.
8.4.3 ABACUS label-output delta
Keep the MD parent's species assets, order, functional, spin treatment and numerical precision. Request the outputs needed by the chosen dataset:
# ABACUS 3.9.0 MD parent delta; no job executed here
cal_force 1
cal_stress 1
md_dumpfreq 1
dump_force 1
dump_virial 1
These entries are conditional on the documented basis/method supporting forces and stress. They do not retroactively generate labels for an old trajectory lacking them. Force/stress quality must be independently converged, including Pulay-related basis effects where relevant. For snapshots from MD, a dedicated static relabeling step at stricter electronic accuracy can be necessary; document whether labels came from online MD or a separate SCF workflow.
Never add forces from one step to coordinates from another. Verify the parser's frame mapping against MD_dump and running_md.log, checking step identifiers, output cadence and any missing unconverged steps. For static SCF data, the official current parser expects the calculation directory containing INPUT, STRU and its OUT directory, not merely a path to a log file.
8.4.4 Proposed dpdata audit and export
The following Python is a not-executed template; replace paths, check your pinned parser and reject failed reference frames before export. The explicit Si type map is appropriate only for the single-species silicon teaching trajectory.
import dpdata
import numpy as np
system = dpdata.LabeledSystem("accepted_md", fmt="abacus/md")
data = system.data
nframe = len(system)
natom = system.get_natoms()
assert nframe > 0 and natom > 0
assert data["coords"].shape == (nframe, natom, 3)
assert data["cells"].shape == (nframe, 3, 3)
assert data["energies"].shape == (nframe,)
assert data["forces"].shape == (nframe, natom, 3)
assert data["atom_names"] == ["Si"]
for name in ("coords", "cells", "energies", "forces"):
assert np.isfinite(data[name]).all(), name
assert np.all(np.abs(np.linalg.det(data["cells"])) > 1e-8)
system.to("deepmd/npy", "audited_si_export")
These shape and finiteness checks are necessary but not sufficient. They do not certify SCF convergence, correct atom mapping or physical units. Manually compare one exported frame's coordinates, energy and every force component to its original ABACUS record. Verify virial units and sign separately before including that label in training. Hash both source files and exported arrays and preserve the parser version, frame-selection list and any exclusions.
For multicomponent systems, establish one explicit element/type map across all dataset groups. Do not rely on alphabetical reordering being harmless: atom type indices, coordinates and force arrays must remain aligned. Separate systems with incompatible atom counts or periodicity into the format grouping expected by the selected DeepMD version.
8.4.5 Blank quality and split record
| Source group | Frame IDs / cadence | Method and asset hashes | SCF accepted count | Units checked? | Force/virial checks | Split group | Exclusions and reason |
|---|---|---|---|---|---|---|---|
| trajectory A | train | ||||||
| independent trajectory B | validation | ||||||
| held-out phase / conditions | test |
Split by independent trajectories, structures or thermodynamic conditions before training, not by randomly shuffling adjacent frames. Reserve a true final test group. Report per-domain errors and failure modes in addition to pooled RMSE. A low average can hide rare large force errors in short interatomic distances, surfaces or phase-changing configurations. Duplicate detection and composition-aware energy comparisons prevent false performance gains.
8.4.6 Pitfalls, exercises and answers
Pitfalls. A converter can parse an unconverged output without knowing your acceptance threshold. A pressure tensor mistaken for eV virial changes label scale and sign. Different functionals in one pool create incompatible target surfaces unless the learning task explicitly models that distinction. Random adjacent-frame splitting leaks correlated information. Rejecting difficult frames without documenting them narrows the applicability domain silently. Treating every low model-error frame as reference-quality overlooks systematic DFT bias.
Exercise 1. All array shapes pass, but forces appear larger than the handwritten reference by a constant factor. What should you test? Answer: source and target units, conversion factors, atom order and whether the parser selected the intended step; check a finite-difference force using consistently converted energies and lengths.
Exercise 2. A random split gives small test RMSE, while an independently sampled trajectory performs poorly. Which is more informative for deployment? Answer: the independent group exposes the intended generalization problem. Keep grouped evaluation and report the domain failure instead of tuning against the held-out test repeatedly.
8.4.7 Sources and course handoff
- Official dpdata ABACUS plugin source: accepted formats and directory inputs; pin its installed version.
- DeePMD system data contract: label shapes and units.
- ABACUS 3.9.0 MD output controls.
- University-linked ABACUS hands-on guide: wider workflow context; do not assume historical examples share current defaults.
Return to the course hub with an auditable workflow: trustworthy electronic references first, documented export second, independent model validation last.