8.3 DeePKS inference and domain validation
8.3.1 Identify what the model actually changes
DeePKS is a machine-learning correction to an electronic energy functional, used self-consistently with electronic structure. It is not a DeePMD interatomic potential that replaces the electronic calculation with an atomic-energy model. The two methods can participate in a larger workflow, but their files, labels, applicability domains and runtime meaning are different.
This unexecuted, conditional exercise follows ABACUS 3.9.0 LCAO DeePKS interfaces. Ordinary LCAO SCF must already work, and the binary must have the documented DeePKS build dependencies, including a compatible LibTorch route. No training is performed, no model is downloaded automatically, and no accuracy improvement is claimed. The advanced deepks_equiv route is described by the pinned manual as internal/development use, so it is excluded here.
The mandatory scientific input is an appropriately trained, traced model with its provenance. You cannot turn the silicon parent into a meaningful DeePKS calculation by supplying an arbitrary neural-network filename. The model must document the elements, electronic baseline, target method, projector/descriptor definition, units and training domain. If no compatible model and descriptor are available, stop at this checklist; the conditional input below is not executable evidence.
8.3.2 Descriptors and self-consistent feedback
The conceptual structure is
Density-dependent local descriptors connect the neural correction to the electronic problem. In self-consistent inference, the derivative changes the Hamiltonian, so the descriptors change again as the density evolves. This differs from adding a number to a completed PBE energy without modifying its density. A model trained only to a post-processing energy correction is not automatically interchangeable with a traced model intended for self-consistent deployment.
The numerical descriptor functions and primary LCAO orbitals have different roles. The first represent the projected information supplied to the model; the second expand the electronic states. Matching one while silently changing the other can alter the model inputs. Record both file hashes and any required cutoff/normalization conventions.
The point cloud illustrates a validation domain, not a trained model's measured uncertainty. Distances in a sketch do not yield a calibrated error bar.
8.3.3 Conditional INPUT and descriptor addition
Begin with a converged LCAO parent whose Hamiltonian matches the model's documented baseline. The first silicon teaching case uses LDA assets; a PBE-trained model therefore requires replacing the pseudopotential and orbitals with a matched PBE family and reconverging PBE before continuing. If the model has another baseline, establish that baseline instead. Keep a separate model-off task at identical geometry and numerical settings.
# ABACUS 3.9.0; valid only with a compatible trained traced model
basis_type lcao
deepks_scf 1
deepks_model compatible_traced_model.pt
compatible_traced_model.pt is a placeholder for a verified asset, not a supplied file. Preserve the complete parent's STRU and add the descriptor section required by that model:
# STRU addition; retain existing NUMERICAL_ORBITAL and geometry
NUMERICAL_DESCRIPTOR
compatible_descriptor.orb
The descriptor filename is likewise conditional. Its contents must match the training construction, not simply have an .orb extension. deepks_out_labels serves label-generation workflows; it is not a synonym for enabling inference. Keep that task distinct and use the release's documented emitted files and unit contracts. Do not apply PW-only descriptor-generation settings to this LCAO inference recipe.
8.3.4 A worked validation workflow
Define the intended domain before opening any output: for example, a particular composition, crystal phase and bounded strain range. Select an equilibrium frame plus modest positive and negative strain frames from that domain, and at least one explicitly excluded structure to test rejection policy. These are sampling categories, not promised model support. For each supported frame, run the baseline, DeePKS and independently specified reference method with consistent geometry and energy accounting.
Hash the traced model and descriptor before inference; record the ABACUS version, model-export framework and supported runtime. Inspect the startup log for model loading, accepted LCAO/DeePKS mode and SCF termination. Save OUT.<suffix>/running_scf.log and all emitted diagnostics. A successful library call does not establish chemical transferability. Check electron count, selected magnetic state, final residual and any unexpectedly unstable iterations.
Report errors relative to the target actually used for training or validation. A reference computed with a different pseudopotential valence or structural constraint may require a carefully justified alignment; comparing unrelated absolute totals can be meaningless. Prefer consistent energy differences within a composition and verify forces only when the deployed method supports the corresponding derivative path. Finite-difference forces require tighter electronic convergence than a qualitative energy comparison.
8.3.5 Blank error table and acceptance policy
For a frame \(i\), define \(\delta E_i=E_i^{\mathrm{model}}-E_i^{\mathrm{ref}}\) and state whether you report eV/frame or eV/atom. For forces, state whether the error is componentwise or a vector norm and use eV/Å. Do not quote a single aggregate metric without its sampling domain.
| Frame / domain group | Baseline and asset hashes | Model / descriptor hashes | SCF accepted? | Baseline error | DeePKS error | Force check if supported | In-domain decision |
|---|---|---|---|---|---|---|---|
| equilibrium, supported | |||||||
| positive strain, supported | |||||||
| negative strain, supported | |||||||
| deliberately outside domain | reject / investigate |
Choose acceptance limits from the application and the independently verified reference uncertainty. A model may improve equilibrium energies while worsening distortions or forces. An extrapolation failure is not fixed by increasing ordinary SCF iterations alone. When a target frame lies outside the documented domain, use a trusted reference workflow or a new validated model development cycle rather than labeling a converged prediction reliable.
8.3.6 Pitfalls, exercises and answers
Pitfalls. A DeePMD frozen potential is not a DeePKS traced functional. PBE-trained correction plus the LDA teaching baseline violates the model contract. Replacing the descriptor file changes feature meaning even if dimensions still match. Reporting model-off energy from one geometry and model-on energy from another confounds the correction. Treating SCF convergence as an uncertainty estimator confuses numerical and statistical questions. An energy-only validation does not establish force, stress or reaction-path accuracy.
Exercise 1. A model loads successfully but its training elements exclude one slab species. May the result be accepted because SCF converges? Answer: no. Loading and convergence do not establish applicability; reject or validate through an explicitly supported development process.
Exercise 2. Model error is small on adjacent MD frames randomly split into train/test sets. Does that prove transferability to a new phase? Answer: no. Correlated frames can make the test set too similar to training. Use independent trajectory/phase groups and evaluate the intended domain, including explicitly challenging structures.
8.3.7 Sources
- ABACUS 3.9.0 DeePKS keyword definitions.
- ABACUS 3.9.0 DeePKS build prerequisites.
- Official DeePKS-kit repository: model development and self-consistent versus perturbative use.
- University-linked ABACUS teaching guide: hands-on context, with release differences checked separately.