1.3 Reliable HPC execution, restart validation and reproducible automation
These are unexecuted teaching inputs and starting settings to test. Original figures are schematics, not computed results. Use licensed VASP and PAW data, replace every placeholder, record the executable version and validate convergence.
1.3.1 Model, units and provenance
Use eV for energy, Å for length and eV/Å for force; 1 kbar = 0.1 GPa. State normalization per atom, molecule, primitive cell or simulation cell. Record PAW identifiers, release, ZVAL, ENMAX and permitted hashes; never redistribute POTCAR. SCF convergence addresses the chosen electronic problem; convergence of the target property requires separate tests.
Original schematic. Curves explain concepts; blank data areas await verified learner results. No calculation is claimed.
1.3.2 Worked case: procedure, interpretation and checks
Intuition
A job ending normally is an operational result; a converged and interpretable trajectory is a scientific result. Automation should make assumptions visible and reject bad states, not merely keep resubmitting until a directory contains output. Long AIMD combines electronic convergence, integrator state, file integrity, queue limits and expensive data. Treat all five as parts of the workflow.
Prerequisites. Know the cluster’s scheduler, module system, licensed VASP build and recommended launch command. Never paste another institution’s account, partition or MPI launcher into a production script. Confirm storage quota and expected output size. Keep each attempt in a new directory with immutable inputs and a parent-run identifier.
Stage 1: benchmark the actual workload
Measure cost per completed ionic step after startup on representative snapshots. Record wall time, allocation size, memory, SCF iterations and numerical results. Compare scientifically equivalent runs. One fast step with a different number of bands or a changed algorithm is not a clean parallel benchmark.
For ordinary MPI CPU builds, NCORE divides work within a band and KPAR groups k-point work. Do not set NPAR and NCORE together; the legacy NPAR can take precedence. OpenMP/GPU paths have different behaviour, including NCORE handling, so follow the current build documentation rather than recycling CPU-only settings. NCORE
A Γ-only calculation has no extra k points to distribute usefully across multiple KPAR groups. Choose divisibility and group sizes compatible with ranks and actual k points; benchmark rather than assuming all available cores help. The number of cores should be selected for efficiency and queue constraints, not prestige. KPAR
This intentionally does not submit a job. Hashes document identity; do not publish restricted potential contents. Record the actual VASP version from program output, not only the filename in the environment variable. Include scheduler job ID, compiler/MPI/GPU environment where relevant, and the resolved input settings printed by VASP.
Stage 2: stop and checkpoint deliberately
Leave a wall-time margin large enough for the next ionic step and file writes; the appropriate margin depends on measured worst-case step time. The STOPCAR option LSTOP=.TRUE. asks VASP to stop at an ionic-step boundary; LABORT=.TRUE. can stop within electronic work and leave nonconverged electronic state. Prefer the orderly route when possible. STOPCAR
After the job exits, verify that restart files are complete, nonempty and parseable, and that their atom count and cell match expectations. A scheduler kill during a write can leave a misleading filename with incomplete contents. Keep the previous verified checkpoint until the new one passes validation.
CONTCAR can contain coordinates, velocities and predictor-corrector information needed for MD continuation. Preserve the state appropriate to the same algorithm. When changing ensemble, VASP documents removal of incompatible predictor-corrector data; use a format-aware procedure, preserve intended velocities and test a short continuation. Copying coordinates alone creates a new velocity initialization if none are provided, so it is not a seamless trajectory restart. CONTCAR
Stage 3: distinguish recovery from changing the experiment
A job that reaches NELM without satisfactory SCF convergence must not silently be accepted as a correct MD force evaluation. Increasing NELM may help a genuinely converging sequence, but it does not repair charge sloshing, pathological geometry or inappropriate electronic settings automatically.
Use a structured triage:
- Identify the first failed ionic step, not only the last error message.
- Inspect that geometry, shortest contacts, prior temperature and force history.
- Check SCF residual/energy history, spin changes, occupation behaviour, available bands and relevant warnings.
- Reproduce the problem on the saved snapshot with a controlled static electronic calculation.
- Test a justified recovery, such as a more robust electronic algorithm or improved initialization, while documenting every change. If the scientific Hamiltonian changes, create a new study branch and re-equilibrate rather than merging results silently.
- Resume from the last trustworthy physical state. If one or more ionic steps used untrustworthy forces, simply deleting their plotted values does not undo their effect on later coordinates.
- Set retry bounds and escalation conditions in automation. Never loop indefinitely with successively looser convergence merely to make the scheduler report success.
This is a diagnosis strategy, not a universal mixing-parameter recipe. A magnetic metal, molecule in vacuum and ionic liquid can fail for different reasons. Retain the failed run so the repair is auditable.
Stage 4: validate continuation and analysis joins
Run a short restart test before planning a very long campaign. Check temperature, kinetic energy, structure, cell, and conserved/ensemble-relevant energy around the boundary. Small chaotic divergence or differences across parallel layouts do not necessarily indicate scientific failure, but an abrupt change from lost velocities does.
Every segment needs a global start time, timestep, frame interval, and whether the first saved frame duplicates the previous final frame. Concatenation must detect duplicates and gaps. Do not glue XDATCAR files by removing a guessed number of header lines if format, cell or species can differ. Use a reader that verifies metadata and preserves per-frame cells. NBLOCK affects MD output cadence, and ML output controls may add another layer; determine the actual timing rather than hardcoding it. NBLOCK
A useful manifest schema is:
The explicit nulls and false flags are deliberate: prepared metadata is not evidence of completion. An analysis pipeline should fail closed until required checks pass. Store software versions, script hashes, plotting parameters and physical unit conversions with outputs.
Traps: overwriting the only restart; leaving STOPCAR in a continuation directory; considering exit code 0 a convergence certificate; mixing parameter changes into one trajectory; deleting failed SCF rows but retaining downstream frames; publishing licensed POTCAR content; assuming a seed alone ensures bitwise reproducibility on different hardware.
Exercise. Stage a short local test campaign, intentionally stop it cleanly, and restart it. Compare the boundary with an uninterrupted short run using appropriate tolerances and ensemble statistics. Then simulate a missing or truncated restart file in a copied test directory and verify that your wrapper refuses to submit. Do not corrupt the original scientific data.
1.3.3 Unexecuted inputs and analysis scaffolds
These are unexecuted teaching inputs and starting settings to test. Original figures are schematics, not computed results. Use licensed VASP and PAW data, replace every placeholder, record the executable version and validate convergence.
1.3.3.1 Input block 1
#!/usr/bin/env bash
# UNEXECUTED scheduler-neutral launch skeleton, NOT a complete batch script
set -euo pipefail
: "${VASP_EXE:?Set the site-approved executable path}"
: "${RUN_ID:?Set a unique run identifier}"
for f in INCAR POSCAR KPOINTS POTCAR; do
test -s "$f" || { echo "Missing or empty $f" >&2; exit 2; }
done
test ! -e STOPCAR || { echo "Review old STOPCAR before launch" >&2; exit 3; }
sha256sum INCAR POSCAR KPOINTS POTCAR > "inputs.${RUN_ID}.sha256"
printf '%s\n' "$VASP_EXE" > "executable.${RUN_ID}.txt"
# Insert the cluster-approved scheduler/MPI launch command HERE.
# Do not run a multi-node executable directly as a substitute for that command.
exit 4 # Intentional fail-closed placeholder until launcher is configured.
1.3.3.2 Input block 2
{
"run_id": "REPLACE_WITH_UNIQUE_ID",
"parent_run_id": null,
"purpose": "NVT equilibration pilot",
"status": "prepared_not_run",
"vasp_version": "REPLACE_AFTER_READING_OUTPUT",
"input_hash_manifest": "inputs.RUN_ID.sha256",
"timestep_fs": 0.5,
"nominal_steps": 400,
"completed_steps": null,
"global_start_time_ps": 0.0,
"saved_frame_interval_fs": null,
"energy_definition": "REPLACE_AFTER_VALIDATING_PARSER",
"electronic_convergence_checked": false,
"restart_state_checked": false,
"analysis_approved": false
}
1.3.4 Related learning paths
- 1.1 Four input files, one physical question
- 1.2 A convergence laboratory with an error budget
- 3.6 Adsorption energies with consistent references, coverage and dispersion