8.1 Benchmark MPI, threads and solvers
8.1.1 Measure the same scientific work
A faster ABACUS run is useful only if it performs the same accepted calculation. Changing the orbital basis, precision, SCF threshold or k mesh while adding processors is not a strong-scaling benchmark. Nor does a short silicon teaching cell predict the best layout for a large metallic surface. This lesson designs measurements; it supplies no invented speedup and launches no real cluster job.
The templates are unexecuted and pinned to the ABACUS 3.9.0 CPU interfaces. Start from first LCAO silicon or first PW silicon, already numerically converged for the observable you need. The original LDA teaching assets may remain unchanged for this performance experiment because the Hamiltonian is fixed. Choose a workload large enough to expose the intended bottleneck, then freeze its INPUT, STRU, KPT and species asset hashes.
Record executable version, compiler, MPI implementation, BLAS/ScaLAPACK/ELPA versions, CPU model, memory, node count and launcher binding. GPU paths and mixed precision are separate, version- and build-specific experiments, not interchangeable CPU layouts. A supported solver name is not evidence that its dependency was compiled into your executable.
8.1.2 Parallel decomposition and scaling definitions
For fixed work with a reference layout \(p_0\),
The reference must be stated: comparing 16 ranks with 4 ranks is not the same speedup normalization as comparison with one rank. Node-hours or core-hours add a resource-cost perspective. A shorter wall time can consume more total compute, and either may be the right objective depending on the scheduling constraint.
PW work distributes wavefunction, FFT and k-point tasks; LCAO work includes numerical-grid assembly and generalized diagonalization. A local-orbital generalized eigenproblem uses an overlap matrix, so solver choices must match that representation. Small matrices can make communication dominate; too few useful k points can leave k pools imbalanced.
The blocks depict decomposition. The curve is conceptual, not a measured hardware result or an advertised ABACUS speedup.
8.1.3 Worked layout delta and launcher template
Duplicate the same converged parent and change only the tested layout parameter:
For the second folder, replace kpar 1 with kpar 2, but first check the accepted physical k-point count and the MPI rank count. The documented upper bound is the smaller of those counts. Choose a rank allocation that the MPI implementation and ABACUS decomposition can distribute evenly for a clean comparison. Gamma-only workloads cannot create two useful k pools by requesting them.
The following Linux shell commands are illustrative launcher syntax, not executed commands or a scheduler submission:
# Run separately inside a fresh, identically prepared benchmark folder.
export OMP_NUM_THREADS=1
mpirun -np 4 abacus > stdout.log 2>&1
Before an actual run, use your institution's launcher and allocation rules. Four ranks is a teaching layout, not a recommendation for a material. For a threading study, hold allocated CPU cores fixed and compare rank/thread decompositions only where the build and linear-algebra libraries honor that configuration. Record BLAS and ELPA threading too. Setting OMP threads without controlling nested library threads can create oversubscription. Do not introduce speculative ABACUS environment variables to hide this.
8.1.4 Separate solver studies from layout studies
For PW, the pinned documentation lists cg, dav, dav_subspace and bpcg paths with different memory behavior and build/device restrictions. For LCAO it documents generalized solvers such as genelpa and scalapack_gvx, with availability dependent on compilation; a sequential build has a different default path. Do not use genelpa with PW merely because it accelerates a LCAO run.
First benchmark layouts with one fixed solver. Then, in a separate experiment, change the solver while retaining the same physical problem and convergence target. Inspect residuals and accepted eigenvalue accuracy, not only iteration count. A Davidson subspace expansion can change memory and per-iteration time. A different initialization may alter SCF iteration count even when the final solution agrees. Retain warm-start and cold-start timing as separate categories.
Do not set bndpar as if it were a general LCAO band-parallel flag: the pinned documentation limits that decomposition to stochastic orbitals. Avoid claiming that every k-pool option, GPU backend and hybrid-exchange path has the same scalability. Validate the particular supported combination you measure.
8.1.5 Outputs, blank timing table and acceptance
Save stdout.log, the full OUT.<suffix>/running_scf.log, accepted solver/layout messages and timing summaries. Measure elapsed time with a tool appropriate to your platform and record whether initialization and I/O are included. Repeat measurements under comparable load, preferably several times, and report spread rather than a single favorable minimum. A repeat that reuses old output is not equivalent to a fresh first run.
| Problem hash | Ranks × threads | kpar | Solver / libraries | SCF iterations | Energy / force check | Elapsed times (s) | Peak memory | Core-hours |
|---|---|---|---|---|---|---|---|---|
| 4 × 1 | 1 | recorded | ||||||
| 4 × 1 | 2, if valid | same | ||||||
| fixed cores, alternate split | valid | same |
Reject a fast layout that fails convergence or changes the selected magnetic/electronic branch. Numerically small reduction-order differences can be expected, but they must remain below the scientific tolerance. If iteration counts differ, report both total time and time per comparable step; do not claim the latter alone represents end-to-end throughput.
8.1.6 Pitfalls, exercises and answers
Pitfalls. Counting ranks rather than allocated cores hides threading costs. Ignoring affinity can make a solver comparison a placement comparison. Requesting more pools than useful k points wastes resources. Changing numerical precision during strong scaling changes the work. Selecting a layout on a two-atom insulator and applying it to a 500-atom metal skips the bottleneck analysis. Timing a failed SCF produces an attractive but meaningless result.
Exercise 1. Relative to four ranks, eight ranks reduce elapsed time by a factor of 1.5. What is the efficiency? Answer: \(1.5/(8/4)=0.75\). This is an arithmetic exercise, not measured performance.
Exercise 2. A threaded layout is faster but uses twice the allocated cores. Is it more efficient? Answer: not necessarily. Compare core-hours, memory and the relevant throughput objective, while retaining equal scientific accuracy and clearly stating the resource allocation.
8.1.7 Sources
- ABACUS 3.9.0 acceleration overview.
- ABACUS 3.9.0 solver and kpar definitions.
- ABACUS 3.9.0 build options.
Next: safe restarts and provenance.