Skip to content

Reproducible software and scratch workflows

Audience: Both users and administrators · Reviewed: 2026-10-04

Version and teaching scope

Use documentation matching your installed release and local policy. Check sinfo --version; upstream documentation identified itself as 26.05 at review. Paths, partitions, accounts, modules, memory, MPI, GPUs and services are site-specific. Example programs beginning ./ must already exist. These are teaching examples, not a validated deployment recipe or evidence that a cluster was modified.

Inspect the software environment

A module changes the shell environment; it is not necessarily a software installation or an isolation boundary. Lmod hierarchies can hide applications until a compatible compiler and MPI are loaded. Inspect before loading:

module list
module avail
module spider openmpi
# Replace the placeholder before running this line:
module show "example-module/1.0"

The last line is a placeholder: replace example-module/1.0 with the exact module name before running. In a batch script, load exact site-supported versions after the site's required module initialization. A deliberate module purge can remove required defaults, so use it only when the site recommends it. Record module output, program version, input checksum, code revision, and job resource layout alongside results. Do not dump the complete environment into a public log because it can contain secrets.

Preserve and test a coherent build

For builds, preserve compiler, MPI, math-library, build flags, and test results as one coherent stack. A successful link is not evidence of numerical correctness. Run a small trusted reference calculation and compare energies or observables with documented tolerances. Use a build allocation for resource-intensive compilation; do not run unrestricted parallel builds on a shared login node.

Understand scratch visibility and lifetime

Scratch is a workflow choice, not a magic variable. $SLURM_TMPDIR is provided by some sites, not guaranteed by Slurm everywhere. Determine the documented scratch path, quota, cleanup policy, visibility, and backup status. In a multi-node job, local scratch on node A is not automatically present on node B.

Stage, validate, persist, then clean

A robust sequence is: create a job-specific directory, stage only required inputs, run with explicit exit-status handling, copy validated results/checkpoints to persistent storage, verify the copy, and then clean up. Preserve failed-job diagnostics. A shell trap cannot recover data after every node crash or forced kill; write periodic checkpoints to a suitable persistent destination when the application supports it.

Exercise: Repeat a small calculation and identify the metadata needed to reconstruct its environment.

Official sources

Previous: Match MPI, OpenMP, and GPU execution to the application · Learning path · Next: Monitor and maintain without surprising users