HPC & Slurm
From workload and cluster architecture to your first job, parallel layouts, reproducible workflows and day-to-day operations. Ten standalone guides distinguish users, administrators and shared concepts, with release boundaries, teaching examples and official sources.
Choose a starting point
- Users: start with guides 05–08 and 10 for resource policy, submission, parallel execution, environments and diagnosis.
- Administrators: begin with guides 01–04 and 09 for architecture, isolated labs, authentication, configuration and maintenance.
- Shared foundations: distinguish login from compute nodes, allocation from application parallelism, and scheduler success from scientific validity.
Teaching examples and production requirements
Reviewed 2026-10-04. Use the documentation matching installed Slurm and local policy. A lab does not establish production interconnect, storage or security behavior. Administrator examples require authorization and review. All examples are educational; no actual cluster operations were performed.
Ten-guide learning path
Plan a cluster around the workload
Roles, capacity and representative benchmarks
Read the guide → 02Build a safe learning lab; provision production deliberately
Isolated VMs, versioned provisioning and canary rollout
Read the guide → 03Identity, time, network, and authentication
UID/GID, time, protected keys and authentication scope
Read the guide → 04Read Slurm configuration as a system
Measured nodes, resource containment and persistent accounting
Read the guide → 05Partitions, QOS, fair share, and honest requests
Partitions, QOS, priority, memory and CPU requests
Read the guide → 06Submit, inspect, and stop a job
Batch, interactive allocation, accounting and cancellation
Read the guide → 07Match MPI, OpenMP, and GPU execution to the application
Thread/rank layouts, compatible launchers and GPU validation
Read the guide → 08Reproducible software and scratch workflows
Modules, coherent builds, checkpoints and safe scratch lifecycle
Read the guide → 09Monitor and maintain without surprising users
Monitoring, drain/resume, backups and tested upgrades
Read the guide → 10Diagnose, measure, and automate scientific workflows
Job states, memory metrics, arrays, dependencies and throughput
Read the guide →Running quantum chemistry software
- VASP - ORCA - Quantum ESPRESSO
First check fit, licence and versioned documentation in the software directory, then use the parallel-layout guide to choose a compatible launch method. Generic Slurm examples do not replace an application’s own release-specific execution requirements.