Running Jobs¶
ARC clusters use SLURM (Simple Linux Utility for Resource Management) as the job scheduler and workload manager. SLURM allocates compute resources to users, manages job queues, and ensures fair sharing across the cluster.
Job Modes¶
Slurm has 2 modes of operation: batch jobs and interactive jobs.
Batch Jobs¶
Batch jobs are SLURM’s primary mode of operation, users write run-scripts, then submit them to the scheduler to be run on the requested hardware as soon as it’s available. This mode is preferred when the workflow is known, the cluster is has a high utilization, or jobs are expected to take a long time. Read more about batch jobs here
Interactive Jobs¶
Interactive jobs are special jobs that allow users to use an allocation as if it was their local machine, they can issue commands and see outputs immediately on their terminal. Interactive jobs are preferred when workflows are unknown (such as testing or building software), cluster utilization is low, or jobs are bursty in nature. Read more about interactive jobs [here]
Life-cycle of a Job¶
Whether submitting batch jobs or interactive ones, each SLURM job goes through a life-cycle.
Submission: jobs start with the submission.
sbatch,salloc, orsruncommands are combined with resource request directivesQueue: once submitted, slurm places a job in the queue. here it waits until the requested resources are available. If resources are available upon submission, the job is launched immediately
Launch: once the necessary resources are available, SLURM launches the job. It’s at this stage that allocations are bound and the environment is setup for the job.
Runtime: whether interactive or batch, the allocation is handed over to the job itself. in the case of batch, execution of the run-script begins. in the case of interactive, control of the nodes is handed over to the user, ready to run their commands.
Completion: the job has completed, either due to the end of a run-script, the exit command, or time limit. the allocation is freed and added to the available pool of resources for new jobs.
Why Use a Scheduler?¶
Clusters are shared resources; while 100, 200, even thousands of nodes may seem like a near-infinite amount of resources, then can quickly be consumed by a few individuals. A scheduler allows the system to be fairly shared between many users without any of them hogging resources. These guarantees also flow both ways, not only is the cluster prevented from being hogged by a few individuals, but users’ jobs are also guaranteed to be running on their own, dedicated allocation.
A scheduler also has benefits to the user. With a scheduler, a user can do powerful things such as queue hyper-parameter sweeps in parallel across jobs, script workflows regardless of which nodes are available, and much more. Read more about advanced SLURM usage here
Slurm Directives¶
SLURM Allocations are defined by directives–flags that specify a specific constraint. sbatch, salloc, and srun use these directives to resolve the correct allocation
Essential Directives¶
Most SLURM directives are optional, here are the few that are required
Directive |
Description |
|---|---|
|
Specify which account to charge to. %%Learn more [here]%% |
|
Specify the maximum runtime (or WallTime) |
Common Directives¶
While very few directives are required, it’s advised to provide as many are required to adequately describe your desired allocation
Directive |
Description |
|---|---|
|
Desired nodes |
|
Desired tasks (processes) to launch |
|
Similar to |
|
Desired Cores to allocate each task. Necessary for multi-threaded jobs |
|
Desired memory per node |
|
Path to write program output to (stdout, stderr) |
|
Name that shows up in slurm accounting ( |
|
Whether to send emails for certain states (NONE, START, END, FAIL, ALL) |
|
Which email to send the above emails to |
|
Desired partition. See which are available for each Cluster |
|
Desired QoS. See which are available for each Cluster |