todo.academy Try a free chapter

True zero to confident Slurm user

Slurm Core Operations

Operate Slurm from first request to evidence-backed diagnosis.

Learn Slurm work from first request to verified history: inspect resources, submit jobs, diagnose waits and failures, build workflows, and recover safely.

No credit card for the free chapter. Preview what the chapter covers.

By the end

Observable skills you will practice.

  • Explain what Slurm is and how scheduling decisions flow
  • Submit, monitor, cancel, and verify batch and interactive work
  • Right-size CPU, memory, time, partition, array, dependency, MPI, and GPU requests
  • Diagnose failures from queue state, reason codes, accounting, exit codes, and node evidence
  • Recognize policy, QOS, fairshare, reservation, and operator-level causes

Performance-assessed credential

Earn the todo.academy Skills Credential: Slurm Core Operations.

Complete the focused path, then pass a hands-on final assessment whose required actions are replayed and checked before eligibility is created.

  1. 01Complete the focused path

    Clear every required course checkpoint. Advanced practice remains optional.

  2. 02Pass the final assessment

    Produce the required results without using optional readiness scenarios as a shortcut.

  3. 03Share only when you choose

    Download the artifact, add it to LinkedIn, or share a public status page with verifiable evidence.

Independent skills credential issued solely by todo.academy. Not affiliated with, sponsored by, endorsed by, or issued by SchedMD LLC. Slurm is a registered trademark of SchedMD LLC.

Syllabus

26 chapters from the published course structure.

The focused path contains 42 required lessons. The advanced library remains available when your role needs more depth.

Focused path chapters11 chapters contain the required 42-lesson path.
Chapter 00Free

What Slurm is

A complete free foundation chapter. Learn what Slurm is, why shared clusters need it, how jobs enter the system, and how to prove what happened.

10 focused lessons + 1 advanced lesson
Chapter 01Full course

Batch scripts as contracts

Learn how a shell file becomes a Slurm job request, how directives shape the record, and how to prove output, errors, paths, and accounting.

5 focused lessons + 6 advanced lessons
Chapter 02Full course

Live queue triage

Read live queue state, decode pending reasons, filter noisy output, inspect detailed job records, and choose the next action from evidence.

5 focused lessons + 6 advanced lessons
Chapter 03Full course

Resource requests and fit

Understand partitions, time, CPUs, tasks, memory, GPUs, and why a request fits, waits, or needs repair.

6 focused lessons + 5 advanced lessons
Chapter 04Full course

Interactive work and job steps

Use srun, salloc, job steps, short interactive shells, queue proof, and cleanup without bypassing the scheduler.

4 focused lessons + 7 advanced lessons
Chapter 05Full course

Software environment and reproducibility

Use modules, runtime version proof, Slurm job variables, clean logs, and accounting to make jobs reproducible.

4 focused lessons + 7 advanced lessons
Chapter 06Full course

Accounting and postmortem evidence

Use sacct fields, step rows, MaxRSS, DerivedExitCode, sstat, seff, stderr, and job detail to diagnose completed or failed work.

3 focused lessons + 9 advanced lessons
Chapter 07Full course

Failure diagnosis and recovery decisions

Classify cancellation, signals, node failure, preemption, requeue, timeout, OOM, and application failure before choosing the next action.

1 focused lesson + 11 advanced lessons
Chapter 08Full course

CPU and memory troubleshooting

Diagnose CPU waste, threaded shape, memory pressure, memory per CPU, live stats, efficiency summaries, and right-sized fixes from evidence.

1 focused lesson + 11 advanced lessons
Chapter 09Full course

The incident gauntlet

Classify Slurm incidents, gather proof, rank hypotheses, choose safe action, verify outcomes, and escalate with exact evidence.

2 focused lessons + 14 advanced lessons
Chapter 10Full course

Final professional capstone

Own a deadline workflow from cluster snapshot through preprocess, array repair, GPU, MPI, hidden failure diagnosis, fixed report, and final evidence packet.

1 focused lesson + 12 advanced lessons
Explore 15 advanced library chapters

These chapters provide optional drills and reference depth. They are not required to complete the focused path.

Library 07Full course

Chapter 7: Job arrays and polite scaling

Scale repeated work with array indexes, task variables, safe per-task logs, throttles, targeted reruns, and accounting proof.

11 advanced lessons
Library 08Full course

Chapter 8: Dependencies and workflow graphs

Build workflow edges with afterok, afterany, afternotok, fan-in, OR release, never-satisfied diagnosis, singleton guards, and accounting proof.

12 advanced lessons
Library 12Full course

Chapter 12: Accounts, QOS, and priority

Read account associations, QOS limits, policy holds, priority factors, fairshare context, and repaired debug QOS requests from evidence.

12 advanced lessons
Library 13Full course

Chapter 13: MPI, GPU, and placement

Prove rank layout, multi-node distribution, CPU binding, exclusive allocations, GPU GRES, GPU visibility, and TRES accounting.

12 advanced lessons
Library 14Full course

Chapter 14: Operator evidence and safe state changes

Read daemon health, config intent, partitions, nodes, reservations, GRES, cgroups, SlurmDBD policy data, and safe scontrol update behavior.

12 advanced lessons
Library 15Full course

Chapter 15: Federation, heterogeneity, and site-scale policy

Read federation scope, heterogeneous components, submit plugins, file staging, usage rollups, site variance, and production runbook discipline.

12 advanced lessons
Library 16Full course

Chapter 16: Scheduler internals and backfill diagnosis

Explain scheduler behavior from config, diagnostics, start estimates, priority, preemption, topology, and final diagnosis packets.

12 advanced lessons
Library 17Full course

Chapter 17: Reservations, maintenance windows, and planned capacity

Read reservation inventory, access scope, reservation submissions, pending reasons, maintenance flags, safe updates, deletion guards, and planned down evidence.

12 advanced lessons
Library 18Full course

Chapter 18: Preemption, requeue, and restart-safe work

Prove preemption policy, read requeued and suspended states, handle signals, protect output, checkpoint progress, and repair long jobs for safe restart.

12 advanced lessons
Library 19Full course

Chapter 19: Containers and reproducible environments

Use Slurm container requests with site-aware OCI config, clean environment boundaries, bind paths, GPU exposure, policy diagnosis, and proof-backed reproducibility.

12 advanced lessons
Library 20Full course

Chapter 20: Operator view, daemons and control plane

Read controller health, node daemon evidence, accounting delay, authentication trust, logs, diagnostics, and incident classification without unsafe guesses.

12 advanced lessons
Library 21Full course

Chapter 21: Configuration as policy code

Read slurm.conf, node and partition definitions, includes, configless distribution, GRES mapping, protected SlurmDBD config, reconfigure safety, and rollback evidence.

12 advanced lessons
Library 22Full course

Chapter 22: Accounts, associations, limits, and billing

Read accounting policy data, association tuples, users, accounts, QOS allowed lists, TRES-minute limits, fairshare context, billing evidence, onboarding, and access failures.

12 advanced lessons
Library 23Full course

Chapter 23: Throughput, scheduler load, and good citizenship

Refactor noisy repeated work into polite high-throughput arrays with throttles, targeted polling, clean output, responsible chunking, and controller diagnostics.

12 advanced lessons
Library 24Full course

Chapter 24: Advanced placement, dynamic capacity, and federation

Reason about topology-aware placement, node features, dynamic and cloud-style capacity, heterogeneous components, federation scope, and final advanced-placement proof.

12 advanced lessons

Full access

Choose access for Slurm Core Operations

Every available offer is shown with its exact CAD price and billing model. Checkout opens only after you choose an offer and enter the receipt email.

Available access options

The selected offer and exact total remain visible before payment.

Already purchased? Restore access

Payments are processed by Lemon Squeezy for Phoenix Soft Inc. Paid access can be restored after secure sign-in.

Questions

Know what to expect before you start.

What can I try for free?

Chapter 1: What Slurm is is free and contains 11 lessons. No credit card is requested before the free workspace opens.

What experience do I need?

No Slurm experience is required. Basic command-line familiarity is helpful.

What does the focused path include?

42 required lessons, with 10 free and 32 included in paid access.

Is there material beyond the focused path?

Yes. The course also includes 268 advanced practice and reference resources. They are available when you need more depth, but they do not lengthen the required path.

Which paid options are available?

Slurm Founder Pass: CA$59 one-time. Academy Founder Annual: CA$139 per year. Founder Vault: CA$279 one-time.

How is paid access billed?

Slurm Founder Pass: No recurring charge. Access remains available unless the purchase is refunded or reversed. Academy Founder Annual: Access continues while the annual subscription remains active. Founder Vault: No recurring charge. Coverage includes courses launched during the first 24 months.

How long does it take?

The published estimate is 6 to 10 active hours for Slurm Core Operations. Your pace will depend on how much you repeat the practice.

Do I need to install anything?

No installation is required to start the guided free chapter. Later lessons explain the real tools, files, and operating boundaries relevant to the skill.

What technology is covered?

The syllabus and these stated outcomes are the source of truth: Explain what Slurm is and how scheduling decisions flow; Submit, monitor, cancel, and verify batch and interactive work; Right-size CPU, memory, time, partition, array, dependency, MPI, and GPU requests; Diagnose failures from queue state, reason codes, accounting, exit codes, and node evidence; Recognize policy, QOS, fairshare, reservation, and operator-level causes.

Which browsers are supported?

Use a current browser with JavaScript enabled. The free chapter is the quickest compatibility check for your device.

Does the workspace work on mobile?

The reading pages reflow for small screens. Command-heavy practice is more comfortable with a physical keyboard and a larger display.

What happens when I make a mistake?

The practice state is isolated from production systems. Read the resulting evidence, revise the action, and try again.

How do I restore access?

Use the receipt email on the restore-access page. The sign-in link verifies the account before paid entitlements are loaded.