In 20 years, you will be more dissapointed by what you didn't do than by what you did.

Slurm Job Pending with Idle GPUs: A 256-GPU Sizing and Troubleshooting Checklist

A Slurm job can remain pending while a dashboard shows dozens of idle GPUs. Before buying more accelerators or changing the fabric, ask a narrower question: can the scheduler assemble the exact node, GPU, CPU, memory, policy and locality shape requested by this job? A cluster-wide free-GPU total is not enough to answer it.

This guide provides a 256-GPU capacity worksheet, a pending-job troubleshooting matrix, two original diagrams and a deployment checklist. The design and queue snapshots are hypothetical; the arithmetic was executed in Python. Configuration and diagnostic commands are unexecuted examples, not results from a production cluster. They require adaptation to your installed Slurm release and site policy.

1. Start with the job shape, not the idle-GPU graph

Separate three questions in the incident ticket: did the job receive an allocation, did its tasks receive the expected devices, and did the running workload use them effectively? Do not troubleshoot an unallocated job with collective-communication tuning. Conversely, a running allocation with low application throughput needs evidence beyond scheduler state.

Slurm's Generic RESources (GRES) framework manages GPUs only when configured, and jobs do not receive GRES unless they request them.[1] The GPU options also have different scopes: --gpus is per job, --gpus-per-node is per node, and --gpus-per-task is per task.[1] The --gpu* request options require select/cons_tres; the documented --gres mechanism has a different compatibility boundary.[1]

Write down this request tuple before changing anything:

  • Exact node count and whether the application can use a smaller allocation.
  • GPUs per node, GPU type, and whether whole devices are required.
  • Tasks per node and the intended process-to-GPU mapping.
  • CPUs and host RAM required on each node, separately from GPU memory.
  • Partition, account, quality of service (QOS), reservation and dependencies.
  • Placement constraints, time limit and requested topology locality.

Treat that tuple as the input to the worksheet. “I need 64 GPUs” is incomplete if the actual application requires eight nodes with eight GPUs each and cannot run across sixteen half-empty nodes. This article deliberately uses whole-GPU, homogeneous servers; it is not a model for NVIDIA Multi-Instance GPU (MIG), fractional sharing or elastic training.

2. The 256-GPU capacity worksheet

Hypothetical infrastructure: 32 servers, eight whole GPUs per server, four placement groups of eight servers each. Assume every eligible server can meet the example's CPU and host-memory request. The placement groups represent scheduler locality domains, not a vendor rack blueprint or a claim about a particular GPU, switch or NIC generation.

For the snapshot below, assume the occupied GPUs belong to unrelated jobs that remain running. Ignore policy limits initially so that fragmentation is visible. The target is an eight-node, 64-GPU job requesting every GPU on each allocated node.

Placement group Servers Free GPUs per server Total free GPUs Fully free servers
A 8 8 64 8
B 8 4 32 0
C 8 4 32 0
D 8 0 0 0
Total 32 Mixed 128 8

Executed worksheet results: installed capacity is 32 × 8 = 256 GPUs. The snapshot has 64 + 32 + 32 = 128 free GPUs, but only eight fully free nodes. Thus only one job with this 64-GPU shape fits simultaneously, not the two suggested by 128 / 64. This is an upper bound before eligibility, topology and policy checks, not a predicted start time.

If one server in group A becomes unavailable, the eligible free-GPU total becomes 120, with seven fully free servers. The same job now has zero feasible placements even though 120 GPUs remain free. In contrast, a four-node job requiring four GPUs per node could have a capacity-only upper bound of six simultaneous placements in the original snapshot, assuming its other resources and sharing policy permit that layout. Do not silently substitute this smaller shape for a training job that requires whole nodes.

Four server groups showing free and occupied GPUs
Original diagram by Network freak. Hypothetical whole-GPU fragmentation example; arithmetic executed in Python.

Reusable calculator: capacity bound, not a scheduler emulator

This dependency-free Python worksheet was executed locally. It counts feasible nodes for a homogeneous per-node GPU request, then applies a node-count bound. It intentionally excludes CPU/RAM fragmentation, topology, reservations, fairness and time. Enter only nodes that have passed your eligibility checks; otherwise the result overstates usable capacity.

# Hypothetical snapshot; whole GPUs, one job allocation per selected node.
free = <a href="#source-8">[8]</a> * 8 + <a href="#source-4">[4]</a> * 16 + <a href="#source-0">[0]</a> * 8

def placement_bound(free_gpus, nodes_per_job, gpus_per_node):
    if nodes_per_job < 1 or gpus_per_node < 1:
        raise ValueError("Requests must be positive")
    eligible = sum(g >= gpus_per_node for g in free_gpus)
    return eligible // nodes_per_job

print(sum(free))                         # 128 free GPUs
print(placement_bound(free, 8, 8))       # 1 full-node job
print(placement_bound(free[1:], 8, 8))   # 0 after one free node is lost
print(placement_bound(free, 4, 4))       # 6 smaller placements, upper bound

For planning, maintain both free-device totals and counts of nodes that can satisfy your common job shapes. Add separate rows for an eight-node training job, a two-node experiment and a single-node evaluation job. A useful dashboard answers “what fits?” rather than simply “how many GPUs are idle?”

3. Pending-job troubleshooting matrix

Slurm may have several reasons a job cannot start, but the displayed reason is the one encountered by the attempted scheduling method.[3] Fixing one blocker can therefore reveal another; a changing reason is not automatically evidence of inconsistent scheduling.

The following interpretations are grounded in Slurm's reason-code reference; the investigation and change-control steps are this article's recommended workflow.[3]

Displayed reason or observation Meaning or hypothesis First investigation Avoid
Resources Requested resources are unavailable Compare the complete request with per-node availability Counting free GPUs alone
Priority Higher-priority jobs exist for the relevant partition or reservation Ask the scheduler owner to explain queue order Changing node or fabric settings
QOSGrpGRES QOS aggregate GRES limit reached Compare aggregate usage with the approved QOS limit Raising quotas without authorization
AssocMaxGRESPerJob Request exceeds association per-job GRES maximum Compare requested GPU count with account policy Repeatedly resubmitting unchanged jobs
Dependency An upstream dependency is unsatisfied Inspect the dependency chain and predecessor result Treating this as GPU exhaustion
ReqNodeNotAvail Required node unavailable, possibly reserved, down or drained Inspect required node list and state reason Resuming a node without fixing its fault
PartitionTimeLimit Requested time exceeds partition limit Check the job time request and approved queue Shortening runtime below safe checkpoint needs
Running, low GPU activity Allocation succeeded; application diagnosis needed Examine rank mapping, data wait and step logs Treating allocation as proof of useful work

Use a small, timestamped evidence bundle. The commands below are unexecuted diagnostic examples; replace identifiers, confirm permissions, and check local command documentation. Run node-local commands only on the intended compute node, not indiscriminately across the cluster.

squeue -j JOB_ID -o "%.18i %.9P %.8T %.10M %.6D %R"
scontrol show job JOB_ID
scontrol show node NODE_NAME
sinfo -N -o "%N %t %G %C %m"
# Administrator: inspect discovered hardware and GRES on one compute node.
slurmd -C
slurmd -G

The final two checks are documented discovery and GRES-validation tools; slurmd -G also reports autodetected GRES that are ignored.[1] Save the job submission script alongside the job record. A wrapper may have requested a GPU type, memory amount or node list that the user did not realize was restrictive.

4. Make physical inventory agree with scheduler inventory

Slurm documents that nodes with fewer resources than configured are placed in DRAIN state.[1] GPU autodetection does not remove the controller-side inventory requirement: Gres= in slurm.conf still tells the controller how many resources to expect.[1] Investigate disagreement as an inventory or configuration problem before attempting to return the node to service.

A useful reconciliation sheet has one row per server: physical GPU count, detected names, configured GPU type/count, visible device files, GPU health result, Slurm state and the last approved configuration revision. Record device UUIDs privately when needed, but do not publish real node identifiers or inventory exports in troubleshooting articles.

The following is an incomplete, unexecuted configuration sketch, not a replacement slurm.conf. It assumes eight supported NVIDIA GPUs on each synthetic node and an existing working Slurm installation. Preserve validated CPU, memory, authentication, task isolation and accounting settings from your site configuration.

# slurm.conf excerpt only
SelectType=select/cons_tres
GresTypes=gpu
NodeName=gpu[01-32] Gres=gpu:8
TopologyPlugin=topology/tree

# gres.conf on the compute nodes: separate file
AutoDetect=nvml

The official guide distinguishes AutoDetect=nvml, which depends on NVML support, from AutoDetect=nvidia, which does not require NVML but does not detect MIGs or NVLinks.[1] Do not interchange these settings merely to silence an error. Pick the method supported by the installed build and the features actually required.

If using typed GPUs, keep typed versus untyped requests consistent between the job allocation and its steps; Slurm explicitly documents this requirement.[1] Type strings must match the detected naming rules rather than a procurement spreadsheet nickname.[1] A disciplined rollout changes one drained test node first, compares discovery output, then restores service only after a scheduled test allocation passes.

5. Model locality without pretending to model every rail

Slurm's tree topology plugin attempts to identify a low-level switch that can satisfy a request and then uses best-fit selection beneath it.[2] It is not a guarantee of the mathematically optimal allocation: the official guide notes that available resources can produce allocations spanning more switches than the optimum.[2]

For this hypothetical layout, use four leaf groups and a common parent. The names below are synthetic. This is a simplified placement model, not an as-built representation of eight independent GPU rails. Validate any abstraction against your actual fabric and Slurm version before deploying it.

# topology.conf sketch only
SwitchName=leaf1 Nodes=gpu[01-08]
SwitchName=leaf2 Nodes=gpu[09-16]
SwitchName=leaf3 Nodes=gpu[17-24]
SwitchName=leaf4 Nodes=gpu[25-32]
SwitchName=fabricRoot Switches=leaf[1-4]

The guide recommends that a simplified leaf-plus-top-level representation can be preferable to enumerating every connection, because detailed models can slow scheduling without much application benefit.[2] It also states that tree leaves without a common parent cannot be spanned by a job unless the relevant optional-topology behavior is enabled.[2] Check the common-parent model when a larger job fails despite apparently sufficient nodes.

With tree topology, --switches=count[@time] expresses a desired maximum leaf-switch count and willingness to wait for it; an administrator can limit that waiting period.[2] Treat it as a deliberate queue-time versus locality choice, not an unconditional promise that every job will remain on one leaf forever. For the 64-GPU example, only a completely free eight-server group can meet a one-group placement preference.

Do not use the example's LinkSpeed omission as missing bandwidth engineering. Slurm's topology guide says that its optional LinkSpeed information is currently unused.[2] Likewise, configured node weights can override topology-based selection.[2] Inspect those policy choices before assuming the topology plugin ignored the fabric.

Allocation gates and a four-leaf scheduler placement model
Original diagram by Network freak, based on the workflow in this article and Slurm tree-topology concepts.[2]

6. Verify task-to-GPU mapping inside the allocation

Allocation correctness and process placement are separate acceptance gates. Slurm sets CUDA_VISIBLE_DEVICES for job steps, and device numbering can differ between a job's constrained environment and a Prolog/Epilog outside it.[1] An environment value of 0 is therefore not proof that every task uses the same physical GPU.[1] Verify mapping from inside the same execution environment used by the application.

Here is an unexecuted one-node smoke-test batch script for a site with select/cons_tres, a GPU partition and suitable task/device-isolation configuration. It requests eight tasks, one GPU and four CPUs per task. Replace the partition name. This proves environment assignment only; it does not perform GPU compute or validate a distributed launcher.

#!/bin/bash
#SBATCH --partition=gpu
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=8
#SBATCH --gpus-per-task=1
#SBATCH --cpus-per-task=4
#SBATCH --time=00:05:00

srun --gpu-bind=single:1 bash -c \
 'printf "rank=%s local=%s visible=%s\n" "$SLURM_PROCID" "$SLURM_LOCALID" "$CUDA_VISIBLE_DEVICES"'

Slurm documents GPU-per-task requests and GPU binding among its supported GPU scheduling options.[1] Follow this environment check with your site's approved tiny GPU workload and device-access test. Do not set CUDA_VISIBLE_DEVICES globally to a fixed list to force a pass; that obscures the scheduler's allocation and may conflict with isolation.

Only after the one-node test passes should you repeat a validated launcher test across two nodes, then the target job shape. Preserve the framework's documented launch model. Starting an extra worker-spawning launcher from every GPU task can invalidate the intended rank layout, so review the generated process tree before measuring throughput.

7. Bottlenecks, failure domains and alternative choices

Use four separate ownership lanes. The scheduler team owns eligibility and allocation policy; the node team owns device inventory and local health; the fabric team owns connectivity and congestion; the application team owns ranks, input pipelines and useful progress. A single incident can cross these lanes, but evidence should identify the first failing boundary.

The worked failure illustrates a placement-domain problem: losing one fully free node removes the only feasible eight-node placement. Buying additional GPUs for already partially occupied nodes is not automatically the remedy. Compare reserving a whole-node pool, consolidating smaller workloads under an approved policy, reducing the application's required shape, or expanding a compatible placement group. Each alternative changes a different constraint.

Whole-node allocation is a simple baseline for reproducible distributed training, but may waste resources for small jobs. Shared nodes may improve packing while increasing CPU, memory and device-isolation testing requirements. Slurm's sharding mechanism permits GPU sharing but does not fence the processes on the GPU; do not treat sharding as tenant isolation.[1]

For a workload mixing batch training and long-lived inference, compare operational needs before choosing another orchestrator. Require an explicit owner for reservations, access control, device allocation and failure recovery. This article does not claim that moving the same fragmented inventory to Kubernetes creates additional capacity; benchmark the intended workload and policy, not just the control-plane brand.

8. Deployment and acceptance checklist

The following gates are proposed tests, not reported passes. Establish local thresholds before commissioning. Use a maintenance window and a reversible configuration revision for any inventory or topology change.

Gate Test to perform Evidence to retain Pass condition
Inventory Compare physical, discovered and configured GPUs Discovery outputs and private asset mapping Every test node matches approved inventory
Basic allocation Run the one-node environment smoke test Job request, allocation and step output Requested tasks and device assignment agree
Isolation Run approved concurrent-job access checks Per-job device-access results Jobs cannot access devices outside policy
Fragmentation Reproduce mixed occupancy in a test partition Free GPUs, eligible nodes and job reason Operators can explain why the target shape fits or waits
Locality Compare single-group and cross-group requests Node list, topology configuration, wait time Placement follows documented policy
Policy Submit a controlled over-limit request QOS/account limit and reason Rejection or pending behavior matches policy
Recovery Remove one idle test node from eligibility State transition and placement recalculation New allocations exclude it until validated recovery
Workload Run a fixed, representative distributed test Versioned launcher, placement and performance logs Agreed baseline and recovery criteria are met

Do not use a universal GPU-utilization percentage as the sole acceptance threshold. Prefer workload progress, correctness, placement reproducibility and a documented allocation-to-first-useful-work interval. For the performance gate, reuse the deeper 256-GPU NCCL and DCGM acceptance plan rather than inventing a benchmark score here.

Retain an incident bundle containing the exact submission, timestamp, job record, node inventory, applicable limits, topology revision and application logs. Sanitize it before external sharing. A reproducible failing request plus an eligibility worksheet is more actionable than a screenshot of idle GPUs.

Practical takeaways

Start with the job shape. Separate free GPUs from eligible nodes, and eligible nodes from policy-approved placements. Reconcile physical devices with GRES before resuming drained nodes. Validate locality and task mapping independently, then measure the real application. The goal is not a greener utilization chart: it is a predictable path from request to correct, useful work.

Continue with the AI Infrastructure and Automation hub, Data Center Networking hub, or Start Here. The official sources below are rolling documentation retrieved on 2 October 2026; check release-matched manuals before applying syntax or relying on specific behavior.

Sources

  1. [1] https://slurm.schedmd.com/gres.html
  2. [2] https://slurm.schedmd.com/topology.html
  3. [3] https://slurm.schedmd.com/job_reason_codes.html

Comments

0 Responses to "Slurm Job Pending with Idle GPUs: A 256-GPU Sizing and Troubleshooting Checklist"

Post a Comment

Popular Posts