npx skills add ...
npx skills add nvidia/skills --skill tao-run-on-slurm
Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed
npx skills add nvidia/skills --skill tao-run-on-slurm
Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Remote GPU compute platform for clusters managed by SLURM. Jobs are submitted
from the TAO service or SDK host to a login node over SSH, staged on a shared
filesystem, submitted with sbatch, and executed with srun container support.
Use SLURM when the user has access to a managed GPU cluster, shared Lustre storage, and scheduler-owned GPU allocation. Do not use SLURM for local files that exist only on the agent machine; data and outputs must be reachable from the cluster.
Run all five steps in order before generating any launcher or submitting any job:
SLURM_USER and SLURM_HOSTNAME are set and
passwordless SSH to the login host works (ssh -o BatchMode=yes).nvidia-tao-sdk[slurm], on public PyPI).SLURM_ACCOUNT via
sacctmgr show associations before generating any scripts. If unset and
only one account exists for the user, auto-select it. If multiple exist, list
them and require the user to export one. Never submit a job without a verified
account — an invalid account causes "Invalid account or account/partition
combination" after SQSH conversion has already run.SLURM_PARTITION exists
via sinfo. The packaged default polar,polar3,polar4,grizzly is valid on
CS-OCI-ORD but may not exist on other clusters.nvcr.io images, install
~/.config/enroot/.credentials on the cluster once per (cluster, user).
Pyxis/Enroot does not read NGC_KEY from the job env; without persistent
credentials, auth-gated pulls fail with "Could not process JSON input". Use
the printf | ssh heredoc so the NGC_KEY value never lands in shell
history, intermediate files, or chat output; never cat/echo the value.If a preflight check fails, the agent prompts the user to authorize the install/fix via Bash. Pip-installable Python requirements are the exception: install them automatically, then rerun preflight.
See references/slurm-ssh-credentials.md for the full preflight script,
account/partition discovery commands, the enroot-credentials heredoc,
prerequisite key setup (keypair, ssh-copy-id, known_hosts, container key
mounts, 2FA handling), and the SSH failure remediation prompt.
Use shared-filesystem URIs, not local or file:// paths; tao-core rejects
local/file paths for remote backends.
lustre:///absolute/path for user-provided datasets on Lustre.slurm:// paths may appear in microservices metadata and are converted to
Lustre paths before the container starts.Accept either dataset roots (model skills map them to required files) or direct
spec-key paths. After SSH succeeds and before generating scripts, test -e each
required dataset path from the login host; if it fails, stop and ask for
corrected paths or staged data rather than producing scripts that fail in the
first training job. See references/slurm-ssh-credentials.md for root vs.
direct-spec modes, backend details, and the results-dir default.
tao-core runs TAO containers through Pyxis/Enroot:
<job_dir>/specs, <job_dir>/env, and <job_dir>/meta.srun -n1 -p <conversion_partition> enroot import. Do NOT use the cpu
partition for this step — cpu has a ~30 min wall-time limit that is
shorter than the conversion time for large TAO images (9+ layers, >30 min).
Use cpu_long (or another partition with ≥2 h limit) and set
SLURM_CONVERSION_PARTITION=cpu_long and
SLURM_CONVERSION_TIMEOUT_MINUTES=120 before constructing SlurmSDK.
The SDK validates the SQSH via SquashFS magic bytes before reusing it, so
partial files from failed conversions are automatically rejected and
reconverted — no manual cleanup needed. See the SQSH Conversion And
Caching section of references/slurm-container-execution.md for the full
env-knob table (SLURM_ENROOT_TEMP_PATH for xattr-restricted filesystems,
memory, force-reconvert), cache/dedup semantics, live-monitoring commands,
and manual pre-staging.<job_dir>/sbatch/job_<job_id>.sbatch.sbatch --export=ALL <script>.srun --container-image=<image> --container-mounts=/lustre.Accepted image formats: /path/to/image.sqsh, registry#image:tag,
docker://registry#image:tag, and ordinary registry/image:tag (converted to
Pyxis form when needed). SQSH conversion is cached by image name; for :latest
images the cached SQSH is reused unless force_reconvert_latest is enabled.
squeue/sacct;
TAO terminal status comes from status.json in the shared results folder.PENDING, RUNNING, or otherwise). Do not stop after a
fixed elapsed time such as 30 minutes; long queue waits are normal on shared
GPU partitions.<job_dir>/slurm-logs/<slurm_job_name>-<slurm_job_id>/main.out and .err.backend_details.slurm_metadata.slurm_job_id and running
scancel <slurm_job_id> over SSH. Treat missing or already terminated jobs as
successful cancellation.Status mapping:
PENDING -> PendingRUNNING or COMPLETING -> RunningCOMPLETED -> check status.jsonFAILED, BOOT_FAIL, DEADLINE, OUT_OF_MEMORY, NODE_FAIL -> retry if
logs match retriable infrastructure patterns, otherwise ErrorCANCELLED, PREEMPTED, REVOKED -> CanceledTIMEOUT -> ErrorSUSPENDED, STOPPED -> PausedAsk for these in the SLURM intake; see references/slurm-ssh-credentials.md
for the full credential list, microservices schema keys, and defaults.
polar,polar3,polar4,grizzly, treated as 4-hour queues.SSH_AUTH_SOCK agent-socket fallback./lustre/fsw/portfolios/edgeai/users/<your-dir> (your per-user Lustre dir).#SBATCH --account. Auto-discovered via sacctmgr during preflight step 3;
only ask the user if multiple accounts are found.Do not ask for SLURM_BASE_RESULTS_DIR in the initial intake unless the user
wants a custom results root.
Defaults from tao-core:
num_nodes: 1num_gpus: 4max_num_gpus_per_node: 8cpus_per_task: 16time_hours: 4timeout_hours: 3.8max_time_hours: 4container_mounts: /lustreuse_requeue: trueuse_sqsh: trueWhen generating launchers or wrapper scripts for SLURM, set the wall-time defaults explicitly from the packaged platform resource defaults:
Do not default to 12 hours on SLURM. If the user supplies a longer
SLURM_TIME_HOURS, verify that the selected partition supports it before
submitting. For the packaged default partition list
polar,polar3,polar4,grizzly, reject requests above 4 hours and ask for a
different partition only if the user actually wants a longer wall time.
When num_gpus is greater than or equal to max_num_gpus_per_node, the
handler treats the request as exclusive per node and computes additional nodes
from total GPU count when necessary.
For multi-node jobs (num_nodes > 1), the SDK builds the sbatch directives and
exports the PyTorch-distributed rendezvous env vars automatically: WORLD_SIZE,
NUM_GPU_PER_NODE, NODE_RANK, MASTER_ADDR, and MASTER_PORT (29500).
TAO entrypoints read WORLD_SIZE + NUM_GPU_PER_NODE and build torchrun
internally. Cosmos-RL has special multi-node role handling for controller,
policy, and rollout workers.
Use Lustre, not S3, for SLURM job inputs. The GPU allocation starts the
moment the job is dispatched, so a long s3:// download at the top of the
script burns the allocation, can get the job killed for GPU-idle, and is billed
either way. Stage training data on the shared filesystem first and reference it
as lustre:///.... S3/HF/NGC pre-fetch is fine for small auxiliary inputs
(checkpoints, configs), not training datasets. K8s/Brev do not share this
scheduler-idle constraint.
Auto-retry of infrastructure failures (NODE_FAIL, BOOT_FAIL, NCCL transport
timeouts, CUDA driver init failures, GPU/IB link-down, OOM-killer node reaping,
Xid errors) is automatic in the SDK, with a stable user-facing Job.id across
retries. Plain training failures surface immediately so a broken spec does not
consume the retry budget. #SBATCH --requeue is enabled by default via
SLURM_USE_REQUEUE=true.
See references/slurm-container-execution.md for the full multi-node
env-var/sbatch directive detail and table, cluster requirements, the optional
TAO SDK path (SlurmSDK, build_entrypoint, ActionWorkflow) with code, the
Lustre-not-S3 rule in full, and the failure-mode checklist;
references/slurm-execution-sdk.md covers the MAX_JOB_RETRIES retry budget.
When the SDK is in scope, read tao-skill-bank:tao-run-platform for the
SlurmSDK kwarg reference.
references/slurm-ssh-credentials.md — preflight script, SSH/key setup,
enroot credentials, full credential list, backend details, storage rules,
SSH remediation prompt.references/slurm-container-execution.md — container execution steps,
monitoring, status mapping, cancellation, multi-node detail, SDK use,
Lustre-not-S3, auto-retry, failure modes.references/slurm-preflight-storage.md — extended preflight/storage notes.references/slurm-execution-sdk.md — extended execution/SDK notes.references/detailed-guide.md — navigation map for the split references.