npx skills add ...
npx skills add nvidia/nvidia-resiliency-ext --skill fr-analysis
Analyze PyTorch NCCL flight-recorder (FR) dumps to identify collective operation hangs and isolate the responsible ranks using CollectiveAnalyzer. Use when a distributed training job hangs due to an NCCL collective timeout and FR dump files are available. Detects the wavefront process group where collectives diverge and returns the root-cause suspect ranks.
npx skills add nvidia/nvidia-resiliency-ext --skill fr-analysis
Analyze PyTorch NCCL flight-recorder (FR) dumps to identify the collective operation hang
and isolate the ranks responsible, using CollectiveAnalyzer.
Script: scripts/fr_attribution.py → attribution/trace_analyzer/fr_attribution.py
--fr-path.Collective records (op type, ranks, process group, timing, state).| Flag | Default | Description |
|---|---|---|
--fr-path | required | Path to a directory (or single file) containing FR dump files |
--pattern, -p | _dump_* | Glob pattern for dump files within --fr-path |
--verbose, -v | off | Print detailed per-rank collective tables |
--health-check, -c | off | Include node health check results in output |
--debug | off | Convert binary trace files to JSON for inspection |
Returns (result, AttributionState) where result is the FR analysis table and describes:
--verbose)--health-check)AttributionState.STOP indicates the hang is unrecoverable; CONTINUE indicates the job
may be restartable after isolating the identified ranks.
| Format | Notes |
|---|---|
_dump_* files | PyTorch FR dump prefix pattern used by the feedback loop |
| Binary pickle / JSON payloads | Detected automatically; use --debug to convert binary traces to JSON |
FR dumps are typically written to the directory specified by TORCH_NCCL_DEBUG_INFO_TEMP_FILE
or triggered automatically on NCCL timeout.
TORCH_NCCL_TRACE_BUFFER_SIZE > 0)FR_DEBUG=1 env var enables verbose debug logging in the script