npx skills add ...
npx skills add nvidia/megatron-lm --skill mcore-testing
Test system for Megatron-LM. Covers test layout, recipe YAML structure, adding and running unit and functional tests, golden values, marker filters, and CI parity.
npx skills add nvidia/megatron-lm --skill mcore-testing
For questions about disabling tests without deleting them:
-broken, for example scope: [mr-github] -> scope: [mr-github-broken].@pytest.mark.flaky_in_dev skips
in the default dev environment, and @pytest.mark.flaky skips in LTS.The GitHub Actions runner invokes launch_nemo_run_workload.py, which uses
nemo-run to launch a DockerExecutor container. The repo is bind-mounted
at /opt/megatron-lm; training data is mounted at /mnt/artifacts.
Unit tests are dispatched through torch.distributed.run:
{assets_dir}/logs/1/ and are uploaded as a
GitHub artifact after the run.Functional tests are driven by
tests/functional_tests/shell_test_utils/run_ci_test.sh. Only rank 0 runs the
pytest validation step; training output from all ranks is uploaded as an artifact.
Flaky-failure auto-retry: launch_nemo_run_workload.py retries up to
3 times for known transient patterns (NCCL timeout, ECC error, segfault,
HuggingFace connectivity, …) before declaring a genuine failure.
Recipes live in tests/test_utils/recipes/ and are parsed by
tests/test_utils/python_scripts/recipe_parser.py. Each file expands a
cartesian products block into individual workload specs:
Key runtime placeholders: {assets_dir}, {artifacts_dir}, {test_case},
{environment}, {platforms}, {n_repeat}.
To temporarily disable a test case in a recipe YAML, suffix its scope value
with -broken — do not delete the entry:
All unit tests initialize a torch.distributed group, so every invocation
requires GPU access and must go through torch.distributed.run:
Use tests/unit_tests/run_ci_test.sh to reproduce a CI bucket failure exactly.
For ad-hoc runs, prefer the direct torch.distributed.run invocations above.
pyproject.toml sets addopts = --durations=15 -s -rA — stdout is not
captured (-s), so ranks interleave during multi-rank runs. Override with
--capture=fd when debugging a specific rank.tests/unit_tests/conftest.py looks for test data under /opt/data and
attempts a download if missing. Supply it manually or skip data-dependent
tests when running outside the canonical container.tests/unit_tests/<category>/test_<name>.py.tests/unit_tests/conftest.py.@pytest.mark.internal — skipped on legacy tag@pytest.mark.flaky_in_dev — skipped in dev environment (CI default; use this to disable a flaky test without blocking the standard pipeline)@pytest.mark.flaky — skipped in lts environment@pytest.mark.experimental — latest tag onlytests/test_utils/recipes/h100/unit-tests.yaml.jit_fuser /
torch.compile, CUDA extension, TE or external-library dispatch, or a
scatter/index accumulation), add or update its bit-exact replay test under
tests/unit_tests/determinism/kernels/ and register it in
tests/unit_tests/determinism/kernels/manifest.py. The linting CI job
(tools/check_kernel_determinism_coverage.py) fails kernel PRs without
this; the determinism-exempt label overrides it for non-numeric edits.
See docs/developer/determinism/testing.md.Create tests/functional_tests/test_cases/<model>/<test_name>/.
Write model_config.yaml with MODEL_ARGS, ENV_VARS, and TEST_TYPE.
Add a YAML recipe under tests/test_utils/recipes/h100/ (and gb200/ if
needed). Required fields: scope, environment, platform, n_repeat,
time_limit.
Push the PR, add the label "Run functional tests" to trigger a full run.
After a successful run, download golden values:
Commit the downloaded golden values.
Golden values keep the full float32 precision of the TensorBoard scalars,
record it as "value_precision": "full", and deterministic test cases are
compared bit-exactly against them. Never round or hand-edit them:
tools/check_golden_values.py (run by the linting CI job on changed golden
files) rejects golden files of deterministically compared cases whose metrics
lack the full marker. Legacy files (no marker) are still compared at five
decimals until they are regenerated from a CI run.
| Problem | Cause | Fix |
|---|---|---|
| Test passes locally but fails in CI | Different environment or data path | Check DATA_PATH, DATA_CACHE_PATH, and the environment tag (dev vs lts) |
| Golden value mismatch after a code change | Numerical regression | Download new golden values via download_golden_values.py after a clean run |
cicd-integration-tests-gb200 not triggered | GB200 jobs require maintainer status | Ask a maintainer to trigger, or add the Run functional tests label |
type: basic
format_version: 1
maintainers: [mcore]
loggers: [stdout]
spec:
name: "{test_case}_{environment}_{platforms}"
model: gpt # maps to tests/functional_tests/test_cases/{model}/
build: mcore-pyt-{environment}
nodes: 1
gpus: 8
n_repeat: 5
platforms: dgx_h100
time_limit: 1800
script_setup: |
...
script: |-
bash tests/functional_tests/shell_test_utils/run_ci_test.sh ...
products:
- test_case: [my_test]
products:
- environment: [dev, lts]
scope: [mr-github]
platforms: [dgx_h100]# before (test runs in CI)
scope: [mr-github]
# after (test is skipped; entry preserved for easy re-enable)
scope: [mr-github-broken]# Full suite
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests
# Single file
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests/models/test_gpt_model.py
# Single test
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests/models/test_gpt_model.py::TestGPTModel::test_constructor
# Filter by name substring
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests -k optimizer# Exclude flaky tests during development
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests -m "not flaky and not flaky_in_dev"
# Include experimental tests
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests --experimental