npx skills add ...
npx skills add nvidia/skills --skill mcore-testing
Test system for Megatron-LM. Covers test layout, recipe YAML structure, adding and running unit and functional tests, golden values, marker filters, and CI parity.
npx skills add nvidia/skills --skill mcore-testing
For questions about disabling tests without deleting them:
-broken, for example scope: [mr-github] -> scope: [mr-github-broken].@pytest.mark.flaky_in_dev skips
in the default dev environment, and @pytest.mark.flaky skips in LTS.The GitHub Actions runner invokes launch_nemo_run_workload.py, which uses
nemo-run to launch a DockerExecutor container. The repo is bind-mounted
at /opt/megatron-lm; training data is mounted at /mnt/artifacts.
Unit tests are dispatched through torch.distributed.run:
{assets_dir}/logs/1/ and are uploaded as a
GitHub artifact after the run.Functional tests are driven by
tests/functional_tests/shell_test_utils/run_ci_test.sh. Only rank 0 runs the
pytest validation step; training output from all ranks is uploaded as an artifact.
Flaky-failure auto-retry: launch_nemo_run_workload.py retries up to
3 times for known transient patterns (NCCL timeout, ECC error, segfault,
HuggingFace connectivity, …) before declaring a genuine failure.
Recipes live in tests/test_utils/recipes/ and are parsed by
tests/test_utils/python_scripts/recipe_parser.py. Each file expands a
cartesian products block into individual workload specs:
Key runtime placeholders: {assets_dir}, {artifacts_dir}, {test_case},
{environment}, {platforms}, {n_repeat}.
To temporarily disable a test case in a recipe YAML, suffix its scope value
with -broken — do not delete the entry:
All unit tests initialize a torch.distributed group, so every invocation
requires GPU access and must go through torch.distributed.run:
Use tests/unit_tests/run_ci_test.sh to reproduce a CI bucket failure exactly.
For ad-hoc runs, prefer the direct torch.distributed.run invocations above.
pyproject.toml sets addopts = --durations=15 -s -rA — stdout is not
captured (-s), so ranks interleave during multi-rank runs. Override with
--capture=fd when debugging a specific rank.tests/unit_tests/conftest.py looks for test data under /opt/data and
attempts a download if missing. Supply it manually or skip data-dependent
tests when running outside the canonical container.tests/unit_tests/<category>/test_<name>.py.tests/unit_tests/conftest.py.@pytest.mark.internal — skipped on legacy tag@pytest.mark.flaky_in_dev — skipped in dev environment (CI default; use this to disable a flaky test without blocking the standard pipeline)@pytest.mark.flaky — skipped in lts environment@pytest.mark.experimental — latest tag onlytests/test_utils/recipes/h100/unit-tests.yaml.Create tests/functional_tests/test_cases/<model>/<test_name>/.
Write model_config.yaml with MODEL_ARGS, ENV_VARS, and TEST_TYPE.
Add a YAML recipe under tests/test_utils/recipes/h100/ (and gb200/ if
needed). Required fields: scope, environment, platform, n_repeat,
time_limit.
Push the PR, add the label "Run functional tests" to trigger a full run.
After a successful run, download golden values:
Commit the downloaded golden values.
| Problem | Cause | Fix |
|---|---|---|
| Test passes locally but fails in CI | Different environment or data path | Check DATA_PATH, DATA_CACHE_PATH, and the environment tag (dev vs lts) |
| Golden value mismatch after a code change | Numerical regression | Download new golden values via download_golden_values.py after a clean run |
cicd-integration-tests-gb200 not triggered | GB200 jobs require maintainer status | Ask a maintainer to trigger, or add the Run functional tests label |
type: basic
format_version: 1
maintainers: [mcore]
loggers: [stdout]
spec:
name: "{test_case}_{environment}_{platforms}"
model: gpt # maps to tests/functional_tests/test_cases/{model}/
build: mcore-pyt-{environment}
nodes: 1
gpus: 8
n_repeat: 5
platforms: dgx_h100
time_limit: 1800
script_setup: |
...
script: |-
bash tests/functional_tests/shell_test_utils/run_ci_test.sh ...
products:
- test_case: [my_test]
products:
- environment: [dev, lts]
scope: [mr-github]
platforms: [dgx_h100]# before (test runs in CI)
scope: [mr-github]
# after (test is skipped; entry preserved for easy re-enable)
scope: [mr-github-broken]# Full suite
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests
# Single file
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests/models/test_gpt_model.py
# Single test
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests/models/test_gpt_model.py::TestGPTModel::test_constructor
# Filter by name substring
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests -k optimizer# Exclude flaky tests during development
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests -m "not flaky and not flaky_in_dev"
# Include experimental tests
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
tests/unit_tests --experimental