Skill-Details
mcore-testing
Testing guidance limited to Megatron-LM.
Vor Nutzung prüfen
Die automatische Prüfung bewertet Relevanz, nicht Sicherheit oder Empfehlung. Lies vor der Nutzung die Quellanweisungen.
SKILL.md
Dieser Auszug wurde bei der Prüfung gespeichert. Die externe Quelle enthält die vollständige und aktuelle Version.
--- name: mcore-testing description: Test system for Megatron-LM. Covers test layout, recipe YAML structure, adding and running unit and functional tests, golden values, marker filters, and CI parity. license: Apache-2.0 when_to_use: Adding or running a unit or functional test; understanding the test layout; writing a recipe YAML; downloading or updating golden values; reproducing a test failure locally; 'how do I add a test', 'run unit tests', 'pytest fails', 'test layout', 'golden values', 'recipe YAML', 'marker filter'. metadata: author: Philip Petrakian <[email protected]> --- # Testing Guide --- ## Answer-First Testing Facts For questions about disabling tests without deleting them: - Functional recipe entries stay in YAML; disable by suffixing scope with `-broken`, for example `scope: [mr-github]` -> `scope: [mr-github-broken]`. - Unit-test skips use pytest markers instead: `@pytest.mark.flaky_in_dev` skips in the default dev environment, and `@pytest.mark.flaky` skips in LTS. - Do not delete the test case or recipe entry when the goal is discoverability and easy re-enable. --- ## Test Layout ```text tests/ ├── unit_tests/ # pytest, 1 node × 8 GPUs, torch.distributed runner ├── functional_tests/ # end-to-end shell + training scripts │ └── test_cases/ │ └── {model}/{test_case}/ │ ├── model_config.yaml # training args │ └── golden_values_{env}_{platform}.json └── test_utils/ ├── recipes/ │ ├── h100/ # YAML recipes for H100 jobs │ └── gb200/ # YAML recipes for GB200 jobs └── python_scripts/ # helpers (recipe_parser, golden-value download, …) ``` --- ## How Tests Execute The GitHub Actions runner invokes `launch_nemo_run_workload.py`, which uses **nemo-run** to launch a `DockerExecutor` container. The repo is bind-mounted at `/opt/megatron-lm`; training data is mounted at `/mnt/artifacts`. **Unit tests** are dispatched through `torch.distributed.run`: - Ranks 0 and 3 are tee-d to stdout; all other ranks write only to log files. - Per-rank log files land at `{assets_dir}/logs/1/` and are uploaded as a GitHub artifact after the run. **Functional tests** are driven by `tests/functional_tests/shell_test_utils/run_ci_test.sh`. Only rank 0 runs the pytest validation step; training output from all ranks is uploaded as an artifact. **Flaky-failure auto-retry**: `launch_nemo_run_workload.py` retries up to **3 times** for known transient patterns (NCCL timeout, ECC error, segfault, HuggingFace connectivity, …) before declaring a genuine failure. --- ## Recipe YAML Structure Recipes live in `tests/test_utils/recipes/` and are parsed by `tests/test_utils/python_scripts/recipe_parser.py`. Each file expands a cartesian `products` block into individual workload specs: ```yaml type: basic format_version: 1 maintainers: [mcore] loggers: [stdout] spec: name: "{test_case}_{environment}_{platforms}" model: gpt # maps to tests/functional_tests/test_cases/{model}/ build: mcore-pyt-{environment} nodes: 1 gpus: 8 n_repeat: 5 platforms: dgx_h100 time_limit: 1800 script_setup: | ... script: |- bash tests/functional_tests/shell_test_utils/run_ci_test.sh ... products: - test_case: [my_test] products: - environment: [dev, lts] scope: [mr-github] platforms: [dgx_h100] ``` Key runtime placeholders: `{assets_dir}`, `{artifacts_dir}`, `{test_case}`, `{environment}`, `{platforms}`, `{n_repeat}`. ### Disabling a Test Without Deleting It To temporarily disable a test case in a recipe YAML, suffix its `scope` value with `-broken` — **do not delete the entry**: ```yaml # before (test runs in CI) scope: [mr-github] # after (test is skipped; entry preserved for easy re-enable) scope: [mr-github-broken] ``` --- ## Running Unit Tests Locally All unit tests initialize a `torch.distributed` group, so every invocation requires GPU access and must go through `torch.distributed.runVollständige Quelle auf GitHub lesen (öffnet externe Seite)