NVIDIA/Megatron-Bridge/resiliency A Agent Skill Active
Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.
Browse SKILL.md-based agent skills for Claude Code, Codex, Cursor, Gemini CLI and more.
Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.
Testing reference for Megatron Bridge — unit and functional test layout, tier semantics (L0/L1/L2/flaky), script conventions, running tests locally, adding/moving/disabling tests,...
External verl end-to-end validation workflow for Megatron-Bridge model/provider changes.
Container-based dev environment setup and dependency management for Megatron-LM.
Bump the NVIDIA PyTorch base image (`nvcr.io/nvidia/pytorch:-py3`) used by Megatron-LM CI.
CI/CD reference for Megatron-LM.
Investigate a failing GitHub Actions run or job and create a GitHub issue for the failure.
Linting and formatting for Megatron-LM.
Domain knowledge for the nightly main-to-dev sync workflow.
Onboard 1-node GitHub MR functional tests for GB200 from existing mr-scoped 2-node tests.
Research and draft a response to a GitHub issue or question from an external contributor.
How to launch distributed Megatron-LM training jobs on a SLURM cluster.
Split a PR into multiple PRs to reduce the number of required CODEOWNERS reviewer groups.
Test system for Megatron-LM.
Refresh golden values from a GitHub Actions workflow run (failing-only or all jobs), score the change with average normalized relative differences, and produce a PR-ready summary.
Query and browse evaluation results stored in MLflow.
Run commands inside a remote Docker container via the file-based command relay (tools/debugger).
Serve a quantized or unquantized LLM checkpoint as an OpenAI-compatible API endpoint using vLLM, SGLang, or TRT-LLM.
Evaluates accuracy of quantized or unquantized LLMs using NeMo Evaluator Launcher (NEL).
Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher.
Monitor submitted jobs (PTQ, evaluation, deployment) on SLURM clusters.
This skill should be used when the user asks to "quantize a model", "run PTQ", "post-training quantization", "NVFP4 quantization", "FP8 quantization", "INT8 quantization", "INT4 AW...
Cherry-pick merged PRs labeled for a release branch into that branch, then open a PR and apply the cherry-pick-done label.
Create custom LLM evaluation benchmarks using the BYOB decorator framework.
Query and browse evaluation results stored in MLflow.
Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher.
Interactive config wizard for NeMo Evaluator Launcher (NEL).
Guide for adding a new benchmark or training environment to NeMo-Gym.
Use when debugging a Nemo Gym run or reward profiling job.
Maintain the NeMo Gym Fern docs site — add, update, move, or remove pages under fern/.
Use when creating, validating, or documenting Nemo Gym pivot datasets from rollout, trajectory, chat-completion, Responses API, or tool-call artifacts.
Use to help users get started with Nemo Gym reward profiling.
Autonomous NeMo-RL research agent workflow for directed hypothesis testing and open-ended discovery.
Brev instance operating guidance for NeMo-RL agents working in /home/ubuntu/RL with limited workspace disk, a larger /ephemeral volume, and optional /home/ubuntu/RL/.env secrets.
Build and dependency management for NeMo-RL.
CI/CD reference for NeMo-RL.
Configuration conventions for NeMo-RL.
Contribution conventions for NeMo-RL.
NVIDIA copyright header requirements for NeMo-RL.
Documentation conventions for NeMo-RL.
Error handling guidelines for NeMo-RL.
Playbook for launching, monitoring, stopping, and debugging NeMo-RL recipes on a Kubernetes cluster via the nrl-k8s CLI.
Code style guidelines for NeMo-RL (Python and shell).
Interactive code review for NVIDIA-NeMo/RL pull requests.
Manage durable working-session memory for coding agents.
Testing conventions for NeMo-RL.
Create GitHub pull requests that follow the NemoClaw PR template.
Scan recent git commits for changes that affect user-facing behavior, then draft or update the corresponding documentation pages and refresh generated user skills for release prep.