Megatron-Bridge
29 MCP servers and agent skills in the Megatron-Bridge category, ranked by quality score — 29 results.
NVIDIA/Megatron-Bridge/adding-model-support
Guide for adding support for new LLM or VLM models in Megatron-Bridge.
NVIDIA/Megatron-Bridge/build-and-dependency
Dev environment setup for Megatron Bridge — container-based development, uv package management, lockfile regeneration, adding dependencies, Slurm container usage, and common build...
NVIDIA/Megatron-Bridge/bump-dependency
Bump a pinned dependency (TransformerEngine, Megatron-LM, NRX, etc.), regenerate the lockfile, open a PR, and drive it to green by attaching a watchdog to the "CICD NeMo" workflow...
NVIDIA/Megatron-Bridge/cicd
CI/CD reference for Megatron Bridge — pipeline structure, commit and PR workflow, CI failure investigation, and common failure patterns.
NVIDIA/Megatron-Bridge/linting-and-formatting
Code style and quality rules for Megatron Bridge — ruff configuration, naming conventions, type hints, mypy rules, docstrings, copyright headers, logging, and the code review check...
NVIDIA/Megatron-Bridge/mlm-bridge-training
Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data.
NVIDIA/Megatron-Bridge/multi-node-slurm
Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.
NVIDIA/Megatron-Bridge/nemo-rl-e2e-testing
External NeMo-RL end-to-end validation workflow for Megatron-Bridge model/provider changes, including downstream compatibility checks, external RL lifecycle behavior, Megatron poli...
NVIDIA/Megatron-Bridge/parity-testing
Structured framework for verifying numerical parity of HFMCore weight conversions.
NVIDIA/Megatron-Bridge/perf-activation-recompute
Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.
NVIDIA/Megatron-Bridge/perf-cpu-offloading
Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.
NVIDIA/Megatron-Bridge/perf-cuda-graphs
Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
NVIDIA/Megatron-Bridge/perf-expert-parallel-overlap
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as...
NVIDIA/Megatron-Bridge/perf-hierarchical-context-parallel
Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
NVIDIA/Megatron-Bridge/perf-megatron-fsdp
Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
NVIDIA/Megatron-Bridge/perf-memory-tuning
Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.
NVIDIA/Megatron-Bridge/perf-moe-comm-overlap
MoE expert-parallel communication overlap in Megatron Bridge.
NVIDIA/Megatron-Bridge/perf-moe-dispatcher-selection
Choose the right MoE token dispatcher (`alltoall`, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage.
NVIDIA/Megatron-Bridge/perf-moe-hardware-configs
Representative MoE training playbooks by hardware platform and model family.
NVIDIA/Megatron-Bridge/perf-moe-long-context
Long-context MoE training guidance for Megatron Bridge.
NVIDIA/Megatron-Bridge/perf-moe-optimization-workflow
Systematic workflow for MoE training optimization in Megatron Bridge, based on the Megatron-Core MoE paper.
NVIDIA/Megatron-Bridge/perf-moe-vlm-training
Practical guidance for training MoE VLMs in Megatron Bridge.
NVIDIA/Megatron-Bridge/perf-parallelism-strategies
Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.
NVIDIA/Megatron-Bridge/perf-sequence-packing
Validate and use packed sequences and long-context training in Megatron-Bridge, distinguishing offline packed SFT for LLMs from in-batch packing for VLMs, and applying the right CP...
NVIDIA/Megatron-Bridge/perf-tp-dp-comm-overlap
Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
NVIDIA/Megatron-Bridge/recipe-recommender
Recommend and customize Megatron Bridge recipes for a user's model, GPU count, and training goal.
NVIDIA/Megatron-Bridge/resiliency
Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.
NVIDIA/Megatron-Bridge/testing
Testing reference for Megatron Bridge — unit and functional test layout, tier semantics (L0/L1/L2/flaky), script conventions, running tests locally, adding/moving/disabling tests,...
NVIDIA/Megatron-Bridge/verl-e2e-testing
External verl end-to-end validation workflow for Megatron-Bridge model/provider changes.