Home › Search: r

Search: r

Search MCP servers and agent skills by name, description, category or topic — 4,174 results.

Agent Skill Active

NVIDIA/Megatron-Bridge/bump-dependency

Bump a pinned dependency (TransformerEngine, Megatron-LM, NRX, etc.), regenerate the lockfile, open a PR, and drive it to green by attaching a watchdog to the "CICD NeMo" workflow...

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/cicd

CI/CD reference for Megatron Bridge — pipeline structure, commit and PR workflow, CI failure investigation, and common failure patterns.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/linting-and-formatting

Code style and quality rules for Megatron Bridge — ruff configuration, naming conventions, type hints, mypy rules, docstrings, copyright headers, logging, and the code review check...

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/mlm-bridge-training

Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/multi-node-slurm

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/nemo-rl-e2e-testing

External NeMo-RL end-to-end validation workflow for Megatron-Bridge model/provider changes, including downstream compatibility checks, external RL lifecycle behavior, Megatron poli...

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/parity-testing

Structured framework for verifying numerical parity of HFMCore weight conversions.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-activation-recompute

Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-cpu-offloading

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-cuda-graphs

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-expert-parallel-overlap

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as...

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-hierarchical-context-parallel

Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-megatron-fsdp

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-memory-tuning

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-moe-comm-overlap

MoE expert-parallel communication overlap in Megatron Bridge.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-moe-dispatcher-selection

Choose the right MoE token dispatcher (`alltoall`, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-moe-hardware-configs

Representative MoE training playbooks by hardware platform and model family.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-moe-long-context

Long-context MoE training guidance for Megatron Bridge.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-moe-optimization-workflow

Systematic workflow for MoE training optimization in Megatron Bridge, based on the Megatron-Core MoE paper.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-moe-vlm-training

Practical guidance for training MoE VLMs in Megatron Bridge.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-parallelism-strategies

Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-sequence-packing

Validate and use packed sequences and long-context training in Megatron-Bridge, distinguishing offline packed SFT for LLMs from in-batch packing for VLMs, and applying the right CP...

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/perf-tp-dp-comm-overlap

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/recipe-recommender

Recommend and customize Megatron Bridge recipes for a user's model, GPU count, and training goal.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/resiliency

Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/testing

Testing reference for Megatron Bridge — unit and functional test layout, tier semantics (L0/L1/L2/flaky), script conventions, running tests locally, adding/moving/disabling tests,...

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Bridge/verl-e2e-testing

External verl end-to-end validation workflow for Megatron-Bridge model/provider changes.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/build-and-dependency

Container-based dev environment setup and dependency management for Megatron-LM.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/bump-base-image

Bump the NVIDIA PyTorch base image (`nvcr.io/nvidia/pytorch:-py3`) used by Megatron-LM CI.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/cicd

CI/CD reference for Megatron-LM.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/create-issue

Investigate a failing GitHub Actions run or job and create a GitHub issue for the failure.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/linting-and-formatting

Linting and formatting for Megatron-LM.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/nightly-sync

Domain knowledge for the nightly main-to-dev sync workflow.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/onboard-gb200-1node-tests

Onboard 1-node GitHub MR functional tests for GB200 from existing mr-scoped 2-node tests.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/respond-to-issue

Research and draft a response to a GitHub issue or question from an external contributor.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/run-on-slurm

How to launch distributed Megatron-LM training jobs on a SLURM cluster.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/split-pr

Split a PR into multiple PRs to reduce the number of required CODEOWNERS reviewer groups.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/testing

Test system for Megatron-LM.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Megatron-Core/update-golden-values

Refresh golden values from a GitHub Actions workflow run (failing-only or all jobs), score the change with average normalized relative differences, and produce a PR-ready summary.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/accessing-mlflow

Query and browse evaluation results stored in MLflow.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/debug

Run commands inside a remote Docker container via the file-based command relay (tools/debugger).

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/deployment

Serve a quantized or unquantized LLM checkpoint as an OpenAI-compatible API endpoint using vLLM, SGLang, or TRT-LLM.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/evaluation

Evaluates accuracy of quantized or unquantized LLMs using NeMo Evaluator Launcher (NEL).

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/launching-evals

Run, monitor, analyze, and debug LLM evaluations via nemo-evaluator-launcher.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/monitor

Monitor submitted jobs (PTQ, evaluation, deployment) on SLURM clusters.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/ptq

This skill should be used when the user asks to "quantize a model", "run PTQ", "post-training quantization", "NVFP4 quantization", "FP8 quantization", "INT8 quantization", "INT4 AW...

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/Model-Optimizer/release-cherry-pick

Cherry-pick merged PRs labeled for a release branch into that branch, then open a PR and apply the cherry-pick-done label.

2.8k Python Updated today Score 84
Agent Skill Active

NVIDIA/NeMo-Evaluator/byob

Create custom LLM evaluation benchmarks using the BYOB decorator framework.

2.8k Python Updated today Score 84
‹ Prev1891087Next ›