Search: c
Search MCP servers and agent skills by name, description, category or topic — 4,153 results.
NVIDIA/Megatron-Bridge/parity-testing
Structured framework for verifying numerical parity of HFMCore weight conversions.
NVIDIA/Megatron-Bridge/perf-expert-parallel-overlap
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as...
NVIDIA/Megatron-Bridge/perf-megatron-fsdp
Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
NVIDIA/Megatron-Bridge/perf-memory-tuning
Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.
NVIDIA/Megatron-Bridge/perf-moe-optimization-workflow
Systematic workflow for MoE training optimization in Megatron Bridge, based on the Megatron-Core MoE paper.
NVIDIA/Megatron-Bridge/perf-moe-vlm-training
Practical guidance for training MoE VLMs in Megatron Bridge.
NVIDIA/Megatron-Bridge/perf-parallelism-strategies
Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.
NVIDIA/Megatron-Bridge/testing
Testing reference for Megatron Bridge — unit and functional test layout, tier semantics (L0/L1/L2/flaky), script conventions, running tests locally, adding/moving/disabling tests,...
NVIDIA/Megatron-Bridge/verl-e2e-testing
External verl end-to-end validation workflow for Megatron-Bridge model/provider changes.
NVIDIA/Model-Optimizer/debug
Run commands inside a remote Docker container via the file-based command relay (tools/debugger).
NVIDIA/Model-Optimizer/deployment
Serve a quantized or unquantized LLM checkpoint as an OpenAI-compatible API endpoint using vLLM, SGLang, or TRT-LLM.
NVIDIA/Model-Optimizer/evaluation
Evaluates accuracy of quantized or unquantized LLMs using NeMo Evaluator Launcher (NEL).
NVIDIA/Model-Optimizer/monitor
Monitor submitted jobs (PTQ, evaluation, deployment) on SLURM clusters.
NVIDIA/Model-Optimizer/ptq
This skill should be used when the user asks to "quantize a model", "run PTQ", "post-training quantization", "NVFP4 quantization", "FP8 quantization", "INT8 quantization", "INT4 AW...
NVIDIA/NeMo-Evaluator/byob
Create custom LLM evaluation benchmarks using the BYOB decorator framework.
NVIDIA/NeMo-Gym/nemo-gym-debugging
Use when debugging a Nemo Gym run or reward profiling job.
NVIDIA/NeMo-Gym/nemo-gym-pivot-datasets
Use when creating, validating, or documenting Nemo Gym pivot datasets from rollout, trajectory, chat-completion, Responses API, or tool-call artifacts.
NVIDIA/NeMo-Gym/nemo-gym-reward-profiling
Use to help users get started with Nemo Gym reward profiling.
NVIDIA/NeMo-RL/brev-etiquette
Brev instance operating guidance for NeMo-RL agents working in /home/ubuntu/RL with limited workspace disk, a larger /ephemeral volume, and optional /home/ubuntu/RL/.env secrets.
NVIDIA/NeMo-RL/linting-and-formatting
Code style guidelines for NeMo-RL (Python and shell).
NVIDIA/NeMo-RL/review-pr
Interactive code review for NVIDIA-NeMo/RL pull requests.
NVIDIA/NeMo-RL/session-memory
Manage durable working-session memory for coding agents.
NVIDIA/TensorRT-LLM/ad-add-fusion-transformation
Claude Code skill (trtllm-agent-toolkit): implement or extend TensorRT-LLM AutoDeploy fusion transforms under transform/library/ in a TensorRT-LLM checkout.
NVIDIA/TensorRT-LLM/ad-graph-dump
Enable and interpret TensorRT-LLM AutoDeploy FX graph text dumps via AD_DUMP_GRAPHS_DIR.
NVIDIA/TensorRT-LLM/ad-layer-visualizer
Visualize a specific transformer decoder layer from an AutoDeploy FX graph text dump as a hierarchical DOT/PNG diagram.
NVIDIA/TensorRT-LLM/ad-model-onboard
Translates a HuggingFace model into a prefill-only AutoDeploy custom model using reference custom ops, validates with hierarchical equivalence tests.
NVIDIA/TensorRT-LLM/kernel-tileir-optimization
Optimize existing Triton kernels for NVIDIA TileIR backend on Blackwell GPUs (sm_100+).
NVIDIA/TensorRT-LLM/kernel-triton-writing
ONLY for OpenAI Triton (@triton.jit) kernel development.
NVIDIA/TensorRT-LLM/perf-host-analysis
Analyze host/CPU overhead in TensorRT-LLM inference from nsys traces.
NVIDIA/TensorRT-LLM/perf-host-optimization
Profiles and optimizes TensorRT-LLM host/CPU overhead using line_profiler (with nsys support planned).
NVIDIA/TensorRT-LLM/perf-nsight-systems
Nsight Systems (nsys) CLI for system-level timeline profiling.
NVIDIA/TensorRT-LLM/perf-optimization
Performance optimization coordination playbook.
NVIDIA/TensorRT-LLM/perf-workload-profiling
Code instrumentation for timing workloads.
NVIDIA/TensorRT-LLM/trtllm-flashinfer-upgrade
Upgrade flashinfer-python version in TensorRT-LLM.
NVIDIA/TensorRT-LLM/trtllm-moe-develop
Review, design, and refactor TensorRT-LLM PyTorch MoE code for architecture fit, clean code, maintainability, and testability.
NVIDIA/deepstream/deepstream-dev
NVIDIA DeepStream SDK 9.0 development with Python pyservicemaker API.
NVIDIA/deepstream/deepstream-import-vision-model
Use this skill to bring any vision model from HuggingFace or NVIDIA NGC into an NVIDIA DeepStream pipeline with end-to-end automation: ONNX download, SafeTensors export, TRT engi...
NVIDIA/rag/rag-blueprint
"NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage.
foryourhealth111-pixel/Vibe-Skills
A skills governed plug-and-play harness for staged, test-driven skill orchestration
aaron-he-zhu/aaron-marketing-skills
69 marketing skills across SEO/GEO, influencer, paid ads, and email on one shared contract, with 5 benchmark-driven auditor gates (CORE-EEAT, CITE, C³, ROAS, SEND) and keyless data connectors
zarazhangrui/frontend-slides
Generate animation-rich HTML presentations with visual style previews
trailofbits/differential-review
Security-focused diff review with git history analysis
trailofbits/entry-point-analyzer
Identify state-changing entry points in smart contracts
trailofbits/modern-python
Modern Python tooling with uv, ruff, ty, and pytest best practices