Home › Agent Skills

Agent Skills

Browse SKILL.md-based agent skills for Claude Code, Codex, Cursor, Gemini CLI and more.

Filters (1 active)
Language
Python25
Activity
Active25
Install method
SKILL.md25
25 results (<1ms) · data updated 2026-08-14

Results

NVIDIA/TensorRT-LLM/ad-accuracy-debug A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/ad-accuracy-debug

Debug AutoDeploy accuracy regressions vs a reference score (PyTorch backend or published baseline).

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/ad-add-fusion-transformation A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/ad-add-fusion-transformation

Claude Code skill (trtllm-agent-toolkit): implement or extend TensorRT-LLM AutoDeploy fusion transforms under transform/library/ in a TensorRT-LLM checkout.

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/ad-conf-check A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/ad-conf-check

Check whether AutoDeploy YAML configs were actually applied by analyzing server logs and optionally graph dumps (AD_DUMP_GRAPHS_DIR).

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/ad-graph-dump A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/ad-graph-dump

Enable and interpret TensorRT-LLM AutoDeploy FX graph text dumps via AD_DUMP_GRAPHS_DIR.

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/ad-layer-visualizer A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/ad-layer-visualizer

Visualize a specific transformer decoder layer from an AutoDeploy FX graph text dump as a hierarchical DOT/PNG diagram.

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/ad-model-onboard A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/ad-model-onboard

Translates a HuggingFace model into a prefill-only AutoDeploy custom model using reference custom ops, validates with hierarchical equivalence tests.

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/kernel-cute-writing A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/kernel-cute-writing

Write and implement GPU kernels using NVIDIA CuTe DSL (CUTLASS 4.x Python API) — NOT for Triton, CUDA C++, or conceptual explanations.

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/perf-host-optimization A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/perf-host-optimization

Profiles and optimizes TensorRT-LLM host/CPU overhead using line_profiler (with nsys support planned).

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/perf-nsight-compute-analysis A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/perf-nsight-compute-analysis

Analyze ncu (NVIDIA Nsight Compute) profiling output: SOL% bottleneck classification, roofline analysis, occupancy diagnosis, memory hierarchy analysis, warp stall analysis, metr...

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/perf-torch-cuda-graphs A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/perf-torch-cuda-graphs

Apply CUDA Graphs to PyTorch workloads — API selection (torch.compile, PyTorch make_graphed_callables, TE make_graphed_callables, MCore CudaGraphManager, FullCudaGraphWrapper, m...

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/trtllm-codebase-exploration A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/trtllm-codebase-exploration

Systematic approach to exploring the TensorRT-LLM codebase before implementing new features or optimizations.

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/trtllm-moe-develop A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/trtllm-moe-develop

Review, design, and refactor TensorRT-LLM PyTorch MoE code for architecture fit, clean code, maintainability, and testability.

★ 2.9k Python Updated today Score 85 TensorRT-LLM

NVIDIA/TensorRT-LLM/trtllm-serve-config-guide A Agent Skill Active

NVIDIA/skills/skills/TensorRT-LLM/trtllm-serve-config-guide

Generate a source-backed starting `trtllm-serve --config` YAML for basic aggregate single-node PyTorch serving, aligned with checked-in TensorRT-LLM configs and deployment docs.

★ 2.9k Python Updated today Score 85 TensorRT-LLM