Home › Skills › Model-Optimizer › NVIDIA/Model-Optimizer/deployment

NVIDIA/Model-Optimizer/deployment

NVIDIA/skills/skills/Model-Optimizer/deployment
Agent Skill Active Apache-2.0

Serve a quantized or unquantized LLM checkpoint as an OpenAI-compatible API endpoint using vLLM, SGLang, or TRT-LLM.

A Quality 85/100 ★ 2.9k stars Updated today Python
View on GitHub →

Installation

Install this skill (Claude Code)

# Clone and copy the skill into your project
git clone https://github.com/NVIDIA/skills.git
mkdir -p .claude/skills
cp -r skills/skills/Model-Optimizer/deployment .claude/skills/
# Or for personal use: ~/.claude/skills/

Quality score breakdown

Transparent heuristic — same formula for every entry. Total 85/100.

GitHub stars (log scale)35/40
Maintenance activity25/25
License present10/10
Official project0/10
Not archived5/5
Meaningful description5/5
Repo topics set5/5

Related in Model-Optimizer

Agent Skill Active

Run commands inside a remote Docker container via the file-based command relay (tools/debugger).

★ 2.9k Python today Score 85
Agent Skill Active

Evaluates accuracy of quantized or unquantized LLMs using NeMo Evaluator Launcher (NEL).

★ 2.9k Python today Score 85
Agent Skill Active

Monitor submitted jobs (PTQ, evaluation, deployment) on SLURM clusters.

★ 2.9k Python today Score 85
Agent Skill Active

This skill should be used when the user asks to "quantize a model", "run PTQ", "post-training quantization", "NVFP4 quantization", "FP8 quantization", "INT8 quantization", "INT4 AW...

★ 2.9k Python today Score 85