NVIDIA/NemoClaw/nemoclaw-maintainer-cut-release-tag A Agent Skill Active
Cut a new semver release — bump all version strings via bump-version.ts, open a release PR, and after merge tag main and push.
Browse SKILL.md-based agent skills for Claude Code, Codex, Cursor, Gemini CLI and more.
Cut a new semver release — bump all version strings via bump-version.ts, open a release PR, and after merge tag main and push.
Runs the daytime maintainer loop for NemoClaw, prioritizing items labeled with the current version target.
Runs the end-of-day maintainer handoff for NemoClaw.
Finds open GitHub PRs with security and priority-high labels, links each to its issue, detects duplicates (multiple PRs fixing the same issue), and presents a table of review candi...
Runs the morning maintainer standup for NemoClaw.
Normalizes GitHub issue and PR titles by removing any bracketed [NemoClaw] tag case-insensitively, even when the tag appears later in the title.
Compares competing PRs that target the same issue and recommends which one to merge.
Performs a comprehensive security review of code changes in a GitHub PR or issue.
AI-assisted label triage for NVIDIA/NemoClaw issues and PRs.
Start here.
Describes the agent skills shipped with NemoClaw and how to access them by cloning the repository.
Connects NemoClaw to a local inference server.
Presents a risk framework for every configurable security control in NemoClaw.
Explains how to run NemoClaw on a remote GPU instance, including the deprecated Brev compatibility path and the preferred installer plus onboard flow.
Installs NemoClaw, launches a sandbox, and runs the first agent prompt.
Adds, removes, or modifies allowed endpoints in the sandbox policy.
Explains operational tasks after the quickstart: listing sandboxes, status and health checks, logs, diagnostics, port forwards, multiple sandboxes, credential reset, rebuilds, netw...
Inspects sandbox health, traces agent behavior, and diagnoses problems.
Explains how OpenClaw, OpenShell, and NemoClaw form the ecosystem, NemoClaw's position in the stack, what NemoClaw adds beyond the community sandbox, and when to prefer NemoClaw ve...
Describes the NemoClaw plugin and blueprint architecture and how they orchestrate the OpenClaw sandbox.
Debug AutoDeploy accuracy regressions vs a reference score (PyTorch backend or published baseline).
Claude Code skill (trtllm-agent-toolkit): implement or extend TensorRT-LLM AutoDeploy fusion transforms under transform/library/ in a TensorRT-LLM checkout.
Check whether AutoDeploy YAML configs were actually applied by analyzing server logs and optionally graph dumps (AD_DUMP_GRAPHS_DIR).
Enable and interpret TensorRT-LLM AutoDeploy FX graph text dumps via AD_DUMP_GRAPHS_DIR.
Visualize a specific transformer decoder layer from an AutoDeploy FX graph text dump as a hierarchical DOT/PNG diagram.
Translates a HuggingFace model into a prefill-only AutoDeploy custom model using reference custom ops, validates with hierarchical equivalence tests.
Compile TensorRT-LLM on a compute node inside a Docker container.
Compile TensorRT-LLM on a SLURM cluster.
Write and implement GPU kernels using NVIDIA CuTe DSL (CUTLASS 4.x Python API) — NOT for Triton, CUDA C++, or conceptual explanations.
Optimize existing Triton kernels for NVIDIA TileIR backend on Blackwell GPUs (sm_100+).
ONLY for OpenAI Triton (@triton.jit) kernel development.
Performance analysis coordination workflow.
Analyze host/CPU overhead in TensorRT-LLM inference from nsys traces.
Profiles and optimizes TensorRT-LLM host/CPU overhead using line_profiler (with nsys support planned).
Analyze ncu (NVIDIA Nsight Compute) profiling output: SOL% bottleneck classification, roofline analysis, occupancy diagnosis, memory hierarchy analysis, warp stall analysis, metr...
Nsight Systems (nsys) CLI for system-level timeline profiling.
Performance optimization coordination playbook.
Apply CUDA Graphs to PyTorch workloads — API selection (torch.compile, PyTorch make_graphed_callables, TE make_graphed_callables, MCore CudaGraphManager, FullCudaGraphWrapper, m...
Identify and eliminate host-device synchronizations in PyTorch code.
Code instrumentation for timing workloads.
Best practices for contributing code to TensorRT-LLM.
Systematic approach to exploring the TensorRT-LLM codebase before implementing new features or optimizations.
Upgrade flashinfer-python version in TensorRT-LLM.
Review, design, and refactor TensorRT-LLM PyTorch MoE code for architecture fit, clean code, maintainability, and testability.
Generate a source-backed starting `trtllm-serve --config` YAML for basic aggregate single-node PyTorch serving, aligned with checked-in TensorRT-LLM configs and deployment docs.
Add a new cuTile GPU kernel operator to TileGym.
Converts cuTile Python GPU kernels (@ct.kernel) to cuTile.jl Julia equivalents.
Converts cuTile GPU kernels (@ct.kernel) to Triton (@triton.jit).