muratcankoylan/context-optimization A Agent Skill Active
Apply compaction, masking, and caching strategies
Browse SKILL.md-based agent skills for Claude Code, Codex, Cursor, Gemini CLI and more.
Apply compaction, masking, and caching strategies
Master orchestrator, peer-to-peer, and hierarchical multi-agent architectures
Design short-term, long-term, and graph-based memory architectures
Build tools that agents can use effectively, including architectural reduction patterns
Build evaluation frameworks for agent systems
Browse and query HF datasets with the Dataset Viewer API
Create and manage datasets with configs and SQL querying
Model evaluation with vLLM/lighteval and eval tables
Run compute jobs and Python scripts on HF infrastructure
Train models with TRL: SFT, DPO, GRPO, GGUF conversion
Create and manage paper pages on HF Hub
Publish papers on HF Hub with model/dataset links
Build reusable scripts for HF API operations
Track ML experiments with real-time dashboards
Train vision models on HF infrastructure
Build Gradio apps and deploy to HF Spaces
Run ML models in the browser with Transformers.js
SEO, GEO, Google Ads, and Meta Ads skills with live data
Browser automation with Playwright
CUDA-Q onboarding guide for installation, test programs, GPU simulation, QPU hardware, and quantum applications.
Use when writing DALI data loading or preprocessing code with `nvidia.dali.experimental.dynamic` (ndd), or when converting DALI pipeline-mode code to dynamic mode, or when the user...
Guide for adding support for new LLM or VLM models in Megatron-Bridge.
Dev environment setup for Megatron Bridge — container-based development, uv package management, lockfile regeneration, adding dependencies, Slurm container usage, and common build...
Bump a pinned dependency (TransformerEngine, Megatron-LM, NRX, etc.), regenerate the lockfile, open a PR, and drive it to green by attaching a watchdog to the "CICD NeMo" workflow...
CI/CD reference for Megatron Bridge — pipeline structure, commit and PR workflow, CI failure investigation, and common failure patterns.
Code style and quality rules for Megatron Bridge — ruff configuration, naming conventions, type hints, mypy rules, docstrings, copyright headers, logging, and the code review check...
Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data.
Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.
External NeMo-RL end-to-end validation workflow for Megatron-Bridge model/provider changes, including downstream compatibility checks, external RL lifecycle behavior, Megatron poli...
Structured framework for verifying numerical parity of HFMCore weight conversions.
Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.
Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.
Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as...
Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.
MoE expert-parallel communication overlap in Megatron Bridge.
Choose the right MoE token dispatcher (`alltoall`, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage.
Representative MoE training playbooks by hardware platform and model family.
Long-context MoE training guidance for Megatron Bridge.
Systematic workflow for MoE training optimization in Megatron Bridge, based on the Megatron-Core MoE paper.
Practical guidance for training MoE VLMs in Megatron Bridge.
Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.
Validate and use packed sequences and long-context training in Megatron-Bridge, distinguishing offline packed SFT for LLMs from in-batch packing for VLMs, and applying the right CP...
Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
Recommend and customize Megatron Bridge recipes for a user's model, GPU count, and training goal.