skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
slowlyc/agent-gpu-skills136 installs

triton-skill

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. Use when the task explicitly involves triton.jit, triton.language, tl.*, Gluon, TensorDescriptor, Triton autotune, TritonGPU/MLIR lowering, triton_kernels, or converting a CUDA kernel to Triton. Use cuda-skill for raw CUDA/PTX and NVIDIA architecture facts, and cutlass-skill for CUTLASS, CuTe, or CuTeDSL work.

How do I install this agent skill?

npx skills add https://github.com/slowlyc/agent-gpu-skills --skill triton-skill
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    This skill provides a development environment and reference guide for Triton and Gluon GPU kernel programming. It includes documentation, code snippets, and a utility script to maintain a local copy of the Triton source code from its official repository.

  • Socketpass

    No alerts

  • Snykwarn

    Risk: MEDIUM · 1 issue

  • Runlayerpass

    1/3 files flagged

What does this agent skill do?

Triton and Gluon development

Use the local Triton checkout as the primary source for APIs and implementation patterns. Prefer current tutorials and source over remembered signatures because Triton and Gluon evolve quickly.

Locate the checkout

Resolve the directory containing this SKILL.md, then use its repos/triton/ child. The installer links that path to agent-gpu-skills/third_party/triton/ or to the checkout supplied through TRITON_REPO.

In commands below, replace TRITON_REPO with the resolved absolute path:

TRITON_REPO=/absolute/path/to/triton-skill/repos/triton

If the checkout is missing, run this from the agent-gpu-skills repository and reinstall the Skill:

bash update-repos.sh triton
bash install.sh --skill triton-skill

Choose the source surface

TaskStart here
Triton language syntax and introductory patternspython/tutorials/
Gluon layout and architecture-level patternspython/tutorials/gluon/
Complete example kernelspython/examples/
Production matmul, reduction, top-k and SwiGLUpython/triton_kernels/triton_kernels/
tl.* definitions and semanticspython/triton/language/
JIT, autotuning and runtime behaviorpython/triton/runtime/
Python compiler entry pointspython/triton/compiler/
Triton and GPU dialect definitionsinclude/triton/Dialect/
Compiler analyses, transforms and loweringlib/

Read quick-reference.md when choosing a tutorial, a complete example, or a production-kernel implementation.

Query workflow

  1. Identify whether the task is Triton, Gluon, or compiler internals.
  2. Find the closest current tutorial or implementation for the operation and architecture.
  3. Verify every API used against its definition or another current call site.
  4. Preserve the target workload's shape, dtype, stride, layout, masking, and numerical contract.
  5. Establish correctness before changing launch parameters or optimization strategy.

Discover current examples before relying on a remembered filename:

find "$TRITON_REPO/python/tutorials" -maxdepth 2 -type f | sort
find "$TRITON_REPO/python/examples" -type f | sort

Query Triton language usage and definitions:

rg -n 'tl\.dot|tl\.dot_scaled' "$TRITON_REPO/python/tutorials"
rg -n '@triton\.autotune' "$TRITON_REPO/python/tutorials"
rg -n '^def (load|store|dot|dot_scaled)' \
  "$TRITON_REPO/python/triton/language"

Query Gluon architecture patterns:

rg -n '@gluon\.jit' "$TRITON_REPO/python/tutorials/gluon"
rg -n 'wgmma|tcgen05|mbarrier|tma' \
  "$TRITON_REPO/python/tutorials/gluon" \
  "$TRITON_REPO/python/examples"

Trace production kernels:

rg -n 'persistent|TensorDescriptor' \
  "$TRITON_REPO/python/triton_kernels/triton_kernels/matmul_details"

rg -n 'mxfp|flexpoint' \
  "$TRITON_REPO/python/triton_kernels/triton_kernels/numerics_details"

Trace compiler definitions and lowering:

rg -n 'def.*Op' "$TRITON_REPO/include/triton/Dialect/Triton/IR"
rg -n 'Encoding' "$TRITON_REPO/include/triton/Dialect/TritonGPU/IR"
rg -n 'wgmma|tma|tcgen05' \
  "$TRITON_REPO/include/triton/Dialect/TritonNvidiaGPU"
rg -n 'Pattern|Rewrite' "$TRITON_REPO/lib/Conversion/TritonGPUToLLVM"

Implementation discipline

Keep these layers separate during diagnosis:

Python kernel and launch metadata
  → Triton/Gluon IR and compiler transforms
  → generated GPU code on the selected target

A source-level pattern does not prove that the compiled kernel uses the intended instruction or memory path. Inspect compiler output or profile data when that distinction matters. Add cuda-skill for PTX semantics, compute capability, Nsight, or Compute Sanitizer details.

For correctness work:

  • compare against an independent reference over representative and boundary shapes;
  • test masked tails, non-power-of-two dimensions, strides and dtype conversions;
  • separate compilation failures from runtime correctness and numerical tolerance;
  • reproduce the original dispatch and launch metadata before reducing the case.

For performance work:

  • freeze the benchmark shape set and measurement method;
  • confirm autotune keys cover every dimension that changes the best configuration;
  • change one tile, warp, stage, persistence, or specialization choice at a time;
  • verify the generated path before attributing a result to TMA, WGMMA, tcgen05, or warp specialization.

Updating the source

From the agent-gpu-skills repository:

bash update-repos.sh triton
python3 scripts/validate_repo.py --require-sources

The checkout follows Triton main, while third_party/UPSTREAMS.toml records the commit last accepted by this Skill. Review source-map drift before updating that record.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/slowlyc/agent-gpu-skills/triton-skill">View triton-skill on skillZs</a>