tilegym-adding-cutile-kernel
NVIDIA/skills
TileGym に新しい cuTile GPU カーネル演算子を追加し、ディスパッチ登録、バックエンドの実装、エクスポート、テスト、およびベンチマークを網羅する。
...すべて拡張しますTileGymへのcuTileカーネルの追加
cuTile バックエンドを使用して新しい演算子(例:my_op)を追加するためのエンドツーエンドのワークフロー。
実行ルール
以下のルールを厳守してください:
- コードを記述する前に、TodoWrite を使用して以下のチェックリストを作成してください
- 手順は順番通りに実行してください。先へ飛ばしたり、手順を組み合わせたりしないでください
- 各タスクは、完了したら「
完了」、開始したら「進行中」とマークしてください - 手順が該当しない場合(例:cuTileの実装がない場合など)は、注釈を付けて「
完了」とマークしてください。黙ってスキップしてはいけません - 各手順は、ファイルへの書き込みまたは明示的なスキップ決定を必ず行わなければなりません。黙って省略することは許されません
手順
最初に、このチェックリストをTodoWriteに必ずコピーしてください:
- [ ] ステップ 1: ops.py にディスパッチインターフェースを登録する
- [ ] ステップ 2: cuTile バックエンドを実装する
- [ ] ステップ 3: __init__.py (cutile) への登録
- [ ] ステップ 4: テストの追加
- [ ] ステップ 5: tests/benchmark へのベンチマークの追加
- [ ] ステップ 6: 検証(pytest + lint を実行)
ステップ 1: ディスパッチインターフェースを登録する
ファイル:src/tilegym/ops/ops.py
@dispatch関数を追加します。これは、すべてのバックエンドに対する唯一のエントリポイントです。
@dispatch(
"my_op",
)
def my_op(
input: torch.Tensor,
out: Optional[torch.Tensor] = None,
**kwargs: Any,
):
"""
my_op の説明。
引数:
input: 入力テンソル
out: 事前割り当て済みの出力テンソル(Optional)
**kwargs: バックエンド固有の設定のための追加引数
戻り値:
torch.Tensor
"""
raise NotImplementedError(f"{get_current_backend()} では my_op は実装されていません")
主なルール:
- 関数本体では
NotImplementedErrorのみを発生させる - バックエンド固有のパラメータには
**kwargs を指定する
参照:src/tilegym/ops/ops.pyにある既存のオペレーションを参照してください(例:silu_and_mul、softmax)
ステップ 2: cuTile バックエンドの実装
ファイル:src/tilegym/ops/cutile/my_op.py
ファイル構造は次のテンプレートに従います:
import torch
import cuda.tile as ct
from tilegym.backend import register_impl
@ct.kernel
def my_op_kernel_ct(x, output, n_elements: ct.Constant[int], BLOCK_SIZE: ct.Constant[int]):
bid = ct.bid(0)
indices = bid * BLOCK_SIZE + ct.arange(0, BLOCK_SIZE)
x_val = ct.gather(x, indices)
# ... 計算 ...
ct.scatter(output, indices, result)
@register_impl("my_op", backend="cutile")
def my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs) -> torch.Tensor:
n = input.numel()
if out is None:
out = torch.empty_like(input)
grid = ((n + 1023) // 1024,)
ct.launch(stream, grid, kernel, (some args, ...))
return out
参照:src/tilegym/ops/cutile/silu_and_mul.py
ステップ 3:__init__.pyへの登録(重要)
この手順を省略すると、cuTile バックエンドの実装が読み込まれなくなります。
ファイル:src/tilegym/ops/cutile/__init__.py
if is_backend_available("cutile"):ブロック内に以下を追加(アルファベット順):
from . import my_op
また、関数のインポートセクションには以下を追加してください:
from .my_op import my_op
さらに、__all__ に「my_op」を追加します。
ステップ 4: テストを追加する
ファイル:tests/ops/test_my_op.py
重要: 常にtilegym.ops からインポートし、tilegym.ops.cutile.my_op からは絶対にインポートしないでください。
import pytest
import torch
from tilegym.backend import is_backend_available, set_backend
from .. import common
_backends = ["cutile"]
class Test_MY_OP(common.PyTestCase):
@staticmethod
def reference(input):
"""PyTorch を使用した参照実装。"""
return torch.some_reference(input)
@pytest.mark.parametrize("shape, dtype", [
((1024,), torch.float16),
((1024, 512), torch.float32),
((64, 64, 64), torch.bfloat16),
])
@pytest.mark.parametrize("backend", _backends)
def test_op(self, shape, dtype, backend, arch):
if backend == "cutile" and not is_backend_available("cutile"):
pytest.skip("Cutile バックエンドが利用できません")
try:
set_backend(backend)
except Exception as e:
pytest.skip(f"バックエンドがサポートされていません: {e}")
self.setUp()
from tilegym.ops import my_op
A = torch.randn(*shape, dtype=dtype, device="cuda")
self.assertCorrectness(
my_op, self.reference, {"input": A},
atol=1e-3, rtol=1e-3,
)
主要なパターン:
_backends = ["cutile"]test_op: try-except ブロック内でset_backend(backend)を使用し、self.setUp()を呼び出す
参照:tests/ops/test_silu_and_mul.py
以下に一般的なエラーを記載します。
1. _backends リストの欠落(クラス内)
2. test_op / test_op_xxx — @pytest.mark.parametrize("backend", _backends)、backend パラメータ、および tilegym.is_backend_available / tilegym.set_backend のパターンが欠落している
ステップ 5: tests/benchmark にベンチマークを追加する
ファイル:tests/benchmark/bench_my_op.py
benchmark_rules.md の主なルール:
tilegym.ops.my_op(a, b, ..., backend=backend)を使用してオペレーションを呼び出してください。set_backendは使用しないでください。ALL_BACKENDSを定義し(少なくともcutileとtorchを含める)、get_supported_backends()でフィルタリングする。reference_my_op(...)を実装し、登録します:register_impl("my_op", "torch")(reference_my_op)。create_benchmark_config()を使用して、triton.testing.Benchmarkの設定を構築します(例: 形状/データ型別)。bench_my_op(...)内で@triton.testing.perf_report([...])を使用します。 ベンチマーク関数内で、torch.testing.assert_close(fn(), ref(), ...)を使用して正しさを確認し、次にms = triton.testing.do_bench(fn)(またはdo_bench_cudagraph) を実行して GB/s または TFLOPS を計算し、メトリックを返します。- エントリポイント:
if __name__ == "__main__": bench_my_op.run(print_data=True)。
テンプレートの構造:
import torch
import triton
import triton.testing
import tilegym
from tilegym.backend import is_backend_available, register_impl
ALL_BACKENDS = [
("cutile", "cuTile", ("orange", "-")) if is_backend_available("cutile") else None,
("torch", "PyTorch", ("green", "-")),
]
def get_supported_backends():
return [p for p in ALL_BACKENDS if p is not None]
def reference_my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs):
"""PyTorch を使用した参照実装。"""
...
register_impl("my_op", "torch")(reference_my_op)
def create_benchmark_config(datatype, ...):
available_backends = get_supported_backends()
if not available_backends:
return None
backends, names, styles = zip(*available_backends)
return triton.testing.Benchmark(
x_names=["M"], # またはその他の次元名
x_vals=[...],
line_arg="backend",
line_vals=list(backends),
line_names=list(names),
styles=list(styles),
ylabel="GB/s", # または TFLOPS
plot_name="my-op-...",
args={"datatype": datatype, ...},
)
@triton.testing.perf_report([
create_benchmark_config(datatype, ...)
for datatype in [torch.float16, torch.float32]
for ... in [...]
])
def bench_my_op(M, backend, datatype, ..., device="cuda"):
x = torch.randn(..., dtype=datatype, device=device)
fn = lambda: tilegym.ops.my_op(x, backend=backend)
ref = lambda: reference_my_op(x)
torch.testing.assert_close(fn(), ref(), rtol=1e-2, atol=1e-2)
ms = triton.testing.do_bench(fn) # または do_bench_cudagraph(fn)
# ms と問題サイズからメトリック(例:GB/s や TFLOPS)を計算
return metric
if __name__ == "__main__":
bench_my_op.run(print_data=True)
ベンチマークプロットの名前:末尾に-TFLOPSまたは-GBpsを必ず含めること
- 例:
plot_name=f"persistent-layer-norm-M{num_rows}-{dtype_name}-GBps"
ステップ 6: 検証
# テストを実行
pytest tests/ops/test_my_op.py -v
# ベンチマークを実行(オプション)
python tests/benchmark/bench_my_op.py
# Lint
pre-commit run -a
---
name: tilegym-adding-cutile-kernel
description: Add a new cuTile GPU kernel operator to TileGym, covering dispatch registration, backend implementation, exports, tests, and benchmarks.
license: CC-BY-4.0 AND Apache-2.0
---
# Adding a cuTile Kernel to TileGym
End-to-end workflow for adding a new operator (e.g., `my_op`) with cuTile backend.
## Execution Rules
**MUST follow these rules strictly:**
1. Use TodoWrite to create the checklist below BEFORE writing any code
2. Execute steps **in order** — do NOT skip ahead or combine steps
3. Mark each todo as `completed` after finishing, `in_progress` when starting
4. If a step is not applicable (e.g., no cuTile impl), mark it `completed` with a note, do NOT silently skip
5. Each step MUST result in a file write or explicit skip decision — no silent omissions
## Instructions
MUST copy this checklist to TodoWrite at the start:
```
- [ ] Step 1: Register dispatch interface in ops.py
- [ ] Step 2: Implement cuTile backend
- [ ] Step 3: Register in __init__.py (cutile)
- [ ] Step 4: Add tests
- [ ] Step 5: Add benchmark to tests/benchmark
- [ ] Step 6: Verify (run pytest + lint)
```
## Step 1: Register dispatch interface
**File**: `src/tilegym/ops/ops.py`
Add a `@dispatch` function — this is the **single entry point** for all backends.
```python
@dispatch(
"my_op",
)
def my_op(
input: torch.Tensor,
out: Optional[torch.Tensor] = None,
**kwargs: Any,
):
"""
Description of my_op.
Args:
input: Input tensor
out: Optional preallocated output tensor
**kwargs: Additional arguments for backend-specific configurations
Returns:
torch.Tensor
"""
raise NotImplementedError(f"my_op is not implemented for {get_current_backend()}")
```
**Key rules:**
- Function body only raises `NotImplementedError`
- Include `**kwargs` for backend-specific parameters
**Reference**: See existing ops in `src/tilegym/ops/ops.py` (e.g., `silu_and_mul`, `softmax`)
## Step 2: Implement cuTile backend
**File**: `src/tilegym/ops/cutile/my_op.py`
The file structure follows this template:
```python
import torch
import cuda.tile as ct
from tilegym.backend import register_impl
@ct.kernel
def my_op_kernel_ct(x, output, n_elements: ct.Constant[int], BLOCK_SIZE: ct.Constant[int]):
bid = ct.bid(0)
indices = bid * BLOCK_SIZE + ct.arange(0, BLOCK_SIZE)
x_val = ct.gather(x, indices)
# ... compute ...
ct.scatter(output, indices, result)
@register_impl("my_op", backend="cutile")
def my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs) -> torch.Tensor:
n = input.numel()
if out is None:
out = torch.empty_like(input)
grid = ((n + 1023) // 1024,)
ct.launch(stream, grid, kernel, (some args, ...))
return out
```
**Reference**: `src/tilegym/ops/cutile/silu_and_mul.py`
## Step 3: Register in `__init__.py` (CRITICAL)
Missing this step means the cuTile backend implementation never gets loaded.
**File**: `src/tilegym/ops/cutile/__init__.py`
Add inside `if is_backend_available("cutile"):` block (alphabetically):
```python
from . import my_op
```
And in the function import section:
```python
from .my_op import my_op
```
And add `"my_op"` to `__all__`.
## Step 4: Add tests
**File**: `tests/ops/test_my_op.py`
**CRITICAL**: Always import from `tilegym.ops`, NEVER from `tilegym.ops.cutile.my_op`.
```python
import pytest
import torch
from tilegym.backend import is_backend_available, set_backend
from .. import common
_backends = ["cutile"]
class Test_MY_OP(common.PyTestCase):
@staticmethod
def reference(input):
"""Reference implementation using PyTorch."""
return torch.some_reference(input)
@pytest.mark.parametrize("shape, dtype", [
((1024,), torch.float16),
((1024, 512), torch.float32),
((64, 64, 64), torch.bfloat16),
])
@pytest.mark.parametrize("backend", _backends)
def test_op(self, shape, dtype, backend, arch):
if backend == "cutile" and not is_backend_available("cutile"):
pytest.skip("Cutile backend not available")
try:
set_backend(backend)
except Exception as e:
pytest.skip(f"Backend is not supported: {e}")
self.setUp()
from tilegym.ops import my_op
A = torch.randn(*shape, dtype=dtype, device="cuda")
self.assertCorrectness(
my_op, self.reference, {"input": A},
atol=1e-3, rtol=1e-3,
)
```
**Key patterns:**
- `_backends = ["cutile"]`
- `test_op`: use `set_backend(backend)` with try-except, call `self.setUp()`
**Reference**: `tests/ops/test_silu_and_mul.py`
Below is the common errors.
```
1. Missing _backends list (inside class)
2. test_op / test_op_xxx — missing @pytest.mark.parametrize("backend", _backends), backend parameter, and tilegym.is_backend_available / tilegym.set_backend pattern
```
## Step 5: Add benchmark to tests/benchmark
**File**: `tests/benchmark/bench_my_op.py`
**Key rules from benchmark_rules.md:**
- Call the op via `tilegym.ops.my_op(a, b, ..., backend=backend)` — do **not** use `set_backend`.
- Define `ALL_BACKENDS` (include at least `cutile` and `torch`), filter with `get_supported_backends()`.
- Implement `reference_my_op(...)` and register it: `register_impl("my_op", "torch")(reference_my_op)`.
- Use `create_benchmark_config()` to build `triton.testing.Benchmark` configs (e.g. by shape/dtype).
- Use `@triton.testing.perf_report([...])` on `bench_my_op(...)`; inside the bench function: correctness check with `torch.testing.assert_close(fn(), ref(), ...)`, then `ms = triton.testing.do_bench(fn)` (or `do_bench_cudagraph`), compute GB/s or TFLOPS, and return the metric.
- Entry point: `if __name__ == "__main__": bench_my_op.run(print_data=True)`.
Template structure:
```python
import torch
import triton
import triton.testing
import tilegym
from tilegym.backend import is_backend_available, register_impl
ALL_BACKENDS = [
("cutile", "cuTile", ("orange", "-")) if is_backend_available("cutile") else None,
("torch", "PyTorch", ("green", "-")),
]
def get_supported_backends():
return [p for p in ALL_BACKENDS if p is not None]
def reference_my_op(input: torch.Tensor, out: torch.Tensor = None, **kwargs):
"""Reference implementation using PyTorch."""
...
register_impl("my_op", "torch")(reference_my_op)
def create_benchmark_config(datatype, ...):
available_backends = get_supported_backends()
if not available_backends:
return None
backends, names, styles = zip(*available_backends)
return triton.testing.Benchmark(
x_names=["M"], # or other dimension names
x_vals=[...],
line_arg="backend",
line_vals=list(backends),
line_names=list(names),
styles=list(styles),
ylabel="GB/s", # or TFLOPS
plot_name="my-op-...",
args={"datatype": datatype, ...},
)
@triton.testing.perf_report([
create_benchmark_config(datatype, ...)
for datatype in [torch.float16, torch.float32]
for ... in [...]
])
def bench_my_op(M, backend, datatype, ..., device="cuda"):
x = torch.randn(..., dtype=datatype, device=device)
fn = lambda: tilegym.ops.my_op(x, backend=backend)
ref = lambda: reference_my_op(x)
torch.testing.assert_close(fn(), ref(), rtol=1e-2, atol=1e-2)
ms = triton.testing.do_bench(fn) # or do_bench_cudagraph(fn)
# Compute metric (e.g. GB/s or TFLOPS) from ms and problem size
return metric
if __name__ == "__main__":
bench_my_op.run(print_data=True)
```
**Benchmark Plot Names**: Must include `-TFLOPS` or `-GBps` suffix
- Example: `plot_name=f"persistent-layer-norm-M{num_rows}-{dtype_name}-GBps"`
## Step 6: Verify
```bash
# Run tests
pytest tests/ops/test_my_op.py -v
# Run benchmark (optional)
python tests/benchmark/bench_my_op.py
# Lint
pre-commit run -a
```
すべてのファイル
5件のファイルtilegym-adding-cutile-kernelをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/tilegym-adding-cutile-kernel # Copy SKILL.md to your .claude/skills/ directory
コピー





家
