nemo-mbridge-perf-moe-long-context
NVIDIA/skills
長いコンテキストウィンドウを持つMixture-of-Expertsモデルの学習に関する指針を提供し、コンテキスト並列処理の規模設定、選択的な再計算、ディスパッチャの選定、および最近の実験から得られた実用的なパターンについて解説する。
...すべて拡張しますMoEの長文コンテキスト学習
Stable ドキュメント: @docs/training/moe-optimization.md カード: @skills/nemo-mbridge-perf-moe-long-context/card.yaml
長文コンテキストにおける変化
シーケンスの長さが4Kクラスを大幅に超えると、アテンションメモリと 活性化値の保持が主要な制約となります。MoEモデルでは、 通常、以下の組み合わせが必要になります:
- コンテキスト並列化
- 選択的再計算
- 精度の低下
- オプティマイザ状態のCPUオフロード
- 残りのDPバジェットを無駄にしないディスパッチャおよびPPレイアウト
丸め処理によるスケーリングパターン
H100上のDSV3
DSV3のロングコンテキスト実行では、安定したパターンが見られます:
- 最も短いコンテキストを過ぎると、 完全な再計算よりも選択的な再計算の方が効果的である
- CPを適切に増加させれば、中程度の長さから非常に長い コンテキストに至るまで、スループットは比較的狭い範囲に収まる
- CPが増加するにつれて、トレードオフは「メモリ収まり」から「GPU枚数の実現可能性」へと移行する
言い換えれば、 レイアウトが適切に選択されていれば、長いコンテキストであっても利用率が即座に低下することはないが、DP予算は極めて急速に消費される。
GB200上のQwen3-Next
Qwen3-Nextの挙動は、メモリに敏感な中規模モデルに近くなります:
- 適度なCPであれば、8Kおよび32Kも実用的な範囲に収まる
- 64Kも可能だが、スループットの低下が顕著になり、メモリが かなり逼迫する
- パイプラインレイアウトとグループ化GEMMの改良は、CPとほぼ同等の重要性を持ちます
GB200上のQwen3 235B
Qwen3 235Bは、TP、CP、およびHybridEPが適切に調整されていれば、NVL72システム上でも 長いコンテキストが依然として効率的であることを示しています。最適な128Kクラスの構成は、 単に「収まるだけ」のレシピというわけではありません。配線、 並列性、再計算のバランスが取れていれば、高い効率を維持できます。
CPのサイズ設定に関する経験則
4Kシャードを目標として開始する:良い最初の目安は
CP ≈ seq_len / 4096 であり、その後、実用的な2の冪のレイアウトに丸める。可能であればDPを維持する:CP、 EP、TP、PPが相まってDPを限界値まで圧迫すると、ロングコンテキストのスケーラビリティが脆弱になる。
選択的な再計算を優先する:
up_proj、norm、moe、moe_act、mlpなどのモジュールは、完全な再計算を行う前に再計算する。非常に長いコンテキストでは、SDPAを多用する再計算を避ける:アテンションの 内部構造を再計算すると、より小さなMoEやMLP側のモジュールを再計算するよりも メモリ上のメリットが少なく、負荷が大幅に増加する可能性がある。
NVL72システムではTPを別の調整手段として活用する:GB200およびGB300の実行では、 効率を維持しつつ、CPを多少犠牲にしてTPを優先できる場合がある。
GBSの縮小が必要になると想定する:CPが上昇しDPが低下するにつれて、 グローバルバッチサイズを縮小するか、より高いGAを受け入れる必要が生じる可能性がある。
代表的な構成ファミリー
H100上の128KのDSV3
TP=1 CP=32 EP=32 PP=8 VPP=4
精度:FP8クラス
ディスパッチャ:DeepEP
再計算:up_proj、norm、moe、mlp
追加のメモリ補助:オプティマイザによるCPUオフロード
H100上のDSV3(256K)
TP=1 CP=64 EP=32 PP=8 EDP=2 VPP=4
精度:FP8クラス
ディスパッチャ:DeepEP
再計算:up_proj、norm、moe、mlp
追加のメモリ支援:オプティマイザによるCPUオフロード
GB200上で128KのQwen3 235B
TP=4 CP=4 EP=32 PP=4 VPP=12
精度:BF16 または MXFP8
ディスパッチャ:HybridEP
再計算:moe_act、norm
CUDAグラフ:attn + moe_router + moe_preprocess
再計算およびCUDAグラフに関するガイダンス
ロングコンテキストのMoEトレーニングの場合:
- まずは選択的再計算から開始
- 形状とルーティングパスが安定してから、CUDAグラフを追加する
- CUDAグラフを使用する際は、シーケンス長とMBSを固定する
- 実行が極めて動的なバッチに依存する場合は、イーガー実行を優先する
参考資料:
- @docs/training/activation-recomputation.md
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md
落とし穴
CPはEPやPPに取って代わるものではありません。CPは新たな次元を追加するものであり、 他の手法を無効にするわけではありません。
優れた4Kベースラインであっても、ロングコンテキストのベースラインとしては不適切な場合があります。ルーティングモード、 再計算の選択、およびオフロード戦略を変更する必要があることがよくあります。
GPU数の制約が真のボトルネックとなります。非常に長いコンテキストは、 単一の処理手順では問題なく見えることがあっても、モデル全体にわたりEPやPPを 適切に適用すると、実現不可能になる場合があります。
CUDAグラフには静的な形状が必要です:可変長のバッチや 機会主義的なパディング戦略は、気付かないうちにパスを破壊する可能性があります。
128K以上では、コンテナとカーネルのサポートがより重要になります。ロングコンテキストのパスは、 ショートコンテキストの立ち上げよりも、新しいカーネルやバグ修正に依存する傾向があるからです。
---
name: nemo-mbridge-perf-moe-long-context
description: Provides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments.
license: Apache-2.0
---
# MoE Long-Context Training
Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-long-context/card.yaml
## What Changes At Long Context
Once sequence length moves well past the 4K-class regime, attention memory and
activation residency become the dominant constraints. For MoE models, that
usually means you need some combination of:
- context parallelism
- selective recompute
- lower precision
- CPU offload for optimizer state
- a dispatcher and PP layout that do not waste the smaller remaining DP budget
## Rounded Scaling Patterns
### DSV3 on H100
The DSV3 long-context runs show a stable pattern:
- selective recompute works better than full recompute once you move past the
shortest contexts
- throughput stays in a fairly narrow band from mid-length through very long
contexts if CP is increased appropriately
- the trade shifts from "memory fit" to "GPU-count feasibility" as CP grows
In other words, long context does not immediately collapse utilization if the
layout is chosen well, but it does consume the DP budget very quickly.
### Qwen3-Next on GB200
Qwen3-Next behaves more like a memory-sensitive medium-scale model:
- 8K and 32K remain practical with moderate CP
- 64K is possible, but the throughput drop is noticeable and memory becomes
much tighter
- pipeline layout and grouped-GEMM improvements matter almost as much as CP
### Qwen3 235B on GB200
Qwen3 235B shows that long context can still be efficient on NVL72 systems when
TP, CP, and HybridEP are coordinated. The best 128K-class configurations are
not just "fit-only" recipes; they can remain highly efficient if routing,
parallelism, and recompute are balanced.
## CP Sizing Rules Of Thumb
1. **Start from a 4K shard target**: a good first guess is
`CP ~= seq_len / 4096`, then round to a practical power-of-two layout.
2. **Keep DP alive if possible**: long-context scaling becomes brittle once CP,
EP, TP, and PP together squeeze DP down to the floor.
3. **Prefer selective recompute**: recompute modules such as `up_proj`, `norm`,
`moe`, `moe_act`, or `mlp` before reaching for full recompute.
4. **Avoid SDPA-heavy recompute at very long context**: recomputing attention
internals can add a lot of work for less memory benefit than recomputing
smaller MoE and MLP-side modules.
5. **Use TP as another lever on NVL72 systems**: GB200 and GB300 runs can
sometimes trade some CP for TP while still staying efficient.
6. **Assume GBS will need to shrink**: as CP rises and DP falls, you may need
to reduce global batch size or accept higher GA.
## Representative Config Families
### DSV3 at 128K on H100
```text
TP=1 CP=32 EP=32 PP=8 VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```
### DSV3 at 256K on H100
```text
TP=1 CP=64 EP=32 PP=8 EDP=2 VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```
### Qwen3 235B at 128K on GB200
```text
TP=4 CP=4 EP=32 PP=4 VPP=12
Precision: BF16 or MXFP8
Dispatcher: HybridEP
Recompute: moe_act, norm
CUDA Graph: attn + moe_router + moe_preprocess
```
## Recompute And CUDA Graph Guidance
For long-context MoE training:
- start with selective recompute
- add CUDA graphs only after the shapes and routing path are stable
- keep sequence length and MBS fixed when using CUDA graphs
- if the run depends on highly dynamic batches, prefer eager execution
Useful references:
- @docs/training/activation-recomputation.md
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md
## Pitfalls
1. **CP does not replace EP or PP**: it adds another dimension; it does not make
the others disappear.
2. **A good 4K baseline can still be a bad long-context baseline**: routing mode,
recompute choice, and offload strategy often need to change.
3. **GPU-count feasibility becomes the real constraint**: very long context can
look fine in a single recipe, then become impossible once EP and PP are added
honestly across the full model.
4. **CUDA graphs need static shapes**: variable-length batches and opportunistic
padding strategies can silently break the path.
5. **Container and kernel support matters more at 128K+**: long-context paths
tend to rely on newer kernels and bug fixes than short-context bring-up does.
すべてのファイル
1件のファイルnemo-mbridge-perf-moe-long-contextをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-long-context # Copy SKILL.md to your .claude/skills/ directory
コピー





家
