選項
首頁首頁 Skill 數據科學與機器學習 nemo-mbridge-perf-moe-long-context

nemo-mbridge-perf-moe-long-context

NVIDIA/skills NVIDIA/skills

針對使用長上下文視窗訓練「專家混合」模型提供指引,內容涵蓋上下文並行處理規模設定、選擇性重新計算、調度器選擇,以及近期實驗中得出的實用模式。

...展開全部
3
更新時間 2026-09-28

MoE 長上下文訓練

Stable 文件:@docs/training/moe-optimization.md 卡片:@skills/nemo-mbridge-perf-moe-long-context/card.yaml

長上下文中的變化

一旦序列長度遠遠超過 4K 級別,注意力記憶與 激發函數駐留便成為主要限制因素。對於 MoE 模型而言,這 通常意味著您需要結合以下幾種方式:

  • 上下文並行處理
  • 選擇性重新計算
  • 降低精確度
  • 將優化器狀態卸載至 CPU
  • 不浪費剩餘較小 DP 預算的調度器與 PP 佈局

四捨五入的擴展模式

H100 上的 DSV3

DSV3 的長上下文執行結果顯示出穩定的模式:

  • 一旦超越 最短執行上下文後,選擇性重新計算的表現便優於全量重新計算
  • 若適當增加 CP,從中等長度到極長 上下文期間,吞吐量將維持在相當狹窄的區間內
  • 隨著 CP 增加,權衡重點會從「記憶體容納度」轉移至「GPU 數量可行性」

換言之,若 佈局選擇得當,長上下文並不會立即導致利用率急遽下降,但會極快地耗盡 DP 預算。

GB200 上的 Qwen3-Next

Qwen3-Next 的表現更像是一個對記憶體敏感的中規模模型:

  • 在適中的 CP 下,8K 和 32K 仍具實用性
  • 64K 雖可實現,但吞吐量下降明顯,且記憶體空間變得 更加緊繃
  • 管線佈局與分組 GEMM 的優化,其重要性幾乎可與 CP 相提並論

Qwen3 235B 在 GB200 上的表現

Qwen3 235B 顯示,當 TP、CP 和 HybridEP 協調運作時,長上下文在 NVL72 系統上仍能保持高效能。最佳的 128K 級配置 並非僅是「僅適合特定情況」的方案;若能平衡佈線、 並行度與重新運算,它們仍能維持極高效率。

CP 尺寸設定經驗法則

  1. 從 4K 碎片目標開始:一個不錯的初步估算值是 CP ≈ seq_len / 4096,然後四捨五入至實用的 2 的冪次佈局。

  2. 若可能,請盡量維持 DP 的運作:一旦 CP、 EP、TP 和 PP 共同將 DP 壓縮至下限,長上下文的擴展性便會變得脆弱。

  3. 優先採用選擇性重新計算:在進行全域重新計算之前,應先重新計算up_proj、norm、 moe、moe_act 或mlp等模組。

  4. 在極長上下文中避免大量使用 SDPA 的重新計算:重新計算注意力 內部結構可能會增加大量工作量,但相較於重新計算 較小的 MoE 和 MLP 側模組,其記憶體效益卻較低。

  5. 在 NVL72 系統上將 TP 作為另一項調控手段:GB200 和 GB300 的執行過程 有時可在維持效率的同時,以部分 CP 換取 TP。

  6. 預設 GBS 需要縮減:隨著 CP 增加而 DP 減少,您可能需要 減少全域批次大小,或接受更高的 GA。

代表性配置家族

H100 上的 128K DSV3

TP=1  CP=32  EP=32  PP=8  VPP=4
精確度:FP8 級
調度器:DeepEP
重新計算:up_proj、norm、moe、mlp
額外記憶體支援:優化器 CPU 卸載

H100 上 256K 的 DSV3

TP=1  CP=64  EP=32  PP=8  EDP=2  VPP=4
精確度:FP8 級
調度器:DeepEP
重新計算:up_proj、norm、moe、mlp
額外記憶體支援:優化器 CPU 卸載

Qwen3 235B 於 GB200 上設定為 128K

TP=4  CP=4  EP=32  PP=4  VPP=12
精確度:BF16 或 MXFP8
調度器:HybridEP
重新計算:moe_act、norm
CUDA 圖:attn + moe_router + moe_preprocess

重新計算與 CUDA 圖指引

針對長上下文 MoE 訓練:

  • 請先從選擇性重新計算開始
  • 僅在形狀與路由路徑穩定後才新增 CUDA 圖
  • 使用 CUDA 圖時,請保持序列長度和 MBS 不變
  • 若訓練運行取決於高度動態的批次,應優先採用貪婪執行

實用參考資料:

  • @docs/training/activation-recomputation.md
  • @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md

常見陷阱

  1. CP 並不會取代 EP 或 PP:它只是增添了另一層次;並不會讓 其他方法消失。

  2. 一個良好的 4K 基準模型,在長上下文情境下仍可能表現不佳:路由模式、 重新計算選項以及卸載策略往往需要調整。

  3. GPU 數量可行性成為真正的限制:極長的上下文在單一方案中 可能看起來沒問題,但一旦將 EP 和 PP 誠實地應用於 整個模型時,便會變得無法實現。

  4. CUDA 圖需要靜態結構:可變長度的批次和機會性 填充策略可能會在不知不覺中破壞路徑。

  5. 在 128K+ 的情境下,容器與核心的支援更為關鍵:長上下文路徑 相較於短上下文的啟動過程,往往更依賴較新的核心版本與錯誤修復。

在 GitHub 上查看
---
name: nemo-mbridge-perf-moe-long-context
description: Provides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments.
license: Apache-2.0
---

# MoE Long-Context Training

Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-long-context/card.yaml

## What Changes At Long Context

Once sequence length moves well past the 4K-class regime, attention memory and
activation residency become the dominant constraints. For MoE models, that
usually means you need some combination of:

- context parallelism
- selective recompute
- lower precision
- CPU offload for optimizer state
- a dispatcher and PP layout that do not waste the smaller remaining DP budget

## Rounded Scaling Patterns

### DSV3 on H100

The DSV3 long-context runs show a stable pattern:

- selective recompute works better than full recompute once you move past the
  shortest contexts
- throughput stays in a fairly narrow band from mid-length through very long
  contexts if CP is increased appropriately
- the trade shifts from "memory fit" to "GPU-count feasibility" as CP grows

In other words, long context does not immediately collapse utilization if the
layout is chosen well, but it does consume the DP budget very quickly.

### Qwen3-Next on GB200

Qwen3-Next behaves more like a memory-sensitive medium-scale model:

- 8K and 32K remain practical with moderate CP
- 64K is possible, but the throughput drop is noticeable and memory becomes
  much tighter
- pipeline layout and grouped-GEMM improvements matter almost as much as CP

### Qwen3 235B on GB200

Qwen3 235B shows that long context can still be efficient on NVL72 systems when
TP, CP, and HybridEP are coordinated. The best 128K-class configurations are
not just "fit-only" recipes; they can remain highly efficient if routing,
parallelism, and recompute are balanced.

## CP Sizing Rules Of Thumb

1. **Start from a 4K shard target**: a good first guess is
   `CP ~= seq_len / 4096`, then round to a practical power-of-two layout.

2. **Keep DP alive if possible**: long-context scaling becomes brittle once CP,
   EP, TP, and PP together squeeze DP down to the floor.

3. **Prefer selective recompute**: recompute modules such as `up_proj`, `norm`,
   `moe`, `moe_act`, or `mlp` before reaching for full recompute.

4. **Avoid SDPA-heavy recompute at very long context**: recomputing attention
   internals can add a lot of work for less memory benefit than recomputing
   smaller MoE and MLP-side modules.

5. **Use TP as another lever on NVL72 systems**: GB200 and GB300 runs can
   sometimes trade some CP for TP while still staying efficient.

6. **Assume GBS will need to shrink**: as CP rises and DP falls, you may need
   to reduce global batch size or accept higher GA.

## Representative Config Families

### DSV3 at 128K on H100

```text
TP=1  CP=32  EP=32  PP=8  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```

### DSV3 at 256K on H100

```text
TP=1  CP=64  EP=32  PP=8  EDP=2  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```

### Qwen3 235B at 128K on GB200

```text
TP=4  CP=4  EP=32  PP=4  VPP=12
Precision: BF16 or MXFP8
Dispatcher: HybridEP
Recompute: moe_act, norm
CUDA Graph: attn + moe_router + moe_preprocess
```

## Recompute And CUDA Graph Guidance

For long-context MoE training:

- start with selective recompute
- add CUDA graphs only after the shapes and routing path are stable
- keep sequence length and MBS fixed when using CUDA graphs
- if the run depends on highly dynamic batches, prefer eager execution

Useful references:

- @docs/training/activation-recomputation.md
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md

## Pitfalls

1. **CP does not replace EP or PP**: it adds another dimension; it does not make
   the others disappear.

2. **A good 4K baseline can still be a bad long-context baseline**: routing mode,
   recompute choice, and offload strategy often need to change.

3. **GPU-count feasibility becomes the real constraint**: very long context can
   look fine in a single recipe, then become impossible once EP and PP are added
   honestly across the full model.

4. **CUDA graphs need static shapes**: variable-length batches and opportunistic
   padding strategies can silently break the path.

5. **Container and kernel support matters more at 128K+**: long-context paths
   tend to rely on newer kernels and bug fixes than short-context bring-up does.

所有檔案

1 個檔案

安裝 nemo-mbridge-perf-moe-long-context

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-long-context # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

web-search
更新時間 2026-06-29
webapp-testing
更新時間 2026-06-29
lark-base
更新時間 2026-07-05
agentmail
更新時間 2026-06-29
OR