选项
首页首页 Skill 数据科学与机器学习 nemo-mbridge-perf-moe-long-context

nemo-mbridge-perf-moe-long-context

NVIDIA/skills NVIDIA/skills

本文为使用长上下文窗口训练“专家混合”模型提供了指导,内容涵盖上下文并行度的确定、选择性重新计算、调度器选择,以及近期实验中总结出的实用模式。

...展开全部
3
更新时间 2026-09-28

MoE 长上下文训练

Stable 文档:@docs/training/moe-optimization.md 卡片:@skills/nemo-mbridge-perf-moe-long-context/card.yaml

长上下文环境下的变化

一旦序列长度远超4K量级,注意力内存和 激活驻留便成为主要限制因素。对于MoE模型而言,这 通常意味着需要结合使用以下方法:

  • 上下文并行化
  • 选择性重计算
  • 降低精度
  • 优化器状态的CPU卸载
  • 调度器和PP布局应避免浪费剩余的较小DP预算

舍入缩放模式

H100上的DSV3

DSV3的长上下文运行显示出稳定的模式:

  • 一旦超过 最短上下文,选择性重计算的效果优于全重计算
  • 如果适当增加CP, 从中等长度到非常长的上下文中,吞吐量将保持在相当窄的范围内
  • 随着 CP 的增加,权衡点将从“内存适配”转向“GPU 数量可行性”

换言之,如果 布局选择得当,长上下文并不会立即导致利用率骤降,但会非常迅速地消耗DP预算。

GB200 上的 Qwen3-Next

Qwen3-Next 的表现更像是一个对内存敏感的中等规模模型:

  • 在适度的 CP 条件下,8K 和 32K 仍具实用性
  • 64K 虽然可行,但吞吐量下降明显,且内存空间变得 非常紧张
  • 流水线布局和分组GEMM的优化效果几乎与CP同等重要

GB200 上的 Qwen3 235B

Qwen3 235B 表明,当 TP、CP 和 HybridEP 协调配合时,长上下文在 NVL72 系统上仍能保持高效。最佳的 128K 级配置 并非仅仅是“仅适配”的方案;只要在布线、 并行度和重新计算之间取得平衡,它们仍能保持极高的效率。

CP 规模确定经验法则

  1. 从 4K 分片目标开始:一个合理的初步估算值是 CP ≈ seq_len / 4096,然后四舍五入到实用的 2 的幂布局。

  2. 尽可能保持 DP 存活:一旦 CP、 EP、TP 和 PP 共同将 DP 压缩至最低阈值,长上下文的可扩展性就会变得脆弱。

  3. 优先采用选择性重计算:在进行全量重计算之前,先对up_proj、norm、 moe、moe_act 或mlp等模块进行重计算。

  4. 在非常长的上下文中避免 SDPA 密集型重计算:重计算注意力 内部结构可能会增加大量工作量,而其带来的内存收益却不如重计算 较小的 MoE 和 MLP 侧模块。

  5. 在 NVL72 系统上将 TP 作为另一项调节手段:GB200 和 GB300 的运行 有时可以在保持高效的同时,通过牺牲部分 CP 来换取 TP。

  6. 假设 GBS 需要缩小:随着 CP 增加和 DP 减少,您可能需要 减小全局批量大小或接受更高的 GA。

代表性配置系列

H100上的128K DSV3

TP=1  CP=32  EP=32  PP=8  VPP=4
精度:FP8 级别
调度器:DeepEP
重新计算:up_proj、norm、moe、mlp
额外内存辅助:优化器 CPU 卸载

H100 上的 DSV3(256K)

TP=1  CP=64  EP=32  PP=8  EDP=2  VPP=4
精度:FP8级
调度器:DeepEP
重计算:up_proj、norm、moe、mlp
额外内存辅助:优化器 CPU 卸载

Qwen3 235B 在 GB200 上以 128K 运行

TP=4  CP=4  EP=32  PP=4  VPP=12
精度:BF16 或 MXFP8
调度器:HybridEP
重计算:moe_act、norm
CUDA 图:attn + moe_router + moe_preprocess

重计算与 CUDA 图指导

针对长上下文MoE训练:

  • 从选择性重新计算开始
  • 仅在上下文形状和路由路径稳定后才添加 CUDA 图
  • 使用 CUDA 图时,保持序列长度和 MBS 固定
  • 如果训练依赖于高度动态的批次,则优先采用即时执行

有用参考资料:

  • @docs/training/activation-recomputation.md
  • @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md

注意事项

  1. CP 并不能取代 EP 或 PP:它只是增添了另一维度,并不会让 其他方法消失。

  2. 一个优秀的4K基线模型,在长上下文场景下仍可能表现不佳:路由模式、 重计算方式的选择以及卸载策略通常需要调整。

  3. GPU数量的可行性成为真正的限制因素:非常长的上下文在 单一方案中可能表现良好,但一旦在整个模型中如实添加EP和PP后, 便会变得不可行。

  4. CUDA 图需要静态结构:可变长度批次和机会主义 填充策略可能会悄无声息地破坏路径。

  5. 在 128K+ 的情况下,容器和内核的支持更为重要:长上下文路径 往往比短上下文的启动更依赖于较新的内核和错误修复。

在 GitHub 上查看
---
name: nemo-mbridge-perf-moe-long-context
description: Provides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments.
license: Apache-2.0
---

# MoE Long-Context Training

Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-long-context/card.yaml

## What Changes At Long Context

Once sequence length moves well past the 4K-class regime, attention memory and
activation residency become the dominant constraints. For MoE models, that
usually means you need some combination of:

- context parallelism
- selective recompute
- lower precision
- CPU offload for optimizer state
- a dispatcher and PP layout that do not waste the smaller remaining DP budget

## Rounded Scaling Patterns

### DSV3 on H100

The DSV3 long-context runs show a stable pattern:

- selective recompute works better than full recompute once you move past the
  shortest contexts
- throughput stays in a fairly narrow band from mid-length through very long
  contexts if CP is increased appropriately
- the trade shifts from "memory fit" to "GPU-count feasibility" as CP grows

In other words, long context does not immediately collapse utilization if the
layout is chosen well, but it does consume the DP budget very quickly.

### Qwen3-Next on GB200

Qwen3-Next behaves more like a memory-sensitive medium-scale model:

- 8K and 32K remain practical with moderate CP
- 64K is possible, but the throughput drop is noticeable and memory becomes
  much tighter
- pipeline layout and grouped-GEMM improvements matter almost as much as CP

### Qwen3 235B on GB200

Qwen3 235B shows that long context can still be efficient on NVL72 systems when
TP, CP, and HybridEP are coordinated. The best 128K-class configurations are
not just "fit-only" recipes; they can remain highly efficient if routing,
parallelism, and recompute are balanced.

## CP Sizing Rules Of Thumb

1. **Start from a 4K shard target**: a good first guess is
   `CP ~= seq_len / 4096`, then round to a practical power-of-two layout.

2. **Keep DP alive if possible**: long-context scaling becomes brittle once CP,
   EP, TP, and PP together squeeze DP down to the floor.

3. **Prefer selective recompute**: recompute modules such as `up_proj`, `norm`,
   `moe`, `moe_act`, or `mlp` before reaching for full recompute.

4. **Avoid SDPA-heavy recompute at very long context**: recomputing attention
   internals can add a lot of work for less memory benefit than recomputing
   smaller MoE and MLP-side modules.

5. **Use TP as another lever on NVL72 systems**: GB200 and GB300 runs can
   sometimes trade some CP for TP while still staying efficient.

6. **Assume GBS will need to shrink**: as CP rises and DP falls, you may need
   to reduce global batch size or accept higher GA.

## Representative Config Families

### DSV3 at 128K on H100

```text
TP=1  CP=32  EP=32  PP=8  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```

### DSV3 at 256K on H100

```text
TP=1  CP=64  EP=32  PP=8  EDP=2  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```

### Qwen3 235B at 128K on GB200

```text
TP=4  CP=4  EP=32  PP=4  VPP=12
Precision: BF16 or MXFP8
Dispatcher: HybridEP
Recompute: moe_act, norm
CUDA Graph: attn + moe_router + moe_preprocess
```

## Recompute And CUDA Graph Guidance

For long-context MoE training:

- start with selective recompute
- add CUDA graphs only after the shapes and routing path are stable
- keep sequence length and MBS fixed when using CUDA graphs
- if the run depends on highly dynamic batches, prefer eager execution

Useful references:

- @docs/training/activation-recomputation.md
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md

## Pitfalls

1. **CP does not replace EP or PP**: it adds another dimension; it does not make
   the others disappear.

2. **A good 4K baseline can still be a bad long-context baseline**: routing mode,
   recompute choice, and offload strategy often need to change.

3. **GPU-count feasibility becomes the real constraint**: very long context can
   look fine in a single recipe, then become impossible once EP and PP are added
   honestly across the full model.

4. **CUDA graphs need static shapes**: variable-length batches and opportunistic
   padding strategies can silently break the path.

5. **Container and kernel support matters more at 128K+**: long-context paths
   tend to rely on newer kernels and bug fixes than short-context bring-up does.

所有文件

1 个文件

安装 nemo-mbridge-perf-moe-long-context

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-long-context # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 会自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

web-search
更新时间 2026-06-29
webapp-testing
更新时间 2026-06-29
lark-base
更新时间 2026-07-05
agentmail
更新时间 2026-06-29
OR