nemo-mbridge-perf-moe-long-context
NVIDIA/skills
긴 컨텍스트 윈도우를 사용하는 Mixture-of-Experts 모델 훈련에 대한 지침을 제공하며, 여기에는 컨텍스트 병렬 처리 규모 설정, 선택적 재계산, 디스패처 선택, 그리고 최근 실험에서 도출된 실용적인 패턴 등이 포함됩니다.
...모든 것을 확장하십시오MoE 장문 맥락 학습
Stable 문서: @docs/training/moe-optimization.md 카드: @skills/nemo-mbridge-perf-moe-long-context/card.yaml
긴 컨텍스트에서 달라지는 점
시퀀스 길이가 4K 단위를 훨씬 넘어가면, 어텐션 메모리와 활성화 값의 상주 시간이 주요 제약 요인이 됩니다. MoE 모델의 경우, 이는 일반적으로 다음 요소들의 조합이 필요함을 의미합니다:
- 컨텍스트 병렬 처리
- 선택적 재계산
- 정밀도 낮추기
- 최적화기 상태에 대한 CPU 오프로드
- 남은 DP 예산을 낭비하지 않는 디스패처 및 PP 레이아웃
반올림된 스케일링 패턴
H100에서의 DSV3
DSV3의 긴 컨텍스트 실행 결과는 안정적인 패턴을 보여줍니다:
- 가장 짧은 컨텍스트를 넘어서는 시점부터는 전체 재계산보다 선택적 재계산이 더 효과적입니다
- CP를 적절히 증가시키면 중간 길이에서 매우 긴 컨텍스트에 이르기까지 처리량이 상당히 좁은 범위 내에서 유지됩니다
- CP가 증가함에 따라 고려해야 할 요소가 “메모리 수용 가능 여부”에서 “GPU 수 실현 가능성”으로 전환됩니다
즉, 레이아웃이 적절히 선택된다면 긴 컨텍스트가 이용률을 즉시 떨어뜨리지는 않지만, DP 예산은 매우 빠르게 소모됩니다.
GB200에서의 Qwen3-Next
Qwen3-Next는 메모리에 민감한 중규모 모델과 유사한 특성을 보입니다:
- 적당한 CP 조건에서는 8K와 32K가 여전히 실용적입니다
- 64K도 가능하지만 처리량 저하가 눈에 띄며 메모리 여유가 훨씬 더 좁아집니다
- 파이프라인 레이아웃과 그룹화된 GEMM 개선 사항은 CP만큼이나 중요합니다
GB200에서의 Qwen3 235B
Qwen3 235B는 TP, CP 및 HybridEP가 조화롭게 작동할 때 NVL72 시스템에서도 긴 컨텍스트가 여전히 효율적일 수 있음을 보여줍니다. 최상의 128K급 구성은 단순히 "적합성만 고려한" 레시피가 아닙니다. 라우팅, 병렬 처리 및 재계산이 균형을 이룬다면 높은 효율성을 유지할 수 있습니다.
CP 크기 결정의 경험적 규칙
4K 샤드 목표값에서 시작하세요: 좋은 초기 추정치는
CP ~= seq_len / 4096이며, 이를 실용적인 2의 제곱 배수로 반올림합니다.가능하면 DP를 유지하십시오: CP, EP, TP 및 PP가 함께 DP를 최저 한계까지 압박하면, 긴 컨텍스트에 대한 확장성이 취약해집니다.
선택적 재계산을 우선시하십시오: 전체 재계산을 시도하기 전에
up_proj,norm,moe,moe_act또는mlp와같은 모듈을 재계산하십시오.매우 긴 컨텍스트에서는 SDPA 집약적인 재계산을 피하십시오: 어텐션 내부 구조를 재계산하는 것은 더 작은 MoE 및 MLP 측면의 모듈을 재계산하는 것보다 메모리 이득이 적으면서도 작업 부하를 크게 증가시킬 수 있습니다.
NVL72 시스템에서 TP를 또 다른 조정 수단으로 활용하십시오: GB200 및 GB300 실행 시 효율성을 유지하면서 CP를 일부 TP로 교환할 수 있는 경우가 있습니다.
GBS를 축소해야 할 것으로 가정하십시오: CP가 증가하고 DP가 감소함에 따라, 전역 배치 크기를 줄이거나 더 높은 GA를 수용해야 할 수 있습니다.
대표적인 구성 계열
H100에서 128K 규모의 DSV3
TP=1 CP=32 EP=32 PP=8 VPP=4
정밀도: FP8급
디스패처: DeepEP
재계산: up_proj, norm, moe, mlp
추가 메모리 지원: 최적화기 CPU 오프로드
H100에서 256K로 설정된 DSV3
TP=1 CP=64 EP=32 PP=8 EDP=2 VPP=4
정밀도: FP8급
디스패처: DeepEP
재계산: up_proj, norm, moe, mlp
추가 메모리 지원: 최적화기 CPU 오프로드
GB200에서 128K로 실행된 Qwen3 235B
TP=4 CP=4 EP=32 PP=4 VPP=12
정밀도: BF16 또는 MXFP8
디스패처: HybridEP
재계산: moe_act, norm
CUDA 그래프: attn + moe_router + moe_preprocess
재계산 및 CUDA 그래프 지침
긴 컨텍스트 MoE 훈련의 경우:
- 선택적 재계산으로 시작
- 셰이프와 라우팅 경로가 안정화된 후에만 CUDA 그래프를 추가하십시오
- CUDA 그래프를 사용할 때는 시퀀스 길이와 MBS를 고정하십시오
- 실행이 매우 동적인 배치에 의존하는 경우, 이거(eager) 실행을 우선적으로 선택하십시오
유용한 참고 자료:
- @docs/training/activation-recomputation.md
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md
주의 사항
CP는 EP나 PP를 대체하지 않습니다: CP는 또 다른 차원을 추가할 뿐, 다른 방식들을 없애지는 않습니다.
훌륭한 4K 베이스라인이라 해도 긴 컨텍스트에서는 부적절한 베이스라인이 될 수 있습니다. 라우팅 모드, 재계산 방식, 오프로드 전략은 종종 변경해야 합니다.
GPU 수에 따른 실현 가능성이 진정한 제약 조건이 됩니다. 매우 긴 컨텍스트는 단일 레시피에서는 괜찮아 보일 수 있지만, 전체 모델에 걸쳐 EP와 PP를 정확히 적용하면 불가능해질 수 있습니다.
CUDA 그래프는 정적인 형태를 가져야 합니다: 가변 길이 배치와 기회주의적 패딩 전략은 경로를 은밀하게 손상시킬 수 있습니다.
128K 이상에서는 컨테이너 및 커널 지원이 더욱 중요합니다: 긴 컨텍스트 경로는 짧은 컨텍스트의 초기 구동 단계보다 최신 커널과 버그 수정 사항에 의존하는 경향이 있습니다.
---
name: nemo-mbridge-perf-moe-long-context
description: Provides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments.
license: Apache-2.0
---
# MoE Long-Context Training
Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-long-context/card.yaml
## What Changes At Long Context
Once sequence length moves well past the 4K-class regime, attention memory and
activation residency become the dominant constraints. For MoE models, that
usually means you need some combination of:
- context parallelism
- selective recompute
- lower precision
- CPU offload for optimizer state
- a dispatcher and PP layout that do not waste the smaller remaining DP budget
## Rounded Scaling Patterns
### DSV3 on H100
The DSV3 long-context runs show a stable pattern:
- selective recompute works better than full recompute once you move past the
shortest contexts
- throughput stays in a fairly narrow band from mid-length through very long
contexts if CP is increased appropriately
- the trade shifts from "memory fit" to "GPU-count feasibility" as CP grows
In other words, long context does not immediately collapse utilization if the
layout is chosen well, but it does consume the DP budget very quickly.
### Qwen3-Next on GB200
Qwen3-Next behaves more like a memory-sensitive medium-scale model:
- 8K and 32K remain practical with moderate CP
- 64K is possible, but the throughput drop is noticeable and memory becomes
much tighter
- pipeline layout and grouped-GEMM improvements matter almost as much as CP
### Qwen3 235B on GB200
Qwen3 235B shows that long context can still be efficient on NVL72 systems when
TP, CP, and HybridEP are coordinated. The best 128K-class configurations are
not just "fit-only" recipes; they can remain highly efficient if routing,
parallelism, and recompute are balanced.
## CP Sizing Rules Of Thumb
1. **Start from a 4K shard target**: a good first guess is
`CP ~= seq_len / 4096`, then round to a practical power-of-two layout.
2. **Keep DP alive if possible**: long-context scaling becomes brittle once CP,
EP, TP, and PP together squeeze DP down to the floor.
3. **Prefer selective recompute**: recompute modules such as `up_proj`, `norm`,
`moe`, `moe_act`, or `mlp` before reaching for full recompute.
4. **Avoid SDPA-heavy recompute at very long context**: recomputing attention
internals can add a lot of work for less memory benefit than recomputing
smaller MoE and MLP-side modules.
5. **Use TP as another lever on NVL72 systems**: GB200 and GB300 runs can
sometimes trade some CP for TP while still staying efficient.
6. **Assume GBS will need to shrink**: as CP rises and DP falls, you may need
to reduce global batch size or accept higher GA.
## Representative Config Families
### DSV3 at 128K on H100
```text
TP=1 CP=32 EP=32 PP=8 VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```
### DSV3 at 256K on H100
```text
TP=1 CP=64 EP=32 PP=8 EDP=2 VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```
### Qwen3 235B at 128K on GB200
```text
TP=4 CP=4 EP=32 PP=4 VPP=12
Precision: BF16 or MXFP8
Dispatcher: HybridEP
Recompute: moe_act, norm
CUDA Graph: attn + moe_router + moe_preprocess
```
## Recompute And CUDA Graph Guidance
For long-context MoE training:
- start with selective recompute
- add CUDA graphs only after the shapes and routing path are stable
- keep sequence length and MBS fixed when using CUDA graphs
- if the run depends on highly dynamic batches, prefer eager execution
Useful references:
- @docs/training/activation-recomputation.md
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md
## Pitfalls
1. **CP does not replace EP or PP**: it adds another dimension; it does not make
the others disappear.
2. **A good 4K baseline can still be a bad long-context baseline**: routing mode,
recompute choice, and offload strategy often need to change.
3. **GPU-count feasibility becomes the real constraint**: very long context can
look fine in a single recipe, then become impossible once EP and PP are added
honestly across the full model.
4. **CUDA graphs need static shapes**: variable-length batches and opportunistic
padding strategies can silently break the path.
5. **Container and kernel support matters more at 128K+**: long-context paths
tend to rely on newer kernels and bug fixes than short-context bring-up does.
모든 파일
1개 파일nemo-mbridge-perf-moe-long-context 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-long-context # Copy SKILL.md to your .claude/skills/ directory
복사





집
