옵션
집집 Skill 데이터 과학 및 ML nemo-mbridge-perf-moe-long-context

nemo-mbridge-perf-moe-long-context

NVIDIA/skills NVIDIA/skills

긴 컨텍스트 윈도우를 사용하는 Mixture-of-Experts 모델 훈련에 대한 지침을 제공하며, 여기에는 컨텍스트 병렬 처리 규모 설정, 선택적 재계산, 디스패처 선택, 그리고 최근 실험에서 도출된 실용적인 패턴 등이 포함됩니다.

...모든 것을 확장하십시오
3
업데이트 된 시간 2026년 9월 28일

MoE 장문 맥락 학습

Stable 문서: @docs/training/moe-optimization.md 카드: @skills/nemo-mbridge-perf-moe-long-context/card.yaml

긴 컨텍스트에서 달라지는 점

시퀀스 길이가 4K 단위를 훨씬 넘어가면, 어텐션 메모리와 활성화 값의 상주 시간이 주요 제약 요인이 됩니다. MoE 모델의 경우, 이는 일반적으로 다음 요소들의 조합이 필요함을 의미합니다:

  • 컨텍스트 병렬 처리
  • 선택적 재계산
  • 정밀도 낮추기
  • 최적화기 상태에 대한 CPU 오프로드
  • 남은 DP 예산을 낭비하지 않는 디스패처 및 PP 레이아웃

반올림된 스케일링 패턴

H100에서의 DSV3

DSV3의 긴 컨텍스트 실행 결과는 안정적인 패턴을 보여줍니다:

  • 가장 짧은 컨텍스트를 넘어서는 시점부터는 전체 재계산보다 선택적 재계산이 더 효과적입니다
  • CP를 적절히 증가시키면 중간 길이에서 매우 긴 컨텍스트에 이르기까지 처리량이 상당히 좁은 범위 내에서 유지됩니다
  • CP가 증가함에 따라 고려해야 할 요소가 “메모리 수용 가능 여부”에서 “GPU 수 실현 가능성”으로 전환됩니다

즉, 레이아웃이 적절히 선택된다면 긴 컨텍스트가 이용률을 즉시 떨어뜨리지는 않지만, DP 예산은 매우 빠르게 소모됩니다.

GB200에서의 Qwen3-Next

Qwen3-Next는 메모리에 민감한 중규모 모델과 유사한 특성을 보입니다:

  • 적당한 CP 조건에서는 8K와 32K가 여전히 실용적입니다
  • 64K도 가능하지만 처리량 저하가 눈에 띄며 메모리 여유가 훨씬 더 좁아집니다
  • 파이프라인 레이아웃과 그룹화된 GEMM 개선 사항은 CP만큼이나 중요합니다

GB200에서의 Qwen3 235B

Qwen3 235B는 TP, CP 및 HybridEP가 조화롭게 작동할 때 NVL72 시스템에서도 긴 컨텍스트가 여전히 효율적일 수 있음을 보여줍니다. 최상의 128K급 구성은 단순히 "적합성만 고려한" 레시피가 아닙니다. 라우팅, 병렬 처리 및 재계산이 균형을 이룬다면 높은 효율성을 유지할 수 있습니다.

CP 크기 결정의 경험적 규칙

  1. 4K 샤드 목표값에서 시작하세요: 좋은 초기 추정치는 CP ~= seq_len / 4096이며, 이를 실용적인 2의 제곱 배수로 반올림합니다.

  2. 가능하면 DP를 유지하십시오: CP, EP, TP 및 PP가 함께 DP를 최저 한계까지 압박하면, 긴 컨텍스트에 대한 확장성이 취약해집니다.

  3. 선택적 재계산을 우선시하십시오: 전체 재계산을 시도하기 전에 up_proj, norm, moe, moe_act 또는 mlp와 같은 모듈을 재계산하십시오.

  4. 매우 긴 컨텍스트에서는 SDPA 집약적인 재계산을 피하십시오: 어텐션 내부 구조를 재계산하는 것은 더 작은 MoE 및 MLP 측면의 모듈을 재계산하는 것보다 메모리 이득이 적으면서도 작업 부하를 크게 증가시킬 수 있습니다.

  5. NVL72 시스템에서 TP를 또 다른 조정 수단으로 활용하십시오: GB200 및 GB300 실행 시 효율성을 유지하면서 CP를 일부 TP로 교환할 수 있는 경우가 있습니다.

  6. GBS를 축소해야 할 것으로 가정하십시오: CP가 증가하고 DP가 감소함에 따라, 전역 배치 크기를 줄이거나 더 높은 GA를 수용해야 할 수 있습니다.

대표적인 구성 계열

H100에서 128K 규모의 DSV3

TP=1  CP=32  EP=32  PP=8  VPP=4
정밀도: FP8급
디스패처: DeepEP
재계산: up_proj, norm, moe, mlp
추가 메모리 지원: 최적화기 CPU 오프로드

H100에서 256K로 설정된 DSV3

TP=1  CP=64  EP=32  PP=8  EDP=2  VPP=4
정밀도: FP8급
디스패처: DeepEP
재계산: up_proj, norm, moe, mlp
추가 메모리 지원: 최적화기 CPU 오프로드

GB200에서 128K로 실행된 Qwen3 235B

TP=4  CP=4  EP=32  PP=4  VPP=12
정밀도: BF16 또는 MXFP8
디스패처: HybridEP
재계산: moe_act, norm
CUDA 그래프: attn + moe_router + moe_preprocess

재계산 및 CUDA 그래프 지침

긴 컨텍스트 MoE 훈련의 경우:

  • 선택적 재계산으로 시작
  • 셰이프와 라우팅 경로가 안정화된 후에만 CUDA 그래프를 추가하십시오
  • CUDA 그래프를 사용할 때는 시퀀스 길이와 MBS를 고정하십시오
  • 실행이 매우 동적인 배치에 의존하는 경우, 이거(eager) 실행을 우선적으로 선택하십시오

유용한 참고 자료:

  • @docs/training/activation-recomputation.md
  • @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md

주의 사항

  1. CP는 EP나 PP를 대체하지 않습니다: CP는 또 다른 차원을 추가할 뿐, 다른 방식들을 없애지는 않습니다.

  2. 훌륭한 4K 베이스라인이라 해도 긴 컨텍스트에서는 부적절한 베이스라인이 될 수 있습니다. 라우팅 모드, 재계산 방식, 오프로드 전략은 종종 변경해야 합니다.

  3. GPU 수에 따른 실현 가능성이 진정한 제약 조건이 됩니다. 매우 긴 컨텍스트는 단일 레시피에서는 괜찮아 보일 수 있지만, 전체 모델에 걸쳐 EP와 PP를 정확히 적용하면 불가능해질 수 있습니다.

  4. CUDA 그래프는 정적인 형태를 가져야 합니다: 가변 길이 배치와 기회주의적 패딩 전략은 경로를 은밀하게 손상시킬 수 있습니다.

  5. 128K 이상에서는 컨테이너 및 커널 지원이 더욱 중요합니다: 긴 컨텍스트 경로는 짧은 컨텍스트의 초기 구동 단계보다 최신 커널과 버그 수정 사항에 의존하는 경향이 있습니다.

GitHub에서 보기
---
name: nemo-mbridge-perf-moe-long-context
description: Provides guidance for training Mixture-of-Experts models with long context windows, covering context parallelism sizing, selective recomputation, dispatcher choices, and practical patterns from recent experiments.
license: Apache-2.0
---

# MoE Long-Context Training

Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-long-context/card.yaml

## What Changes At Long Context

Once sequence length moves well past the 4K-class regime, attention memory and
activation residency become the dominant constraints. For MoE models, that
usually means you need some combination of:

- context parallelism
- selective recompute
- lower precision
- CPU offload for optimizer state
- a dispatcher and PP layout that do not waste the smaller remaining DP budget

## Rounded Scaling Patterns

### DSV3 on H100

The DSV3 long-context runs show a stable pattern:

- selective recompute works better than full recompute once you move past the
  shortest contexts
- throughput stays in a fairly narrow band from mid-length through very long
  contexts if CP is increased appropriately
- the trade shifts from "memory fit" to "GPU-count feasibility" as CP grows

In other words, long context does not immediately collapse utilization if the
layout is chosen well, but it does consume the DP budget very quickly.

### Qwen3-Next on GB200

Qwen3-Next behaves more like a memory-sensitive medium-scale model:

- 8K and 32K remain practical with moderate CP
- 64K is possible, but the throughput drop is noticeable and memory becomes
  much tighter
- pipeline layout and grouped-GEMM improvements matter almost as much as CP

### Qwen3 235B on GB200

Qwen3 235B shows that long context can still be efficient on NVL72 systems when
TP, CP, and HybridEP are coordinated. The best 128K-class configurations are
not just "fit-only" recipes; they can remain highly efficient if routing,
parallelism, and recompute are balanced.

## CP Sizing Rules Of Thumb

1. **Start from a 4K shard target**: a good first guess is
   `CP ~= seq_len / 4096`, then round to a practical power-of-two layout.

2. **Keep DP alive if possible**: long-context scaling becomes brittle once CP,
   EP, TP, and PP together squeeze DP down to the floor.

3. **Prefer selective recompute**: recompute modules such as `up_proj`, `norm`,
   `moe`, `moe_act`, or `mlp` before reaching for full recompute.

4. **Avoid SDPA-heavy recompute at very long context**: recomputing attention
   internals can add a lot of work for less memory benefit than recomputing
   smaller MoE and MLP-side modules.

5. **Use TP as another lever on NVL72 systems**: GB200 and GB300 runs can
   sometimes trade some CP for TP while still staying efficient.

6. **Assume GBS will need to shrink**: as CP rises and DP falls, you may need
   to reduce global batch size or accept higher GA.

## Representative Config Families

### DSV3 at 128K on H100

```text
TP=1  CP=32  EP=32  PP=8  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```

### DSV3 at 256K on H100

```text
TP=1  CP=64  EP=32  PP=8  EDP=2  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
```

### Qwen3 235B at 128K on GB200

```text
TP=4  CP=4  EP=32  PP=4  VPP=12
Precision: BF16 or MXFP8
Dispatcher: HybridEP
Recompute: moe_act, norm
CUDA Graph: attn + moe_router + moe_preprocess
```

## Recompute And CUDA Graph Guidance

For long-context MoE training:

- start with selective recompute
- add CUDA graphs only after the shapes and routing path are stable
- keep sequence length and MBS fixed when using CUDA graphs
- if the run depends on highly dynamic batches, prefer eager execution

Useful references:

- @docs/training/activation-recomputation.md
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md

## Pitfalls

1. **CP does not replace EP or PP**: it adds another dimension; it does not make
   the others disappear.

2. **A good 4K baseline can still be a bad long-context baseline**: routing mode,
   recompute choice, and offload strategy often need to change.

3. **GPU-count feasibility becomes the real constraint**: very long context can
   look fine in a single recipe, then become impossible once EP and PP are added
   honestly across the full model.

4. **CUDA graphs need static shapes**: variable-length batches and opportunistic
   padding strategies can silently break the path.

5. **Container and kernel support matters more at 128K+**: long-context paths
   tend to rely on newer kernels and bug fixes than short-context bring-up does.

모든 파일

1개 파일

nemo-mbridge-perf-moe-long-context 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-long-context # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용합니다.
저장소 NVIDIA/skills

관련 스킬

web-search
업데이트 된 시간 2026년 6월 29일
webapp-testing
업데이트 된 시간 2026년 6월 29일
lark-base
업데이트 된 시간 2026년 7월 5일
agentmail
업데이트 된 시간 2026년 6월 29일
OR