accelerated-computing-cudf
NVIDIA/skills
cuDF 및 dask-cuDF을 사용하여 GPU DataFrames로 pandas 워크플로우를 가속화하고, ETL, 조인, 그룹바이 및 대규모 데이터 처리를 수행하세요.
...모든 것을 확장하십시오cuDF 및 dask-cuDF 구현자 가이드
호환성
- 이 스킬에서 추적하는 릴리스: 26.04.
- CUDA 12 기준 NVIDIA 볼타 이상 또는 CUDA 13 기준 터밍 이상 필요. 릴리스 26.04는 드라이버 535+ 환경에서 CUDA 12.2-12.9, 드라이버 580+ 환경에서 CUDA 13.0-13.1, 그리고 Python 3.11-3.14를 지원합니다. cuDF 최적 성능 구간: 행 수 10만 개 이상.
명명 규칙
사용자 대상 답변에서는 NVIDIA 라이브러리 우선 문구를 사용하십시오. 출처를 인용할 때 literal RAPIDS/rapidsai URL, 패키지 이름 및 릴리스 메타데이터는 그대로 유지하십시오.
역할
귀하는 GPU 데이터프레임과 함께 작업하는 구현자를 돕는 cuDF 전문가입니다. 사용자는 pandas와 자신의 데이터를 이해하고 있으며, 귀하의 임무는 최소한의 마찰로 정확하고 빠른 GPU 코드를 작성하도록 돕는 것입니다. 사용자의 의도에 따라 경로를 선택하십시오: 광범위한 호환성 또는 최소 변경 가속화를 위한 cudf.pandas, 명시적 DataFrame 마이그레이션을 위한 명시적 cuDF, 핫 ETL 경로 및 패리티 민감 작업용 명시적 cuDF. 소스 스키마, 행 수, null 위치, 순서 및 숫자 허용오차는 사용자 가시적 동작으로 취급하십시오.
중요 규칙
- 올바른 cuDF 경로 선택. 광범위한 호환성 또는 최소 변경 가속화를 위해
cudf.pandas를 사용하십시오. 사용자가 DataFrame 코드 마이그레이션을 요청하거나, 패리티를 검사하거나, 가시적인 ETL 핫 경로를 최적화하거나, 지원되지 않는 작업을 제어할 때 명시적 cuDF를 사용하십시오. - 크기 제한: 최소 10만 행. 그 이하에서는 GPU 전송 오버헤드가 일반적으로 속도 향상보다 큽니다. 정확성을 위해 작은 데이터를 사용하고, 성능을 위해 더 큰 작업 집합을 벤치마킹하십시오.
- 변환은 경계에서만 수행. 표시, 플롯팅, CPU 전용 라이브러리 또는 최종 출력 경계를 위해
.to_pandas(),.values또는.numpy()를 사용하십시오. 중간 ETL 데이터는 GPU에 유지하십시오. - Float32은 당신의 친구입니다. float64에 대한 cuDF 연산은 느립니다. 정밀도가 허용되면 조기에 캐스팅하십시오.
- 대표적인 슬라이스에서 의미론을 검증하십시오. null 처리, 조인, 시계열, 재형성 또는 그룹화 논리의 경우, 패리티를 주장하기 전에 작은 pandas 참조 경로를 유지하고 형상, 레이블, null 수, 순서 및 대표 값을 비교하십시오.
- GPU 메모리 초과 데이터의 경우,
enable_cudf_spill=True로 dask-cuDF로 이동하십시오.references/dask-cudf-patterns.md를 참조하십시오.
GPU 데이터프레임으로 가는 세 가지 경로
경로 1: cudf.pandas 가속기 (호환성 / 최소 변경)
사용자가 작은 코드 변경, 서드파티 pandas 호환성 또는 지원되지 않는 작업이 폴백되는 동안 계속 실행될 수 있는 하나의 코드 경로가 필요한 경우에 사용하십시오.
Jupyter/IPython:
%load_ext cudf.pandas
import pandas as pd # 이제 GPU 백업; 지원되지 않는 작업에 대해 조용히 폴백
스크립트:
python -m cudf.pandas my_script.py
멀티프로세싱 사용 시:
import cudf.pandas
cudf.pandas.install() # pandas import 전, Pool 생성 전에 반드시 호출해야 함
from multiprocessing import Pool
속도 향상을 주장하기 전에 cudf.pandas 프로파일러로 가속을 확인하십시오. 노트북, CLI 및 통계 예제의 경우 references/cudf-pandas-accelerator.md를 읽으십시오. 프로파일에서 핫 경로가 CPU에서 실행되는 경우, 명시적 cuDF 제어를 위해 경로 2를 사용하십시오.
경로 2: 명시적 cuDF API
전체 제어, 핫 경로 최적화, 명시적 DataFrame 마이그레이션 및 패리티 민감 작업을 위해:
import cudf
# 데이터를 GPU에 직접 읽기
df = cudf.read_parquet("data.parquet")
# 연산은 pandas를 반영함
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# 문자열 연산
df["clean"] = df["name"].str.strip().str.lower()
# 마이그레이션에 확정하기 전에 API 커버리지를 확인하려면:
# 알려진 격차 및 우회 방법에 대해서는 references/api-patterns.md를 참조
데이터를 끝까지 GPU에 유지하십시오. 표시 또는 CPU 또는 비-GPU 이관을 위해 매우 마지막 단계에서만 .to_pandas()를 호출하십시오.
read_csv/read_parquet, 조인, groupby, 재형성, nullable 타입, fillna/where, 시간 버킷, 롤링 윈도우 또는 CPU/GPU 패리티 검사와 관련된 작업에는 명시적 cuDF를 선호하십시오. 성공적인 실행에만 의존하기보다 의미론이 중요한 경우 작은 CPU/GPU 검증 경로를 추가하십시오.
null 처리, 재형성 또는 시계열 동작이 있는 pandas 코드의 경우, 다시 작성하기 전에 관련 의미론 체크리스트를 위해 references/api-patterns.md를 읽으십시오. 최소 변경 요청에는 cudf.pandas 부트스트랩으로 충분하지만, 구현 요청은 핫 경로를 명시적이고 관찰 가능하게 만들어야 합니다.
재형성이 많은 pandas 코드(pivot_table, melt, stack/unstack, crosstab)의 경우, 인덱스 레이블, 열 레이블 또는 레벨, fill_value, aggfunc, 마진 및 정규화를 포함하여 소스 스키마를 계약의 일부로 유지하십시오. 동일한 기능이 지원되는 경우 명시적 cuDF를 사용하고, 모든 연산을 다시 작성하는 것보다 정확한 pandas 재형성 의미론이 더 중요한 경우 cudf.pandas 또는 좁은 호환성 경계를 사용하십시오. 최종화하기 전에 형상, 레이블 및 대표 값에 대한 작은 pandas 참조 패리티 검사를 추가하십시오. references/api-patterns.md를 참조하십시오.
경로 3: dask-cuDF (멀티-GPU / 대용량 데이터)
데이터셋이 GPU 메모리를 초과하는 경우. 전체 패턴에 대해서는 references/dask-cudf-patterns.md를 참조하십시오.
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # GPU당 하나의 워커
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
메모리 관리
OOM 발생 전에 스플을 활성화하십시오 (발생 이후가 아님):
import cudf
cudf.set_option("spill", True) # GPU가 가득 차면 호스트 RAM으로 스플
RMM 풀 할당자 (많은 할당이 있는 파이프라인에서 cudaMalloc 오버헤드 감소):
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# 모든 cuDF 연산 전에 호출해야 함
| GPU 여유 공간 vs 데이터셋 | 전략 |
|---|---|
| 여유 > 데이터셋의 2배 | 단일 GPU cuDF |
| 여유 1~2배 데이터셋 | cuDF + `cudf.set_option("spill", True)` |
| 데이터셋 > GPU 메모리 | dask-cuDF |
| 데이터셋 > 노드 메모리 | dask-cuDF + 멀티 노드 (accelerated-computing-mpf 참조) |
문제 해결
pandas 대비 속도 향상 없음:
- 데이터
%%cudf.pandas.profile실행 — 높은 CPU %는 많은 폴백을 의미합니다. 해당 연산을 식별하고 수정하십시오.- 알려진 격차에 대해서는
references/api-patterns.md를 확인하십시오.
OOM (CUDA 메모리 부족):
- 스플 활성화:
cudf.set_option("spill", True) - 할당자 단편화 또는 반복 할당 오버헤드가 보이는 경우, GPU 할당 전에
accelerated-computing-rmm메모리 리소스 설정 가이드를 사용하십시오. - 여전히 실패하는 경우: dask-cuDF로 이동
AttributeError / NotImplementedError:
- 특정 연동에 대한
references/api-patterns.md확인 - 해당 연산 하나를 좁은 경계에서 CPU에 유지하고 지원되는 파이프라인은 GPU에서 계속 실행
- 지원되지 않는 연산에만
.to_pandas()를 사용하고, 다시.from_pandas()로 복원
pandas 대비 잘못된 결과:
- null/NaN 처리가 다름: cuDF는 기본적으로
<na></na>(nullable)를 사용하고, pandas는NaN을 사용합니다.references/api-patterns.md를 참조하십시오. - 정렬 안정성:
stable=True가 전달되지 않는 한 cuDF 정렬은 안정성이 보장되지 않음 - 차이가 부동 소수점 차이로 인한 경우, 더 높은 정밀도의 부동 소수점(예:
float32대신float64)으로 캐스팅해 보십시오. 결과가 여전히 다른 경우 중단하십시오. GPU와 CPU 알고리즘은 부동 소수점 산술의 비결합성으로 인해 부동 소수점 숫자에 대해 항상 다른 결과를 생성하며, 이는 수정할 수 없습니다.
Nullable 및 Fill 의미론
사용자가 명시적으로 pandas nullable dtypes, fillna, where/mask 또는 그룹화된 null 동작에 관심이 있는 경우, 패리티 검사를 구현의 일부로 취급하십시오. nullable dtype 예제는 references/api-patterns.md를 참조하십시오.
- 소스 코드가 이미 그렇게 하지 않는 한, 센티널 값으로 채우는 대신 nullable integer/string 열을 보존하십시오.
- 조건을 인코딩할 때
where/mask의미론을 유지하십시오. 조건이 정확히 null-only인 경우에만 광범위한fillna를 사용하십시오. - pandas 참조가 nullable 확장 dtypes를 사용하는 경우
to_pandas(nullable=True)로 비교하십시오. - 미래의 변경 사항이 동일한 nullable 변환 및 집계 검사를 수행하도록, GPU 경로 옆에 재사용 가능한 도우미에 패리티 검사를 배치하십시오.
- 의미론적 패리티를 주장하기 전에 행 수, null 수, 마스크 진리표, 그룹화 집계 및 대표 dtypes를 검증하십시오.
참고 파일
references/cudf-pandas-accelerator.md— 프로파일링, 폴백 감지, cudf.pandas 심층 분석references/api-patterns.md— 알려진 API 격차, 우회 방법, 의미론적 차이references/dask-cudf-patterns.md— 멀티-GPU 패턴, 모범 사례, 파티션 튜닝
외부 문서
WebFetch를 사용하여 자세한 API 서명, 매개변수 설명 및 예제를 필요에 따라 검색하십시오.
---
name: accelerated-computing-cudf
description: Accelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
license: CC-BY-4.0 AND Apache-2.0
---
# cuDF & dask-cuDF Implementer's Guide
## Compatibility
- Release tracked by this skill: 26.04.
- Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.
## Naming
Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.
## Role
You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: `cudf.pandas` for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.
## Critical Rules
1. **Choose the right cuDF path.** Use `cudf.pandas` for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
2. **Size gate: 100K rows minimum.** Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
3. **Keep conversions at boundaries.** Use `.to_pandas()`, `.values`, or `.numpy()` for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
4. **Float32 is your friend.** cuDF operations on float64 are slower; cast early when precision allows.
5. **Validate semantics on representative slices.** For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
6. **For data > GPU memory**, move to dask-cuDF with `enable_cudf_spill=True`. See `references/dask-cudf-patterns.md`.
## Three Paths to GPU DataFrames
### Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)
Use when the user needs a small code change, third-party pandas compatibility,
or one code path that can keep running while unsupported operations fall back.
**Jupyter/IPython:**
```python
%load_ext cudf.pandas
import pandas as pd # now GPU-backed; falls back silently for unsupported ops
```
**Script:**
```bash
python -m cudf.pandas my_script.py
```
**With multiprocessing:**
```python
import cudf.pandas
cudf.pandas.install() # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool
```
Confirm acceleration with the cudf.pandas profiler before claiming speedup.
For notebook, CLI, and stats examples, read
`references/cudf-pandas-accelerator.md`. If the profile shows the hot path
running on CPU, use Path 2 for explicit cuDF control.
### Path 2: Explicit cuDF API
For full control, hot-path optimization, named DataFrame migrations, and
parity-sensitive operations:
```python
import cudf
# Read data directly to GPU
df = cudf.read_parquet("data.parquet")
# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# String operations
df["clean"] = df["name"].str.strip().str.lower()
# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds
```
**Keep data on GPU end-to-end.** Only call `.to_pandas()` at the very end for display or CPU or non-GPU handoff.
Prefer explicit cuDF for tasks involving `read_csv`/`read_parquet`, joins,
groupby, reshape, nullable types, `fillna`/`where`, time buckets, rolling
windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when
semantics matter instead of relying on successful execution alone.
For pandas code with null handling, reshape, or time-series behavior, read
`references/api-patterns.md` for the relevant semantic checklist before
rewriting. A `cudf.pandas` bootstrap is enough for a minimal-change request; an
implementation request should make the hot path explicit and observable.
For reshape-heavy pandas code (`pivot_table`, `melt`, `stack`/`unstack`,
`crosstab`), keep the source schema as part of the contract: index labels,
column labels or levels, `fill_value`, `aggfunc`, margins, and normalization.
Use explicit cuDF where the equivalent is supported; use `cudf.pandas` or a
narrow compatibility boundary when exact pandas reshape semantics matter more
than rewriting every operation. Add a small pandas-reference parity check for
shape, labels, and representative values before finalizing. See
`references/api-patterns.md`.
### Path 3: dask-cuDF (Multi-GPU / Large Data)
When dataset exceeds GPU memory. See `references/dask-cudf-patterns.md` for full patterns.
```python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # one worker per GPU
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
```
## Memory Management
**Enable spill before OOM happens** (not after):
```python
import cudf
cudf.set_option("spill", True) # spill to host RAM when GPU is full
```
**RMM pool allocator** (reduces cudaMalloc overhead in pipelines with many allocations):
```python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
```
| GPU Free vs Dataset | Strategy |
|---|---|
| Free > 2× dataset | Single GPU cuDF |
| Free 1–2× dataset | cuDF + `cudf.set_option("spill", True)` |
| Dataset > GPU mem | dask-cuDF |
| Dataset > node mem | dask-cuDF + multi-node (see accelerated-computing-mpf) |
## Troubleshooting
**No speedup vs pandas:**
- Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
- Run `%%cudf.pandas.profile` — high CPU % means many fallbacks. Identify and fix those ops.
- Check `references/api-patterns.md` for known gaps.
**OOM (CUDA out of memory):**
1. Enable spill: `cudf.set_option("spill", True)`
2. If allocator fragmentation or repeated allocation overhead is visible, use the `accelerated-computing-rmm` memory-resource setup guidance before GPU allocations
3. Still failing: move to dask-cuDF
**AttributeError / NotImplementedError:**
- Check `references/api-patterns.md` for the specific operation
- Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
- Use `.to_pandas()` only for the unsupported op, then `.from_pandas()` back
**Wrong results vs pandas:**
- Null/NaN handling differs: cuDF uses `<NA>` (nullable) by default, pandas uses `NaN`. See `references/api-patterns.md`.
- Sort stability: cuDF sort is not guaranteed stable unless `stable=True` is passed
- If the difference is due to floating point differences, try casting to higher precision floats (e.g. `float64` instead of `float32`). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.
## Nullable and Fill Semantics
When the user explicitly cares about pandas nullable dtypes, `fillna`,
`where`/`mask`, or grouped null behavior, treat parity checks as part of the
implementation. See `references/api-patterns.md` for nullable dtype examples.
- Preserve nullable integer/string columns instead of filling them with sentinel
values unless the source code already did that.
- Keep `where`/`mask` semantics when they encode a condition. Use broad
`fillna` only when the condition is exactly null-only.
- Compare with `to_pandas(nullable=True)` when the pandas reference uses
nullable extension dtypes.
- Put the parity check in a reusable helper next to the GPU path, so future
changes exercise the same nullable conversion and aggregation checks.
- Validate row counts, null counts, mask truth tables, grouped aggregates, and
representative dtypes before claiming semantic parity.
## Reference Files
- `references/cudf-pandas-accelerator.md` — Profiling, fallback detection, cudf.pandas deep dive
- `references/api-patterns.md` — Known API gaps, workarounds, semantic differences
- `references/dask-cudf-patterns.md` — Multi-GPU patterns, best practices, partition tuning
## External Documentation
Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.
- **cuDF Documentation:** https://docs.rapids.ai/api/cudf/stable/
- **dask-cuDF API Reference:** https://docs.rapids.ai/api/dask-cudf/stable/api/
- **GitHub:** https://github.com/rapidsai/cudf
- **CHANGELOG:** https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md
accelerated-computing-cudf 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉토리에 추출하세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/accelerated-computing-cudf # Copy SKILL.md to your .claude/skills/ directory
복사





집
