選項
首頁首頁 Skill 資料庫管理 accelerated-computing-cudf

accelerated-computing-cudf

NVIDIA/skills NVIDIA/skills

使用 cuDF 和 dask-cuDF 加速基於 GPU 資料幀的 pandas 工作流,適用於 ETL、連線、groupby 以及大規模資料處理。

...展開全部
2
更新時間 2026-09-27

cuDF 與 dask-cuDF 實現者指南

相容性

  • 本技能跟蹤的版本:26.04。
  • 在 CUDA 12 上需要 NVIDIA Volta 或更新架構,或在 CUDA 13 上需要 Turing 或更新架構。26.04 版本支援 CUDA 12.2-12.9(驅動版本 535+)或 CUDA 13.0-13.1(驅動版本 580+),以及 Python 3.11-3.14。cuDF 的最佳適用場景:資料行數大於 10 萬行。

命名規範

在面向使用者的回答中,優先使用 NVIDIA 庫的術語。在引用來源時,保留字面的 RAPIDS/rapidsai URL、包名和版本後設資料。

角色定義

你是 cuDF 專家,協助實現者處理 GPU DataFrame。使用者已理解 pandas 及其資料——你的任務是以最小的摩擦幫助他們編寫正確、高效的 GPU 程式碼。根據使用者意圖選擇路徑:若需廣泛相容性或最小變更加速,使用 cudf.pandas;若需顯式遷移 DataFrame 程式碼、最佳化 ETL 關鍵路徑或處理對一致性敏感的工作,使用顯式 cuDF。將源模式、行數、空值位置、排序順序和數值容差視為使用者可見的行為。

關鍵規則

  1. 選擇合適的 cuDF 路徑。 若需廣泛相容性或最小變更加速,使用 cudf.pandas。若使用者要求遷移 DataFrame 程式碼、檢查一致性、最佳化可見的 ETL 關鍵路徑或控制不支援的操作,則使用顯式 cuDF。
  2. 大小限制:最低 10 萬行。 低於此行數時,GPU 資料傳輸開銷通常會抵消加速收益;使用小資料進行正確性驗證,並對更大的工作集進行效能基準測試。
  3. 轉換保持在邊界處。 對於顯示、繪圖、僅 CPU 庫或最終輸出邊界,使用 .to_pandas()、.values 或 .numpy()。中間 ETL 資料應保留在 GPU 上。
  4. Float32 是你的朋友。 cuDF 對 float64 的操作較慢;當精度允許時,請儘早轉換型別。
  5. 在代表性切片上驗證語義。 對於空值處理、連線、時間序列、重塑或分組邏輯,保留一個小型 pandas 參考路徑,並在聲稱一致性之前比較形狀、標籤、空值計數、排序順序和代表性值。
  6. 對於超過 GPU 記憶體的資料,請遷移到 dask-cuDF 並設定 enable_cudf_spill=True。參見 references/dask-cudf-patterns.md。

通往 GPU DataFrame 的三條路徑

路徑 1:cudf.pandas 加速器(相容性 / 最小變更)

當使用者需要少量程式碼更改、第三方 pandas 相容性,或需要一個在支援的操作回退時仍能執行的單一程式碼路徑時使用。

Jupyter/IPython:

%load_ext cudf.pandas
import pandas as pd   # 現在由 GPU 支援;不支援的操作會靜默回退

指令碼:

python -m cudf.pandas my_script.py

使用多程序:

import cudf.pandas
cudf.pandas.install()   # 必須在匯入 pandas 之前、建立 Pool 之前呼叫
from multiprocessing import Pool

在聲稱加速之前,請使用 cudf.pandas 分析器進行確認。對於筆記本、CLI 和統計示例,請閱讀 references/cudf-pandas-accelerator.md。如果分析顯示關鍵路徑在 CPU 上執行,請使用路徑 2 進行顯式 cuDF 控制。

路徑 2:顯式 cuDF API

如需完全控制、關鍵路徑最佳化、命名 DataFrame 遷移以及對一致性敏感的操作:

import cudf

# 直接將資料讀取到 GPU
df = cudf.read_parquet("data.parquet")

# 操作與 pandas 類似
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]

# 字串操作
df["clean"] = df["name"].str.strip().str.lower()

# 在提交遷移之前檢查 API 覆蓋範圍:
# 參見 references/api-patterns.md 以獲取已知缺陷和變通方法

使資料端到端保留在 GPU 上。 僅在最後呼叫 .to_pandas() 用於顯示、CPU 處理或非 GPU 交接。

對於涉及 read_csv/read_parquet、連線、groupby、重塑、可空型別、fillna/where、時間桶、滾動視窗或 CPU/GPU 一致性檢查的任務,優先使用顯式 cuDF。當語義重要時,新增小型 CPU/GPU 驗證路徑,而不是僅依賴執行成功。

對於涉及空值處理、重塑或時間序列行為的 pandas 程式碼,在重寫之前請閱讀 references/api-patterns.md 以獲取相關語義檢查清單。對於最小變更請求,cudf.pandas 載入程式已足夠;對於實現請求,應使關鍵路徑顯式且可觀察。

對於重塑密集的 pandas 程式碼(pivot_table、melt、stack/unstack、crosstab),將源模式作為合同的一部分:索引標籤、列標籤或層級、fill_value、aggfunc、邊距和歸一化。在等效操作受支援的地方使用顯式 cuDF;當精確的 pandas 重塑語義比重寫每個操作更重要時,使用 cudf.pandas 或狹窄的相容性邊界。在最終確定之前,新增小型 pandas 參考一致性檢查以驗證形狀、標籤和代表性值。參見 references/api-patterns.md。

路徑 3:dask-cuDF(多 GPU / 大資料)

當資料集超過 GPU 記憶體時。參見 references/dask-cudf-patterns.md 以獲取完整模式。

from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf

cluster = LocalCUDACluster(enable_cudf_spill=True)  # 每個 GPU 一個工作程序
client = Client(cluster)

ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()

記憶體管理

在發生 OOM(記憶體溢位)之前啟用溢位(而不是之後):

import cudf
cudf.set_option("spill", True)   # 當 GPU 滿時溢位到主機 RAM

RMM 池分配器(減少具有大量分配的操作流程中的 cudaMalloc 開銷):

import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# 必須在任何 cuDF 操作之前呼叫
GPU 空閒空間與資料集大小對比策略
空閒空間 > 資料集大小的 2 倍單 GPU cuDF
空閒空間為資料集大小的 1-2 倍cuDF + `cudf.set_option("spill", True)`
資料集大小 > GPU 記憶體dask-cuDF
資料集大小 > 節點記憶體dask-cuDF + 多節點(參見 accelerated-computing-mpf)

故障排除

與 pandas 相比無加速:

  • 資料
  • 執行 %%cudf.pandas.profile — 高 CPU % 意味著存在大量回退。識別並修復這些操作。
  • 檢查 references/api-patterns.md 以獲取已知缺陷。

OOM(CUDA 記憶體溢位):

  1. 啟用溢位:cudf.set_option("spill", True)
  2. 如果可見分配器碎片或重複分配開銷,請在 GPU 分配之前使用 accelerated-computing-rmm 記憶體資源設定指南
  3. 仍然失敗:遷移到 dask-cuDF

AttributeError / NotImplementedError:

  • 檢查 references/api-patterns.md 以獲取特定操作
  • 在狹窄邊界上將一個操作保留在 CPU 上,並在 GPU 上繼續支援的操作流程
  • 僅對不支援的操作使用 .to_pandas(),然後使用 .from_pandas() 轉換回來

與 pandas 相比結果錯誤:

  • 空值/NaN 處理不同:cuDF 預設使用 <na></na>(可空),pandas 使用 NaN。參見 references/api-patterns.md。
  • 排序穩定性:除非傳遞 stable=True,否則 cuDF 排序不保證穩定
  • 如果差異是由於浮點數差異引起的,請嘗試轉換為更高精度的浮點數(例如 float64 而不是 float32)。如果結果仍然不同,請停止。由於浮點運算的非結合性,GPU 和 CPU 演算法在浮點數上始終會產生不同的結果,這是無法修復的。

可空值與填充語義

當使用者明確關注 pandas 可空資料型別、fillna、where/mask 或分組空值行為時,將一致性檢查視為實現的一部分。參見 references/api-patterns.md 以獲取可空資料型別示例。

  • 保留可空整數/字串列,除非原始碼已經用哨兵值填充它們。
  • 當 where/mask 編碼條件時,請保留其語義。僅當條件恰好為空時,才使用廣泛的 fillna。
  • 當 pandas 參考使用可空擴充套件資料型別時,請使用 to_pandas(nullable=True) 進行比較。
  • 在 GPU 路徑旁邊放置一個可重用的輔助函式以進行一致性檢查,以便未來的更改能夠執行相同的可空轉換和聚合檢查。
  • 在聲稱語義一致性之前,驗證行數、空值計數、掩碼真值表、分組聚合和代表性資料型別。

參考檔案

  • references/cudf-pandas-accelerator.md — 分析、回退檢測、cudf.pandas 深入解析
  • references/api-patterns.md — 已知 API 缺陷、變通方法、語義差異
  • references/dask-cudf-patterns.md — 多 GPU 模式、最佳實踐、分割槽調優

外部文件

使用 WebFetch 按需檢索詳細的 API 簽名、引數描述和示例。

在 GitHub 上查看
---
name: accelerated-computing-cudf
description: Accelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
license: CC-BY-4.0 AND Apache-2.0
---

# cuDF & dask-cuDF Implementer's Guide

## Compatibility

- Release tracked by this skill: 26.04.
- Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.

## Naming

Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.

## Role

You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: `cudf.pandas` for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.

## Critical Rules

1. **Choose the right cuDF path.** Use `cudf.pandas` for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
2. **Size gate: 100K rows minimum.** Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
3. **Keep conversions at boundaries.** Use `.to_pandas()`, `.values`, or `.numpy()` for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
4. **Float32 is your friend.** cuDF operations on float64 are slower; cast early when precision allows.
5. **Validate semantics on representative slices.** For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
6. **For data > GPU memory**, move to dask-cuDF with `enable_cudf_spill=True`. See `references/dask-cudf-patterns.md`.

## Three Paths to GPU DataFrames

### Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)

Use when the user needs a small code change, third-party pandas compatibility,
or one code path that can keep running while unsupported operations fall back.

**Jupyter/IPython:**
```python
%load_ext cudf.pandas
import pandas as pd   # now GPU-backed; falls back silently for unsupported ops
```

**Script:**
```bash
python -m cudf.pandas my_script.py
```

**With multiprocessing:**
```python
import cudf.pandas
cudf.pandas.install()   # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool
```

Confirm acceleration with the cudf.pandas profiler before claiming speedup.
For notebook, CLI, and stats examples, read
`references/cudf-pandas-accelerator.md`. If the profile shows the hot path
running on CPU, use Path 2 for explicit cuDF control.

### Path 2: Explicit cuDF API

For full control, hot-path optimization, named DataFrame migrations, and
parity-sensitive operations:

```python
import cudf

# Read data directly to GPU
df = cudf.read_parquet("data.parquet")

# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]

# String operations
df["clean"] = df["name"].str.strip().str.lower()

# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds
```

**Keep data on GPU end-to-end.** Only call `.to_pandas()` at the very end for display or CPU or non-GPU handoff.

Prefer explicit cuDF for tasks involving `read_csv`/`read_parquet`, joins,
groupby, reshape, nullable types, `fillna`/`where`, time buckets, rolling
windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when
semantics matter instead of relying on successful execution alone.

For pandas code with null handling, reshape, or time-series behavior, read
`references/api-patterns.md` for the relevant semantic checklist before
rewriting. A `cudf.pandas` bootstrap is enough for a minimal-change request; an
implementation request should make the hot path explicit and observable.

For reshape-heavy pandas code (`pivot_table`, `melt`, `stack`/`unstack`,
`crosstab`), keep the source schema as part of the contract: index labels,
column labels or levels, `fill_value`, `aggfunc`, margins, and normalization.
Use explicit cuDF where the equivalent is supported; use `cudf.pandas` or a
narrow compatibility boundary when exact pandas reshape semantics matter more
than rewriting every operation. Add a small pandas-reference parity check for
shape, labels, and representative values before finalizing. See
`references/api-patterns.md`.

### Path 3: dask-cuDF (Multi-GPU / Large Data)

When dataset exceeds GPU memory. See `references/dask-cudf-patterns.md` for full patterns.

```python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf

cluster = LocalCUDACluster(enable_cudf_spill=True)  # one worker per GPU
client = Client(cluster)

ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
```

## Memory Management

**Enable spill before OOM happens** (not after):
```python
import cudf
cudf.set_option("spill", True)   # spill to host RAM when GPU is full
```

**RMM pool allocator** (reduces cudaMalloc overhead in pipelines with many allocations):
```python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
```

| GPU Free vs Dataset | Strategy |
|---|---|
| Free > 2× dataset | Single GPU cuDF |
| Free 1–2× dataset | cuDF + `cudf.set_option("spill", True)` |
| Dataset > GPU mem | dask-cuDF |
| Dataset > node mem | dask-cuDF + multi-node (see accelerated-computing-mpf) |

## Troubleshooting

**No speedup vs pandas:**
- Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
- Run `%%cudf.pandas.profile` — high CPU % means many fallbacks. Identify and fix those ops.
- Check `references/api-patterns.md` for known gaps.

**OOM (CUDA out of memory):**
1. Enable spill: `cudf.set_option("spill", True)`
2. If allocator fragmentation or repeated allocation overhead is visible, use the `accelerated-computing-rmm` memory-resource setup guidance before GPU allocations
3. Still failing: move to dask-cuDF

**AttributeError / NotImplementedError:**
- Check `references/api-patterns.md` for the specific operation
- Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
- Use `.to_pandas()` only for the unsupported op, then `.from_pandas()` back

**Wrong results vs pandas:**
- Null/NaN handling differs: cuDF uses `<NA>` (nullable) by default, pandas uses `NaN`. See `references/api-patterns.md`.
- Sort stability: cuDF sort is not guaranteed stable unless `stable=True` is passed
- If the difference is due to floating point differences, try casting to higher precision floats (e.g. `float64` instead of `float32`). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.

## Nullable and Fill Semantics

When the user explicitly cares about pandas nullable dtypes, `fillna`,
`where`/`mask`, or grouped null behavior, treat parity checks as part of the
implementation. See `references/api-patterns.md` for nullable dtype examples.

- Preserve nullable integer/string columns instead of filling them with sentinel
  values unless the source code already did that.
- Keep `where`/`mask` semantics when they encode a condition. Use broad
  `fillna` only when the condition is exactly null-only.
- Compare with `to_pandas(nullable=True)` when the pandas reference uses
  nullable extension dtypes.
- Put the parity check in a reusable helper next to the GPU path, so future
  changes exercise the same nullable conversion and aggregation checks.
- Validate row counts, null counts, mask truth tables, grouped aggregates, and
  representative dtypes before claiming semantic parity.

## Reference Files

- `references/cudf-pandas-accelerator.md` — Profiling, fallback detection, cudf.pandas deep dive
- `references/api-patterns.md` — Known API gaps, workarounds, semantic differences
- `references/dask-cudf-patterns.md` — Multi-GPU patterns, best practices, partition tuning

## External Documentation

Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.

- **cuDF Documentation:** https://docs.rapids.ai/api/cudf/stable/
- **dask-cuDF API Reference:** https://docs.rapids.ai/api/dask-cudf/stable/api/
- **GitHub:** https://github.com/rapidsai/cudf
- **CHANGELOG:** https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md

所有檔案

34 個檔案

安裝 accelerated-computing-cudf

將技能檔案下載並解壓至你的 .claude/skills/ 目錄。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/accelerated-computing-cudf # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ 目錄。Claude 將自動檢測並使用該技能。
儲存庫 NVIDIA/skills

相關技能

microservices-patterns
更新時間 2026-06-29
jpa-patterns
更新時間 2026-06-30
fabric-lakehouse
更新時間 2026-06-30
prisma-expert
更新時間 2026-06-29
OR