accelerated-computing-cudf
NVIDIA/skills
使用 cuDF 和 dask-cuDF 加速基於 GPU 資料幀的 pandas 工作流,適用於 ETL、連線、groupby 以及大規模資料處理。
...展開全部cuDF 與 dask-cuDF 實現者指南
相容性
- 本技能跟蹤的版本:26.04。
- 在 CUDA 12 上需要 NVIDIA Volta 或更新架構,或在 CUDA 13 上需要 Turing 或更新架構。26.04 版本支援 CUDA 12.2-12.9(驅動版本 535+)或 CUDA 13.0-13.1(驅動版本 580+),以及 Python 3.11-3.14。cuDF 的最佳適用場景:資料行數大於 10 萬行。
命名規範
在面向使用者的回答中,優先使用 NVIDIA 庫的術語。在引用來源時,保留字面的 RAPIDS/rapidsai URL、包名和版本後設資料。
角色定義
你是 cuDF 專家,協助實現者處理 GPU DataFrame。使用者已理解 pandas 及其資料——你的任務是以最小的摩擦幫助他們編寫正確、高效的 GPU 程式碼。根據使用者意圖選擇路徑:若需廣泛相容性或最小變更加速,使用 cudf.pandas;若需顯式遷移 DataFrame 程式碼、最佳化 ETL 關鍵路徑或處理對一致性敏感的工作,使用顯式 cuDF。將源模式、行數、空值位置、排序順序和數值容差視為使用者可見的行為。
關鍵規則
- 選擇合適的 cuDF 路徑。 若需廣泛相容性或最小變更加速,使用
cudf.pandas。若使用者要求遷移 DataFrame 程式碼、檢查一致性、最佳化可見的 ETL 關鍵路徑或控制不支援的操作,則使用顯式 cuDF。 - 大小限制:最低 10 萬行。 低於此行數時,GPU 資料傳輸開銷通常會抵消加速收益;使用小資料進行正確性驗證,並對更大的工作集進行效能基準測試。
- 轉換保持在邊界處。 對於顯示、繪圖、僅 CPU 庫或最終輸出邊界,使用
.to_pandas()、.values或.numpy()。中間 ETL 資料應保留在 GPU 上。 - Float32 是你的朋友。 cuDF 對 float64 的操作較慢;當精度允許時,請儘早轉換型別。
- 在代表性切片上驗證語義。 對於空值處理、連線、時間序列、重塑或分組邏輯,保留一個小型 pandas 參考路徑,並在聲稱一致性之前比較形狀、標籤、空值計數、排序順序和代表性值。
- 對於超過 GPU 記憶體的資料,請遷移到 dask-cuDF 並設定
enable_cudf_spill=True。參見references/dask-cudf-patterns.md。
通往 GPU DataFrame 的三條路徑
路徑 1:cudf.pandas 加速器(相容性 / 最小變更)
當使用者需要少量程式碼更改、第三方 pandas 相容性,或需要一個在支援的操作回退時仍能執行的單一程式碼路徑時使用。
Jupyter/IPython:
%load_ext cudf.pandas
import pandas as pd # 現在由 GPU 支援;不支援的操作會靜默回退
指令碼:
python -m cudf.pandas my_script.py
使用多程序:
import cudf.pandas
cudf.pandas.install() # 必須在匯入 pandas 之前、建立 Pool 之前呼叫
from multiprocessing import Pool
在聲稱加速之前,請使用 cudf.pandas 分析器進行確認。對於筆記本、CLI 和統計示例,請閱讀 references/cudf-pandas-accelerator.md。如果分析顯示關鍵路徑在 CPU 上執行,請使用路徑 2 進行顯式 cuDF 控制。
路徑 2:顯式 cuDF API
如需完全控制、關鍵路徑最佳化、命名 DataFrame 遷移以及對一致性敏感的操作:
import cudf
# 直接將資料讀取到 GPU
df = cudf.read_parquet("data.parquet")
# 操作與 pandas 類似
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# 字串操作
df["clean"] = df["name"].str.strip().str.lower()
# 在提交遷移之前檢查 API 覆蓋範圍:
# 參見 references/api-patterns.md 以獲取已知缺陷和變通方法
使資料端到端保留在 GPU 上。 僅在最後呼叫 .to_pandas() 用於顯示、CPU 處理或非 GPU 交接。
對於涉及 read_csv/read_parquet、連線、groupby、重塑、可空型別、fillna/where、時間桶、滾動視窗或 CPU/GPU 一致性檢查的任務,優先使用顯式 cuDF。當語義重要時,新增小型 CPU/GPU 驗證路徑,而不是僅依賴執行成功。
對於涉及空值處理、重塑或時間序列行為的 pandas 程式碼,在重寫之前請閱讀 references/api-patterns.md 以獲取相關語義檢查清單。對於最小變更請求,cudf.pandas 載入程式已足夠;對於實現請求,應使關鍵路徑顯式且可觀察。
對於重塑密集的 pandas 程式碼(pivot_table、melt、stack/unstack、crosstab),將源模式作為合同的一部分:索引標籤、列標籤或層級、fill_value、aggfunc、邊距和歸一化。在等效操作受支援的地方使用顯式 cuDF;當精確的 pandas 重塑語義比重寫每個操作更重要時,使用 cudf.pandas 或狹窄的相容性邊界。在最終確定之前,新增小型 pandas 參考一致性檢查以驗證形狀、標籤和代表性值。參見 references/api-patterns.md。
路徑 3:dask-cuDF(多 GPU / 大資料)
當資料集超過 GPU 記憶體時。參見 references/dask-cudf-patterns.md 以獲取完整模式。
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # 每個 GPU 一個工作程序
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
記憶體管理
在發生 OOM(記憶體溢位)之前啟用溢位(而不是之後):
import cudf
cudf.set_option("spill", True) # 當 GPU 滿時溢位到主機 RAM
RMM 池分配器(減少具有大量分配的操作流程中的 cudaMalloc 開銷):
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# 必須在任何 cuDF 操作之前呼叫
| GPU 空閒空間與資料集大小對比 | 策略 |
|---|---|
| 空閒空間 > 資料集大小的 2 倍 | 單 GPU cuDF |
| 空閒空間為資料集大小的 1-2 倍 | cuDF + `cudf.set_option("spill", True)` |
| 資料集大小 > GPU 記憶體 | dask-cuDF |
| 資料集大小 > 節點記憶體 | dask-cuDF + 多節點(參見 accelerated-computing-mpf) |
故障排除
與 pandas 相比無加速:
- 資料
- 執行
%%cudf.pandas.profile— 高 CPU % 意味著存在大量回退。識別並修復這些操作。 - 檢查
references/api-patterns.md以獲取已知缺陷。
OOM(CUDA 記憶體溢位):
- 啟用溢位:
cudf.set_option("spill", True) - 如果可見分配器碎片或重複分配開銷,請在 GPU 分配之前使用
accelerated-computing-rmm記憶體資源設定指南 - 仍然失敗:遷移到 dask-cuDF
AttributeError / NotImplementedError:
- 檢查
references/api-patterns.md以獲取特定操作 - 在狹窄邊界上將一個操作保留在 CPU 上,並在 GPU 上繼續支援的操作流程
- 僅對不支援的操作使用
.to_pandas(),然後使用.from_pandas()轉換回來
與 pandas 相比結果錯誤:
- 空值/NaN 處理不同:cuDF 預設使用
<na></na>(可空),pandas 使用NaN。參見references/api-patterns.md。 - 排序穩定性:除非傳遞
stable=True,否則 cuDF 排序不保證穩定 - 如果差異是由於浮點數差異引起的,請嘗試轉換為更高精度的浮點數(例如
float64而不是float32)。如果結果仍然不同,請停止。由於浮點運算的非結合性,GPU 和 CPU 演算法在浮點數上始終會產生不同的結果,這是無法修復的。
可空值與填充語義
當使用者明確關注 pandas 可空資料型別、fillna、where/mask 或分組空值行為時,將一致性檢查視為實現的一部分。參見 references/api-patterns.md 以獲取可空資料型別示例。
- 保留可空整數/字串列,除非原始碼已經用哨兵值填充它們。
- 當
where/mask編碼條件時,請保留其語義。僅當條件恰好為空時,才使用廣泛的fillna。 - 當 pandas 參考使用可空擴充套件資料型別時,請使用
to_pandas(nullable=True)進行比較。 - 在 GPU 路徑旁邊放置一個可重用的輔助函式以進行一致性檢查,以便未來的更改能夠執行相同的可空轉換和聚合檢查。
- 在聲稱語義一致性之前,驗證行數、空值計數、掩碼真值表、分組聚合和代表性資料型別。
參考檔案
references/cudf-pandas-accelerator.md— 分析、回退檢測、cudf.pandas 深入解析references/api-patterns.md— 已知 API 缺陷、變通方法、語義差異references/dask-cudf-patterns.md— 多 GPU 模式、最佳實踐、分割槽調優
外部文件
使用 WebFetch 按需檢索詳細的 API 簽名、引數描述和示例。
- cuDF 文件: https://docs.rapids.ai/api/cudf/stable/
- dask-cuDF API 參考: https://docs.rapids.ai/api/dask-cudf/stable/api/
- GitHub: https://github.com/rapidsai/cudf
- CHANGELOG: https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md
---
name: accelerated-computing-cudf
description: Accelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
license: CC-BY-4.0 AND Apache-2.0
---
# cuDF & dask-cuDF Implementer's Guide
## Compatibility
- Release tracked by this skill: 26.04.
- Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.
## Naming
Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.
## Role
You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: `cudf.pandas` for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.
## Critical Rules
1. **Choose the right cuDF path.** Use `cudf.pandas` for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
2. **Size gate: 100K rows minimum.** Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
3. **Keep conversions at boundaries.** Use `.to_pandas()`, `.values`, or `.numpy()` for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
4. **Float32 is your friend.** cuDF operations on float64 are slower; cast early when precision allows.
5. **Validate semantics on representative slices.** For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
6. **For data > GPU memory**, move to dask-cuDF with `enable_cudf_spill=True`. See `references/dask-cudf-patterns.md`.
## Three Paths to GPU DataFrames
### Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)
Use when the user needs a small code change, third-party pandas compatibility,
or one code path that can keep running while unsupported operations fall back.
**Jupyter/IPython:**
```python
%load_ext cudf.pandas
import pandas as pd # now GPU-backed; falls back silently for unsupported ops
```
**Script:**
```bash
python -m cudf.pandas my_script.py
```
**With multiprocessing:**
```python
import cudf.pandas
cudf.pandas.install() # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool
```
Confirm acceleration with the cudf.pandas profiler before claiming speedup.
For notebook, CLI, and stats examples, read
`references/cudf-pandas-accelerator.md`. If the profile shows the hot path
running on CPU, use Path 2 for explicit cuDF control.
### Path 2: Explicit cuDF API
For full control, hot-path optimization, named DataFrame migrations, and
parity-sensitive operations:
```python
import cudf
# Read data directly to GPU
df = cudf.read_parquet("data.parquet")
# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# String operations
df["clean"] = df["name"].str.strip().str.lower()
# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds
```
**Keep data on GPU end-to-end.** Only call `.to_pandas()` at the very end for display or CPU or non-GPU handoff.
Prefer explicit cuDF for tasks involving `read_csv`/`read_parquet`, joins,
groupby, reshape, nullable types, `fillna`/`where`, time buckets, rolling
windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when
semantics matter instead of relying on successful execution alone.
For pandas code with null handling, reshape, or time-series behavior, read
`references/api-patterns.md` for the relevant semantic checklist before
rewriting. A `cudf.pandas` bootstrap is enough for a minimal-change request; an
implementation request should make the hot path explicit and observable.
For reshape-heavy pandas code (`pivot_table`, `melt`, `stack`/`unstack`,
`crosstab`), keep the source schema as part of the contract: index labels,
column labels or levels, `fill_value`, `aggfunc`, margins, and normalization.
Use explicit cuDF where the equivalent is supported; use `cudf.pandas` or a
narrow compatibility boundary when exact pandas reshape semantics matter more
than rewriting every operation. Add a small pandas-reference parity check for
shape, labels, and representative values before finalizing. See
`references/api-patterns.md`.
### Path 3: dask-cuDF (Multi-GPU / Large Data)
When dataset exceeds GPU memory. See `references/dask-cudf-patterns.md` for full patterns.
```python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # one worker per GPU
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
```
## Memory Management
**Enable spill before OOM happens** (not after):
```python
import cudf
cudf.set_option("spill", True) # spill to host RAM when GPU is full
```
**RMM pool allocator** (reduces cudaMalloc overhead in pipelines with many allocations):
```python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
```
| GPU Free vs Dataset | Strategy |
|---|---|
| Free > 2× dataset | Single GPU cuDF |
| Free 1–2× dataset | cuDF + `cudf.set_option("spill", True)` |
| Dataset > GPU mem | dask-cuDF |
| Dataset > node mem | dask-cuDF + multi-node (see accelerated-computing-mpf) |
## Troubleshooting
**No speedup vs pandas:**
- Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
- Run `%%cudf.pandas.profile` — high CPU % means many fallbacks. Identify and fix those ops.
- Check `references/api-patterns.md` for known gaps.
**OOM (CUDA out of memory):**
1. Enable spill: `cudf.set_option("spill", True)`
2. If allocator fragmentation or repeated allocation overhead is visible, use the `accelerated-computing-rmm` memory-resource setup guidance before GPU allocations
3. Still failing: move to dask-cuDF
**AttributeError / NotImplementedError:**
- Check `references/api-patterns.md` for the specific operation
- Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
- Use `.to_pandas()` only for the unsupported op, then `.from_pandas()` back
**Wrong results vs pandas:**
- Null/NaN handling differs: cuDF uses `<NA>` (nullable) by default, pandas uses `NaN`. See `references/api-patterns.md`.
- Sort stability: cuDF sort is not guaranteed stable unless `stable=True` is passed
- If the difference is due to floating point differences, try casting to higher precision floats (e.g. `float64` instead of `float32`). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.
## Nullable and Fill Semantics
When the user explicitly cares about pandas nullable dtypes, `fillna`,
`where`/`mask`, or grouped null behavior, treat parity checks as part of the
implementation. See `references/api-patterns.md` for nullable dtype examples.
- Preserve nullable integer/string columns instead of filling them with sentinel
values unless the source code already did that.
- Keep `where`/`mask` semantics when they encode a condition. Use broad
`fillna` only when the condition is exactly null-only.
- Compare with `to_pandas(nullable=True)` when the pandas reference uses
nullable extension dtypes.
- Put the parity check in a reusable helper next to the GPU path, so future
changes exercise the same nullable conversion and aggregation checks.
- Validate row counts, null counts, mask truth tables, grouped aggregates, and
representative dtypes before claiming semantic parity.
## Reference Files
- `references/cudf-pandas-accelerator.md` — Profiling, fallback detection, cudf.pandas deep dive
- `references/api-patterns.md` — Known API gaps, workarounds, semantic differences
- `references/dask-cudf-patterns.md` — Multi-GPU patterns, best practices, partition tuning
## External Documentation
Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.
- **cuDF Documentation:** https://docs.rapids.ai/api/cudf/stable/
- **dask-cuDF API Reference:** https://docs.rapids.ai/api/dask-cudf/stable/api/
- **GitHub:** https://github.com/rapidsai/cudf
- **CHANGELOG:** https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md
安裝 accelerated-computing-cudf
將技能檔案下載並解壓至你的 .claude/skills/ 目錄。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/accelerated-computing-cudf # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
