accelerated-computing-cudf
NVIDIA/skills
使用 cuDF 和 dask-cuDF 加速基于 GPU 数据帧的 pandas 工作流,适用于 ETL、连接、groupby 以及大规模数据处理。
...展开全部cuDF 与 dask-cuDF 实现者指南
兼容性
- 本技能跟踪的版本:26.04。
- 在 CUDA 12 上需要 NVIDIA Volta 或更新架构,或在 CUDA 13 上需要 Turing 或更新架构。26.04 版本支持 CUDA 12.2-12.9(驱动版本 535+)或 CUDA 13.0-13.1(驱动版本 580+),以及 Python 3.11-3.14。cuDF 的最佳适用场景:数据行数大于 10 万行。
命名规范
在面向用户的回答中,优先使用 NVIDIA 库的术语。在引用来源时,保留字面的 RAPIDS/rapidsai URL、包名和版本元数据。
角色定义
你是 cuDF 专家,协助实现者处理 GPU DataFrame。用户已理解 pandas 及其数据——你的任务是以最小的摩擦帮助他们编写正确、高效的 GPU 代码。根据用户意图选择路径:若需广泛兼容性或最小变更加速,使用 cudf.pandas;若需显式迁移 DataFrame 代码、优化 ETL 关键路径或处理对一致性敏感的工作,使用显式 cuDF。将源模式、行数、空值位置、排序顺序和数值容差视为用户可见的行为。
关键规则
- 选择合适的 cuDF 路径。 若需广泛兼容性或最小变更加速,使用
cudf.pandas。若用户要求迁移 DataFrame 代码、检查一致性、优化可见的 ETL 关键路径或控制不支持的操作,则使用显式 cuDF。 - 大小限制:最低 10 万行。 低于此行数时,GPU 数据传输开销通常会抵消加速收益;使用小数据进行正确性验证,并对更大的工作集进行性能基准测试。
- 转换保持在边界处。 对于显示、绘图、仅 CPU 库或最终输出边界,使用
.to_pandas()、.values或.numpy()。中间 ETL 数据应保留在 GPU 上。 - Float32 是你的朋友。 cuDF 对 float64 的操作较慢;当精度允许时,请尽早转换类型。
- 在代表性切片上验证语义。 对于空值处理、连接、时间序列、重塑或分组逻辑,保留一个小型 pandas 参考路径,并在声称一致性之前比较形状、标签、空值计数、排序顺序和代表性值。
- 对于超过 GPU 内存的数据,请迁移到 dask-cuDF 并设置
enable_cudf_spill=True。参见references/dask-cudf-patterns.md。
通往 GPU DataFrame 的三条路径
路径 1:cudf.pandas 加速器(兼容性 / 最小变更)
当用户需要少量代码更改、第三方 pandas 兼容性,或需要一个在支持的操作回退时仍能运行的单一代码路径时使用。
Jupyter/IPython:
%load_ext cudf.pandas
import pandas as pd # 现在由 GPU 支持;不支持的操作会静默回退
脚本:
python -m cudf.pandas my_script.py
使用多进程:
import cudf.pandas
cudf.pandas.install() # 必须在导入 pandas 之前、创建 Pool 之前调用
from multiprocessing import Pool
在声称加速之前,请使用 cudf.pandas 分析器进行确认。对于笔记本、CLI 和统计示例,请阅读 references/cudf-pandas-accelerator.md。如果分析显示关键路径在 CPU 上运行,请使用路径 2 进行显式 cuDF 控制。
路径 2:显式 cuDF API
如需完全控制、关键路径优化、命名 DataFrame 迁移以及对一致性敏感的操作:
import cudf
# 直接将数据读取到 GPU
df = cudf.read_parquet("data.parquet")
# 操作与 pandas 类似
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# 字符串操作
df["clean"] = df["name"].str.strip().str.lower()
# 在提交迁移之前检查 API 覆盖范围:
# 参见 references/api-patterns.md 以获取已知缺陷和变通方法
使数据端到端保留在 GPU 上。 仅在最后调用 .to_pandas() 用于显示、CPU 处理或非 GPU 交接。
对于涉及 read_csv/read_parquet、连接、groupby、重塑、可空类型、fillna/where、时间桶、滚动窗口或 CPU/GPU 一致性检查的任务,优先使用显式 cuDF。当语义重要时,添加小型 CPU/GPU 验证路径,而不是仅依赖执行成功。
对于涉及空值处理、重塑或时间序列行为的 pandas 代码,在重写之前请阅读 references/api-patterns.md 以获取相关语义检查清单。对于最小变更请求,cudf.pandas 引导程序已足够;对于实现请求,应使关键路径显式且可观察。
对于重塑密集的 pandas 代码(pivot_table、melt、stack/unstack、crosstab),将源模式作为合同的一部分:索引标签、列标签或层级、fill_value、aggfunc、边距和归一化。在等效操作受支持的地方使用显式 cuDF;当精确的 pandas 重塑语义比重写每个操作更重要时,使用 cudf.pandas 或狭窄的兼容性边界。在最终确定之前,添加小型 pandas 参考一致性检查以验证形状、标签和代表性值。参见 references/api-patterns.md。
路径 3:dask-cuDF(多 GPU / 大数据)
当数据集超过 GPU 内存时。参见 references/dask-cudf-patterns.md 以获取完整模式。
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # 每个 GPU 一个工作进程
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
内存管理
在发生 OOM(内存溢出)之前启用溢出(而不是之后):
import cudf
cudf.set_option("spill", True) # 当 GPU 满时溢出到主机 RAM
RMM 池分配器(减少具有大量分配的操作流程中的 cudaMalloc 开销):
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# 必须在任何 cuDF 操作之前调用
| GPU 空闲空间与数据集大小对比 | 策略 |
|---|---|
| 空闲空间 > 数据集大小的 2 倍 | 单 GPU cuDF |
| 空闲空间为数据集大小的 1-2 倍 | cuDF + `cudf.set_option("spill", True)` |
| 数据集大小 > GPU 内存 | dask-cuDF |
| 数据集大小 > 节点内存 | dask-cuDF + 多节点(参见 accelerated-computing-mpf) |
故障排除
与 pandas 相比无加速:
- 数据
- 运行
%%cudf.pandas.profile— 高 CPU % 意味着存在大量回退。识别并修复这些操作。 - 检查
references/api-patterns.md以获取已知缺陷。
OOM(CUDA 内存溢出):
- 启用溢出:
cudf.set_option("spill", True) - 如果可见分配器碎片或重复分配开销,请在 GPU 分配之前使用
accelerated-computing-rmm内存资源设置指南 - 仍然失败:迁移到 dask-cuDF
AttributeError / NotImplementedError:
- 检查
references/api-patterns.md以获取特定操作 - 在狭窄边界上将一个操作保留在 CPU 上,并在 GPU 上继续支持的操作流程
- 仅对不支持的操作使用
.to_pandas(),然后使用.from_pandas()转换回来
与 pandas 相比结果错误:
- 空值/NaN 处理不同:cuDF 默认使用
<na></na>(可空),pandas 使用NaN。参见references/api-patterns.md。 - 排序稳定性:除非传递
stable=True,否则 cuDF 排序不保证稳定 - 如果差异是由于浮点数差异引起的,请尝试转换为更高精度的浮点数(例如
float64而不是float32)。如果结果仍然不同,请停止。由于浮点运算的非结合性,GPU 和 CPU 算法在浮点数上始终会产生不同的结果,这是无法修复的。
可空值与填充语义
当用户明确关注 pandas 可空数据类型、fillna、where/mask 或分组空值行为时,将一致性检查视为实现的一部分。参见 references/api-patterns.md 以获取可空数据类型示例。
- 保留可空整数/字符串列,除非源代码已经用哨兵值填充它们。
- 当
where/mask编码条件时,请保留其语义。仅当条件恰好为空时,才使用广泛的fillna。 - 当 pandas 参考使用可空扩展数据类型时,请使用
to_pandas(nullable=True)进行比较。 - 在 GPU 路径旁边放置一个可重用的辅助函数以进行一致性检查,以便未来的更改能够执行相同的可空转换和聚合检查。
- 在声称语义一致性之前,验证行数、空值计数、掩码真值表、分组聚合和代表性数据类型。
参考文件
references/cudf-pandas-accelerator.md— 分析、回退检测、cudf.pandas 深入解析references/api-patterns.md— 已知 API 缺陷、变通方法、语义差异references/dask-cudf-patterns.md— 多 GPU 模式、最佳实践、分区调优
外部文档
使用 WebFetch 按需检索详细的 API 签名、参数描述和示例。
- cuDF 文档: https://docs.rapids.ai/api/cudf/stable/
- dask-cuDF API 参考: https://docs.rapids.ai/api/dask-cudf/stable/api/
- GitHub: https://github.com/rapidsai/cudf
- CHANGELOG: https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md
---
name: accelerated-computing-cudf
description: Accelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
license: CC-BY-4.0 AND Apache-2.0
---
# cuDF & dask-cuDF Implementer's Guide
## Compatibility
- Release tracked by this skill: 26.04.
- Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.
## Naming
Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.
## Role
You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: `cudf.pandas` for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.
## Critical Rules
1. **Choose the right cuDF path.** Use `cudf.pandas` for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
2. **Size gate: 100K rows minimum.** Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
3. **Keep conversions at boundaries.** Use `.to_pandas()`, `.values`, or `.numpy()` for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
4. **Float32 is your friend.** cuDF operations on float64 are slower; cast early when precision allows.
5. **Validate semantics on representative slices.** For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
6. **For data > GPU memory**, move to dask-cuDF with `enable_cudf_spill=True`. See `references/dask-cudf-patterns.md`.
## Three Paths to GPU DataFrames
### Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)
Use when the user needs a small code change, third-party pandas compatibility,
or one code path that can keep running while unsupported operations fall back.
**Jupyter/IPython:**
```python
%load_ext cudf.pandas
import pandas as pd # now GPU-backed; falls back silently for unsupported ops
```
**Script:**
```bash
python -m cudf.pandas my_script.py
```
**With multiprocessing:**
```python
import cudf.pandas
cudf.pandas.install() # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool
```
Confirm acceleration with the cudf.pandas profiler before claiming speedup.
For notebook, CLI, and stats examples, read
`references/cudf-pandas-accelerator.md`. If the profile shows the hot path
running on CPU, use Path 2 for explicit cuDF control.
### Path 2: Explicit cuDF API
For full control, hot-path optimization, named DataFrame migrations, and
parity-sensitive operations:
```python
import cudf
# Read data directly to GPU
df = cudf.read_parquet("data.parquet")
# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# String operations
df["clean"] = df["name"].str.strip().str.lower()
# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds
```
**Keep data on GPU end-to-end.** Only call `.to_pandas()` at the very end for display or CPU or non-GPU handoff.
Prefer explicit cuDF for tasks involving `read_csv`/`read_parquet`, joins,
groupby, reshape, nullable types, `fillna`/`where`, time buckets, rolling
windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when
semantics matter instead of relying on successful execution alone.
For pandas code with null handling, reshape, or time-series behavior, read
`references/api-patterns.md` for the relevant semantic checklist before
rewriting. A `cudf.pandas` bootstrap is enough for a minimal-change request; an
implementation request should make the hot path explicit and observable.
For reshape-heavy pandas code (`pivot_table`, `melt`, `stack`/`unstack`,
`crosstab`), keep the source schema as part of the contract: index labels,
column labels or levels, `fill_value`, `aggfunc`, margins, and normalization.
Use explicit cuDF where the equivalent is supported; use `cudf.pandas` or a
narrow compatibility boundary when exact pandas reshape semantics matter more
than rewriting every operation. Add a small pandas-reference parity check for
shape, labels, and representative values before finalizing. See
`references/api-patterns.md`.
### Path 3: dask-cuDF (Multi-GPU / Large Data)
When dataset exceeds GPU memory. See `references/dask-cudf-patterns.md` for full patterns.
```python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # one worker per GPU
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
```
## Memory Management
**Enable spill before OOM happens** (not after):
```python
import cudf
cudf.set_option("spill", True) # spill to host RAM when GPU is full
```
**RMM pool allocator** (reduces cudaMalloc overhead in pipelines with many allocations):
```python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
```
| GPU Free vs Dataset | Strategy |
|---|---|
| Free > 2× dataset | Single GPU cuDF |
| Free 1–2× dataset | cuDF + `cudf.set_option("spill", True)` |
| Dataset > GPU mem | dask-cuDF |
| Dataset > node mem | dask-cuDF + multi-node (see accelerated-computing-mpf) |
## Troubleshooting
**No speedup vs pandas:**
- Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
- Run `%%cudf.pandas.profile` — high CPU % means many fallbacks. Identify and fix those ops.
- Check `references/api-patterns.md` for known gaps.
**OOM (CUDA out of memory):**
1. Enable spill: `cudf.set_option("spill", True)`
2. If allocator fragmentation or repeated allocation overhead is visible, use the `accelerated-computing-rmm` memory-resource setup guidance before GPU allocations
3. Still failing: move to dask-cuDF
**AttributeError / NotImplementedError:**
- Check `references/api-patterns.md` for the specific operation
- Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
- Use `.to_pandas()` only for the unsupported op, then `.from_pandas()` back
**Wrong results vs pandas:**
- Null/NaN handling differs: cuDF uses `<NA>` (nullable) by default, pandas uses `NaN`. See `references/api-patterns.md`.
- Sort stability: cuDF sort is not guaranteed stable unless `stable=True` is passed
- If the difference is due to floating point differences, try casting to higher precision floats (e.g. `float64` instead of `float32`). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.
## Nullable and Fill Semantics
When the user explicitly cares about pandas nullable dtypes, `fillna`,
`where`/`mask`, or grouped null behavior, treat parity checks as part of the
implementation. See `references/api-patterns.md` for nullable dtype examples.
- Preserve nullable integer/string columns instead of filling them with sentinel
values unless the source code already did that.
- Keep `where`/`mask` semantics when they encode a condition. Use broad
`fillna` only when the condition is exactly null-only.
- Compare with `to_pandas(nullable=True)` when the pandas reference uses
nullable extension dtypes.
- Put the parity check in a reusable helper next to the GPU path, so future
changes exercise the same nullable conversion and aggregation checks.
- Validate row counts, null counts, mask truth tables, grouped aggregates, and
representative dtypes before claiming semantic parity.
## Reference Files
- `references/cudf-pandas-accelerator.md` — Profiling, fallback detection, cudf.pandas deep dive
- `references/api-patterns.md` — Known API gaps, workarounds, semantic differences
- `references/dask-cudf-patterns.md` — Multi-GPU patterns, best practices, partition tuning
## External Documentation
Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.
- **cuDF Documentation:** https://docs.rapids.ai/api/cudf/stable/
- **dask-cuDF API Reference:** https://docs.rapids.ai/api/dask-cudf/stable/api/
- **GitHub:** https://github.com/rapidsai/cudf
- **CHANGELOG:** https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md





首页
