选项
首页首页 Skill 数据库管理 accelerated-computing-cudf

accelerated-computing-cudf

NVIDIA/skills NVIDIA/skills

使用 cuDF 和 dask-cuDF 加速基于 GPU 数据帧的 pandas 工作流,适用于 ETL、连接、groupby 以及大规模数据处理。

...展开全部
2
更新时间 2026-09-27

cuDF 与 dask-cuDF 实现者指南

兼容性

  • 本技能跟踪的版本:26.04。
  • 在 CUDA 12 上需要 NVIDIA Volta 或更新架构,或在 CUDA 13 上需要 Turing 或更新架构。26.04 版本支持 CUDA 12.2-12.9(驱动版本 535+)或 CUDA 13.0-13.1(驱动版本 580+),以及 Python 3.11-3.14。cuDF 的最佳适用场景:数据行数大于 10 万行。

命名规范

在面向用户的回答中,优先使用 NVIDIA 库的术语。在引用来源时,保留字面的 RAPIDS/rapidsai URL、包名和版本元数据。

角色定义

你是 cuDF 专家,协助实现者处理 GPU DataFrame。用户已理解 pandas 及其数据——你的任务是以最小的摩擦帮助他们编写正确、高效的 GPU 代码。根据用户意图选择路径:若需广泛兼容性或最小变更加速,使用 cudf.pandas;若需显式迁移 DataFrame 代码、优化 ETL 关键路径或处理对一致性敏感的工作,使用显式 cuDF。将源模式、行数、空值位置、排序顺序和数值容差视为用户可见的行为。

关键规则

  1. 选择合适的 cuDF 路径。 若需广泛兼容性或最小变更加速,使用 cudf.pandas。若用户要求迁移 DataFrame 代码、检查一致性、优化可见的 ETL 关键路径或控制不支持的操作,则使用显式 cuDF。
  2. 大小限制:最低 10 万行。 低于此行数时,GPU 数据传输开销通常会抵消加速收益;使用小数据进行正确性验证,并对更大的工作集进行性能基准测试。
  3. 转换保持在边界处。 对于显示、绘图、仅 CPU 库或最终输出边界,使用 .to_pandas()、.values 或 .numpy()。中间 ETL 数据应保留在 GPU 上。
  4. Float32 是你的朋友。 cuDF 对 float64 的操作较慢;当精度允许时,请尽早转换类型。
  5. 在代表性切片上验证语义。 对于空值处理、连接、时间序列、重塑或分组逻辑,保留一个小型 pandas 参考路径,并在声称一致性之前比较形状、标签、空值计数、排序顺序和代表性值。
  6. 对于超过 GPU 内存的数据,请迁移到 dask-cuDF 并设置 enable_cudf_spill=True。参见 references/dask-cudf-patterns.md。

通往 GPU DataFrame 的三条路径

路径 1:cudf.pandas 加速器(兼容性 / 最小变更)

当用户需要少量代码更改、第三方 pandas 兼容性,或需要一个在支持的操作回退时仍能运行的单一代码路径时使用。

Jupyter/IPython:

%load_ext cudf.pandas
import pandas as pd   # 现在由 GPU 支持;不支持的操作会静默回退

脚本:

python -m cudf.pandas my_script.py

使用多进程:

import cudf.pandas
cudf.pandas.install()   # 必须在导入 pandas 之前、创建 Pool 之前调用
from multiprocessing import Pool

在声称加速之前,请使用 cudf.pandas 分析器进行确认。对于笔记本、CLI 和统计示例,请阅读 references/cudf-pandas-accelerator.md。如果分析显示关键路径在 CPU 上运行,请使用路径 2 进行显式 cuDF 控制。

路径 2:显式 cuDF API

如需完全控制、关键路径优化、命名 DataFrame 迁移以及对一致性敏感的操作:

import cudf

# 直接将数据读取到 GPU
df = cudf.read_parquet("data.parquet")

# 操作与 pandas 类似
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]

# 字符串操作
df["clean"] = df["name"].str.strip().str.lower()

# 在提交迁移之前检查 API 覆盖范围:
# 参见 references/api-patterns.md 以获取已知缺陷和变通方法

使数据端到端保留在 GPU 上。 仅在最后调用 .to_pandas() 用于显示、CPU 处理或非 GPU 交接。

对于涉及 read_csv/read_parquet、连接、groupby、重塑、可空类型、fillna/where、时间桶、滚动窗口或 CPU/GPU 一致性检查的任务,优先使用显式 cuDF。当语义重要时,添加小型 CPU/GPU 验证路径,而不是仅依赖执行成功。

对于涉及空值处理、重塑或时间序列行为的 pandas 代码,在重写之前请阅读 references/api-patterns.md 以获取相关语义检查清单。对于最小变更请求,cudf.pandas 引导程序已足够;对于实现请求,应使关键路径显式且可观察。

对于重塑密集的 pandas 代码(pivot_table、melt、stack/unstack、crosstab),将源模式作为合同的一部分:索引标签、列标签或层级、fill_value、aggfunc、边距和归一化。在等效操作受支持的地方使用显式 cuDF;当精确的 pandas 重塑语义比重写每个操作更重要时,使用 cudf.pandas 或狭窄的兼容性边界。在最终确定之前,添加小型 pandas 参考一致性检查以验证形状、标签和代表性值。参见 references/api-patterns.md。

路径 3:dask-cuDF(多 GPU / 大数据)

当数据集超过 GPU 内存时。参见 references/dask-cudf-patterns.md 以获取完整模式。

from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf

cluster = LocalCUDACluster(enable_cudf_spill=True)  # 每个 GPU 一个工作进程
client = Client(cluster)

ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()

内存管理

在发生 OOM(内存溢出)之前启用溢出(而不是之后):

import cudf
cudf.set_option("spill", True)   # 当 GPU 满时溢出到主机 RAM

RMM 池分配器(减少具有大量分配的操作流程中的 cudaMalloc 开销):

import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# 必须在任何 cuDF 操作之前调用
GPU 空闲空间与数据集大小对比策略
空闲空间 > 数据集大小的 2 倍单 GPU cuDF
空闲空间为数据集大小的 1-2 倍cuDF + `cudf.set_option("spill", True)`
数据集大小 > GPU 内存dask-cuDF
数据集大小 > 节点内存dask-cuDF + 多节点(参见 accelerated-computing-mpf)

故障排除

与 pandas 相比无加速:

  • 数据
  • 运行 %%cudf.pandas.profile — 高 CPU % 意味着存在大量回退。识别并修复这些操作。
  • 检查 references/api-patterns.md 以获取已知缺陷。

OOM(CUDA 内存溢出):

  1. 启用溢出:cudf.set_option("spill", True)
  2. 如果可见分配器碎片或重复分配开销,请在 GPU 分配之前使用 accelerated-computing-rmm 内存资源设置指南
  3. 仍然失败:迁移到 dask-cuDF

AttributeError / NotImplementedError:

  • 检查 references/api-patterns.md 以获取特定操作
  • 在狭窄边界上将一个操作保留在 CPU 上,并在 GPU 上继续支持的操作流程
  • 仅对不支持的操作使用 .to_pandas(),然后使用 .from_pandas() 转换回来

与 pandas 相比结果错误:

  • 空值/NaN 处理不同:cuDF 默认使用 <na></na>(可空),pandas 使用 NaN。参见 references/api-patterns.md。
  • 排序稳定性:除非传递 stable=True,否则 cuDF 排序不保证稳定
  • 如果差异是由于浮点数差异引起的,请尝试转换为更高精度的浮点数(例如 float64 而不是 float32)。如果结果仍然不同,请停止。由于浮点运算的非结合性,GPU 和 CPU 算法在浮点数上始终会产生不同的结果,这是无法修复的。

可空值与填充语义

当用户明确关注 pandas 可空数据类型、fillna、where/mask 或分组空值行为时,将一致性检查视为实现的一部分。参见 references/api-patterns.md 以获取可空数据类型示例。

  • 保留可空整数/字符串列,除非源代码已经用哨兵值填充它们。
  • 当 where/mask 编码条件时,请保留其语义。仅当条件恰好为空时,才使用广泛的 fillna。
  • 当 pandas 参考使用可空扩展数据类型时,请使用 to_pandas(nullable=True) 进行比较。
  • 在 GPU 路径旁边放置一个可重用的辅助函数以进行一致性检查,以便未来的更改能够执行相同的可空转换和聚合检查。
  • 在声称语义一致性之前,验证行数、空值计数、掩码真值表、分组聚合和代表性数据类型。

参考文件

  • references/cudf-pandas-accelerator.md — 分析、回退检测、cudf.pandas 深入解析
  • references/api-patterns.md — 已知 API 缺陷、变通方法、语义差异
  • references/dask-cudf-patterns.md — 多 GPU 模式、最佳实践、分区调优

外部文档

使用 WebFetch 按需检索详细的 API 签名、参数描述和示例。

在 GitHub 上查看
---
name: accelerated-computing-cudf
description: Accelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
license: CC-BY-4.0 AND Apache-2.0
---

# cuDF & dask-cuDF Implementer's Guide

## Compatibility

- Release tracked by this skill: 26.04.
- Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.

## Naming

Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.

## Role

You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: `cudf.pandas` for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.

## Critical Rules

1. **Choose the right cuDF path.** Use `cudf.pandas` for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
2. **Size gate: 100K rows minimum.** Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
3. **Keep conversions at boundaries.** Use `.to_pandas()`, `.values`, or `.numpy()` for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
4. **Float32 is your friend.** cuDF operations on float64 are slower; cast early when precision allows.
5. **Validate semantics on representative slices.** For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
6. **For data > GPU memory**, move to dask-cuDF with `enable_cudf_spill=True`. See `references/dask-cudf-patterns.md`.

## Three Paths to GPU DataFrames

### Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)

Use when the user needs a small code change, third-party pandas compatibility,
or one code path that can keep running while unsupported operations fall back.

**Jupyter/IPython:**
```python
%load_ext cudf.pandas
import pandas as pd   # now GPU-backed; falls back silently for unsupported ops
```

**Script:**
```bash
python -m cudf.pandas my_script.py
```

**With multiprocessing:**
```python
import cudf.pandas
cudf.pandas.install()   # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool
```

Confirm acceleration with the cudf.pandas profiler before claiming speedup.
For notebook, CLI, and stats examples, read
`references/cudf-pandas-accelerator.md`. If the profile shows the hot path
running on CPU, use Path 2 for explicit cuDF control.

### Path 2: Explicit cuDF API

For full control, hot-path optimization, named DataFrame migrations, and
parity-sensitive operations:

```python
import cudf

# Read data directly to GPU
df = cudf.read_parquet("data.parquet")

# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]

# String operations
df["clean"] = df["name"].str.strip().str.lower()

# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds
```

**Keep data on GPU end-to-end.** Only call `.to_pandas()` at the very end for display or CPU or non-GPU handoff.

Prefer explicit cuDF for tasks involving `read_csv`/`read_parquet`, joins,
groupby, reshape, nullable types, `fillna`/`where`, time buckets, rolling
windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when
semantics matter instead of relying on successful execution alone.

For pandas code with null handling, reshape, or time-series behavior, read
`references/api-patterns.md` for the relevant semantic checklist before
rewriting. A `cudf.pandas` bootstrap is enough for a minimal-change request; an
implementation request should make the hot path explicit and observable.

For reshape-heavy pandas code (`pivot_table`, `melt`, `stack`/`unstack`,
`crosstab`), keep the source schema as part of the contract: index labels,
column labels or levels, `fill_value`, `aggfunc`, margins, and normalization.
Use explicit cuDF where the equivalent is supported; use `cudf.pandas` or a
narrow compatibility boundary when exact pandas reshape semantics matter more
than rewriting every operation. Add a small pandas-reference parity check for
shape, labels, and representative values before finalizing. See
`references/api-patterns.md`.

### Path 3: dask-cuDF (Multi-GPU / Large Data)

When dataset exceeds GPU memory. See `references/dask-cudf-patterns.md` for full patterns.

```python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf

cluster = LocalCUDACluster(enable_cudf_spill=True)  # one worker per GPU
client = Client(cluster)

ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
```

## Memory Management

**Enable spill before OOM happens** (not after):
```python
import cudf
cudf.set_option("spill", True)   # spill to host RAM when GPU is full
```

**RMM pool allocator** (reduces cudaMalloc overhead in pipelines with many allocations):
```python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
```

| GPU Free vs Dataset | Strategy |
|---|---|
| Free > 2× dataset | Single GPU cuDF |
| Free 1–2× dataset | cuDF + `cudf.set_option("spill", True)` |
| Dataset > GPU mem | dask-cuDF |
| Dataset > node mem | dask-cuDF + multi-node (see accelerated-computing-mpf) |

## Troubleshooting

**No speedup vs pandas:**
- Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
- Run `%%cudf.pandas.profile` — high CPU % means many fallbacks. Identify and fix those ops.
- Check `references/api-patterns.md` for known gaps.

**OOM (CUDA out of memory):**
1. Enable spill: `cudf.set_option("spill", True)`
2. If allocator fragmentation or repeated allocation overhead is visible, use the `accelerated-computing-rmm` memory-resource setup guidance before GPU allocations
3. Still failing: move to dask-cuDF

**AttributeError / NotImplementedError:**
- Check `references/api-patterns.md` for the specific operation
- Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
- Use `.to_pandas()` only for the unsupported op, then `.from_pandas()` back

**Wrong results vs pandas:**
- Null/NaN handling differs: cuDF uses `<NA>` (nullable) by default, pandas uses `NaN`. See `references/api-patterns.md`.
- Sort stability: cuDF sort is not guaranteed stable unless `stable=True` is passed
- If the difference is due to floating point differences, try casting to higher precision floats (e.g. `float64` instead of `float32`). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.

## Nullable and Fill Semantics

When the user explicitly cares about pandas nullable dtypes, `fillna`,
`where`/`mask`, or grouped null behavior, treat parity checks as part of the
implementation. See `references/api-patterns.md` for nullable dtype examples.

- Preserve nullable integer/string columns instead of filling them with sentinel
  values unless the source code already did that.
- Keep `where`/`mask` semantics when they encode a condition. Use broad
  `fillna` only when the condition is exactly null-only.
- Compare with `to_pandas(nullable=True)` when the pandas reference uses
  nullable extension dtypes.
- Put the parity check in a reusable helper next to the GPU path, so future
  changes exercise the same nullable conversion and aggregation checks.
- Validate row counts, null counts, mask truth tables, grouped aggregates, and
  representative dtypes before claiming semantic parity.

## Reference Files

- `references/cudf-pandas-accelerator.md` — Profiling, fallback detection, cudf.pandas deep dive
- `references/api-patterns.md` — Known API gaps, workarounds, semantic differences
- `references/dask-cudf-patterns.md` — Multi-GPU patterns, best practices, partition tuning

## External Documentation

Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.

- **cuDF Documentation:** https://docs.rapids.ai/api/cudf/stable/
- **dask-cuDF API Reference:** https://docs.rapids.ai/api/dask-cudf/stable/api/
- **GitHub:** https://github.com/rapidsai/cudf
- **CHANGELOG:** https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md

所有文件

34 个文件

安装 accelerated-computing-cudf

将技能文件下载并解压至你的 .claude/skills/ 目录。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/accelerated-computing-cudf # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ 目录。Claude 将自动检测并使用该技能。
仓库 NVIDIA/skills

相关技能

microservices-patterns
更新时间 2026-06-29
jpa-patterns
更新时间 2026-06-30
fabric-lakehouse
更新时间 2026-06-30
prisma-expert
更新时间 2026-06-29
OR