Option
HeimHeim Skill Datenbankverwaltung accelerated-computing-cudf

accelerated-computing-cudf

NVIDIA/skills NVIDIA/skills

Beschleunigen Sie pandas-Workflows mit GPU-DataFrames unter Verwendung von cuDF und dask-cuDF für ETL, Joins, Groupby und die Verarbeitung großer Datenmengen.

...Alle erweitern
2
Zeit aktualisiert 27. September 2026

cuDF & dask-cuDF Implementer's Guide

Kompatibilität

  • Von dieser Funktion verfolgte Version: 26.04.
  • Erfordert NVIDIA Volta oder neuer bei CUDA 12, oder Turing oder neuer bei CUDA 13. Version 26.04 unterstützt CUDA 12.2-12.9 mit Treiber 535+ oder CUDA 13.0-13.1 mit Treiber 580+ sowie Python 3.11-3.14. Der optimale Bereich für cuDF liegt bei >100.000 Zeilen.

Namensgebung

Verwenden Sie in benutzerorientierten Antworten die von NVIDIA bevorzugte Bibliotheksbenennung. Behalten Sie wörtliche RAPIDS/rapidsai-URLs, Paketnamen und Versionsmetadaten bei der Quellenangabe bei.

Rolle

Sie sind ein cuDF-Experte, der einen Implementierer bei der Arbeit mit GPU-Datenrahmen unterstützt. Der Benutzer versteht pandas und seine Daten – Ihre Aufgabe ist es, ihn mit minimalem Aufwand zu korrektem, schnellem GPU-Code zu führen. Wählen Sie den Pfad basierend auf der Absicht des Benutzers: cudf.pandas für breite Kompatibilität oder Beschleunigung bei minimalen Änderungen, explizites cuDF für benannte Datenrahmenmigrationen, kritische ETL-Pfade und arbeitslasten mit Paritätsanforderungen. Behandeln Sie Quellschema, Zeilenanzahl, Nullplatzierung, Reihenfolge und numerische Toleranzen als für den Benutzer sichtbares Verhalten.

Kritische Regeln

  1. Wählen Sie den richtigen cuDF-Pfad. Verwenden Sie cudf.pandas für breite Kompatibilität oder Beschleunigung bei minimalen Änderungen. Verwenden Sie explizites cuDF, wenn der Benutzer den Datenrahmencode migrieren, die Parität prüfen, einen sichtbaren ETL-Hotpath optimieren oder nicht unterstützte Operationen steuern möchte.
  2. Größenbeschränkung: Mindestens 100.000 Zeilen. Unterhalb dieses Werts überwiegt normalerweise der GPU-Übertragungs overhead den Geschwindigkeitsvorteil; verwenden Sie kleine Daten für die Korrektheit und testen Sie größere Arbeitsmengen für die Leistung.
  3. Konvertierungen an den Grenzen halten. Verwenden Sie .to_pandas(), .values oder .numpy() für Anzeige, Plotting, CPU-nur-Bibliotheken oder endgültige Ausgabegrenzen. Halten Sie die zwischengeschalteten ETL-Daten auf der GPU.
  4. Float32 ist Ihr Freund. cuDF-Operationen mit float64 sind langsamer; casten Sie frühzeitig, wenn die Genauigkeit es zulässt.
  5. Semantik an repräsentativen Ausschnitten validieren. Für Nullbehandlung, Joins, Zeitreihen, Umformung oder gruppierte Logik halten Sie einen kleinen pandas-Referenzpfad bereit und vergleichen Sie Form, Beschriftungen, Nullanzahlen, Reihenfolge und repräsentative Werte, bevor Sie Parität behaupten.
  6. Für Daten > GPU-Speicher wechseln Sie zu dask-cuDF mit enable_cudf_spill=True. Siehe references/dask-cudf-patterns.md.

Drei Wege zu GPU-Datenrahmen

Pfad 1: cudf.pandas-Beschleuniger (Kompatibilität / Minimale Änderung)

Verwenden Sie diesen Pfad, wenn der Benutzer eine kleine Codeänderung, die Kompatibilität mit Drittanbieter-pandas oder einen einzigen Codepfad benötigt, der weiterläuft, während nicht unterstützte Operationen zurückfallen.

Jupyter/IPython:

%load_ext cudf.pandas
import pandas as pd   # jetzt GPU-gestützt; fällt bei nicht unterstützten Operationen stillschweigend zurück

Skript:

python -m cudf.pandas my_script.py

Mit Multiprocessing:

import cudf.pandas
cudf.pandas.install()   # muss VOR dem pandas-Import und vor der Pool-Erstellung erfolgen
from multiprocessing import Pool

Bestätigen Sie die Beschleunigung mit dem cudf.pandas-Profiler, bevor Sie einen Geschwindigkeitsvorteil behaupten. Für Notebook-, CLI- und Statistikbeispiele lesen Sie references/cudf-pandas-accelerator.md. Wenn der Profilierungsbericht zeigt, dass der Hotpath auf der CPU läuft, verwenden Sie Pfad 2 für explizite cuDF-Steuerung.

Pfad 2: Explizite cuDF-API

Für volle Kontrolle, Hotpath-Optimierung, benannte Datenrahmenmigrationen und paritätssensitive Operationen:

import cudf

# Daten direkt auf die GPU lesen
df = cudf.read_parquet("data.parquet")

# Operationen spiegeln pandas wider
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]

# String-Operationen
df["clean"] = df["name"].str.strip().str.lower()

# Um die API-Abdeckung vor der Migration zu prüfen:
# Siehe references/api-patterns.md für bekannte Lücken und Workarounds

Halten Sie die Daten durchgängig auf der GPU. Rufen Sie .to_pandas() erst ganz am Ende für die Anzeige oder CPU- oder Nicht-GPU-Übertragung auf.

Bevorzugen Sie explizites cuDF für Aufgaben, die read_csv/read_parquet, Joins, groupby, Umformung, nullable Typen, fillna/where, Zeitbuckets, Rolling-Fenster oder CPU/GPU-Paritätsprüfungen umfassen. Fügen Sie einen kleinen CPU/GPU-Validierungspfad hinzu, wenn die Semantik wichtig ist, anstatt sich ausschließlich auf eine erfolgreiche Ausführung zu verlassen.

Für pandas-Code mit Nullbehandlung, Umformung oder Zeitreihenverhalten lesen Sie references/api-patterns.md für die relevante semantische Checkliste, bevor Sie neu schreiben. Ein cudf.pandas-Bootstrap reicht für eine Anfrage mit minimalen Änderungen aus; eine Implementierungsanfrage sollte den Hotpath explizit und beobachtbar machen.

Für umformungsschweren pandas-Code (pivot_table, melt, stack/unstack, crosstab) halten Sie das Quellschema als Teil des Vertrags: Indexbeschriftungen, Spaltenbeschriftungen oder -ebenen, fill_value, aggfunc, Marginalen und Normalisierung. Verwenden Sie explizites cuDF, wo das Äquivalent unterstützt wird; verwenden Sie cudf.pandas oder eine schmale Kompatibilitätsgrenze, wenn exakte pandas-Umformungssemantiken wichtiger sind als das Neuschreiben jeder Operation. Fügen Sie eine kleine pandas-Referenz-Paritätsprüfung für Form, Beschriftungen und repräsentative Werte hinzu, bevor Sie abschließen. Siehe references/api-patterns.md.

Pfad 3: dask-cuDF (Multi-GPU / Große Daten)

Wenn der Datensatz den GPU-Speicher überschreitet. Siehe references/dask-cudf-patterns.md für vollständige Muster.

from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf

cluster = LocalCUDACluster(enable_cudf_spill=True)  # ein Worker pro GPU
client = Client(cluster)

ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()

Speicherverwaltung

Spill aktivieren, bevor OOM auftritt (nicht danach):

import cudf
cudf.set_option("spill", True)   # Spill auf Host-RAM, wenn GPU voll ist

RMM-Poolzuordner (reduziert cudaMalloc-Overhead in Pipelines mit vielen Zuordnungen):

import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Muss VOR allen cuDF-Operationen aufgerufen werden
GPU-Freier Speicher vs. DatensatzStrategie
Frei > 2× DatensatzEinzeln-GPU-cuDF
Frei 1–2× DatensatzcuDF + `cudf.set_option("spill", True)`
Datensatz > GPU-Speicherdask-cuDF
Datensatz > Knoten-Speicherdask-cuDF + Multi-Knoten (siehe accelerated-computing-mpf)

Fehlerbehebung

Kein Geschwindigkeitsvorteil gegenüber pandas:

  • Daten
  • Führen Sie %%cudf.pandas.profile aus – hoher CPU-Anteil bedeutet viele Rückfälle. Identifizieren und beheben Sie diese Operationen.
  • Prüfen Sie references/api-patterns.md auf bekannte Lücken.

OOM (CUDA out of memory):

  1. Spill aktivieren: cudf.set_option("spill", True)
  2. Wenn Allocator-Fragmentierung oder wiederholter Zuordnungs-Overhead sichtbar ist, verwenden Sie die accelerated-computing-rmm-Speicherressourcen-Einrichtungsempfehlung vor GPU-Zuordnungen.
  3. Immer noch fehlschlagend: Wechseln Sie zu dask-cuDF.

AttributeError / NotImplementedError:

  • Prüfen Sie references/api-patterns.md für die spezifische Operation.
  • Halten Sie diese eine Operation an einer schmalen Grenze auf der CPU und führen Sie den unterstützten Pipeline auf der GPU fort.
  • Verwenden Sie .to_pandas() nur für die nicht unterstützte Operation und dann .from_pandas() zurück.

Falsche Ergebnisse gegenüber pandas:

  • Null-/NaN-Behandlung unterscheidet sich: cuDF verwendet standardmäßig <na></na> (nullable), pandas verwendet NaN. Siehe references/api-patterns.md.
  • Sortierstabilität: Die cuDF-Sortierung ist nicht stabil garantiert, es sei denn, stable=True wird übergeben.
  • Wenn der Unterschied auf Gleitkommaunterschiede zurückzuführen ist, versuchen Sie, auf höhere Genauigkeits-Gleitkommazahlen zu casten (z. B. float64 statt float32). Wenn die Ergebnisse immer noch unterschiedlich sind, stoppen Sie. GPU- und CPU-Algorithmen werden aufgrund der Nicht-Assoziativität der Gleitkommaarithmetik immer unterschiedliche Ergebnisse bei Gleitkommazahlen produzieren, und das kann nicht behoben werden.

Semantik für Nullable und Füllen

Wenn der Benutzer explizit pandas-nullable-Datentypen, fillna, where/mask oder gruppiertes Nullverhalten betrifft, behandeln Sie Paritätsprüfungen als Teil der Implementierung. Siehe references/api-patterns.md für Beispiele zu nullable-Datentypen.

  • Bewahren Sie nullable Integer-/String-Spalten bei, anstatt sie mit Sentinelwerten zu füllen, es sei denn, der Quelrcode hat dies bereits getan.
  • Behalten Sie die where/mask-Semantik bei, wenn sie eine Bedingung kodieren. Verwenden Sie breites fillna nur, wenn die Bedingung genau null-nur ist.
  • Vergleichen Sie mit to_pandas(nullable=True), wenn der pandas-Referenz nullable Extension-Datentypen verwendet.
  • Platzieren Sie die Paritätsprüfung in einem wiederverwendbaren Helfer neben dem GPU-Pfad, damit zukünftige Änderungen dieselben nullable-Konvertierungs- und Aggregationsprüfungen ausüben.
  • Validieren Sie Zeilenanzahlen, Nullanzahlen, Masken-Wahrheitstabellen, gruppierte Aggregationen und repräsentative Datentypen, bevor Sie semantische Parität behaupten.

Referenzdateien

  • references/cudf-pandas-accelerator.md – Profiling, Rückfallerkennung, tiefgehende Analyse von cudf.pandas
  • references/api-patterns.md – Bekannte API-Lücken, Workarounds, semantische Unterschiede
  • references/dask-cudf-patterns.md – Multi-GPU-Muster, Best Practices, Partitionstuning

Externe Dokumentation

Verwenden Sie WebFetch, um auf Anfrage detaillierte API-Signaturen, Parameterbeschreibungen und Beispiele abzurufen.

Auf GitHub ansehen
---
name: accelerated-computing-cudf
description: Accelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
license: CC-BY-4.0 AND Apache-2.0
---

# cuDF & dask-cuDF Implementer's Guide

## Compatibility

- Release tracked by this skill: 26.04.
- Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.

## Naming

Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.

## Role

You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: `cudf.pandas` for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.

## Critical Rules

1. **Choose the right cuDF path.** Use `cudf.pandas` for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
2. **Size gate: 100K rows minimum.** Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
3. **Keep conversions at boundaries.** Use `.to_pandas()`, `.values`, or `.numpy()` for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
4. **Float32 is your friend.** cuDF operations on float64 are slower; cast early when precision allows.
5. **Validate semantics on representative slices.** For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
6. **For data > GPU memory**, move to dask-cuDF with `enable_cudf_spill=True`. See `references/dask-cudf-patterns.md`.

## Three Paths to GPU DataFrames

### Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)

Use when the user needs a small code change, third-party pandas compatibility,
or one code path that can keep running while unsupported operations fall back.

**Jupyter/IPython:**
```python
%load_ext cudf.pandas
import pandas as pd   # now GPU-backed; falls back silently for unsupported ops
```

**Script:**
```bash
python -m cudf.pandas my_script.py
```

**With multiprocessing:**
```python
import cudf.pandas
cudf.pandas.install()   # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool
```

Confirm acceleration with the cudf.pandas profiler before claiming speedup.
For notebook, CLI, and stats examples, read
`references/cudf-pandas-accelerator.md`. If the profile shows the hot path
running on CPU, use Path 2 for explicit cuDF control.

### Path 2: Explicit cuDF API

For full control, hot-path optimization, named DataFrame migrations, and
parity-sensitive operations:

```python
import cudf

# Read data directly to GPU
df = cudf.read_parquet("data.parquet")

# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]

# String operations
df["clean"] = df["name"].str.strip().str.lower()

# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds
```

**Keep data on GPU end-to-end.** Only call `.to_pandas()` at the very end for display or CPU or non-GPU handoff.

Prefer explicit cuDF for tasks involving `read_csv`/`read_parquet`, joins,
groupby, reshape, nullable types, `fillna`/`where`, time buckets, rolling
windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when
semantics matter instead of relying on successful execution alone.

For pandas code with null handling, reshape, or time-series behavior, read
`references/api-patterns.md` for the relevant semantic checklist before
rewriting. A `cudf.pandas` bootstrap is enough for a minimal-change request; an
implementation request should make the hot path explicit and observable.

For reshape-heavy pandas code (`pivot_table`, `melt`, `stack`/`unstack`,
`crosstab`), keep the source schema as part of the contract: index labels,
column labels or levels, `fill_value`, `aggfunc`, margins, and normalization.
Use explicit cuDF where the equivalent is supported; use `cudf.pandas` or a
narrow compatibility boundary when exact pandas reshape semantics matter more
than rewriting every operation. Add a small pandas-reference parity check for
shape, labels, and representative values before finalizing. See
`references/api-patterns.md`.

### Path 3: dask-cuDF (Multi-GPU / Large Data)

When dataset exceeds GPU memory. See `references/dask-cudf-patterns.md` for full patterns.

```python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf

cluster = LocalCUDACluster(enable_cudf_spill=True)  # one worker per GPU
client = Client(cluster)

ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
```

## Memory Management

**Enable spill before OOM happens** (not after):
```python
import cudf
cudf.set_option("spill", True)   # spill to host RAM when GPU is full
```

**RMM pool allocator** (reduces cudaMalloc overhead in pipelines with many allocations):
```python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
```

| GPU Free vs Dataset | Strategy |
|---|---|
| Free > 2× dataset | Single GPU cuDF |
| Free 1–2× dataset | cuDF + `cudf.set_option("spill", True)` |
| Dataset > GPU mem | dask-cuDF |
| Dataset > node mem | dask-cuDF + multi-node (see accelerated-computing-mpf) |

## Troubleshooting

**No speedup vs pandas:**
- Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
- Run `%%cudf.pandas.profile` — high CPU % means many fallbacks. Identify and fix those ops.
- Check `references/api-patterns.md` for known gaps.

**OOM (CUDA out of memory):**
1. Enable spill: `cudf.set_option("spill", True)`
2. If allocator fragmentation or repeated allocation overhead is visible, use the `accelerated-computing-rmm` memory-resource setup guidance before GPU allocations
3. Still failing: move to dask-cuDF

**AttributeError / NotImplementedError:**
- Check `references/api-patterns.md` for the specific operation
- Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
- Use `.to_pandas()` only for the unsupported op, then `.from_pandas()` back

**Wrong results vs pandas:**
- Null/NaN handling differs: cuDF uses `<NA>` (nullable) by default, pandas uses `NaN`. See `references/api-patterns.md`.
- Sort stability: cuDF sort is not guaranteed stable unless `stable=True` is passed
- If the difference is due to floating point differences, try casting to higher precision floats (e.g. `float64` instead of `float32`). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.

## Nullable and Fill Semantics

When the user explicitly cares about pandas nullable dtypes, `fillna`,
`where`/`mask`, or grouped null behavior, treat parity checks as part of the
implementation. See `references/api-patterns.md` for nullable dtype examples.

- Preserve nullable integer/string columns instead of filling them with sentinel
  values unless the source code already did that.
- Keep `where`/`mask` semantics when they encode a condition. Use broad
  `fillna` only when the condition is exactly null-only.
- Compare with `to_pandas(nullable=True)` when the pandas reference uses
  nullable extension dtypes.
- Put the parity check in a reusable helper next to the GPU path, so future
  changes exercise the same nullable conversion and aggregation checks.
- Validate row counts, null counts, mask truth tables, grouped aggregates, and
  representative dtypes before claiming semantic parity.

## Reference Files

- `references/cudf-pandas-accelerator.md` — Profiling, fallback detection, cudf.pandas deep dive
- `references/api-patterns.md` — Known API gaps, workarounds, semantic differences
- `references/dask-cudf-patterns.md` — Multi-GPU patterns, best practices, partition tuning

## External Documentation

Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.

- **cuDF Documentation:** https://docs.rapids.ai/api/cudf/stable/
- **dask-cuDF API Reference:** https://docs.rapids.ai/api/dask-cudf/stable/api/
- **GitHub:** https://github.com/rapidsai/cudf
- **CHANGELOG:** https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md

Alle Dateien

34 Dateien

accelerated-computing-cudf installieren

Laden Sie die Skill-Dateien herunter und extrahieren Sie diese in Ihr .claude/skills/-Verzeichnis.

ZIP herunterladen

Klonen Sie das Repository und kopieren Sie die Skill-Dateien in Ihr Projekt.

git clone https://github.com/NVIDIA/skills/tree/main/skills/accelerated-computing-cudf # Copy SKILL.md to your .claude/skills/ directory

Kopieren Kopieren
Schnelle Einrichtung: Kopieren Sie den Ordner „skill“ nach .claude/skills/. Claude erkennt und verwendet die Fähigkeit automatisch.
Repository NVIDIA/skills

Ähnliche Skills

microservices-patterns
Zeit aktualisiert 29. Juni 2026
jpa-patterns
Zeit aktualisiert 30. Juni 2026
fabric-lakehouse
Zeit aktualisiert 30. Juni 2026
prisma-expert
Zeit aktualisiert 29. Juni 2026
OR