accelerated-computing-cudf
NVIDIA/skills
Beschleunigen Sie pandas-Workflows mit GPU-DataFrames unter Verwendung von cuDF und dask-cuDF für ETL, Joins, Groupby und die Verarbeitung großer Datenmengen.
...Alle erweiterncuDF & dask-cuDF Implementer's Guide
Kompatibilität
- Von dieser Funktion verfolgte Version: 26.04.
- Erfordert NVIDIA Volta oder neuer bei CUDA 12, oder Turing oder neuer bei CUDA 13. Version 26.04 unterstützt CUDA 12.2-12.9 mit Treiber 535+ oder CUDA 13.0-13.1 mit Treiber 580+ sowie Python 3.11-3.14. Der optimale Bereich für cuDF liegt bei >100.000 Zeilen.
Namensgebung
Verwenden Sie in benutzerorientierten Antworten die von NVIDIA bevorzugte Bibliotheksbenennung. Behalten Sie wörtliche RAPIDS/rapidsai-URLs, Paketnamen und Versionsmetadaten bei der Quellenangabe bei.
Rolle
Sie sind ein cuDF-Experte, der einen Implementierer bei der Arbeit mit GPU-Datenrahmen unterstützt. Der Benutzer versteht pandas und seine Daten – Ihre Aufgabe ist es, ihn mit minimalem Aufwand zu korrektem, schnellem GPU-Code zu führen. Wählen Sie den Pfad basierend auf der Absicht des Benutzers: cudf.pandas für breite Kompatibilität oder Beschleunigung bei minimalen Änderungen, explizites cuDF für benannte Datenrahmenmigrationen, kritische ETL-Pfade und arbeitslasten mit Paritätsanforderungen. Behandeln Sie Quellschema, Zeilenanzahl, Nullplatzierung, Reihenfolge und numerische Toleranzen als für den Benutzer sichtbares Verhalten.
Kritische Regeln
- Wählen Sie den richtigen cuDF-Pfad. Verwenden Sie
cudf.pandasfür breite Kompatibilität oder Beschleunigung bei minimalen Änderungen. Verwenden Sie explizites cuDF, wenn der Benutzer den Datenrahmencode migrieren, die Parität prüfen, einen sichtbaren ETL-Hotpath optimieren oder nicht unterstützte Operationen steuern möchte. - Größenbeschränkung: Mindestens 100.000 Zeilen. Unterhalb dieses Werts überwiegt normalerweise der GPU-Übertragungs overhead den Geschwindigkeitsvorteil; verwenden Sie kleine Daten für die Korrektheit und testen Sie größere Arbeitsmengen für die Leistung.
- Konvertierungen an den Grenzen halten. Verwenden Sie
.to_pandas(),.valuesoder.numpy()für Anzeige, Plotting, CPU-nur-Bibliotheken oder endgültige Ausgabegrenzen. Halten Sie die zwischengeschalteten ETL-Daten auf der GPU. - Float32 ist Ihr Freund. cuDF-Operationen mit float64 sind langsamer; casten Sie frühzeitig, wenn die Genauigkeit es zulässt.
- Semantik an repräsentativen Ausschnitten validieren. Für Nullbehandlung, Joins, Zeitreihen, Umformung oder gruppierte Logik halten Sie einen kleinen pandas-Referenzpfad bereit und vergleichen Sie Form, Beschriftungen, Nullanzahlen, Reihenfolge und repräsentative Werte, bevor Sie Parität behaupten.
- Für Daten > GPU-Speicher wechseln Sie zu dask-cuDF mit
enable_cudf_spill=True. Siehereferences/dask-cudf-patterns.md.
Drei Wege zu GPU-Datenrahmen
Pfad 1: cudf.pandas-Beschleuniger (Kompatibilität / Minimale Änderung)
Verwenden Sie diesen Pfad, wenn der Benutzer eine kleine Codeänderung, die Kompatibilität mit Drittanbieter-pandas oder einen einzigen Codepfad benötigt, der weiterläuft, während nicht unterstützte Operationen zurückfallen.
Jupyter/IPython:
%load_ext cudf.pandas
import pandas as pd # jetzt GPU-gestützt; fällt bei nicht unterstützten Operationen stillschweigend zurück
Skript:
python -m cudf.pandas my_script.py
Mit Multiprocessing:
import cudf.pandas
cudf.pandas.install() # muss VOR dem pandas-Import und vor der Pool-Erstellung erfolgen
from multiprocessing import Pool
Bestätigen Sie die Beschleunigung mit dem cudf.pandas-Profiler, bevor Sie einen Geschwindigkeitsvorteil behaupten. Für Notebook-, CLI- und Statistikbeispiele lesen Sie references/cudf-pandas-accelerator.md. Wenn der Profilierungsbericht zeigt, dass der Hotpath auf der CPU läuft, verwenden Sie Pfad 2 für explizite cuDF-Steuerung.
Pfad 2: Explizite cuDF-API
Für volle Kontrolle, Hotpath-Optimierung, benannte Datenrahmenmigrationen und paritätssensitive Operationen:
import cudf
# Daten direkt auf die GPU lesen
df = cudf.read_parquet("data.parquet")
# Operationen spiegeln pandas wider
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# String-Operationen
df["clean"] = df["name"].str.strip().str.lower()
# Um die API-Abdeckung vor der Migration zu prüfen:
# Siehe references/api-patterns.md für bekannte Lücken und Workarounds
Halten Sie die Daten durchgängig auf der GPU. Rufen Sie .to_pandas() erst ganz am Ende für die Anzeige oder CPU- oder Nicht-GPU-Übertragung auf.
Bevorzugen Sie explizites cuDF für Aufgaben, die read_csv/read_parquet, Joins, groupby, Umformung, nullable Typen, fillna/where, Zeitbuckets, Rolling-Fenster oder CPU/GPU-Paritätsprüfungen umfassen. Fügen Sie einen kleinen CPU/GPU-Validierungspfad hinzu, wenn die Semantik wichtig ist, anstatt sich ausschließlich auf eine erfolgreiche Ausführung zu verlassen.
Für pandas-Code mit Nullbehandlung, Umformung oder Zeitreihenverhalten lesen Sie references/api-patterns.md für die relevante semantische Checkliste, bevor Sie neu schreiben. Ein cudf.pandas-Bootstrap reicht für eine Anfrage mit minimalen Änderungen aus; eine Implementierungsanfrage sollte den Hotpath explizit und beobachtbar machen.
Für umformungsschweren pandas-Code (pivot_table, melt, stack/unstack, crosstab) halten Sie das Quellschema als Teil des Vertrags: Indexbeschriftungen, Spaltenbeschriftungen oder -ebenen, fill_value, aggfunc, Marginalen und Normalisierung. Verwenden Sie explizites cuDF, wo das Äquivalent unterstützt wird; verwenden Sie cudf.pandas oder eine schmale Kompatibilitätsgrenze, wenn exakte pandas-Umformungssemantiken wichtiger sind als das Neuschreiben jeder Operation. Fügen Sie eine kleine pandas-Referenz-Paritätsprüfung für Form, Beschriftungen und repräsentative Werte hinzu, bevor Sie abschließen. Siehe references/api-patterns.md.
Pfad 3: dask-cuDF (Multi-GPU / Große Daten)
Wenn der Datensatz den GPU-Speicher überschreitet. Siehe references/dask-cudf-patterns.md für vollständige Muster.
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # ein Worker pro GPU
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
Speicherverwaltung
Spill aktivieren, bevor OOM auftritt (nicht danach):
import cudf
cudf.set_option("spill", True) # Spill auf Host-RAM, wenn GPU voll ist
RMM-Poolzuordner (reduziert cudaMalloc-Overhead in Pipelines mit vielen Zuordnungen):
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Muss VOR allen cuDF-Operationen aufgerufen werden
| GPU-Freier Speicher vs. Datensatz | Strategie |
|---|---|
| Frei > 2× Datensatz | Einzeln-GPU-cuDF |
| Frei 1–2× Datensatz | cuDF + `cudf.set_option("spill", True)` |
| Datensatz > GPU-Speicher | dask-cuDF |
| Datensatz > Knoten-Speicher | dask-cuDF + Multi-Knoten (siehe accelerated-computing-mpf) |
Fehlerbehebung
Kein Geschwindigkeitsvorteil gegenüber pandas:
- Daten
- Führen Sie
%%cudf.pandas.profileaus – hoher CPU-Anteil bedeutet viele Rückfälle. Identifizieren und beheben Sie diese Operationen. - Prüfen Sie
references/api-patterns.mdauf bekannte Lücken.
OOM (CUDA out of memory):
- Spill aktivieren:
cudf.set_option("spill", True) - Wenn Allocator-Fragmentierung oder wiederholter Zuordnungs-Overhead sichtbar ist, verwenden Sie die
accelerated-computing-rmm-Speicherressourcen-Einrichtungsempfehlung vor GPU-Zuordnungen. - Immer noch fehlschlagend: Wechseln Sie zu dask-cuDF.
AttributeError / NotImplementedError:
- Prüfen Sie
references/api-patterns.mdfür die spezifische Operation. - Halten Sie diese eine Operation an einer schmalen Grenze auf der CPU und führen Sie den unterstützten Pipeline auf der GPU fort.
- Verwenden Sie
.to_pandas()nur für die nicht unterstützte Operation und dann.from_pandas()zurück.
Falsche Ergebnisse gegenüber pandas:
- Null-/NaN-Behandlung unterscheidet sich: cuDF verwendet standardmäßig
<na></na>(nullable), pandas verwendetNaN. Siehereferences/api-patterns.md. - Sortierstabilität: Die cuDF-Sortierung ist nicht stabil garantiert, es sei denn,
stable=Truewird übergeben. - Wenn der Unterschied auf Gleitkommaunterschiede zurückzuführen ist, versuchen Sie, auf höhere Genauigkeits-Gleitkommazahlen zu casten (z. B.
float64stattfloat32). Wenn die Ergebnisse immer noch unterschiedlich sind, stoppen Sie. GPU- und CPU-Algorithmen werden aufgrund der Nicht-Assoziativität der Gleitkommaarithmetik immer unterschiedliche Ergebnisse bei Gleitkommazahlen produzieren, und das kann nicht behoben werden.
Semantik für Nullable und Füllen
Wenn der Benutzer explizit pandas-nullable-Datentypen, fillna, where/mask oder gruppiertes Nullverhalten betrifft, behandeln Sie Paritätsprüfungen als Teil der Implementierung. Siehe references/api-patterns.md für Beispiele zu nullable-Datentypen.
- Bewahren Sie nullable Integer-/String-Spalten bei, anstatt sie mit Sentinelwerten zu füllen, es sei denn, der Quelrcode hat dies bereits getan.
- Behalten Sie die
where/mask-Semantik bei, wenn sie eine Bedingung kodieren. Verwenden Sie breitesfillnanur, wenn die Bedingung genau null-nur ist. - Vergleichen Sie mit
to_pandas(nullable=True), wenn der pandas-Referenz nullable Extension-Datentypen verwendet. - Platzieren Sie die Paritätsprüfung in einem wiederverwendbaren Helfer neben dem GPU-Pfad, damit zukünftige Änderungen dieselben nullable-Konvertierungs- und Aggregationsprüfungen ausüben.
- Validieren Sie Zeilenanzahlen, Nullanzahlen, Masken-Wahrheitstabellen, gruppierte Aggregationen und repräsentative Datentypen, bevor Sie semantische Parität behaupten.
Referenzdateien
references/cudf-pandas-accelerator.md– Profiling, Rückfallerkennung, tiefgehende Analyse von cudf.pandasreferences/api-patterns.md– Bekannte API-Lücken, Workarounds, semantische Unterschiedereferences/dask-cudf-patterns.md– Multi-GPU-Muster, Best Practices, Partitionstuning
Externe Dokumentation
Verwenden Sie WebFetch, um auf Anfrage detaillierte API-Signaturen, Parameterbeschreibungen und Beispiele abzurufen.
- cuDF-Dokumentation: https://docs.rapids.ai/api/cudf/stable/
- dask-cuDF-API-Referenz: https://docs.rapids.ai/api/dask-cudf/stable/api/
- GitHub: https://github.com/rapidsai/cudf
- CHANGELOG: https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md
---
name: accelerated-computing-cudf
description: Accelerate pandas workflows with GPU DataFrames using cuDF and dask-cuDF for ETL, joins, groupby, and large-scale data processing.
license: CC-BY-4.0 AND Apache-2.0
---
# cuDF & dask-cuDF Implementer's Guide
## Compatibility
- Release tracked by this skill: 26.04.
- Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.
## Naming
Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.
## Role
You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: `cudf.pandas` for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.
## Critical Rules
1. **Choose the right cuDF path.** Use `cudf.pandas` for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
2. **Size gate: 100K rows minimum.** Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
3. **Keep conversions at boundaries.** Use `.to_pandas()`, `.values`, or `.numpy()` for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
4. **Float32 is your friend.** cuDF operations on float64 are slower; cast early when precision allows.
5. **Validate semantics on representative slices.** For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
6. **For data > GPU memory**, move to dask-cuDF with `enable_cudf_spill=True`. See `references/dask-cudf-patterns.md`.
## Three Paths to GPU DataFrames
### Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)
Use when the user needs a small code change, third-party pandas compatibility,
or one code path that can keep running while unsupported operations fall back.
**Jupyter/IPython:**
```python
%load_ext cudf.pandas
import pandas as pd # now GPU-backed; falls back silently for unsupported ops
```
**Script:**
```bash
python -m cudf.pandas my_script.py
```
**With multiprocessing:**
```python
import cudf.pandas
cudf.pandas.install() # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool
```
Confirm acceleration with the cudf.pandas profiler before claiming speedup.
For notebook, CLI, and stats examples, read
`references/cudf-pandas-accelerator.md`. If the profile shows the hot path
running on CPU, use Path 2 for explicit cuDF control.
### Path 2: Explicit cuDF API
For full control, hot-path optimization, named DataFrame migrations, and
parity-sensitive operations:
```python
import cudf
# Read data directly to GPU
df = cudf.read_parquet("data.parquet")
# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]
# String operations
df["clean"] = df["name"].str.strip().str.lower()
# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds
```
**Keep data on GPU end-to-end.** Only call `.to_pandas()` at the very end for display or CPU or non-GPU handoff.
Prefer explicit cuDF for tasks involving `read_csv`/`read_parquet`, joins,
groupby, reshape, nullable types, `fillna`/`where`, time buckets, rolling
windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when
semantics matter instead of relying on successful execution alone.
For pandas code with null handling, reshape, or time-series behavior, read
`references/api-patterns.md` for the relevant semantic checklist before
rewriting. A `cudf.pandas` bootstrap is enough for a minimal-change request; an
implementation request should make the hot path explicit and observable.
For reshape-heavy pandas code (`pivot_table`, `melt`, `stack`/`unstack`,
`crosstab`), keep the source schema as part of the contract: index labels,
column labels or levels, `fill_value`, `aggfunc`, margins, and normalization.
Use explicit cuDF where the equivalent is supported; use `cudf.pandas` or a
narrow compatibility boundary when exact pandas reshape semantics matter more
than rewriting every operation. Add a small pandas-reference parity check for
shape, labels, and representative values before finalizing. See
`references/api-patterns.md`.
### Path 3: dask-cuDF (Multi-GPU / Large Data)
When dataset exceeds GPU memory. See `references/dask-cudf-patterns.md` for full patterns.
```python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf
cluster = LocalCUDACluster(enable_cudf_spill=True) # one worker per GPU
client = Client(cluster)
ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
```
## Memory Management
**Enable spill before OOM happens** (not after):
```python
import cudf
cudf.set_option("spill", True) # spill to host RAM when GPU is full
```
**RMM pool allocator** (reduces cudaMalloc overhead in pipelines with many allocations):
```python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
```
| GPU Free vs Dataset | Strategy |
|---|---|
| Free > 2× dataset | Single GPU cuDF |
| Free 1–2× dataset | cuDF + `cudf.set_option("spill", True)` |
| Dataset > GPU mem | dask-cuDF |
| Dataset > node mem | dask-cuDF + multi-node (see accelerated-computing-mpf) |
## Troubleshooting
**No speedup vs pandas:**
- Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
- Run `%%cudf.pandas.profile` — high CPU % means many fallbacks. Identify and fix those ops.
- Check `references/api-patterns.md` for known gaps.
**OOM (CUDA out of memory):**
1. Enable spill: `cudf.set_option("spill", True)`
2. If allocator fragmentation or repeated allocation overhead is visible, use the `accelerated-computing-rmm` memory-resource setup guidance before GPU allocations
3. Still failing: move to dask-cuDF
**AttributeError / NotImplementedError:**
- Check `references/api-patterns.md` for the specific operation
- Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
- Use `.to_pandas()` only for the unsupported op, then `.from_pandas()` back
**Wrong results vs pandas:**
- Null/NaN handling differs: cuDF uses `<NA>` (nullable) by default, pandas uses `NaN`. See `references/api-patterns.md`.
- Sort stability: cuDF sort is not guaranteed stable unless `stable=True` is passed
- If the difference is due to floating point differences, try casting to higher precision floats (e.g. `float64` instead of `float32`). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.
## Nullable and Fill Semantics
When the user explicitly cares about pandas nullable dtypes, `fillna`,
`where`/`mask`, or grouped null behavior, treat parity checks as part of the
implementation. See `references/api-patterns.md` for nullable dtype examples.
- Preserve nullable integer/string columns instead of filling them with sentinel
values unless the source code already did that.
- Keep `where`/`mask` semantics when they encode a condition. Use broad
`fillna` only when the condition is exactly null-only.
- Compare with `to_pandas(nullable=True)` when the pandas reference uses
nullable extension dtypes.
- Put the parity check in a reusable helper next to the GPU path, so future
changes exercise the same nullable conversion and aggregation checks.
- Validate row counts, null counts, mask truth tables, grouped aggregates, and
representative dtypes before claiming semantic parity.
## Reference Files
- `references/cudf-pandas-accelerator.md` — Profiling, fallback detection, cudf.pandas deep dive
- `references/api-patterns.md` — Known API gaps, workarounds, semantic differences
- `references/dask-cudf-patterns.md` — Multi-GPU patterns, best practices, partition tuning
## External Documentation
Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.
- **cuDF Documentation:** https://docs.rapids.ai/api/cudf/stable/
- **dask-cuDF API Reference:** https://docs.rapids.ai/api/dask-cudf/stable/api/
- **GitHub:** https://github.com/rapidsai/cudf
- **CHANGELOG:** https://github.com/rapidsai/cudf/blob/main/CHANGELOG.md
Alle Dateien
34 Dateienaccelerated-computing-cudf installieren
Laden Sie die Skill-Dateien herunter und extrahieren Sie diese in Ihr .claude/skills/-Verzeichnis.
ZIP herunterladenKlonen Sie das Repository und kopieren Sie die Skill-Dateien in Ihr Projekt.
git clone https://github.com/NVIDIA/skills/tree/main/skills/accelerated-computing-cudf # Copy SKILL.md to your .claude/skills/ directory
Kopieren





Heim
