Option
HeimHeim Skill Datenbankverwaltung cupynumeric-install

cupynumeric-install

NVIDIA/skills NVIDIA/skills

Installieren und überprüfen Sie cuPyNumeric für Python mithilfe von conda oder pip, einschließlich der Überprüfung der GPU-Nutzung.

...Alle erweitern
32
Zeit aktualisiert 23. September 2026

cuPyNumeric-Installation (Benutzer)

Zweck

Nutzen Sie diese Anleitung, um cuPyNumeric für die Verwendung in Python zu installieren und zu überprüfen, ob die Installation tatsächlich funktioniert (einschließlich der GPU-Nutzung). Wenden Sie sie an, wenn ein Benutzer cuPyNumeric über conda oder pip ausführen möchte. Verwenden Sie sie nicht, um aus dem Quellcode zu kompilieren (um Änderungen vorzunehmen oder Beiträge zu leisten) – dies liegt außerhalb des Anwendungsbereichs.

Verbindliche Regeln

  • Führen Sie niemals Installationen aus. Führen Sie weder „pip install“ noch „conda install“ oder andere Installationsprogramme aus. Geben Sie den Befehl aus; lassen Sie den Benutzer ihn ausführen.
  • Stets isoliert vorgehen. Keine Installationen in das Basis-conda, das System-Python oder gemeinsam genutzte globale Umgebungen.
  • Erst prüfen, dann empfehlen. Lesezugriff auf „--version“ ist in Ordnung.

Voraussetzungen

Überprüfen Sie diese Systemanforderungen, bevor Sie eine Installation empfehlen:

  • GPU: Rechenleistung ≥ 7.0 (Volta+). Auch reine CPU-Nutzung wird unterstützt.
  • CUDA: 12.2+.
  • Betriebssystem: Linux (x86_64 / aarch64), macOS aarch64 (nur pip-Wheels), Windows über WSL.
  • Python: 3.11 bis 3.14 unter Linux; 3.11 bis 3.13 unter macOS aarch64.
  • conda: ≥ 24.1 (nur conda-Pfad).
  • Paketmanager: conda (vom Upstream empfohlen) oder pip. Falls keiner der beiden vorhanden ist, installieren Sie zunächst einen (siehe Anleitung).

Anleitung

Führen Sie diese Schritte der Reihe nach aus: Überprüfen Sie die Voraussetzungen, beantworten Sie die Fragen zum Anwendungsbereich, führen Sie die Installation über den gewählten Pfad durch und überprüfen Sie anschließend das Ergebnis.

Vor der Installation abklären

  1. Paketmanager? Überprüfen Sie `conda --version ` und `pip --version`. Bevorzugen Sie conda (vom Upstream empfohlen); greifen Sie andernfalls auf pip zurück.
  2. Zielumgebung? GPU-Rechner, reiner CPU-Laptop, Cloud, Container oder Remote-/Serverumgebung.
  3. CUDA-Version? Nur abfragen, wenn die GPU-Variante auf einem Host ohne sichtbare GPU erzwungen wird. Überprüfen Sie dies mit `nvidia-smi ` bzw. `nvcc --version`.

Bootstrap – Installieren Sie zuerst einen Paketmanager

Falls weder conda noch pip verfügbar sind, installieren Sie einen der beiden. Geben Sie den Befehl und den Link zur Dokumentation an; führen Sie ihn nicht aus – „curl | bash“ erfordert das Vertrauen des Benutzers.

Empfohlen: Miniforge (vollständiges conda, conda-forge als Standard)

curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash „Miniforge3-$(uname)-$(uname -m).sh“

Dokumentation: https://github.com/conda-forge/miniforge

Alternative: Python + pip

Installieren Sie Python über den Paketmanager Ihres Betriebssystems (apt/dnf/brew) oder unter https://www.python.org/downloads/. Falls pip bei einer bereits installierten Python-Version fehlt: python -m ensurepip --upgrade.

Öffnen Sie nach der Installation eine neue Shell, damit die Binärdatei im PATH steht.

Installieren – conda-Pfad

conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric

In eine bestehende Umgebung: ` conda install -c conda-forge -c legate cupynumeric`.

conda wählt automatisch die GPU- oder CPU-Variante aus, je nachdem, ob „nvidia-smi“ zum Zeitpunkt der Installation funktioniert. Um dies zu überschreiben, siehe unten.

GPU-Variante erzwingen

Setzen Sie CONDA_OVERRIDE_CUDA nur, wenn zum Zeitpunkt der Installation keine GPU sichtbar ist (z. B. beim Erstellen eines Containers für einen GPU-Host). Verwenden Sie die CUDA-Version des Laufzeit-Hosts:

CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric

Nightly (weniger geprüft)

conda install -c conda-forge -c legate-nightly cupynumeric

Installieren — pip-Pfad

python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric

Überprüfen

Smoke-Test (immer ausführen)

Führen Sie ein eigenständiges Skript über den Legate -Launcher aus – kein Auschecken des Repositories erforderlich.

TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum())            # erwarteter Wert: 45
print("matmul:", (b @ b).sum())   # erwarteter Wert: 64,0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"

Erwartet werden sum: 45 und matmul: 64,0. Fehlt „legate“, ist die Umgebung nicht aktiviert – siehe Fehlerbehebung.

Überprüfung der GPU-Nutzung (obligatorisch, wenn eine unterstützte GPU vorhanden ist)

Ein bestandener Smoke-Test ist kein Beweis für die GPU-Nutzung – auch eine Installation der CPU-Variante auf einem GPU-Rechner liefert korrekte Ergebnisse. Führen Sie beide Schritte aus.

1. Erzwinge einen GPU-Start. „legate --gpus N“ fordert N GPUs an; schlägt schnell fehl, wenn keine GPU sichtbar ist oder die CPU-Variante installiert ist.

TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"

Erwartet wird der Wert 16777216,0. Wenn CUDA-Treiber, libcudart oder keine verfügbaren GPUs angezeigt werden, ist die CPU-Variante installiert; führen Sie die Neuinstallation mit CONDA_OVERRIDE_CUDA durch.

2. Überprüfen Sie, ob die GPU genutzt wurde. Führen Sie eine zeitlich begrenzte Matmul-Schleife parallel zu ` nvidia-smi` aus, alles in einer Shell – kein Wettlauf um ein zweites Terminal:

TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
    b = a @ a
    _ = float(b.sum())   # Synchronisation erzwingen, damit die Matmul-Operation tatsächlich ausgeführt wird
    iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5                                     # Puffer für den Start von Legate
for _ in $(seq 10); do                      # 10 Stichproben im Abstand von 1 s – deckt langsamen Start ab
  nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
  sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"

Es ist zu erwarten, dass „memory.used“ bei den meisten Samples im GiB-Bereich liegt und „utilization.gpu“ bei einigen nicht unerheblich ist. Wenn beide Werte bei jedem Sample auf dem Basiswert bleiben, ist die GPU-Variante nicht installiert – überprüfe „conda list cupynumeric“ auf *_gpu (nicht *_cpu).

Erweiterte Anleitungen

Siehe „verification_examples.md“ für Multi-GPU-Prüfungen, CPU-Fallback, Container und Fehlerbehebung.

Einschränkungen

  • Verwenden Sie conda und pip nicht in derselben Umgebung. Eine gemischte Verwendung überschreibt die erste Installation und führt beim Import zu Fehlern. Um zu wechseln, führen Sie zunächst „pip uninstall nvidia-cupynumeric“ oder „conda remove cupynumeric“ aus.
  • Verwenden Sie den „legate “-Launcher für Multi-GPU-/Multi-Rank-Läufe. Reines Python läuft im Einzelprozess: „legate --gpus 2 script.py“.
  • Erzwinge die GPU-Variante auf einem reinen CPU-Host mit `CONDA_OVERRIDE_CUDA`. Andernfalls wählt conda bei der Installation automatisch die CPU- oder GPU-Variante anhand von ` nvidia-smi ` aus.
  • Volta oder neuer erforderlich. Pascal (GTX 10xx / P100) wird nicht unterstützt.
  • Stellen Sie sicher, dass conda --version ≥ 24.1 ist. Ältere Versionen führen zu stillschweigenden Fehlern bei der Variantenauswahl.
  • Behandeln Sie Multi-Node-/MPI-/UCX-Szenarien als außerhalb des Geltungsbereichs. Verweisen Sie auf https://docs.nvidia.com/legate/latest/networking-wheels.html und https://docs.nvidia.com/legate/latest/mpi-wrapper.html.

Fehlerbehebung

  • ModuleNotFoundError: Kein Modul namens „cupynumeric“ → Führen Sie „which python“ und „pip list | grep cupynumeric“ (oder „conda list | grep cupynumeric“) in derselben Shell aus, um die Umgebungsinkongruenz zu ermitteln.
  • ImportError mit Hinweis auf CUDA / libcudart → Neuinstallation mit CONDA_OVERRIDE_CUDA=""; die CPU-Variante befindet sich auf einem GPU-Rechner oder die CUDA-Versionen stimmen nicht überein.
  • legate: Befehl nicht gefunden → Aktivieren Sie die Umgebung und führen Sie anschließend „which legate“ aus, um dies zu überprüfen.
  • Langsamer als NumPy auf einem Laptop → Dies ist bei kleinen Problemen zu erwarten (Legate-Overhead pro Aufgabe). Siehe die cuPyNumeric-FAQ.

Siehe auch

  • references/verification_examples.md – Anleitungen zur Verifizierung und Fehlerbehebung.
  • Upstream-Dokumentation: https://docs.nvidia.com/cupynumeric/latest/installation.html
  • Legate-Anforderungen: https://docs.nvidia.com/legate/latest/installation.html
Auf GitHub ansehen
---
name: cupynumeric-install
description: Install and verify cuPyNumeric for Python using conda or pip, including GPU usage checks.
license: CC-BY-4.0 OR Apache-2.0
---

# cuPyNumeric Install (user)

## Purpose

Use this skill to install cuPyNumeric for *use* from Python and to verify the install actually works (including GPU usage). Apply it whenever a user wants cuPyNumeric running via conda or pip. Do not use it to build from source (to modify or contribute) — that is out of scope.

## Mandatory rules

- **Never run installs.** Do not run `pip install`, `conda install`, or any installer. Print the command; let the user run it.
- **Always isolate.** No installs into base conda, system Python, or shared global envs.
- **Detect before recommending.** Read-only `--version` checks are fine.

## Prerequisites

Confirm these system requirements before recommending any install:

- **GPU**: Compute Capability ≥ 7.0 (Volta+). CPU-only also supported.
- **CUDA**: 12.2+.
- **OS**: Linux (x86_64 / aarch64), macOS aarch64 (pip wheels only), Windows via WSL.
- **Python**: 3.11 through 3.14 on Linux; 3.11 through 3.13 on macOS aarch64.
- **conda**: ≥ 24.1 (conda path only).
- **Package manager**: conda (upstream-recommended) or pip. If neither is present, bootstrap one first (see Instructions).

## Instructions

Follow these steps in order: confirm the prerequisites, ask the scoping questions, install via the chosen path, then verify.

### Ask before installing

1. **Package manager?** Check `conda --version` and `pip --version`. Prefer conda (upstream-recommended); fall back to pip.
1. **Env target?** GPU machine, CPU-only laptop, cloud, container, or remote/server.
1. **CUDA version?** Ask only when forcing the GPU variant on a host without a visible GPU. Check with `nvidia-smi` / `nvcc --version`.

### Bootstrap — install a package manager first

If neither `conda` nor `pip` is available, install one. **Provide the command and the docs link; do not run it** — `curl | bash` requires user trust.

#### Recommended: Miniforge (full conda, conda-forge default)

```bash
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"
```

Docs: https://github.com/conda-forge/miniforge

#### Alternative: Python + pip

Install Python from your OS package manager (apt/dnf/brew) or https://www.python.org/downloads/. If pip is missing on an existing Python: `python -m ensurepip --upgrade`.

After installing, **open a new shell** so the binary is on PATH.

### Install — conda path

```bash
conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric
```

Into an existing env: `conda install -c conda-forge -c legate cupynumeric`.

conda auto-selects the GPU vs CPU variant from whether `nvidia-smi` works at install time. To override that, see below.

#### Force the GPU variant

Set `CONDA_OVERRIDE_CUDA` only when no GPU is visible at install time (e.g. building a container for a GPU host). Use the runtime host's CUDA version:

```bash
CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric
```

#### Nightly (less validated)

```bash
conda install -c conda-forge -c legate-nightly cupynumeric
```

### Install — pip path

```bash
python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric
```

### Verify

#### Smoke test (always run)

Run a self-contained script through the `legate` launcher — no repo checkout needed.

```bash
TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum())            # expect 45
print("matmul:", (b @ b).sum())   # expect 64.0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"
```

Expect `sum: 45` and `matmul: 64.0`. If `legate` is missing, the env is not activated — see Troubleshooting.

#### GPU usage check (mandatory when a supported GPU is present)

A passing smoke test does **not** prove GPU usage — a CPU-variant install on a GPU box produces correct results too. Run both steps.

**1. Force a GPU launch.** `legate --gpus N` requests N GPUs; fails fast if no GPU is visible or the CPU variant is installed.

```bash
TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"
```

Expect `16777216.0`. If you see `CUDA driver`, `libcudart`, or `no GPUs available`, the CPU variant is installed; reinstall with `CONDA_OVERRIDE_CUDA`.

**2. Confirm the GPU was touched.** Run a deadline-bounded matmul loop alongside `nvidia-smi`, all from one shell — no second-terminal race:

```bash
TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
    b = a @ a
    _ = float(b.sum())   # force sync so the matmul actually runs
    iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5                                     # buffer for Legate startup
for _ in $(seq 10); do                      # 10 samples at 1s — covers slow startup
  nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
  sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"
```

Expect `memory.used` in the GiB range across most samples and non-trivial `utilization.gpu` in several. If both stay at baseline across every sample, the GPU variant is not installed — check `conda list cupynumeric` for `*_gpu` (not `*_cpu`).

#### Deeper recipes

See [verification_examples.md](references/verification_examples.md) for multi-GPU checks, CPU fallback, container, and troubleshooting.

## Limitations

- **Don't mix conda and pip in one env.** Mixing overrides the first install and breaks at import. To switch, run `pip uninstall nvidia-cupynumeric` or `conda remove cupynumeric` first.
- **Use the `legate` launcher for multi-GPU / multi-rank runs.** Plain `python` runs single-process: `legate --gpus 2 script.py`.
- **Force the GPU variant on a CPU-only host with `CONDA_OVERRIDE_CUDA`.** conda otherwise auto-selects the CPU or GPU variant from `nvidia-smi` at install time.
- **Require Volta or newer.** Pascal (GTX 10xx / P100) is unsupported.
- **Verify `conda --version` ≥ 24.1.** Older releases silently break variant selection.
- **Treat multi-node / MPI / UCX as out of scope.** Defer to https://docs.nvidia.com/legate/latest/networking-wheels.html and https://docs.nvidia.com/legate/latest/mpi-wrapper.html.

## Troubleshooting

- **`ModuleNotFoundError: No module named 'cupynumeric'`** → Run `which python` and `pip list | grep cupynumeric` (or `conda list | grep cupynumeric`) from the same shell to find the env mismatch.
- **`ImportError` mentioning CUDA / `libcudart`** → Reinstall with `CONDA_OVERRIDE_CUDA="<your-cuda-version>"`; the CPU variant is on a GPU box, or CUDA versions are mismatched.
- **`legate: command not found`** → Activate the env, then run `which legate` to confirm.
- **Slower than NumPy on a laptop** → Expect this for small problems (Legate per-task overhead). See the cuPyNumeric FAQ.

## See also

- [references/verification_examples.md](references/verification_examples.md) — verification + troubleshooting recipes.
- Upstream docs: https://docs.nvidia.com/cupynumeric/latest/installation.html
- Legate requirements: https://docs.nvidia.com/legate/latest/installation.html

cupynumeric-install installieren

Laden Sie die Skill-Dateien herunter und entpacken Sie sie in Ihr Verzeichnis „.claude/skills/“.

ZIP herunterladen

Klonen Sie das Repository und kopieren Sie die Skill-Dateien in Ihr Projekt.

git clone https://github.com/NVIDIA/skills/tree/main/skills/cupynumeric-install # Copy SKILL.md to your .claude/skills/ directory

Kopieren Kopieren
Schnelle Einrichtung: Kopiere den Skill-Ordner nach .claude/skills/. Claude erkennt den Skill automatisch und nutzt ihn.
Repository NVIDIA/skills

Ähnliche Skills

microservices-patterns
Zeit aktualisiert 29. Juni 2026
jpa-patterns
Zeit aktualisiert 30. Juni 2026
fabric-lakehouse
Zeit aktualisiert 30. Juni 2026
prisma-expert
Zeit aktualisiert 29. Juni 2026
OR