cupynumeric-install
NVIDIA/skills
使用 conda 或 pip 安裝並驗證 Python 版的 cuPyNumeric,包括 GPU 使用情況的檢查。
...展開全部cuPyNumeric 安裝(使用者)
目的
運用此技能安裝 cuPyNumeric,以便在 Python 中使用,並驗證安裝是否確實運作(包括 GPU 使用情況)。當使用者希望透過 conda 或 pip 執行 cuPyNumeric 時,即可套用此技能。請勿將其用於從原始碼編譯(以進行修改或貢獻)——此操作超出本技能範圍。
強制性規則
- 切勿執行安裝程序。請勿執行
pip install、conda install或任何安裝程式。僅需輸出指令,並讓使用者自行執行。 - 務必進行隔離。切勿將程式安裝至基礎 conda、系統 Python 或共用全域環境中。
- 先檢測再建議。僅讀取式的
--version檢查是可接受的。
先決條件
在建議任何安裝之前,請確認以下系統需求:
- GPU:運算能力 ≥ 7.0(Volta 及以上)。亦支援僅使用 CPU。
- CUDA:12.2 以上。
- 作業系統:Linux (x86_64 / aarch64)、macOS aarch64(僅限 pip wheel 套件)、Windows(透過 WSL 執行)。
- Python:Linux 系統上為 3.11 至 3.14 版本;macOS aarch64 系統上為 3.11 至 3.13 版本。
- conda:≥ 24.1(僅限 conda path)。
- 套件管理員:conda(上游建議)或 pip。若兩者皆未安裝,請先安裝其中一種(參見說明)。
操作說明
請依序執行以下步驟:確認先決條件、回答範圍確認問題、透過選定的路徑進行安裝,然後進行驗證。
安裝前請確認
- 使用哪種套件管理工具?請檢查 `
conda --version` 和`pip --version`。優先使用 conda(上游推薦);若不可用則改用 pip。 - 目標環境?GPU 機器、僅 CPU 的筆記型電腦、雲端、容器,或遠端/伺服器。
- CUDA 版本?僅在無可見 GPU 的主機上強制使用 GPU 版本時才需確認。請透過
nvidia-smi/nvcc --version進行檢查。
初始化 — 先安裝套件管理器
若既無conda也無pip,請先安裝其中一種。請提供指令及文件連結;切勿直接執行—curl | bash需要使用者信任。
推薦:Miniforge(完整 conda,conda-forge 預設值)
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"
文件:https://github.com/conda-forge/miniforge
替代方案:Python + pip
透過作業系統套件管理員(apt/dnf/brew)或 https://www.python.org/downloads/ 安裝 Python。若現有 Python 環境中缺少 pip,請執行:python -m ensurepip --upgrade。
安裝完成後,請開啟一個新的終端機,確保二進位檔已加入 PATH 環境變數。
安裝 — conda 路徑
conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric
若要安裝至現有環境:conda install -c conda-forge -c legate cupynumeric。
conda 會根據安裝時nvidia-smi是否能正常運作,自動選擇 GPU 或 CPU 版本。若要覆寫此設定,請參閱下方說明。
強制使用 GPU 版本
僅當安裝時未偵測到 GPU 時,才設定CONDA_OVERRIDE_CUDA(例如:為 GPU 主機建置容器)。使用執行時主機的 CUDA 版本:
CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric
夜建版(驗證較少)
conda install -c conda-forge -c legate-nightly cupynumeric
安裝 — pip 路徑
python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric
驗證
基本測試(務必執行)
透過legacy啟動器執行自包含腳本 — 無需檢出儲存庫。
TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum()) # 預期結果為 45
print("matmul:", (b @ b).sum()) # 預期結果為 64.0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"
預期結果:sum: 45及matmul: 64.0。若缺少「legate」指令,表示環境未被啟用 — 請參閱「疑難排解」。
GPU 使用情況檢查(當系統配備受支援的 GPU 時為必選步驟)
煙霧測試通過並不代表一定使用 GPU —— 在配備 GPU 的系統上安裝 CPU 版本同樣會產生正確結果。請執行以下兩個步驟。
1. 強制啟動 GPU。 legate --gpus N會請求 N 個 GPU;若未偵測到 GPU 或安裝的是 CPU 版本,程式將立即失敗。
TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"
預期結果為16777216.0。若偵測到CUDA 驅動程式、libcudart,或無可用 GPU,則表示已安裝 CPU 版本;請使用CONDA_OVERRIDE_CUDA 重新安裝。
2. 確認 GPU 是否已被調用。在同一個終端機中同時執行具有時間限制的 matmul 迴圈與nvidia-smi,避免因開啟第二個終端機而產生的競態條件:
TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
b = a @ a
_ = float(b.sum()) # 強制同步以確保矩陣乘法確實執行
iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5 # Legate 啟動緩衝時間
for _ in $(seq 10); do # 每 1 秒 10 個樣本 — 涵蓋緩慢的啟動過程
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"
預期多數樣本中的memory.used 值應在 GiB 範圍內,且部分樣本的utilization.gpu 值應有顯著變化。若兩者於所有樣本中均維持在基準值,則表示未安裝 GPU 版本 — 請執行conda list cupynumeric並檢查是否存在*_gpu(而非*_cpu)。
進階操作指南
請參閱 verification_examples.md,了解多 GPU 檢查、CPU 備用機制、容器運作及疑難排解。
限制事項
- 請勿在同一環境中混用 conda 與 pip。混用會覆寫首次安裝的設定,並在匯入時導致錯誤。若要切換,請先執行
pip uninstall nvidia-cupynumeric或conda remove cupynumeric。 - 進行多 GPU/多排名執行時,請使用
legate啟動程式。純Python執行為單一進程模式:legate --gpus 2 script.py。 - 在僅有 CPU 的主機上,可透過設定
`CONDA_OVERRIDE_CUDA`強制使用 GPU 版本。否則,conda 會在安裝時根據 `nvidia-smi`的資訊自動選擇 CPU 或 GPU 版本。 - 需 Volta 或更新版本。Pascal(GTX 10xx / P100)不受支援。
- 請確認
`conda --version` 版本≥ 24.1。較舊的版本會無預警地導致變體選擇功能失效。 - 將多節點/MPI/UCX 視為本文件範圍外。相關內容請參閱 https://docs.nvidia.com/legate/latest/networking-wheels.html 及 https://docs.nvidia.com/legate/latest/mpi-wrapper.html。
疑難排解
ModuleNotFoundError:未找到名為 'cupynumeric' 的模組→ 請在同一終端機中執行which python和pip list | grep cupynumeric(或conda list | grep cupynumeric),以找出環境不匹配的問題。- 出現提及 CUDA /
libcudart的 ImportError→ 請使用CONDA_OVERRIDE_CUDA="重新安裝;這表示在配備 GPU 的系統上安裝了 CPU 版本,或 CUDA 版本不匹配。" legate:找不到指令→ 先啟用該環境,然後執行`which legate` 來確認。- 在筆記型電腦上運行速度慢於 NumPy→ 處理小型問題時應預期此情況(Legate 的每項任務開銷)。請參閱 cuPyNumeric 常見問題解答。
另請參閱
- references/verification_examples.md — 驗證與疑難排解指南。
- 上游文件:https://docs.nvidia.com/cupynumeric/latest/installation.html
- Legate 需求:https://docs.nvidia.com/legate/latest/installation.html
---
name: cupynumeric-install
description: Install and verify cuPyNumeric for Python using conda or pip, including GPU usage checks.
license: CC-BY-4.0 OR Apache-2.0
---
# cuPyNumeric Install (user)
## Purpose
Use this skill to install cuPyNumeric for *use* from Python and to verify the install actually works (including GPU usage). Apply it whenever a user wants cuPyNumeric running via conda or pip. Do not use it to build from source (to modify or contribute) — that is out of scope.
## Mandatory rules
- **Never run installs.** Do not run `pip install`, `conda install`, or any installer. Print the command; let the user run it.
- **Always isolate.** No installs into base conda, system Python, or shared global envs.
- **Detect before recommending.** Read-only `--version` checks are fine.
## Prerequisites
Confirm these system requirements before recommending any install:
- **GPU**: Compute Capability ≥ 7.0 (Volta+). CPU-only also supported.
- **CUDA**: 12.2+.
- **OS**: Linux (x86_64 / aarch64), macOS aarch64 (pip wheels only), Windows via WSL.
- **Python**: 3.11 through 3.14 on Linux; 3.11 through 3.13 on macOS aarch64.
- **conda**: ≥ 24.1 (conda path only).
- **Package manager**: conda (upstream-recommended) or pip. If neither is present, bootstrap one first (see Instructions).
## Instructions
Follow these steps in order: confirm the prerequisites, ask the scoping questions, install via the chosen path, then verify.
### Ask before installing
1. **Package manager?** Check `conda --version` and `pip --version`. Prefer conda (upstream-recommended); fall back to pip.
1. **Env target?** GPU machine, CPU-only laptop, cloud, container, or remote/server.
1. **CUDA version?** Ask only when forcing the GPU variant on a host without a visible GPU. Check with `nvidia-smi` / `nvcc --version`.
### Bootstrap — install a package manager first
If neither `conda` nor `pip` is available, install one. **Provide the command and the docs link; do not run it** — `curl | bash` requires user trust.
#### Recommended: Miniforge (full conda, conda-forge default)
```bash
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"
```
Docs: https://github.com/conda-forge/miniforge
#### Alternative: Python + pip
Install Python from your OS package manager (apt/dnf/brew) or https://www.python.org/downloads/. If pip is missing on an existing Python: `python -m ensurepip --upgrade`.
After installing, **open a new shell** so the binary is on PATH.
### Install — conda path
```bash
conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric
```
Into an existing env: `conda install -c conda-forge -c legate cupynumeric`.
conda auto-selects the GPU vs CPU variant from whether `nvidia-smi` works at install time. To override that, see below.
#### Force the GPU variant
Set `CONDA_OVERRIDE_CUDA` only when no GPU is visible at install time (e.g. building a container for a GPU host). Use the runtime host's CUDA version:
```bash
CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric
```
#### Nightly (less validated)
```bash
conda install -c conda-forge -c legate-nightly cupynumeric
```
### Install — pip path
```bash
python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric
```
### Verify
#### Smoke test (always run)
Run a self-contained script through the `legate` launcher — no repo checkout needed.
```bash
TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum()) # expect 45
print("matmul:", (b @ b).sum()) # expect 64.0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"
```
Expect `sum: 45` and `matmul: 64.0`. If `legate` is missing, the env is not activated — see Troubleshooting.
#### GPU usage check (mandatory when a supported GPU is present)
A passing smoke test does **not** prove GPU usage — a CPU-variant install on a GPU box produces correct results too. Run both steps.
**1. Force a GPU launch.** `legate --gpus N` requests N GPUs; fails fast if no GPU is visible or the CPU variant is installed.
```bash
TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"
```
Expect `16777216.0`. If you see `CUDA driver`, `libcudart`, or `no GPUs available`, the CPU variant is installed; reinstall with `CONDA_OVERRIDE_CUDA`.
**2. Confirm the GPU was touched.** Run a deadline-bounded matmul loop alongside `nvidia-smi`, all from one shell — no second-terminal race:
```bash
TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
b = a @ a
_ = float(b.sum()) # force sync so the matmul actually runs
iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5 # buffer for Legate startup
for _ in $(seq 10); do # 10 samples at 1s — covers slow startup
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"
```
Expect `memory.used` in the GiB range across most samples and non-trivial `utilization.gpu` in several. If both stay at baseline across every sample, the GPU variant is not installed — check `conda list cupynumeric` for `*_gpu` (not `*_cpu`).
#### Deeper recipes
See [verification_examples.md](references/verification_examples.md) for multi-GPU checks, CPU fallback, container, and troubleshooting.
## Limitations
- **Don't mix conda and pip in one env.** Mixing overrides the first install and breaks at import. To switch, run `pip uninstall nvidia-cupynumeric` or `conda remove cupynumeric` first.
- **Use the `legate` launcher for multi-GPU / multi-rank runs.** Plain `python` runs single-process: `legate --gpus 2 script.py`.
- **Force the GPU variant on a CPU-only host with `CONDA_OVERRIDE_CUDA`.** conda otherwise auto-selects the CPU or GPU variant from `nvidia-smi` at install time.
- **Require Volta or newer.** Pascal (GTX 10xx / P100) is unsupported.
- **Verify `conda --version` ≥ 24.1.** Older releases silently break variant selection.
- **Treat multi-node / MPI / UCX as out of scope.** Defer to https://docs.nvidia.com/legate/latest/networking-wheels.html and https://docs.nvidia.com/legate/latest/mpi-wrapper.html.
## Troubleshooting
- **`ModuleNotFoundError: No module named 'cupynumeric'`** → Run `which python` and `pip list | grep cupynumeric` (or `conda list | grep cupynumeric`) from the same shell to find the env mismatch.
- **`ImportError` mentioning CUDA / `libcudart`** → Reinstall with `CONDA_OVERRIDE_CUDA="<your-cuda-version>"`; the CPU variant is on a GPU box, or CUDA versions are mismatched.
- **`legate: command not found`** → Activate the env, then run `which legate` to confirm.
- **Slower than NumPy on a laptop** → Expect this for small problems (Legate per-task overhead). See the cuPyNumeric FAQ.
## See also
- [references/verification_examples.md](references/verification_examples.md) — verification + troubleshooting recipes.
- Upstream docs: https://docs.nvidia.com/cupynumeric/latest/installation.html
- Legate requirements: https://docs.nvidia.com/legate/latest/installation.html
安裝 cupynumeric-install
請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。
下載 ZIP複製儲存庫並將技能檔案複製到您的專案中。
git clone https://github.com/NVIDIA/skills/tree/main/skills/cupynumeric-install # Copy SKILL.md to your .claude/skills/ directory
複製





首頁
