cupynumeric-install
NVIDIA/skills
使用 conda 或 pip 安装并验证 Python 版的 cuPyNumeric,包括 GPU 使用情况的检查。
...展开全部cuPyNumeric 安装(用户)
目的
使用本技能安装 cuPyNumeric 以便在 Python 中使用,并验证安装是否正常(包括 GPU 使用情况)。当用户希望通过 conda 或 pip 运行 cuPyNumeric 时,请应用本技能。请勿将其用于从源代码构建(以进行修改或贡献)——这超出了本技能的范围。
强制性规则
- 切勿直接执行安装操作。不要运行
`pip install`、`conda install` 或任何安装程序。仅输出命令行,由用户自行执行。 - 始终保持隔离。禁止将软件安装到基础 conda、系统 Python 或共享的全局环境中。
- 先检测后推荐。仅读取
--version版本信息的检查是允许的。
先决条件
在推荐任何安装之前,请确认以下系统要求:
- GPU:计算能力 ≥ 7.0(Volta 及以上)。也支持仅 CPU 环境。
- CUDA:12.2 及以上。
- 操作系统:Linux(x86_64 / aarch64)、macOS aarch64(仅限 pip 轮子包)、通过 WSL 运行的 Windows。
- Python:Linux 系统上为 3.11 至 3.14 版本;macOS aarch64 系统上为 3.11 至 3.13 版本。
- conda:≥ 24.1(仅限 conda 路径)。
- 包管理器:conda(上游推荐)或 pip。若两者均未安装,请先安装其中之一(参见操作指南)。
操作指南
请按以下步骤依次操作:确认先决条件、回答范围相关问题、通过所选路径进行安装,然后进行验证。
安装前请确认
- 使用哪个包管理器?检查
conda --version和pip --version。优先使用 conda(上游推荐);若不可用则退而求其次使用 pip。 - 目标环境?GPU 机器、仅 CPU 的笔记本电脑、云环境、容器,或远程/服务器。
- CUDA 版本?仅当在无可见 GPU 的主机上强制使用 GPU 版本时才需询问。请通过
`nvidia-smi`或`nvcc --version`进行检查。
初始化 — 先安装包管理器
如果既没有conda也没有pip,请安装其中一个。提供命令及文档链接;请勿直接运行该命令 ——curl | bash需要用户信任。
推荐:Miniforge(完整 conda,conda-forge 默认)
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"
文档:https://github.com/conda-forge/miniforge
替代方案:Python + pip
通过操作系统包管理器(apt/dnf/brew)或 https://www.python.org/downloads/ 安装 Python。如果现有 Python 环境中缺少 pip,请执行:python -m ensurepip --upgrade。
安装完成后,请打开一个新的终端,以确保二进制文件已加入 PATH 环境变量。
安装 — conda 路径
conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric
在现有环境中:conda install -c conda-forge -c legate cupynumeric。
conda 会根据安装时nvidia-smi是否正常运行,自动选择 GPU 或 CPU 版本。如需覆盖此设置,请参见下文。
强制使用 GPU 版本
仅当安装时未检测到 GPU 时(例如为 GPU 主机构建容器),才需设置CONDA_OVERRIDE_CUDA。使用运行时主机的 CUDA 版本:
CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric
夜间构建版(验证较少)
conda install -c conda-forge -c legate-nightly cupynumeric
安装 — pip 路径
python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric
验证
快速测试(始终运行)
通过Legate启动器运行一个自包含脚本——无需检出代码库。
TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum()) # 预期结果为 45
print("matmul:", (b @ b).sum()) # 预期结果为 64.0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"
预期结果:sum 为 45 ,matmul 为 64.0。如果缺少"legate"命令,则环境未激活——请参阅“故障排除”。
GPU 使用情况检查(当存在受支持的 GPU 时为必选步骤)
通过烟雾测试并不证明使用了 GPU —— 在配备 GPU 的机器上安装 CPU 版本同样会产生正确结果。请运行以下两个步骤。
1. 强制启动 GPU。 legate --gpus N会请求 N 个 GPU;若未检测到 GPU 或安装了 CPU 版本,则会立即失败。
TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"
预期结果为16777216.0。若检测到CUDA 驱动程序、libcudart 或无可用 GPU,则表示已安装 CPU 版本;请使用CONDA_OVERRIDE_CUDA 重新安装。
2. 确认是否使用了 GPU。在同一终端中同时运行一个具有时间限制的矩阵乘法循环和nvidia-smi,避免因使用第二个终端而产生的竞争条件:
TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
b = a @ a
_ = float(b.sum()) # 强制同步以确保矩阵乘法实际运行
iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5 # 为 Legate 启动预留缓冲时间
for _ in $(seq 10); do # 每秒 10 个采样 — 覆盖启动延迟
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"
预计大多数样本中的memory.used 值应在 GiB 量级,且部分样本的utilization.gpu 值不低。如果这两个指标在所有样本中均保持在基准水平,则说明未安装 GPU 版本——请运行conda list cupynumeric检查是否包含*_gpu(而非*_cpu)。
更深入的配置方法
有关多GPU检测、CPU回退、容器及故障排除的详细信息,请参阅 verification_examples.md。
限制
- 请勿在同一环境中混合使用 conda 和 pip。混合使用会覆盖首次安装的结果,并在导入时导致错误。若需切换,请先运行
`pip uninstall nvidia-cupynumeric` 或`conda remove cupynumeric`。 - 进行多GPU/多rank运行时,请
使用legate启动器。普通Python运行采用单进程模式:legate --gpus 2 script.py。 - 在仅支持 CPU 的主机上,可通过
CONDA_OVERRIDE_CUDA强制使用 GPU 版本。否则,conda 会在安装时根据nvidia-smi自动选择 CPU 或 GPU 版本。 - 要求 Volta 或更新版本。Pascal(GTX 10xx / P100)不受支持。
- 请确认
conda --version≥ 24.1。旧版本会静默破坏变体选择功能。 - 将多节点 / MPI / UCX 视为超出本指南范围。相关内容请参见 https://docs.nvidia.com/legate/latest/networking-wheels.html 和 https://docs.nvidia.com/legate/latest/mpi-wrapper.html。
故障排除
ModuleNotFoundError:未找到名为 'cupynumeric' 的模块→ 在同一终端中运行`which python` 和`pip list | grep cupynumeric`(或 `conda list | grep cupynumeric`)以查找环境不匹配的问题。- 提及 CUDA /
libcudart的 ImportError→ 使用CONDA_OVERRIDE_CUDA="重新安装;这可能是 CPU 版本在 GPU 机器上运行,或 CUDA 版本不匹配。" legate:未找到命令→ 激活该环境,然后运行`which legate` 进行确认。- 在笔记本电脑上运行速度慢于 NumPy→ 处理小规模问题时可能会出现这种情况(Legate 的每任务开销)。请参阅 cuPyNumeric 常见问题解答。
另请参阅
- references/verification_examples.md — 验证与故障排除指南。
- 上游文档:https://docs.nvidia.com/cupynumeric/latest/installation.html
- Legate 依赖项:https://docs.nvidia.com/legate/latest/installation.html
---
name: cupynumeric-install
description: Install and verify cuPyNumeric for Python using conda or pip, including GPU usage checks.
license: CC-BY-4.0 OR Apache-2.0
---
# cuPyNumeric Install (user)
## Purpose
Use this skill to install cuPyNumeric for *use* from Python and to verify the install actually works (including GPU usage). Apply it whenever a user wants cuPyNumeric running via conda or pip. Do not use it to build from source (to modify or contribute) — that is out of scope.
## Mandatory rules
- **Never run installs.** Do not run `pip install`, `conda install`, or any installer. Print the command; let the user run it.
- **Always isolate.** No installs into base conda, system Python, or shared global envs.
- **Detect before recommending.** Read-only `--version` checks are fine.
## Prerequisites
Confirm these system requirements before recommending any install:
- **GPU**: Compute Capability ≥ 7.0 (Volta+). CPU-only also supported.
- **CUDA**: 12.2+.
- **OS**: Linux (x86_64 / aarch64), macOS aarch64 (pip wheels only), Windows via WSL.
- **Python**: 3.11 through 3.14 on Linux; 3.11 through 3.13 on macOS aarch64.
- **conda**: ≥ 24.1 (conda path only).
- **Package manager**: conda (upstream-recommended) or pip. If neither is present, bootstrap one first (see Instructions).
## Instructions
Follow these steps in order: confirm the prerequisites, ask the scoping questions, install via the chosen path, then verify.
### Ask before installing
1. **Package manager?** Check `conda --version` and `pip --version`. Prefer conda (upstream-recommended); fall back to pip.
1. **Env target?** GPU machine, CPU-only laptop, cloud, container, or remote/server.
1. **CUDA version?** Ask only when forcing the GPU variant on a host without a visible GPU. Check with `nvidia-smi` / `nvcc --version`.
### Bootstrap — install a package manager first
If neither `conda` nor `pip` is available, install one. **Provide the command and the docs link; do not run it** — `curl | bash` requires user trust.
#### Recommended: Miniforge (full conda, conda-forge default)
```bash
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"
```
Docs: https://github.com/conda-forge/miniforge
#### Alternative: Python + pip
Install Python from your OS package manager (apt/dnf/brew) or https://www.python.org/downloads/. If pip is missing on an existing Python: `python -m ensurepip --upgrade`.
After installing, **open a new shell** so the binary is on PATH.
### Install — conda path
```bash
conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric
```
Into an existing env: `conda install -c conda-forge -c legate cupynumeric`.
conda auto-selects the GPU vs CPU variant from whether `nvidia-smi` works at install time. To override that, see below.
#### Force the GPU variant
Set `CONDA_OVERRIDE_CUDA` only when no GPU is visible at install time (e.g. building a container for a GPU host). Use the runtime host's CUDA version:
```bash
CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric
```
#### Nightly (less validated)
```bash
conda install -c conda-forge -c legate-nightly cupynumeric
```
### Install — pip path
```bash
python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric
```
### Verify
#### Smoke test (always run)
Run a self-contained script through the `legate` launcher — no repo checkout needed.
```bash
TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum()) # expect 45
print("matmul:", (b @ b).sum()) # expect 64.0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"
```
Expect `sum: 45` and `matmul: 64.0`. If `legate` is missing, the env is not activated — see Troubleshooting.
#### GPU usage check (mandatory when a supported GPU is present)
A passing smoke test does **not** prove GPU usage — a CPU-variant install on a GPU box produces correct results too. Run both steps.
**1. Force a GPU launch.** `legate --gpus N` requests N GPUs; fails fast if no GPU is visible or the CPU variant is installed.
```bash
TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"
```
Expect `16777216.0`. If you see `CUDA driver`, `libcudart`, or `no GPUs available`, the CPU variant is installed; reinstall with `CONDA_OVERRIDE_CUDA`.
**2. Confirm the GPU was touched.** Run a deadline-bounded matmul loop alongside `nvidia-smi`, all from one shell — no second-terminal race:
```bash
TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
b = a @ a
_ = float(b.sum()) # force sync so the matmul actually runs
iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5 # buffer for Legate startup
for _ in $(seq 10); do # 10 samples at 1s — covers slow startup
nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"
```
Expect `memory.used` in the GiB range across most samples and non-trivial `utilization.gpu` in several. If both stay at baseline across every sample, the GPU variant is not installed — check `conda list cupynumeric` for `*_gpu` (not `*_cpu`).
#### Deeper recipes
See [verification_examples.md](references/verification_examples.md) for multi-GPU checks, CPU fallback, container, and troubleshooting.
## Limitations
- **Don't mix conda and pip in one env.** Mixing overrides the first install and breaks at import. To switch, run `pip uninstall nvidia-cupynumeric` or `conda remove cupynumeric` first.
- **Use the `legate` launcher for multi-GPU / multi-rank runs.** Plain `python` runs single-process: `legate --gpus 2 script.py`.
- **Force the GPU variant on a CPU-only host with `CONDA_OVERRIDE_CUDA`.** conda otherwise auto-selects the CPU or GPU variant from `nvidia-smi` at install time.
- **Require Volta or newer.** Pascal (GTX 10xx / P100) is unsupported.
- **Verify `conda --version` ≥ 24.1.** Older releases silently break variant selection.
- **Treat multi-node / MPI / UCX as out of scope.** Defer to https://docs.nvidia.com/legate/latest/networking-wheels.html and https://docs.nvidia.com/legate/latest/mpi-wrapper.html.
## Troubleshooting
- **`ModuleNotFoundError: No module named 'cupynumeric'`** → Run `which python` and `pip list | grep cupynumeric` (or `conda list | grep cupynumeric`) from the same shell to find the env mismatch.
- **`ImportError` mentioning CUDA / `libcudart`** → Reinstall with `CONDA_OVERRIDE_CUDA="<your-cuda-version>"`; the CPU variant is on a GPU box, or CUDA versions are mismatched.
- **`legate: command not found`** → Activate the env, then run `which legate` to confirm.
- **Slower than NumPy on a laptop** → Expect this for small problems (Legate per-task overhead). See the cuPyNumeric FAQ.
## See also
- [references/verification_examples.md](references/verification_examples.md) — verification + troubleshooting recipes.
- Upstream docs: https://docs.nvidia.com/cupynumeric/latest/installation.html
- Legate requirements: https://docs.nvidia.com/legate/latest/installation.html





首页
