옵션
집집 Skill 데이터베이스 관리 cupynumeric-install

cupynumeric-install

NVIDIA/skills NVIDIA/skills

conda 또는 pip를 사용하여 Python용 cuPyNumeric을 설치하고, GPU 사용 여부 확인을 포함하여 정상 작동 여부를 검증합니다.

...모든 것을 확장하십시오
32
업데이트 된 시간 2026년 9월 23일

cuPyNumeric 설치 (사용자)

목적

이 스킬을 사용하여 Python에서 사용할 수 있도록 cuPyNumeric을 설치하고, 설치가 실제로 정상적으로 작동하는지(GPU 사용 포함) 확인합니다. 사용자가 conda 또는 pip를 통해 cuPyNumeric을 실행하고자 할 때마다 이 스킬을 적용하십시오. 소스 코드 빌드(수정 또는 기여 목적)에는 이 스킬을 사용하지 마십시오. 이는 본 스킬의 적용 범위를 벗어납니다.

필수 규칙

  • 절대 직접 설치 명령을 실행하지 마십시오. pip install, conda install 또는 기타 설치 프로그램을 실행하지 마십시오. 명령어를 화면에 표시하고, 사용자가 직접 실행하도록 하십시오.
  • 항상 격리하십시오. 기본 conda, 시스템 Python 또는 공유 글로벌 환경에 설치하지 마십시오.
  • 권장하기 전에 감지하십시오. 읽기 전용 --version 확인은 괜찮습니다.

필수 조건

설치를 권장하기 전에 다음 시스템 요구 사항을 확인하십시오:

  • GPU: 컴퓨트 캐파빌리티(Compute Capability) ≥ 7.0 (Volta 이상). CPU 전용도 지원됩니다.
  • CUDA: 12.2 이상.
  • OS: Linux (x86_64 / aarch64), macOS aarch64 (pip 휠만 지원), WSL을 통한 Windows.
  • Python: Linux의 경우 3.11~3.14; macOS aarch64의 경우 3.11~3.13.
  • conda: 24.1 이상 (conda 경로만 해당).
  • 패키지 관리자: conda(업스트림 권장) 또는 pip. 둘 다 설치되어 있지 않은 경우, 먼저 하나를 부트스트랩해야 합니다(지침 참조).

지침

다음 단계를 순서대로 따르십시오: 필수 구성 요건을 확인하고, 범위 관련 질문을 확인한 후, 선택한 경로를 통해 설치한 다음, 결과를 확인하십시오.

설치 전 확인 사항

  1. 패키지 관리자? conda --version 및 pip --version을 확인하십시오. conda(업스트림 권장)를 우선적으로 사용하며, 그렇지 않은 경우 pip를 대체로 사용하십시오.
  2. 대상 환경? GPU 머신, CPU 전용 노트북, 클라우드, 컨테이너 또는 원격/서버.
  3. CUDA 버전? GPU가 명시적으로 표시되지 않는 호스트에서 GPU 변형을 강제로 적용할 때만 확인하십시오. nvidia-smi / nvcc --version으로 확인하십시오.

부트스트랩 — 먼저 패키지 관리자를 설치하세요

conda나 pip 중 어느 것도 사용할 수 없는 경우, 하나를 설치하십시오. 명령어와 문서 링크를 제공하되, 직접 실행하지 마십시오. — curl | bash는 사용자의 신뢰가 필요합니다.

권장: Miniforge (전체 conda, conda-forge 기본값)

curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"

문서: https://github.com/conda-forge/miniforge

대안: Python + pip

운영 체제의 패키지 관리자(apt/dnf/brew)나 https://www.python.org/downloads/를 통해 Python을 설치하십시오. 기존 Python에 pip가 없는 경우: python -m ensurepip --upgrade.

설치가 완료되면 바이너리가 PATH에 포함되도록 새 셸을 엽니다.

설치 — conda 경로

conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric

기존 환경에 설치하려면: conda install -c conda-forge -c legate cupynumeric.

conda는 설치 시 nvidia-smi가 작동하는지 여부에 따라 GPU 또는 CPU 변형을 자동으로 선택합니다. 이를 무시하려면 아래 내용을 참조하십시오.

GPU 변형 강제 적용

설치 시점에 GPU가 감지되지 않는 경우(예: GPU 호스트용 컨테이너 빌드 시)에만 CONDA_OVERRIDE_CUDA를 설정하십시오. 런타임 호스트의 CUDA 버전을 사용하십시오:

CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric

나이트리(검증 수준이 낮음)

conda install -c conda-forge -c legate-nightly cupynumeric

설치 — pip 경로

python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric

확인

스모크 테스트 (항상 실행)

Legate 런처를 통해 독립 실행형 스크립트를 실행합니다. — 리포지토리 체크아웃이 필요하지 않습니다.

TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("합계:", a.sum())            # 45가 출력되어야 함
print("벡터 곱:", (b @ b).sum())   # 64.0이 출력되어야 함
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"

합계: 45, 행렬 곱셈: 64.0이 출력되어야 합니다. 'legate' 명령이 누락된 경우 환경이 활성화되지 않은 것입니다 — 문제 해결 방법을 참조하십시오.

GPU 사용 여부 확인 (지원되는 GPU가 있을 경우 필수)

스모크 테스트를 통과했다고 해서 GPU가 사용되고 있다는 증거는 아닙니다. GPU 시스템에 CPU 버전으로 설치된 경우에도 올바른 결과가 나올 수 있습니다. 두 단계를 모두 실행하십시오.

1. GPU 실행을 강제합니다. ` legate --gpus N `은 N개의 GPU를 요청합니다. 감지 가능한 GPU가 없거나 CPU 버전이 설치된 경우 즉시 실패합니다.

TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"

결과는 16777216.0이어야 합니다. CUDA 드라이버나 libcudart가 표시되거나 사용 가능한 GPU가 없는 경우, CPU 버전이 설치된 것이므로 CONDA_OVERRIDE_CUDA를 사용하여 재설치하십시오.

2. GPU가 사용되었는지 확인합니다. 하나의 셸에서 nvidia-smi와 함께 데드라인이 지정된 matmul 루프를 실행하세요. 별도의 터미널을 열지 않아도 됩니다:

TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
    b = a @ a
    _ = float(b.sum())   # matmul이 실제로 실행되도록 강제 동기화
    iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5                                     # Legate 시작을 위한 버퍼
for _ in $(seq 10); do                      # 1초 간격으로 10개 샘플 — 느린 시작 시간 대비
  nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
  sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"

대부분의 샘플에서 memory.used가 GiB 단위로 나타나고, 몇몇 샘플에서는 utilization.gpu 값이 무시할 수 없는 수준일 것으로 예상됩니다. 모든 샘플에서 두 값이 모두 기준치에 머무른다면, GPU 변형이 설치되지 않은 것입니다. conda list cupynumeric을 확인하여 *_gpu ( *_cpu가 아님)가 있는지 확인하십시오.

더 심층적인 레시피

다중 GPU 확인, CPU 대체, 컨테이너 및 문제 해결에 대해서는 verification_examples.md를 참조하십시오.

제한 사항

  • 하나의 환경에서 conda와 pip를 혼용하지 마십시오. 혼용 시 첫 번째 설치가 재정의되어 import 시 오류가 발생합니다. 전환하려면 먼저 pip uninstall nvidia-cupynumeric 또는 conda remove cupynumeric을 실행하십시오.
  • 다중 GPU/다중 랭크 실행 시 레가테(legate) 런처를 사용하십시오. 일반 파이썬은 단일 프로세스로 실행됩니다: ` legate --gpus 2 script.py`.
  • CPU 전용 호스트에서 CONDA_OVERRIDE_CUDA를 사용하여 GPU 변형을 강제 적용하십시오. 그렇지 않으면 conda가 설치 시 nvidia-smi를 통해 CPU 또는 GPU 변형을 자동으로 선택합니다.
  • Volta 이상 버전이 필요합니다. Pascal(GTX 10xx / P100)은 지원되지 않습니다.
  • conda --version이 24.1 이상인지 확인하십시오. 이전 릴리스에서는 변형 선택 기능이 무음으로 오작동할 수 있습니다.
  • 다중 노드 / MPI / UCX는 지원 범위에서 제외합니다. https://docs.nvidia.com/legate/latest/networking-wheels.html 및 https://docs.nvidia.com/legate/latest/mpi-wrapper.html를 참조하십시오.

문제 해결

  • ModuleNotFoundError: 'cupynumeric'이라는 모듈이 없습니다 → 동일한 셸에서 which python 및 pip list | grep cupynumeric (또는 conda list | grep cupynumeric) 을 실행하여 환경 불일치를 확인하십시오.
  • CUDA / libcudart와 관련된ImportError → CONDA_OVERRIDE_CUDA=""를 사용하여 재설치하십시오. GPU 시스템에서 CPU 버전을 사용 중이거나 CUDA 버전이 일치하지 않는 경우입니다.
  • legate: 명령을 찾을 수 없음 → 환경을 활성화한 후, which legate를 실행하여 확인하십시오.
  • 노트북에서 NumPy보다 느림 → 작은 규모의 문제에서는 이러한 현상이 예상됩니다(Legate의 작업별 오버헤드). cuPyNumeric FAQ를 참조하십시오.

또한 다음을 참조하십시오.

  • references/verification_examples.md — 검증 및 문제 해결 방법.
  • 업스트림 문서: https://docs.nvidia.com/cupynumeric/latest/installation.html
  • Legate 요구 사항: https://docs.nvidia.com/legate/latest/installation.html
GitHub에서 보기
---
name: cupynumeric-install
description: Install and verify cuPyNumeric for Python using conda or pip, including GPU usage checks.
license: CC-BY-4.0 OR Apache-2.0
---

# cuPyNumeric Install (user)

## Purpose

Use this skill to install cuPyNumeric for *use* from Python and to verify the install actually works (including GPU usage). Apply it whenever a user wants cuPyNumeric running via conda or pip. Do not use it to build from source (to modify or contribute) — that is out of scope.

## Mandatory rules

- **Never run installs.** Do not run `pip install`, `conda install`, or any installer. Print the command; let the user run it.
- **Always isolate.** No installs into base conda, system Python, or shared global envs.
- **Detect before recommending.** Read-only `--version` checks are fine.

## Prerequisites

Confirm these system requirements before recommending any install:

- **GPU**: Compute Capability ≥ 7.0 (Volta+). CPU-only also supported.
- **CUDA**: 12.2+.
- **OS**: Linux (x86_64 / aarch64), macOS aarch64 (pip wheels only), Windows via WSL.
- **Python**: 3.11 through 3.14 on Linux; 3.11 through 3.13 on macOS aarch64.
- **conda**: ≥ 24.1 (conda path only).
- **Package manager**: conda (upstream-recommended) or pip. If neither is present, bootstrap one first (see Instructions).

## Instructions

Follow these steps in order: confirm the prerequisites, ask the scoping questions, install via the chosen path, then verify.

### Ask before installing

1. **Package manager?** Check `conda --version` and `pip --version`. Prefer conda (upstream-recommended); fall back to pip.
1. **Env target?** GPU machine, CPU-only laptop, cloud, container, or remote/server.
1. **CUDA version?** Ask only when forcing the GPU variant on a host without a visible GPU. Check with `nvidia-smi` / `nvcc --version`.

### Bootstrap — install a package manager first

If neither `conda` nor `pip` is available, install one. **Provide the command and the docs link; do not run it** — `curl | bash` requires user trust.

#### Recommended: Miniforge (full conda, conda-forge default)

```bash
curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
bash "Miniforge3-$(uname)-$(uname -m).sh"
```

Docs: https://github.com/conda-forge/miniforge

#### Alternative: Python + pip

Install Python from your OS package manager (apt/dnf/brew) or https://www.python.org/downloads/. If pip is missing on an existing Python: `python -m ensurepip --upgrade`.

After installing, **open a new shell** so the binary is on PATH.

### Install — conda path

```bash
conda create -n cupynumeric -c conda-forge -c legate cupynumeric
conda activate cupynumeric
```

Into an existing env: `conda install -c conda-forge -c legate cupynumeric`.

conda auto-selects the GPU vs CPU variant from whether `nvidia-smi` works at install time. To override that, see below.

#### Force the GPU variant

Set `CONDA_OVERRIDE_CUDA` only when no GPU is visible at install time (e.g. building a container for a GPU host). Use the runtime host's CUDA version:

```bash
CONDA_OVERRIDE_CUDA="12.2" conda install -c conda-forge -c legate cupynumeric
```

#### Nightly (less validated)

```bash
conda install -c conda-forge -c legate-nightly cupynumeric
```

### Install — pip path

```bash
python -m venv .venv
source .venv/bin/activate
pip install nvidia-cupynumeric
```

### Verify

#### Smoke test (always run)

Run a self-contained script through the `legate` launcher — no repo checkout needed.

```bash
TMP=$(mktemp -d)
cat > "$TMP/smoke.py" <<'EOF'
import cupynumeric as np
a = np.arange(10)
b = np.ones((4, 4))
print("sum:", a.sum())            # expect 45
print("matmul:", (b @ b).sum())   # expect 64.0
EOF
legate "$TMP/smoke.py"
rm -rf "$TMP"
```

Expect `sum: 45` and `matmul: 64.0`. If `legate` is missing, the env is not activated — see Troubleshooting.

#### GPU usage check (mandatory when a supported GPU is present)

A passing smoke test does **not** prove GPU usage — a CPU-variant install on a GPU box produces correct results too. Run both steps.

**1. Force a GPU launch.** `legate --gpus N` requests N GPUs; fails fast if no GPU is visible or the CPU variant is installed.

```bash
TMP=$(mktemp -d)
cat > "$TMP/check.py" <<'EOF'
import cupynumeric as np
print(np.ones((4096, 4096)).sum())
EOF
legate --gpus 1 "$TMP/check.py"
rm -rf "$TMP"
```

Expect `16777216.0`. If you see `CUDA driver`, `libcudart`, or `no GPUs available`, the CPU variant is installed; reinstall with `CONDA_OVERRIDE_CUDA`.

**2. Confirm the GPU was touched.** Run a deadline-bounded matmul loop alongside `nvidia-smi`, all from one shell — no second-terminal race:

```bash
TMPDIR_GPU=$(mktemp -d)
SCRIPT="$TMPDIR_GPU/cupynumeric_gpu_check.py"
cat > "$SCRIPT" <<'EOF'
import cupynumeric as np, time
a = np.ones((10000, 10000))
deadline = time.time() + 20
iters = 0
while time.time() < deadline:
    b = a @ a
    _ = float(b.sum())   # force sync so the matmul actually runs
    iters += 1
print("iters:", iters)
EOF
legate --gpus 1 "$SCRIPT" &
WORKLOAD=$!
sleep 5                                     # buffer for Legate startup
for _ in $(seq 10); do                      # 10 samples at 1s — covers slow startup
  nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv,noheader
  sleep 1
done
wait "$WORKLOAD"
rm -rf "$TMPDIR_GPU"
```

Expect `memory.used` in the GiB range across most samples and non-trivial `utilization.gpu` in several. If both stay at baseline across every sample, the GPU variant is not installed — check `conda list cupynumeric` for `*_gpu` (not `*_cpu`).

#### Deeper recipes

See [verification_examples.md](references/verification_examples.md) for multi-GPU checks, CPU fallback, container, and troubleshooting.

## Limitations

- **Don't mix conda and pip in one env.** Mixing overrides the first install and breaks at import. To switch, run `pip uninstall nvidia-cupynumeric` or `conda remove cupynumeric` first.
- **Use the `legate` launcher for multi-GPU / multi-rank runs.** Plain `python` runs single-process: `legate --gpus 2 script.py`.
- **Force the GPU variant on a CPU-only host with `CONDA_OVERRIDE_CUDA`.** conda otherwise auto-selects the CPU or GPU variant from `nvidia-smi` at install time.
- **Require Volta or newer.** Pascal (GTX 10xx / P100) is unsupported.
- **Verify `conda --version` ≥ 24.1.** Older releases silently break variant selection.
- **Treat multi-node / MPI / UCX as out of scope.** Defer to https://docs.nvidia.com/legate/latest/networking-wheels.html and https://docs.nvidia.com/legate/latest/mpi-wrapper.html.

## Troubleshooting

- **`ModuleNotFoundError: No module named 'cupynumeric'`** → Run `which python` and `pip list | grep cupynumeric` (or `conda list | grep cupynumeric`) from the same shell to find the env mismatch.
- **`ImportError` mentioning CUDA / `libcudart`** → Reinstall with `CONDA_OVERRIDE_CUDA="<your-cuda-version>"`; the CPU variant is on a GPU box, or CUDA versions are mismatched.
- **`legate: command not found`** → Activate the env, then run `which legate` to confirm.
- **Slower than NumPy on a laptop** → Expect this for small problems (Legate per-task overhead). See the cuPyNumeric FAQ.

## See also

- [references/verification_examples.md](references/verification_examples.md) — verification + troubleshooting recipes.
- Upstream docs: https://docs.nvidia.com/cupynumeric/latest/installation.html
- Legate requirements: https://docs.nvidia.com/legate/latest/installation.html

cupynumeric-install 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/cupynumeric-install # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 NVIDIA/skills

관련 스킬

microservices-patterns
업데이트 된 시간 2026년 6월 29일
jpa-patterns
업데이트 된 시간 2026년 6월 30일
fabric-lakehouse
업데이트 된 시간 2026년 6월 30일
prisma-expert
업데이트 된 시간 2026년 6월 29일
OR