tao-setup-nvidia-gpu-host
NVIDIA/skills
GPU 가속 Docker 및 Kubernetes 호스트용 NVIDIA 드라이버, CUDA 툴킷, NVIDIA 컨테이너 툴킷을 확인하고 설치합니다. 자동 설치 및 읽기 전용 확인 모드를 통해 다양한 Linux 배포판을 지원합니다.
...모든 것을 확장하십시오NVIDIA GPU 호스트 설정
Docker, local-Docker 또는
Kubernetes 백엔드에서 TAO 워크플로가 실행되기 전에 이 설정 절차를 수행하십시오. 이 절차는 다음에서 호스트 GPU 런타임을 표준화합니다:
- NVIDIA 드라이버 브랜치
580(오픈 커널 모듈 권장) - CUDA 툴킷 패키지
cuda-toolkit-13-0 - NVIDIA 컨테이너 툴킷
1.19.0 - Docker 엔진 —
docker/local-docker백엔드용으로만 설치되며, Docker가 설치되어 있지 않은 경우에만 설치됩니다. 선택되는 패키지는 배포판 계열에 따라 다릅니다(기본적으로 Debian 계열에서는docker.io, RHEL 계열에서는download.docker.com의moby-engine/docker-ce, SUSE 계열에서는docker). 설치하지 않으려면--skip-docker-install옵션을 전달하십시오.
이 확인 과정은 기본적으로 안전하며 읽기 전용입니다. 즉, 모든 Linux
배포판에서 작동합니다. 이는 nvidia-smi, CUDA 툴킷 경로,
설치된 container-toolkit 패키지 버전(dpkg/rpm 또는
nvidia-ctk 바이너리 버전을 통해 확인), 그리고 Docker 데몬의 NVIDIA 런타임을 탐색하기만 하기 때문입니다.
설치는 사용자가 명시적으로 승인해야 하며,
--install 옵션을 사용하여 다시 실행해야 합니다. 다음 배포판 계열의 경우 설치 경로가 자동으로 설정됩니다:
| 디스트로 계열 | 테스트된 배포판 | 관리자 | 비고 |
|---|---|---|---|
| debian | Ubuntu 22.04 / 24.04, Debian 12 (및 파생 배포판인 Pop!_OS, Mint, Zorin, Raspbian, KDE Neon 등. UBUNTU_CODENAME / VERSION_CODENAME을 통해) |
apt-get |
NVIDIA cuda-keyring + Container Toolkit .list를 추가합니다. docker.io를 통한 Docker ( $DOCKER_PACKAGE_DEBIAN 재정의). |
| rhel | Fedora 39 이상, RHEL / Rocky / AlmaLinux 9 및 10 | dnf (또는 yum) |
NVIDIA cuda- 및 Container Toolkit .repo를 추가합니다. 가능한 경우 Fedora moby-engine을 통해 Docker를 설치하고, 그렇지 않은 경우 download.docker.com에서 docker-ce를 설치합니다. |
| suse | openSUSE Leap 15, SLES 15 | zypper |
동일한 NVIDIA .repo 파일을 추가합니다. 배포판의 Docker 패키지를 통해 Docker를 사용합니다. |
| 기타 (Arch, Alpine, Gentoo, NixOS, FreeBSD, …) | 해당 없음 | 해당 없음 | --install 명령을 실행하면 버전 대상 및 NVIDIA 설치 가이드 URL이 명시된 명확한 오류 메시지와 함께 종료됩니다. 수동으로 설치한 후 --check-only를 다시 실행하십시오. |
빠른 시작
스킬 뱅크 루트 디렉터리에서:
# 로컬 Docker 백엔드 호스트를 확인합니다.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --check-only
# 사용자의 승인을 받은 후 설치하거나 복구합니다.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --install
# Kubernetes GPU 워커 호스트를 확인합니다.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend kubernetes --check-only
⚠️ 참고 — 비대화형 실행(에이전트/스킬 실행): 스킬 실행 시 터미널이 없으므로 설치 프로그램의
“Continue? [y/N]” 프롬프트에 응답할 수 없습니다.--check-only를실행하여 미리 확인하고 사용자의 승인을 받은 후,--install명령어에 assume-yes 플래그(--yes)를 추가하여 프롬프트 없이 진행되도록 하십시오. 이렇게 하면 시스템 패키지(NVIDIA 드라이버, CUDA Toolkit, NVIDIA Container Toolkit 및 Docker 백엔드용 Docker)의 설치를 자동으로 확인하고 호스트를 수정하므로, 반드시 본인이 관리하는 호스트에서만 이 작업을 수행하십시오. 터미널에서 직접--install을실행하는 사용자에게는 대신 프롬프트가 표시됩니다.
워크플로 계약
Docker 및 Kubernetes 워크플로는 GPU 작업을 제출하기 전에 다음 검사를 실행해야 합니다:
SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT:-$PWD}/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"
bash "$SETUP_SCRIPT" --backend docker --check-only || {
echo "MISSING: TAO GPU 호스트 런타임이 준비되지 않았습니다."
echo "사용자 승인 후 다음 명령을 실행하십시오(비대화형 에이전트 실행 시 --yes를 추가하십시오):"
echo " bash \"$SETUP_SCRIPT\" --backend docker --install"
exit 1
}
절대 무음 모드로 설치하지 마십시오. 확인에 실패하면 누락된 항목을 설명하고, 사용자에게 수정 승인 요청을 한 다음, 설치 명령을 실행하고 확인을 다시 수행하십시오.
설치 프로그램의 작동 방식
설치 프로그램은 감지된 배포판 계열에 따라 작업을 수행합니다. 지원되는 모든
계열에서 NVIDIA의 CUDA 및 Container Toolkit 저장소를
(없는 경우) 추가하고, 고정된 런타임 패키지를 설치하며, 선택적으로
Docker를 설치하고, NVIDIA Docker 런타임을 연결하며, 호출한 사용자를
docker 그룹에 추가합니다.
공통 단계(모든 계열):
- NVIDIA의 CUDA 저장소가 없는 경우 추가합니다(apt의 경우
cuda-keyringdeb, dnf/zypper의 경우cuda-)..repo - NVIDIA의 Container Toolkit 저장소가 없는 경우 추가합니다(apt의 경우
.list, dnf/zypper의 경우.repo). - 실행 중인 커널에 맞는 커널 헤더/개발 패키지를 설치합니다.
- 드라이버 브랜치 580 패키지,
cuda-toolkit-13-0및1.19.0으로고정된 컨테이너 툴킷을 설치합니다(dpkg 접미사가 붙은1.19.0-1은apt용으로 표현된 동일한 업스트림 버전입니다). - Docker 백엔드의 경우, 또는 Docker가 설치되어 있지 않은 경우, Docker를
설치하고(아래의 재정의/제외 플래그 참조), 데몬을 활성화/시작한 다음,
nvidia-ctk runtime configure --runtime=docker를실행하며,systemctl을사용할 수 있는 경우 Docker를 재시작합니다. - 호출한 사용자(사용 가능한 경우
$SUDO_USER, 그렇지 않으면$USER)를docker그룹에 추가하여 이후 셸에서sudo없이docker를실행할 수 있도록 합니다 —--skip-docker-group 옵션을사용하여 이 기능을 비활성화할 수 있습니다. 새로운 그룹 멤버십은 현재 셸에서는 적용되지 않습니다: 로그아웃 후 다시 로그인하거나, 새로운 셸을 열 때마다 `newgrp docker`를실행하십시오. - 재부팅 전에 검증 단계가 통과될 수 있도록
modprobe nvidia를실행합니다.
패밀리별 패키지 선택:
| 단계 | debian 계열 | rhel-family | suse-family |
|---|---|---|---|
| 커널 헤더 | linux-headers-$(uname -r) |
kernel-devel-$(uname -r), kernel-headers-$(uname -r) |
kernel-default-devel |
| 드라이버 | nvidia-driver-pinning-580, nvidia-open-580 (재정의: $NVIDIA_DRIVER_PACKAGE_DEBIAN) |
nvidia-driver-cuda, kmod-nvidia-open-dkms (재정의: $NVIDIA_DRIVER_PACKAGE_RHEL, $NVIDIA_DRIVER_KMOD_RHEL) |
nvidia-open-driver-G06-signed-kmp-default (재정의: $NVIDIA_DRIVER_PACKAGE_SUSE) |
| CUDA 툴킷 | cuda-toolkit-13-0 |
cuda-toolkit-13-0 |
cuda-toolkit-13-0 |
| 컨테이너 툴킷 | nvidia-container-toolkit=1.19.0-1 + base/tools/libs |
nvidia-container-toolkit-1.19.0 + base/tools/libs |
RHEL과 동일 |
| Docker | docker.io (재정의: $DOCKER_PACKAGE_DEBIAN) |
Fedora에서는 가능한 경우moby-engine+moby-cli, 그렇지 않으면 docker-ce docker-ce-cli download.docker.com의 containerd.io |
docker |
검증
설치 후 다음을 확인하십시오:
nvidia-smi
/usr/local/cuda-13.0/bin/nvcc --version
docker info --format '{{json .Runtimes}}' | grep nvidia
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
예상되는 nvidia-smi 출력 결과에는 드라이버 580.x 및 CUDA 버전 13.0이 포함됩니다.
예상되는 nvcc 출력 결과에는 릴리스 13.0이 포함됩니다.
Kubernetes 참고 사항
자체 관리형 쿠버네티스 클러스터의 경우, NVIDIA GPU Operator 또는 장치 플러그인을 설치하기 전에 모든 GPU 워커 노드에서 호스트 설치 프로그램을 실행하거나, 동일한 패키지 세트를 노드 이미지에 미리 포함시켜야 합니다.
kubectl을 사용할 수 있지만 클러스터가
nvidia.com/gpu 할당 가능한 용량이 없다고 보고하는 경우에도 워크플로 검사에서 경고가 표시됩니다. 이 경우,
워커 호스트 런타임이 준비된 후에 NVIDIA GPU Operator를 설치/구성하십시오:
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install --wait gpu-operator -n gpu-operator --create-namespace nvidia/gpu-operator
관리형 쿠버네티스 제공업체는 노드 이미지나 GPU Operator 정책을 통해 드라이버 설치를 직접 관리할 수 있습니다. 사용자의 승인과 롤백 계획이 없는 상태에서 제공업체가 관리하는 GPU 노드를 덮어쓰지 마십시오.
오류 유형
지원되지 않는 배포판 계열: --install 옵션은 debian-, rhel-,
및 suse-family 호스트를 자동으로 처리합니다. Arch, Alpine, Gentoo, NixOS, FreeBSD 또는
/etc/os-release 파일이 없는 환경(예: macOS)에서는 스크립트가 종료되며,
네 가지 버전 대상과 업스트림 NVIDIA 설치 가이드
URL을 나열하는 명확한 오류 메시지가 표시됩니다:
https://docs.nvidia.com/cuda/cuda-installation-guide-linux/https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.htmlhttps://docs.docker.com/engine/install/
사용 중인 배포판의 패키지 관리자를 사용하여 이 네 가지 구성 요소를 설치한 후
--check-only 옵션을 지정하여 스크립트를 다시 실행해 확인해 보십시오. 이 확인 과정은
바이너리 및 패키지 데이터베이스만 조회하므로 보편적으로 호환됩니다. 따라서
런타임이 제대로 설치되어 있다면, 기본 배포판에 관계없이
워크플로 계약 조건이 충족됩니다.
지원되지 않는 Ubuntu/Debian 파생 배포판: ID가 예를 들어 pop, mint,
zorin, raspbian 또는 기타 Debian 계열 파생 배포판인 경우, 스크립트는
UBUNTU_CODENAME /
VERSION_CODENAME을 통해 업스트림 우분투/데비안 CUDA 저장소로 매핑합니다(focal/jammy/noble → 우분투 20.04/22.04/24.04;
bullseye/bookworm/trixie → Debian 11/12/12). 호스트의 코드명이
알려진 업스트림 릴리스와 일치하지 않는 경우, --install은 앞서 설명한 것과 동일한
수동 설치 안내와 함께 종료됩니다.
Docker가 설치되지 않은 경우: --check-only는 MISSING: Docker is not installed를 보고하고, 감지된
디스트로 계열에 적합한 정확한 재실행 명령을 출력합니다. 기본 --install 경로는 Docker(docker.io /
moby-engine / docker-ce / docker )을 설치하고, 데몬을 활성화 및 시작하며,
NVIDIA 런타임을 구성하고, 스크립트를 실행한 사용자를
docker 그룹에 추가합니다. Docker를 직접 관리하고 싶다면, 스크립트를 다시 실행하기 전에
Docker를 설치하거나 --skip-docker-install 옵션을 전달하십시오.
Docker는 설치되었으나 docker run을 실행할 때 여전히 sudo가 필요한 경우: 이 스크립트는
실행 사용자를 docker 그룹에 추가하지만, Linux는 새로운 로그인 세션이 시작될 때만
그룹 멤버십을 갱신합니다. 로그아웃 후 다시 로그인하거나, 새로운 멤버십이 활성화될 때까지
각 새 셸에서newgrp docker를 실행하십시오.
Docker 런타임이 여전히 누락된 경우: Docker를 다시 시작한 다음,
nvidia-ctk runtime configure --runtime=docker를 다시 실행하십시오.
감지된 드라이버 브랜치가 580과 다름: 드라이버 브랜치 고정값은
debian-family(nvidia-open-580)에서 정확히 지정됩니다. rhel-/suse-family에서는 스크립트가
감지된 배포판에 대해 NVIDIA의 CUDA 13.0 저장소에 포함된 최신 오픈 드라이버를
설치하며, 이 버전은 항상 580 이상입니다. 호스트에 더 엄격한
핀 설정이 필요한 경우, --install을 실행하기 전에 $NVIDIA_DRIVER_PACKAGE_RHEL / $NVIDIA_DRIVER_KMOD_RHEL /
$NVIDIA_DRIVER_PACKAGE_SUSE를 원하는 정확한 패키지 이름으로 설정하십시오.
드라이버는 설치되었으나 nvidia-smi가 실패하는 경우:
sudo modprobe nvidia 명령어로 모듈을 로드하거나 시스템을 재부팅하십시오. 시큐어 부트가 활성화된 시스템에서는
MOK 등록이 필요할 수 있습니다.
Kubernetes에 여전히 GPU 용량이 없는 경우: nvidia-smi를 사용하여 각 GPU
노드에서 드라이버가 작동하는지 확인한 다음, GPU Operator/디바이스 플러그인 파드와 노드
라벨을 확인하십시오.
---
name: tao-setup-nvidia-gpu-host
description: Checks and installs NVIDIA driver, CUDA Toolkit, and NVIDIA Container Toolkit for GPU-accelerated Docker and Kubernetes hosts. Supports multiple Linux distributions with automated install and read-only check modes.
license: Apache-2.0
---
# NVIDIA GPU Host Setup
Use this setup skill before TAO workflows run on the `docker`, `local-docker`,
or `kubernetes` backend. It standardizes the host GPU runtime on:
- NVIDIA driver branch `580` (open kernel module preferred)
- CUDA Toolkit package `cuda-toolkit-13-0`
- NVIDIA Container Toolkit `1.19.0`
- Docker engine — only installed for `docker` / `local-docker` backends and
only when Docker is missing. The package picked depends on the distro
family (`docker.io` on Debian-family by default, `moby-engine` /
`docker-ce` from `download.docker.com` on RHEL-family, `docker` on
SUSE-family). Pass `--skip-docker-install` to opt out.
The check is safe and read-only by default — it works on any Linux
distribution because it only probes `nvidia-smi`, the CUDA toolkit path,
the installed container-toolkit package version (via `dpkg`/`rpm`/the
`nvidia-ctk` binary version), and the Docker daemon's NVIDIA runtime.
Installation must be explicitly authorized by the user and rerun with
`--install`. The install path is automated for these distro families:
| Family | Tested distros | Manager | Notes |
|---|---|---|---|
| debian | Ubuntu 22.04 / 24.04, Debian 12 (and derivatives Pop!_OS, Mint, Zorin, Raspbian, KDE Neon, etc. via `UBUNTU_CODENAME` / `VERSION_CODENAME`) | `apt-get` | Adds NVIDIA `cuda-keyring` + Container Toolkit `.list`. Docker via `docker.io` (override `$DOCKER_PACKAGE_DEBIAN`). |
| rhel | Fedora 39+, RHEL / Rocky / AlmaLinux 9 and 10 | `dnf` (or `yum`) | Adds NVIDIA `cuda-<distro>.repo` + Container Toolkit `.repo`. Docker via Fedora `moby-engine` when available, otherwise `docker-ce` from `download.docker.com`. |
| suse | openSUSE Leap 15, SLES 15 | `zypper` | Adds the same NVIDIA `.repo` files. Docker via the distribution `docker` package. |
| other (Arch, Alpine, Gentoo, NixOS, FreeBSD, …) | n/a | n/a | `--install` exits with a clear error listing the version targets and the NVIDIA install-guide URLs. Install manually, then rerun `--check-only`. |
## Quick Start
From the skill bank root:
```bash
# Check the local Docker backend host.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --check-only
# Install or repair after user approval.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --install
# Check a Kubernetes GPU worker host.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend kubernetes --check-only
```
> ⚠️ **Note — running non-interactively (agent/skill runs):** a skill run has no terminal, so the
> installer's `Continue? [y/N]` prompt cannot be answered. After running `--check-only` to preview and
> getting the user's approval, append the assume-yes flag (`--yes`) to the `--install` command so it
> proceeds without a prompt — this auto-confirms installation of system packages (NVIDIA driver, CUDA
> Toolkit, NVIDIA Container Toolkit, and Docker for Docker backends) and modifies the host, so only do
> this on a host you control. A person running `--install` directly at a terminal gets the prompt instead.
## Workflow Contract
Docker and Kubernetes workflows must run the check before submitting GPU work:
```bash
SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT:-$PWD}/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"
bash "$SETUP_SCRIPT" --backend docker --check-only || {
echo "MISSING: TAO GPU host runtime is not ready."
echo "After user approval, run (append --yes for non-interactive agent runs):"
echo " bash \"$SETUP_SCRIPT\" --backend docker --install"
exit 1
}
```
Never install silently. If the check fails, explain what is missing, ask the
user to authorize the fix, then run the install command and rerun the check.
## What The Installer Does
The installer dispatches on the detected distribution family. On every
supported family it adds NVIDIA's CUDA and Container Toolkit repositories
(if missing), installs the pinned runtime packages, optionally installs
Docker, wires the NVIDIA Docker runtime, and adds the invoking user to
the `docker` group.
Common steps (all families):
1. Adds NVIDIA's CUDA repository if missing (apt `cuda-keyring` deb,
`cuda-<distro>.repo` for dnf/zypper).
2. Adds NVIDIA's Container Toolkit repository if missing (`.list` for apt,
`.repo` for dnf/zypper).
3. Installs the matching kernel header / devel package for the running
kernel.
4. Installs the driver branch 580 packages, `cuda-toolkit-13-0`, and the
Container Toolkit pinned to `1.19.0` (the dpkg-suffixed `1.19.0-1` is
the same upstream version expressed for apt).
5. For Docker backends and when Docker is missing, installs Docker
(override / opt-out flags below), enables/starts the daemon, then runs
`nvidia-ctk runtime configure --runtime=docker` and restarts Docker
when `systemctl` is available.
6. Adds the invoking user (`$SUDO_USER` if available, else `$USER`) to the
`docker` group so subsequent shells can run `docker` without `sudo` —
opt out with `--skip-docker-group`. **The new group membership does not
take effect in the current shell**: log out and back in, or run
`newgrp docker` in each new shell.
7. Attempts `modprobe nvidia` so verification can pass before reboot.
Family-specific package selections:
| Step | debian-family | rhel-family | suse-family |
|---|---|---|---|
| Kernel headers | `linux-headers-$(uname -r)` | `kernel-devel-$(uname -r)`, `kernel-headers-$(uname -r)` | `kernel-default-devel` |
| Driver | `nvidia-driver-pinning-580`, `nvidia-open-580` (override: `$NVIDIA_DRIVER_PACKAGE_DEBIAN`) | `nvidia-driver-cuda`, `kmod-nvidia-open-dkms` (override: `$NVIDIA_DRIVER_PACKAGE_RHEL`, `$NVIDIA_DRIVER_KMOD_RHEL`) | `nvidia-open-driver-G06-signed-kmp-default` (override: `$NVIDIA_DRIVER_PACKAGE_SUSE`) |
| CUDA toolkit | `cuda-toolkit-13-0` | `cuda-toolkit-13-0` | `cuda-toolkit-13-0` |
| Container Toolkit | `nvidia-container-toolkit=1.19.0-1` + base/tools/libs | `nvidia-container-toolkit-1.19.0` + base/tools/libs | same as rhel |
| Docker | `docker.io` (override: `$DOCKER_PACKAGE_DEBIAN`) | `moby-engine`+`moby-cli` on Fedora when available, else `docker-ce docker-ce-cli containerd.io` from `download.docker.com` | `docker` |
## Verification
After installation, verify:
```bash
nvidia-smi
/usr/local/cuda-13.0/bin/nvcc --version
docker info --format '{{json .Runtimes}}' | grep nvidia
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
```
Expected `nvidia-smi` output includes driver `580.x` and CUDA Version `13.0`.
Expected `nvcc` output includes `release 13.0`.
## Kubernetes Notes
For self-managed Kubernetes clusters, run the host installer on every GPU
worker node or bake the same package set into the node image before installing
the NVIDIA GPU Operator or device plugin.
The workflow check also warns if `kubectl` is available but the cluster reports
no `nvidia.com/gpu` allocatable capacity. In that case, install/configure the
NVIDIA GPU Operator after the worker host runtime is ready:
```bash
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install --wait gpu-operator -n gpu-operator --create-namespace nvidia/gpu-operator
```
Managed Kubernetes providers may own driver installation through node images or
GPU Operator policy. Do not overwrite a provider-managed GPU node without user
approval and a rollback plan.
## Failure Modes
**Unsupported distribution family**: `--install` automates debian-, rhel-,
and suse-family hosts. On Arch, Alpine, Gentoo, NixOS, FreeBSD, or anything
without `/etc/os-release` (e.g. macOS), the script exits with a clear error
that lists the four version targets and the upstream NVIDIA install-guide
URLs:
- `https://docs.nvidia.com/cuda/cuda-installation-guide-linux/`
- `https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html`
- `https://docs.docker.com/engine/install/`
Install those four pieces using your distribution's package manager and
rerun the script with `--check-only` to verify. The check is universally
portable — it only queries the binaries / package databases — so once the
runtime is in place the workflow contract is satisfied regardless of the
underlying distro.
**Unsupported Ubuntu/Debian derivative**: When `ID` is e.g. `pop`, `mint`,
`zorin`, `raspbian`, or another debian-family derivative, the script maps
the host onto the upstream Ubuntu/Debian CUDA repo via `UBUNTU_CODENAME` /
`VERSION_CODENAME` (`focal`/`jammy`/`noble` → Ubuntu 20.04/22.04/24.04;
`bullseye`/`bookworm`/`trixie` → Debian 11/12/12). If the host's codename
doesn't match a known upstream release, `--install` exits with the same
manual-install guidance described above.
**Docker not installed**: `--check-only` reports `MISSING: Docker is not
installed` and prints the exact rerun command appropriate to the detected
distro family. The default `--install` path installs Docker (`docker.io` /
`moby-engine` / `docker-ce` / `docker` depending on family), enables/starts
the daemon, configures the NVIDIA runtime, and adds the invoking user to
the `docker` group. If you prefer to manage Docker yourself, install it
before rerunning the script or pass `--skip-docker-install`.
**Docker installed but `docker run` still needs sudo**: The script adds the
invoking user to the `docker` group, but Linux only refreshes group
membership on a new login session. Log out and back in, or run
`newgrp docker` in each new shell, until the new membership is active.
**Docker runtime still missing**: Restart Docker, then rerun
`nvidia-ctk runtime configure --runtime=docker`.
**Driver branch detected != 580**: The driver-branch pin is exact on
debian-family (`nvidia-open-580`). On rhel-/suse-family the script
installs the latest open driver shipped in NVIDIA's CUDA 13.0 repo for
the detected distro, which is always ≥ 580. If your host needs a stricter
pin, set `$NVIDIA_DRIVER_PACKAGE_RHEL` / `$NVIDIA_DRIVER_KMOD_RHEL` /
`$NVIDIA_DRIVER_PACKAGE_SUSE` to the exact package names you want before
running `--install`.
**Driver installed but `nvidia-smi` fails**: Load the module with
`sudo modprobe nvidia` or reboot. Secure Boot may require MOK enrollment on
systems where it is enabled.
**Kubernetes still has no GPU capacity**: Confirm the driver works on each GPU
node with `nvidia-smi`, then check the GPU Operator/device plugin pods and node
labels.
tao-setup-nvidia-gpu-host 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-setup-nvidia-gpu-host # Copy SKILL.md to your .claude/skills/ directory
복사





집
