選項
首頁首頁 Skill 開發營運和 CI/CD tao-setup-nvidia-gpu-host

tao-setup-nvidia-gpu-host

NVIDIA/skills NVIDIA/skills

檢查並安裝 NVIDIA 驅動程式、CUDA 工具包以及 NVIDIA 容器工具包,適用於支援 GPU 加速的 Docker 和 Kubernetes 主機。支援多種 Linux 發行版,並提供自動安裝與唯讀檢查模式。

...展開全部
4
更新時間 2026-09-27

NVIDIA GPU 主機設定

在 TAO 工作流程於Docker、本地 Docker 或Kubernetes後端執行之前,請先執行此設定步驟。此步驟將主機 GPU 執行環境標準化為:

  • NVIDIA 驅動程式分支580(建議使用開放式核心模組)
  • CUDA 工具包cuda-toolkit-13-0
  • NVIDIA 容器工具包1.19.0
  • Docker 引擎 — 僅針對docker/local-docker後端安裝,且 僅在缺少 Docker 時安裝。 所選用的套件取決於發行版 家族(預設在 Debian 家族上使用docker.io,在 RHEL 家族上使用來自download.docker.com 的 moby-engine/ docker-ce,在 SUSE 家族上使用docker)。若要跳過安裝,請傳入--skip-docker-install 參數。

此檢查預設為安全且唯讀模式 — 適用於任何 Linux 發行版,因為它僅會檢測nvidia-smi、CUDA 工具包路徑、 已安裝的 container-toolkit 套件版本(透過dpkg/rpm/ nvidia-ctk二進位檔版本),以及 Docker 守護程式的 NVIDIA 執行環境。

安裝必須由使用者明確授權,並使用 --install 選項重新執行。對於以下發行版系列,安裝路徑會自動設定:

發行版家族 已測試的發行版 管理工具 備註
debian Ubuntu 22.04 / 24.04、Debian 12(及其衍生版本 Pop!_OS、Mint、Zorin、Raspbian、KDE Neon 等,透過UBUNTU_CODENAME/VERSION_CODENAME 指定) apt-get 新增 NVIDIAcuda-keyring及 Container Toolkit 的.list 檔案。透過docker.io安裝 Docker(覆寫$DOCKER_PACKAGE_DEBIAN)。
rhel Fedora 39 以上、RHEL / Rocky / AlmaLinux 9 及 10 dnf(或yum) 新增 NVIDIAcuda-.repo及 Container Toolkit.repo。若 Fedoramoby-engine可用,則透過其安裝 Docker;否則從download.docker.com 安裝docker-ce。
suse openSUSE Leap 15、SLES 15 zypper 新增相同的 NVIDIA.repo檔案。透過發行版的Docker套件使用 Docker。
其他(Arch、Alpine、Gentoo、NixOS、FreeBSD 等) 不適用 不適用 --install會以明確的錯誤訊息結束執行,其中列出版本目標及 NVIDIA 安裝指南的網址。請手動安裝,然後重新執行--check-only。

快速入門

從技能庫根目錄開始:

# 檢查本機 Docker 後端主機。
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --check-only

# 經使用者核准後進行安裝或修復。
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --install

# 檢查 Kubernetes GPU 工作節點主機。
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend kubernetes --check-only

⚠️注意 — 非互動式執行(代理程式/技能執行):技能執行時沒有終端機,因此 安裝程式提出的「繼續嗎?[y/N]」提示無法回應。 在執行--check-only進行預覽並 獲得使用者批准後,請在--install指令後附加 assume-yes 標誌 (--yes),使其 無需提示即可繼續執行 — 此操作會自動確認安裝系統套件(NVIDIA 驅動程式、CUDA 工具包、NVIDIA 容器工具包,以及用於 Docker 後端的 Docker)並修改主機,因此請僅在 您所控制的主機上執行此操作。若在終端機上直接執行--install指令,則會顯示該提示。

工作流程合約

Docker 和 Kubernetes 工作流程在提交 GPU 工作前,必須執行以下檢查:

SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT:-$PWD}/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"

bash "$SETUP_SCRIPT" --backend docker --check-only || {
  echo "缺失:TAO GPU 主機執行環境尚未就緒。"
  echo "經使用者批准後,請執行(若為非互動式代理程式執行,請在後面加上 --yes):"
  echo "  bash \"$SETUP_SCRIPT\" --backend docker --install"
  exit 1
}

切勿以靜默模式安裝。若檢查失敗,請說明缺少哪些項目,並請 使用者授權進行修復,然後執行安裝命令並重新執行檢查。

安裝程式的工作內容

安裝程式會根據偵測到的發行版家族進行部署。在每個 受支援的家族中,它會新增 NVIDIA 的 CUDA 和 Container Toolkit 儲存庫 (若尚未存在),安裝固定版本的執行環境套件,可選地安裝 Docker,配置 NVIDIA Docker 執行環境,並將呼叫使用者加入 docker群組。

共通步驟(所有系列):

  1. 若 NVIDIA 的 CUDA 儲存庫缺失,則新增該儲存庫(對於 apt 為cuda-keyringdeb, 對於 dnf/zypper 為cuda-.repo)。
  2. 若未存在,則新增 NVIDIA 的 Container Toolkit 儲存庫(apt 格式為.list, dnf/zypper 格式為.repo)。
  3. 安裝與執行中 核心相容的核心標頭/開發套件。
  4. 安裝驅動程式分支 580 的套件、cuda-toolkit-13-0,以及 版本固定為1.19.0的 Container Toolkit(帶有 dpkg 後綴的1.19.0-1 即為 apt 格式所表示的相同上游版本)。
  5. 針對 Docker 後端,以及當系統未安裝 Docker 時,會安裝 Docker (可透過下方標誌覆寫或選擇不安裝),啟用並啟動 daemon,接著執行 nvidia-ctk runtime configure --runtime=docker,並在systemctl可用時重新啟動 Docker。
  6. 將呼叫使用者(若可用則為$SUDO_USER,否則為$USER)加入 docker群組,以便後續的 shell 能不需sudo即可執行docker— 若要跳過此步驟,請使用--skip-docker-group。新的群組成員資格不會 在當前 shell 中生效:請登出並重新登入,或在每個新 shell 中執行 newgrp docker。
  7. 嘗試執行modprobe nvidia,以便在重新開機前通過驗證。

針對特定處理器家族的套件選項:

步驟 debian-family rhel-family suse-家族
核心標頭 linux-headers-$(uname -r) kernel-devel-$(uname -r),kernel-headers-$(uname -r) 核心預設開發套件
驅動程式 nvidia-driver-pinning-580,nvidia-open-580(覆寫:$NVIDIA_DRIVER_PACKAGE_DEBIAN) nvidia-driver-cuda,kmod-nvidia-open-dkms(覆寫:$NVIDIA_DRIVER_PACKAGE_RHEL,$NVIDIA_DRIVER_KMOD_RHEL) nvidia-open-driver-G06-signed-kmp-default(覆寫:$NVIDIA_DRIVER_PACKAGE_SUSE)
CUDA 工具包 cuda-toolkit-13-0 cuda-toolkit-13-0 cuda-toolkit-13-0
容器工具包 nvidia-container-toolkit=1.19.0-1+ base/tools/libs nvidia-container-toolkit-1.19.0+ base/tools/libs 與 RHEL 相同
Docker docker.io(覆寫:$DOCKER_PACKAGE_DEBIAN) 若 Fedora 系統中可用,則使用moby-engine+moby-cli;否則使用docker-ce、docker-ce-cli及來自download.docker.com的containerd.io docker

驗證

安裝完成後,請驗證:

nvidia-smi
/usr/local/cuda-13.0/bin/nvcc --version
docker info --format '{{json .Runtimes}}' | grep nvidia
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi

預期nvidia-smi的輸出應包含驅動程式580.x及 CUDA 版本13.0。 預期nvcc的輸出應包含版本 13.0。

Kubernetes 注意事項

對於自主管理的 Kubernetes 叢集,請在每個 GPU 工作節點上執行主機安裝程式,或在安裝 NVIDIA GPU Operator 或裝置外掛程式之前,將相同的套件集整合至節點映像檔中。

若kubectl可用,但叢集報告 無可分配的nvidia.com/gpu容量,工作流程檢查也會發出警告。在此情況下,請於 工作節點主機執行環境就緒後,安裝並設定 NVIDIA GPU Operator:

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install --wait gpu-operator -n gpu-operator --create-namespace nvidia/gpu-operator

受管 Kubernetes 提供者可能會透過節點映像或 GPU Operator 政策來管理驅動程式安裝。在未經使用者 批准且未制定回滾計畫的情況下,請勿覆寫由提供者管理的 GPU 節點。

故障模式

不支援的發行版家族:--install會自動處理 debian-、rhel- 及 suse- 家族的主機。 在 Arch、Alpine、Gentoo、NixOS、FreeBSD 或任何 未包含/etc/os-release的系統(例如 macOS)上,腳本會以明確的錯誤訊息退出, 並列出四個版本目標以及上游 NVIDIA 安裝指南 的網址:

  • https://docs.nvidia.com/cuda/cuda-installation-guide-linux/
  • https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html
  • https://docs.docker.com/engine/install/

請使用您的發行版套件管理員安裝這四個元件,並 使用--check-only 參數重新執行腳本以進行驗證。此檢查具有普遍 的可移植性——它僅查詢二進位檔/套件資料庫——因此一旦 執行時環境就位,無論底層 發行版為何,工作流程契約即視為滿足。

不受支援的 Ubuntu/Debian 衍生版本:當ID例如為pop、mint、 zorin、raspbian 或其他 Debian 家族衍生版本時,該腳本會透過UBUNTU_CODENAME/ VERSION_CODENAME(focal/jammy/noble→ Ubuntu 20.04/22.04/24.04; bullseye/bookworm/trixie→ Debian 11/12/12)。若主機的代號 與已知的上游發行版不符,--install會以上述相同的 手動安裝指引結束執行。

未安裝 Docker:--check-only會回報「MISSING: Docker 未安裝」,並輸出適用於偵測到的 發行版家族的精確重新執行指令。預設的--install路徑會安裝 Docker(docker.io/ moby-engine/docker-ce/docker,視發行版家族而定),啟用並啟動 守護程序,配置 NVIDIA 執行環境,並將呼叫使用者加入 docker群組。 若您希望自行管理 Docker,請在 重新執行腳本前先安裝 Docker,或傳入--skip-docker-install 參數。

已安裝 Docker 但執行 `docker run`仍需 `sudo`:此腳本會將 呼叫使用者加入 `docker` 群組,但 Linux 僅會在新的登入 會話中更新群組成員資格。請登出並重新登入,或在每個新的 終端機中執行`newgrp docker`,直到新的群組成員資格生效為止。

仍缺少 Docker 執行環境:重新啟動 Docker,然後重新執行 nvidia-ctk runtime configure --runtime=docker。

偵測到的驅動程式分支 ≠ 580:在 debian-family 系統上,驅動程式分支的固定版本為 nvidia-open-580。 在 rhel-/suse-family 上,該腳本 會為偵測到的發行版安裝 NVIDIA CUDA 13.0 儲存庫中 提供的最新開放式驅動程式,其版本始終 ≥ 580。 若您的主機需要更嚴格的 版本鎖定,請在執行 --install 之前,將$NVIDIA_DRIVER_PACKAGE_RHEL/$NVIDIA_DRIVER_KMOD_RHEL/ $NVIDIA_DRIVER_PACKAGE_SUSE設定為您所需的精確套件名稱。

驅動程式已安裝但nvidia-smi執行失敗:請使用 sudo modprobe nvidia載入模組,或重新開機。若系統已啟用安全開機 (Secure Boot), 可能需要進行 MOK 註冊。

Kubernetes 仍無 GPU 容量:請使用nvidia-smi 確認每個 GPU 節點上的驅動程式是否正常運作,然後檢查 GPU Operator/裝置外掛程式 Pod 以及節點 標籤。

在 GitHub 上查看
---
name: tao-setup-nvidia-gpu-host
description: Checks and installs NVIDIA driver, CUDA Toolkit, and NVIDIA Container Toolkit for GPU-accelerated Docker and Kubernetes hosts. Supports multiple Linux distributions with automated install and read-only check modes.
license: Apache-2.0
---

# NVIDIA GPU Host Setup

Use this setup skill before TAO workflows run on the `docker`, `local-docker`,
or `kubernetes` backend. It standardizes the host GPU runtime on:

- NVIDIA driver branch `580` (open kernel module preferred)
- CUDA Toolkit package `cuda-toolkit-13-0`
- NVIDIA Container Toolkit `1.19.0`
- Docker engine — only installed for `docker` / `local-docker` backends and
  only when Docker is missing. The package picked depends on the distro
  family (`docker.io` on Debian-family by default, `moby-engine` /
  `docker-ce` from `download.docker.com` on RHEL-family, `docker` on
  SUSE-family). Pass `--skip-docker-install` to opt out.

The check is safe and read-only by default — it works on any Linux
distribution because it only probes `nvidia-smi`, the CUDA toolkit path,
the installed container-toolkit package version (via `dpkg`/`rpm`/the
`nvidia-ctk` binary version), and the Docker daemon's NVIDIA runtime.

Installation must be explicitly authorized by the user and rerun with
`--install`. The install path is automated for these distro families:

| Family | Tested distros | Manager | Notes |
|---|---|---|---|
| debian | Ubuntu 22.04 / 24.04, Debian 12 (and derivatives Pop!_OS, Mint, Zorin, Raspbian, KDE Neon, etc. via `UBUNTU_CODENAME` / `VERSION_CODENAME`) | `apt-get` | Adds NVIDIA `cuda-keyring` + Container Toolkit `.list`. Docker via `docker.io` (override `$DOCKER_PACKAGE_DEBIAN`). |
| rhel | Fedora 39+, RHEL / Rocky / AlmaLinux 9 and 10 | `dnf` (or `yum`) | Adds NVIDIA `cuda-<distro>.repo` + Container Toolkit `.repo`. Docker via Fedora `moby-engine` when available, otherwise `docker-ce` from `download.docker.com`. |
| suse | openSUSE Leap 15, SLES 15 | `zypper` | Adds the same NVIDIA `.repo` files. Docker via the distribution `docker` package. |
| other (Arch, Alpine, Gentoo, NixOS, FreeBSD, …) | n/a | n/a | `--install` exits with a clear error listing the version targets and the NVIDIA install-guide URLs. Install manually, then rerun `--check-only`. |

## Quick Start

From the skill bank root:

```bash
# Check the local Docker backend host.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --check-only

# Install or repair after user approval.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --install

# Check a Kubernetes GPU worker host.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend kubernetes --check-only
```

> ⚠️ **Note — running non-interactively (agent/skill runs):** a skill run has no terminal, so the
> installer's `Continue? [y/N]` prompt cannot be answered. After running `--check-only` to preview and
> getting the user's approval, append the assume-yes flag (`--yes`) to the `--install` command so it
> proceeds without a prompt — this auto-confirms installation of system packages (NVIDIA driver, CUDA
> Toolkit, NVIDIA Container Toolkit, and Docker for Docker backends) and modifies the host, so only do
> this on a host you control. A person running `--install` directly at a terminal gets the prompt instead.

## Workflow Contract

Docker and Kubernetes workflows must run the check before submitting GPU work:

```bash
SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT:-$PWD}/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"

bash "$SETUP_SCRIPT" --backend docker --check-only || {
  echo "MISSING: TAO GPU host runtime is not ready."
  echo "After user approval, run (append --yes for non-interactive agent runs):"
  echo "  bash \"$SETUP_SCRIPT\" --backend docker --install"
  exit 1
}
```

Never install silently. If the check fails, explain what is missing, ask the
user to authorize the fix, then run the install command and rerun the check.

## What The Installer Does

The installer dispatches on the detected distribution family. On every
supported family it adds NVIDIA's CUDA and Container Toolkit repositories
(if missing), installs the pinned runtime packages, optionally installs
Docker, wires the NVIDIA Docker runtime, and adds the invoking user to
the `docker` group.

Common steps (all families):

1. Adds NVIDIA's CUDA repository if missing (apt `cuda-keyring` deb,
   `cuda-<distro>.repo` for dnf/zypper).
2. Adds NVIDIA's Container Toolkit repository if missing (`.list` for apt,
   `.repo` for dnf/zypper).
3. Installs the matching kernel header / devel package for the running
   kernel.
4. Installs the driver branch 580 packages, `cuda-toolkit-13-0`, and the
   Container Toolkit pinned to `1.19.0` (the dpkg-suffixed `1.19.0-1` is
   the same upstream version expressed for apt).
5. For Docker backends and when Docker is missing, installs Docker
   (override / opt-out flags below), enables/starts the daemon, then runs
   `nvidia-ctk runtime configure --runtime=docker` and restarts Docker
   when `systemctl` is available.
6. Adds the invoking user (`$SUDO_USER` if available, else `$USER`) to the
   `docker` group so subsequent shells can run `docker` without `sudo` —
   opt out with `--skip-docker-group`. **The new group membership does not
   take effect in the current shell**: log out and back in, or run
   `newgrp docker` in each new shell.
7. Attempts `modprobe nvidia` so verification can pass before reboot.

Family-specific package selections:

| Step | debian-family | rhel-family | suse-family |
|---|---|---|---|
| Kernel headers | `linux-headers-$(uname -r)` | `kernel-devel-$(uname -r)`, `kernel-headers-$(uname -r)` | `kernel-default-devel` |
| Driver | `nvidia-driver-pinning-580`, `nvidia-open-580` (override: `$NVIDIA_DRIVER_PACKAGE_DEBIAN`) | `nvidia-driver-cuda`, `kmod-nvidia-open-dkms` (override: `$NVIDIA_DRIVER_PACKAGE_RHEL`, `$NVIDIA_DRIVER_KMOD_RHEL`) | `nvidia-open-driver-G06-signed-kmp-default` (override: `$NVIDIA_DRIVER_PACKAGE_SUSE`) |
| CUDA toolkit | `cuda-toolkit-13-0` | `cuda-toolkit-13-0` | `cuda-toolkit-13-0` |
| Container Toolkit | `nvidia-container-toolkit=1.19.0-1` + base/tools/libs | `nvidia-container-toolkit-1.19.0` + base/tools/libs | same as rhel |
| Docker | `docker.io` (override: `$DOCKER_PACKAGE_DEBIAN`) | `moby-engine`+`moby-cli` on Fedora when available, else `docker-ce docker-ce-cli containerd.io` from `download.docker.com` | `docker` |

## Verification

After installation, verify:

```bash
nvidia-smi
/usr/local/cuda-13.0/bin/nvcc --version
docker info --format '{{json .Runtimes}}' | grep nvidia
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
```

Expected `nvidia-smi` output includes driver `580.x` and CUDA Version `13.0`.
Expected `nvcc` output includes `release 13.0`.

## Kubernetes Notes

For self-managed Kubernetes clusters, run the host installer on every GPU
worker node or bake the same package set into the node image before installing
the NVIDIA GPU Operator or device plugin.

The workflow check also warns if `kubectl` is available but the cluster reports
no `nvidia.com/gpu` allocatable capacity. In that case, install/configure the
NVIDIA GPU Operator after the worker host runtime is ready:

```bash
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install --wait gpu-operator -n gpu-operator --create-namespace nvidia/gpu-operator
```

Managed Kubernetes providers may own driver installation through node images or
GPU Operator policy. Do not overwrite a provider-managed GPU node without user
approval and a rollback plan.

## Failure Modes

**Unsupported distribution family**: `--install` automates debian-, rhel-,
and suse-family hosts. On Arch, Alpine, Gentoo, NixOS, FreeBSD, or anything
without `/etc/os-release` (e.g. macOS), the script exits with a clear error
that lists the four version targets and the upstream NVIDIA install-guide
URLs:

- `https://docs.nvidia.com/cuda/cuda-installation-guide-linux/`
- `https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html`
- `https://docs.docker.com/engine/install/`

Install those four pieces using your distribution's package manager and
rerun the script with `--check-only` to verify. The check is universally
portable — it only queries the binaries / package databases — so once the
runtime is in place the workflow contract is satisfied regardless of the
underlying distro.

**Unsupported Ubuntu/Debian derivative**: When `ID` is e.g. `pop`, `mint`,
`zorin`, `raspbian`, or another debian-family derivative, the script maps
the host onto the upstream Ubuntu/Debian CUDA repo via `UBUNTU_CODENAME` /
`VERSION_CODENAME` (`focal`/`jammy`/`noble` → Ubuntu 20.04/22.04/24.04;
`bullseye`/`bookworm`/`trixie` → Debian 11/12/12). If the host's codename
doesn't match a known upstream release, `--install` exits with the same
manual-install guidance described above.

**Docker not installed**: `--check-only` reports `MISSING: Docker is not
installed` and prints the exact rerun command appropriate to the detected
distro family. The default `--install` path installs Docker (`docker.io` /
`moby-engine` / `docker-ce` / `docker` depending on family), enables/starts
the daemon, configures the NVIDIA runtime, and adds the invoking user to
the `docker` group. If you prefer to manage Docker yourself, install it
before rerunning the script or pass `--skip-docker-install`.

**Docker installed but `docker run` still needs sudo**: The script adds the
invoking user to the `docker` group, but Linux only refreshes group
membership on a new login session. Log out and back in, or run
`newgrp docker` in each new shell, until the new membership is active.

**Docker runtime still missing**: Restart Docker, then rerun
`nvidia-ctk runtime configure --runtime=docker`.

**Driver branch detected != 580**: The driver-branch pin is exact on
debian-family (`nvidia-open-580`). On rhel-/suse-family the script
installs the latest open driver shipped in NVIDIA's CUDA 13.0 repo for
the detected distro, which is always ≥ 580. If your host needs a stricter
pin, set `$NVIDIA_DRIVER_PACKAGE_RHEL` / `$NVIDIA_DRIVER_KMOD_RHEL` /
`$NVIDIA_DRIVER_PACKAGE_SUSE` to the exact package names you want before
running `--install`.

**Driver installed but `nvidia-smi` fails**: Load the module with
`sudo modprobe nvidia` or reboot. Secure Boot may require MOK enrollment on
systems where it is enabled.

**Kubernetes still has no GPU capacity**: Confirm the driver works on each GPU
node with `nvidia-smi`, then check the GPU Operator/device plugin pods and node
labels.

安裝 tao-setup-nvidia-gpu-host

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-setup-nvidia-gpu-host # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

klingai-upgrade-migration
更新時間 2026-07-03
Verification &amp; Quality Assurance
更新時間 2026-06-29
base44-cli
更新時間 2026-06-29
Railway CLI Management
更新時間 2026-07-02
OR