选项
首页首页 Skill 开发运营和 CI/CD tao-setup-nvidia-gpu-host

tao-setup-nvidia-gpu-host

NVIDIA/skills NVIDIA/skills

检查并安装 NVIDIA 驱动程序、CUDA 工具包以及 NVIDIA 容器工具包,用于支持 GPU 加速的 Docker 和 Kubernetes 主机。支持多种 Linux 发行版,并提供自动安装和只读检查模式。

...展开全部
4
更新时间 2026-09-27

NVIDIA GPU 主机配置

在 TAO 工作流于Docker、本地 Docker 或Kubernetes后端上运行之前,请使用此配置技能。它将主机 GPU 运行时标准化为:

  • NVIDIA 驱动分支580(建议使用开放内核模块)
  • CUDA 工具包cuda-toolkit-13-0
  • NVIDIA 容器工具包1.19.0
  • Docker 引擎 — 仅针对docker/local-docker后端安装,且 仅在缺少 Docker 时安装。 所选软件包取决于发行版 系列(Debian 系列默认使用docker.io,RHEL 系列使用来自download.docker.com 的 moby-engine/ docker-ce,SUSE 系列使用docker)。若要跳过安装,请传递--skip-docker-install 参数。

该检查默认是安全的且仅读取数据——它适用于任何 Linux 发行版,因为它仅检测nvidia-smi、CUDA 工具包路径、 已安装的 container-toolkit 软件包版本(通过dpkg/rpm 或 nvidia-ctk二进制文件版本),以及 Docker 守护进程的 NVIDIA 运行时。

安装必须由用户明确授权,并使用 --install 选项重新运行。以下发行版家族的安装路径已实现自动化:

发行版系列 已测试的发行版 管理器 备注
debian Ubuntu 22.04 / 24.04、Debian 12(及其衍生版本 Pop!_OS、Mint、Zorin、Raspbian、KDE Neon 等,可通过UBUNTU_CODENAME/VERSION_CODENAME 实现) apt-get 添加 NVIDIAcuda-keyring以及 Container Toolkit 的.list 文件。通过docker.io安装 Docker(覆盖$DOCKER_PACKAGE_DEBIAN)。
rhel Fedora 39 及以上版本、RHEL / Rocky / AlmaLinux 9 和 10 dnf(或yum) 添加 NVIDIAcuda-.repo及 Container Toolkit.repo。若 Fedoramoby-engine可用,则通过其安装 Docker;否则从download.docker.com 下载 docker-ce进行安装。
suse openSUSE Leap 15、SLES 15 zypper 添加相同的 NVIDIA.repo文件。通过发行版的Docker软件包安装 Docker。
其他(Arch、Alpine、Gentoo、NixOS、FreeBSD 等) 不适用 不适用 --install命令会以明确的错误信息退出,其中列出了目标版本和 NVIDIA 安装指南的 URL。请手动安装,然后重新运行--check-only 命令。

快速入门

从技能库根目录开始:

# 检查本地 Docker 后端主机。
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --check-only

# 经用户批准后进行安装或修复。
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --install

# 检查 Kubernetes GPU 工作节点。
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend kubernetes --check-only

⚠️注意 — 非交互式运行(代理/技能运行):技能运行时没有终端,因此无法 对安装程序的“继续?[y/N]”提示进行响应。 在运行--check-only进行预览并 获得用户批准后,请在--install命令后添加 assume-yes 标志(--yes),以便其 不经提示直接继续 —— 这将自动确认安装系统软件包(NVIDIA 驱动程序、CUDA 工具包、NVIDIA 容器工具包以及用于 Docker 后端的 Docker)并修改主机,因此请仅在 您控制的主机上执行此操作。在终端上直接运行--install命令的用户则会看到该提示。

工作流契约

Docker 和 Kubernetes 工作流在提交 GPU 任务前必须运行以下检查:

SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT:-$PWD}/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"

bash "$SETUP_SCRIPT" --backend docker --check-only || {
  echo "缺失:TAO GPU 主机运行时尚未就绪。"
  echo "用户批准后,请运行(若为非交互式代理运行,请在命令后添加 --yes):"
  echo "  bash \"$SETUP_SCRIPT\" --backend docker --install"
  exit 1
}

切勿进行静默安装。如果检查失败,请说明缺失的内容,请求 用户授权修复,然后运行安装命令并重新执行检查。

安装程序的功能

安装程序会根据检测到的发行版家族进行部署。对于每个 受支持的家族,它会添加 NVIDIA 的 CUDA 和容器工具包软件源 (若缺失),安装固定的运行时软件包,可选地安装 Docker,配置 NVIDIA Docker 运行时,并将调用该程序的用户添加到 docker组中。

通用步骤(所有系列):

  1. 若缺失,则添加 NVIDIA 的 CUDA 软件源(对于 apt,使用cuda-keyringdeb; 对于 dnf/zypper,使用cuda-.repo)。
  2. 若未添加,则添加 NVIDIA 的容器工具包软件源(apt 格式为.list, dnf/zypper 格式为.repo)。
  3. 安装与运行中内核 匹配的内核头文件/开发包。
  4. 安装驱动程序分支 580 的软件包、cuda-toolkit-13-0 以及 固定版本为1.19.0的 Container Toolkit(后缀为 dpkg 的1.19.0-1是 针对 apt 表达的相同上游版本)。
  5. 对于 Docker 后端以及当 Docker 缺失时,会安装 Docker (可通过下文的标志进行覆盖或选择不安装),启用/启动守护进程,然后运行 nvidia-ctk runtime configure --runtime=docker,并在系统支持systemctl时重启 Docker。
  6. 将调用用户(如有则为$SUDO_USER,否则为$USER)添加到 docker组,以便后续 shell 可以不使用sudo运行docker— 使用--skip-docker-group 可选择不执行此操作。新的组成员身份不会 在当前 shell 中生效:请注销并重新登录,或在每个新 shell 中运行 newgrp docker。
  7. 尝试执行modprobe nvidia,以便在重启前通过验证。

针对特定系列的软件包选择:

步骤 debian-family rhel-family suse-family
内核头文件 linux-headers-$(uname -r) kernel-devel-$(uname -r),kernel-headers-$(uname -r) 内核默认开发包
驱动程序 nvidia-driver-pinning-580,nvidia-open-580(覆盖:$NVIDIA_DRIVER_PACKAGE_DEBIAN) nvidia-driver-cuda、kmod-nvidia-open-dkms(覆盖:$NVIDIA_DRIVER_PACKAGE_RHEL、$NVIDIA_DRIVER_KMOD_RHEL) nvidia-open-driver-G06-signed-kmp-default(覆盖:$NVIDIA_DRIVER_PACKAGE_SUSE)
CUDA 工具包 cuda-toolkit-13-0 cuda-toolkit-13-0 cuda-toolkit-13-0
容器工具包 nvidia-container-toolkit=1.19.0-1+ base/tools/libs nvidia-container-toolkit-1.19.0+ base/tools/libs 与 RHEL 相同
Docker docker.io(覆盖:$DOCKER_PACKAGE_DEBIAN) 在 Fedora 上若可用则使用moby-engine+moby-cli,否则使用docker-ce、docker-ce-cli及从download.docker.com获取的containerd.io docker

验证

安装完成后,请验证:

nvidia-smi
/usr/local/cuda-13.0/bin/nvcc --version
docker info --format '{{json .Runtimes}}' | grep nvidia
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi

预期nvidia-smi的输出应包含驱动程序580.x和 CUDA 版本13.0。 预期nvcc 的输出应包含版本 13.0。

Kubernetes 注意事项

对于自管式 Kubernetes 集群,请在每个 GPU 工作节点上运行主机安装程序,或在安装 NVIDIA GPU Operator 或设备插件之前,将同一套软件包预集成到节点镜像中。

如果kubectl可用,但集群报告 没有可分配的nvidia.com/gpu容量,工作流检查也会发出警告。在这种情况下,请在 工作节点主机运行时准备就绪后安装/配置 NVIDIA GPU Operator:

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install --wait gpu-operator -n gpu-operator --create-namespace nvidia/gpu-operator

托管 Kubernetes 提供商可能会通过节点镜像或 GPU Operator 策略来管理驱动程序的安装。未经用户 批准且未制定回滚计划,请勿覆盖由提供商管理的 GPU 节点。

故障模式

不支持的发行版家族:--install选项会自动处理 debian-、rhel- 和 suse- 家族的主机。 在 Arch、Alpine、Gentoo、NixOS、FreeBSD 或任何 没有/etc/os-release的系统(例如 macOS)上,脚本会以明确的错误信息退出, 该信息会列出四个版本目标以及上游 NVIDIA 安装指南的 URL:

  • https://docs.nvidia.com/cuda/cuda-installation-guide-linux/
  • https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html
  • https://docs.docker.com/engine/install/

请使用您所用发行版的包管理器安装这四个组件, 并使用--check-only 参数重新运行脚本进行验证。该检查具有通用的 可移植性——它仅查询二进制文件/包数据库——因此一旦 运行时环境就位,无论底层 发行版为何,工作流契约均视为满足。

不受支持的 Ubuntu/Debian 衍生版本:当ID例如为pop、mint、 zorin、raspbian 或其他 Debian 家族衍生版本时,脚本会通过UBUNTU_CODENAME/ VERSION_CODENAME将主机映射到上游的 Ubuntu/Debian CUDA 仓库(focal/jammy/noble→ Ubuntu 20.04/22.04/24.04; bullseye/bookworm/trixie→ Debian 11/12/12)。如果主机的代号 与已知的上游版本不匹配,--install选项将退出并显示与 上述相同的手动安装指南。

未安装 Docker:--check-only会报告“MISSING: Docker 未安装”,并输出适用于检测到的 发行版家族的具体重跑命令。默认的--install路径会安装 Docker(docker.io/ moby-engine/docker-ce/docker,具体取决于发行版家族),启用并启动 守护进程,配置 NVIDIA 运行时,并将调用用户添加到 docker组中。 如果您希望自行管理 Docker,请在 重新运行脚本前先安装 Docker,或传递--skip-docker-install 参数。

Docker 已安装但`docker run`仍需 `sudo`:该脚本会将 调用用户添加到`docker` 组中,但 Linux 仅在新的登录会话中 刷新组成员资格。请注销并重新登录,或在每个新终端中运行 `newgrp docker`,直到新成员资格生效。

仍缺少 Docker 运行时:重启 Docker,然后重新运行 nvidia-ctk runtime configure --runtime=docker。

检测到的驱动程序分支 ≠ 580:在 debian-family 系统上,驱动程序分支的固定版本为 nvidia-open-580。 在 rhel-/suse-family 上,该脚本 会为检测到的发行版安装 NVIDIA CUDA 13.0 仓库中提供的 最新开源驱动,其版本始终≥580。 如果您的主机需要更严格的 固定版本,请在 运行--install 之前,将$NVIDIA_DRIVER_PACKAGE_RHEL/$NVIDIA_DRIVER_KMOD_RHEL/ $NVIDIA_DRIVER_PACKAGE_SUSE设置为您所需的精确软件包名称。

驱动程序已安装但nvidia-smi运行失败:请使用 sudo modprobe nvidia加载模块,或重启系统。在启用了安全启动(Secure Boot)的 系统上,可能需要进行 MOK 注册。

Kubernetes 仍无 GPU 算力:使用nvidia-smi 确认每个 GPU 节点上的驱动程序是否正常工作,然后检查 GPU Operator/设备插件的 Pod 及节点 标签。

在 GitHub 上查看
---
name: tao-setup-nvidia-gpu-host
description: Checks and installs NVIDIA driver, CUDA Toolkit, and NVIDIA Container Toolkit for GPU-accelerated Docker and Kubernetes hosts. Supports multiple Linux distributions with automated install and read-only check modes.
license: Apache-2.0
---

# NVIDIA GPU Host Setup

Use this setup skill before TAO workflows run on the `docker`, `local-docker`,
or `kubernetes` backend. It standardizes the host GPU runtime on:

- NVIDIA driver branch `580` (open kernel module preferred)
- CUDA Toolkit package `cuda-toolkit-13-0`
- NVIDIA Container Toolkit `1.19.0`
- Docker engine — only installed for `docker` / `local-docker` backends and
  only when Docker is missing. The package picked depends on the distro
  family (`docker.io` on Debian-family by default, `moby-engine` /
  `docker-ce` from `download.docker.com` on RHEL-family, `docker` on
  SUSE-family). Pass `--skip-docker-install` to opt out.

The check is safe and read-only by default — it works on any Linux
distribution because it only probes `nvidia-smi`, the CUDA toolkit path,
the installed container-toolkit package version (via `dpkg`/`rpm`/the
`nvidia-ctk` binary version), and the Docker daemon's NVIDIA runtime.

Installation must be explicitly authorized by the user and rerun with
`--install`. The install path is automated for these distro families:

| Family | Tested distros | Manager | Notes |
|---|---|---|---|
| debian | Ubuntu 22.04 / 24.04, Debian 12 (and derivatives Pop!_OS, Mint, Zorin, Raspbian, KDE Neon, etc. via `UBUNTU_CODENAME` / `VERSION_CODENAME`) | `apt-get` | Adds NVIDIA `cuda-keyring` + Container Toolkit `.list`. Docker via `docker.io` (override `$DOCKER_PACKAGE_DEBIAN`). |
| rhel | Fedora 39+, RHEL / Rocky / AlmaLinux 9 and 10 | `dnf` (or `yum`) | Adds NVIDIA `cuda-<distro>.repo` + Container Toolkit `.repo`. Docker via Fedora `moby-engine` when available, otherwise `docker-ce` from `download.docker.com`. |
| suse | openSUSE Leap 15, SLES 15 | `zypper` | Adds the same NVIDIA `.repo` files. Docker via the distribution `docker` package. |
| other (Arch, Alpine, Gentoo, NixOS, FreeBSD, …) | n/a | n/a | `--install` exits with a clear error listing the version targets and the NVIDIA install-guide URLs. Install manually, then rerun `--check-only`. |

## Quick Start

From the skill bank root:

```bash
# Check the local Docker backend host.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --check-only

# Install or repair after user approval.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend docker --install

# Check a Kubernetes GPU worker host.
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend kubernetes --check-only
```

> ⚠️ **Note — running non-interactively (agent/skill runs):** a skill run has no terminal, so the
> installer's `Continue? [y/N]` prompt cannot be answered. After running `--check-only` to preview and
> getting the user's approval, append the assume-yes flag (`--yes`) to the `--install` command so it
> proceeds without a prompt — this auto-confirms installation of system packages (NVIDIA driver, CUDA
> Toolkit, NVIDIA Container Toolkit, and Docker for Docker backends) and modifies the host, so only do
> this on a host you control. A person running `--install` directly at a terminal gets the prompt instead.

## Workflow Contract

Docker and Kubernetes workflows must run the check before submitting GPU work:

```bash
SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT:-$PWD}/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"

bash "$SETUP_SCRIPT" --backend docker --check-only || {
  echo "MISSING: TAO GPU host runtime is not ready."
  echo "After user approval, run (append --yes for non-interactive agent runs):"
  echo "  bash \"$SETUP_SCRIPT\" --backend docker --install"
  exit 1
}
```

Never install silently. If the check fails, explain what is missing, ask the
user to authorize the fix, then run the install command and rerun the check.

## What The Installer Does

The installer dispatches on the detected distribution family. On every
supported family it adds NVIDIA's CUDA and Container Toolkit repositories
(if missing), installs the pinned runtime packages, optionally installs
Docker, wires the NVIDIA Docker runtime, and adds the invoking user to
the `docker` group.

Common steps (all families):

1. Adds NVIDIA's CUDA repository if missing (apt `cuda-keyring` deb,
   `cuda-<distro>.repo` for dnf/zypper).
2. Adds NVIDIA's Container Toolkit repository if missing (`.list` for apt,
   `.repo` for dnf/zypper).
3. Installs the matching kernel header / devel package for the running
   kernel.
4. Installs the driver branch 580 packages, `cuda-toolkit-13-0`, and the
   Container Toolkit pinned to `1.19.0` (the dpkg-suffixed `1.19.0-1` is
   the same upstream version expressed for apt).
5. For Docker backends and when Docker is missing, installs Docker
   (override / opt-out flags below), enables/starts the daemon, then runs
   `nvidia-ctk runtime configure --runtime=docker` and restarts Docker
   when `systemctl` is available.
6. Adds the invoking user (`$SUDO_USER` if available, else `$USER`) to the
   `docker` group so subsequent shells can run `docker` without `sudo` —
   opt out with `--skip-docker-group`. **The new group membership does not
   take effect in the current shell**: log out and back in, or run
   `newgrp docker` in each new shell.
7. Attempts `modprobe nvidia` so verification can pass before reboot.

Family-specific package selections:

| Step | debian-family | rhel-family | suse-family |
|---|---|---|---|
| Kernel headers | `linux-headers-$(uname -r)` | `kernel-devel-$(uname -r)`, `kernel-headers-$(uname -r)` | `kernel-default-devel` |
| Driver | `nvidia-driver-pinning-580`, `nvidia-open-580` (override: `$NVIDIA_DRIVER_PACKAGE_DEBIAN`) | `nvidia-driver-cuda`, `kmod-nvidia-open-dkms` (override: `$NVIDIA_DRIVER_PACKAGE_RHEL`, `$NVIDIA_DRIVER_KMOD_RHEL`) | `nvidia-open-driver-G06-signed-kmp-default` (override: `$NVIDIA_DRIVER_PACKAGE_SUSE`) |
| CUDA toolkit | `cuda-toolkit-13-0` | `cuda-toolkit-13-0` | `cuda-toolkit-13-0` |
| Container Toolkit | `nvidia-container-toolkit=1.19.0-1` + base/tools/libs | `nvidia-container-toolkit-1.19.0` + base/tools/libs | same as rhel |
| Docker | `docker.io` (override: `$DOCKER_PACKAGE_DEBIAN`) | `moby-engine`+`moby-cli` on Fedora when available, else `docker-ce docker-ce-cli containerd.io` from `download.docker.com` | `docker` |

## Verification

After installation, verify:

```bash
nvidia-smi
/usr/local/cuda-13.0/bin/nvcc --version
docker info --format '{{json .Runtimes}}' | grep nvidia
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
```

Expected `nvidia-smi` output includes driver `580.x` and CUDA Version `13.0`.
Expected `nvcc` output includes `release 13.0`.

## Kubernetes Notes

For self-managed Kubernetes clusters, run the host installer on every GPU
worker node or bake the same package set into the node image before installing
the NVIDIA GPU Operator or device plugin.

The workflow check also warns if `kubectl` is available but the cluster reports
no `nvidia.com/gpu` allocatable capacity. In that case, install/configure the
NVIDIA GPU Operator after the worker host runtime is ready:

```bash
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install --wait gpu-operator -n gpu-operator --create-namespace nvidia/gpu-operator
```

Managed Kubernetes providers may own driver installation through node images or
GPU Operator policy. Do not overwrite a provider-managed GPU node without user
approval and a rollback plan.

## Failure Modes

**Unsupported distribution family**: `--install` automates debian-, rhel-,
and suse-family hosts. On Arch, Alpine, Gentoo, NixOS, FreeBSD, or anything
without `/etc/os-release` (e.g. macOS), the script exits with a clear error
that lists the four version targets and the upstream NVIDIA install-guide
URLs:

- `https://docs.nvidia.com/cuda/cuda-installation-guide-linux/`
- `https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html`
- `https://docs.docker.com/engine/install/`

Install those four pieces using your distribution's package manager and
rerun the script with `--check-only` to verify. The check is universally
portable — it only queries the binaries / package databases — so once the
runtime is in place the workflow contract is satisfied regardless of the
underlying distro.

**Unsupported Ubuntu/Debian derivative**: When `ID` is e.g. `pop`, `mint`,
`zorin`, `raspbian`, or another debian-family derivative, the script maps
the host onto the upstream Ubuntu/Debian CUDA repo via `UBUNTU_CODENAME` /
`VERSION_CODENAME` (`focal`/`jammy`/`noble` → Ubuntu 20.04/22.04/24.04;
`bullseye`/`bookworm`/`trixie` → Debian 11/12/12). If the host's codename
doesn't match a known upstream release, `--install` exits with the same
manual-install guidance described above.

**Docker not installed**: `--check-only` reports `MISSING: Docker is not
installed` and prints the exact rerun command appropriate to the detected
distro family. The default `--install` path installs Docker (`docker.io` /
`moby-engine` / `docker-ce` / `docker` depending on family), enables/starts
the daemon, configures the NVIDIA runtime, and adds the invoking user to
the `docker` group. If you prefer to manage Docker yourself, install it
before rerunning the script or pass `--skip-docker-install`.

**Docker installed but `docker run` still needs sudo**: The script adds the
invoking user to the `docker` group, but Linux only refreshes group
membership on a new login session. Log out and back in, or run
`newgrp docker` in each new shell, until the new membership is active.

**Docker runtime still missing**: Restart Docker, then rerun
`nvidia-ctk runtime configure --runtime=docker`.

**Driver branch detected != 580**: The driver-branch pin is exact on
debian-family (`nvidia-open-580`). On rhel-/suse-family the script
installs the latest open driver shipped in NVIDIA's CUDA 13.0 repo for
the detected distro, which is always ≥ 580. If your host needs a stricter
pin, set `$NVIDIA_DRIVER_PACKAGE_RHEL` / `$NVIDIA_DRIVER_KMOD_RHEL` /
`$NVIDIA_DRIVER_PACKAGE_SUSE` to the exact package names you want before
running `--install`.

**Driver installed but `nvidia-smi` fails**: Load the module with
`sudo modprobe nvidia` or reboot. Secure Boot may require MOK enrollment on
systems where it is enabled.

**Kubernetes still has no GPU capacity**: Confirm the driver works on each GPU
node with `nvidia-smi`, then check the GPU Operator/device plugin pods and node
labels.

安装 tao-setup-nvidia-gpu-host

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-setup-nvidia-gpu-host # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 会自动检测并使用该技能
仓库 NVIDIA/skills

相关技能

klingai-upgrade-migration
更新时间 2026-07-03
Verification &amp; Quality Assurance
更新时间 2026-06-29
base44-cli
更新时间 2026-06-29
Railway CLI Management
更新时间 2026-07-02
OR