選項
首頁首頁 Skill 開發營運和 CI/CD tao-run-inference-service

tao-run-inference-service

NVIDIA/skills NVIDIA/skills

透過將容器執行任務委派給適當的平台技能,為特定網路架構啟動、查詢及停止 TAO 推論微服務。

...展開全部
1
更新時間 2026-09-27

TAO 推論微服務

操作說明

要啟動推論服務:

  1. 收集所需輸入(第 1 節)並解析容器映像檔(第 2 節)。
  2. 建構工作負載與內部指令(第 3–4.1 節);使用 references/code-templates.yaml → job_payload_builder.
  3. 「讀取」 skills/platform//SKILL.md 並啟動容器(第 4.2 節)。
  4. 寫入服務註冊表並輪詢就緒狀態(第 4.3 節);使用 references/code-templates.yaml → registry_write. 和 readiness_check.

要傳送推論請求:

  1. 根據第 6.0 節的說明,解析應由哪個服務接收該請求(透過 job_id、透過 network_arch,或當多個服務同時運行時由使用者明確選擇 — 絕不能在存在多個服務時,默默預設為 "latest"),接著從 references/code-templates.yaml → request.registry_read 中讀取端點,並使用已確定的 job_id.
  2. 在建構請求正文之前,應提示使用者輸入 vLLM 風格的取樣參數(第 6.1 節)。呈現 max_tokens, top_p, temperature (以及任何針對特定架構的附加參數)及其預設值;讓使用者可覆寫或跳過各參數以接受預設值。切勿在未經提示的情況下自動採用預設值。
  3. 根據第 6.2 節的說明建構並傳送請求正文;根據第 6.3 節的說明處理回應。

若要停止服務:讀取 references/code-templates.yaml → stop.registry_read 以解析 job_id,讀取 skills/platform//SKILL.md,然後依照第 5 節的說明操作。

參考資料(模式、映射、有效值 — 不包含操作說明):

  • references/service.yaml — 映像映射、有效 network_arch 名稱、工作負載資料結構、環境變數名稱、機密資訊分類。
  • references/request.yaml — 端點定義、請求欄位架構、回應結構、程式碼範例。
  • references/code-templates.yaml — 用於建構有效載荷、註冊表寫入、就緒檢查以及停止/請求流程的 Python 範本。

機密規則(適用於此技能中生成的每個程式碼區塊)

切勿要求使用者在提示中輸入機密值。針對每個機密值:

  1. 告知使用者應設定哪個環境變數(例如 export HF_TOKEN=...).
  2. 生成使用 os.environ["VAR_NAME"] — 切勿將該值硬編碼、插值或透過提示字元索取。

機密環境變數(完整清單請參閱 references/service.yaml → secrets_handling): HF_TOKEN, WANDB_API_KEY, CLEARML_API_ACCESS_KEY, CLEARML_API_SECRET_KEY, TAO_API_KEY, TAO_USER_KEY.

可在提示中安全收集的項目: network_arch, model_path, num_gpus、提示文字、 WANDB_* 設定檔 URL、 CLEARML_*_HOST URL 中收集是安全的。

1. 應從使用者處收集哪些資訊

輸入 角色
network_arch 選擇容器映像檔、各架構的內部指令結構(references/service.yaml → container_commands.),以及 neural_network_name 在工作 JSON 中(如適用)。必須與 valid_network_arch_config_basenames 中 references/service.yaml 中的基名(例如 cosmos-rl, cosmos-predict2.5).
model_path 已訓練的模型檢查點。有效格式: hf_model:/// (HuggingFace Hub — 設定 HF_TOKEN 用於門控模型)或本機容器檔案系統路徑。雲端 URI(s3://, gs://, az://) 不受支援 — 推論服務不依賴雲端儲存。務必詢問使用者;切勿以佔位符替代。參見 references/service.yaml → model_path_protocols.
platform 運算平台: local-docker, brev, slurm,或 kubernetes.
num_gpus 預設值為 1;推論的最小值為 1。

2. 影像解析度

每個 network_arch 皆有一個名為 {network_arch}.config.json。請依下列方式解析容器映像檔:

  1. 請參閱 {network_arch}.config.json 並取用 api_params.image (例如 COSMOS_RL)。此為存取 docker_image_defaults.mapping in references/service.yaml.
  2. 在映射中查詢該鍵。若主機環境變數 IMAGE_ 已設定(例如 IMAGE_COSMOS_RL),則會覆寫映射的預設值。
  3. 映射值通常是存放於儲存庫根目錄 versions.yaml 清單中的點分隔鍵(例如 tao_toolkit.cosmos_rl)。將其解析為具體的 nvcr.io/... 映像 URI,方法是查詢 versions.yaml → images..。絕對 URI 會原樣傳遞,因此 IMAGE_ 包含完整 URI 的環境變數覆寫仍會生效。對應的 Python 輔助程式位於 references/code-templates.yaml.
  4. 若配置檔不存在或 api_params.image 為空,則回退至 COSMOS_RL 鍵。

該設定檔還包含 spec_params.inference.model_path ,該設定用於區分資料夾與檔案路徑的語義:若值中包含子字串 folder,容器會將該路徑視為目錄。

3. 環境變數(無回呼函式)

請在 env_payload 編碼前 env_json前,請在中設定這些變數。請勿設定 TAO_LOGGING_SERVER_URL 或 TAO_ADMIN_KEY.

TAO_EXECUTION_BACKEND — 必須與平台相符:

平台 TAO_EXECUTION_BACKEND 值
local-docker local-docker
brev local-docker
slurm slurm
Kubernetes local-k8s

CLOUD_BASED — 始終 "False" 針對此技能(停用發送回調至 TAO_LOGGING_SERVER_URL).

GPU 環境變數 — 僅在平台技能無法自動處理 GPU 注入時才需使用:

  • Tegra / Jetson: --runtime=nvidia 搭配 NVIDIA_DRIVER_CAPABILITIES=all 以及 NVIDIA_VISIBLE_DEVICES=.
  • 標準 x86 + nvidia-container-toolkit:請使用 Docker device_requests。平台技能會自動處理此事項。

4. 跨平台執行

工作負載與內部指令(第 1–3 節)不依賴特定平台。針對每個平台,請在產生任何執行程式碼之前,參閱 skills/platform//SKILL.md 以進行預檢與憑證驗證。

4.1 建置內部指令(依架構而定)

內部指令的結構詳見 network_arch —— 並無統一的範本。請查閱 references/service.yaml → container_commands.中查閱;若未列出,則表示該架構不被支援——請停止並諮詢。從 references/code-templates.yaml → job_payload_builder.中選擇相應的子區塊。請在指令前加上前綴 umask 0 && ,並確保其在各平台(local-docker、brev、slurm、kubernetes)上保持完全一致。

各架構共通項目:

  • job_id: fresh uuid.uuid4() — 將成為容器名稱與註冊表金鑰。
  • image: 依照第 2 節進行解析。
  • 機密資訊(access_key, secret_key, HF_TOKEN等)會在執行時從環境變數中讀取 — 切勿硬編碼,切勿記錄或輸出。

架構特定注意事項(完整詳情請參閱 references/service.yaml → container_commands):

  • cosmos-rl — 單一 --job '' --docker_env_vars '' 二進位檔; json.dumps(...) + shlex.quote(...). env_payload 包含 TAO_EXECUTION_BACKEND (參見第 3 節表格), TAO_API_JOB_ID, CLOUD_BASED=False。推論服務不依賴雲端儲存; HF_TOKEN 是唯一會適用的憑證環境變數(適用於受控的 HuggingFace 模型)。
  • cosmos-predict2.5 — 標誌式 cosmos_predict inference_microservice start ... --port 8080 (無 setup. 前綴;不接受 tyro.conf.OmitArgPrefixes). --job/--docker_env_vars 均不被接受。翻譯 model_path 為 --checkpoint-path (本機路徑) 或 --model (hf_model://);雲端 URI 將被拒絕。唯一適用之憑證環境變數為 HF_TOKEN 僅適用於受限的 HuggingFace 模型。每次請求的參數(prompt、inference_type、num_output_frames、guidance、seed、num_steps、negative_prompt)應置於請求正文中,而非在啟動時設定。 TAO_EXECUTION_BACKEND/TAO_API_JOB_ID/CLOUD_BASED 這些參數未被使用,可省略。

4.2 將執行權限委派給平台技能

請閱讀 skills/platform//SKILL.md 並依照說明啟動容器。

基礎參數(所有平台):

參數 值
image 已解析的容器映像檔(第 2 節)
command inner — 第 4.1 節中建構的 shell 字串
gpu_count num_gpus
env_vars env_payload
工作/容器名稱 job_id — 必須等同於第 4.1 節中的 UUID,以便註冊表能參照該容器
host_port (local-docker, brev) 用於將主機端埠與容器埠 8080 綁定的埠號。預設值 8080,但每個並行服務必須具有唯一性 — 請參閱下方的端口分配規則。

特定於平台的額外輸入參數:

平台 額外輸入
local-docker 除基礎設定外無其他
brev instance_id (可選 — 重複使用現有實例);若為多憑證/多工作區帳戶,則亦需 cloud_cred_id 以及 workspace_group_id 首次建立時 — 參見 skills/platform/tao-run-on-brev/SKILL.md
slurm partition 以及 account — 檢查 SLURM_PARTITION/SLURM_ACCOUNT 環境變數;若未設定則詢問使用者
Kubernetes namespace (預設: default); image_pull_secret ( nvcr.io 映像檔)

埠綁定(local-docker 和 brev):使用直接的 docker run(而非 DockerSDK),以便 -p :8080 參數得以傳遞,且容器名稱與 job_id 完全相符。

埠號分配規則(local-docker 和 brev,對並行服務而言為必備條件):在啟動服務之前,請讀取註冊表(/tmp/tao-inf-ms-state.json)並收集 host_port 值集合。對於 brev,則需包含同一 instance_id)。從 8080 開始,選取該集合中尚未被佔用的最低可用埠號——例如 host_port = next(p for p in range(8080, 8200) if p not in used_ports)。預設 8080 僅在沒有其他服務正在運行時才適用。這正是「啟動 3 個服務,每個服務皆可透過不同的 host_url」的機制得以運作;若無此機制,服務 2 和 3 將會因 bind: address already in use。SLURM 和 Kubernetes 會透過其自身的平台機制取得不同的端點,因此無需此步驟。

4.3 啟動後:服務註冊表與端點

在平台確認容器正在運行後,請立即寫入服務註冊表。該註冊表(/tmp/tao-inf-ms-state.json) 以 job_id; "latest" 始終指向最近啟動的服務。

請參閱 references/code-templates.yaml → registry_write. 以取得 Python 範本。

平台 host_url platform_job_id 撰寫前需執行的額外步驟
local-docker http://localhost:{host_port} — 無
brev http://{brev_ip}:{host_port} — brev ls → 取得實例 IP(localhost 在遠端虛擬機器上無效)
slurm http://localhost:{host_port} SLURM 排程器工作 ID 等待進入「Running」狀態;透過 SSH 進行埠轉發 localhost:{host_port}→{node}:8080
Kubernetes http://{external_ip}:8080 k8s 工作名稱 kubectl expose job … --type=LoadBalancer; 等待取得外部 IP

寫入註冊表後,輸出 job_id 和 URL:

print(f"Inference service started.")
print(f"  Job ID : {job_id}")
print(f"  Arch   : {network_arch}")
print(f"  URL    : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")

接著輪詢就緒狀態 — 詳見 references/code-templates.yaml → readiness_check。容器會在後台載入模型;在返回 200 狀態碼前,請勿發送請求。

5. 停止推論服務

請使用者提供 job_id 以停止服務。若未提供,則預設使用 state["latest"] ,並確認要停止的 job_id。透過 references/code-templates.yaml → stop.registry_read讀取註冊表,接著讀取 `skills/platform//SKILL.md` 並使用其取消/停止機制。

平台 識別碼以傳遞 額外清理
local-docker job_id_to_stop — 容器名稱 無
brev job_id_to_stop — 容器名稱 無
slurm entry["platform_job_id"] — SLURM 工作 ID pkill -f "ssh.*-L.*{entry['host_port']}"
kubernetes entry["platform_job_id"] — k8s 工作名稱 kubectl delete svc {entry["platform_job_id"]} -n

其中 entry = state[job_id_to_stop]. 停止後,請清理註冊表: references/code-templates.yaml → stop.registry_cleanup.

6. 傳送推論請求

6.0 確定由哪個服務接收此請求(必填)

每個請求都必須路由至運行相應模型的特定服務。路由是透過 job_id — 註冊表會針對 network_arch 每個條目,因此當使用者指定模型名稱而非 job_id。請依序套用以下規則:

  1. 使用者提供了明確的job_id → 使用該。驗證其是否存在於 state.
  2. 使用者指定了 network_arch(例如「將此請求傳送至 cosmos-rl 服務」)→ 查詢相符的條目: candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch].
    • 僅有一項完全匹配 → 使用該項目。
    • 有多個匹配項 → 向使用者提示候選 job_id及其 started_at;切勿自動選取。
    • 無匹配項目 → 停止並告知使用者該架構下沒有正在運行的服務。
  3. 若無job_id且無network_arch → 統計非"latest" 的條目數量 state:
    • 恰好有一個正在運作的服務 → 使用該服務。
    • 兩個或以上 → 不要默默預設為 state["latest"]。向使用者顯示完整清單(job_id, network_arch, host_url),並要求明確選擇。該 "latest" 指標僅是單一服務工作流程的便利功能,而非多項服務共存時的路由備用方案。
    • 零 → 暫停並告知使用者需先啟動服務。

解析完成後,從註冊表中讀取端點(references/code-templates.yaml → request.registry_read),並將解析後的 job_id 作為 user_provided_job_id。向使用者確認:「正在傳送至 job_id=… arch=… url=…」。若服務可能仍在載入中,請先輪詢就緒狀態(references/code-templates.yaml → readiness_check).

傳送前進行交叉檢查:若使用者提供的請求正文包含架構特定欄位(例如 guidance / num_steps / seed / negative_prompt → cosmos-predict2.5;必填 image_url/video_url 內容項目 → cosmos-rl),請驗證其是否與 state[job_id]["network_arch"]。若不符,則停止並詢問 — 將 cosmos-predict2.5 的請求內容傳送至 cosmos-rl 服務時,會在容器端因 4xx/5xx 錯誤而失敗,此情況比在此處攔截更難診斷。

6.1 採樣參數 — 每次請求前必須向使用者提示

在建構請求正文之前,您必須明確提示使用者輸入 vLLM 風格的抽樣參數。請勿在未經提示的情況下自動套用預設值。請使用結構化的提示,每個欄位一個問題,且該提示須:

  1. 列出所有適用欄位及其類型與預設值。
  2. 允許使用者跳過/接受任何欄位以採用該欄位的預設值 — 絕不強制要求輸入值。
  3. 在單一輪次中收集所有欄位。

提示結束後,請原樣套用每個使用者輸入的值,並以預設值取代任何被跳過的欄位。切勿自行編造數值或默默進行限制。

欄位清單、預設值及各架構適用性: references/request.yaml → chat_completions_request_body (基礎採樣欄位: max_tokens, top_p, temperature) 以及 network_arch_constraints. (各架構的覆寫設定及額外項目,例如 guidance/num_steps/seed/negative_prompt 針對 cosmos-predict2.5)。若某欄位被標記為當前架構不支援,則勿提示該欄位,亦勿將其納入正文中。

6.2 請求格式

請傳送一則 POST 至 {BASE_URL}/v1/chat/completions ,並設定 Content-Type: application/json ,並設定至少 300 秒的超時時間。正文須符合 OpenAI 規格(vLLM 聊天補全);請參閱 references/request.yaml → chat_completions_request_body 以查看完整的欄位架構及內容項型態(文字 / image_url / video_url),並參閱 code_examples 此處可取得可直接執行的 Python 與 curl 範例。

限制條件:僅處理第一則使用者訊息。請求正文中不得包含機密值。各網路的特定限制(例如:cosmos-rl 要求每則請求都必須包含圖片或影片;cosmos-rl 會拒絕 data: URI)。相關限制詳見 references/request.yaml → network_arch_constraints.

6.3 回應處理

HTTP 狀態碼 含義 動作
200 成功 — choices[0].message.content 包含生成的文字 讀取結果
202 伺服器仍在初始化中,或模型仍在載入中 延遲後重試
503 初始化失敗、模型載入失敗,或模型尚未準備就緒 檢查 error.type: model_not_ready → 重試; initialization_error / model_load_error → 放棄並檢查日誌
400 JSON 內容缺失或為空 修正請求
500 推論過程中發生未處理的例外狀況 檢查容器日誌

對於 202 和 503 狀態碼,請求體包含 {"error": {"type": "", "message": ""}}。詳見 container_response_shapes 在 references/request.yaml 中查閱錯誤類型字串。

在 GitHub 上查看
---
name: tao-run-inference-service
description: Start, query, and stop a TAO inference microservice for a specific network architecture by delegating container execution to the appropriate platform skill.
license: Apache-2.0
---

# TAO Inference Microservice

## Instructions

**To start an inference service:**
1. Collect required inputs (Section 1) and resolve the container image (Section 2).
2. Build the job payload and inner command (Sections 3–4.1); use `references/code-templates.yaml` → `job_payload_builder`.
3. Read `skills/platform/<platform>/SKILL.md` and start the container (Section 4.2).
4. Write the service registry and poll readiness (Section 4.3); use `references/code-templates.yaml` → `registry_write.<platform>` and `readiness_check`.

**To send an inference request:**
1. Resolve which service receives the request per Section 6.0 (by `job_id`, by `network_arch`, or by explicit user choice when multiple services run — **never silently default to `"latest"` when more than one service exists**), then read the endpoint from `references/code-templates.yaml` → `request.registry_read` with the resolved `job_id`.
2. **Before building the request body, prompt the user for the vLLM-style sampling parameters (Section 6.1).** Present `max_tokens`, `top_p`, `temperature` (and any per-arch extras) with their defaults; let the user override or skip each one to accept the default. Never silently use defaults.
3. Build and send the body per Section 6.2; handle the response per Section 6.3.

**To stop a service:** Read `references/code-templates.yaml` → `stop.registry_read` to resolve the job_id, read `skills/platform/<platform>/SKILL.md`, then follow Section 5.

**Reference data** (schemas, mappings, valid values — no instructions):
- **`references/service.yaml`** — image mappings, valid `network_arch` names, job payload schema, env var names, secrets classification.
- **`references/request.yaml`** — endpoint definition, request field schema, response shapes, code examples.
- **`references/code-templates.yaml`** — Python templates for payload building, registry writes, readiness checks, and stop/request flows.

---

## Secrets rule (applies to every generated code block in this skill)

**Never ask the user to type a secret value into a prompt.** For every secret value:
1. Tell the user which environment variable to set (e.g. `export HF_TOKEN=...`).
2. Generate code that reads it with `os.environ["VAR_NAME"]` — never hard-code, interpolate, or prompt for the value.

**Secret env vars** (full list in `references/service.yaml` → `secrets_handling`):
`HF_TOKEN`, `WANDB_API_KEY`, `CLEARML_API_ACCESS_KEY`, `CLEARML_API_SECRET_KEY`, `TAO_API_KEY`, `TAO_USER_KEY`.

**Safe to collect in the prompt:** `network_arch`, `model_path`, `num_gpus`, prompt text, `WANDB_*` config URLs, `CLEARML_*_HOST` URLs.

---

## 1. What to collect from the user

| Input | Role |
|--------|------|
| **`network_arch`** | Chooses container image, the per-arch inner command shape (`references/service.yaml` → `container_commands.<network_arch>`), and `neural_network_name` in the job JSON when applicable. Must match a basename in `valid_network_arch_config_basenames` in `references/service.yaml` (e.g. `cosmos-rl`, `cosmos-predict2.5`). |
| **`model_path`** | The trained model checkpoint. Valid forms: `hf_model://<org>/<model>` (HuggingFace Hub — set `HF_TOKEN` for gated models) or a local container filesystem path. Cloud URIs (`s3://`, `gs://`, `az://`) are NOT supported — the inference service has no cloud-storage dependency. Always ask the user; never substitute a placeholder. See `references/service.yaml` → `model_path_protocols`. |
| **`platform`** | Compute platform: `local-docker`, `brev`, `slurm`, or `kubernetes`. |
| **`num_gpus`** | Defaults to **1**; minimum **1** for inference. |

---

## 2. Image resolution

Each `network_arch` has a sidecar config file named `{network_arch}.config.json`. Resolve the container image as follows:

1. Read `{network_arch}.config.json` and take `api_params.image` (e.g. `COSMOS_RL`). This is a key into `docker_image_defaults.mapping` in `references/service.yaml`.
2. Look up that key in the mapping. If the host env var `IMAGE_<KEY>` is set (e.g. `IMAGE_COSMOS_RL`), it overrides the mapped default.
3. The mapped value is normally a dotted key into the repo-root `versions.yaml` manifest (e.g. `tao_toolkit.cosmos_rl`). Resolve it to a concrete `nvcr.io/...` image URI by looking up `versions.yaml` → `images.<group>.<name>`. Absolute URIs pass through unchanged, so an `IMAGE_<KEY>` env-var override that contains a full URI still works. The Python helper for this lives in `references/code-templates.yaml`.
4. If the config file is missing or `api_params.image` is empty, fall back to the `COSMOS_RL` key.

The config file also has `spec_params.inference.model_path` which drives **folder vs file** path semantics: if the value contains the substring `folder`, the container treats the path as a directory.

---

## 3. Environment variables (no callbacks)

Set these in `env_payload` before encoding `env_json`. Do **not** set `TAO_LOGGING_SERVER_URL` or `TAO_ADMIN_KEY`.

**`TAO_EXECUTION_BACKEND`** — must match the platform:

| Platform | `TAO_EXECUTION_BACKEND` value |
|----------|-------------------------------|
| local-docker | `local-docker` |
| brev | `local-docker` |
| slurm | `slurm` |
| kubernetes | `local-k8s` |

**`CLOUD_BASED`** — always `"False"` for this skill (disables callback posting to `TAO_LOGGING_SERVER_URL`).

**GPU env vars** — only needed when the platform skill does not handle GPU injection automatically:
- Tegra / Jetson: `--runtime=nvidia` with `NVIDIA_DRIVER_CAPABILITIES=all` and `NVIDIA_VISIBLE_DEVICES=<ids>`.
- Standard x86 + nvidia-container-toolkit: use Docker `device_requests`. The platform skill handles this.

---

## 4. Executing across platforms

The job payload and inner command (Sections 1–3) are **platform-agnostic**. For each platform, read **`skills/platform/<name>/SKILL.md`** for preflight checks and credentials **before** generating any execution code.

### 4.1 Build the inner command (per arch)

The inner-command shape is **per `network_arch`** — there is no uniform template. Look up the per-arch entry in `references/service.yaml` → `container_commands.<network_arch>`; if not present, the arch is unsupported — stop and ask. Pick the matching sub-block in `references/code-templates.yaml` → `job_payload_builder.<network_arch>`. Prefix the command with `umask 0 &&` and keep it **identical across platforms** (local-docker, brev, slurm, kubernetes).

Common across arches:

- `job_id`: fresh `uuid.uuid4()` — becomes the container name and registry key.
- `image`: resolve per Section 2.
- Secrets (`access_key`, `secret_key`, `HF_TOKEN`, etc.) are read from env vars at runtime — never hard-code, never log or print.

Arch-specific notes (full details in `references/service.yaml` → `container_commands`):

- **`cosmos-rl`** — single `--job '<JOB_JSON>' --docker_env_vars '<ENV_JSON>'` blob; `json.dumps(...)` + `shlex.quote(...)`. `env_payload` carries `TAO_EXECUTION_BACKEND` (per Section 3 table), `TAO_API_JOB_ID`, `CLOUD_BASED=False`. The inference service has no cloud-storage dependency; `HF_TOKEN` is the only cred env var that ever applies (for gated HuggingFace models).
- **`cosmos-predict2.5`** — flag-style `cosmos_predict inference_microservice start ... --port 8080` (no `setup.` prefix; uses `tyro.conf.OmitArgPrefixes`). `--job`/`--docker_env_vars` are **not** accepted. Translate `model_path` to `--checkpoint-path` (local path) or `--model <registered_key>` (`hf_model://`); cloud URIs are rejected. The only cred env var that ever applies is `HF_TOKEN` for gated HuggingFace models. Per-request params (prompt, inference_type, num_output_frames, guidance, seed, num_steps, negative_prompt) go in the request body, not at startup. `TAO_EXECUTION_BACKEND`/`TAO_API_JOB_ID`/`CLOUD_BASED` are unused and may be omitted.

### 4.2 Delegate execution to the platform skill

Read **`skills/platform/<platform>/SKILL.md`** and follow it to start the container.

**Base parameters (all platforms):**

| Parameter | Value |
|-----------|-------|
| `image` | resolved container image (Section 2) |
| `command` | `inner` — the shell string built in Section 4.1 |
| `gpu_count` | `num_gpus` |
| `env_vars` | `env_payload` |
| job / container name | `job_id` — must equal the UUID from 4.1 so the registry can reference it |
| `host_port` *(local-docker, brev)* | host-side port to bind to container port 8080. Default `8080`, but **must be unique per concurrent service** — see the port-allocation rule below. |

**Platform-specific additional inputs:**

| Platform | Additional inputs |
|----------|------------------|
| **local-docker** | None beyond base |
| **brev** | `instance_id` (optional — reuse an existing instance); on multi-credential / multi-workspace accounts also `cloud_cred_id` and `workspace_group_id` for first-create — see `skills/platform/tao-run-on-brev/SKILL.md` |
| **slurm** | `partition` and `account` — check `SLURM_PARTITION`/`SLURM_ACCOUNT` env vars; ask user if unset |
| **kubernetes** | `namespace` (default: `default`); `image_pull_secret` (required for `nvcr.io` images) |

**Port binding (local-docker and brev):** use **direct docker run** (not DockerSDK) so that `-p <host_port>:8080` can be passed and the container name equals `job_id` exactly.

**Port allocation rule (local-docker and brev, REQUIRED for concurrent services):** Before starting a service, read the registry (`/tmp/tao-inf-ms-state.json`) and collect the set of `host_port` values from every existing entry on the same platform (and, for brev, the same `instance_id`). Pick the **lowest free port starting from 8080** that is not in that set — e.g. `host_port = next(p for p in range(8080, 8200) if p not in used_ports)`. The default `8080` only applies when no other service is running. This is what makes "start 3 services, each reachable at a distinct `host_url`" work; without it, services 2 and 3 fail with `bind: address already in use`. SLURM and kubernetes get distinct endpoints from their own platform mechanisms and do not need this step.

### 4.3 After start: service registry and endpoint

Write the service registry immediately after the platform confirms the container is running. The registry (`/tmp/tao-inf-ms-state.json`) is keyed by `job_id`; `"latest"` always points to the most recently started service.

See `references/code-templates.yaml` → `registry_write.<platform>` for the Python template.

| Platform | `host_url` | `platform_job_id` | Extra step before writing |
|----------|-----------|-------------------|--------------------------|
| **local-docker** | `http://localhost:{host_port}` | — | None |
| **brev** | `http://{brev_ip}:{host_port}` | — | `brev ls` → get instance IP (`localhost` is invalid on remote VM) |
| **slurm** | `http://localhost:{host_port}` | SLURM scheduler job ID | Wait until Running; SSH port-forward `localhost:{host_port}→{node}:8080` |
| **kubernetes** | `http://{external_ip}:8080` | k8s job name | `kubectl expose job … --type=LoadBalancer`; wait for external IP |

After writing the registry, print the job_id and URL:

```python
print(f"Inference service started.")
print(f"  Job ID : {job_id}")
print(f"  Arch   : {network_arch}")
print(f"  URL    : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")
```

Then poll for readiness — see `references/code-templates.yaml` → `readiness_check`. The container loads the model in the background; do not send requests before it returns 200.

---

## 5. Stopping the inference service

Ask the user for the `job_id` to stop. If they don't provide one, default to `state["latest"]` and confirm which job_id is being stopped. Read the registry using `references/code-templates.yaml` → `stop.registry_read`, then read **`skills/platform/<platform>/SKILL.md`** and use its cancellation / stop mechanism.

| Platform | Identifier to pass | Extra cleanup |
|----------|--------------------|---------------|
| **local-docker** | `job_id_to_stop` — container name | None |
| **brev** | `job_id_to_stop` — container name | None |
| **slurm** | `entry["platform_job_id"]` — SLURM job ID | `pkill -f "ssh.*-L.*{entry['host_port']}"` |
| **kubernetes** | `entry["platform_job_id"]` — k8s job name | `kubectl delete svc {entry["platform_job_id"]} -n <namespace>` |

where `entry = state[job_id_to_stop]`. After stopping, clean up the registry: `references/code-templates.yaml` → `stop.registry_cleanup`.

---

## 6. Sending inference requests

### 6.0 Resolve which service receives this request (REQUIRED)

Each request must be routed to the **specific** service that runs the matching model. Routing happens by `job_id` — the registry stores `network_arch` per entry, so you can resolve a target by arch when the user names a model instead of a `job_id`. Apply these rules in order:

1. **User provided an explicit `job_id`** → use it. Verify it exists in `state`.
2. **User named a `network_arch`** (e.g. "send this to the cosmos-rl service") → look up matching entries: `candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch]`.
   - Exactly one match → use it.
   - Multiple matches → **prompt the user** with the candidate `job_id`s and their `started_at`; do not auto-pick.
   - No match → stop and tell the user no service for that arch is running.
3. **No `job_id` and no `network_arch`** → count non-`"latest"` entries in `state`:
   - Exactly one running service → use it.
   - Two or more → **do not silently default to `state["latest"]`**. Prompt the user with the full list (`job_id`, `network_arch`, `host_url`) and require an explicit choice. The `"latest"` pointer is a convenience for single-service workflows, not a routing fallback when multiple services coexist.
   - Zero → stop and tell the user to start a service first.

After resolving, read the endpoint from the registry (`references/code-templates.yaml` → `request.registry_read`), passing the resolved `job_id` as `user_provided_job_id`. Confirm to the user: "Sending to job_id=… arch=… url=…". If the service may still be loading, poll readiness first (`references/code-templates.yaml` → `readiness_check`).

**Cross-check before sending:** if the user-supplied request body contains arch-specific fields (e.g. `guidance` / `num_steps` / `seed` / `negative_prompt` → cosmos-predict2.5; required `image_url`/`video_url` content items → cosmos-rl), verify they are consistent with `state[job_id]["network_arch"]`. On mismatch, stop and ask — sending a cosmos-predict2.5 body to a cosmos-rl service will fail at the container with a 4xx/5xx that is harder to diagnose than catching it here.

### 6.1 Sampling parameters — REQUIRED user prompt before each request

Before constructing the request body, you **MUST** explicitly prompt the user for the vLLM-style sampling parameters. Do **not** silently apply defaults. Use a structured prompt, one question per field, that:

1. Lists every applicable field with its **type** and **default value**.
2. Lets the user skip / accept any field to take that field's default — entering a value is never required.
3. Collects all fields in one round.

After the prompt, apply each user-entered value verbatim and substitute the default for any skipped field. Do not invent values or silently clamp.

**Field list, defaults, and per-arch applicability:** `references/request.yaml` → `chat_completions_request_body` (base sampling fields: `max_tokens`, `top_p`, `temperature`) and `network_arch_constraints.<network_arch>` (per-arch overrides and extras such as `guidance`/`num_steps`/`seed`/`negative_prompt` for `cosmos-predict2.5`). If a field is marked unsupported for the active arch, do **not** prompt for it and do **not** include it in the body.

### 6.2 Request format

Send a `POST` to `{BASE_URL}/v1/chat/completions` with `Content-Type: application/json` and a timeout of **at least 300 s**. The body is OpenAI-compatible (vLLM chat completions); see `references/request.yaml` → `chat_completions_request_body` for the full field schema and content-item shapes (text / image_url / video_url), and `code_examples` for ready-to-run Python and curl samples.

**Constraints:** only the first user message is processed. No secret values in request bodies. **Per-network constraints** (e.g. cosmos-rl requires every request to include an image or video; cosmos-rl rejects `data:` URIs) are in `references/request.yaml` → `network_arch_constraints`.

### 6.3 Response handling

| HTTP status | Meaning | Action |
|-------------|---------|--------|
| **200** | Success — `choices[0].message.content` has the generated text | Read result |
| **202** | Server still initializing or model still loading | Retry after a delay |
| **503** | Initialization failed, model load failed, **or model not yet ready** | Inspect `error.type`: `model_not_ready` → retry; `initialization_error` / `model_load_error` → give up and check logs |
| **400** | Missing or empty JSON body | Fix request |
| **500** | Unhandled exception during inference | Check container logs |

For 202 and 503, the body contains `{"error": {"type": "<error_type>", "message": "<reason>"}}`. See `container_response_shapes` in `references/request.yaml` for error type strings.

安裝 tao-run-inference-service

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

klingai-upgrade-migration
更新時間 2026-07-03
Verification &amp; Quality Assurance
更新時間 2026-06-29
base44-cli
更新時間 2026-06-29
Railway CLI Management
更新時間 2026-07-02
OR