オプション
家家 Skill DevOps と CI/CD tao-run-inference-service

tao-run-inference-service

NVIDIA/skills NVIDIA/skills

特定のネットワークアーキテクチャ向けのTAO推論マイクロサービスを、コンテナの実行を適切なプラットフォームスキルに委任することで、起動、クエリ実行、停止を行います。

...すべて拡張します
1
更新された時間 2026年9月27日

TAO推論マイクロサービス

手順

推論サービスを開始するには:

  1. 必要な入力(セクション 1)を収集し、コンテナイメージを解決します(セクション 2)。
  2. ジョブのペイロードと内部コマンドを構築します(セクション3~4.1)。 references/code-templates.yaml → job_payload_builder.
  3. Read skills/platform//SKILL.md を使用してコンテナを起動します(セクション 4.2)。
  4. サービスレジストリへの書き込みとレディネス状態のポーリングを行います(セクション 4.3)。 references/code-templates.yaml → registry_write. および readiness_check.

推論リクエストを送信するには:

  1. セクション6.0に従って、どのサービスがリクエストを受信するかを決定します( job_id、またはによって、どのサービスがリクエストを受け取るかを決定し(セクション6.0に従って)、その後、エンドポイントを読み取 network_arch、または複数のサービスが実行されている場合はユーザーの明示的な選択によって決定する(複数のサービスが存在する場合、決して黙って"latest"をデフォルトとして使用してはならない)。その後、 references/code-templates.yaml → request.registry_read からエンドポイントを読み取る job_id.
  2. リクエスト本体を構築する前に、vLLM 形式のサンプリングパラメータ(第 6.1 節)についてユーザーに確認を求める。 max_tokens, top_p, temperature (およびアーキテクチャごとの追加パラメータ)をデフォルト値と共に提示し、ユーザーが各パラメータを上書きするか、スキップしてデフォルトを受け入れるかを選択できるようにする。決して黙ってデフォルト値を使用してはならない。
  3. セクション 6.2 に従ってリクエスト・ボディを構築・送信し、セクション 6.3 に従ってレスポンスを処理する。

サービスを停止するには: references/code-templates.yaml → stop.registry_read を読み込んで job_id を解決し、 skills/platform//SKILL.mdを読み取り、セクション5の手順に従う。

参照データ(スキーマ、マッピング、有効な値 — 指示は含まない):

  • references/service.yaml — イメージのマッピング、有効な network_arch 名、ジョブペイロードスキーマ、環境変数名、シークレットの分類。
  • references/request.yaml — エンドポイントの定義、リクエストフィールドのスキーマ、レスポンスのシェイプ、コード例。
  • references/code-templates.yaml — ペイロード構築、レジストリへの書き込み、レディネスチェック、および停止/リクエストフロー用のPythonテンプレート。

シークレットに関するルール(このスキルで生成されるすべてのコードブロックに適用されます)

ユーザーにプロンプトでシークレット値を入力するよう求めてはなりません。すべてのシークレット値について:

  1. ユーザーに、どの環境変数を設定すべきかを明示してください(例: export HF_TOKEN=...).
  2. それを読み取るコードを生成し、 os.environ["VAR_NAME"] — 値をハードコーディング、補間、またはプロンプトで要求してはなりません。

シークレット環境変数(完全なリストは references/service.yaml → secrets_handling): HF_TOKEN, WANDB_API_KEY, CLEARML_API_ACCESS_KEY, CLEARML_API_SECRET_KEY, TAO_API_KEY, TAO_USER_KEY.

プロンプトで収集しても安全な項目: network_arch, model_path, num_gpus、プロンプトテキスト、 WANDB_* 設定用URL、 CLEARML_*_HOST URL。

1. ユーザーから収集すべき情報

入力 ロール
network_arch コンテナイメージ、アーキテクチャごとの内部コマンドの形式(references/service.yaml → container_commands.)、および neural_network_name 該当する場合はジョブJSON内の項目を選択します。 valid_network_arch_config_basenames 内の references/service.yaml 内のベース名と一致する必要があります(例:トレーニング済みモデルのチェックポイント。有効な形式: cosmos-rl, cosmos-predict2.5).
model_path 学習済みモデルのチェックポイント。有効な形式: hf_model:/// (HuggingFace Hub — ゲート付きモデルの場合は HF_TOKEN )またはローカルのコンテナファイルシステムのパス。クラウドURI(s3://, gs://, az://)はサポートされていません — 推論サービスにはクラウドストレージへの依存関係がありません。必ずユーザーに確認し、プレースホルダーで置き換えないでください。 references/service.yaml → model_path_protocols.
platform コンピューティングプラットフォーム: local-docker, brev, slurm、または kubernetes.
num_gpus デフォルトは 1 です。推論には最小値 1 が必要です。

2. 画像の解像度

各 network_arch には、 {network_arch}.config.jsonという名前のサイドカー設定ファイルを持っています。コンテナイメージは次のように解決します:

  1. 読み込み {network_arch}.config.json を読み込み、 api_params.image (例: COSMOS_RL)。これは docker_image_defaults.mapping in references/service.yaml.
  2. そのキーをマッピングで検索します。ホスト環境変数 IMAGE_ が設定されている場合(例: IMAGE_COSMOS_RL)が設定されている場合、マッピングされたデフォルト値よりも優先されます。
  3. マッピングされた値は通常、リポジトリルートの versions.yaml マニフェストへのドット区切りのキーです(例: tao_toolkit.cosmos_rl)。これを具体的な nvcr.io/... イメージURIに解決します versions.yaml → images..ことで、具体的なイメージURIに解決します。絶対URIは変更されずに透過的に渡されるため、完全なURIを含む IMAGE_ 完全なURIを含む環境変数による上書きは引き続き機能します。これに対応するPythonヘルパーは references/code-templates.yaml.
  4. 設定ファイルが存在しないか、 api_params.image 空の場合、 COSMOS_RL キーにフォールバックします。

また、設定ファイルには spec_params.inference.model_path が含まれており、これによりフォルダパスとファイルパスの区別が決まります。値に folderが含まれている場合、コンテナはそのパスをディレクトリとして扱います。

3. 環境変数(コールバックなし)

これらを env_payload で設定してください env_json前にこれらを設定してください。 TAO_LOGGING_SERVER_URL または TAO_ADMIN_KEY.

TAO_EXECUTION_BACKEND — プラットフォームと一致させる必要があります:

プラットフォーム TAO_EXECUTION_BACKEND 値
local-docker local-docker
brev local-docker
slurm slurm
Kubernetes local-k8s

CLOUD_BASED — 常に "False" このスキルでは常に( TAO_LOGGING_SERVER_URL).

GPU環境変数へのコールバック投稿を無効化 — プラットフォームスキルがGPUインジェクションを自動的に処理しない場合にのみ必要:

  • Tegra / Jetson: --runtime=nvidia と NVIDIA_DRIVER_CAPABILITIES=all および NVIDIA_VISIBLE_DEVICES=.
  • 標準 x86 + nvidia-container-toolkit: Docker を使用 device_requestsを使用します。プラットフォームスキルがこれを処理します。

4. プラットフォームをまたいだ実行

ジョブのペイロードおよび内部コマンド(セクション1~3)は、プラットフォームに依存しません。各プラットフォームについて、実行コードを生成する前に、事前チェックおよび認証情報については skills/platform//SKILL.md を参照してください。

4.1 内部コマンドの構築(アーキテクチャごと)

内部コマンドの構成は network_arch に準拠しており、統一されたテンプレートはありません。 references/service.yaml → container_commands.を参照してください。該当するエントリがない場合、そのアーキテクチャはサポートされていません。その場合は作業を中止し、問い合わせを行ってください。 references/code-templates.yaml → job_payload_builder.から適切なサブブロックを選択してください。コマンドの先頭に umask 0 && を付け、すべてのプラットフォーム(local-docker、brev、slurm、kubernetes)で同一のものを使用してください。

すべてのアーキテクチャに共通する項目:

  • job_id: fresh uuid.uuid4() — これがコンテナ名およびレジストリキーになります。
  • image: セクション2に従って解決します。
  • シークレット(access_key, secret_key, HF_TOKENなど)は実行時に環境変数から読み込まれます。— 決してハードコーディングせず、ログや出力も行わないでください。

アーキテクチャ固有の注意事項(詳細は references/service.yaml → container_commands):

  • cosmos-rl — 単一の --job '' --docker_env_vars '' blob; json.dumps(...) + shlex.quote(...). env_payload には TAO_EXECUTION_BACKEND (セクション3の表参照)、 TAO_API_JOB_ID, CLOUD_BASED=False。推論サービスにはクラウドストレージへの依存関係はありません。 HF_TOKEN (ゲート付きHuggingFaceモデルの場合)適用される唯一の認証環境変数である。
  • cosmos-predict2.5 — フラグ形式 cosmos_predict inference_microservice start ... --port 8080 ( setup. 接頭辞なし; tyro.conf.OmitArgPrefixes). --job/--docker_env_vars は受け入れられません。翻訳 model_path を --checkpoint-path (ローカルパス) または --model (hf_model://)に翻訳してください。クラウドURIは拒否されます。適用されるcred環境変数は、 HF_TOKEN ゲート付きHuggingFaceモデルに対してのみ適用される。リクエストごとのパラメータ(prompt、inference_type、num_output_frames、guidance、seed、num_steps、negative_prompt)は、起動時ではなくリクエスト本文に記述する。 TAO_EXECUTION_BACKEND/TAO_API_JOB_ID/CLOUD_BASED は未使用であり、省略可能です。

4.2 プラットフォームスキルへの実行の委譲

skills/platform//SKILL.md を参照し、その手順に従ってコンテナを起動してください。

基本パラメータ(すべてのプラットフォーム):

パラメータ 値
image 解決済みのコンテナイメージ(セクション 2)
command inner — セクション4.1で構築されたシェル文字列
gpu_count num_gpus
env_vars env_payload
ジョブ/コンテナ名 job_id — レジストリが参照できるように、4.1で生成されたUUIDと一致している必要がある
host_port (local-docker, brev) コンテナのポート 8080 にバインドするホスト側のポート。デフォルト 8080ですが、同時実行されるサービスごとに一意である必要があります — 以下のポート割り当てルールを参照してください。

プラットフォーム固有の追加入力項目:

プラットフォーム 追加の入力
local-docker 基本設定以外になし
brev instance_id (オプション — 既存のインスタンスを再利用);複数の認証情報/複数のワークスペースを持つアカウントの場合は、以下も指定 cloud_cred_id および workspace_group_id 初回作成時 — 参照 skills/platform/tao-run-on-brev/SKILL.md
slurm partition および初回作成時の場合は — 参照slurm account — SLURM_PARTITION/SLURM_ACCOUNT 環境変数を確認し、設定されていない場合はユーザーに確認を求める
kubernetes namespace (デフォルト: default); image_pull_secret ( nvcr.io イメージ)

ポートバインディング(local-docker および brev): -p :8080 引数を渡せるようにし、コンテナ名が job_id ようにします。

ポート割り当てルール(local-docker および brev、同時実行サービスでは必須):サービスの起動前に、レジストリ(/tmp/tao-inf-ms-state.json)を読み込み、 host_port 値のセットを収集する。サービスを開始する前に、レジストリ()を読み込み、同じプラットフォーム上のすべての既存エントリ(また、brevの場合は同じ instance_id)から、一連の値を収集します。そのセットに含まれていない、8080 以降で最も低い空きポートを選択します — 例: host_port = next(p for p in range(8080, 8200) if p not in used_ports)。デフォルトの 8080 設定は、他のサービスが実行されていない場合にのみ適用されます。これにより、「3つのサービスを起動し、それぞれが異なる host_url」という仕組みが機能する理由です。これがなければ、サービス2と3は bind: address already in useで失敗します。SLURM や Kubernetes は、それぞれのプラットフォームの仕組みから固有のエンドポイントを取得するため、この手順は必要ありません。

4.3 起動後:サービスレジストリとエンドポイント

プラットフォームがコンテナの実行を確認したら、直ちにサービスレジストリを書き込みます。レジストリ(/tmp/tao-inf-ms-state.json)のキーは job_id; "latest" 常に、最も最近起動されたサービスを指すように設定されます。

Pythonテンプレートについては references/code-templates.yaml → registry_write. を参照してください。

プラットフォーム host_url platform_job_id 記述前の追加手順
local-docker http://localhost:{host_port} — なし
brev http://{brev_ip}:{host_port} — brev ls → インスタンスのIPを取得(localhost リモートVMでは無効)
slurm http://localhost:{host_port} SLURMスケジューラのジョブID 「Running」状態になるまで待機;SSHポートフォワード localhost:{host_port}→{node}:8080
Kubernetes http://{external_ip}:8080 k8sジョブ名 kubectl expose job … --type=LoadBalancer; 外部IPが取得できるまで待機

レジストリへの書き込み後、job_id と URL を出力します:

print(f"Inference service started.")
print(f"  Job ID : {job_id}")
print(f"  Arch   : {network_arch}")
print(f"  URL    : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")

その後、レディネスをポーリングする — 詳細は references/code-templates.yaml → readiness_checkを参照。コンテナはバックグラウンドでモデルを読み込むため、200が返されるまではリクエストを送信しないでください。

5. 推論サービスの停止

ユーザーに停止を依頼します。 job_id を尋ねてください。指定がない場合は、デフォルトで state["latest"] をデフォルトとして設定し、停止される job_id を確認します。 references/code-templates.yaml → stop.registry_readを使用してレジストリを読み取り、skills/platform//SKILL.md を参照して、そのキャンセル/停止メカニズムを使用してください。

プラットフォーム 識別子を渡す 追加のクリーンアップ
local-docker job_id_to_stop — コンテナ名 なし
brev job_id_to_stop — コンテナ名 なし
slurm entry["platform_job_id"] — SLURM ジョブ ID pkill -f "ssh.*-L.*{entry['host_port']}"
kubernetes entry["platform_job_id"] — k8sジョブ名 kubectl delete svc {entry["platform_job_id"]} -n

ここで entry = state[job_id_to_stop]。停止後は、レジストリをクリーンアップしてください: references/code-templates.yaml → stop.registry_cleanup.

6. 推論リクエストの送信

6.0 このリクエストを受け取るサービスを特定する(必須)

各リクエストは、該当するモデルを実行している特定のサービスにルーティングされなければなりません。ルーティングは job_id — レジストリは network_arch エントリごとに保存されるため、ユーザーが job_id。以下のルールを順に適用します:

  1. ユーザーが明示的にjob_idを指定した場合 → それを使用する。それが state.
  2. ユーザーがnetwork_archを指定した場合(例:「これをcosmos-rlサービスに送信」など)→ 一致するエントリを検索: candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch].
    • 完全に一致するエントリが1つだけ → それを使用する。
    • 一致するエントリが複数ある場合 → 候補となる job_id候補と started_at;自動選択は行わない。
    • 一致するエントリがない → 処理を中止し、そのアーキテクチャに対応するサービスが実行されていないことをユーザーに通知する。
  3. job_id も network_arch もない場合 → 該当しないエントリの数を数える"latest" エントリの数を数える state:
    • 実行中のサービスがちょうど1つ → それを使用する。
    • 2つ以上 → 黙ってstate["latest"]をデフォルトにしない。完全なリストを表示してユーザーに確認を求める(job_id, network_arch, host_url)を表示してユーザーに明示的な選択を求めます。この "latest" ポインタは、単一サービスのワークフローにおける便宜上の機能であり、複数のサービスが共存する場合のルーティングのフォールバックではない。
    • 0 → 処理を停止し、ユーザーにまずサービスを起動するよう指示する。

解決後、レジストリからエンドポイントを読み取り(references/code-templates.yaml → request.registry_read)からエンドポイントを読み取り、解決された job_id をとして渡す。 user_provided_job_idとして渡す。ユーザーに「job_id=… arch=… url=… に送信中」と確認を表示する。サービスがまだ読み込み中の可能性がある場合は、まず準備完了状態をポーリングする(references/code-templates.yaml → readiness_check).

送信前に照合を行う:ユーザーが指定したリクエスト本文にアーキテクチャ固有のフィールド(例: guidance / num_steps / seed / negative_prompt → cosmos-predict2.5; 必須の image_url/video_url コンテンツ項目 → cosmos-rl)が含まれている場合、それらが state[job_id]["network_arch"]。不一致の場合は処理を中止し、確認を求めること — cosmos-predict2.5のリクエストボディをcosmos-rlサービスに送信すると、コンテナ側で4xx/5xxエラーが発生し、ここで検知するよりも診断が困難になる。

6.1 サンプリングパラメータ — 各リクエスト前の必須ユーザープロンプト

リクエスト・ボディを構築する前に、vLLM 形式のサンプリングパラメータについて、ユーザーに明示的に入力を求める必要があります。デフォルト値を黙って適用してはなりません。構造化されたプロンプト(フィールドごとに 1 つの質問)を使用し、以下の要件を満たすようにしてください:

  1. 該当するすべてのフィールドを、その型とデフォルト値とともに一覧表示すること。
  2. ユーザーが任意のフィールドをスキップまたは承諾することで、そのフィールドのデフォルト値を採用できるようにすること — 値の入力は決して必須ではありません。
  3. すべてのフィールドを一度に収集すること。

プロンプト終了後、ユーザーが入力した各値をそのまま適用し、スキップされたフィールドについてはデフォルト値を代入してください。値を勝手に生成したり、黙ってクリッピングを行ったりしてはなりません。

フィールド一覧、デフォルト値、およびアーキテクチャごとの適用範囲: references/request.yaml → chat_completions_request_body (基本サンプリングフィールド: max_tokens, top_p, temperature) および network_arch_constraints. (アーキテクチャごとの上書き設定および guidance/num_steps/seed/negative_prompt for cosmos-predict2.5)。フィールドがアクティブなアーキテクチャに対して「非対応」とマークされている場合、そのフィールドについてプロンプトを表示せず、本文にも含めないでください。

6.2 リクエスト形式

以下の形式で送信してください POST を {BASE_URL}/v1/chat/completions に送信してください。 Content-Type: application/json を指定し、タイムアウトを少なくとも300秒に設定して送信してください。本文はOpenAI互換(vLLMチャット補完)です。完全なフィールドスキーマおよびコンテンツアイテムのシェイプ(text / image_url / video_url)については、 references/request.yaml → chat_completions_request_body を参照してください。また、 code_examples を参照してください。

制約事項:処理されるのは最初のユーザーメッセージのみです。リクエスト本文に秘密値を含めないでください。ネットワークごとの制約があります(例:cosmos-rl では、すべてのリクエストに画像または動画を含める必要があります。cosmos-rl は data: URIを拒否します)は references/request.yaml → network_arch_constraints.

6.3 レスポンスの処理

HTTPステータス 意味 アクション
200 成功 — choices[0].message.content 生成されたテキストがある 結果の読み取り
202 サーバーの初期化中、またはモデルの読み込み中 しばらく待ってから再試行してください
503 初期化に失敗した、モデルの読み込みに失敗した、またはモデルがまだ準備できていない 確認 error.type: model_not_ready → 再試行; initialization_error / model_load_error → 処理を中止し、ログを確認する
400 JSON本文が欠落しているか空です リクエストを修正する
500 推論中に未処理の例外が発生しました コンテナのログを確認してください

202 および 503 の場合、本文には {"error": {"type": "", "message": ""}}が含まれています。詳細は container_response_shapes を参照 references/request.yaml を参照してください。

GitHubで見る
---
name: tao-run-inference-service
description: Start, query, and stop a TAO inference microservice for a specific network architecture by delegating container execution to the appropriate platform skill.
license: Apache-2.0
---

# TAO Inference Microservice

## Instructions

**To start an inference service:**
1. Collect required inputs (Section 1) and resolve the container image (Section 2).
2. Build the job payload and inner command (Sections 3–4.1); use `references/code-templates.yaml` → `job_payload_builder`.
3. Read `skills/platform/<platform>/SKILL.md` and start the container (Section 4.2).
4. Write the service registry and poll readiness (Section 4.3); use `references/code-templates.yaml` → `registry_write.<platform>` and `readiness_check`.

**To send an inference request:**
1. Resolve which service receives the request per Section 6.0 (by `job_id`, by `network_arch`, or by explicit user choice when multiple services run — **never silently default to `"latest"` when more than one service exists**), then read the endpoint from `references/code-templates.yaml` → `request.registry_read` with the resolved `job_id`.
2. **Before building the request body, prompt the user for the vLLM-style sampling parameters (Section 6.1).** Present `max_tokens`, `top_p`, `temperature` (and any per-arch extras) with their defaults; let the user override or skip each one to accept the default. Never silently use defaults.
3. Build and send the body per Section 6.2; handle the response per Section 6.3.

**To stop a service:** Read `references/code-templates.yaml` → `stop.registry_read` to resolve the job_id, read `skills/platform/<platform>/SKILL.md`, then follow Section 5.

**Reference data** (schemas, mappings, valid values — no instructions):
- **`references/service.yaml`** — image mappings, valid `network_arch` names, job payload schema, env var names, secrets classification.
- **`references/request.yaml`** — endpoint definition, request field schema, response shapes, code examples.
- **`references/code-templates.yaml`** — Python templates for payload building, registry writes, readiness checks, and stop/request flows.

---

## Secrets rule (applies to every generated code block in this skill)

**Never ask the user to type a secret value into a prompt.** For every secret value:
1. Tell the user which environment variable to set (e.g. `export HF_TOKEN=...`).
2. Generate code that reads it with `os.environ["VAR_NAME"]` — never hard-code, interpolate, or prompt for the value.

**Secret env vars** (full list in `references/service.yaml` → `secrets_handling`):
`HF_TOKEN`, `WANDB_API_KEY`, `CLEARML_API_ACCESS_KEY`, `CLEARML_API_SECRET_KEY`, `TAO_API_KEY`, `TAO_USER_KEY`.

**Safe to collect in the prompt:** `network_arch`, `model_path`, `num_gpus`, prompt text, `WANDB_*` config URLs, `CLEARML_*_HOST` URLs.

---

## 1. What to collect from the user

| Input | Role |
|--------|------|
| **`network_arch`** | Chooses container image, the per-arch inner command shape (`references/service.yaml` → `container_commands.<network_arch>`), and `neural_network_name` in the job JSON when applicable. Must match a basename in `valid_network_arch_config_basenames` in `references/service.yaml` (e.g. `cosmos-rl`, `cosmos-predict2.5`). |
| **`model_path`** | The trained model checkpoint. Valid forms: `hf_model://<org>/<model>` (HuggingFace Hub — set `HF_TOKEN` for gated models) or a local container filesystem path. Cloud URIs (`s3://`, `gs://`, `az://`) are NOT supported — the inference service has no cloud-storage dependency. Always ask the user; never substitute a placeholder. See `references/service.yaml` → `model_path_protocols`. |
| **`platform`** | Compute platform: `local-docker`, `brev`, `slurm`, or `kubernetes`. |
| **`num_gpus`** | Defaults to **1**; minimum **1** for inference. |

---

## 2. Image resolution

Each `network_arch` has a sidecar config file named `{network_arch}.config.json`. Resolve the container image as follows:

1. Read `{network_arch}.config.json` and take `api_params.image` (e.g. `COSMOS_RL`). This is a key into `docker_image_defaults.mapping` in `references/service.yaml`.
2. Look up that key in the mapping. If the host env var `IMAGE_<KEY>` is set (e.g. `IMAGE_COSMOS_RL`), it overrides the mapped default.
3. The mapped value is normally a dotted key into the repo-root `versions.yaml` manifest (e.g. `tao_toolkit.cosmos_rl`). Resolve it to a concrete `nvcr.io/...` image URI by looking up `versions.yaml` → `images.<group>.<name>`. Absolute URIs pass through unchanged, so an `IMAGE_<KEY>` env-var override that contains a full URI still works. The Python helper for this lives in `references/code-templates.yaml`.
4. If the config file is missing or `api_params.image` is empty, fall back to the `COSMOS_RL` key.

The config file also has `spec_params.inference.model_path` which drives **folder vs file** path semantics: if the value contains the substring `folder`, the container treats the path as a directory.

---

## 3. Environment variables (no callbacks)

Set these in `env_payload` before encoding `env_json`. Do **not** set `TAO_LOGGING_SERVER_URL` or `TAO_ADMIN_KEY`.

**`TAO_EXECUTION_BACKEND`** — must match the platform:

| Platform | `TAO_EXECUTION_BACKEND` value |
|----------|-------------------------------|
| local-docker | `local-docker` |
| brev | `local-docker` |
| slurm | `slurm` |
| kubernetes | `local-k8s` |

**`CLOUD_BASED`** — always `"False"` for this skill (disables callback posting to `TAO_LOGGING_SERVER_URL`).

**GPU env vars** — only needed when the platform skill does not handle GPU injection automatically:
- Tegra / Jetson: `--runtime=nvidia` with `NVIDIA_DRIVER_CAPABILITIES=all` and `NVIDIA_VISIBLE_DEVICES=<ids>`.
- Standard x86 + nvidia-container-toolkit: use Docker `device_requests`. The platform skill handles this.

---

## 4. Executing across platforms

The job payload and inner command (Sections 1–3) are **platform-agnostic**. For each platform, read **`skills/platform/<name>/SKILL.md`** for preflight checks and credentials **before** generating any execution code.

### 4.1 Build the inner command (per arch)

The inner-command shape is **per `network_arch`** — there is no uniform template. Look up the per-arch entry in `references/service.yaml` → `container_commands.<network_arch>`; if not present, the arch is unsupported — stop and ask. Pick the matching sub-block in `references/code-templates.yaml` → `job_payload_builder.<network_arch>`. Prefix the command with `umask 0 &&` and keep it **identical across platforms** (local-docker, brev, slurm, kubernetes).

Common across arches:

- `job_id`: fresh `uuid.uuid4()` — becomes the container name and registry key.
- `image`: resolve per Section 2.
- Secrets (`access_key`, `secret_key`, `HF_TOKEN`, etc.) are read from env vars at runtime — never hard-code, never log or print.

Arch-specific notes (full details in `references/service.yaml` → `container_commands`):

- **`cosmos-rl`** — single `--job '<JOB_JSON>' --docker_env_vars '<ENV_JSON>'` blob; `json.dumps(...)` + `shlex.quote(...)`. `env_payload` carries `TAO_EXECUTION_BACKEND` (per Section 3 table), `TAO_API_JOB_ID`, `CLOUD_BASED=False`. The inference service has no cloud-storage dependency; `HF_TOKEN` is the only cred env var that ever applies (for gated HuggingFace models).
- **`cosmos-predict2.5`** — flag-style `cosmos_predict inference_microservice start ... --port 8080` (no `setup.` prefix; uses `tyro.conf.OmitArgPrefixes`). `--job`/`--docker_env_vars` are **not** accepted. Translate `model_path` to `--checkpoint-path` (local path) or `--model <registered_key>` (`hf_model://`); cloud URIs are rejected. The only cred env var that ever applies is `HF_TOKEN` for gated HuggingFace models. Per-request params (prompt, inference_type, num_output_frames, guidance, seed, num_steps, negative_prompt) go in the request body, not at startup. `TAO_EXECUTION_BACKEND`/`TAO_API_JOB_ID`/`CLOUD_BASED` are unused and may be omitted.

### 4.2 Delegate execution to the platform skill

Read **`skills/platform/<platform>/SKILL.md`** and follow it to start the container.

**Base parameters (all platforms):**

| Parameter | Value |
|-----------|-------|
| `image` | resolved container image (Section 2) |
| `command` | `inner` — the shell string built in Section 4.1 |
| `gpu_count` | `num_gpus` |
| `env_vars` | `env_payload` |
| job / container name | `job_id` — must equal the UUID from 4.1 so the registry can reference it |
| `host_port` *(local-docker, brev)* | host-side port to bind to container port 8080. Default `8080`, but **must be unique per concurrent service** — see the port-allocation rule below. |

**Platform-specific additional inputs:**

| Platform | Additional inputs |
|----------|------------------|
| **local-docker** | None beyond base |
| **brev** | `instance_id` (optional — reuse an existing instance); on multi-credential / multi-workspace accounts also `cloud_cred_id` and `workspace_group_id` for first-create — see `skills/platform/tao-run-on-brev/SKILL.md` |
| **slurm** | `partition` and `account` — check `SLURM_PARTITION`/`SLURM_ACCOUNT` env vars; ask user if unset |
| **kubernetes** | `namespace` (default: `default`); `image_pull_secret` (required for `nvcr.io` images) |

**Port binding (local-docker and brev):** use **direct docker run** (not DockerSDK) so that `-p <host_port>:8080` can be passed and the container name equals `job_id` exactly.

**Port allocation rule (local-docker and brev, REQUIRED for concurrent services):** Before starting a service, read the registry (`/tmp/tao-inf-ms-state.json`) and collect the set of `host_port` values from every existing entry on the same platform (and, for brev, the same `instance_id`). Pick the **lowest free port starting from 8080** that is not in that set — e.g. `host_port = next(p for p in range(8080, 8200) if p not in used_ports)`. The default `8080` only applies when no other service is running. This is what makes "start 3 services, each reachable at a distinct `host_url`" work; without it, services 2 and 3 fail with `bind: address already in use`. SLURM and kubernetes get distinct endpoints from their own platform mechanisms and do not need this step.

### 4.3 After start: service registry and endpoint

Write the service registry immediately after the platform confirms the container is running. The registry (`/tmp/tao-inf-ms-state.json`) is keyed by `job_id`; `"latest"` always points to the most recently started service.

See `references/code-templates.yaml` → `registry_write.<platform>` for the Python template.

| Platform | `host_url` | `platform_job_id` | Extra step before writing |
|----------|-----------|-------------------|--------------------------|
| **local-docker** | `http://localhost:{host_port}` | — | None |
| **brev** | `http://{brev_ip}:{host_port}` | — | `brev ls` → get instance IP (`localhost` is invalid on remote VM) |
| **slurm** | `http://localhost:{host_port}` | SLURM scheduler job ID | Wait until Running; SSH port-forward `localhost:{host_port}→{node}:8080` |
| **kubernetes** | `http://{external_ip}:8080` | k8s job name | `kubectl expose job … --type=LoadBalancer`; wait for external IP |

After writing the registry, print the job_id and URL:

```python
print(f"Inference service started.")
print(f"  Job ID : {job_id}")
print(f"  Arch   : {network_arch}")
print(f"  URL    : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")
```

Then poll for readiness — see `references/code-templates.yaml` → `readiness_check`. The container loads the model in the background; do not send requests before it returns 200.

---

## 5. Stopping the inference service

Ask the user for the `job_id` to stop. If they don't provide one, default to `state["latest"]` and confirm which job_id is being stopped. Read the registry using `references/code-templates.yaml` → `stop.registry_read`, then read **`skills/platform/<platform>/SKILL.md`** and use its cancellation / stop mechanism.

| Platform | Identifier to pass | Extra cleanup |
|----------|--------------------|---------------|
| **local-docker** | `job_id_to_stop` — container name | None |
| **brev** | `job_id_to_stop` — container name | None |
| **slurm** | `entry["platform_job_id"]` — SLURM job ID | `pkill -f "ssh.*-L.*{entry['host_port']}"` |
| **kubernetes** | `entry["platform_job_id"]` — k8s job name | `kubectl delete svc {entry["platform_job_id"]} -n <namespace>` |

where `entry = state[job_id_to_stop]`. After stopping, clean up the registry: `references/code-templates.yaml` → `stop.registry_cleanup`.

---

## 6. Sending inference requests

### 6.0 Resolve which service receives this request (REQUIRED)

Each request must be routed to the **specific** service that runs the matching model. Routing happens by `job_id` — the registry stores `network_arch` per entry, so you can resolve a target by arch when the user names a model instead of a `job_id`. Apply these rules in order:

1. **User provided an explicit `job_id`** → use it. Verify it exists in `state`.
2. **User named a `network_arch`** (e.g. "send this to the cosmos-rl service") → look up matching entries: `candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch]`.
   - Exactly one match → use it.
   - Multiple matches → **prompt the user** with the candidate `job_id`s and their `started_at`; do not auto-pick.
   - No match → stop and tell the user no service for that arch is running.
3. **No `job_id` and no `network_arch`** → count non-`"latest"` entries in `state`:
   - Exactly one running service → use it.
   - Two or more → **do not silently default to `state["latest"]`**. Prompt the user with the full list (`job_id`, `network_arch`, `host_url`) and require an explicit choice. The `"latest"` pointer is a convenience for single-service workflows, not a routing fallback when multiple services coexist.
   - Zero → stop and tell the user to start a service first.

After resolving, read the endpoint from the registry (`references/code-templates.yaml` → `request.registry_read`), passing the resolved `job_id` as `user_provided_job_id`. Confirm to the user: "Sending to job_id=… arch=… url=…". If the service may still be loading, poll readiness first (`references/code-templates.yaml` → `readiness_check`).

**Cross-check before sending:** if the user-supplied request body contains arch-specific fields (e.g. `guidance` / `num_steps` / `seed` / `negative_prompt` → cosmos-predict2.5; required `image_url`/`video_url` content items → cosmos-rl), verify they are consistent with `state[job_id]["network_arch"]`. On mismatch, stop and ask — sending a cosmos-predict2.5 body to a cosmos-rl service will fail at the container with a 4xx/5xx that is harder to diagnose than catching it here.

### 6.1 Sampling parameters — REQUIRED user prompt before each request

Before constructing the request body, you **MUST** explicitly prompt the user for the vLLM-style sampling parameters. Do **not** silently apply defaults. Use a structured prompt, one question per field, that:

1. Lists every applicable field with its **type** and **default value**.
2. Lets the user skip / accept any field to take that field's default — entering a value is never required.
3. Collects all fields in one round.

After the prompt, apply each user-entered value verbatim and substitute the default for any skipped field. Do not invent values or silently clamp.

**Field list, defaults, and per-arch applicability:** `references/request.yaml` → `chat_completions_request_body` (base sampling fields: `max_tokens`, `top_p`, `temperature`) and `network_arch_constraints.<network_arch>` (per-arch overrides and extras such as `guidance`/`num_steps`/`seed`/`negative_prompt` for `cosmos-predict2.5`). If a field is marked unsupported for the active arch, do **not** prompt for it and do **not** include it in the body.

### 6.2 Request format

Send a `POST` to `{BASE_URL}/v1/chat/completions` with `Content-Type: application/json` and a timeout of **at least 300 s**. The body is OpenAI-compatible (vLLM chat completions); see `references/request.yaml` → `chat_completions_request_body` for the full field schema and content-item shapes (text / image_url / video_url), and `code_examples` for ready-to-run Python and curl samples.

**Constraints:** only the first user message is processed. No secret values in request bodies. **Per-network constraints** (e.g. cosmos-rl requires every request to include an image or video; cosmos-rl rejects `data:` URIs) are in `references/request.yaml` → `network_arch_constraints`.

### 6.3 Response handling

| HTTP status | Meaning | Action |
|-------------|---------|--------|
| **200** | Success — `choices[0].message.content` has the generated text | Read result |
| **202** | Server still initializing or model still loading | Retry after a delay |
| **503** | Initialization failed, model load failed, **or model not yet ready** | Inspect `error.type`: `model_not_ready` → retry; `initialization_error` / `model_load_error` → give up and check logs |
| **400** | Missing or empty JSON body | Fix request |
| **500** | Unhandled exception during inference | Check container logs |

For 202 and 503, the body contains `{"error": {"type": "<error_type>", "message": "<reason>"}}`. See `container_response_shapes` in `references/request.yaml` for error type strings.

tao-run-inference-serviceをインストール

スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。

ZIPをダウンロード

リポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service # Copy SKILL.md to your .claude/skills/ directory

コピー コピー
クイックセットアップ: スキルフォルダを .claude/skills/ にコピーしてください。 Claude が自動的にそのスキルを検出して使用します。
リポジトリ NVIDIA/skills

関連スキル

klingai-upgrade-migration
更新された時間 2026年7月3日
Verification &amp; Quality Assurance
更新された時間 2026年6月29日
base44-cli
更新された時間 2026年6月29日
Railway CLI Management
更新された時間 2026年7月2日
OR