tao-run-inference-service
NVIDIA/skills
通过将容器执行任务委托给相应的平台技能,启动、查询和停止针对特定网络架构的 TAO 推理微服务。
...展开全部TAO 推理微服务
使用说明
要启动推理服务:
- 收集所需输入(第 1 节)并解析容器镜像(第 2 节)。
- 构建作业有效载荷和内部命令(第 3–4.1 节);使用
references/code-templates.yaml→job_payload_builder. - Read
skills/platform/并启动容器(第 4.2 节)。/SKILL.md - 写入服务注册表并轮询就绪状态(第 4.3 节);使用
references/code-templates.yaml→registry_write.和readiness_check.
要发送推理请求:
- 根据第 6.0 节确定哪个服务接收该请求(通过
job_id、通过network_arch,或在多个服务运行时由用户显式选择——当存在多个服务时,切勿默认默认使用"latest"),然后从references/code-templates.yaml→request.registry_read中读取端点,使用已确定的job_id. - 在构建请求正文之前,应提示用户输入 vLLM 风格的采样参数(第 6.1 节)。将
max_tokens,top_p,temperature(以及任何针对特定架构的附加参数)及其默认值;允许用户覆盖或跳过每个参数以接受默认值。切勿在用户不知情的情况下使用默认值。 - 按照第 6.2 节构建并发送请求主体;按照第 6.3 节处理响应。
要停止服务:读取 references/code-templates.yaml → stop.registry_read 以解析 job_id,读取 skills/platform/,然后按照第 5 节操作。
参考数据(模式、映射、有效值——不含说明):
references/service.yaml— 镜像映射、有效network_arch名称、作业有效载荷模式、环境变量名称、密钥分类。references/request.yaml— 端点定义、请求字段模式、响应结构、代码示例。references/code-templates.yaml— 用于有效载荷构建、注册表写入、就绪性检查以及停止/请求流的 Python 模板。
密钥规则(适用于该技能中生成的每个代码块)
切勿要求用户在提示中输入机密值。对于每个机密值:
- 告知用户应设置哪个环境变量(例如
export HF_TOKEN=...). - 生成使用以下方式读取该值的代码:
os.environ["VAR_NAME"]— 切勿硬编码、插值或通过提示符获取该值。
机密环境变量(完整列表见 references/service.yaml → secrets_handling):
HF_TOKEN, WANDB_API_KEY, CLEARML_API_ACCESS_KEY, CLEARML_API_SECRET_KEY, TAO_API_KEY, TAO_USER_KEY.
可在提示中安全收集的内容: network_arch, model_path, num_gpus、提示文本、 WANDB_* 配置 URL、 CLEARML_*_HOST URL中。
1. 需要从用户处收集哪些信息
| 输入 | 角色 |
|---|---|
network_arch |
选择容器镜像、针对各架构的内部命令格式(references/service.yaml → container_commands.)以及 neural_network_name 在作业 JSON 中(如适用)。必须与 valid_network_arch_config_basenames 中 references/service.yaml 中的基名(例如 cosmos-rl, cosmos-predict2.5). |
model_path |
已训练模型的检查点。有效格式: hf_model:// (HuggingFace Hub — 设置 HF_TOKEN 用于门控模型)或本地容器文件系统路径。云 URI(s3://, gs://, az://) 不受支持——推理服务不依赖云存储。请务必征得用户同意;切勿使用占位符替代。参见 references/service.yaml → model_path_protocols. |
platform |
计算平台: local-docker, brev, slurm,或 kubernetes. |
num_gpus |
默认值为 1;推理的最小值为 1。 |
2. 图像分辨率
每个 network_arch 都有一个名为 {network_arch}.config.json。请按以下方式解析容器镜像:
- 阅读
{network_arch}.config.json并获取api_params.image(例如COSMOS_RL)。这是docker_image_defaults.mappinginreferences/service.yaml. - 在映射中查找该键。如果主机环境变量
IMAGE_已设置(例如IMAGE_COSMOS_RL),则该值将覆盖映射的默认值。 - 映射值通常是指向仓库根目录
versions.yaml清单中的点分隔键(例如tao_toolkit.cosmos_rl)。将其解析为具体的nvcr.io/...镜像URI,方法是查询versions.yaml→images.。绝对URI会原样传递,因此包含完整URI的. IMAGE_包含完整 URI 的环境变量覆盖仍可正常工作。相关的 Python 辅助函数位于references/code-templates.yaml. - 如果配置文件缺失或
api_params.image为空,则回退到COSMOS_RL键。
配置文件中还包含 spec_params.inference.model_path ,该键用于区分文件夹路径与文件路径的语义:如果值中包含子字符串 folder,容器将该路径视为目录。
3. 环境变量(无回调)
请在 env_payload 编码前 env_json之前进行设置。请勿设置 TAO_LOGGING_SERVER_URL 或 TAO_ADMIN_KEY.
TAO_EXECUTION_BACKEND ——必须与平台匹配:
| 平台 | TAO_EXECUTION_BACKEND 值 |
|---|---|
| local-docker | local-docker |
| brev | local-docker |
| slurm | slurm |
| Kubernetes | local-k8s |
CLOUD_BASED — 始终 "False" 针对此技能(禁用向 TAO_LOGGING_SERVER_URL).
GPU 环境变量 — 仅当平台技能无法自动处理 GPU 注入时才需要:
- Tegra / Jetson:
--runtime=nvidia配合NVIDIA_DRIVER_CAPABILITIES=all以及NVIDIA_VISIBLE_DEVICES=. - 标准 x86 + nvidia-container-toolkit:请使用 Docker
device_requests。平台技能会自动处理此操作。
4. 跨平台执行
作业有效载荷和内部命令(第 1–3 节)与平台无关。对于每个平台,在生成任何执行代码之前,请阅读 skills/platform/ 以了解预检查和凭据信息。
4.1 构建内部命令(按架构)
内部命令的结构应遵循 network_arch —— 没有统一的模板。请在 references/service.yaml → container_commands.中查找该架构的条目;若不存在,则表示该架构不受支持——请停止操作并咨询。从 references/code-templates.yaml → job_payload_builder.中选择匹配的子模块。在命令前添加前缀 umask 0 && ,并确保其在所有平台(local-docker、brev、slurm、kubernetes)上保持一致。
各架构通用的部分:
job_id: freshuuid.uuid4()——将作为容器名称和注册表键。image: 按第 2 节进行解析。- 密钥(
access_key,secret_key,HF_TOKEN等)在运行时从环境变量中读取——切勿硬编码,切勿记录或打印。
架构特定说明(完整细节见 references/service.yaml → container_commands):
cosmos-rl— 单个--job '二进制数据块;' --docker_env_vars ' ' json.dumps(...)+shlex.quote(...).env_payload包含TAO_EXECUTION_BACKEND(参见第 3 节表格),TAO_API_JOB_ID,CLOUD_BASED=False。推理服务不依赖云存储;HF_TOKEN是唯一适用的凭证环境变量(适用于受限访问的 HuggingFace 模型)。cosmos-predict2.5— 标志式cosmos_predict inference_microservice start ... --port 8080(无setup.前缀;不接受tyro.conf.OmitArgPrefixes).--job/--docker_env_vars不被接受。翻译model_path为--checkpoint-path(本地路径) 或--model(hf_model://);云端URI将被拒绝。唯一适用的凭证环境变量是HF_TOKEN仅适用于受限访问的 HuggingFace 模型。每次请求的参数(prompt、inference_type、num_output_frames、guidance、seed、num_steps、negative_prompt)应放入请求正文中,而非在启动时设置。TAO_EXECUTION_BACKEND/TAO_API_JOB_ID/CLOUD_BASED未被使用,可省略。
4.2 将执行委托给平台技能
请阅读《skills/platform/》并按照其中说明启动容器。
基础参数(所有平台):
| 参数 | 值 |
|---|---|
image |
已解析的容器镜像(第 2 节) |
command |
inner — 第 4.1 节中构建的 shell 字符串 |
gpu_count |
num_gpus |
env_vars |
env_payload |
| 任务/容器名称 | job_id — 必须与第 4.1 节中的 UUID 一致,以便注册表能够引用它 |
host_port (本地 Docker,简写) |
用于将主机端口绑定至容器端口 8080 的端口。默认值 8080,但每个并发服务必须唯一 — 请参阅下方的端口分配规则。 |
特定于平台的附加输入:
| 平台 | 附加输入 |
|---|---|
| local-docker | 除基础参数外无其他 |
| brev | instance_id (可选 — 复用现有实例);对于支持多凭证/多工作区的账户,还需指定 cloud_cred_id 以及 workspace_group_id 首次创建时 — 参见 skills/platform/tao-run-on-brev/SKILL.md |
| slurm | partition 和 account — 检查 SLURM_PARTITION/SLURM_ACCOUNT 环境变量;若未设置则询问用户 |
| Kubernetes | namespace (默认: default); image_pull_secret ( nvcr.io 镜像) |
端口绑定(local-docker 和 brev):使用直接的 docker run(而非 DockerSDK),以便 -p 参数能够被传递,且容器名称与 job_id 完全一致。
端口分配规则(local-docker 和 brev,并发服务必须遵守):在启动服务之前,读取注册表(/tmp/tao-inf-ms-state.json)并收集 host_port 值集合。对于 brev,还需从同一 instance_id)。从8080开始,选择该集合中不包含的最低可用端口——例如 host_port = next(p for p in range(8080, 8200) if p not in used_ports)。默认 8080 仅在没有其他服务运行时才适用。这正是“启动 3 个服务,每个服务均可通过不同 host_url”这一机制得以实现;若无此机制,服务2和3将因 bind: address already in use。SLURM 和 Kubernetes 通过各自的平台机制获取独立的端点,因此无需此步骤。
4.3 启动后:服务注册表与端点
在平台确认容器正在运行后,立即写入服务注册表。该注册表(/tmp/tao-inf-ms-state.json)以 job_id; "latest" 始终指向最近启动的服务。
参见 references/code-templates.yaml → registry_write. 中的 Python 模板。
| 平台 | host_url |
platform_job_id |
编写前需执行的额外步骤 |
|---|---|---|---|
| local-docker | http://localhost:{host_port} |
— | 无 |
| brev | http://{brev_ip}:{host_port} |
— | brev ls → 获取实例 IP(localhost 在远程虚拟机上无效) |
| slurm | http://localhost:{host_port} |
SLURM调度器作业ID | 等待进入“运行”状态;SSH 端口转发 localhost:{host_port}→{node}:8080 |
| Kubernetes | http://{external_ip}:8080 |
k8s 作业名称 | kubectl expose job … --type=LoadBalancer; 等待获取外部 IP |
写入注册表后,打印 job_id 和 URL:
print(f"Inference service started.")
print(f" Job ID : {job_id}")
print(f" Arch : {network_arch}")
print(f" URL : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")
然后轮询就绪状态 — 参见 references/code-templates.yaml → readiness_check。容器会在后台加载模型;在返回 200 状态码之前,请勿发送请求。
5. 停止推理服务
向用户询问 job_id 以停止服务。若用户未提供,则默认使用 state["latest"] ,并确认要停止的 job_id。使用 references/code-templates.yaml → stop.registry_read读取注册表,随后读取skills/platform/并使用其取消/停止机制。
| 平台 | 标识符,用于传递 | 额外清理 |
|---|---|---|
| local-docker | job_id_to_stop — 容器名称 |
无 |
| brev | job_id_to_stop — 容器名称 |
无 |
| slurm | entry["platform_job_id"] — SLURM 作业 ID |
pkill -f "ssh.*-L.*{entry['host_port']}" |
| kubernetes | entry["platform_job_id"] — k8s 作业名称 |
kubectl delete svc {entry["platform_job_id"]} -n |
其中 entry = state[job_id_to_stop]。停止后,清理注册表: references/code-templates.yaml → stop.registry_cleanup.
6. 发送推理请求
6.0 确定哪个服务接收此请求(必填)
每个请求都必须路由到运行匹配模型的特定服务。路由通过 job_id — 注册表为 network_arch 每个条目,因此当用户指定模型名称而非 job_id。按以下顺序应用这些规则:
- 用户提供了显式的
job_id→ 使用该服务。验证该服务是否存在于state. - 用户指定了
network_arch(例如“将此请求发送至cosmos-rl服务”)→ 查找匹配的条目:candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch].- 仅有一个匹配项 → 使用该项。
- 有多个匹配项 → 向用户提示候选
job_id及其started_at;不要自动选择。 - 无匹配项 → 停止并告知用户该架构下没有正在运行的服务。
- 既没有
job_id也没有network_arch→ 统计非"latest"条目state:- 恰好有一个正在运行的服务 → 使用该服务。
- 两个或更多 → 不要默认静默使用
state["latest"]。向用户提示完整列表(job_id,network_arch,host_url),并要求用户明确选择。该"latest"指针仅为单服务工作流提供便利,而非多服务共存时的路由后备方案。 - 零 → 停止并提示用户先启动一个服务。
解析完成后,从注册表中读取端点(references/code-templates.yaml → request.registry_read),并将解析后的 job_id 作为 user_provided_job_id。向用户确认:“正在发送至 job_id=… arch=… url=…”。如果服务可能仍在加载中,请先轮询就绪状态(references/code-templates.yaml → readiness_check).
发送前进行交叉核对:如果用户提供的请求正文包含架构特定字段(例如 guidance / num_steps / seed / negative_prompt → cosmos-predict2.5;必需的 image_url/video_url 内容项 → cosmos-rl),请验证它们是否与 state[job_id]["network_arch"]。若不匹配,则停止并询问——将 cosmos-predict2.5 请求体发送至 cosmos-rl 服务将在容器端因 4xx/5xx 错误而失败,此时的故障排查难度远高于在此处捕获该错误。
6.1 采样参数 — 每次请求前必须向用户提示
在构建请求正文之前,您必须明确提示用户输入 vLLM 风格的采样参数。切勿在用户不知情的情况下应用默认值。请使用结构化提示,每个字段一个问题,该提示应:
- 列出所有适用字段及其类型和默认值。
- 允许用户跳过或接受任何字段,以采用该字段的默认值——绝不强制要求输入值。
- 在同一轮内收集所有字段。
提示完成后,应原样应用用户输入的每个值,并对任何被跳过的字段使用其默认值。切勿自行生成值或默认进行限制。
字段列表、默认值及各架构适用性: references/request.yaml → chat_completions_request_body (基础采样字段: max_tokens, top_p, temperature) 以及 network_arch_constraints. (针对各架构的覆盖项和额外项,例如 guidance/num_steps/seed/negative_prompt for cosmos-predict2.5)。如果某个字段被标记为当前架构不支持,则不要提示该字段,也不要将其包含在请求主体中。
6.2 请求格式
请发送一个 POST 发送至 {BASE_URL}/v1/chat/completions , Content-Type: application/json ,并设置至少 300 秒的超时时间。正文应兼容 OpenAI 格式(vLLM 聊天补全);请参阅 references/request.yaml → chat_completions_request_body 以获取完整的字段模式和内容项结构(text / image_url / video_url),并 code_examples 可获取可直接运行的 Python 和 curl 示例。
限制:仅处理用户的第一个消息。请求正文中不得包含机密值。各网络有特定限制(例如,cosmos-rl 要求每个请求都包含图片或视频;cosmos-rl 会拒绝 data: URI)。 references/request.yaml → network_arch_constraints.
6.3 响应处理
| HTTP 状态码 | 含义 | 操作 |
|---|---|---|
| 200 | 成功 — choices[0].message.content 包含生成的文本 |
读取结果 |
| 202 | 服务器仍在初始化或模型仍在加载中 | 延迟后重试 |
| 503 | 初始化失败、模型加载失败或模型尚未准备就绪 | 检查 error.type: model_not_ready → 重试; initialization_error / model_load_error → 放弃并检查日志 |
| 400 | JSON 请求体缺失或为空 | 修复请求 |
| 500 | 推理过程中发生未处理的异常 | 检查容器日志 |
对于 202 和 503 状态码,请求体包含 {"error": {"type": "。参见 container_response_shapes 在 references/request.yaml 中查看错误类型字符串。
---
name: tao-run-inference-service
description: Start, query, and stop a TAO inference microservice for a specific network architecture by delegating container execution to the appropriate platform skill.
license: Apache-2.0
---
# TAO Inference Microservice
## Instructions
**To start an inference service:**
1. Collect required inputs (Section 1) and resolve the container image (Section 2).
2. Build the job payload and inner command (Sections 3–4.1); use `references/code-templates.yaml` → `job_payload_builder`.
3. Read `skills/platform/<platform>/SKILL.md` and start the container (Section 4.2).
4. Write the service registry and poll readiness (Section 4.3); use `references/code-templates.yaml` → `registry_write.<platform>` and `readiness_check`.
**To send an inference request:**
1. Resolve which service receives the request per Section 6.0 (by `job_id`, by `network_arch`, or by explicit user choice when multiple services run — **never silently default to `"latest"` when more than one service exists**), then read the endpoint from `references/code-templates.yaml` → `request.registry_read` with the resolved `job_id`.
2. **Before building the request body, prompt the user for the vLLM-style sampling parameters (Section 6.1).** Present `max_tokens`, `top_p`, `temperature` (and any per-arch extras) with their defaults; let the user override or skip each one to accept the default. Never silently use defaults.
3. Build and send the body per Section 6.2; handle the response per Section 6.3.
**To stop a service:** Read `references/code-templates.yaml` → `stop.registry_read` to resolve the job_id, read `skills/platform/<platform>/SKILL.md`, then follow Section 5.
**Reference data** (schemas, mappings, valid values — no instructions):
- **`references/service.yaml`** — image mappings, valid `network_arch` names, job payload schema, env var names, secrets classification.
- **`references/request.yaml`** — endpoint definition, request field schema, response shapes, code examples.
- **`references/code-templates.yaml`** — Python templates for payload building, registry writes, readiness checks, and stop/request flows.
---
## Secrets rule (applies to every generated code block in this skill)
**Never ask the user to type a secret value into a prompt.** For every secret value:
1. Tell the user which environment variable to set (e.g. `export HF_TOKEN=...`).
2. Generate code that reads it with `os.environ["VAR_NAME"]` — never hard-code, interpolate, or prompt for the value.
**Secret env vars** (full list in `references/service.yaml` → `secrets_handling`):
`HF_TOKEN`, `WANDB_API_KEY`, `CLEARML_API_ACCESS_KEY`, `CLEARML_API_SECRET_KEY`, `TAO_API_KEY`, `TAO_USER_KEY`.
**Safe to collect in the prompt:** `network_arch`, `model_path`, `num_gpus`, prompt text, `WANDB_*` config URLs, `CLEARML_*_HOST` URLs.
---
## 1. What to collect from the user
| Input | Role |
|--------|------|
| **`network_arch`** | Chooses container image, the per-arch inner command shape (`references/service.yaml` → `container_commands.<network_arch>`), and `neural_network_name` in the job JSON when applicable. Must match a basename in `valid_network_arch_config_basenames` in `references/service.yaml` (e.g. `cosmos-rl`, `cosmos-predict2.5`). |
| **`model_path`** | The trained model checkpoint. Valid forms: `hf_model://<org>/<model>` (HuggingFace Hub — set `HF_TOKEN` for gated models) or a local container filesystem path. Cloud URIs (`s3://`, `gs://`, `az://`) are NOT supported — the inference service has no cloud-storage dependency. Always ask the user; never substitute a placeholder. See `references/service.yaml` → `model_path_protocols`. |
| **`platform`** | Compute platform: `local-docker`, `brev`, `slurm`, or `kubernetes`. |
| **`num_gpus`** | Defaults to **1**; minimum **1** for inference. |
---
## 2. Image resolution
Each `network_arch` has a sidecar config file named `{network_arch}.config.json`. Resolve the container image as follows:
1. Read `{network_arch}.config.json` and take `api_params.image` (e.g. `COSMOS_RL`). This is a key into `docker_image_defaults.mapping` in `references/service.yaml`.
2. Look up that key in the mapping. If the host env var `IMAGE_<KEY>` is set (e.g. `IMAGE_COSMOS_RL`), it overrides the mapped default.
3. The mapped value is normally a dotted key into the repo-root `versions.yaml` manifest (e.g. `tao_toolkit.cosmos_rl`). Resolve it to a concrete `nvcr.io/...` image URI by looking up `versions.yaml` → `images.<group>.<name>`. Absolute URIs pass through unchanged, so an `IMAGE_<KEY>` env-var override that contains a full URI still works. The Python helper for this lives in `references/code-templates.yaml`.
4. If the config file is missing or `api_params.image` is empty, fall back to the `COSMOS_RL` key.
The config file also has `spec_params.inference.model_path` which drives **folder vs file** path semantics: if the value contains the substring `folder`, the container treats the path as a directory.
---
## 3. Environment variables (no callbacks)
Set these in `env_payload` before encoding `env_json`. Do **not** set `TAO_LOGGING_SERVER_URL` or `TAO_ADMIN_KEY`.
**`TAO_EXECUTION_BACKEND`** — must match the platform:
| Platform | `TAO_EXECUTION_BACKEND` value |
|----------|-------------------------------|
| local-docker | `local-docker` |
| brev | `local-docker` |
| slurm | `slurm` |
| kubernetes | `local-k8s` |
**`CLOUD_BASED`** — always `"False"` for this skill (disables callback posting to `TAO_LOGGING_SERVER_URL`).
**GPU env vars** — only needed when the platform skill does not handle GPU injection automatically:
- Tegra / Jetson: `--runtime=nvidia` with `NVIDIA_DRIVER_CAPABILITIES=all` and `NVIDIA_VISIBLE_DEVICES=<ids>`.
- Standard x86 + nvidia-container-toolkit: use Docker `device_requests`. The platform skill handles this.
---
## 4. Executing across platforms
The job payload and inner command (Sections 1–3) are **platform-agnostic**. For each platform, read **`skills/platform/<name>/SKILL.md`** for preflight checks and credentials **before** generating any execution code.
### 4.1 Build the inner command (per arch)
The inner-command shape is **per `network_arch`** — there is no uniform template. Look up the per-arch entry in `references/service.yaml` → `container_commands.<network_arch>`; if not present, the arch is unsupported — stop and ask. Pick the matching sub-block in `references/code-templates.yaml` → `job_payload_builder.<network_arch>`. Prefix the command with `umask 0 &&` and keep it **identical across platforms** (local-docker, brev, slurm, kubernetes).
Common across arches:
- `job_id`: fresh `uuid.uuid4()` — becomes the container name and registry key.
- `image`: resolve per Section 2.
- Secrets (`access_key`, `secret_key`, `HF_TOKEN`, etc.) are read from env vars at runtime — never hard-code, never log or print.
Arch-specific notes (full details in `references/service.yaml` → `container_commands`):
- **`cosmos-rl`** — single `--job '<JOB_JSON>' --docker_env_vars '<ENV_JSON>'` blob; `json.dumps(...)` + `shlex.quote(...)`. `env_payload` carries `TAO_EXECUTION_BACKEND` (per Section 3 table), `TAO_API_JOB_ID`, `CLOUD_BASED=False`. The inference service has no cloud-storage dependency; `HF_TOKEN` is the only cred env var that ever applies (for gated HuggingFace models).
- **`cosmos-predict2.5`** — flag-style `cosmos_predict inference_microservice start ... --port 8080` (no `setup.` prefix; uses `tyro.conf.OmitArgPrefixes`). `--job`/`--docker_env_vars` are **not** accepted. Translate `model_path` to `--checkpoint-path` (local path) or `--model <registered_key>` (`hf_model://`); cloud URIs are rejected. The only cred env var that ever applies is `HF_TOKEN` for gated HuggingFace models. Per-request params (prompt, inference_type, num_output_frames, guidance, seed, num_steps, negative_prompt) go in the request body, not at startup. `TAO_EXECUTION_BACKEND`/`TAO_API_JOB_ID`/`CLOUD_BASED` are unused and may be omitted.
### 4.2 Delegate execution to the platform skill
Read **`skills/platform/<platform>/SKILL.md`** and follow it to start the container.
**Base parameters (all platforms):**
| Parameter | Value |
|-----------|-------|
| `image` | resolved container image (Section 2) |
| `command` | `inner` — the shell string built in Section 4.1 |
| `gpu_count` | `num_gpus` |
| `env_vars` | `env_payload` |
| job / container name | `job_id` — must equal the UUID from 4.1 so the registry can reference it |
| `host_port` *(local-docker, brev)* | host-side port to bind to container port 8080. Default `8080`, but **must be unique per concurrent service** — see the port-allocation rule below. |
**Platform-specific additional inputs:**
| Platform | Additional inputs |
|----------|------------------|
| **local-docker** | None beyond base |
| **brev** | `instance_id` (optional — reuse an existing instance); on multi-credential / multi-workspace accounts also `cloud_cred_id` and `workspace_group_id` for first-create — see `skills/platform/tao-run-on-brev/SKILL.md` |
| **slurm** | `partition` and `account` — check `SLURM_PARTITION`/`SLURM_ACCOUNT` env vars; ask user if unset |
| **kubernetes** | `namespace` (default: `default`); `image_pull_secret` (required for `nvcr.io` images) |
**Port binding (local-docker and brev):** use **direct docker run** (not DockerSDK) so that `-p <host_port>:8080` can be passed and the container name equals `job_id` exactly.
**Port allocation rule (local-docker and brev, REQUIRED for concurrent services):** Before starting a service, read the registry (`/tmp/tao-inf-ms-state.json`) and collect the set of `host_port` values from every existing entry on the same platform (and, for brev, the same `instance_id`). Pick the **lowest free port starting from 8080** that is not in that set — e.g. `host_port = next(p for p in range(8080, 8200) if p not in used_ports)`. The default `8080` only applies when no other service is running. This is what makes "start 3 services, each reachable at a distinct `host_url`" work; without it, services 2 and 3 fail with `bind: address already in use`. SLURM and kubernetes get distinct endpoints from their own platform mechanisms and do not need this step.
### 4.3 After start: service registry and endpoint
Write the service registry immediately after the platform confirms the container is running. The registry (`/tmp/tao-inf-ms-state.json`) is keyed by `job_id`; `"latest"` always points to the most recently started service.
See `references/code-templates.yaml` → `registry_write.<platform>` for the Python template.
| Platform | `host_url` | `platform_job_id` | Extra step before writing |
|----------|-----------|-------------------|--------------------------|
| **local-docker** | `http://localhost:{host_port}` | — | None |
| **brev** | `http://{brev_ip}:{host_port}` | — | `brev ls` → get instance IP (`localhost` is invalid on remote VM) |
| **slurm** | `http://localhost:{host_port}` | SLURM scheduler job ID | Wait until Running; SSH port-forward `localhost:{host_port}→{node}:8080` |
| **kubernetes** | `http://{external_ip}:8080` | k8s job name | `kubectl expose job … --type=LoadBalancer`; wait for external IP |
After writing the registry, print the job_id and URL:
```python
print(f"Inference service started.")
print(f" Job ID : {job_id}")
print(f" Arch : {network_arch}")
print(f" URL : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")
```
Then poll for readiness — see `references/code-templates.yaml` → `readiness_check`. The container loads the model in the background; do not send requests before it returns 200.
---
## 5. Stopping the inference service
Ask the user for the `job_id` to stop. If they don't provide one, default to `state["latest"]` and confirm which job_id is being stopped. Read the registry using `references/code-templates.yaml` → `stop.registry_read`, then read **`skills/platform/<platform>/SKILL.md`** and use its cancellation / stop mechanism.
| Platform | Identifier to pass | Extra cleanup |
|----------|--------------------|---------------|
| **local-docker** | `job_id_to_stop` — container name | None |
| **brev** | `job_id_to_stop` — container name | None |
| **slurm** | `entry["platform_job_id"]` — SLURM job ID | `pkill -f "ssh.*-L.*{entry['host_port']}"` |
| **kubernetes** | `entry["platform_job_id"]` — k8s job name | `kubectl delete svc {entry["platform_job_id"]} -n <namespace>` |
where `entry = state[job_id_to_stop]`. After stopping, clean up the registry: `references/code-templates.yaml` → `stop.registry_cleanup`.
---
## 6. Sending inference requests
### 6.0 Resolve which service receives this request (REQUIRED)
Each request must be routed to the **specific** service that runs the matching model. Routing happens by `job_id` — the registry stores `network_arch` per entry, so you can resolve a target by arch when the user names a model instead of a `job_id`. Apply these rules in order:
1. **User provided an explicit `job_id`** → use it. Verify it exists in `state`.
2. **User named a `network_arch`** (e.g. "send this to the cosmos-rl service") → look up matching entries: `candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch]`.
- Exactly one match → use it.
- Multiple matches → **prompt the user** with the candidate `job_id`s and their `started_at`; do not auto-pick.
- No match → stop and tell the user no service for that arch is running.
3. **No `job_id` and no `network_arch`** → count non-`"latest"` entries in `state`:
- Exactly one running service → use it.
- Two or more → **do not silently default to `state["latest"]`**. Prompt the user with the full list (`job_id`, `network_arch`, `host_url`) and require an explicit choice. The `"latest"` pointer is a convenience for single-service workflows, not a routing fallback when multiple services coexist.
- Zero → stop and tell the user to start a service first.
After resolving, read the endpoint from the registry (`references/code-templates.yaml` → `request.registry_read`), passing the resolved `job_id` as `user_provided_job_id`. Confirm to the user: "Sending to job_id=… arch=… url=…". If the service may still be loading, poll readiness first (`references/code-templates.yaml` → `readiness_check`).
**Cross-check before sending:** if the user-supplied request body contains arch-specific fields (e.g. `guidance` / `num_steps` / `seed` / `negative_prompt` → cosmos-predict2.5; required `image_url`/`video_url` content items → cosmos-rl), verify they are consistent with `state[job_id]["network_arch"]`. On mismatch, stop and ask — sending a cosmos-predict2.5 body to a cosmos-rl service will fail at the container with a 4xx/5xx that is harder to diagnose than catching it here.
### 6.1 Sampling parameters — REQUIRED user prompt before each request
Before constructing the request body, you **MUST** explicitly prompt the user for the vLLM-style sampling parameters. Do **not** silently apply defaults. Use a structured prompt, one question per field, that:
1. Lists every applicable field with its **type** and **default value**.
2. Lets the user skip / accept any field to take that field's default — entering a value is never required.
3. Collects all fields in one round.
After the prompt, apply each user-entered value verbatim and substitute the default for any skipped field. Do not invent values or silently clamp.
**Field list, defaults, and per-arch applicability:** `references/request.yaml` → `chat_completions_request_body` (base sampling fields: `max_tokens`, `top_p`, `temperature`) and `network_arch_constraints.<network_arch>` (per-arch overrides and extras such as `guidance`/`num_steps`/`seed`/`negative_prompt` for `cosmos-predict2.5`). If a field is marked unsupported for the active arch, do **not** prompt for it and do **not** include it in the body.
### 6.2 Request format
Send a `POST` to `{BASE_URL}/v1/chat/completions` with `Content-Type: application/json` and a timeout of **at least 300 s**. The body is OpenAI-compatible (vLLM chat completions); see `references/request.yaml` → `chat_completions_request_body` for the full field schema and content-item shapes (text / image_url / video_url), and `code_examples` for ready-to-run Python and curl samples.
**Constraints:** only the first user message is processed. No secret values in request bodies. **Per-network constraints** (e.g. cosmos-rl requires every request to include an image or video; cosmos-rl rejects `data:` URIs) are in `references/request.yaml` → `network_arch_constraints`.
### 6.3 Response handling
| HTTP status | Meaning | Action |
|-------------|---------|--------|
| **200** | Success — `choices[0].message.content` has the generated text | Read result |
| **202** | Server still initializing or model still loading | Retry after a delay |
| **503** | Initialization failed, model load failed, **or model not yet ready** | Inspect `error.type`: `model_not_ready` → retry; `initialization_error` / `model_load_error` → give up and check logs |
| **400** | Missing or empty JSON body | Fix request |
| **500** | Unhandled exception during inference | Check container logs |
For 202 and 503, the body contains `{"error": {"type": "<error_type>", "message": "<reason>"}}`. See `container_response_shapes` in `references/request.yaml` for error type strings.





首页
