vss-deploy-detection-tracking-2d
NVIDIA/skills
部署、调试和运行 RTVI-CV 2D 检测/跟踪微服务,并调用其 REST API 进行流管理、健康检查和指标监控。
...展开全部目的
部署、调试和运行 RTVI-CV 检测/跟踪 2D 微服务,并调用其 REST API。
先决条件
- 可通过
$HOST_IP访问的活跃 VSS 部署(参见vss-deploy-profile和references/)。 $NGC_CLI_API_KEY和$NVIDIA_API_KEY中需包含 NGC 凭据,用于拉取任何镜像。- 调用方需具备
curl、jq和 Docker。
操作指南
请遵循下方的路由表和分步工作流。每个以“工作流”、“快速入门”或“流程”结尾的章节均应从上至下依次执行。详细参考资料位于references/ 目录中,辅助脚本位于scripts/目录中——当技能通过名称指向某个脚本时,请通过run_script调用它们。
示例
已验证的端到端示例保存在evals/目录下(每个*.json清单文件包含一个可运行的场景),并内嵌在下文各工作流对应的curl代码块中。运行 Tier-3 评估时,请使用nv-base validate命令来重现这些示例。
限制
- 需要已部署相应的 VSS 配置文件/微服务,且调用方可访问该资源。
- NGC 托管的模型和 NIM 可能受速率限制、GPU 内存要求以及许可限制的约束。
- 并发数、GPU 内存和存储限制取决于主机硬件以及配置文件的 compose 文件。
故障排除
- 错误:REST 调用返回“连接被拒绝”。原因:目标微服务未运行。解决方案:探测
/docs或/health;通过vss-deploy-profile或相应的vss-deploy-*技能重新部署。 - 错误:NGC 拉取返回 HTTP 401/403 错误。原因:缺少或过期的
NGC_CLI_API_KEY。解决方案:执行 `docker login nvcr.io`,并在重试前重新导出该密钥。 - 错误:容器内存不足(OOM)或模型加载失败。原因:所选配置文件所需的 GPU 内存不足。解决方案:切换至较小规格,或通过 `
docker compose down` 释放 GPU 资源。
RTVI-CV — 检测与跟踪(统一技能)
实时视频智能计算机视觉(RTVI-CV)微服务的统一技能。一个技能包含两个操作界面:
- 在本地部署/运行/调试/卸载RTVI-CV 容器 → 参见
references/deploy-vss-detection-tracking-2d.md - 在运行中的实例上调用 RTVI-CV REST API(流、健康状态、指标、嵌入向量)→ 参见
references/usage-vss-detection-tracking-2d.md
服务:
rtvi-cv(metropolis_perception_app) 镜像:nvcr.io/— 部署时由用户提供 REST 端口:/ : 9000(/api/v1—/live、/ready、/startup、/metrics、/stream/add、/stream/remove、embeddings) 硬件:x86/aarch64 独立显卡 (T4、A100、L40、H100、B200、RTX)、SBSA(Spark、Grace-Hopper)、Jetson(Thor、Orin、Xavier)
操作路由 — 每次调用仅选择一次
| 用户意图(示例表述) | 流程 | 加载此参考 |
|---|---|---|
部署 rtvi-cv warehouse 2d ,运行 rtvicv warehouse-3d(4 个数据流),启动 smartcity gdino,启动感知应用,加载 sparse4d |
部署 | references/deploy-vss-detection-tracking-2d.md |
停止 rtvi-cv,释放资源,终止感知容器,清理 rtvicv-perception-docker |
拆卸(由部署文档处理 → “模式选择”) | references/deploy-vss-detection-tracking-2d.md+references/teardown-flow.md |
检查 rtvi-cv 日志, 诊断 rtvi-cv 崩溃问题 ,排查健康检查失败问题,以及rtvi-cv 无法启动的问题 |
DEBUG | references/deploy-vss-detection-tracking-2d.md+references/troubleshooting.md |
添加流、移除摄像头、列出流、健康检查、rtvi-cv是否就绪、获取指标、帧率是多少、检查GPU使用率、生成文本嵌入、调用rtvi-cv API |
API 使用方法 | references/usage-vss-detection-tracking-2d.md+references/api-reference.md |
选择规则:将用户的表述与上表进行匹配,并立即加载相应的参考文件。请勿混淆工作流——“部署”假设尚未有正在运行的容器;“API 使用”则假设容器已在http:// 上运行。
如果意图确实模棱两可(例如,用户仅说“我想使用 rtvi-cv”),则提出一个AskQuestion:是部署新实例,还是调用已运行的实例?
文件结构
vss-deploy-detection-tracking-2d/
├── SKILL.md # 本文件(路由 + 合约)
├── assets/ # 数据文件(deploy-defaults.yml — 标签/引用/路径/GPU 的唯一可信来源)
├── evals/ # Tier-3 评估清单(deploy-evals.json、usage-evals.json)
├── scripts/ # 23 个 bash + python 辅助脚本(完整列表请参见 `scripts/`)
└── references/ # 工作流运行手册(部署 / API 使用 / 清理 / 故障排除 / …)
有关完整的按文件分类清单以及每个参考文档涵盖的内容,请参阅
references/workflow-reference.md。
所有脚本均通过$SKILL_DIR/scripts/路径从技能根目录调用
可用脚本
辅助脚本位于scripts/目录下,可通过名称从技能根目录调用——
请通过run_script("scripts/调用每个脚本,以便代理记录
正确的工具调用。
要查看所有辅助脚本(缓存、GPU 检查、初始化)的完整列表,请浏览
scripts/;每个脚本的--help 选项会说明其参数。
如何使用此技能
- 请先阅读此文件。它仅负责路由——不包含工作流。
- 根据上方的路由表匹配用户的意图。
- 仅加载一份参考文档(DEPLOY 或 API USAGE)。不要同时预加载两份——每份参考文档体积庞大,且包含其自身的完整契约。
- 请严格遵循已加载的参考文档。这些参考文档是前身技能
vss-deploy-detection-tracking-2d(deploy/teardown/debug)和rtvicv-api(REST API)中字节级精确保留的合约——每个步骤的顺序不变量、bash 批处理规则、框渲染规则以及AskQuestion合约均被完整保留。 - 对于 DEPLOY,参考文档强制执行其自身的启动契约:一行确认 → 调用规划工具(
TodoWrite数组包含 5 个待办事项,或在较新的 Claude Code 上连续调用 5 次TaskCreate)→ 第 1 步问题。 不要进行叙述,不要进行预检,也绝不要打印“正在加载 TodoWrite/TaskCreate”或任何关于延迟工具解析的说明性文字——规划工具将静默加载。
输出契约 — DEPLOY 流程
在运行 DEPLOY / TEARDOWN / DEBUG 流程时,代理必须在每次成功部署后 严格遵守以下四项要求。这些是用户 在各步骤之间唯一的反馈渠道;跳过其中任何一项都属于 行为退化。
- 将每个步骤的结束结果渲染在固定宽度的框中——步骤 1部署
目标、步骤 2管道配置、步骤 3容器、步骤 4
应用配置、步骤 5计划+结果。而不仅仅是最终
摘要。 该框即用户的步骤收据。其几何形状是固定的(参见
下文§“通用框格式”)。各步骤的内容规则(即
哪些行应放入每个框中)位于
references/deploy-vss-detection-tracking-2d.md文件中的“步骤 N 框内容规则”部分。 - 在“步骤 5 结果”框之后,调用来自
references/next-steps.md§ “11.c”中的步骤 6AskUserQuestion——切勿将其替换为自由格式的“下一步”项目符号列表。该 菜单是部署的退出入口:它允许用户一键运行指标、 管理流、查看日志尾部或终止部署,而无需 记住 curl URL。 - 在用户选择第 6 步的存储桶后,执行来自
references/next-steps.md§ “11.d”中的后续AskUserQuestion指令——切勿用散文说明 + 可直接复制的 curl 示例 + 一个 自由文本形式的“要我运行 X 吗?”问题来替代。 每个存储桶都有其专属的 具体操作菜单;用户选择操作后,技能 便会弹出 API 输入框并执行 curl 命令。各存储桶的后续操作:- 管理流→ 添加 / 移除 / 列出。移除操作会从
/stream/get-stream-info动态生成选项——每个活跃流对应一个 选项,标记为此外 当· , ACTIVE > 1时会显示“移除全部”(完整规范:§“remove_streams子流程”)。 - 停止部署→ 停止应用 / 停止容器 / 完全拆除。
- 检查指标与 FPS→ 无后续操作;在打印
/api/v1/metricsAPI 框后 直接运行collect_metrics.sh。 - 检查存活状态 / 就绪状态→ 无需后续操作;在打印所有三个 健康检查端点的 API 框后,对它们进行探测。
- 管理流→ 添加 / 移除 / 列出。移除操作会从
- 渲染完整的每步内容,而非概述行——
渲染框是必要的,但还不够。每一步都有一个
行组成规范,位于
references/deploy-vss-detection-tracking-2d.md中的“步骤 N 框内容规则”下。步骤 4(应用配置)是 代理最常发生折叠的环节——其规范的 按用例划分的键列表位于references/apply-config.md§ “按用例划分的完整编辑列表”中,且代理必须生成一个✔ [section] key=value — 注释行,针对该表中每个键,对应 当前活动用例及设置。包含 5 个键的章节 → 5 行;包含 6 个键的章节 → 6 行。切勿为每个章节生成一行概述行。
禁止(这些是代理在 压力下退而求其次的捷径,且会破坏用户的用户体验):
- ❌内部工具加载说明。绝不能显示“我需要加载
TodoWrite(技能为任务小部件调用的延迟工具)”,
“正在加载 TaskCreate…”,“正在调用 ToolSearch 以获取规划工具…”,
或任何其他关于解析/加载/获取延迟工具的文本。
智能代理应静默加载工具。用户仅会看到
✔摘要行及其后的控件——绝不能 显示任何关于工具解析的辅助信息。 - ❌将全部 5 个部署步骤合并到单个
TaskCreate 的描述字段中。当TaskCreate是可用的规划 工具时,应连续发出5 次独立的TaskCreate调用(每个 步骤一次)。 请参阅references/task-list.md§ “InitialTaskCreatecalls” 以获取原样模板。TodoWrite也遵循相同规则——通过单次调用将 全部 5 个待办事项放入todos:[…]数组中;绝不允许单个待办事项的内容是多行列表。 - ❌默认选择
动态流模式。该技能的默认设置为stream_mode=static—— 代理会在应用启动前,将自动发现的file://URL 嵌入 DS 主配置的[source-list]块中。 仅当用户明确要求(“稍后通过 REST 添加流”、“使用动态流模式”)或在第 2 步 AskQuestion 中选择动态模式时,才切换为动态模式。 对于“部署 rtvi-cv 并启用 N 个流”这类通用查询,若选择动态模式,将破坏部署规范并 违背用户对/metrics的预期。详见references/pipeline-config.md§ “默认设置 — 技能默认处于静态模式”以了解完整的 理由。 - ❌ 仅显示一行内容
✔ 应用在 Ns 内就绪,N 个流,总帧率 Y,取代 第 5 步的“结果”框。 - ❌ 使用 ASCII 绘图字符 (
+,-,=,*) 代替轻量级 绘图字符 (┌ ─ ┐ │ └ ┘)。 - ❌ 基于“用户知道下一步该做什么”的假设而跳过步骤 6。
- ❌ 在第 6 步之后,直接输出一段 Markdown 格式的长篇大论 + 多个 curl 代码块 + 结尾的“要我运行其中任何一个吗?”——这就是 该智能代理的默认回退形式,它既绕过了 11.d 菜单 ,也跳过了每个 API 调用的确认框。 用户从菜单中选择;技能 显示已解析的API对话框;技能执行该操作。不允许自由文本提问。
- ❌ 第 4 步概览折叠——这被
部署文档中第 4 步的内容规则明确禁止:
✔ 批处理大小为 3(分块网格:1×3)→ 要求:5 行独立显示 ([streammux] batch-size=3,[primary-gie] batch-size=3,[source-list] max-batch-size=3,[tiled-display] rows=1,[tiled-display] columns=3)。✔ 输出接收器 eglsink→ 要求:每个接收器键对应一行 (eglsink 有 4 个键,例如[sink0] enable=1,type=2,sync=0,qos=0— 具体列表请参阅 apply-config.md)。✔ 静态源(3 个流,http-port=9000)→ 要求:六个 带注释的[source-list]行。✔ 拼接网格 1 行 × 3 列(单行)→ 要求:两 行,即[tiled-display] rows=1和[tiled-display] columns=3。
通用框格式
每个步骤输出框(步骤 1 至步骤 5 结果)的几何结构约定。所有框的形状相同;仅标题和 正文行会随每个步骤而变化。
- 宽度:对角线长度为128 个字符— 第 1 列为
┌,第 128 列为┐。 较宽的终端字符使框左对齐;不要拉伸 它。内部内容区域为124 个字符(在│边框内两侧各留一个空格边距)。 - 仅使用轻量级绘制框的字符:
┌ ─ ┐ │ └ ┘。不使用+、-、=、* 等ASCII 替代字符。 - 顶部边框 — 标题居中:
┌+ N₁ 个连字符 +␣+ 标题 +␣- N₂ 个连字符 +
┐,其中N₁ + N₂ + len(标题) + 2 = 126。 分配 填充符:N₁ = floor((126 − 标题长度 − 2) / 2),N₂ = 126 − 标题长度 − 2 − N₁。N₁ 和 N₂ 的差值至多为 1。
- N₂ 个连字符 +
- 正文:每个事实对应一行
│。 每行事实采用│ ✔格式(向内缩进两个 空格,符号,键值向右对齐至13,两个空格,数值)。 - 组间空白行:在逻辑组之间
(例如步骤 1 中的“身份 / 模型 / 视频”)渲染
│<124 spaces> │,以便 用户能一目了然地浏览该框。 - 底部边框:
└+ 126 个连线 +┘— 实线边框,无标题。
标准步骤标题(用于每个步骤框的顶部):
┌─────────────────────────────────────────────────────── 部署目标 ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── 管道配置 ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── 容器 ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── 应用配置 ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── 感知应用 — 计划 ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── 感知应用 — 结果 ──────────────────────────────────────────────┐
每步内容规则(哪些行放入哪个框、基于模式的行
隐藏、apply-config 分段布局、第 5 步“计划-结果”
模式、第 3 步docker run合成要求)位于
references/deploy-vss-detection-tracking-2d.md
文件中的“第 N 步框内容规则”部分——在渲染
相应步骤时请阅读这些内容。
快速触发器(助记符)
| 短语 | 流程 |
|---|---|
部署具有 4 个流的 rtvicv 仓库 2d 并显示 |
DEPLOY |
在 GPU 1 上运行 smartcity gdino |
DEPLOY |
停止感知容器 |
TEARDOWN(部署文档) |
rtvi-cv 健康检查失败 |
调试(部署文档 + 故障排除) |
向 rtvi-cv 添加流 |
API 使用方法 |
rtvi-cv 在 localhost:9000 上是否已就绪 |
API 使用 |
获取 rtvi-cv 指标 |
API 使用方法 |
通过 rtvi-cv 生成文本嵌入向量 |
API 使用方法 |
bump:1
---
name: vss-deploy-detection-tracking-2d
description: Deploy, debug, and operate the RTVI-CV 2D detection/tracking microservice and call its REST API for stream management, health checks, and metrics.
license: Apache-2.0
---
## Purpose
Deploy, debug, and operate the RTVI-CV detection / tracking 2D microservice and drive its REST API.
## Prerequisites
- Active VSS deployment reachable on `$HOST_IP` (see `vss-deploy-profile` and `references/`).
- NGC credentials in `$NGC_CLI_API_KEY` and `$NVIDIA_API_KEY` for any image pulls.
- `curl`, `jq`, and Docker available on the caller.
## Instructions
Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/` and helper scripts live in `scripts/` — call them via `run_script` when the skill points to a script by name.
## Examples
Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.
## Limitations
- Requires the matching VSS profile / microservice to be deployed and reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
## Troubleshooting
- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.
# RTVI-CV — Detection & Tracking (Unified Skill)
Unified skill for the **Real Time Video Intelligence CV (RTVI-CV)** microservice. Two action surfaces in one skill:
- **Deploy / operate / debug / tear down** the RTVI-CV container locally → see [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
- **Call the RTVI-CV REST API** (streams, health, metrics, embeddings) on a running instance → see [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md)
> **Service**: `rtvi-cv` (`metropolis_perception_app`)
> **Image**: `nvcr.io/<org>/<repo>:<tag>` — user-supplied at deploy time
> **REST port**: `9000` (`/api/v1` — `/live`, `/ready`, `/startup`, `/metrics`, `/stream/add`, `/stream/remove`, embeddings)
> **Hardware**: x86/aarch64 dGPU (T4, A100, L40, H100, B200, RTX), SBSA (Spark, Grace-Hopper), Jetson (Thor, Orin, Xavier)
---
## Action routing — pick once per invocation
| User intent (sample phrasing) | Flow | Load this reference |
|-------------------------------|------|---------------------|
| `deploy rtvi-cv warehouse 2d`, `run rtvicv warehouse-3d with 4 streams`, `start smartcity gdino`, `launch perception app`, `bring up sparse4d` | **DEPLOY** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) |
| `stop rtvi-cv`, `tear down`, `kill the perception container`, `cleanup rtvicv-perception-docker` | **TEARDOWN** (handled by deploy doc → "Mode Selection") | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/teardown-flow.md`](references/teardown-flow.md) |
| `check rtvi-cv logs`, `diagnose rtvi-cv crashing`, `troubleshoot healthcheck failing`, `rtvi-cv won't start` | **DEBUG** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/troubleshooting.md`](references/troubleshooting.md) |
| `add a stream`, `remove camera`, `list streams`, `health check`, `is rtvi-cv ready`, `get metrics`, `what's the FPS`, `check GPU usage`, `generate text embeddings`, `call rtvi-cv api` | **API USAGE** | [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md) + [`references/api-reference.md`](references/api-reference.md) |
**Selection rule:** match the user's phrasing against the table above and immediately load the corresponding reference file. Do not mix the flows — DEPLOY assumes no running container yet; API USAGE assumes the container is already running on `http://<host>:9000`.
If intent is genuinely ambiguous (e.g., the user says just "I want to use rtvi-cv"), ask one `AskQuestion`: deploy a new instance, or call an already-running one?
---
## What lives where
```
vss-deploy-detection-tracking-2d/
├── SKILL.md # this file (routing + contracts)
├── assets/ # data files (deploy-defaults.yml — single source of truth for tags / refs / paths / GPU)
├── evals/ # Tier-3 eval manifests (deploy-evals.json, usage-evals.json)
├── scripts/ # 23 bash + python helpers (see `scripts/` for the full inventory)
└── references/ # workflow runbooks (deploy / api-usage / teardown / troubleshooting / …)
```
For the full per-file inventory and what each reference covers, see
[`references/workflow-reference.md`](references/workflow-reference.md).
All scripts are invoked from the skill root via `$SKILL_DIR/scripts/<name>` — paths inside the deploy reference doc are preserved verbatim and resolve correctly when the agent runs from skill root.
---
## Available Scripts
Helpers live in `scripts/` and are invoked from the skill root by name —
call each via `run_script("scripts/<name>")` so the agent records a
proper tool invocation.
| Script | Purpose | Arguments |
| --- | --- | --- |
| `load_defaults.sh` | Detect platform (x86 dGPU / SBSA / Jetson) and resolve YAML defaults from `assets/deploy-defaults.yml`. | `--usecase <name>` |
| `fetch_resources.sh` | Download + extract NGC resources, scan for layout. | `--ngc-ref <ref>` (optional) |
| `apply_in_container.sh` | Host-side wrapper for Step 4 (`apply_config.sh` inside the running container). | `<container_name>` |
| `apply_config.sh` | In-container path-substitution, batch, sink, sources, engine cache. | `<usecase> <stream_count> <sink_type>` |
| `start_app_in_container.sh` | Host-side wrapper for Step 5 (`run_app_and_wait.sh`). | `<container_name>` |
| `run_app_and_wait.sh` | In-container app launch + readiness + metrics + log. | `<config_path>` |
| `add_streams.sh` / `update_stream_sources.sh` | REST stream lifecycle for Step 6. | `<rtsp_or_file_uri>...` |
| `collect_metrics.sh` | Pull `/api/v1/metrics` snapshot. | none |
| `discover_streams.sh` | Enumerate active streams via `/stream/get-stream-info`. | none |
| `synthesize_docker_run.sh` | Print the platform-correct `docker run` line for the resolved env. | none |
| `render_box.sh` | Render the fixed-width step receipt. | `<step_label>` |
| `calibration_manager.py` | Manage calibration artefacts + per-use-case engine cache invalidation. | `--usecase <name> --reset` |
For the full inventory of helpers (cache, GPU checks, setup) browse
`scripts/`; each script's `--help` describes its arguments.
## How to use this skill
1. **Read this file first.** It only routes — it does not contain workflows.
2. **Match the user's intent** against the routing table above.
3. **Load exactly one reference doc** (DEPLOY or API USAGE). Don't preload both — each reference is large and contains its own full contract.
4. **Follow the loaded reference exactly.** The reference docs are the byte-for-byte preserved contracts from the predecessor skills `vss-deploy-detection-tracking-2d` (deploy/teardown/debug) and `rtvicv-api` (REST API) — every step ordering invariant, bash-batching rule, box-rendering rule, and `AskQuestion` contract is retained.
5. **For DEPLOY**, the reference doc enforces its own startup contract: one-line acknowledgement → planning-tool call (`TodoWrite` array of 5 todos, OR 5 successive `TaskCreate` calls on newer Claude Code) → Step 1 question. Do not narrate, do not pre-flight, and never print "loading TodoWrite/TaskCreate" or any deferred-tool resolution prose — the planning tool is loaded silently.
---
## Output contract — DEPLOY flow
When running the DEPLOY / TEARDOWN / DEBUG flow, the agent MUST honour
all four items below on every successful deploy. These are the user's
only feedback channel between steps; skipping any of them is a
behaviour regression.
1. **Render every step's exit in a fixed-width box** — Step 1 *Deploy
targets*, Step 2 *Pipeline configuration*, Step 3 *Container*, Step 4
*Apply configuration*, Step 5 *Plan* + *Results*. Not just the final
summary. The box is the user's step receipt. Geometry is fixed (see
§ "Universal box format" below). Per-step **content** rules (what
rows go inside each box) live in [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule".
2. **After the Step 5 Results box, issue the Step 6 `AskUserQuestion`**
from [`references/next-steps.md`](references/next-steps.md) § "11.c"
— never replace it with a free-form *Next steps* bullet list. The
menu is the deploy's exit handle: it lets the user run metrics,
manage streams, tail logs, or tear down with one click instead of
having to remember curl URLs.
3. **After the user picks a Step 6 bucket, issue the follow-up
`AskUserQuestion`** from [`references/next-steps.md`](references/next-steps.md)
§ "11.d" — never substitute prose + ready-to-copy curl examples + a
free-text "want me to run X?" question. Each bucket has its own
menu of concrete actions; the user picks the action, then the skill
emits the API box and runs the curl. Per-bucket follow-ups:
- **Manage streams** → Add / Remove / List. **Remove builds its
options dynamically from `/stream/get-stream-info`** — one option
per active stream labelled `<camera_id> · <camera_url>` plus
"Remove ALL" when `ACTIVE > 1` (full spec: § "`remove_streams`
sub-flow").
- **Stop the deployment** → Stop app / Stop container / Full teardown.
- **Check metrics & FPS** → no follow-up; run `collect_metrics.sh`
directly after printing the `/api/v1/metrics` API box.
- **Check liveness / readiness** → no follow-up; probe all three
health endpoints after printing their API boxes.
4. **Render the FULL per-step content, not an overview row** —
rendering the box is necessary but not sufficient. Each step has a
row composition spec in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule". **Step 4 (Apply configuration) is
where the agent collapses most often** — its canonical
per-use-case key list lives in
[`references/apply-config.md`](references/apply-config.md)
§ "Per-use-case complete edit list", and the agent MUST emit one
`✔ [section] key=value — annotation` row per key in that table for
the active use case + settings. A section with 5 keys → 5 rows; a
section with 6 keys → 6 rows. Never one overview row per section.
Forbidden (these are the shortcuts the agent falls back to under
pressure, and they break the user's UX):
- ❌ **Internal tool-loading narration.** Never print "I need to load
TodoWrite (a deferred tool the skill calls for the task widget)",
"Loading TaskCreate…", "Calling ToolSearch for the planning tool…",
or any other text about resolving / loading / fetching deferred tools.
The agent loads tools **silently**. The user only ever sees the `✔
<pinned-values>` summary line followed by the widget — never any
scaffolding around tool resolution.
- ❌ **Collapsing all 5 deploy steps into a single `TaskCreate`'s
`description` field.** When `TaskCreate` is the available planning
tool, issue **5 separate `TaskCreate` calls** back-to-back (one per
step). See `references/task-list.md` § "Initial `TaskCreate` calls"
for the verbatim template. Same rule for `TodoWrite` — one call with
all 5 todos in the `todos:[…]` array; never one todo whose `content`
is a multi-line list.
- ❌ **Silently choosing `dynamic` stream-mode.** The skill default is
`stream_mode=static` — the agent bakes auto-discovered `file://` URLs
into the DS main config's `[source-list]` block before app start.
Switch to `dynamic` only when the user explicitly asks ("add streams
later via REST", "use dynamic stream mode") OR when they pick `dynamic`
in the Step 2 AskQuestion. Picking `dynamic` for a generic "deploy
rtvi-cv with N streams" query breaks the deploy rubric and the
user's `/metrics` expectations. See
[`references/pipeline-config.md`](references/pipeline-config.md)
§ "Defaults — the skill is static-mode by default" for the full
rationale.
- ❌ A one-line `✔ App ready in Ns, N streams, fps total Y` in place of
the Step 5 Results box.
- ❌ ASCII box-drawing chars (`+`, `-`, `=`, `*`) instead of light
box-drawing chars (`┌ ─ ┐ │ └ ┘`).
- ❌ Skipping Step 6 on the assumption "the user knows what to do next".
- ❌ After Step 6, dumping a markdown wall of prose + multiple curl
blocks + a closing "want me to run any of these?" — that's the
shape the agent falls back to and it bypasses both the 11.d menu
and the per-API-call box. The user picks from a menu; the skill
shows the resolved API box; the skill runs it. No free-text Q.
- ❌ Step 4 overview collapses — these are explicitly banned by the
deploy doc's Step 4 content rule:
- `✔ Batch size 3 (tile grid: 1×3)` → required: 5 separate rows
(`[streammux] batch-size=3`, `[primary-gie] batch-size=3`,
`[source-list] max-batch-size=3`, `[tiled-display] rows=1`,
`[tiled-display] columns=3`).
- `✔ Output sink eglsink` → required: one row per sink key
(4 keys for eglsink, e.g. `[sink0] enable=1`, `type=2`,
`sync=0`, `qos=0` — read apply-config.md for the exact list).
- `✔ Sources static (3 streams, http-port=9000)` → required: six
annotated `[source-list]` rows.
- `✔ Tile grid 1 row × 3 cols` (single row) → required: two
rows, `[tiled-display] rows=1` and `[tiled-display] columns=3`.
## Universal box format
The geometry contract for every step-exit box (Step 1 through Step 5
Results). The same shape across every box; only the **title** and the
**body rows** change per step.
- **Width: 128 chars** corner-to-corner — `┌` at column 1, `┐` at
column 128. Wider terminals leave the box flush-left; do not stretch
it. Inner content area is **124 chars** (with one space margin on
each side inside the `│` borders).
- **Light box-drawing chars only**: `┌ ─ ┐ │ └ ┘`. No `+`, `-`, `=`,
`*` ASCII fallbacks.
- **Top border — title CENTERED**: `┌` + N₁ dashes + `␣` + title + `␣`
+ N₂ dashes + `┐`, where `N₁ + N₂ + len(title) + 2 = 126`. Distribute
the pad: `N₁ = floor((126 − len(title) − 2) / 2)`,
`N₂ = 126 − len(title) − 2 − N₁`. N₁ and N₂ differ by at most 1.
- **Body**: one `│ <content padded to inner-content 124> │` per fact.
Each fact line uses the ` ✔ <key-padded-to-13> <value>` form (two
spaces in, glyph, key right-padded to 13, two spaces, value).
- **Blank lines between groups**: render `│ <124 spaces> │` between
logical groups (e.g. Identity / Model / Videos in Step 1) so the
user can scan the box at a glance.
- **Bottom border**: `└` + 126 dashes + `┘` — solid border, no title.
Standard step titles (used at the top of each step's box):
```
┌─────────────────────────────────────────────────────── Deploy targets ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── Pipeline configuration ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── Container ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── Apply configuration ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── Perception Application — Plan ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── Perception Application — Results ──────────────────────────────────────────────┐
```
Per-step content rules (which rows go in which box, mode-aware row
hiding, the apply-config sectioned layout, the Step 5 PLAN-then-RESULT
pattern, the Step 3 `docker run` synthesis requirement) live in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule" — read those when rendering the
corresponding step.
## Quick triggers (mnemonic)
| Phrase | Flow |
|--------|------|
| `deploy rtvicv warehouse 2d with 4 streams and display` | DEPLOY |
| `run smartcity gdino on gpu 1` | DEPLOY |
| `stop the perception container` | TEARDOWN (deploy doc) |
| `rtvi-cv healthcheck failing` | DEBUG (deploy doc + troubleshooting) |
| `add a stream to rtvi-cv` | API USAGE |
| `is rtvi-cv ready on localhost:9000` | API USAGE |
| `get rtvi-cv metrics` | API USAGE |
| `generate text embeddings via rtvi-cv` | API USAGE |
bump:1





首页
