vss-deploy-detection-tracking-2d
NVIDIA/skills
RTVI-CV 2D 検出・追跡マイクロサービスをデプロイ、デバッグ、運用し、ストリーム管理、ヘルスチェック、メトリクスのためにその REST API を呼び出します。
...すべて拡張します目的
RTVI-CV検出/追跡2Dマイクロサービスをデプロイ、デバッグ、運用し、そのREST APIを操作する。
前提条件
$HOST_IPからアクセス可能なアクティブな VSS デプロイメント(vss-deploy-profileおよびreferences/を参照)。- イメージのプルを行うための
$NGC_CLI_API_KEYおよび$NVIDIA_API_KEYに格納された NGC 認証情報。 - 呼び出し元で
curl、jq、および Docker が利用可能であること。
手順
以下のルーティングテーブルおよびステップバイステップのワークフローに従ってください。「workflow」、「quick start」、または「flow」で終わる各セクションは、上から順に実行することを想定しています。詳細なリファレンス資料はreferences/に、ヘルパースクリプトはscripts/にあります。スキルがスクリプト名を指定している場合は、run_scriptを通じてそれらを呼び出してください。
例
動作確認済みのエンドツーエンドの例は、evals/ディレクトリ(各*.jsonマニフェストには実行可能なシナリオが含まれています)および以下のワークフローごとのcurlブロック内に記載されています。これらを再現するには、nv-base validateを使用して Tier-3 評価を実行してください。
制限事項
- 対応する VSS プロファイル/マイクロサービスがデプロイされており、呼び出し元からアクセス可能である必要があります。
- NGCでホストされているモデルおよびNIMは、レート制限、GPUメモリ要件、およびライセンス制限の対象となる場合があります。
- 同時実行数、GPU メモリ、およびストレージの制限は、ホストのハードウェアとプロファイルの compose ファイルによって異なります。
トラブルシューティング
- エラー: REST 呼び出しで「接続拒否」が返されました。原因: ターゲットのマイクロサービスが実行されていません。解決策:
/docsまたは/healthをプローブする。vss-deploy-profileまたは対応するvss-deploy-*スキルを使用して再デプロイする。 - エラー:NGC プルから HTTP 401/403 が返されました。原因:
NGC_CLI_API_KEYが存在しないか、有効期限が切れています。解決策:docker login nvcr.ioを実行し、キーを再エクスポートしてから再試行してください。 - エラー:コンテナの OOM またはモデルの読み込みに失敗しました。原因:選択したプロファイルに対して GPU メモリが不足しています。解決策:より小さなバリアントに切り替えるか、
`docker compose down` を実行して GPU を解放してください。
RTVI-CV — 検出および追跡(統合スキル)
Real Time Video Intelligence CV (RTVI-CV)マイクロサービス用の統合スキルです。1つのスキル内に2つのアクションサーフェスが含まれます:
- RTVI-CV コンテナをローカルでデプロイ/運用/デバッグ/終了→
references/deploy-vss-detection-tracking-2d.md を参照 - 実行中のインスタンスでRTVI-CV REST API(ストリーム、ヘルス、メトリクス、埋め込み)を呼び出す→
references/usage-vss-detection-tracking-2d.md を参照
サービス:
rtvi-cv(metropolis_perception_app) イメージ:nvcr.io/— デプロイ時にユーザーが指定 RESTポート:/ : 9000(/api/v1—/live,/ready,/startup,/metrics,/stream/add,/stream/remove, embeddings) ハードウェア: x86/aarch64 dGPU (T4、A100、L40、H100、B200、RTX)、SBSA (Spark、Grace-Hopper)、Jetson (Thor、Orin、Xavier)
アクションルーティング — 呼び出しごとに1回選択
| ユーザーの意図(例文) | フロー | このリファレンスをロード |
|---|---|---|
rtvi-cv warehouse 2d をデプロイ、rtvicv warehouse-3d を 4 ストリームで実行、smartcity gdino を起動、知覚アプリを起動、sparse4d を起動 |
DEPLOY | references/deploy-vss-detection-tracking-2d.md |
rtvi-cvを停止し、テアダウンを行い、perceptionコンテナを強制終了し、rtvicv-perception-dockerをクリーンアップする |
TEARDOWN(deploy ドキュメントの「モードの選択」で処理) | references/deploy-vss-detection-tracking-2d.md+references/teardown-flow.md |
rtvi-cvのログを確認し、rtvi-cvのクラッシュを診断し、ヘルスチェックの失敗をトラブルシューティングし、rtvi-cvが起動しない問題を解決する |
DEBUG | references/deploy-vss-detection-tracking-2d.md+references/troubleshooting.md |
ストリームの追加、カメラの削除、ストリームの一覧表示、ヘルスチェック、rtvi-cv の準備状況、メトリクスの取得、FPS の確認、GPU 使用率の確認、テキスト埋め込みの生成、rtvi-cv API の呼び出し |
API の使用方法 | references/usage-vss-detection-tracking-2d.md+references/api-reference.md |
選択ルール:ユーザーの表現を上記の表と照合し、対応するリファレンスファイルを直ちに読み込みます。フローを混同しないでください。DEPLOY はコンテナがまだ実行されていないことを前提としています。API USAGE は、コンテナがすでにhttp:// で実行されていることを前提としています。
意図が真に曖昧な場合(例:ユーザーが単に「rtvi-cvを使いたい」と言った場合)、1つのAskQuestionを尋ねる:新しいインスタンスをデプロイするか、それとも既に実行中のインスタンスを呼び出すか?
各ファイルの配置場所
vss-deploy-detection-tracking-2d/
├── SKILL.md # このファイル(ルーティング + 契約)
├── assets/ # データファイル(deploy-defaults.yml — タグ / リファレンス / パス / GPU に関する唯一の信頼できる情報源)
├── evals/ # Tier-3 評価マニフェスト (deploy-evals.json, usage-evals.json)
├── scripts/ # 23個のbashおよびpythonヘルパー(完全な一覧は `scripts/` を参照)
└── references/ # ワークフロー・ランブック(デプロイ / API使用 / テアダウン / トラブルシューティング / …)
ファイルごとの完全な一覧および各リファレンスのカバー範囲については、
references/workflow-reference.md を参照してください。
すべてのスクリプトは、$SKILL_DIR/scripts/経由でスキルルートから呼び出されます。deploy リファレンスドキュメント内のパスはそのまま保持されており、エージェントがスキルルートから実行された場合でも正しく解決されます。
利用可能なスクリプト
ヘルパースクリプトはscripts/ディレクトリにあり、スキルルートから名前を指定して呼び出されます。
エージェントが
適切なツール呼び出しを記録できるよう、それぞれrun_script("scripts/を使用して呼び出してください。
ヘルパー(キャッシュ、GPUチェック、セットアップ)の完全な一覧については、
scripts/ を参照してください。各スクリプトの--helpオプションで引数が説明されています。
このスキルの使い方
- まずこのファイルをお読みください。本ツールはルーティングのみを行い、ワークフローは含まれていません。
- ユーザーの意図を上記のルーティングテーブルと照合してください。
- リファレンスドキュメント(DEPLOY または API USAGE)を正確に 1 つだけ読み込みます。両方を同時に読み込まないでください。各リファレンスは容量が大きく、独自の完全な契約が含まれています。
- 読み込まれたリファレンスに正確に従ってください。リファレンス文書は、先行スキル
vss-deploy-detection-tracking-2d(deploy/teardown/debug)およびrtvicv-api(REST API)からバイト単位で忠実に保存された契約であり、各ステップの順序不変条件、bash バッチ処理ルール、ボックスレンダリングルール、AskQuestion契約がすべて保持されています。 - DEPLOY の場合、リファレンスドキュメントは独自の起動契約を強制します:1行の承認 → プランニングツールの呼び出し(5つのToDoからなる
TodoWrite配列、または新しい Claude Code では 5 回連続のTaskCreate呼び出し) → ステップ 1 の質問。 ナレーションを行わず、事前チェックもせず、「loading TodoWrite/TaskCreate」や、ツール解決が先送りされていることを示す文章を絶対に表示しないでください — 計画ツールは静かに読み込まれます。
出力契約 — DEPLOY フロー
DEPLOY / TEARDOWN / DEBUG フローを実行する際、エージェントは デプロイが成功するたびに、以下の 4 項目すべてを遵守しなければなりません。これらは、ステップ間の ユーザーへの唯一のフィードバックチャネルであり、いずれかを省略することは 動作の退行となります。
- 各ステップの終了画面を固定幅のボックスで表示すること— ステップ1デプロイ
対象、ステップ2パイプライン構成、ステップ3コンテナ、ステップ4
構成の適用、ステップ5計画+結果。最終的な
要約だけでなく。 このボックスは、ユーザーにとってのステップの「領収書」です。サイズや形状は固定されています(以下の
§「ユニバーサルボックス形式」を参照)。ステップごとのコンテンツルール(各ボックス内に
どの行を表示するか)は、
references/deploy-vss-detection-tracking-2d.mdの 「Step N box content rule」に記載されています。 - ステップ5「結果」ボックスの後、
references/next-steps.mdの§「11.c」にある ステップ6「AskUserQuestion」を実行してください — これを自由形式の「次のステップ」箇条書きリストで置き換えてはなりません。この メニューはデプロイの終了ハンドルです。これにより、ユーザーは curl の URL を 覚える必要がなく、ワンクリックでメトリクスの実行、 ストリームの管理、ログのテール表示、または環境のクリーンアップを行えます。 - ユーザーがステップ6のバケットを選択した後、
references/next-steps.md§「11.d」にあるフォローアップのAskUserQuestionを発行してください。— 文章やコピー可能なcurlの例、 および「Xを実行しますか?」といった自由形式の質問で代用してはいけません。 各バケットには独自の 具体的なアクションのメニューがあります。ユーザーがアクションを選択すると、スキルが APIボックスを発行し、curlを実行します。バケットごとのフォローアップ:- ストリームの管理→ 追加 / 削除 / 一覧表示。「削除」は
/stream/get-stream-infoから動的にオプションを生成します。アクティブなストリームごとにというラベルのオプションが1つずつ表示され、· ACTIVE > 1の場合は「すべて削除」が追加されます(詳細仕様:§「remove_streamsサブフロー」)。 - デプロイを停止→ アプリの停止 / コンテナの停止 / 完全なクリーンアップ。
- メトリクスとFPSを確認→ 追跡処理なし;
/api/v1/metricsAPIボックスの表示直後にcollect_metrics.shを実行。 - 稼働状態/準備状態を確認→ 追跡なし;APIボックスを出力した後、3つの ヘルスエンドポイントすべてをプローブする。
- ストリームの管理→ 追加 / 削除 / 一覧表示。「削除」は
- 概要行ではなく、ステップごとの完全なコンテンツをレンダリングする —
ボックスのレンダリングは必要だが十分ではない。各ステップには、
references/deploy-vss-detection-tracking-2d.md内の「Step N box content rule」に 行構成の仕様がある。ステップ 4(構成の適用)は、 エージェントが最も頻繁に折りたたまれる箇所です。その標準的な ユースケースごとのキーリストは、references/apply-config.mdの§「ユースケースごとの完全な編集リスト」に記載されており、エージェントは必ず 1 つを✔ [section] key=value — の注釈行を1つ出力しなければなりません。 アクティブなユースケースと設定に対応する、そのテーブル内の各キーごとに。キーが5つのセクション → 5行、 キーが6つのセクション → 6行。セクションごとに1つの概要行を作成してはなりません。
禁止事項(これらはエージェントが逼迫時に頼ってしまう近道であり、 ユーザーのUXを損なうものです):
- ❌内部ツール読み込みに関するナレーション。「TodoWrite(タスクウィジェットのためにスキルが呼び出す遅延ツール)を
読み込む必要があります」、「TaskCreateを読み込み中…」、「計画ツール用にToolSearchを呼び出し中…」、
あるいは遅延ツールの解決/読み込み/取得に関するその他のテキストを、決して表示してはならない。
エージェントはツールを黙って読み込みます。ユーザーには、ウィジェットの直後に表示される「
✔」という要約行のみが表示され、ツールの解決に関する 説明的な記述は一切表示されません。 - ❌5つのデプロイ手順すべてを、
1つのTaskCreateの説明フィールドにまとめ込むこと。TaskCreateが利用可能な計画 ツールである場合、5つの別々のTaskCreate呼び出しを連続して発行する(各 手順につき1回)。 テンプレートの原文については、references/task-list.md の§「InitialTaskCreatecalls」を 参照してください。TodoWriteについても同様のルールが適用されます。つまり、 todos:[…]配列に5つのToDoをすべて含めた1回の呼び出しを行い、内容が多行リストである単一のToDoを作成してはなりません。 - ❌
動的ストリームモードを黙って選択すること。スキルのデフォルトはstream_mode=staticです。エージェントは、アプリ起動前に自動検出されたfile://URL を DS メイン設定の[source-list]ブロックに組み込みます。 ユーザーが明示的に要求した場合(「REST経由で後でストリームを追加する」、 「動的ストリームモードを使用する」)またはステップ2のAskQuestionで動的モードを選択した場合にのみ、動的モードに切り替えます。 一般的な「Nストリームのrtvi-cvを デプロイ」というクエリに対して動的モードを選択すると、デプロイの基準が破られ、ユーザーの/metricsに対する期待も裏切られます。詳細な 根拠については、references/pipeline-config.mdの§「デフォルト — スキルはデフォルトで静的モード」を参照してください。 - ❌ 1行の
✔ 「N秒でアプリ準備完了、Nストリーム、合計Y fps」を ステップ5の結果ボックスの代わりに表示。 - ❌ ライトな
ボックス描画文字(
┌ ─ ┐ │ └ ┘)の代わりに、ASCIIのボックス描画文字(+,-,=,*)を使用。 - ❌ 「ユーザーは次に何をすべきか分かっている」という前提で、ステップ6をスキップする。
- ❌ ステップ6の後、長文のマークダウン + 複数のcurl ブロック + 締めくくりの「これらを実行しますか?」を 一気に出力する — これがエージェントがフォールバックする形式であり、11.dのメニュー とAPI呼び出しごとのボックスの両方をバイパスしてしまう。 ユーザーがメニューから選択し、スキルが 解決済みのAPIボックスを表示し、スキルがそれを実行します。自由入力形式の質問はありません。
- ❌ ステップ4の概要が折りたたまれる — これらはデプロイドキュメントの
ステップ4コンテンツルールによって明示的に禁止されています:
✔ バッチサイズ 3(タイルグリッド:1×3)→ 必須:5 つの別々の行 ([streammux] batch-size=3,[primary-gie] batch-size=3,[source-list] max-batch-size=3,[tiled-display] rows=1,[tiled-display] columns=3)。✔ 出力シンク eglsink→ 必須:シンクキーごとに1行 (eglsinkの場合は4つのキー、例:[sink0] enable=1,type=2,sync=0,qos=0— 正確なリストについては apply-config.md を参照)。✔ ソース static(3 ストリーム、http-port=9000)→ 必須:6 つの 注釈付き[source-list]行。✔ タイルグリッド 1 行 × 3 列(1 行のみ) → 必須:2 行、[tiled-display] rows=1および[tiled-display] columns=3。
ユニバーサルボックス形式
各ステップ終了ボックス(ステップ1からステップ5の 結果)のジオメトリ契約。すべてのボックスで形状は同一で、ステップごとに タイトル行と本文行のみが変更されます。
- 幅:対角線間で128文字—
1列目に┌、128列目に┐。 幅の広い端末では、ボックスを左揃えにし、 ボックスを横方向に伸ばさない。内部のコンテンツ領域は124文字(│の境界線の内側、 両側に1文字分の余白を含む)。 - ボックス描画には軽めの文字のみを使用:
┌ ─ ┐ │ └ ┘。+、-、=、*といったASCIIの代替文字は使用しない。 - 上枠 — タイトルを中央揃え:
┌+ N₁ ダッシュ +␣+ タイトル +␣- N₂ 本のダッシュ +
┐。ここで、N₁ + N₂ + len(title) + 2 = 126となる。 パディングを 配分する:N₁ = floor((126 − タイトルの長さ − 2) / 2)、N₂ = 126 − タイトルの長さ − 2 − N₁。N₁ と N₂ の差は最大で 1 である。
- N₂ 本のダッシュ +
- 本文:事実ごとに 1 行
│を配置する。 各事実行は、│ ✔の形式を使用する(2 スペース空けて、記号、13 文字になるよう右詰めされたキー、2 スペース、値)。 - グループ間の空白行:論理的なグループ間(例:ステップ1の「Identity」/「Model」/「Videos」)に
│<124 spaces> │を配置し、 ユーザーがボックスを一目で把握できるようにする。 - 下部の境界線:
└+ 126 ダッシュ +┘— 実線、タイトルなし。
標準的なステップタイトル(各ステップのボックス上部に使用):
┌─────────────────────────────────────────────────────── デプロイ先 ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── パイプライン構成 ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── コンテナ ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── 構成を適用 ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── 知覚アプリケーション — 計画 ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── 知覚アプリケーション — 結果 ──────────────────────────────────────────────┐
ステップごとのコンテンツルール(どの行をどのボックスに入れるか、モードに応じた行の
非表示、apply-config のセクション化されたレイアウト、ステップ 5 の PLAN-then-RESULT
パターン、ステップ3のdocker run合成要件)は、
references/deploy-vss-detection-tracking-2d.md
内の「ステップNのボックス内容ルール」に記載されています。対応するステップを
レンダリングする際は、それらを参照してください。
クイックトリガー(ニーモニック)
| フレーズ | フロー |
|---|---|
4つのストリームを持つrtvicv warehouse 2dをデプロイし、結果を表示 |
DEPLOY |
GPU 1上でsmartcity gdinoを実行する |
DEPLOY |
Perceptionコンテナを停止する |
TEARDOWN(ドキュメントのデプロイ) |
rtvi-cv のヘルスチェックに失敗 |
DEBUG (デプロイドキュメント + トラブルシューティング) |
rtvi-cvにストリームを追加 |
APIの使用方法 |
rtvi-cvはlocalhost:9000で利用可能か |
APIの使用方法 |
rtvi-cvのメトリクスを取得する |
APIの使用方法 |
rtvi-cv を使用してテキストの埋め込みを生成する |
APIの使用方法 |
bump:1
---
name: vss-deploy-detection-tracking-2d
description: Deploy, debug, and operate the RTVI-CV 2D detection/tracking microservice and call its REST API for stream management, health checks, and metrics.
license: Apache-2.0
---
## Purpose
Deploy, debug, and operate the RTVI-CV detection / tracking 2D microservice and drive its REST API.
## Prerequisites
- Active VSS deployment reachable on `$HOST_IP` (see `vss-deploy-profile` and `references/`).
- NGC credentials in `$NGC_CLI_API_KEY` and `$NVIDIA_API_KEY` for any image pulls.
- `curl`, `jq`, and Docker available on the caller.
## Instructions
Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/` and helper scripts live in `scripts/` — call them via `run_script` when the skill points to a script by name.
## Examples
Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.
## Limitations
- Requires the matching VSS profile / microservice to be deployed and reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
## Troubleshooting
- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.
# RTVI-CV — Detection & Tracking (Unified Skill)
Unified skill for the **Real Time Video Intelligence CV (RTVI-CV)** microservice. Two action surfaces in one skill:
- **Deploy / operate / debug / tear down** the RTVI-CV container locally → see [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
- **Call the RTVI-CV REST API** (streams, health, metrics, embeddings) on a running instance → see [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md)
> **Service**: `rtvi-cv` (`metropolis_perception_app`)
> **Image**: `nvcr.io/<org>/<repo>:<tag>` — user-supplied at deploy time
> **REST port**: `9000` (`/api/v1` — `/live`, `/ready`, `/startup`, `/metrics`, `/stream/add`, `/stream/remove`, embeddings)
> **Hardware**: x86/aarch64 dGPU (T4, A100, L40, H100, B200, RTX), SBSA (Spark, Grace-Hopper), Jetson (Thor, Orin, Xavier)
---
## Action routing — pick once per invocation
| User intent (sample phrasing) | Flow | Load this reference |
|-------------------------------|------|---------------------|
| `deploy rtvi-cv warehouse 2d`, `run rtvicv warehouse-3d with 4 streams`, `start smartcity gdino`, `launch perception app`, `bring up sparse4d` | **DEPLOY** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) |
| `stop rtvi-cv`, `tear down`, `kill the perception container`, `cleanup rtvicv-perception-docker` | **TEARDOWN** (handled by deploy doc → "Mode Selection") | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/teardown-flow.md`](references/teardown-flow.md) |
| `check rtvi-cv logs`, `diagnose rtvi-cv crashing`, `troubleshoot healthcheck failing`, `rtvi-cv won't start` | **DEBUG** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/troubleshooting.md`](references/troubleshooting.md) |
| `add a stream`, `remove camera`, `list streams`, `health check`, `is rtvi-cv ready`, `get metrics`, `what's the FPS`, `check GPU usage`, `generate text embeddings`, `call rtvi-cv api` | **API USAGE** | [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md) + [`references/api-reference.md`](references/api-reference.md) |
**Selection rule:** match the user's phrasing against the table above and immediately load the corresponding reference file. Do not mix the flows — DEPLOY assumes no running container yet; API USAGE assumes the container is already running on `http://<host>:9000`.
If intent is genuinely ambiguous (e.g., the user says just "I want to use rtvi-cv"), ask one `AskQuestion`: deploy a new instance, or call an already-running one?
---
## What lives where
```
vss-deploy-detection-tracking-2d/
├── SKILL.md # this file (routing + contracts)
├── assets/ # data files (deploy-defaults.yml — single source of truth for tags / refs / paths / GPU)
├── evals/ # Tier-3 eval manifests (deploy-evals.json, usage-evals.json)
├── scripts/ # 23 bash + python helpers (see `scripts/` for the full inventory)
└── references/ # workflow runbooks (deploy / api-usage / teardown / troubleshooting / …)
```
For the full per-file inventory and what each reference covers, see
[`references/workflow-reference.md`](references/workflow-reference.md).
All scripts are invoked from the skill root via `$SKILL_DIR/scripts/<name>` — paths inside the deploy reference doc are preserved verbatim and resolve correctly when the agent runs from skill root.
---
## Available Scripts
Helpers live in `scripts/` and are invoked from the skill root by name —
call each via `run_script("scripts/<name>")` so the agent records a
proper tool invocation.
| Script | Purpose | Arguments |
| --- | --- | --- |
| `load_defaults.sh` | Detect platform (x86 dGPU / SBSA / Jetson) and resolve YAML defaults from `assets/deploy-defaults.yml`. | `--usecase <name>` |
| `fetch_resources.sh` | Download + extract NGC resources, scan for layout. | `--ngc-ref <ref>` (optional) |
| `apply_in_container.sh` | Host-side wrapper for Step 4 (`apply_config.sh` inside the running container). | `<container_name>` |
| `apply_config.sh` | In-container path-substitution, batch, sink, sources, engine cache. | `<usecase> <stream_count> <sink_type>` |
| `start_app_in_container.sh` | Host-side wrapper for Step 5 (`run_app_and_wait.sh`). | `<container_name>` |
| `run_app_and_wait.sh` | In-container app launch + readiness + metrics + log. | `<config_path>` |
| `add_streams.sh` / `update_stream_sources.sh` | REST stream lifecycle for Step 6. | `<rtsp_or_file_uri>...` |
| `collect_metrics.sh` | Pull `/api/v1/metrics` snapshot. | none |
| `discover_streams.sh` | Enumerate active streams via `/stream/get-stream-info`. | none |
| `synthesize_docker_run.sh` | Print the platform-correct `docker run` line for the resolved env. | none |
| `render_box.sh` | Render the fixed-width step receipt. | `<step_label>` |
| `calibration_manager.py` | Manage calibration artefacts + per-use-case engine cache invalidation. | `--usecase <name> --reset` |
For the full inventory of helpers (cache, GPU checks, setup) browse
`scripts/`; each script's `--help` describes its arguments.
## How to use this skill
1. **Read this file first.** It only routes — it does not contain workflows.
2. **Match the user's intent** against the routing table above.
3. **Load exactly one reference doc** (DEPLOY or API USAGE). Don't preload both — each reference is large and contains its own full contract.
4. **Follow the loaded reference exactly.** The reference docs are the byte-for-byte preserved contracts from the predecessor skills `vss-deploy-detection-tracking-2d` (deploy/teardown/debug) and `rtvicv-api` (REST API) — every step ordering invariant, bash-batching rule, box-rendering rule, and `AskQuestion` contract is retained.
5. **For DEPLOY**, the reference doc enforces its own startup contract: one-line acknowledgement → planning-tool call (`TodoWrite` array of 5 todos, OR 5 successive `TaskCreate` calls on newer Claude Code) → Step 1 question. Do not narrate, do not pre-flight, and never print "loading TodoWrite/TaskCreate" or any deferred-tool resolution prose — the planning tool is loaded silently.
---
## Output contract — DEPLOY flow
When running the DEPLOY / TEARDOWN / DEBUG flow, the agent MUST honour
all four items below on every successful deploy. These are the user's
only feedback channel between steps; skipping any of them is a
behaviour regression.
1. **Render every step's exit in a fixed-width box** — Step 1 *Deploy
targets*, Step 2 *Pipeline configuration*, Step 3 *Container*, Step 4
*Apply configuration*, Step 5 *Plan* + *Results*. Not just the final
summary. The box is the user's step receipt. Geometry is fixed (see
§ "Universal box format" below). Per-step **content** rules (what
rows go inside each box) live in [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule".
2. **After the Step 5 Results box, issue the Step 6 `AskUserQuestion`**
from [`references/next-steps.md`](references/next-steps.md) § "11.c"
— never replace it with a free-form *Next steps* bullet list. The
menu is the deploy's exit handle: it lets the user run metrics,
manage streams, tail logs, or tear down with one click instead of
having to remember curl URLs.
3. **After the user picks a Step 6 bucket, issue the follow-up
`AskUserQuestion`** from [`references/next-steps.md`](references/next-steps.md)
§ "11.d" — never substitute prose + ready-to-copy curl examples + a
free-text "want me to run X?" question. Each bucket has its own
menu of concrete actions; the user picks the action, then the skill
emits the API box and runs the curl. Per-bucket follow-ups:
- **Manage streams** → Add / Remove / List. **Remove builds its
options dynamically from `/stream/get-stream-info`** — one option
per active stream labelled `<camera_id> · <camera_url>` plus
"Remove ALL" when `ACTIVE > 1` (full spec: § "`remove_streams`
sub-flow").
- **Stop the deployment** → Stop app / Stop container / Full teardown.
- **Check metrics & FPS** → no follow-up; run `collect_metrics.sh`
directly after printing the `/api/v1/metrics` API box.
- **Check liveness / readiness** → no follow-up; probe all three
health endpoints after printing their API boxes.
4. **Render the FULL per-step content, not an overview row** —
rendering the box is necessary but not sufficient. Each step has a
row composition spec in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule". **Step 4 (Apply configuration) is
where the agent collapses most often** — its canonical
per-use-case key list lives in
[`references/apply-config.md`](references/apply-config.md)
§ "Per-use-case complete edit list", and the agent MUST emit one
`✔ [section] key=value — annotation` row per key in that table for
the active use case + settings. A section with 5 keys → 5 rows; a
section with 6 keys → 6 rows. Never one overview row per section.
Forbidden (these are the shortcuts the agent falls back to under
pressure, and they break the user's UX):
- ❌ **Internal tool-loading narration.** Never print "I need to load
TodoWrite (a deferred tool the skill calls for the task widget)",
"Loading TaskCreate…", "Calling ToolSearch for the planning tool…",
or any other text about resolving / loading / fetching deferred tools.
The agent loads tools **silently**. The user only ever sees the `✔
<pinned-values>` summary line followed by the widget — never any
scaffolding around tool resolution.
- ❌ **Collapsing all 5 deploy steps into a single `TaskCreate`'s
`description` field.** When `TaskCreate` is the available planning
tool, issue **5 separate `TaskCreate` calls** back-to-back (one per
step). See `references/task-list.md` § "Initial `TaskCreate` calls"
for the verbatim template. Same rule for `TodoWrite` — one call with
all 5 todos in the `todos:[…]` array; never one todo whose `content`
is a multi-line list.
- ❌ **Silently choosing `dynamic` stream-mode.** The skill default is
`stream_mode=static` — the agent bakes auto-discovered `file://` URLs
into the DS main config's `[source-list]` block before app start.
Switch to `dynamic` only when the user explicitly asks ("add streams
later via REST", "use dynamic stream mode") OR when they pick `dynamic`
in the Step 2 AskQuestion. Picking `dynamic` for a generic "deploy
rtvi-cv with N streams" query breaks the deploy rubric and the
user's `/metrics` expectations. See
[`references/pipeline-config.md`](references/pipeline-config.md)
§ "Defaults — the skill is static-mode by default" for the full
rationale.
- ❌ A one-line `✔ App ready in Ns, N streams, fps total Y` in place of
the Step 5 Results box.
- ❌ ASCII box-drawing chars (`+`, `-`, `=`, `*`) instead of light
box-drawing chars (`┌ ─ ┐ │ └ ┘`).
- ❌ Skipping Step 6 on the assumption "the user knows what to do next".
- ❌ After Step 6, dumping a markdown wall of prose + multiple curl
blocks + a closing "want me to run any of these?" — that's the
shape the agent falls back to and it bypasses both the 11.d menu
and the per-API-call box. The user picks from a menu; the skill
shows the resolved API box; the skill runs it. No free-text Q.
- ❌ Step 4 overview collapses — these are explicitly banned by the
deploy doc's Step 4 content rule:
- `✔ Batch size 3 (tile grid: 1×3)` → required: 5 separate rows
(`[streammux] batch-size=3`, `[primary-gie] batch-size=3`,
`[source-list] max-batch-size=3`, `[tiled-display] rows=1`,
`[tiled-display] columns=3`).
- `✔ Output sink eglsink` → required: one row per sink key
(4 keys for eglsink, e.g. `[sink0] enable=1`, `type=2`,
`sync=0`, `qos=0` — read apply-config.md for the exact list).
- `✔ Sources static (3 streams, http-port=9000)` → required: six
annotated `[source-list]` rows.
- `✔ Tile grid 1 row × 3 cols` (single row) → required: two
rows, `[tiled-display] rows=1` and `[tiled-display] columns=3`.
## Universal box format
The geometry contract for every step-exit box (Step 1 through Step 5
Results). The same shape across every box; only the **title** and the
**body rows** change per step.
- **Width: 128 chars** corner-to-corner — `┌` at column 1, `┐` at
column 128. Wider terminals leave the box flush-left; do not stretch
it. Inner content area is **124 chars** (with one space margin on
each side inside the `│` borders).
- **Light box-drawing chars only**: `┌ ─ ┐ │ └ ┘`. No `+`, `-`, `=`,
`*` ASCII fallbacks.
- **Top border — title CENTERED**: `┌` + N₁ dashes + `␣` + title + `␣`
+ N₂ dashes + `┐`, where `N₁ + N₂ + len(title) + 2 = 126`. Distribute
the pad: `N₁ = floor((126 − len(title) − 2) / 2)`,
`N₂ = 126 − len(title) − 2 − N₁`. N₁ and N₂ differ by at most 1.
- **Body**: one `│ <content padded to inner-content 124> │` per fact.
Each fact line uses the ` ✔ <key-padded-to-13> <value>` form (two
spaces in, glyph, key right-padded to 13, two spaces, value).
- **Blank lines between groups**: render `│ <124 spaces> │` between
logical groups (e.g. Identity / Model / Videos in Step 1) so the
user can scan the box at a glance.
- **Bottom border**: `└` + 126 dashes + `┘` — solid border, no title.
Standard step titles (used at the top of each step's box):
```
┌─────────────────────────────────────────────────────── Deploy targets ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── Pipeline configuration ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── Container ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── Apply configuration ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── Perception Application — Plan ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── Perception Application — Results ──────────────────────────────────────────────┐
```
Per-step content rules (which rows go in which box, mode-aware row
hiding, the apply-config sectioned layout, the Step 5 PLAN-then-RESULT
pattern, the Step 3 `docker run` synthesis requirement) live in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule" — read those when rendering the
corresponding step.
## Quick triggers (mnemonic)
| Phrase | Flow |
|--------|------|
| `deploy rtvicv warehouse 2d with 4 streams and display` | DEPLOY |
| `run smartcity gdino on gpu 1` | DEPLOY |
| `stop the perception container` | TEARDOWN (deploy doc) |
| `rtvi-cv healthcheck failing` | DEBUG (deploy doc + troubleshooting) |
| `add a stream to rtvi-cv` | API USAGE |
| `is rtvi-cv ready on localhost:9000` | API USAGE |
| `get rtvi-cv metrics` | API USAGE |
| `generate text embeddings via rtvi-cv` | API USAGE |
bump:1
すべてのファイル
51件のファイルvss-deploy-detection-tracking-2dをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-detection-tracking-2d # Copy SKILL.md to your .claude/skills/ directory
コピー





家
