vss-deploy-detection-tracking-2d
NVIDIA/skills
Implementa, depura y gestiona el microservicio de detección y seguimiento 2D RTVI-CV, y utiliza su API REST para la gestión de flujos, las comprobaciones de estado y las métricas.
...Expandir todoObjetivo
Implementar, depurar y poner en funcionamiento el microservicio 2D de detección y seguimiento RTVI-CV, así como gestionar su API REST.
Requisitos previos
- Una implementación activa de VSS accesible en
$HOST_IP(consultevss-deploy-profileyreferences/). - Credenciales de NGC en
$NGC_CLI_API_KEYy$NVIDIA_API_KEYpara cualquier descarga de imágenes. curl,jqy Docker disponibles en el sistema que realiza la llamada.
Instrucciones
Sigue las tablas de enrutamiento y los flujos de trabajo paso a paso que se indican a continuación. Cada sección que termine en «workflow», «quick start» o «flow» debe ejecutarse de arriba abajo. El material de referencia detallado se encuentra en references/ y los scripts de ayuda en scripts/; ejecútalos mediante run_script cuando la skill apunte a un script por su nombre.
Ejemplos
Los ejemplos completos y funcionales se encuentran en la carpeta «evals/» (cada manifiesto *.json contiene un escenario ejecutable) y, además, integrados en los bloques «curl» de cada flujo de trabajo que se muestran a continuación. Ejecuta una evaluación de Nivel 3 con «nv-base validate para reproducirlos.
Limitaciones
- Requiere que el perfil VSS o microservicio correspondiente esté implementado y sea accesible desde el llamante.
- Los modelos alojados en NGC y los NIM pueden estar sujetos a límites de tasa, requisitos de memoria de la GPU y restricciones de licencia.
- Los límites de concurrencia, memoria de GPU y almacenamiento dependen del hardware del host y del archivo de composición del perfil.
Solución de problemas
- Error: la llamada REST devuelve «conexión rechazada». Causa: el microservicio de destino no se está ejecutando. Solución: comprueba
/docso/health; vuelve a implementar mediantevss-deploy-profileo la habilidadvss-deploy-*correspondiente. - Error: código HTTP 401/403 en las consultas a NGC. Causa: falta
la clave NGC_CLI_API_KEYo ha caducado. Solución:inicia sesión en Docker en nvcr.ioy vuelve a exportar la clave antes de volver a intentarlo. - Error: el contenedor está sin memoria (OOM) o no se carga el modelo. Causa: memoria de GPU insuficiente para el perfil seleccionado. Solución: cambiar a una variante más pequeña o liberar GPU mediante
«docker compose down».
RTVI-CV — Detección y seguimiento (habilidad unificada)
Habilidad unificada para el microservicio Real Time Video Intelligence CV (RTVI-CV). Dos superficies de acción en una sola habilidad:
- Implementar / operar / depurar / desmontar el contenedor RTVI-CV localmente → véase
references/deploy-vss-detection-tracking-2d.md - Llamar a la API REST de RTVI-CV (flujos, estado, métricas, incrustaciones) en una instancia en ejecución → véase
references/usage-vss-detection-tracking-2d.md
Servicio:
rtvi-cv(metropolis_perception_app) Imagen:nvcr.io/— proporcionada por el usuario en el momento de la implementación Puerto REST:/ : 9000(/api/v1—/live,/ready,/startup,/metrics,/stream/add,/stream/remove, incrustaciones) Hardware: GPU dedicada x86/aarch64 (T4, A100, L40, H100, B200, RTX), SBSA (Spark, Grace-Hopper), Jetson (Thor, Orin, Xavier)
Enrutamiento de acciones: elegir una vez por invocación
| Intención del usuario (ejemplos de formulación) | Flujo | Cargar esta referencia |
|---|---|---|
implementar rtvi-cv warehouse 2d, ejecutar rtvicv warehouse-3d con 4 flujos, iniciar smartcity gdino, iniciar la aplicación de percepción, activar sparse4d |
IMPLANTAR | references/deploy-vss-detection-tracking-2d.md |
detener rtvi-cv, desmontar, cerrar el contenedor de percepción, limpiar rtvicv-perception-docker |
DESMONTAJE (gestionado por el documento de implementación → «Selección de modo») | references/deploy-vss-detection-tracking-2d.md + references/teardown-flow.md |
Comprobar los registros de rtvi-cv, diagnosticar el fallo de rtvi-cv, solucionar el fallo de la comprobación de estado, rtvi-cv no se inicia |
DEPURACIÓN | references/deploy-vss-detection-tracking-2d.md + references/troubleshooting.md |
Añadir una transmisión, eliminar una cámara, listar transmisiones, comprobar el estado, ver si rtvi-cv está listo, obtener métricas, comprobar los FPS, comprobar el uso de la GPU, generar incrustaciones de texto, llamar a la API de rtvi-cv |
USO DE LA API | references/usage-vss-detection-tracking-2d.md + references/api-reference.md |
Regla de selección: compara la frase del usuario con la tabla anterior y carga inmediatamente el archivo de referencia correspondiente. No mezcles los flujos: «DEPLOY» asume que aún no hay ningún contenedor en ejecución; «USO DE LA API» asume que el contenedor ya se está ejecutando en http://.
Si la intención es realmente ambigua (por ejemplo, el usuario dice simplemente «Quiero usar rtvi-cv»), plantea una pregunta: ¿desplegar una nueva instancia o llamar a una que ya esté en ejecución?
Qué se encuentra dónde
vss-deploy-detection-tracking-2d/
├── SKILL.md # este archivo (enrutamiento + contratos)
├── assets/ # archivos de datos (deploy-defaults.yml — única fuente de verdad para etiquetas / referencias / rutas / GPU)
├── evals/ # Manifiestos de evaluación de nivel 3 (deploy-evals.json, usage-evals.json)
├── scripts/ # 23 scripts auxiliares en bash y Python (consulta `scripts/` para ver el inventario completo)
└── references/ # manuales de flujo de trabajo (implementación / uso de la API / desmontaje / resolución de problemas / …)
Para consultar el inventario completo por archivo y lo que abarca cada referencia, véase
references/workflow-reference.md.
Todos los scripts se ejecutan desde la raíz de la habilidad a través de $SKILL_DIR/scripts/: las rutas que aparecen en la documentación de referencia de «deploy» se conservan tal cual y se resuelven correctamente cuando el agente se ejecuta desde la raíz de la habilidad.
Scripts disponibles
Las funciones auxiliares se encuentran en scripts/ y se invocan desde la raíz de la skill por su nombre —
invoca cada una mediante run_script("scripts/ para que el agente registre una
invocación correcta de la herramienta.
Para ver el inventario completo de utilidades (caché, comprobaciones de GPU, configuración), consulta
scripts/; la opción --help de cada script describe sus argumentos.
Cómo utilizar esta habilidad
- Lee primero este archivo. Solo se encarga del enrutamiento; no contiene flujos de trabajo.
- Compara la intención del usuario con la tabla de enrutamiento anterior.
- Carga exactamente un documento de referencia (DEPLOY o API USAGE). No precargues ambos: cada referencia es extensa y contiene su propio contrato completo.
- Sigue al pie de la letra la referencia cargada. Los documentos de referencia son los contratos conservados byte a byte de las habilidades predecesoras
vss-deploy-detection-tracking-2d(deploy/teardown/debug) yrtvicv-api(API REST): se conservan todos los pasos, el orden de las reglas de procesamiento por lotes en Bash, las reglas de renderizado de cuadros y el contratoAskQuestion. - Para DEPLOY, el documento de referencia impone su propio contrato de inicio: acuse de recibo de una línea → llamada a la herramienta de planificación (matriz
TodoWritede 5 tareas pendientes, O 5 llamadas sucesivasa TaskCreateen el código Claude más reciente) → pregunta del paso 1. No narres, no realices comprobaciones previas y nunca muestres «cargando TodoWrite/TaskCreate» ni ningún texto sobre la resolución diferida de herramientas: la herramienta de planificación se carga de forma silenciosa.
Contrato de salida — Flujo DEPLOY
Al ejecutar el flujo DEPLOY / TEARDOWN / DEBUG, el agente DEBE cumplir los cuatro puntos siguientes en cada implementación correcta. Estos son el único canal de retroalimentación del usuario entre pasos; omitir cualquiera de ellos constituye una regresión de comportamiento.
- Muestra la salida de cada paso en un recuadro de ancho fijo: Paso 1 : Destinos de implementación,
Paso 2: Configuración del pipeline, Paso 3: Contenedor, Paso 4:
Aplicar configuración, Paso 5: Plan + Resultados. No solo el
resumen final. El recuadro es el resguardo del paso para el usuario. La geometría es fija (véase
el apartado «Formato universal del recuadro» más abajo). Las reglas de contenido por paso (qué
filas van dentro de cada recuadro) se encuentran en
references/deploy-vss-detection-tracking-2d.mdbajo «Regla de contenido del recuadro del paso N». - Tras el recuadro de Resultados del paso 5, ejecuta la pregunta del paso 6
AskUserQuestiondereferences/next-steps.md, § «11.c» — nunca la sustituyas por una lista con viñetas de «Próximos pasos» de formato libre. El menú es el punto de salida de la implementación: permite al usuario ejecutar métricas, gestionar flujos, seguir registros o desmontar el entorno con un solo clic, en lugar de tener que recordar direcciones URL de curl. - Una vez que el usuario haya seleccionado un bucket del Paso 6, ejecuta la pregunta de seguimiento
«AskUserQuestion»dereferences/next-steps.md§ «11.d»; nunca la sustituyas por texto explicativo + ejemplos de curl listos para copiar + una pregunta de texto libre del tipo «¿quieres que ejecute X?». Cada bucket tiene su propio menú de acciones concretas; el usuario elige la acción y, a continuación, la skill muestra el cuadro de la API y ejecuta el comando curl. Acciones de seguimiento por bucket:- Gestionar flujos → Añadir / Eliminar / Listar. La opción «Eliminar» genera sus
opciones dinámicamente a partir de
/stream/get-stream-info: una opción por cada flujo activo etiquetado como, además de «Eliminar TODO» cuando· ACTIVE > 1(especificación completa: §«remove_streamssub-flow»). - Detener la implementación → Detener la aplicación / Detener el contenedor / Desmontaje completo.
- Comprobar métricas y FPS → sin seguimiento; ejecutar
collect_metrics.shinmediatamente después de mostrar el cuadro de la API/api/v1/metrics. - Comprobar estado de actividad / disponibilidad → sin seguimiento; sondear los tres puntos finales de estado tras mostrar sus cuadros de API.
- Gestionar flujos → Añadir / Eliminar / Listar. La opción «Eliminar» genera sus
opciones dinámicamente a partir de
- Mostrar el contenido COMPLETO de cada paso, no una fila de resumen:
mostrar el recuadro es necesario, pero no suficiente. Cada paso tiene una
especificación de composición de filas en
references/deploy-vss-detection-tracking-2d.mdbajo «Regla de contenido del recuadro del paso N». El paso 4 (Aplicar configuración) es donde el agente falla con mayor frecuencia: su lista canónica de claves por caso de uso se encuentra enreferences/apply-config.md§ «Lista completa de edición por caso de uso», y el agente DEBE emitir una✔ [sección] clave=valor —filade anotaciónpor clave en esa tabla para el caso de uso activo + ajustes. Una sección con 5 claves → 5 filas; una sección con 6 claves → 6 filas. Nunca una sola fila de resumen por sección.
Prohibido (estos son los atajos a los que recurre el agente bajo presión, y que perjudican la experiencia de usuario):
- ❌ Narración interna sobre la carga de herramientas. Nunca muestres «Tengo que cargar
TodoWrite (una herramienta diferida que la habilidad solicita para el widget de tareas)»,
«Cargando TaskCreate…», «Llamando a ToolSearch para la herramienta de planificación…»,
ni ningún otro texto sobre la resolución, carga o obtención de herramientas diferidas.
El agente carga las herramientas de forma silenciosa. El usuario solo ve la línea de resumen
✔ «seguida del widget; nunca ningún elemento de apoyo relacionado con la resolución de herramientas.» - ❌ Agrupar los 5 pasos de implementación en un único campo
dedescripciónde TaskCreate. CuandoTaskCreatesea la herramienta de planificación disponible, realiza 5 llamadasa TaskCreatepor separado, una tras otra (una por paso). Consultareferences/task-list.md, apartado «Llamadas inicialesa TaskCreate» para ver la plantilla literal. La misma regla se aplica aTodoWrite: una sola llamada con las 5 tareas pendientes en la matriztodos:[…]; nunca una sola tarea cuyocontenidosea una lista de varias líneas. - ❌ Elegir el modo de flujo
dinámicode forma implícita. El valor predeterminado de la skill esstream_mode=static: el agente integra las URLfile://detectadas automáticamente en el bloque[source-list]de la configuración principal de DS antes de que se inicie la aplicación. Cambia adinámicosolo cuando el usuario lo solicite explícitamente («añadir flujos más tarde vía REST», «usar modo de flujo dinámico») O cuando seleccionedinámicoen el paso 2 de AskQuestion. Seleccionarel modo dinámicopara una consulta genérica del tipo «implementar rtvi-cv con N flujos» incumple las directrices de implementación y las expectativas del usuarioen cuanto a métricas. Véasereferences/pipeline-config.md§ «Valores predeterminados: la skill está en modo estático por defecto» para conocer la justificación completa. - ❌ Una sola línea
✔ «Aplicación lista en Ns, N flujos, Y fps en total»en lugar del cuadro de resultados del paso 5. - ❌ Caracteres ASCII para dibujar cuadros (
+,-,=,*) en lugar de caracteres ligeros para dibujar cuadros (┌ ─ ┐ │ └ ┘). - ❌ Saltar el paso 6 partiendo de la suposición de que «el usuario sabe qué hacer a continuación».
- ❌ Tras el paso 6, mostrar un bloque de texto en Markdown + varios bloques de curl + una pregunta final del tipo «¿quieres que ejecute alguno de estos?»: esa es la forma en la que recurre el agente y que omite tanto el menú 11.d como el cuadro por llamada a la API. El usuario elige de un menú; la skill muestra el cuadro de la API resuelta; la skill la ejecuta. No hay preguntas de texto libre.
- ❌ El resumen del paso 4 se contrae; esto está explícitamente prohibido por la
regla de contenido del paso 4 del documento de implementación:
✔ Tamaño de lote 3 (cuadrícula de mosaicos: 1×3)→ requerido: 5 filas separadas ([streammux] tamaño-lote=3,[primary-gie] tamaño-lote=3,[source-list] tamaño-máximo-del-lote=3,[tiled-display] filas=1,[tiled-display] columnas=3).✔ Destino de salida eglsink→ requerido: una fila por clave de destino (4 claves para eglsink, p. ej.,[sink0] enable=1,type=2,sync=0,qos=0— consulta apply-config.md para ver la lista exacta).✔ Fuentes estáticas (3 flujos, http-port=9000)→ se requieren: seis filas[source-list]anotadas.✔ Cuadrícula de mosaicos: 1 fila × 3 columnas(una sola fila) → se requieren: dos filas,[tiled-display] rows=1y[tiled-display] columns=3.
Formato universal de recuadros
El contrato geométrico para cada cuadro de salida de paso (del paso 1 al paso 5 Resultados). La misma forma en todos los cuadros; solo cambian el título y las filas del cuerpo en cada paso.
- Ancho: 128 caracteres de esquina a esquina —
┌en la columna 1,┐en la columna 128. Los caracteres terminales más anchos dejan el recuadro alineado a la izquierda; no lo estiren . El área de contenido interior es de 124 caracteres (con un margen de un espacio a cada lado dentro de los bordes│). - Solo caracteres ligeros para dibujar el recuadro:
┌ ─ ┐ │ └ ┘. No se utilizan los caracteres ASCII de reserva+,-,=,*. - Borde superior — título CENTRADO:
┌+ N₁ guiones +␣+ título +␣- N₂ guiones +
┐, dondeN₁ + N₂ + len(título) + 2 = 126. Distribuir el relleno:N₁ = floor((126 − longitud(título) − 2) / 2),N₂ = 126 − longitud(título) − 2 − N₁. N₁ y N₂ difieren como máximo en 1.
- N₂ guiones +
- Cuerpo: un
│por dato. Cada línea de datos utiliza el formato│ ✔(dos espacios de separación, glifo, clave alineada a la derecha hasta 13, dos espacios, valor). - Líneas en blanco entre grupos: mostrar
│ <124 spaces> │entre grupos lógicos (p. ej., Identidad / Modelo / Vídeos en el Paso 1) para que el usuario pueda examinar el recuadro de un vistazo. - Borde inferior:
└+ 126 guiones +┘— borde continuo, sin título.
Títulos estándar de los pasos (utilizados en la parte superior del recuadro de cada paso):
┌─────────────────────────────────────────────────────── Destinos de implementación ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── Configuración del pipeline ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── Contenedor ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── Aplicar configuración ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── Aplicación de percepción — Plan ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── Aplicación de percepción — Resultados ──────────────────────────────────────────────┐
Reglas de contenido por paso (qué filas van en cada recuadro, ocultación de filas según el modo,
el diseño seccionado de «apply-config», el patrón del Paso 5 «PLAN-then-RESULT»
, el requisito de síntesis «docker run» del Paso 3) se encuentran en
references/deploy-vss-detection-tracking-2d.md
bajo «Regla de contenido del recuadro del Paso N»; léelas al renderizar el
paso correspondiente.
Desencadenantes rápidos (mnemotécnicos)
| Frase | Flujo |
|---|---|
implementar rtvicv warehouse 2d con 4 flujos y mostrar |
DEPLOY |
ejecutar smartcity gdino en la GPU 1 |
IMPLANTAR |
detener el contenedor de percepción |
DESMONTAR (documentación de implementación) |
Fallo en la comprobación de estado de rtvi-cv |
DEPURAR (documentación de implementación + resolución de problemas) |
Añadir un flujo a rtvi-cv |
USO DE LA API |
¿Está rtvi-cv listo en localhost:9000? |
USO DE LA API |
Obtener métricas de rtvi-cv |
USO DE LA API |
Generar incrustaciones de texto mediante rtvi-cv |
USO DE LA API |
bump:1
---
name: vss-deploy-detection-tracking-2d
description: Deploy, debug, and operate the RTVI-CV 2D detection/tracking microservice and call its REST API for stream management, health checks, and metrics.
license: Apache-2.0
---
## Purpose
Deploy, debug, and operate the RTVI-CV detection / tracking 2D microservice and drive its REST API.
## Prerequisites
- Active VSS deployment reachable on `$HOST_IP` (see `vss-deploy-profile` and `references/`).
- NGC credentials in `$NGC_CLI_API_KEY` and `$NVIDIA_API_KEY` for any image pulls.
- `curl`, `jq`, and Docker available on the caller.
## Instructions
Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/` and helper scripts live in `scripts/` — call them via `run_script` when the skill points to a script by name.
## Examples
Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.
## Limitations
- Requires the matching VSS profile / microservice to be deployed and reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
## Troubleshooting
- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.
# RTVI-CV — Detection & Tracking (Unified Skill)
Unified skill for the **Real Time Video Intelligence CV (RTVI-CV)** microservice. Two action surfaces in one skill:
- **Deploy / operate / debug / tear down** the RTVI-CV container locally → see [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
- **Call the RTVI-CV REST API** (streams, health, metrics, embeddings) on a running instance → see [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md)
> **Service**: `rtvi-cv` (`metropolis_perception_app`)
> **Image**: `nvcr.io/<org>/<repo>:<tag>` — user-supplied at deploy time
> **REST port**: `9000` (`/api/v1` — `/live`, `/ready`, `/startup`, `/metrics`, `/stream/add`, `/stream/remove`, embeddings)
> **Hardware**: x86/aarch64 dGPU (T4, A100, L40, H100, B200, RTX), SBSA (Spark, Grace-Hopper), Jetson (Thor, Orin, Xavier)
---
## Action routing — pick once per invocation
| User intent (sample phrasing) | Flow | Load this reference |
|-------------------------------|------|---------------------|
| `deploy rtvi-cv warehouse 2d`, `run rtvicv warehouse-3d with 4 streams`, `start smartcity gdino`, `launch perception app`, `bring up sparse4d` | **DEPLOY** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) |
| `stop rtvi-cv`, `tear down`, `kill the perception container`, `cleanup rtvicv-perception-docker` | **TEARDOWN** (handled by deploy doc → "Mode Selection") | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/teardown-flow.md`](references/teardown-flow.md) |
| `check rtvi-cv logs`, `diagnose rtvi-cv crashing`, `troubleshoot healthcheck failing`, `rtvi-cv won't start` | **DEBUG** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/troubleshooting.md`](references/troubleshooting.md) |
| `add a stream`, `remove camera`, `list streams`, `health check`, `is rtvi-cv ready`, `get metrics`, `what's the FPS`, `check GPU usage`, `generate text embeddings`, `call rtvi-cv api` | **API USAGE** | [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md) + [`references/api-reference.md`](references/api-reference.md) |
**Selection rule:** match the user's phrasing against the table above and immediately load the corresponding reference file. Do not mix the flows — DEPLOY assumes no running container yet; API USAGE assumes the container is already running on `http://<host>:9000`.
If intent is genuinely ambiguous (e.g., the user says just "I want to use rtvi-cv"), ask one `AskQuestion`: deploy a new instance, or call an already-running one?
---
## What lives where
```
vss-deploy-detection-tracking-2d/
├── SKILL.md # this file (routing + contracts)
├── assets/ # data files (deploy-defaults.yml — single source of truth for tags / refs / paths / GPU)
├── evals/ # Tier-3 eval manifests (deploy-evals.json, usage-evals.json)
├── scripts/ # 23 bash + python helpers (see `scripts/` for the full inventory)
└── references/ # workflow runbooks (deploy / api-usage / teardown / troubleshooting / …)
```
For the full per-file inventory and what each reference covers, see
[`references/workflow-reference.md`](references/workflow-reference.md).
All scripts are invoked from the skill root via `$SKILL_DIR/scripts/<name>` — paths inside the deploy reference doc are preserved verbatim and resolve correctly when the agent runs from skill root.
---
## Available Scripts
Helpers live in `scripts/` and are invoked from the skill root by name —
call each via `run_script("scripts/<name>")` so the agent records a
proper tool invocation.
| Script | Purpose | Arguments |
| --- | --- | --- |
| `load_defaults.sh` | Detect platform (x86 dGPU / SBSA / Jetson) and resolve YAML defaults from `assets/deploy-defaults.yml`. | `--usecase <name>` |
| `fetch_resources.sh` | Download + extract NGC resources, scan for layout. | `--ngc-ref <ref>` (optional) |
| `apply_in_container.sh` | Host-side wrapper for Step 4 (`apply_config.sh` inside the running container). | `<container_name>` |
| `apply_config.sh` | In-container path-substitution, batch, sink, sources, engine cache. | `<usecase> <stream_count> <sink_type>` |
| `start_app_in_container.sh` | Host-side wrapper for Step 5 (`run_app_and_wait.sh`). | `<container_name>` |
| `run_app_and_wait.sh` | In-container app launch + readiness + metrics + log. | `<config_path>` |
| `add_streams.sh` / `update_stream_sources.sh` | REST stream lifecycle for Step 6. | `<rtsp_or_file_uri>...` |
| `collect_metrics.sh` | Pull `/api/v1/metrics` snapshot. | none |
| `discover_streams.sh` | Enumerate active streams via `/stream/get-stream-info`. | none |
| `synthesize_docker_run.sh` | Print the platform-correct `docker run` line for the resolved env. | none |
| `render_box.sh` | Render the fixed-width step receipt. | `<step_label>` |
| `calibration_manager.py` | Manage calibration artefacts + per-use-case engine cache invalidation. | `--usecase <name> --reset` |
For the full inventory of helpers (cache, GPU checks, setup) browse
`scripts/`; each script's `--help` describes its arguments.
## How to use this skill
1. **Read this file first.** It only routes — it does not contain workflows.
2. **Match the user's intent** against the routing table above.
3. **Load exactly one reference doc** (DEPLOY or API USAGE). Don't preload both — each reference is large and contains its own full contract.
4. **Follow the loaded reference exactly.** The reference docs are the byte-for-byte preserved contracts from the predecessor skills `vss-deploy-detection-tracking-2d` (deploy/teardown/debug) and `rtvicv-api` (REST API) — every step ordering invariant, bash-batching rule, box-rendering rule, and `AskQuestion` contract is retained.
5. **For DEPLOY**, the reference doc enforces its own startup contract: one-line acknowledgement → planning-tool call (`TodoWrite` array of 5 todos, OR 5 successive `TaskCreate` calls on newer Claude Code) → Step 1 question. Do not narrate, do not pre-flight, and never print "loading TodoWrite/TaskCreate" or any deferred-tool resolution prose — the planning tool is loaded silently.
---
## Output contract — DEPLOY flow
When running the DEPLOY / TEARDOWN / DEBUG flow, the agent MUST honour
all four items below on every successful deploy. These are the user's
only feedback channel between steps; skipping any of them is a
behaviour regression.
1. **Render every step's exit in a fixed-width box** — Step 1 *Deploy
targets*, Step 2 *Pipeline configuration*, Step 3 *Container*, Step 4
*Apply configuration*, Step 5 *Plan* + *Results*. Not just the final
summary. The box is the user's step receipt. Geometry is fixed (see
§ "Universal box format" below). Per-step **content** rules (what
rows go inside each box) live in [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule".
2. **After the Step 5 Results box, issue the Step 6 `AskUserQuestion`**
from [`references/next-steps.md`](references/next-steps.md) § "11.c"
— never replace it with a free-form *Next steps* bullet list. The
menu is the deploy's exit handle: it lets the user run metrics,
manage streams, tail logs, or tear down with one click instead of
having to remember curl URLs.
3. **After the user picks a Step 6 bucket, issue the follow-up
`AskUserQuestion`** from [`references/next-steps.md`](references/next-steps.md)
§ "11.d" — never substitute prose + ready-to-copy curl examples + a
free-text "want me to run X?" question. Each bucket has its own
menu of concrete actions; the user picks the action, then the skill
emits the API box and runs the curl. Per-bucket follow-ups:
- **Manage streams** → Add / Remove / List. **Remove builds its
options dynamically from `/stream/get-stream-info`** — one option
per active stream labelled `<camera_id> · <camera_url>` plus
"Remove ALL" when `ACTIVE > 1` (full spec: § "`remove_streams`
sub-flow").
- **Stop the deployment** → Stop app / Stop container / Full teardown.
- **Check metrics & FPS** → no follow-up; run `collect_metrics.sh`
directly after printing the `/api/v1/metrics` API box.
- **Check liveness / readiness** → no follow-up; probe all three
health endpoints after printing their API boxes.
4. **Render the FULL per-step content, not an overview row** —
rendering the box is necessary but not sufficient. Each step has a
row composition spec in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule". **Step 4 (Apply configuration) is
where the agent collapses most often** — its canonical
per-use-case key list lives in
[`references/apply-config.md`](references/apply-config.md)
§ "Per-use-case complete edit list", and the agent MUST emit one
`✔ [section] key=value — annotation` row per key in that table for
the active use case + settings. A section with 5 keys → 5 rows; a
section with 6 keys → 6 rows. Never one overview row per section.
Forbidden (these are the shortcuts the agent falls back to under
pressure, and they break the user's UX):
- ❌ **Internal tool-loading narration.** Never print "I need to load
TodoWrite (a deferred tool the skill calls for the task widget)",
"Loading TaskCreate…", "Calling ToolSearch for the planning tool…",
or any other text about resolving / loading / fetching deferred tools.
The agent loads tools **silently**. The user only ever sees the `✔
<pinned-values>` summary line followed by the widget — never any
scaffolding around tool resolution.
- ❌ **Collapsing all 5 deploy steps into a single `TaskCreate`'s
`description` field.** When `TaskCreate` is the available planning
tool, issue **5 separate `TaskCreate` calls** back-to-back (one per
step). See `references/task-list.md` § "Initial `TaskCreate` calls"
for the verbatim template. Same rule for `TodoWrite` — one call with
all 5 todos in the `todos:[…]` array; never one todo whose `content`
is a multi-line list.
- ❌ **Silently choosing `dynamic` stream-mode.** The skill default is
`stream_mode=static` — the agent bakes auto-discovered `file://` URLs
into the DS main config's `[source-list]` block before app start.
Switch to `dynamic` only when the user explicitly asks ("add streams
later via REST", "use dynamic stream mode") OR when they pick `dynamic`
in the Step 2 AskQuestion. Picking `dynamic` for a generic "deploy
rtvi-cv with N streams" query breaks the deploy rubric and the
user's `/metrics` expectations. See
[`references/pipeline-config.md`](references/pipeline-config.md)
§ "Defaults — the skill is static-mode by default" for the full
rationale.
- ❌ A one-line `✔ App ready in Ns, N streams, fps total Y` in place of
the Step 5 Results box.
- ❌ ASCII box-drawing chars (`+`, `-`, `=`, `*`) instead of light
box-drawing chars (`┌ ─ ┐ │ └ ┘`).
- ❌ Skipping Step 6 on the assumption "the user knows what to do next".
- ❌ After Step 6, dumping a markdown wall of prose + multiple curl
blocks + a closing "want me to run any of these?" — that's the
shape the agent falls back to and it bypasses both the 11.d menu
and the per-API-call box. The user picks from a menu; the skill
shows the resolved API box; the skill runs it. No free-text Q.
- ❌ Step 4 overview collapses — these are explicitly banned by the
deploy doc's Step 4 content rule:
- `✔ Batch size 3 (tile grid: 1×3)` → required: 5 separate rows
(`[streammux] batch-size=3`, `[primary-gie] batch-size=3`,
`[source-list] max-batch-size=3`, `[tiled-display] rows=1`,
`[tiled-display] columns=3`).
- `✔ Output sink eglsink` → required: one row per sink key
(4 keys for eglsink, e.g. `[sink0] enable=1`, `type=2`,
`sync=0`, `qos=0` — read apply-config.md for the exact list).
- `✔ Sources static (3 streams, http-port=9000)` → required: six
annotated `[source-list]` rows.
- `✔ Tile grid 1 row × 3 cols` (single row) → required: two
rows, `[tiled-display] rows=1` and `[tiled-display] columns=3`.
## Universal box format
The geometry contract for every step-exit box (Step 1 through Step 5
Results). The same shape across every box; only the **title** and the
**body rows** change per step.
- **Width: 128 chars** corner-to-corner — `┌` at column 1, `┐` at
column 128. Wider terminals leave the box flush-left; do not stretch
it. Inner content area is **124 chars** (with one space margin on
each side inside the `│` borders).
- **Light box-drawing chars only**: `┌ ─ ┐ │ └ ┘`. No `+`, `-`, `=`,
`*` ASCII fallbacks.
- **Top border — title CENTERED**: `┌` + N₁ dashes + `␣` + title + `␣`
+ N₂ dashes + `┐`, where `N₁ + N₂ + len(title) + 2 = 126`. Distribute
the pad: `N₁ = floor((126 − len(title) − 2) / 2)`,
`N₂ = 126 − len(title) − 2 − N₁`. N₁ and N₂ differ by at most 1.
- **Body**: one `│ <content padded to inner-content 124> │` per fact.
Each fact line uses the ` ✔ <key-padded-to-13> <value>` form (two
spaces in, glyph, key right-padded to 13, two spaces, value).
- **Blank lines between groups**: render `│ <124 spaces> │` between
logical groups (e.g. Identity / Model / Videos in Step 1) so the
user can scan the box at a glance.
- **Bottom border**: `└` + 126 dashes + `┘` — solid border, no title.
Standard step titles (used at the top of each step's box):
```
┌─────────────────────────────────────────────────────── Deploy targets ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── Pipeline configuration ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── Container ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── Apply configuration ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── Perception Application — Plan ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── Perception Application — Results ──────────────────────────────────────────────┐
```
Per-step content rules (which rows go in which box, mode-aware row
hiding, the apply-config sectioned layout, the Step 5 PLAN-then-RESULT
pattern, the Step 3 `docker run` synthesis requirement) live in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule" — read those when rendering the
corresponding step.
## Quick triggers (mnemonic)
| Phrase | Flow |
|--------|------|
| `deploy rtvicv warehouse 2d with 4 streams and display` | DEPLOY |
| `run smartcity gdino on gpu 1` | DEPLOY |
| `stop the perception container` | TEARDOWN (deploy doc) |
| `rtvi-cv healthcheck failing` | DEBUG (deploy doc + troubleshooting) |
| `add a stream to rtvi-cv` | API USAGE |
| `is rtvi-cv ready on localhost:9000` | API USAGE |
| `get rtvi-cv metrics` | API USAGE |
| `generate text embeddings via rtvi-cv` | API USAGE |
bump:1
Todos los archivos
51 archivosInstalar vss-deploy-detection-tracking-2d
Descarga y descomprime los archivos de habilidades en tu directorio .claude/skills/.
Descargar ZIPClona el repositorio y copia los archivos de la habilidad a tu proyecto.
git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-detection-tracking-2d # Copy SKILL.md to your .claude/skills/ directory
Copiar





Hogar
