vss-deploy-detection-tracking-2d
NVIDIA/skills
RTVI-CV 2D 감지/추적 마이크로서비스를 배포, 디버그 및 운영하고, 스트림 관리, 상태 확인 및 메트릭을 위해 해당 REST API를 호출합니다.
...모든 것을 확장하십시오목적
RTVI-CV 탐지/추적 2D 마이크로서비스를 배포, 디버그 및 운영하고, 해당 REST API를 구동합니다.
필수 조건
$HOST_IP에서 접근 가능한 활성 VSS 배포 환경이 있어야 합니다(vss-deploy-profile및references/참조).- 이미지 가져오기를 위해
$NGC_CLI_API_KEY및$NVIDIA_API_KEY에NGC 자격 증명이 설정되어 있어야 합니다. - 호출 측에서
curl,jq및 Docker를 사용할 수 있어야 합니다.
지침
아래의 라우팅 테이블과 단계별 워크플로를 따르십시오. ‘워크플로’, ‘빠른 시작’ 또는 ‘흐름’으로 끝나는 각 섹션은 위에서 아래로 순서대로 실행하도록 되어 있습니다. 자세한 참조 자료는 references/에, 헬퍼 스크립트는 scripts/에 있습니다 . 스킬이 스크립트 이름을 지정하여 호출할 때는 run_script를 통해 해당 스크립트를 실행하십시오.
예제
완료된 엔드투엔드 예제는 evals/ 디렉터리에 저장되어 있으며(각 *.json 매니페스트에는 실행 가능한 시나리오가 포함되어 있음), 아래의 워크플로우별 curl 블록 내에도 인라인으로 포함되어 있습니다. nv-base validate 사용하여 Tier-3 평가를 실행하면 해당 예제를 재현할 수 있습니다.
제한 사항
- 호출 측에서 해당 VSS 프로필/마이크로서비스가 배포되어 있고 접근 가능해야 합니다.
- NGC에서 호스팅되는 모델 및 NIM은 사용량 제한, GPU 메모리 요구 사항 및 라이선스 제한의 적용을 받을 수 있습니다.
- 동시 실행, GPU 메모리 및 스토리지 제한은 호스트 하드웨어와 프로필의 컴포즈 파일에 따라 다릅니다.
문제 해결
- 오류: REST 호출에서 연결 거부(connection refused)가 반환됩니다. 원인: 대상 마이크로서비스가 실행 중이 아닙니다. 해결 방법:
/docs또는/health를프로브하고,vss-deploy-profile또는 일치하는vss-deploy-*스킬을 통해 다시 배포하십시오. - 오류: NGC 풀에서 HTTP 401/403 오류가 발생합니다. 원인:
NGC_CLI_API_KEY가없거나 만료되었습니다. 해결 방법:docker login nvcr.io를실행하고 키를 다시 내보낸 후 다시 시도하십시오. - 오류: 컨테이너 OOM 또는 모델 로딩 실패. 원인: 선택한 프로필에 대한 GPU 메모리가 부족합니다. 해결 방법: 더 작은 변형으로 전환하거나 `
docker compose down`을 통해 GPU 리소스를 확보하십시오.
RTVI-CV — 탐지 및 추적 (통합 스킬)
실시간 영상 인텔리전스 CV(RTVI-CV) 마이크로서비스를 위한 통합 스킬입니다. 하나의 스킬에 두 가지 작업 영역이 포함되어 있습니다:
- 로컬에서 RTVI-CV 컨테이너배포/운영/디버깅/종료 →
references/deploy-vss-detection-tracking-2d.md참조 - 실행 중인 인스턴스에서RTVI-CV REST API (스트림, 상태, 메트릭, 임베딩)호출 →
references/usage-vss-detection-tracking-2d.md참조
서비스:
rtvi-cv(metropolis_perception_app) 이미지:nvcr.io/— 배포 시 사용자가 제공 REST 포트:/ : 9000(/api/v1—/live,/ready,/startup,/metrics,/stream/add,/stream/remove, embeddings) 하드웨어: x86/aarch64 dGPU (T4, A100, L40, H100, B200, RTX), SBSA (Spark, Grace-Hopper), Jetson (Thor, Orin, Xavier)
액션 라우팅 — 호출당 한 번만 선택
| 사용자 의도 (예시 문구) | 흐름 | 이 레퍼런스 로드 |
|---|---|---|
rtvi-cv warehouse 2d 배포, 4개의 스트림으로 rtvicv warehouse-3d 실행, smartcity gdino 시작, 인식 앱 실행, sparse4d 불러오기 |
배포 | references/deploy-vss-detection-tracking-2d.md |
rtvi-cv 중지, 테어다운, 퍼셉션 컨테이너 종료, rtvicv-perception-docker 정리 |
종료 (배포 문서 → "모드 선택"에서 처리) | references/deploy-vss-detection-tracking-2d.md + references/teardown-flow.md |
rtvi-cv 로그 확인, rtvi-cv 충돌 진단, 상태 점검 실패 문제 해결, rtvi-cv가 시작되지 않음 |
DEBUG | references/deploy-vss-detection-tracking-2d.md + references/troubleshooting.md |
스트림 추가, 카메라 제거, 스트림 나열, 상태 확인, rtvi-cv 준비 상태 확인, 메트릭 가져오기, FPS 확인, GPU 사용량 확인, 텍스트 임베딩 생성, rtvi-cv API 호출 |
API 사용법 | references/usage-vss-detection-tracking-2d.md + references/api-reference.md |
선택 규칙: 사용자의 표현을 위 표와 대조하여 일치하는 경우, 해당 참조 파일을 즉시 로드합니다. 흐름을 혼동하지 마십시오. 'DEPLOY'는 아직 실행 중인 컨테이너가 없다고 가정하는 반면, 'API USAGE'는 컨테이너가 이미 http:// 실행 중이라고 가정합니다.
의도가 명백히 모호한 경우(예: 사용자가 단순히 “rtvi-cv를 사용하고 싶어요”라고 말하는 경우), AskQuestion을 한 번 요청합니다: 새 인스턴스를 배포할까요, 아니면 이미 실행 중인 인스턴스를 호출할까요?
구성 요소 위치
vss-deploy-detection-tracking-2d/
├── SKILL.md # 이 파일 (라우팅 + 계약)
├── assets/ # 데이터 파일 (deploy-defaults.yml — 태그 / 참조 / 경로 / GPU에 대한 단일 신뢰 소스)
├── evals/ # Tier-3 평가 매니페스트 (deploy-evals.json, usage-evals.json)
├── scripts/ # 23개의 bash 및 python 헬퍼 (전체 목록은 `scripts/` 참조)
└── references/ # 워크플로우 런북 (deploy / api-usage / teardown / troubleshooting / …)
파일별 전체 목록과 각 레퍼런스가 다루는 내용은
references/workflow-reference.md를 참조하십시오.
모든 스크립트는 $SKILL_DIR/scripts/ 경로를 통해 스킬 루트에서 호출됩니다
사용 가능한 스크립트
헬퍼는 scripts/ 디렉터리에 위치하며 스킬 루트에서 이름으로 호출됩니다.
각 헬퍼는 run_script("scripts/ 를 통해 호출해야 에이전트가
올바른 도구 호출 기록을 남길 수 있습니다.
모든 헬퍼(캐시, GPU 확인, 설정) 목록을 보려면
scripts/ 디렉터리를 살펴보세요. 각 스크립트의 --help 옵션에서 해당 인수를 확인할 수 있습니다.
이 스킬 사용 방법
- 먼저 이 파일을 읽어보세요. 이 파일은 라우팅 기능만 제공하며 워크플로는 포함되어 있지 않습니다.
- 사용자의 의도를 위의 라우팅 테이블과대조하십시오.
- 참조 문서 (DEPLOY 또는 API USAGE)를 정확히 하나만 불러오십시오. 두 문서를 모두 미리 불러오지 마십시오. 각 참조 문서는 용량이 크며 자체적인 전체 계약서를 포함하고 있습니다.
- 로드된 참조 문서를 정확히 따르십시오. 참조 문서는 이전 스킬
vss-deploy-detection-tracking-2d(deploy/teardown/debug) 및rtvicv-api(REST API)에서 바이트 단위로 그대로 보존된 계약서입니다. 모든 단계 순서 불변 조건, bash 배치 규칙, 박스 렌더링 규칙,AskQuestion계약이 그대로 유지됩니다. - DEPLOY의 경우, 참조 문서는 자체 시작 계약을 적용합니다: 한 줄의 확인 → 계획 도구 호출(5개의 할 일로 구성된
TodoWrite배열, 또는 최신 Claude Code에서는 5번의 연속적인TaskCreate호출) → 1단계 질문. 설명하지 말고, 사전 점검을 하지 말고, 절대 "loading TodoWrite/TaskCreate"나 지연된 도구 해결에 대한 문구를 출력하지 마십시오 — 계획 도구는 자동으로 로드됩니다.
출력 계약 — DEPLOY 흐름
DEPLOY / TEARDOWN / DEBUG 흐름을 실행할 때, 에이전트는 성공적인 배포 시마다 아래 네 가지 항목을 반드시 준수해야 합니다. 이는 단계 간에 사용자에게 제공되는 유일한 피드백 채널이며, 이 중 하나라도 생략하는 것은 행동 퇴행으로 간주됩니다.
- 각 단계의 종료 결과를 고정 너비 상자에 표시하십시오 — 1단계 배포
대상, 2단계 파이프라인 구성, 3단계 컨테이너, 4단계
구성 적용, 5단계 계획 + 결과. 최종
요약만 표시해서는 안 됩니다. 이 상자는 사용자에게 제공되는 단계별 확인서입니다. 크기는 고정되어 있습니다(아래
§ "범용 상자 형식" 참조). 단계별 콘텐츠 규칙(각 상자 안에 어떤
행이 들어가는지)은
references/deploy-vss-detection-tracking-2d.md의"Step N box content rule" 섹션에 있습니다. - 5단계 결과 상자 다음에는
references/next-steps.md의§ "11.c"에 있는 6단계AskUserQuestion을실행하십시오. — 절대로 자유 형식의 '다음 단계' 글머리 기호 목록으로 대체해서는 안 됩니다. 이 메뉴는 배포 프로세스의 종료 수단입니다. 이를 통해 사용자는 curl URL을 기억할 필요 없이 한 번의 클릭으로 메트릭 실행, 스트림 관리, 로그 확인 또는 시스템 종료를 수행할 수 있습니다. - 사용자가 Step 6 버킷을 선택한 후,
references/next-steps.md§ "11.d"에 있는 후속AskUserQuestion을실행하십시오 — 절대로 설명문 + 복사할 수 있는 curl 예제 + 자유 형식의 "X를 실행할까요?"라는 질문으로 대체해서는 안 됩니다. 각 버킷에는 고유한 구체적인 작업 메뉴가 있습니다. 사용자가 작업을 선택하면 스킬이 API 입력란을 표시하고 curl을 실행합니다. 버킷별 후속 조치:- 스트림 관리 → 추가 / 제거 / 목록. ‘제거’는
/stream/get-stream-info에서옵션을 동적으로 생성합니다 — 활성 스트림당 하나의 옵션으로,라벨은,· 이며 ACTIVE > 1일경우 “모두 제거”가 추가됩니다(전체 사양: §“remove_streams하위 흐름”). - 배포 중지 → 앱 중지 / 컨테이너 중지 / 전체 해체.
- 메트릭 및 FPS 확인 → 후속 조치 없음;
/api/v1/metricsAPI 상자 출력 직후에collect_metrics.sh를직접 실행합니다. - 활성 상태/준비 상태 확인 → 후속 조치 없음; 세 개의 상태 확인 엔드포인트 모두에 대해 API 박스를 출력한 후 프로브 수행.
- 스트림 관리 → 추가 / 제거 / 목록. ‘제거’는
- 개요 행이 아닌 단계별 전체 내용을 렌더링 —
박스 렌더링은 필요하지만 충분조건은 아닙니다. 각 단계에는
references/deploy-vss-detection-tracking-2d.md의 "Step N box content rule" 아래에 행 구성 사양이 있습니다. 4단계(구성 적용)는 에이전트가 가장 자주 오류가 발생하는 단계입니다. 해당 단계의 표준 사용 사례별 키 목록은references/apply-config.md§ "사용 사례별 전체 편집 목록"에 있으며, 에이전트는 반드시 하나를 출력해야 합니다✔ [section] key=value — 주석을 생성해야합니다. 활성 사용 사례 및 설정에 대해 해당 테이블의 키당 한 행씩 생성해야 합니다. 키가 5개인 섹션 → 5행; 키가 6개인 섹션 → 6행. 절대로 섹션당 하나의 개요 행을 생성해서는 안 됩니다.
금지 사항 (다음은 에이전트가 압박을 받을 때 의지하는 지름길이며, 사용자의 UX를 저해합니다):
- ❌ 내부 도구 로딩에 대한 설명. “TodoWrite(스킬이 작업 위젯을 위해 호출하는 지연 도구)를
로딩해야 합니다”,
“TaskCreate 로딩 중…”, “계획 도구를 위해 ToolSearch 호출 중…”,
또는 지연 도구의 해결/로딩/가져오기와 관련된 그 어떤 텍스트도 절대 출력해서는 안 됩니다.
에이전트는 도구를 자동으로 로드합니다. 사용자에게는
✔요약 줄과 그 뒤에 이어지는 위젯만 표시되며, 도구 해결과 관련된 어떤 보조 설명도 절대 표시되지 않습니다. - ❌ 5개의 배포 단계를 모두 하나의
TaskCreate의설명필드에 압축하는 것.TaskCreate가사용 가능한 계획 도구인 경우, 5개의 별도TaskCreate호출을 연속으로 수행하십시오(단계당 하나씩). 정확한 템플릿은references/task-list.md§ "InitialTaskCreatecalls" 를 참조하십시오.TodoWrite의경우에도 동일한 규칙이 적용됩니다.todos:[…]배열에 5개의 할 일 항목을 모두 포함하는 단일 호출을 사용해야 하며,내용이여러 줄로 구성된 목록인 단일 할 일 항목은 절대 사용해서는 안 됩니다. - ❌
동적스트림 모드를 암묵적으로 선택하는 행위. 스킬의 기본값은stream_mode=static입니다. 에이전트는 앱 시작 전에 자동 탐지된file://URL을 DS 메인 구성의[source-list]블록에 포함시킵니다. 사용자가 명시적으로 요청한 경우("나중에 REST를 통해 스트림 추가", "동적 스트림 모드 사용")이거나 2단계 AskQuestion에서동적 모드를선택한 경우에만동적모드로 전환하십시오. 일반적인 "N개의 스트림으로 rtvi-cv 배포" 쿼리에 대해동적 모드를선택하면 배포 기준과 사용자의/metrics에 대한기대치가 무너집니다. 자세한 이유는references/pipeline-config.md§ "기본값 — 스킬은 기본적으로 정적 모드"를 참조하십시오. - ❌ 5단계 결과 상자 대신 한 줄의
✔ "N초 내 앱 준비 완료, N 스트림, 총 Y fps"를 표시하는 경우. - ❌ 밝은 색상의
상자 그리기 문자(
┌ ─ ┐ │ └ ┘) 대신 ASCII 상자 그리기 문자(+,-,=,*)를 사용합니다. - ❌ “사용자가 다음에 무엇을 해야 할지 알고 있다”는 가정 하에 6단계를 건너뜁니다.
- ❌ 6단계 이후, 마크다운으로 작성된 긴 문장 + 여러 개의 curl 블록 + 마지막에 “이 중 실행할 것이 있나요?”라는 문구를 무더기로 쏟아내는 것 — 이것이 에이전트가 대체로 취하는 형태이며, 이는 11.d 메뉴와 API 호출별 상자를 모두 우회합니다. 사용자는 메뉴에서 선택하고, 스킬은 처리된 API 상자를 표시하며, 스킬이 이를 실행합니다. 자유 형식 질문은 없습니다.
- ❌ 4단계 개요가 접혀 있는 경우 — 이는 배포 문서의
4단계 콘텐츠 규칙에서 명시적으로 금지되어 있습니다:
✔ 배치 크기 3 (타일 그리드: 1×3)→ 필수: 5개의 별도 행 ([streammux] batch-size=3,[primary-gie] batch-size=3,[source-list] max-batch-size=3,[tiled-display] rows=1,[tiled-display] columns=3).✔ 출력 싱크 eglsink→ 필수: 싱크 키당 한 행 (eglsink의 경우 4개 키, 예:[sink0] enable=1,type=2,sync=0,qos=0— 정확한 목록은 apply-config.md를 참조하십시오).✔ 소스 static (스트림 3개, http-port=9000)→ 필수: 주석이 달린[source-list]행 6개.✔ 타일 그리드 1행 × 3열(단일 행) → 필수: 두 행,[tiled-display] rows=1및[tiled-display] columns=3.
범용 박스 형식
모든 단계 종료 박스(1단계부터 5단계 결과까지)에 대한 기하학적 계약. 모든 박스의 모양은 동일하며, 단계별로 제목과 본문 행만 변경됩니다.
- 너비: 대각선길이 128자 — 1열에
┌, 128열에┐. 더 넓은 종결자는 박스 좌측에 정렬되며, 박스를 늘리지 않습니다. 내부 콘텐츠 영역은 124자입니다(│ 경계선 안쪽 양쪽에 공백 여백 1자씩 포함). 더 넓은 터미널은 박스를 왼쪽 정렬 상태로 유지하며, 늘리지 않습니다. 내부 콘텐츠 영역은 124자입니다 (│테두리 안쪽 양쪽에 공백 여백 1개씩 포함). - 상자 그리기에 사용되는 가벼운 문자만 사용:
┌ ─ ┐ │ └ ┘.+,-,=,*와 같은ASCII 대체 문자는 사용하지 마십시오. - 상단 테두리 — 제목 중앙 정렬:
┌+ N₁ 개의 대시 +␣+ 제목 +␣- N₂ 개의 대시 +
┐, 여기서N₁ + N₂ + len(제목) + 2 = 126입니다. 채움 공간을 분배:N₁ = floor((126 − len(title) − 2) / 2),N₂ = 126 − len(title) − 2 − N₁. N₁과 N₂의 차이는 최대 1입니다.
- N₂ 개의 대시 +
- 본문: 사실 하나당
│하나씩. 각 사실 행은│ ✔형식을 사용합니다(두 공백, 기호, 오른쪽으로 13까지 채워진 키, 두 공백, 값). - 그룹 사이의 빈 줄: 논리적 그룹(예: 1단계의 ‘정체성 / 모델 / 동영상’) 사이에
│ <124 spaces> │을표시하여 사용자가 한눈에 내용을 파악할 수 있도록 합니다. - 하단 테두리:
└+ 126개의 대시 +┘— 실선 테두리, 제목 없음.
표준 단계 제목(각 단계 상자 상단에 사용):
┌─────────────────────────────────────────────────────── 배포 대상 ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── 파이프라인 구성 ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── 컨테이너 ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── 구성 적용 ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── 지각 적용 — 계획 ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── 지각 적용 — 결과 ──────────────────────────────────────────────┐
단계별 콘텐츠 규칙(어떤 행이 어떤 상자에 들어가는지, 모드에 따른 행
숨기기, apply-config 섹션 레이아웃, 5단계 PLAN-then-RESULT
패턴, 3단계의 Docker 실행 합성 요구 사항)은
references/deploy-vss-detection-tracking-2d.md
파일의 “Step N box content rule” 섹션에 있습니다. 해당 단계를 렌더링할 때
이 내용을 참조하십시오.
간편 트리거(기억법)
| 문구 | 흐름 |
|---|---|
4개의 스트림을 사용하여 rtvicv warehouse 2d를 배포하고 표시 |
DEPLOY |
GPU 1에서 smartcity gdino 실행 |
배포 |
퍼셉션 컨테이너 중지 |
TEARDOWN (문서 배포) |
rtvi-cv 상태 점검 실패 |
디버그 (배포 문서 + 문제 해결) |
rtvi-cv에 스트림 추가 |
API 사용법 |
rtvi-cv가 localhost:9000에서 준비되었는지 확인 |
API 사용법 |
rtvi-cv 메트릭 가져오기 |
API 사용법 |
rtvi-cv를 통해 텍스트 임베딩 생성하기 |
API 사용법 |
bump:1
---
name: vss-deploy-detection-tracking-2d
description: Deploy, debug, and operate the RTVI-CV 2D detection/tracking microservice and call its REST API for stream management, health checks, and metrics.
license: Apache-2.0
---
## Purpose
Deploy, debug, and operate the RTVI-CV detection / tracking 2D microservice and drive its REST API.
## Prerequisites
- Active VSS deployment reachable on `$HOST_IP` (see `vss-deploy-profile` and `references/`).
- NGC credentials in `$NGC_CLI_API_KEY` and `$NVIDIA_API_KEY` for any image pulls.
- `curl`, `jq`, and Docker available on the caller.
## Instructions
Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/` and helper scripts live in `scripts/` — call them via `run_script` when the skill points to a script by name.
## Examples
Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.
## Limitations
- Requires the matching VSS profile / microservice to be deployed and reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
## Troubleshooting
- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.
# RTVI-CV — Detection & Tracking (Unified Skill)
Unified skill for the **Real Time Video Intelligence CV (RTVI-CV)** microservice. Two action surfaces in one skill:
- **Deploy / operate / debug / tear down** the RTVI-CV container locally → see [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
- **Call the RTVI-CV REST API** (streams, health, metrics, embeddings) on a running instance → see [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md)
> **Service**: `rtvi-cv` (`metropolis_perception_app`)
> **Image**: `nvcr.io/<org>/<repo>:<tag>` — user-supplied at deploy time
> **REST port**: `9000` (`/api/v1` — `/live`, `/ready`, `/startup`, `/metrics`, `/stream/add`, `/stream/remove`, embeddings)
> **Hardware**: x86/aarch64 dGPU (T4, A100, L40, H100, B200, RTX), SBSA (Spark, Grace-Hopper), Jetson (Thor, Orin, Xavier)
---
## Action routing — pick once per invocation
| User intent (sample phrasing) | Flow | Load this reference |
|-------------------------------|------|---------------------|
| `deploy rtvi-cv warehouse 2d`, `run rtvicv warehouse-3d with 4 streams`, `start smartcity gdino`, `launch perception app`, `bring up sparse4d` | **DEPLOY** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) |
| `stop rtvi-cv`, `tear down`, `kill the perception container`, `cleanup rtvicv-perception-docker` | **TEARDOWN** (handled by deploy doc → "Mode Selection") | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/teardown-flow.md`](references/teardown-flow.md) |
| `check rtvi-cv logs`, `diagnose rtvi-cv crashing`, `troubleshoot healthcheck failing`, `rtvi-cv won't start` | **DEBUG** | [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md) + [`references/troubleshooting.md`](references/troubleshooting.md) |
| `add a stream`, `remove camera`, `list streams`, `health check`, `is rtvi-cv ready`, `get metrics`, `what's the FPS`, `check GPU usage`, `generate text embeddings`, `call rtvi-cv api` | **API USAGE** | [`references/usage-vss-detection-tracking-2d.md`](references/usage-vss-detection-tracking-2d.md) + [`references/api-reference.md`](references/api-reference.md) |
**Selection rule:** match the user's phrasing against the table above and immediately load the corresponding reference file. Do not mix the flows — DEPLOY assumes no running container yet; API USAGE assumes the container is already running on `http://<host>:9000`.
If intent is genuinely ambiguous (e.g., the user says just "I want to use rtvi-cv"), ask one `AskQuestion`: deploy a new instance, or call an already-running one?
---
## What lives where
```
vss-deploy-detection-tracking-2d/
├── SKILL.md # this file (routing + contracts)
├── assets/ # data files (deploy-defaults.yml — single source of truth for tags / refs / paths / GPU)
├── evals/ # Tier-3 eval manifests (deploy-evals.json, usage-evals.json)
├── scripts/ # 23 bash + python helpers (see `scripts/` for the full inventory)
└── references/ # workflow runbooks (deploy / api-usage / teardown / troubleshooting / …)
```
For the full per-file inventory and what each reference covers, see
[`references/workflow-reference.md`](references/workflow-reference.md).
All scripts are invoked from the skill root via `$SKILL_DIR/scripts/<name>` — paths inside the deploy reference doc are preserved verbatim and resolve correctly when the agent runs from skill root.
---
## Available Scripts
Helpers live in `scripts/` and are invoked from the skill root by name —
call each via `run_script("scripts/<name>")` so the agent records a
proper tool invocation.
| Script | Purpose | Arguments |
| --- | --- | --- |
| `load_defaults.sh` | Detect platform (x86 dGPU / SBSA / Jetson) and resolve YAML defaults from `assets/deploy-defaults.yml`. | `--usecase <name>` |
| `fetch_resources.sh` | Download + extract NGC resources, scan for layout. | `--ngc-ref <ref>` (optional) |
| `apply_in_container.sh` | Host-side wrapper for Step 4 (`apply_config.sh` inside the running container). | `<container_name>` |
| `apply_config.sh` | In-container path-substitution, batch, sink, sources, engine cache. | `<usecase> <stream_count> <sink_type>` |
| `start_app_in_container.sh` | Host-side wrapper for Step 5 (`run_app_and_wait.sh`). | `<container_name>` |
| `run_app_and_wait.sh` | In-container app launch + readiness + metrics + log. | `<config_path>` |
| `add_streams.sh` / `update_stream_sources.sh` | REST stream lifecycle for Step 6. | `<rtsp_or_file_uri>...` |
| `collect_metrics.sh` | Pull `/api/v1/metrics` snapshot. | none |
| `discover_streams.sh` | Enumerate active streams via `/stream/get-stream-info`. | none |
| `synthesize_docker_run.sh` | Print the platform-correct `docker run` line for the resolved env. | none |
| `render_box.sh` | Render the fixed-width step receipt. | `<step_label>` |
| `calibration_manager.py` | Manage calibration artefacts + per-use-case engine cache invalidation. | `--usecase <name> --reset` |
For the full inventory of helpers (cache, GPU checks, setup) browse
`scripts/`; each script's `--help` describes its arguments.
## How to use this skill
1. **Read this file first.** It only routes — it does not contain workflows.
2. **Match the user's intent** against the routing table above.
3. **Load exactly one reference doc** (DEPLOY or API USAGE). Don't preload both — each reference is large and contains its own full contract.
4. **Follow the loaded reference exactly.** The reference docs are the byte-for-byte preserved contracts from the predecessor skills `vss-deploy-detection-tracking-2d` (deploy/teardown/debug) and `rtvicv-api` (REST API) — every step ordering invariant, bash-batching rule, box-rendering rule, and `AskQuestion` contract is retained.
5. **For DEPLOY**, the reference doc enforces its own startup contract: one-line acknowledgement → planning-tool call (`TodoWrite` array of 5 todos, OR 5 successive `TaskCreate` calls on newer Claude Code) → Step 1 question. Do not narrate, do not pre-flight, and never print "loading TodoWrite/TaskCreate" or any deferred-tool resolution prose — the planning tool is loaded silently.
---
## Output contract — DEPLOY flow
When running the DEPLOY / TEARDOWN / DEBUG flow, the agent MUST honour
all four items below on every successful deploy. These are the user's
only feedback channel between steps; skipping any of them is a
behaviour regression.
1. **Render every step's exit in a fixed-width box** — Step 1 *Deploy
targets*, Step 2 *Pipeline configuration*, Step 3 *Container*, Step 4
*Apply configuration*, Step 5 *Plan* + *Results*. Not just the final
summary. The box is the user's step receipt. Geometry is fixed (see
§ "Universal box format" below). Per-step **content** rules (what
rows go inside each box) live in [`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule".
2. **After the Step 5 Results box, issue the Step 6 `AskUserQuestion`**
from [`references/next-steps.md`](references/next-steps.md) § "11.c"
— never replace it with a free-form *Next steps* bullet list. The
menu is the deploy's exit handle: it lets the user run metrics,
manage streams, tail logs, or tear down with one click instead of
having to remember curl URLs.
3. **After the user picks a Step 6 bucket, issue the follow-up
`AskUserQuestion`** from [`references/next-steps.md`](references/next-steps.md)
§ "11.d" — never substitute prose + ready-to-copy curl examples + a
free-text "want me to run X?" question. Each bucket has its own
menu of concrete actions; the user picks the action, then the skill
emits the API box and runs the curl. Per-bucket follow-ups:
- **Manage streams** → Add / Remove / List. **Remove builds its
options dynamically from `/stream/get-stream-info`** — one option
per active stream labelled `<camera_id> · <camera_url>` plus
"Remove ALL" when `ACTIVE > 1` (full spec: § "`remove_streams`
sub-flow").
- **Stop the deployment** → Stop app / Stop container / Full teardown.
- **Check metrics & FPS** → no follow-up; run `collect_metrics.sh`
directly after printing the `/api/v1/metrics` API box.
- **Check liveness / readiness** → no follow-up; probe all three
health endpoints after printing their API boxes.
4. **Render the FULL per-step content, not an overview row** —
rendering the box is necessary but not sufficient. Each step has a
row composition spec in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule". **Step 4 (Apply configuration) is
where the agent collapses most often** — its canonical
per-use-case key list lives in
[`references/apply-config.md`](references/apply-config.md)
§ "Per-use-case complete edit list", and the agent MUST emit one
`✔ [section] key=value — annotation` row per key in that table for
the active use case + settings. A section with 5 keys → 5 rows; a
section with 6 keys → 6 rows. Never one overview row per section.
Forbidden (these are the shortcuts the agent falls back to under
pressure, and they break the user's UX):
- ❌ **Internal tool-loading narration.** Never print "I need to load
TodoWrite (a deferred tool the skill calls for the task widget)",
"Loading TaskCreate…", "Calling ToolSearch for the planning tool…",
or any other text about resolving / loading / fetching deferred tools.
The agent loads tools **silently**. The user only ever sees the `✔
<pinned-values>` summary line followed by the widget — never any
scaffolding around tool resolution.
- ❌ **Collapsing all 5 deploy steps into a single `TaskCreate`'s
`description` field.** When `TaskCreate` is the available planning
tool, issue **5 separate `TaskCreate` calls** back-to-back (one per
step). See `references/task-list.md` § "Initial `TaskCreate` calls"
for the verbatim template. Same rule for `TodoWrite` — one call with
all 5 todos in the `todos:[…]` array; never one todo whose `content`
is a multi-line list.
- ❌ **Silently choosing `dynamic` stream-mode.** The skill default is
`stream_mode=static` — the agent bakes auto-discovered `file://` URLs
into the DS main config's `[source-list]` block before app start.
Switch to `dynamic` only when the user explicitly asks ("add streams
later via REST", "use dynamic stream mode") OR when they pick `dynamic`
in the Step 2 AskQuestion. Picking `dynamic` for a generic "deploy
rtvi-cv with N streams" query breaks the deploy rubric and the
user's `/metrics` expectations. See
[`references/pipeline-config.md`](references/pipeline-config.md)
§ "Defaults — the skill is static-mode by default" for the full
rationale.
- ❌ A one-line `✔ App ready in Ns, N streams, fps total Y` in place of
the Step 5 Results box.
- ❌ ASCII box-drawing chars (`+`, `-`, `=`, `*`) instead of light
box-drawing chars (`┌ ─ ┐ │ └ ┘`).
- ❌ Skipping Step 6 on the assumption "the user knows what to do next".
- ❌ After Step 6, dumping a markdown wall of prose + multiple curl
blocks + a closing "want me to run any of these?" — that's the
shape the agent falls back to and it bypasses both the 11.d menu
and the per-API-call box. The user picks from a menu; the skill
shows the resolved API box; the skill runs it. No free-text Q.
- ❌ Step 4 overview collapses — these are explicitly banned by the
deploy doc's Step 4 content rule:
- `✔ Batch size 3 (tile grid: 1×3)` → required: 5 separate rows
(`[streammux] batch-size=3`, `[primary-gie] batch-size=3`,
`[source-list] max-batch-size=3`, `[tiled-display] rows=1`,
`[tiled-display] columns=3`).
- `✔ Output sink eglsink` → required: one row per sink key
(4 keys for eglsink, e.g. `[sink0] enable=1`, `type=2`,
`sync=0`, `qos=0` — read apply-config.md for the exact list).
- `✔ Sources static (3 streams, http-port=9000)` → required: six
annotated `[source-list]` rows.
- `✔ Tile grid 1 row × 3 cols` (single row) → required: two
rows, `[tiled-display] rows=1` and `[tiled-display] columns=3`.
## Universal box format
The geometry contract for every step-exit box (Step 1 through Step 5
Results). The same shape across every box; only the **title** and the
**body rows** change per step.
- **Width: 128 chars** corner-to-corner — `┌` at column 1, `┐` at
column 128. Wider terminals leave the box flush-left; do not stretch
it. Inner content area is **124 chars** (with one space margin on
each side inside the `│` borders).
- **Light box-drawing chars only**: `┌ ─ ┐ │ └ ┘`. No `+`, `-`, `=`,
`*` ASCII fallbacks.
- **Top border — title CENTERED**: `┌` + N₁ dashes + `␣` + title + `␣`
+ N₂ dashes + `┐`, where `N₁ + N₂ + len(title) + 2 = 126`. Distribute
the pad: `N₁ = floor((126 − len(title) − 2) / 2)`,
`N₂ = 126 − len(title) − 2 − N₁`. N₁ and N₂ differ by at most 1.
- **Body**: one `│ <content padded to inner-content 124> │` per fact.
Each fact line uses the ` ✔ <key-padded-to-13> <value>` form (two
spaces in, glyph, key right-padded to 13, two spaces, value).
- **Blank lines between groups**: render `│ <124 spaces> │` between
logical groups (e.g. Identity / Model / Videos in Step 1) so the
user can scan the box at a glance.
- **Bottom border**: `└` + 126 dashes + `┘` — solid border, no title.
Standard step titles (used at the top of each step's box):
```
┌─────────────────────────────────────────────────────── Deploy targets ───────────────────────────────────────────────────────┐
┌─────────────────────────────────────────────────── Pipeline configuration ───────────────────────────────────────────────────┐
┌───────────────────────────────────────────────────────── Container ──────────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────────── Apply configuration ─────────────────────────────────────────────────────┐
┌──────────────────────────────────────────────── Perception Application — Plan ───────────────────────────────────────────────┐
┌────────────────────────────────────────────── Perception Application — Results ──────────────────────────────────────────────┐
```
Per-step content rules (which rows go in which box, mode-aware row
hiding, the apply-config sectioned layout, the Step 5 PLAN-then-RESULT
pattern, the Step 3 `docker run` synthesis requirement) live in
[`references/deploy-vss-detection-tracking-2d.md`](references/deploy-vss-detection-tracking-2d.md)
under "Step N box content rule" — read those when rendering the
corresponding step.
## Quick triggers (mnemonic)
| Phrase | Flow |
|--------|------|
| `deploy rtvicv warehouse 2d with 4 streams and display` | DEPLOY |
| `run smartcity gdino on gpu 1` | DEPLOY |
| `stop the perception container` | TEARDOWN (deploy doc) |
| `rtvi-cv healthcheck failing` | DEBUG (deploy doc + troubleshooting) |
| `add a stream to rtvi-cv` | API USAGE |
| `is rtvi-cv ready on localhost:9000` | API USAGE |
| `get rtvi-cv metrics` | API USAGE |
| `generate text embeddings via rtvi-cv` | API USAGE |
bump:1
모든 파일
51개 파일vss-deploy-detection-tracking-2d 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-detection-tracking-2d # Copy SKILL.md to your .claude/skills/ directory
복사





집
