vss-deploy-dense-captioning
NVIDIA/skills
독립형 RT-VLM 고밀도 캡션 생성 마이크로서비스를 배포하고, 파일 업로드, 캡션 생성, 스트리밍, 채팅 자동 완성 및 Kafka 통합을 위한 REST API 엔드포인트를 테스트해 봅니다.
...모든 것을 확장하십시오목적
RT-VLM 고밀도 캡션 생성 마이크로서비스를 독립적으로 구동하고, 해당 서비스가 제공하는 모든 엔드포인트(파일 업로드, generate_captions, 스트림 추가/삭제, 채팅 자동 완성, Kafka 토픽)를 테스트합니다.
필수 조건
RT-VLM 독립형 배포를 위해서는 다음이 필요합니다:
- Docker, Docker Compose, NVIDIA Container Toolkit 및 접근 가능한 GPU.
docker login nvcr.io, 이미지 가져오기, 로컬 NGC 모델/아티팩트 다운로드를 위한$NGC_CLI_API_KEY에저장된 NGC 레지스트리 자격 증명.curl,jq및 독립형 Compose 복사본을 위한 쓰기 가능한 작업 디렉터리.
기존 서비스에 대한 API 호출의 경우:
- $BASE_URL에서 접근 가능한 RT-VLM 서비스가 실행 중이어야
합니다. - 서비스 구성 방식에 따라
$RTVI_VLM_API_KEY또는$NGC_CLI_API_KEY에 Bearer 토큰이 설정되어 있어야 합니다.
전체 VSS 프로필 배포의 경우:
../vss-deploy-profile/SKILL.md를사용하십시오. 이 스킬은 전체 VSS 프로필을 배포하지 않습니다.
지침
아래의 라우팅 테이블과 단계별 워크플로를 따르십시오. '워크플로', '빠른 시작' 또는 '흐름'으로 끝나는 각 섹션은 위에서 아래로 순서대로 실행하도록 되어 있습니다. 자세한 참조 자료는 references/ 디렉터리에 있습니다. 향후 개정판에서 구체적인 헬퍼가 명시되지 않는 한, 문서화된 워크플로를 직접 실행하십시오.
예제
완료된 엔드투엔드 예제는 evals/ 디렉터리에 저장되어 있으며(각 *.json 매니페스트 파일에는 실행 가능한 시나리오가 포함되어 있음), 아래의 워크플로우별 curl 블록 내에도 포함되어 있습니다. 이를 재현하려면 nv-base validate 사용하여 Tier-3 평가를 실행하십시오.
제한 사항
- 이 스킬을 통해 배포된 독립형 RT-VLM 서비스 또는 호출자가 접근할 수 있는 기존 RT-VLM 서비스 중 하나가 필요합니다.
- NGC에서 호스팅되는 모델 및 NIM은 사용량 제한, GPU 메모리 요구 사항 및 라이선스 제한의 적용을 받을 수 있습니다.
- 동시 실행, GPU 메모리 및 스토리지 제한은 호스트 하드웨어와 프로필의 compose 파일에 따라 다릅니다.
NGC_CLI_API_KEY,RTVI_VLM_API_KEY및.env파일은 git 및 로그에 포함되지 않도록 하십시오. 자격 증명 값을 출력하거나 최종 응답에 포함하지 마십시오.- Docker 그룹 액세스 및
sudo는사실상 루트 수준의 권한입니다. 비밀번호 없는 sudo를 사용할 수 없는 경우, 배포 참조에서 비상호형sudo -n가드를 사용하고 호스트 소유자의 조치를 위해 중지하십시오.
문제 해결
- 오류: REST 호출에서 연결 거부(connection refused)가 반환됩니다. 원인: 대상 마이크로서비스가 실행 중이 아닙니다. 해결 방법:
/docs또는/health를프로브하고,vss-deploy-profile또는 해당vss-deploy-*스킬을 통해 재배포하십시오. - 오류: NGC 풀에서 HTTP 401/403 오류가 발생했습니다. 원인:
NGC_CLI_API_KEY가없거나 만료되었습니다. 해결 방법:docker login nvcr.io를실행하고 키를 다시 내보낸 후 재시도하십시오. - 오류: 컨테이너 OOM 또는 모델 로딩 실패. 원인: 선택한 프로필에 대한 GPU 메모리 부족. 해결 방법: 더 작은 변형으로 전환하거나
docker compose down을통해 GPU 리소스를 확보하십시오.
RT-VLM Dense Captioning(VSS 3.2) 배포 및 사용
RT-VLM은 NVIDIA의 실시간 비전-언어 마이크로서비스입니다. 비디오(파일 또는
RTSP)를 디코딩하고, 청크로 분할한 후, VLM(cosmos-reason1, cosmos-reason2 또는
OpenAI 호환 모델)을 실행하고, SSE/HTTP를 통해 덴스 캡션을 다시 스트리밍하며,
캡션, 이벤트 알림 및 오류를 Kafka에 게시합니다. 이 스킬을 사용하여 전체 VSS 프로필이 아직 실행 중이지 않을 때
독립형 RT-VLM 서비스를 배포한 다음,
해당 /v1/... API를 호출하여 캡션 생성, 파일 업로드, 라이브 스트림 관리, 상태
점검, NIM 호환 채팅 자동 완성 또는 Prometheus 메트릭을 수행할 수 있습니다. API 참조:
https://docs.nvidia.com/vss/latest/real-time-vlm-api.html.
배포 라우팅
사용자가 전체 VSS 프로필을 배포해 달라고 요청하는 경우,
../vss-deploy-profile/SKILL.md를 사용하십시오. 이 스킬은
프로필 라우팅, generated.env, resolved.yml, 다중 서비스 크기 조정 및
풀스택 배포/종료를 담당합니다.
사용자가 독립형 RT-VLM 고밀도 캡셔닝을 요청하거나,
이미 실행 중인 VSS 프로필이 없는 경우, API를 호출하기 전에
references/deploy-rt-vlm-service.md에있는
독립형 RT-VLM 흐름을 사용하십시오. 이는vss-deploy-profile과 동일한 Compose 중심 패턴을 따릅니다.
즉, 컨텍스트 수집, 사전 검증 실행, 로컬 사본을 기반으로 작업,
Docker Compose 구성으로 드라이 런 수행, 검토, 배포, 그리고 상태 확인을 기다리는 순서입니다.
독립형 배포 흐름
항상 이 순서를 따르십시오. 드라이 런을 절대 건너뛰지 마십시오.
# 1. deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml 파일을
# 쓰기 가능한 독립형 작업 디렉터리에 복사합니다.
# 2. 해당 Compose 복사본에서 RTVI_VLM_IMAGE_TAG를 도출합니다.
# 3. 복사본에서 독립 실행형 전용인 불필요한 depends_on 블록을 제거하십시오.
# 4. 필요한 RT-VLM 값이 포함된 gitignored .env 파일을 생성하십시오.
# 5. $VSS_DATA_DIR/data_log/vst/clip_storage와 같은 호스트 바인드 경로를 준비하십시오.
# 소유권 수정 시 `sudo -n`을 사용하십시오. 비밀번호 없이 sudo를 사용할 수 없는 경우,
# 작업을 중지하고 호스트 소유자에게 출력된 명령을 수동으로 실행해 달라고 요청하십시오.
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. 정확한 RT-VLM 이미지 태그를 docker pull로 가져옵니다.
# 8. docker compose ... up -d rtvi-vlm을 실행하고, 준비 상태가 될 때까지 기다린 후, 기본 테스트를 수행합니다.
pull 또는 up 명령을 실행하기 전에 사전 점검을 수행하십시오. RT-VLM 자체를 디버깅하기 전에
여기서 오류를 중지하고 수정하십시오:
nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
독립형 단일 파일 배포의 경우,
deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml 파일을 직접 실행하지 마십시오. 이 파일에는
전체 VSS/met-blueprints 컴포즈 프로젝트에만 정의된
전체 VSS/met-blueprints compose 프로젝트에서만 정의된 동급 VLM/NIM 서비스에 대한
` depends_on ` 참조가 포함되어 있습니다. 독립형 참조 예시는
compose 파일을 복사하고, 이를 기반으로 현재 이미지 태그를 도출하며,
`depends_on` 블록을 제거하고, `up` 명령을 실행하기 전에 결과를 검증하는 방법을 보여줍니다.
에이전트 기반 유효성 검사의 경우, 절대 sudo 프롬프트가 대화형으로 작동하게 해서는 안 됩니다. 권한이 필요한 소유권 설정이나 Docker 작업 전에,
references/deploy-rt-vlm-service.md에 있는 비대화형 가드 명령을 사용하십시오:
일반 docker를 선호하고, 그렇지 않으면 sudo -n docker를 사용하십시오. sudo -n이 실패할 경우,
대화형 sudo로 재시도하거나 권한을 완화하는 대신, 호스트 소유자를 위한 정확한 수동 명령을 사용하여
중단하십시오.
Docker 28 이상에서 docker pull이 containerd snapshotter/unpack 오류로 실패할 경우,
재시도하기 전에 standalone 참조 문서에 있는 /etc/docker/daemon.json의
containerd-snapshotter=false 수정 사항을 적용하십시오.
최소 독립형 .env 값:
| 호스트 환경 변수 | 다음 경우에 필수 | 목적 |
|---|---|---|
NGC_CLI_API_KEY |
독립형 배포 경로 | NGC 레지스트리 이미지 가져오기 및 NGC 모델/아티팩트 다운로드 |
RTVI_VLM_API_KEY 또는 NGC_CLI_API_KEY |
인증된 API 호출 | 서비스 실행 후 RT-VLM 베어러 인증 |
RTVI_VLM_PORT |
항상 | 컨테이너 8000에 매핑된 호스트 API 포트 |
HOST_IP |
항상 | Kafka 부트스트랩 호스트 (${HOST_IP}:9092) |
VSS_DATA_DIR |
항상 | 필수 clip-storage 바인드 마운트 |
RTVI_VLM_MODEL_TO_USE |
독립 실행형인 경우 항상 | 백엔드 선택기; 기본 로컬 모델의 경우 cosmos-reason2를, 원격/동급 엔드포인트의 경우 openai-compat를 사용하십시오 |
RTVI_VLM_MODEL_PATH |
로컬 자체 호스팅 모델 | 소스 기반 Cosmos Reason 2 경로: ngc:nim/nvidia/cosmos-reason2-8b:hf-1208 |
RTVI_VLM_ENDPOINT |
RTVI_VLM_MODEL_TO_USE=openai-compat |
원격/동급 OpenAI 호환 VLM 엔드포인트 |
VLM_NAME |
RTVI_VLM_MODEL_TO_USE=openai-compat |
해당 엔드포인트에서 노출된 모델/배포 이름 |
설정
export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}" # 호스트 측 RT-VLM 포트
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # 호스트 측 curl 명령어에서 사용하는 베어러 토큰
: "${API_KEY:?인증된 엔드포인트를 호출하기 전에 NGC_CLI_API_KEY 또는 RTVI_VLM_API_KEY를 설정하세요}"
아래의 모든 요청은 Authorization: Bearer $API_KEY를 사용합니다. 상태 확인 엔드포인트
(/v1/health/*, /v1/ready, /v1/live, /v1/startup)는 일반적으로 인증 없이 작동합니다.
사용 전 간단한 테스트:
curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
RTSP 샘플 스트림 가드
작업이나 eval에서 RTSP_SAMPLE_URL을 지정하면, 해당 환경 변수를
필수 입력으로 취급하십시오. 스트림을 탐색하거나
등록하기 전에 해당 변수가 설정되어 있고 비어 있지 않은지 확인하십시오. 변수가 누락된 경우 명확한 오류 메시지와 함께 중단하십시오. NvStreamer, VIOS, 샘플 데이터 번들 또는 기타
대체 수단에서 대체값을 도출하지 마십시오. 이는 호출자가 요청한 것과는 다른 스트림을 검증하게 되기 때문입니다.
: "${RTSP_SAMPLE_URL:?RTSP 유효성 검사 전에 RTSP_SAMPLE_URL을 연결 가능한 RTSP 샘플 스트림으로 설정하십시오}"
case "$RTSP_SAMPLE_URL" in
rtsp://*) ;;
*) echo "RTSP_SAMPLE_URL은 rtsp:// URL이어야 합니다. 입력값: $RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac
if command -v ffprobe >/dev/null 2>&1; then
ffprobe -v error -rtsp_transport tcp \
-select_streams v:0 -show_entries stream=codec_type \
-of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
echo "RTSP 유효성 검사를 수행하기 전에 ffprobe 또는 gst-discoverer-1.0을 설치하십시오." >&2
exit 1
fi
빠른 시작 — 로컬 동영상에서 밀집 자막 추출
# 1. 동영상을 업로드하고 파일 ID를 가져옵니다.
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@/path/to/warehouse.mp4" \
-F "purpose=vision" \
-F "media_type=video" | jq -r '.id')
# 2. 캡션 및 알림 생성 (청크화된 응답의 SSE 스트림)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"id\": \"$FILE_ID\",
\"prompt\": \"이 창고 영상의 10초 단위 세그먼트마다 간결하고 핵심을 담은 캡션을 작성해 주세요.\",
\"model\": \"$MODEL_ID\",
\"chunk_duration\": 10,
\"stream\": true
}"
API 인터페이스
선택적 엔드포인트를 호출하기 전에 라이브 OpenAPI를 신뢰할 수 있는 정보원으로 활용하십시오:
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
VSS 3.2의 핵심 경로는 다음과 같습니다:
- 다중 파트 미디어 업로드를 위한
POST /v1/files; 반환된 파일ID를자막 생성에 전달하고, 완료되면 파일을 삭제하십시오. - 파일 또는 스트림 자막 생성을 위한
POST /v1/generate_captions.GET /v1/models에서반환된 정확한 모델 ID를 사용하십시오.cosmos-reason2와 같은 별칭은 백엔드 선택자일 뿐, 요청 모델 ID가 아닙니다. - RTSP 라이프사이클을 위한
POST /v1/streams/add,GET /v1/streams/get-stream-info및DELETE /v1/streams/delete/{stream_id}.results[0].id에서 스트림 ID를 파싱합니다. - OpenAI 호환 텍스트 및 다중 모드 호출을 위해
POST /v1/chat/completions를 사용하십시오. 현재 26.05 빌드에서는 텍스트 전용/v1/completions에대해 HTTP 400을 반환합니다. 레거시 동작을 검증할 때는 이를 예상된 동작으로 간주하십시오. - 서비스 상태를 확인하려면
GET /v1/health/ready,/v1/models,/v1/assets/stats및/v1/metrics를사용하십시오. OpenAPI에 명시되어 있지 않은 한/v1/license가존재한다고 가정하지 마십시오.
상세한 엔드포인트 스키마, 응답 형식, CV 스타일의 단일 스트림 엔드포인트,
및 26.05 호환성 참고 사항은
references/api-surface-26.05.md에 있습니다.
일반적인 워크플로
- 저장된 파일 자막 생성:
POST /v1/files로업로드하고, 반환된 파일 ID를 사용하여/v1/generate_captions를호출한 후, SSE의 경우stream=true를사용하고, 저장 공간을 확보하기 위해 파일을 삭제합니다. - RTSP 실시간 자막 생성: 호출자가
RTSP_SAMPLE_URL을제공하면, 그 정확한 URL을 사용하고 등록 전에 RTSP Sample Stream Guard를 실행합니다.RTSP_SAMPLE_URL이비어 있을 때는 NvStreamer나 VIOS에서 대체 스트림을 생성하지 말고, 대신 즉시 실패 처리하십시오. 등록 전에 실제 비디오 스트림/자막 항목이 필수여야 하며, 스트림을 추가하고 자막을 생성한 후 등록을 해제하십시오. - 경고 메시지:
'이상 감지: 예/아니오'라는결정론적 문구를 포함하십시오. Kafka 게시 설정은 서버 측 구성으로, HTTP 응답에 추가되며,references/kafka-workflows.md에 문서화되어 있습니다. - Kafka 유효성 검사: 토픽 이름에 대해서는 라이브
vss-rtvi-vlm환경을 신뢰하십시오. 전체 VSS 알림 실시간 프로필에서는 CLI 검사 및 최종 인시던트 소비자 명령을 위해 기존 VSS Kafka 컨테이너mdx-kafka를사용하십시오. 독립형 유효성 검사의 경우,${HOST_IP}:9092를알리는 브로커를 사용하십시오. 사용자의 확인 없이 기존 브로커를 중지하거나 교체해서는 안 됩니다.
오류 참조
일반적인 원인: 400은 잘못된 요청 형식 또는 모델 ID, 401/403은
베어러 토큰 누락 또는 오류, 404는 삭제된 파일/스트림 또는 지원되지 않는 엔드포인트,
413: 업로드 파일 크기 초과, 422: 스키마 유효성 검사 오류, 429: 동시 접속 과다,
500: 추론/런타임 오류, 503: 시작 중일 때 발생. 서비스 측 오류는 Docker 로그 vss-rtvi-vlm을 확인하십시오.
---
name: vss-deploy-dense-captioning
description: Deploy a standalone RT-VLM dense-captioning microservice and exercise its REST API endpoints for file upload, caption generation, streaming, chat completions, and Kafka integration.
license: Apache-2.0
---
## Purpose
Stand up the RT-VLM dense-captioning microservice on its own and exercise every endpoint it exposes (file upload, generate_captions, stream add/delete, chat-completions, Kafka topics).
## Prerequisites
For standalone RT-VLM deployment:
- Docker, Docker Compose, NVIDIA Container Toolkit, and a visible GPU.
- NGC registry credentials in `$NGC_CLI_API_KEY` for `docker login nvcr.io`,
image pulls, and local NGC model/artifact downloads.
- `curl`, `jq`, and any writable working directory for the standalone compose copy.
For API calls against an existing service:
- Running RT-VLM service reachable at `$BASE_URL`.
- Bearer token in `$RTVI_VLM_API_KEY` or `$NGC_CLI_API_KEY`, depending on how the
service was configured.
For full VSS profile deployment:
- Use `../vss-deploy-profile/SKILL.md`; this skill does not deploy full VSS profiles.
## Instructions
Follow the routing tables and step-by-step workflows below. Each section that ends in *workflow*, *quick start*, or *flow* is intended to be executed top-to-bottom. Detailed reference material lives in `references/`; execute the documented workflows directly unless a future revision names a concrete helper.
## Examples
Worked end-to-end examples are kept under `evals/` (each `*.json` manifest contains a runnable scenario) and inline in the per-workflow `curl` blocks below. Run a Tier-3 evaluation with `nv-base validate <this-skill-dir> --agent-eval` to replay them.
## Limitations
- Requires either a standalone RT-VLM service deployed via this skill or an
existing RT-VLM service reachable from the caller.
- NGC-hosted models and NIMs may be subject to rate-limits, GPU memory requirements, and license restrictions.
- Concurrency, GPU memory, and storage limits depend on the host hardware and the profile's compose file.
- Keep `NGC_CLI_API_KEY`, `RTVI_VLM_API_KEY`, and `.env` files out of git and out of logs; do not echo credential values or include them in final responses.
- Docker group access and `sudo` are effectively root-level privileges. Use the non-interactive `sudo -n` guard in the deploy reference and stop for host-owner action when passwordless sudo is unavailable.
## Troubleshooting
- **Error**: REST call returns connection refused. **Cause**: target microservice not running. **Solution**: probe `/docs` or `/health`; redeploy via `vss-deploy-profile` or the matching `vss-deploy-*` skill.
- **Error**: HTTP 401/403 from NGC pulls. **Cause**: missing/expired `NGC_CLI_API_KEY`. **Solution**: `docker login nvcr.io` and re-export the key before retrying.
- **Error**: container OOM or model fails to load. **Cause**: insufficient GPU memory for the selected profile. **Solution**: switch to a smaller variant or free GPUs via `docker compose down`.
# Deploy and Use RT-VLM Dense Captioning (VSS 3.2)
RT-VLM is NVIDIA's real-time vision-language microservice: decode video (file or
RTSP), segment it into chunks, run a VLM (`cosmos-reason1`, `cosmos-reason2`, or any
OpenAI-compatible model), stream dense captions back over SSE/HTTP, and publish
captions, incident alerts, and errors to Kafka. Use this skill to deploy the
standalone RT-VLM service when a full VSS profile is not already running, then call
its `/v1/...` API for caption generation, file upload, live-stream management, health
checks, NIM-compatible chat completions, or Prometheus metrics. API reference:
<https://docs.nvidia.com/vss/latest/real-time-vlm-api.html>.
## Deployment Routing
If the user asks to deploy a full VSS profile, use
[`../vss-deploy-profile/SKILL.md`](../vss-deploy-profile/SKILL.md). That skill
owns profile routing, `generated.env`, `resolved.yml`, multi-service sizing, and
full-stack deploy/teardown.
If the user asks for standalone RT-VLM dense captioning, or no VSS profile is
already running, use the standalone RT-VLM flow in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md)
before calling the API. This follows the same compose-centric pattern as
`vss-deploy-profile`: gather context, run preflights, work from a local copy,
dry-run with `docker compose config`, review, deploy, then wait for health.
## Standalone Deployment Flow
Always follow this sequence. Never skip the dry-run.
```bash
# 1. Copy deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml
# into any writable standalone working directory.
# 2. Derive RTVI_VLM_IMAGE_TAG from that compose copy.
# 3. Strip the standalone-only dangling depends_on block from the copy.
# 4. Create a gitignored .env with the required RT-VLM values.
# 5. Prepare host bind paths such as $VSS_DATA_DIR/data_log/vst/clip_storage.
# Use `sudo -n` for ownership fixes; if passwordless sudo is unavailable,
# stop and ask the host owner to run the printed command manually.
# 6. docker compose --env-file .env -f rtvi-vlm-docker-compose.yml config --quiet
# 7. docker pull the exact RT-VLM image tag.
# 8. docker compose ... up -d rtvi-vlm, wait for ready, then smoke test.
```
Run preflights before any pull or `up`; stop and fix failures here before
debugging RT-VLM itself:
```bash
nvidia-smi --query-gpu=index,name --format=csv,noheader
nvidia-container-cli info
docker compose version
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
```
For standalone single-file deployments, do not run the raw
`deploy/docker/services/rtvi/rtvi-vlm/rtvi-vlm-docker-compose.yml` directly: it
contains `depends_on` references to sibling VLM/NIM services that are only
defined in the full VSS/met-blueprints compose project. The standalone reference
shows how to copy the compose file, derive the current image tag from it, strip
the `depends_on` block, and validate the result before `up`.
For agent-driven validation, never let `sudo` prompt interactively. Before any
privileged ownership or Docker operation, use the non-interactive guard in
[`references/deploy-rt-vlm-service.md`](references/deploy-rt-vlm-service.md):
prefer plain `docker`; otherwise use `sudo -n docker`; if `sudo -n` fails, stop
with the exact manual command for the host owner instead of retrying with
interactive sudo or weakening permissions.
If `docker pull` fails with a containerd snapshotter/unpack error on Docker 28+,
apply the `/etc/docker/daemon.json` `containerd-snapshotter=false` fix in the
standalone reference before retrying.
Minimum standalone `.env` values:
| Host env var | Required when | Purpose |
|---|---|---|
| `NGC_CLI_API_KEY` | Standalone deploy path | NGC registry image pull and NGC model/artifact download |
| `RTVI_VLM_API_KEY` or `NGC_CLI_API_KEY` | Authenticated API calls | RT-VLM bearer auth after the service is running |
| `RTVI_VLM_PORT` | Always | Host API port mapped to container `8000` |
| `HOST_IP` | Always | Kafka bootstrap host (`${HOST_IP}:9092`) |
| `VSS_DATA_DIR` | Always | Required clip-storage bind mount |
| `RTVI_VLM_MODEL_TO_USE` | Always for standalone | Backend selector; use `cosmos-reason2` for the default local model or `openai-compat` for a remote/sibling endpoint |
| `RTVI_VLM_MODEL_PATH` | Local self-hosted model | Source-backed Cosmos Reason 2 path: `ngc:nim/nvidia/cosmos-reason2-8b:hf-1208` |
| `RTVI_VLM_ENDPOINT` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Remote/sibling OpenAI-compatible VLM endpoint |
| `VLM_NAME` | `RTVI_VLM_MODEL_TO_USE=openai-compat` | Model/deployment name exposed by that endpoint |
## Setup
```bash
export BASE_URL="http://localhost:${RTVI_VLM_PORT:-8018}" # host-side RT-VLM port
export API_KEY="${NGC_CLI_API_KEY:-${RTVI_VLM_API_KEY:-}}" # bearer token used by host-side curl commands
: "${API_KEY:?Set NGC_CLI_API_KEY or RTVI_VLM_API_KEY before calling authenticated endpoints}"
```
Every request below uses `Authorization: Bearer $API_KEY`. Health endpoints
(`/v1/health/*`, `/v1/ready`, `/v1/live`, `/v1/startup`) typically work without auth.
**Smoke test before use:**
```bash
curl -fsS "$BASE_URL/v1/health/ready"
MODEL_ID="$(curl -fsS "$BASE_URL/v1/models" -H "Authorization: Bearer $API_KEY" | jq -r '.data[0].id // .id')"
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```
## RTSP Sample Stream Guard
When a task or eval names `RTSP_SAMPLE_URL`, treat that exact environment
variable as a required input. Verify it is set and non-empty before probing or
registering any stream; if it is missing, stop with a clear failure message. Do
not derive a substitute from NvStreamer, VIOS, sample-data bundles, or any other
fallback, because that validates a different stream than the caller requested.
```bash
: "${RTSP_SAMPLE_URL:?Set RTSP_SAMPLE_URL to a reachable RTSP sample stream before RTSP validation}"
case "$RTSP_SAMPLE_URL" in
rtsp://*) ;;
*) echo "RTSP_SAMPLE_URL must be an rtsp:// URL, got: $RTSP_SAMPLE_URL" >&2; exit 1 ;;
esac
if command -v ffprobe >/dev/null 2>&1; then
ffprobe -v error -rtsp_transport tcp \
-select_streams v:0 -show_entries stream=codec_type \
-of csv=p=0 "$RTSP_SAMPLE_URL" | grep -qx video
elif command -v gst-discoverer-1.0 >/dev/null 2>&1; then
gst-discoverer-1.0 "$RTSP_SAMPLE_URL" | grep -qi 'video'
else
echo "Install ffprobe or gst-discoverer-1.0 before RTSP validation." >&2
exit 1
fi
```
## Quick Start — dense captions from a local video
```bash
# 1. Upload the video, capture its file id
FILE_ID=$(curl -fsS -X POST "$BASE_URL/v1/files" \
-H "Authorization: Bearer $API_KEY" \
-F "file=@/path/to/warehouse.mp4" \
-F "purpose=vision" \
-F "media_type=video" | jq -r '.id')
# 2. Generate captions + alerts (SSE stream of chunked responses)
curl -N -X POST "$BASE_URL/v1/generate_captions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"id\": \"$FILE_ID\",
\"prompt\": \"Write a concise dense caption for each 10-second segment of this warehouse video.\",
\"model\": \"$MODEL_ID\",
\"chunk_duration\": 10,
\"stream\": true
}"
```
## API Surface
Use the live OpenAPI as the source of truth before calling optional endpoints:
```bash
curl -fsS "$BASE_URL/openapi.json" | jq -r '.paths | keys[]' | sort
```
Core paths for VSS 3.2 are:
- `POST /v1/files` for multipart media upload; pass the returned file `id` into
caption generation and delete the file when finished.
- `POST /v1/generate_captions` for file or stream captioning. Use the exact
model id returned by `GET /v1/models`; aliases such as `cosmos-reason2` are
backend selectors, not request model ids.
- `POST /v1/streams/add`, `GET /v1/streams/get-stream-info`, and
`DELETE /v1/streams/delete/{stream_id}` for RTSP lifecycle. Parse stream ids
from `results[0].id`.
- `POST /v1/chat/completions` for OpenAI-compatible text and multimodal calls.
Current 26.05 builds return HTTP 400 for text-only `/v1/completions`; treat
that as expected when validating legacy behavior.
- `GET /v1/health/ready`, `/v1/models`, `/v1/assets/stats`, and `/v1/metrics`
for service probes. Do not assume `/v1/license` exists unless OpenAPI lists it.
Detailed endpoint schemas, response shapes, CV-style singular stream endpoints,
and 26.05 compatibility notes live in
[`references/api-surface-26.05.md`](references/api-surface-26.05.md).
## Common Workflows
- Stored file captioning: upload with `POST /v1/files`, call
`/v1/generate_captions` with the returned file id, use `stream=true` for SSE,
then delete the file to release storage.
- RTSP live captioning: when the caller provides `RTSP_SAMPLE_URL`, use that
exact URL and run the **RTSP Sample Stream Guard** before registration. Do not
derive a replacement stream from NvStreamer or VIOS when `RTSP_SAMPLE_URL` is
empty; fail fast instead. Require an actual video stream/caps entry before
registration; add the stream, caption it, then unregister it.
- Alert prompts: include a deterministic `Anomaly Detected: Yes/No` line.
Kafka publication is server-side config, additive to HTTP responses, and
documented in [`references/kafka-workflows.md`](references/kafka-workflows.md).
- Kafka validation: trust the live `vss-rtvi-vlm` environment for topic names.
In a full VSS alerts real-time profile, use the existing VSS Kafka container
`mdx-kafka` for CLI checks and final incident-consumer commands. For
standalone validation, use a broker that advertises `${HOST_IP}:9092`; never
stop or replace a pre-existing broker without user confirmation.
## Error Reference
Common causes: 400 for invalid request shape or model id, 401/403 for missing
or wrong bearer token, 404 for deleted files/streams or unsupported endpoints,
413 for oversized uploads, 422 for schema validation, 429 for too much
concurrency, 500 for inference/runtime failures, and 503 while startup is still
in progress. Inspect `docker logs vss-rtvi-vlm` for service-side failures.
vss-deploy-dense-captioning 설치
스킬 파일을 다운로드한 후 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-deploy-dense-captioning # Copy SKILL.md to your .claude/skills/ directory
복사





집
