옵션
집집 Skill 데이터베이스 관리 vss-generate-video-report

vss-generate-video-report

NVIDIA/skills NVIDIA/skills

클립별 분석을 위해 VLM 백엔드로, 또는 사건 범위 보고서를 위해 분석 백엔드로 라우팅하여 비디오 분석 보고서를 생성하며, 이 과정에서 배포 프로필 검증 및 URL 재작성 기능을 수행합니다.

...모든 것을 확장하십시오
2
업데이트 된 시간 2026년 9월 27일

보고서

두 개의 백엔드 중 하나로 라우팅하여 비디오 분석 보고서를 생성하십시오. VSS 에이전트의 POST /generate 메서드를 통해서는 절대 생성하지 마십시오.

모드 백엔드
A. 동영상 클립 /vss-manage-video-io-storage → 클립 URL → VLM 채팅/완료
B. 인시던트 범위 /vss-query-analytics → 인시던트 목록 → 서술형 보고서

요청이 모호한 경우(예: “시간 범위 및 인시던트 문구가 지정되지 않은 "와 같이 시간 범위나 인시던트 문구가 없는 경우), 기본적으로 모드 A를 적용합니다. 사용자가 센서와 시간 범위를 모두 언급한 경우에만 해당 모드를 적용합니다. 각 모드로 라우팅되는 요청 문구에 대해서는 아래 예시를 참조하십시오.

지침

  1. 모드를 선택하십시오 — 단일 녹화 클립/센서 영상의 경우 모드 A, 요청에 시간 범위나 인시던트/경보가 명시된 경우 모드 B를 선택하십시오( 예시와 대조).
  2. “배포 전제 조건”에서 해당 모드의배포 프로필을 확인하고, 프로브 테스트에 실패하면 /vss-deploy-profile로 인계하십시오.
  3. 해당 모드의 번호가 매겨진 단계를 실행하십시오 — 아래의 모드 A 또는 모드 B.
  4. 보고서에 삽입하기 전에,사용자에게 표시되는 모든 클립 URL을 $VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT 한 줄 형식(브라우저에서 재생 가능한 클립 URL)으로다시 작성하십시오.
  5. 렌더링된 보고서 마크다운을 사용자에게반환합니다.

평가자용 계약서:

  • 모드 A의 최상위 제목은 반드시 # Video Analysis Report여야 합니다.
  • 모드 B의 상단 제목은 반드시 # Incident Range Report 여야 합니다(절대 # Incident Report 나 센서 이름이 포함된 변형 제목은 사용해서는 안 됩니다).
  • 모드 B에는 템플릿에 명시된 필수 행(보고서 식별자, 범위, 적용 범위, 총 사고 수, 확인됨 / 거부됨 / 미확인)이 포함된 ## 기본 정보가 반드시 포함되어야 합니다.

예시

  • "이 비디오에 대한 보고서 생성" / "다음에 대한 보고서 " → 모드 A
  • "warehouse_01.mp4 분석" / "업로드된 동영상에 대한 분석 보고서 생성" → 모드 A
  • "12:31Z부터 12:32Z까지의 사건에 대한 보고서" → 모드 B
  • "오늘의 알림에 대한 보고서" / "지난 지난 1시간 동안 어떤 사고가 발생했는지" → 모드 B
  • "다음 기간의 알림 요약 ~ 부터"까지의 경보 요약" → 모드 B " → 모드 B

부정 트리거

요청이 다음 중 하나에 해당하는 경우 이 스킬을 사용하지 마십시오:

  • 보고서를 명시적으로 요청하지 않는 클립에 대한 즉석 시각적 Q&A("트럭 색깔이 뭐예요?", "00:12에 무슨 일이 일어나나요?") → /vss-ask-video를 사용하십시오.
  • 아카이브/의미적 유사성 검색("지게차 찾기", "모든 동영상에서 과밀한 차간 거리 위반 장면 검색") → /vss-search-archive를 사용하십시오.
  • 보고서 생성이 필요 없는 읽기 전용 사고/지표 조회 → /vss-query-analytics를 사용하십시오.
  • 배포/해체/프로필 변경("알림 배포", "프로필 전환", "베이스 가동") → /vss-deploy-profile을 사용하십시오.
  • 실시간 경보/규칙 관리 요청 → /vss-manage-alerts를 사용하십시오.

VSS-agent POST /generate를 통해 보고서를 라우팅하지 마십시오.

배포 전제 조건

모드 A에는 VSS 기본 프로필(VST + VLM NIM)이 필요합니다. 모드 B에는 VSS 알림 프로필(VA-MCP + Elasticsearch)이 필요합니다.

테스트:

# 모드 A — VST + VLM 연결 상태 확인
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null

# 모드 B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null

프로브가 실패하면 -p base (모드 A) 또는 -p alerts (모드 B) 옵션을 사용하여 /vss-deploy-profile로 처리하십시오. 배포 전에 항상 사용자와 먼저 확인하십시오.

클립 URL: VLM 입력 대 브라우저 보고서 링크

VST는 에이전트 내부 호스트:포트인 ${HOST_IP}:30888을 사용하여 클립 URL을 반환합니다. 로컬 또는 클러스터 내 VLM 프레임 가져오기를 위해 해당 원본 URL을 VIDEO_URL 로 유지하십시오. 단순히 브라우저에서 재생할 수 있도록 하기 위해 VLM 입력 URL을 재작성 하지 마십시오.

렌더링된 보고서에 표시된 URL에 대해서만 BROWSER_CLIP_URL을 생성하십시오. 배포 계층은 브라우저용 호스트:포트를 $VSS_PUBLIC_HOST / $VSS_PUBLIC_PORT (그리고 스키마는 $VSS_PUBLIC_HTTP_PROTOCOL)로 내보냅니다 . 따라서 보고서 링크 재작성 규칙은 다음과 같습니다:

: "${VSS_PUBLIC_HOST:?클립 URL을 재작성하기 전에 VSS_PUBLIC_HOST를 설정하십시오}"
: "${VSS_PUBLIC_PORT:?클립 URL을 재작성하기 전에 VSS_PUBLIC_PORT를 설정하십시오}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")

필요한 공용 호스트 값 중 하나라도 누락된 경우, 보고서에 표시되는 클립 링크를 생략하고 브라우저에서 재생 가능한 URL을 생성할 수 없음을 명시하십시오. 단, 로컬 VLM 분석 경로는 차단하지 마십시오. 렌더링된 보고서에 표시된 모든 클립 URL에 재작성 규칙을 적용합니다(모드 A 4단계 클립 URL 행; 모드 B 사건별 클립 하위 항목). VLM이 로컬/클러스터 내부에 있을 경우, 모드 A 3단계의 VLM video_url 콘텐츠 블록은 원래의 내부 URL로 유지합니다.

모드 A — 녹화된 비디오 클립에 대한 보고서

VSS lvs 프로필이 배포된 경우 — curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready" 명령어가 HTTP 200을 반환하면 — /vss-summarize-video를 실행하여 요약 정보를 생성하고, 그 후 4단계에서 해당 출력을 보고서 템플릿에 붙여넣고 1~3단계(VLM 직접 경로)는 건너뜁니다. /v1/ready의 응답 코드가 200이 아닐 때만 1~3단계를 실행하십시오.

1단계 — 클립 URL 확인

/vss-manage-video-io-storage 로 넘겨서 다음을 수행합니다:

  1. 센서를 나열하고 지정된 존재하는지 확인합니다(존재하지 않으면 먼저 업로드하십시오).

  2. 사용자가 startTime/endTime을 지정하지 않은 경우, 기록된 범위에 대한 /storage//timelines를 가져옵니다.

  3. 클립 URL을 요청합니다:

    curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file//url?startTime=&endTime=&container=mp4&disableAudio=true" | jq -r .videoUrl
    

    이렇게 하면 로컬/클러스터 내 VLM이 프레임을 가져올 수 있는 직접적인 mp4 URL이 생성됩니다. 이 URL을 VIDEO_URL (3단계에서 VLM이 사용)에 바인딩하고, 4단계에서 BROWSER_CLIP_URL을 생성하기 위해 report-link 재작성을 적용하기 전에 RAW_URL="$VIDEO_URL"을 설정합니다. 사용자의 브라우저는 $VIDEO_URL에 직접 접속할 수 없습니다. 모드 A에서는 선택한 VLM 엔드포인트가 VIDEO_URL을 가져올 수 있어야 합니다. 로컬 NIM/RT-VLM 배포 환경에서는 일반적으로 가능하지만, 원격 엔드포인트는 대개 localhost, 사설 HOST_IP 또는 VST 내부 URL을 가져올 수 없습니다. 라이브 VLM_ENDPOINT가 원격인 경우, /v1/models 호출이 성공한 후 실패하게 될 채팅 요청을 보내는 대신 해당 접근 가능성 요구 사항을 명시하십시오.

2단계 — VLM 엔드포인트 및 모델 확인

배포 환경은 두 가지 스택 중 하나를 통해 VLM을 제공할 수 있습니다. 두 스택 모두 OpenAI 호환 채팅/완성 API를 노출하므로, 현재 가동 중인 쪽을 선택하십시오:

백엔드 환경 변수 일반적인 호스트 엔드포인트 다음 경우에 선택됨
NIM Cosmos VLM_BASE_URL, VLM_NAME, VLM_MODE, VLM_MODEL_TYPE ${VLM_BASE_URL}/v1 (환경 변수 끝에 /v1이 없음; 에이전트가 이를 추가함) VLM_MODEL_TYPE != rtvi 이고, VLM_MODE ∈ {local, local_shared, remote} 이며, VLM_BASE_URL이 비어 있지 않은 경우
RT-VLM Cosmos RTVI_VLM_BASE_URL, RTVI_VLM_MODEL_TO_USE, VLM_MODEL_TYPE ${RTVI_VLM_BASE_URL}/v1 — 설정되지 않은 경우 ${HOST_IP} 에서 파생(알림의 경우http://${HOST_IP}:8018/v1, http://${HOST_IP}:30082/v1 (기본값)) VLM_MODEL_TYPE = rtvi, 또는 VLM_MODE=none, 또는 VLM_BASE_URL이 비어 있는 경우; 또한 웨어하우스의 유일한 경로

실행 중인 에이전트 컨테이너에서 실시간 값을 읽어오십시오 — 임의로 추측하지 마십시오:

docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
  v="$(printenv "$k")"
  [ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'

vss-agent 환경에서 RTVI_VLM_ENDPOINT를 필수로 요구하지 마십시오. 여러 프로필에서 이를 주입하지 않습니다.

선택 규칙:

if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1"   # 알림 기본값
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
  VLM_BACKEND="nim_cosmos"
  VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
  VLM_MODEL="${VLM_NAME}"
else
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1"  # 기본값
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi

채팅 요청을 보내기 전에 /v1/models를 확인하여 선택한 엔드포인트가 정상적으로 작동하고 모델이 로드되었는지 확인합니다:

curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'

프로브가 실패하거나 나열된 ID에 ${VLM_MODEL}이 포함되지 않으면 다른 백엔드로 전환하십시오(또는 오류를 표시하십시오 — 서버에 없는 모델을 절대 조용히 선택해서는 안 됩니다).

3단계 — VLM을 직접 호출하기

video_url 콘텐츠 블록과 함께 OpenAI 호환 채팅/완성 엔드포인트를 사용합니다. 이는 src/vss_agents/tools/video_understanding.py에서 video_understanding이 구축하는 것과 동일한 페이로드 구조 및 멀티모달 설정을 따릅니다. (_build_vlm_messages + Cosmos의 base_vlm.bind(...) 호출)과 동일한 페이로드 형식과 멀티모달 설정을 사용합니다.

프레임 샘플링 및 시각 토큰(픽셀) 할당량은 활성 프로필에 대한 실시간 video_understanding 설정을 반영해야 합니다. mm_processor_kwargs 및 media_io_kwargs를 전송하여 직접 호출이 에이전트 내 video_understanding 도구와 동일한 프레임 샘플링 및 픽셀 할당량을 사용하도록 하십시오. 이를 생략하면 VLM이 자체 기본값을 적용하므로 출력이 에이전트 경로와 달라집니다.

PROMPT='동영상에서 일어나는 일을 각 세그먼트나 이벤트별로 타임스탬프(클립 시작 시점부터의 시작–종료 시간, 초 단위)와 함께 상세히 설명하세요. 장면, 사물, 사람, 차량 및 주목할 만한 행동을 포함하세요.'

# 추론 기능은 기본적으로 비활성화되어 있습니다 — 기본 프로필 video_understanding 구성(`reasoning: false`)과 일치합니다.
# video_understanding.py는 호출자가 재정의하지 않는 한 config.reasoning을 사용하므로, 기본적으로 추론 기능을 사용하지 않습니다.
# 사용자가 명시적으로 추론을 요청한 경우에만 Cosmos Reason 2 추론 접미사를 추가합니다
# (cosmos-reason2가 아닌 VLM의 경우 생략). 추론이 비활성화된 경우, 응답에는  블록이  포함되지 않습니다 .
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}

다음 형식을 사용하여 질문에 답변하십시오:


당신의 추론.


`` 태그 바로 뒤에 최종 답변을 작성하십시오 ."
fi

# 3단계를 단독으로 실행하는 경우, 현재 환경/모델에서 누락된 백엔드를 추론합니다.
[ -z "${VLM_BACKEND:-}" ] && {
  if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
    VLM_BACKEND="rtvlm"
  elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
    VLM_BACKEND="nim_cosmos"
  else
    VLM_BACKEND="rtvlm"
  fi
}

# 다중 모드 설정 — 고정된 후보값이 아닌 라이브 에이전트 구성 파일 경로에서 값을 확인합니다.
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
    raise SystemExit("vss-agent에서 VSS_AGENT_CONFIG_FILE이 설정되지 않았습니다")
if not os.path.isabs(p):
    p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
    cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
    "max_fps": int(vu.get("max_fps", 2)),
    "max_frames": int(vu.get("max_frames", 30)),
    "min_pixels": int(vu.get("min_pixels", 3136)),
    "max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "vss-agent에서 video_understanding 구성을 읽는 데 실패했습니다"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "vss-agent에서 받은 JSON 구성이 유효하지 않습니다"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"

# num_frames = min(int(clip_seconds) * max_fps, max_frames), 최소 1 — video_understanding.py와 일치합니다.
# clip_seconds (1단계 endTime-startTime)는 소수 값일 수 있음; 정수 초 단위로 반올림 — bash $((...))
# 은 정수만 허용하며 "15.0"/"1.5"에서는 오류가 발생함. 기본값 15초 -> MAX_FRAMES로 상한 설정.
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1

# NIM Cosmos 경로에서만 Cosmos mm/media kwargs를 적용합니다.
# RT-VLM 모드는 자체 서버 측 전처리 기능을 사용하므로 이러한 kwargs를 받아서는 안 됩니다.
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
  case "$VLM_MODEL" in
    *cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *cosmos*)         MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *)                      MM_KWARGS="" ;;
  esac
fi

curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
  -H "Content-Type: application/json" \
  -d @- <<EOF | jq -r '.choices[0].message.content'
{
  "model": $(jq -Rs . <<< "${VLM_MODEL}"),
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
        {"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
      ]
    }
  ],
  "max_tokens": 1024,
  "temperature": 0.0${MM_KWARGS}
}
EOF

kwargs 블록은 백엔드에 따라 다르게 처리됩니다. nim_cosmos에서 Reason2 변형(nvidia/cosmos-reason2*)은 mm_processor_kwargs.size{shortest_edge,longest_edge} 를 사용하고, 다른 NIM Cosmos 변형(nvidia/cosmos*)은 mm_processor_kwargs.videos_kwargs{min_pixels,max_pixels}; 두 경우 모두 media_io_kwargs.video.num_frames도 전송합니다. rtvlm에서는 Cosmos kwargs가 전송되지 않습니다.

VLM이 … 블록(Cosmos Reason 추론 모드)을 반환하면, 이후의 텍스트만 보고서 본문으로 유지하십시오.

4단계 — 동영상 분석 보고서 템플릿 작성

assets/video-analysis-report.md 파일을 복사하고, 모든 자리 표시자를 채운 다음, 렌더링된 마크다운을 사용자에게 반환하십시오. 원본 자산은 변경하지 마십시오. 렌더링하기 전에 BROWSER_CLIP_URL이 설정되어 있고 비어 있지 않은지 확인한 후, Clip URL 행에서 해당 정확한 값으로 대체하십시오. 출력물에 자리 표시자를 절대 남겨두지 말고, 채워진 셀에 템플릿 지침을 절대 포함하지 않으며, 원시 HOST_IP:30888 URL을 절대 사용하지 마십시오.

모드 B — 특정 기간 내의 인시던트 보고서

1단계 — 시간 범위 및 (선택 사항) 센서

  • start_time / end_time은 반드시 ISO 8601 UTC 형식(YYYY-MM-DDTHH:MM:SS.sssZ)이어야 합니다. "지난 1시간", "오늘"과 같은 상대적 표현은 현재 호스트 시계를 기준으로 해석하십시오.
  • 사용자가 센서 이름을 지정하면, 이를 source + source_type=sensor 형식으로 캡처하십시오. 그렇지 않은 경우, 모든 센서를 대상으로 하는 쿼리를 위해 두 항목 모두 설정하지 마십시오.

2단계 — /vss-query-analytics를 통해 인시던트 가져오기

다음과 같이 /vss-query-analytics (초기화 → tools/call)로 인계합니다:

{
  "jsonrpc": "2.0",
  "method": "tools/call",
  "params": {
    "name": "video_analytics__get_incidents",
    "arguments": {
      "source": "",
      "source_type": "sensor",
      "start_time": "",
      "end_time": "",
      "max_count": 100,
      "includes": ["objectIds", "info"]
    }
  },
  "id": 1
}

읽기 전용 경계 (필수):

  • 모드 B는 엄격히 읽기 전용 분석 데이터 검색입니다. Elasticsearch/VA 데이터에 대한 쓰기, 시드, 백필 또는 변형을 절대 수행해서는 안 됩니다.
  • 금지된 예시: 합성 인시던트 색인 생성, ES에 피처 페이로드 재전송, 보고서를 위해 “데이터를 사용할 수 있게 하기” 위해 write/update/delete API 호출.
  • 요청된 범위/범위에 해당하는 인시던트가 없는 경우, 결과를 비어 있는 것으로 처리하십시오(아래 참조). 데이터를 조작해서는 안 됩니다.

각 인시던트에 대해 다음 항목을 유지하십시오: id, sensorId, timestamp, end, category, place.name, info.verdict, info.reasoning, objectIds 및 클립 URL(일반적으로 info.clip_url, clip_url 또는 응답에 포함된 클립 포인터 필드 중 하나). 보고서에 붙여넣기 전에 모든 클립 URL에 $VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT 재작성(위의 ‘브라우저에서 재생 가능한 클립 URL’ 참조)을 적용하십시오. 원본 값은 사용자의 브라우저가 접근할 수 없는 HOST_IP:30888 URL입니다.

3단계 — 인시던트 범위 보고서 템플릿 작성

assets/incident-range-report.md 파일을 복사한 후, 센서별로(센서 범위가 없는 경우 카테고리별로) 그룹화하고, 판정을 집계하며, 각 인시던트를 타임스탬프/카테고리/판정/이유와 함께 나열하십시오. 원본 자산은 변경하지 마십시오. 모든 인시던트 클립 값은 브라우저에서 재생 가능한 URL로 재작성되어야 합니다. 인시던트에 클립 URL이 없는 경우 클립 행을 생략하십시오. 채워진 셀에는 절대로 템플릿 지침을 포함하지 마십시오.

get_incidents가 0개의 결과를 반환하면, 중단하고 요청된 범위와 범위를 명시한 정확히 한 줄짜리 빈 범위 문장만 반환하십시오. 전체 인시던트 범위 템플릿을 렌더링하지 말고, 인시던트를 임의로 생성하지 말고, 테스트 데이터를 삽입하지 말고, 모드 A로 전환하지 마십시오.

오류 처리

  • 프로브, curl, VLM 호출 또는 /vss-query-analytics 요청이 실패하면 워크플로를 중지하고, 실패한 엔드포인트, HTTP 상태 또는 명령 오류, 그리고 다음으로 취해야 할 유용한 복구 단계를 보고하십시오. 불완전하거나 누락된 데이터를 바탕으로 보고서를 조작해서는 안 됩니다.
  • VLM 응답이 비어 있거나, 형식이 잘못되었거나, 추론 블록만 포함된 경우, 해당 응답 문제를 표면화하고 재시도 전에 모델 준비 상태/로그를 확인할 것을 제안하십시오.
  • 클립 URL을 공개 호스트/포트로 재작성할 수 없는 경우, 렌더링된 보고서에서 이를 생략하고 브라우저에서 재생 가능한 URL을 생성할 수 없음을 명시하십시오.
  • 모드 B의 경우, 누락된 선택적 인시던트 필드(info.reasoning, objectIds, 클립 URL)는 보고서에서 생략된 것으로 취급하되, 누락된 ID, 타임스탬프 또는 카테고리는 보고해야 하는 데이터 품질 오류로 취급하십시오.

참조

  • /vss-manage-video-io-storage — 모드 A 1단계에 대한 센서 목록, 타임라인 및 클립 URL.
  • /vss-query-analytics — 모드 B 2단계에 대한 인시던트 검색(및 판정/추론 보강).
  • /vss-ask-video — 단일 클립에 대한 임시 VLM Q&A(구조화된 보고서가 아님).
  • /vss-summarize-video — 모드 A에서 lvs 프로필이 배포되었을 때 요약 본문을 생성하는 데 사용되며, 보고서 템플릿(4단계)은 여전히 여기에서 채워집니다.
GitHub에서 보기
---
name: vss-generate-video-report
description: Generates video analysis reports by routing to a VLM backend for per-clip analysis or an analytics backend for incident-range reports, with deployment profile verification and URL rewriting.
license: Apache-2.0
---

# Report

Generate a video analysis report by routing to one of two backends — **never via** `POST /generate` on the VSS agent.

| Mode | Backend |
|---|---|
| **A. Video clip** | `/vss-manage-video-io-storage` → clip URL → **VLM chat/completions** |
| **B. Incident range** | `/vss-query-analytics` → incident list → narrative report |

If the request is ambiguous (e.g. "report on `<sensor>`" with no time range and no incident wording), default to **Mode A**. Ask only if the user mentions both a sensor and a time range. See **Examples** below for the request phrasings that route to each mode.

---

## Instructions

1. **Pick the mode** — Mode A for a single recorded clip/sensor video, Mode B when the request names a time range or incidents/alerts (match against *Examples*).
2. **Verify the deployment profile** for that mode under *Deployment prerequisite*; hand off to `/vss-deploy-profile` if its probe fails.
3. **Run that mode's numbered steps** — *Mode A* or *Mode B* below.
4. **Rewrite every user-facing clip URL** with the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` one-liner (*Browser-playable clip URL*) before embedding it in the report.
5. **Return the rendered report markdown** to the user.

Output contract for evaluators:
- Mode A top title MUST be exactly `# Video Analysis Report`.
- Mode B top title MUST be exactly `# Incident Range Report` (never `# Incident Report` or sensor-named variants).
- Mode B MUST include `## Basic Information` with the exact required rows from the template (Report Identifier, Range, Scope, Total Incidents, Confirmed / Rejected / Unverified).

---

## Examples

- "Generate a report for this video" / "report on `<sensor-id>`" → **Mode A**
- "Analyze warehouse_01.mp4" / "create an analysis report on the uploaded video" → **Mode A**
- "Report on incidents from 12:31Z to 12:32Z" → **Mode B**
- "Report on alerts today" / "what incidents happened on `<sensor>` last hour" → **Mode B**
- "Summarize alerts on `<sensor>` between `<t1>` and `<t2>`" → **Mode B**

---

## Negative Triggers

Do **not** use this skill when the request is one of the following:

- Ad-hoc visual Q&A on a clip that do not ask explicitly for a report ("what color is the truck?", "what happens at 00:12?") → use `/vss-ask-video`.
- Archive/semantic similarity retrieval ("find forklifts", "search all videos for tailgating") → use `/vss-search-archive`.
- Read-only incident/metrics lookup without report rendering needs → use `/vss-query-analytics`.
- Deploy/teardown/profile changes ("deploy alerts", "switch profile", "bring up base") → use `/vss-deploy-profile`.
- Real-time alert/rule management requests → use `/vss-manage-alerts`.

Never route reports through VSS-agent `POST /generate`.

---

## Deployment prerequisite

**Mode A** needs the VSS **base** profile (VST + VLM NIM).
**Mode B** needs the VSS **alerts** profile (VA-MCP + Elasticsearch).

Probe:

```bash
# Mode A — VST + VLM reachability
curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/version" >/dev/null

# Mode B — VA-MCP
curl -sf --max-time 5 "http://${HOST_IP}:9901/" >/dev/null
```

If the probe fails, hand off to `/vss-deploy-profile` with `-p base` (Mode A) or `-p alerts` (Mode B). **Always** confirm the deploy with the user first.

---

## Clip URLs: VLM input vs browser report link

VST returns clip URLs using the agent-internal `${HOST_IP}:30888` host:port.
Keep that original URL as `VIDEO_URL` for local / in-cluster VLM frame pulls.
Do **not** rewrite the VLM input URL just to make it browser-playable.

Only create `BROWSER_CLIP_URL` for URLs shown in the rendered report. The
deploy layer exports the browser-facing host:port as `$VSS_PUBLIC_HOST` /
`$VSS_PUBLIC_PORT` (and scheme as `$VSS_PUBLIC_HTTP_PROTOCOL`) in every
profile `.env` — Brev or bare-metal — so the report-link rewrite is:

```bash
: "${VSS_PUBLIC_HOST:?Set VSS_PUBLIC_HOST before rewriting clip URLs}"
: "${VSS_PUBLIC_PORT:?Set VSS_PUBLIC_PORT before rewriting clip URLs}"
VSS_PUBLIC_HTTP_PROTOCOL="${VSS_PUBLIC_HTTP_PROTOCOL:-http}"
BROWSER_CLIP_URL=$(echo "$RAW_URL" | sed -E "s|^https?://[^/]+|${VSS_PUBLIC_HTTP_PROTOCOL}://${VSS_PUBLIC_HOST}:${VSS_PUBLIC_PORT}|")
```

If either required public host value is missing, omit the report-facing clip
link and call out that a browser-playable URL could not be produced; do not
block the local VLM analysis path. Apply the rewrite to **every clip URL
surfaced in the rendered report** (Mode A Step 4 Clip URL row; Mode B
per-incident clip sub-bullet). Leave the VLM `video_url` content block in Mode A
Step 3 on the original internal URL when the VLM is local / in-cluster.

---

## Mode A — Report on a recorded video clip

**If the VSS `lvs` profile is deployed** — `curl -sf --max-time 5 "http://${HOST_IP}:38111/v1/ready"` returns HTTP 200 — run `/vss-summarize-video` to produce the summary, then paste its output into the report template in Step 4 and skip Steps 1–3 (the VLM-direct path). Run Steps 1–3 only when `/v1/ready` is non-200.

### Step 1 — Resolve the clip URL

Hand off to `/vss-manage-video-io-storage` to:

1. List sensors and confirm the named `<sensor-id>` exists (upload first if not).
2. Fetch `/storage/<streamId>/timelines` for the recorded range when the user did not supply `startTime` / `endTime`.
3. Request a clip URL:

   ```bash
   curl -s "http://${HOST_IP}:30888/vst/api/v1/storage/file/<streamId>/url?startTime=<startTime>&endTime=<endTime>&container=mp4&disableAudio=true" | jq -r .videoUrl
   ```

   That gives a direct `mp4` URL that the local / in-cluster VLM can pull frames from. Bind it to `VIDEO_URL` (used by the VLM in Step 3) and set `RAW_URL="$VIDEO_URL"` before applying the report-link rewrite to produce `BROWSER_CLIP_URL` for Step 4 — the user's browser cannot reach `$VIDEO_URL` directly.
   Mode A requires the selected VLM endpoint to be able to fetch `VIDEO_URL`.
   Local NIM/RT-VLM deployments normally can; remote endpoints generally cannot
   fetch `localhost`, private `HOST_IP`, or VST-internal URLs. If the live
   `VLM_ENDPOINT` is remote, surface that reachability requirement instead of
   making a chat request that will fail after `/v1/models` succeeds.

### Step 2 — Resolve VLM endpoint and model

The deploy may serve the VLM through either of two stacks. Both expose an OpenAI-compatible `chat/completions` API — pick whichever is live:

| Backend | Env vars | Typical host endpoint | Picked when |
|---|---|---|---|
| **NIM Cosmos** | `VLM_BASE_URL`, `VLM_NAME`, `VLM_MODE`, `VLM_MODEL_TYPE` | `${VLM_BASE_URL}/v1` (no trailing `/v1` on the env var; the agent appends it) | `VLM_MODEL_TYPE != rtvi` **and** `VLM_MODE` ∈ {`local`, `local_shared`, `remote`} **and** `VLM_BASE_URL` is non-empty |
| **RT-VLM Cosmos** | `RTVI_VLM_BASE_URL`, `RTVI_VLM_MODEL_TO_USE`, `VLM_MODEL_TYPE` | `${RTVI_VLM_BASE_URL}/v1` — if unset, derive from `${HOST_IP}` (`http://${HOST_IP}:8018/v1` for alerts, `http://${HOST_IP}:30082/v1` for base) | `VLM_MODEL_TYPE = rtvi`, or `VLM_MODE=none`, or `VLM_BASE_URL` empty; also the only path for `warehouse` |

Read the live values off the running agent container — do not guess:

```bash
docker exec vss-agent sh -lc '
for k in HOST_IP VLM_MODE VLM_MODEL_TYPE VLM_BASE_URL VLM_NAME RTVI_VLM_BASE_URL RTVI_VLM_MODEL_TO_USE; do
  v="$(printenv "$k")"
  [ -n "$v" ] && printf "%s=%s\n" "$k" "$v"
done
'
```

Do not require `RTVI_VLM_ENDPOINT` from `vss-agent` env; several profiles do not inject it.

Selection rule:

```bash
if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:8018/v1"   # alerts default
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
elif [ -n "${VLM_BASE_URL}" ] && [ "${VLM_MODE}" != "none" ]; then
  VLM_BACKEND="nim_cosmos"
  VLM_ENDPOINT="${VLM_BASE_URL%/}/v1"
  VLM_MODEL="${VLM_NAME}"
else
  VLM_BACKEND="rtvlm"
  VLM_ENDPOINT="${RTVI_VLM_BASE_URL:+${RTVI_VLM_BASE_URL%/}/v1}"
  [ -z "${VLM_ENDPOINT}" ] && VLM_ENDPOINT="http://${HOST_IP}:30082/v1"  # base default
  VLM_MODEL="${RTVI_VLM_MODEL_TO_USE}"
fi
```

Probe `/v1/models` before sending a chat request to confirm the chosen endpoint is alive and the model is loaded:

```bash
curl -sf --max-time 5 "${VLM_ENDPOINT}/models" | jq -r '.data[].id'
```

If the probe fails or the listed ids don't include `${VLM_MODEL}`, fall back to the other backend (or surface the error — never silently pick a model that isn't on the server).

### Step 3 — Call the VLM directly

Use the OpenAI-compatible `chat/completions` endpoint with a `video_url` content block — the same payload shape **and multimodal settings** `video_understanding` builds in `src/vss_agents/tools/video_understanding.py` (`_build_vlm_messages` + the Cosmos `base_vlm.bind(...)` call).

The frame sampling and visual-token (pixel) budget must mirror the **live** `video_understanding` settings for the active profile. **Send `mm_processor_kwargs` and `media_io_kwargs`** so the direct call uses the same frame sampling and pixel budget as the in-agent `video_understanding` tool — omitting them lets the VLM apply its own defaults, so the output diverges from the agent path.

```bash
PROMPT='Describe in detail what happens in the video, with timestamps (start–end in seconds from clip start) for each segment or event. Cover scenes, objects, people, vehicles, and notable actions.'

# Reasoning is OFF by default — matches the base-profile video_understanding config (`reasoning: false`).
# video_understanding.py uses config.reasoning unless the caller overrides it, so default to non-reasoning.
# Append the Cosmos Reason 2 reasoning suffix ONLY when the user explicitly asks for reasoning
# (drop it for non-cosmos-reason2 VLMs). With reasoning off, the response has no <think> block.
if [ "${REASONING:-false}" = "true" ]; then
PROMPT="${PROMPT}

Answer the question using the following format:

<think>
Your reasoning.
</think>

Write your final answer immediately after the </think> tag."
fi

# If Step 3 is run standalone, derive missing backend from current env/model.
[ -z "${VLM_BACKEND:-}" ] && {
  if [ "${VLM_MODEL_TYPE:-}" = "rtvi" ]; then
    VLM_BACKEND="rtvlm"
  elif [[ "${VLM_MODEL:-}" == nvidia/cosmos* ]]; then
    VLM_BACKEND="nim_cosmos"
  else
    VLM_BACKEND="rtvlm"
  fi
}

# Multimodal settings — resolve from the live agent config file path, not hardcoded candidates.
CFG_JSON=$(
docker exec vss-agent python3 -c '
import json, os, yaml
p = os.getenv("VSS_AGENT_CONFIG_FILE")
if not p:
    raise SystemExit("VSS_AGENT_CONFIG_FILE is not set in vss-agent")
if not os.path.isabs(p):
    p = os.path.join("/vss-agent", p.lstrip("./"))
with open(p, encoding="utf-8") as f:
    cfg = yaml.safe_load(f) or {}
vu = (cfg.get("functions", {}) or {}).get("video_understanding", {}) or {}
print(json.dumps({
    "max_fps": int(vu.get("max_fps", 2)),
    "max_frames": int(vu.get("max_frames", 30)),
    "min_pixels": int(vu.get("min_pixels", 3136)),
    "max_pixels": int(vu.get("max_pixels", 8388608)),
}))
')
)
[ -n "${CFG_JSON}" ] || { echo "Failed to read video_understanding config from vss-agent"; exit 1; }
jq -e . >/dev/null <<< "${CFG_JSON}" || { echo "Invalid config JSON from vss-agent"; exit 1; }
MAX_FPS="$(jq -r '.max_fps' <<< "${CFG_JSON}")"
MAX_FRAMES="$(jq -r '.max_frames' <<< "${CFG_JSON}")"
MIN_PIXELS="$(jq -r '.min_pixels' <<< "${CFG_JSON}")"
MAX_PIXELS="$(jq -r '.max_pixels' <<< "${CFG_JSON}")"

# num_frames = min(int(clip_seconds) * max_fps, max_frames), min 1 — matches video_understanding.py.
# clip_seconds (Step 1 endTime-startTime) may be fractional; truncate to integer seconds — bash $((...))
# is integer-only and errors on "15.0"/"1.5". Default 15s -> caps at MAX_FRAMES.
CLIP_SECONDS=$(awk -v s="${CLIP_SECONDS:-15}" 'BEGIN{printf "%d", s}')
NUM_FRAMES=$(( CLIP_SECONDS * MAX_FPS ))
[ "$NUM_FRAMES" -gt "$MAX_FRAMES" ] && NUM_FRAMES=$MAX_FRAMES
[ "$NUM_FRAMES" -lt 1 ] && NUM_FRAMES=1

# Only apply Cosmos mm/media kwargs on the NIM Cosmos path.
# RT-VLM mode uses its own server-side preprocessing and should not receive these kwargs.
MM_KWARGS=""
if [ "${VLM_BACKEND}" = "nim_cosmos" ]; then
  case "$VLM_MODEL" in
    *cosmos-reason2*) MM_KWARGS=", \"mm_processor_kwargs\": {\"size\": {\"shortest_edge\": ${MIN_PIXELS}, \"longest_edge\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *cosmos*)         MM_KWARGS=", \"mm_processor_kwargs\": {\"videos_kwargs\": {\"min_pixels\": ${MIN_PIXELS}, \"max_pixels\": ${MAX_PIXELS}}}, \"media_io_kwargs\": {\"video\": {\"num_frames\": ${NUM_FRAMES}}}" ;;
    *)                      MM_KWARGS="" ;;
  esac
fi

curl -s --connect-timeout 5 --max-time 120 -X POST "${VLM_ENDPOINT}/chat/completions" \
  -H "Content-Type: application/json" \
  -d @- <<EOF | jq -r '.choices[0].message.content'
{
  "model": $(jq -Rs . <<< "${VLM_MODEL}"),
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": $(jq -Rs . <<< "${PROMPT}")},
        {"type": "video_url", "video_url": {"url": $(jq -Rs . <<< "${VIDEO_URL}")}}
      ]
    }
  ],
  "max_tokens": 1024,
  "temperature": 0.0${MM_KWARGS}
}
EOF
```

> The kwargs block is backend-aware: on `nim_cosmos`, Reason2 variants (`nvidia/cosmos-reason2*`) use `mm_processor_kwargs.size{shortest_edge,longest_edge}` and other NIM Cosmos variants (`nvidia/cosmos*`) use `mm_processor_kwargs.videos_kwargs{min_pixels,max_pixels}`; both also send `media_io_kwargs.video.num_frames`. On `rtvlm`, no Cosmos kwargs are sent.

If the VLM returns a `<think>…</think>` block (Cosmos Reason reasoning mode), keep only the text after `</think>` as the report body.

### Step 4 — Fill the Video Analysis Report template

Copy [`assets/video-analysis-report.md`](assets/video-analysis-report.md), fill every placeholder, and return the rendered markdown to the user. Keep the source asset unchanged. Before rendering, verify `BROWSER_CLIP_URL` is set and non-empty, then replace `<BROWSER_CLIP_URL>` with that exact value in the `Clip URL` row. Never leave the placeholder in the output, never include template instructions in a filled cell, and never use the raw `HOST_IP:30888` URL.

---

## Mode B — Report on incidents in a time range

### Step 1 — Resolve the time range and (optionally) sensor

- `start_time` / `end_time` must be ISO 8601 UTC (`YYYY-MM-DDTHH:MM:SS.sssZ`). Resolve relative phrases ("last hour", "today") against the current host clock.
- If the user names a sensor, capture it as `source` + `source_type=sensor`. Otherwise leave both unset for an all-sensors query.

### Step 2 — Fetch incidents via `/vss-query-analytics`

Hand off to `/vss-query-analytics` (initialize → `tools/call`) with:

```json
{
  "jsonrpc": "2.0",
  "method": "tools/call",
  "params": {
    "name": "video_analytics__get_incidents",
    "arguments": {
      "source": "<sensor-id-or-omit>",
      "source_type": "sensor",
      "start_time": "<ISO>",
      "end_time": "<ISO>",
      "max_count": 100,
      "includes": ["objectIds", "info"]
    }
  },
  "id": 1
}
```

Read-only boundary (mandatory):
- Mode B is strictly read-only analytics retrieval. Never write, seed, backfill, or mutate Elasticsearch/VA data.
- Forbidden examples: indexing synthetic incidents, replaying fixture payloads into ES, calling write/update/delete APIs to "make data available" for the report.
- If no incidents exist for the requested range/scope, handle as empty results (see below); do not fabricate data.

For each incident keep: `id`, `sensorId`, `timestamp`, `end`, `category`, `place.name`, `info.verdict`, `info.reasoning`, `objectIds`, and the clip URL (commonly `info.clip_url`, `clip_url`, or whichever clip-pointer field the response carries). **Apply the `$VSS_PUBLIC_HOST:$VSS_PUBLIC_PORT` rewrite (see *Browser-playable clip URL* above) to every clip URL before pasting it into the report** — the raw value is a `HOST_IP:30888` URL the user's browser cannot reach.

### Step 3 — Fill the Incident Range Report template

Copy [`assets/incident-range-report.md`](assets/incident-range-report.md), then group by sensor (or by category if no sensor scope), tally verdicts, and list each incident with timestamp / category / verdict / reasoning. Keep the source asset unchanged. Every incident clip value must be a rewritten browser-playable URL; omit the clip line when the incident carries no clip URL. Never include template instructions in a filled cell.

If `get_incidents` returns zero results, STOP and return exactly a one-line empty-range statement naming the requested range and scope. Do not render the full Incident Range template, do not invent incidents, do not seed test data, and do not fall back to Mode A.

---

## Error Handling

- If a probe, `curl`, VLM call, or `/vss-query-analytics` request fails, stop the workflow and report the failing endpoint, HTTP status or command error, and the next useful recovery step. Do not fabricate a report from partial or missing data.
- If the VLM response is empty, malformed, or contains only a reasoning block, surface that response problem and suggest checking model readiness/logs before retrying.
- If a clip URL cannot be rewritten to the public host/port, omit it from the rendered report and call out that the browser-playable URL could not be produced.
- For Mode B, treat missing optional incident fields (`info.reasoning`, `objectIds`, clip URL) as omissions in the report, but treat missing `id`, `timestamp`, or `category` as a data-quality error that should be reported.

---

## Cross-Reference

- **`/vss-manage-video-io-storage`** — sensor list, timelines, and clip URL for Mode A Step 1.
- **`/vss-query-analytics`** — incident retrieval (and verdict / reasoning enrichment) for Mode B Step 2.
- **`/vss-ask-video`** — ad-hoc VLM Q&A on a single clip (not a structured report).
- **`/vss-summarize-video`** — used by Mode A to produce the summary body when the `lvs` profile is deployed; the report template (Step 4) is still filled here.

vss-generate-video-report 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/vss-generate-video-report # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용합니다.
저장소 NVIDIA/skills

관련 스킬

microservices-patterns
업데이트 된 시간 2026년 6월 29일
jpa-patterns
업데이트 된 시간 2026년 6월 30일
fabric-lakehouse
업데이트 된 시간 2026년 6월 30일
prisma-expert
업데이트 된 시간 2026년 6월 29일
OR