옵션
집집 Skill DevOps 및 CI/CD tao-run-inference-service

tao-run-inference-service

NVIDIA/skills NVIDIA/skills

컨테이너 실행을 적절한 플랫폼 스킬에 위임하여 특정 네트워크 아키텍처에 대한 TAO 추론 마이크로서비스를 시작, 쿼리 및 중지합니다.

...모든 것을 확장하십시오
1
업데이트 된 시간 2026년 9월 27일

TAO 추론 마이크로서비스

사용 방법

추론 서비스를 시작하려면:

  1. 필요한 입력값을 수집하고(1절), 컨테이너 이미지를 해결합니다(2절).
  2. 작업 페이로드와 내부 명령을 빌드합니다(3~4.1절); 다음을 사용합니다 references/code-templates.yaml → job_payload_builder.
  3. Read skills/platform//SKILL.md 를 사용하여 컨테이너를 시작합니다(4.2절).
  4. 서비스 레지스트리에 등록하고 준비 상태를 폴링합니다(4.3절). references/code-templates.yaml → registry_write. 및 readiness_check.

추론 요청을 보내려면:

  1. 6.0절에 따라(다음과 같이) 어떤 서비스가 요청을 수신할지 결정합니다 job_id, network_arch, 또는 여러 서비스가 실행 중인 경우 사용자의 명시적인 선택을 통해 결정합니다. — 두 개 이상의 서비스가 존재할 때 절대 "latest"로 묵시적으로 기본 설정하지 마십시오), 그런 다음 references/code-templates.yaml → request.registry_read 에서 엔드포인트를 읽어옵니다. job_id.
  2. 요청 본문을 구성하기 전에, 사용자에게 vLLM 스타일의 샘플링 매개변수(6.1절)를 입력하도록 요청하십시오. max_tokens, top_p, temperature (및 아키텍처별 추가 항목)을 기본값과 함께 표시하고, 사용자가 각 항목을 재설정하거나 건너뛰어 기본값을 수락할 수 있도록 하십시오. 절대 기본값을 자동으로 사용해서는 안 됩니다.
  3. 6.2절에 따라 본문을 구성하고 전송하며, 6.3절에 따라 응답을 처리합니다.

서비스를 중지하려면: references/code-templates.yaml → stop.registry_read 를 참조하여 job_id를 확인한 후, skills/platform//SKILL.md를 읽고, 5절을 따르십시오.

참조 데이터(스키마, 매핑, 유효한 값 — 지침 없음):

  • references/service.yaml — 이미지 매핑, 유효한 network_arch 이름, 작업 페이로드 스키마, 환경 변수 이름, 시크릿 분류.
  • references/request.yaml — 엔드포인트 정의, 요청 필드 스키마, 응답 형식, 코드 예제.
  • references/code-templates.yaml — 페이로드 생성, 레지스트리 기록, 준비 상태 확인, 중지/요청 흐름을 위한 Python 템플릿.

시크릿 규칙 (이 스킬에서 생성된 모든 코드 블록에 적용됨)

사용자에게 프롬프트에 시크릿 값을 직접 입력하도록 요청해서는 안 됩니다. 모든 시크릿 값에 대해:

  1. 사용자에게 어떤 환경 변수를 설정해야 하는지 알려주십시오(예: export HF_TOKEN=...).
  2. 다음과 같이 해당 값을 읽어들이는 코드를 생성하십시오. os.environ["VAR_NAME"] — 절대로 값을 하드코딩하거나, 보간하거나, 프롬프트를 통해 입력받지 마십시오.

비밀 환경 변수(전체 목록은 references/service.yaml → secrets_handling): HF_TOKEN, WANDB_API_KEY, CLEARML_API_ACCESS_KEY, CLEARML_API_SECRET_KEY, TAO_API_KEY, TAO_USER_KEY.

프롬프트에서 수집해도 안전한 항목: network_arch, model_path, num_gpus, 프롬프트 텍스트, WANDB_* 구성 URL, CLEARML_*_HOST URL.

1. 사용자로부터 수집해야 할 정보

입력 역할
network_arch 컨테이너 이미지를 선택하며, 아키텍처별 내부 명령어 형식(references/service.yaml → container_commands.) 및 neural_network_name 해당되는 경우 작업 JSON에서 선택합니다. valid_network_arch_config_basenames in references/service.yaml (예: cosmos-rl, cosmos-predict2.5).
model_path 훈련된 모델 체크포인트. 유효한 형식: hf_model:/// (HuggingFace Hub — 게이트형 모델의 경우 HF_TOKEN 게이트형 모델용) 또는 로컬 컨테이너 파일 시스템 경로여야 합니다. 클라우드 URI(s3://, gs://, az://)은 지원되지 않습니다 — 추론 서비스는 클라우드 스토리지에 대한 종속성이 없습니다. 항상 사용자에게 확인을 요청하고, 절대 자리 표시자로 대체하지 마십시오. references/service.yaml → model_path_protocols.
platform 컴퓨팅 플랫폼: local-docker, brev, slurm를 참조하거나, kubernetes.
num_gpus 기본값은 1이며, 추론을 위해서는 최소 1이 필요합니다.

2. 이미지 해상도

각 network_arch 에는 {network_arch}.config.json라는 이름의 사이드카 구성 파일을 가지고 있습니다. 컨테이너 이미지를 다음과 같이 지정하십시오:

  1. 읽기 {network_arch}.config.json 파일을 읽고 api_params.image (예: COSMOS_RL). 이는 docker_image_defaults.mapping in references/service.yaml.
  2. 매핑에서 해당 키를 찾아보세요. 호스트 환경 변수 IMAGE_ 가 설정되어 있다면(예: IMAGE_COSMOS_RL), 매핑된 기본값보다 우선합니다.
  3. 매핑된 값은 일반적으로 리포지토리 루트 versions.yaml 매니페스트 내의 점으로 구분된 키입니다(예: tao_toolkit.cosmos_rl). 이를 조회하여 구체적인 nvcr.io/... 이미지 URI로 변환하려면 versions.yaml → images..를 조회하여 구체적인이미지 URI로 변환합니다. 절대 URI는 변경 없이 그대로 전달되므로, 전체 URI를 포함하는 IMAGE_ 전체 URI를 포함하는 환경 변수 재정의도 여전히 작동합니다. 이를 위한 Python 헬퍼는 references/code-templates.yaml.
  4. 설정 파일이 없거나 api_params.image 비어 있는 경우, COSMOS_RL 키로 대체합니다.

또한 구성 파일에는 spec_params.inference.model_path 이 키는 폴더 경로와 파일 경로의 구분을 결정합니다. 값에 folder가 포함되어 있으면, 컨테이너는 해당 경로를 디렉터리로 취급합니다.

3. 환경 변수 (콜백 없음)

다음에서 이를 설정하십시오 env_payload 인코딩 전에 env_json. 다음은 설정하지 마십시오 TAO_LOGGING_SERVER_URL 또는 TAO_ADMIN_KEY.

TAO_EXECUTION_BACKEND — 플랫폼과 일치해야 합니다:

플랫폼 TAO_EXECUTION_BACKEND 값
local-docker local-docker
brev local-docker
slurm slurm
쿠버네티스 local-k8s

CLOUD_BASED — 항상 "False" 이 스킬의 경우 ( TAO_LOGGING_SERVER_URL).

GPU 환경 변수 — 플랫폼 스킬이 GPU 주입을 자동으로 처리하지 않는 경우에만 필요합니다:

  • Tegra / Jetson: --runtime=nvidia 다음과 같이 NVIDIA_DRIVER_CAPABILITIES=all 그리고 NVIDIA_VISIBLE_DEVICES=.
  • 표준 x86 + nvidia-container-toolkit: Docker를 사용하십시오 device_requests. 플랫폼 스킬이 이를 처리합니다.

4. 여러 플랫폼에서 실행

작업 페이로드와 내부 명령어(1~3절)는 플랫폼에 구애받지 않습니다. 각 플랫폼에 대해서는 실행 코드를 생성하기 전에 skills/platform//SKILL.md에서 사전 점검 및 자격 증명에 대한 내용을 확인하십시오.

4.1 내부 명령어 빌드(아키텍처별)

내부 명령어의 형식은 network_arch에 명시된 대로이며, 통일된 템플릿은 없습니다. references/service.yaml → container_commands.; 해당 항목이 없다면 해당 아키텍처는 지원되지 않으므로 작업을 중단하고 문의하십시오. references/code-templates.yaml → job_payload_builder.에서 해당 아키텍처에 맞는 하위 블록을 선택하십시오. 명령어 앞에 umask 0 && 를 붙이고, 모든 플랫폼(local-docker, brev, slurm, kubernetes)에서 동일하게 유지하십시오.

모든 아키텍처 공통:

  • job_id: fresh uuid.uuid4() — 컨테이너 이름 및 레지스트리 키로 사용됩니다.
  • image: 2절에 따라 해결합니다.
  • 시크릿(access_key, secret_key, HF_TOKEN등)은 런타임 시 환경 변수에서 읽어옵니다 — 절대 하드코딩하지 말고, 로그에 기록하거나 출력하지 마십시오.

아키텍처별 참고 사항(자세한 내용은 references/service.yaml → container_commands):

  • cosmos-rl — 단일 --job '' --docker_env_vars '' 블롭; json.dumps(...) + shlex.quote(...). env_payload 다음 내용을 포함합니다 TAO_EXECUTION_BACKEND (3절 표 참조), TAO_API_JOB_ID, CLOUD_BASED=False. 추론 서비스는 클라우드 스토리지에 대한 의존성이 없으며, HF_TOKEN (게이트가 적용된 HuggingFace 모델의 경우) 적용되는 유일한 인증 환경 변수입니다.
  • cosmos-predict2.5 — 플래그 형식 cosmos_predict inference_microservice start ... --port 8080 (접두사 없음 setup. 접두사 없음; tyro.conf.OmitArgPrefixes). --job/--docker_env_vars 는 허용되지 않습니다. 다음을 변환하십시오 model_path 를 --checkpoint-path (로컬 경로) 또는 --model (hf_model://); 클라우드 URI는 거부됩니다. 적용되는 유일한 cred 환경 변수는 HF_TOKEN 게이트가 적용된 HuggingFace 모델에 한해 적용됩니다. 요청별 매개변수(prompt, inference_type, num_output_frames, guidance, seed, num_steps, negative_prompt)는 시작 시가 아닌 요청 본문에 포함됩니다. TAO_EXECUTION_BACKEND/TAO_API_JOB_ID/CLOUD_BASED 는 사용되지 않으므로 생략할 수 있습니다.

4.2 플랫폼 스킬에 실행 위임

skills/platform//SKILL.md를 읽고 안내에 따라 컨테이너를 시작하십시오.

기본 매개변수(모든 플랫폼):

매개변수 값
image 해결된 컨테이너 이미지(2절)
command inner — 4.1절에서 생성된 셸 문자열
gpu_count num_gpus
env_vars env_payload
작업/컨테이너 이름 job_id — 레지스트리가 이를 참조할 수 있도록 4.1절의 UUID와 일치해야 함
host_port (local-docker, brev) 컨테이너 포트 8080에 바인딩할 호스트 측 포트. 기본값 8080이지만, 동시 실행되는 서비스마다 고유해야 합니다. — 아래의 포트 할당 규칙을 참조하십시오.

플랫폼별 추가 입력 사항:

플랫폼 추가 입력
local-docker 기본 항목 외에는 없음
brev instance_id (선택 사항 — 기존 인스턴스 재사용); 다중 자격 증명/다중 작업 공간 계정의 경우 다음도 포함 cloud_cred_id and workspace_group_id 첫 생성 시 — 참조: skills/platform/tao-run-on-brev/SKILL.md
slurm partition 및 account — SLURM_PARTITION/SLURM_ACCOUNT 환경 변수를 확인하십시오; 설정되어 있지 않은 경우 사용자에게 확인을 요청하십시오
kubernetes namespace (기본값: default); image_pull_secret ( nvcr.io 이미지)

포트 바인딩 (local-docker 및 brev): -p :8080 매개변수를 전달할 수 있고 컨테이너 이름이 job_id 정확히 일치하도록 하십시오.

포트 할당 규칙 (local-docker 및 brev, 동시 실행 서비스의 경우 필수): 서비스를 시작하기 전에 레지스트리(/tmp/tao-inf-ms-state.json)을 읽고, host_port 값 집합을 수집합니다. brev의 경우 동일한 instance_id). 해당 집합에 포함되지 않은 8080부터 시작하는 가장 낮은 사용 가능한 포트를 선택합니다 — 예: host_port = next(p for p in range(8080, 8200) if p not in used_ports). 기본값은 8080 설정은 다른 서비스가 실행 중이지 않을 때만 적용됩니다. 이것이 바로 “3개의 서비스를 시작하고, 각 서비스에 서로 다른 host_url"가 작동하는 이유입니다. 이 단계가 없다면, 서비스 2와 3은 bind: address already in use. SLURM과 쿠버네티스는 자체 플랫폼 메커니즘을 통해 고유한 엔드포인트를 확보하므로 이 단계가 필요하지 않습니다.

4.3 시작 후: 서비스 레지스트리 및 엔드포인트

플랫폼에서 컨테이너가 실행 중임을 확인하는 즉시 서비스 레지스트리를 작성합니다. 레지스트리(/tmp/tao-inf-ms-state.json)는 job_id; "latest" 항상 가장 최근에 시작된 서비스를 가리킵니다.

Python 템플릿에 대해서는 references/code-templates.yaml → registry_write. Python 템플릿을 참조하십시오.

플랫폼 host_url platform_job_id 작성하기 전의 추가 단계
local-docker http://localhost:{host_port} — 없음
brev http://{brev_ip}:{host_port} — brev ls → 인스턴스 IP 확인 (localhost 원격 VM에서는 유효하지 않음)
slurm http://localhost:{host_port} SLURM 스케줄러 작업 ID '실행 중' 상태가 될 때까지 대기; SSH 포트 포워딩 localhost:{host_port}→{node}:8080
kubernetes http://{external_ip}:8080 k8s 작업 이름 kubectl expose job … --type=LoadBalancer; 외부 IP가 표시될 때까지 대기

레지스트리에 기록한 후, job_id와 URL을 출력합니다:

print(f"Inference service started.")
print(f"  Job ID : {job_id}")
print(f"  Arch   : {network_arch}")
print(f"  URL    : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")

그런 다음 준비 상태를 폴링합니다 — 참조 references/code-templates.yaml → readiness_check. 컨테이너는 백그라운드에서 모델을 로드하므로, 200이 반환되기 전에는 요청을 보내지 마십시오.

5. 추론 서비스 중지

사용자에게 중지할 job_id 중지할 작업을 요청합니다. 사용자가 지정하지 않으면 기본값으로 state["latest"] 를 기본값으로 설정하고, 중지될 job_id를 확인합니다. references/code-templates.yaml → stop.registry_read를 사용하여 레지스트리를 조회한 후, skills/platform//SKILL.md를 조회하여 해당 취소/중지 메커니즘을 사용하십시오.

플랫폼 전달할 식별자 추가 정리
local-docker job_id_to_stop — 컨테이너 이름 없음
brev job_id_to_stop — 컨테이너 이름 없음
slurm entry["platform_job_id"] — SLURM 작업 ID pkill -f "ssh.*-L.*{entry['host_port']}"
kubernetes entry["platform_job_id"] — k8s 작업 이름 kubectl delete svc {entry["platform_job_id"]} -n

여기서 entry = state[job_id_to_stop]. 중지 후 레지스트리를 정리하십시오: references/code-templates.yaml → stop.registry_cleanup.

6. 추론 요청 전송

6.0 이 요청을 수신할 서비스 결정 (필수)

각 요청은 해당 모델을 실행하는 특정 서비스로 라우팅되어야 합니다. 라우팅은 다음을 통해 이루어집니다. job_id — 레지스트리는 network_arch 항목별로 저장하므로, 사용자가 모델 이름을 지정할 때 아키텍처별로 대상을 결정할 수 있습니다 job_id. 다음 규칙을 순서대로 적용하십시오:

  1. 사용자가 명시적인 job_id를 제공한 경우 → 이를 사용합니다. 해당 항목이 state.
  2. 사용자가 network_arch을 명시한 경우(예: "이것을 cosmos-rl 서비스로 보내라") → 일치하는 항목을 조회합니다: candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch].
    • 정확히 하나의 항목이 일치할 경우 → 해당 항목을 사용합니다.
    • 일치하는 항목이 여러 개인 경우 → 후보 항목 job_id및 해당 started_at; 자동으로 선택하지 마십시오.
    • 일치하는 항목 없음 → 중단하고 사용자에게 해당 아키텍처용 서비스가 실행 중이지 않다고 알립니다.
  3. job_id와 network_arch가 모두 없는 경우 → 해당 아키텍처가 아닌 항목의 수를 계산합니다."latest" 항목 수를 세고 state:
    • 정확히 하나의 서비스가 실행 중일 경우 → 해당 서비스를 사용합니다.
    • 두 개 이상인 경우 → state["latest"]를 기본값으로 조용히 설정하지 마십시오. 전체 목록을 사용자에게 표시하고(job_id, network_arch, host_url)을 표시하고 명시적인 선택을 요구하십시오. "latest" 포인터는 단일 서비스 워크플로우를 위한 편의 기능일 뿐, 여러 서비스가 공존할 때의 라우팅 대체 수단이 아닙니다.
    • 0개 → 중지하고 사용자에게 먼저 서비스를 시작하라고 알립니다.

해결이 완료된 후, 레지스트리에서 엔드포인트를 읽어오되 (references/code-templates.yaml → request.registry_read)에서 엔드포인트를 읽어들이되, 해결된 job_id 를 user_provided_job_id로 전달합니다. 사용자에게 "job_id=… arch=… url=…로 전송 중"이라고 확인 메시지를 표시합니다. 서비스가 아직 로딩 중일 수 있으므로, 먼저 준비 상태를 폴링합니다(references/code-templates.yaml → readiness_check).

전송 전에 교차 확인: 사용자가 제공한 요청 본문에 아키텍처별 필드(예: guidance / num_steps / seed / negative_prompt → cosmos-predict2.5; 필수 image_url/video_url 콘텐츠 항목 → cosmos-rl), 해당 필드가 state[job_id]["network_arch"]. 불일치 시 중단하고 확인을 요청하십시오 — cosmos-predict2.5 본문을 cosmos-rl 서비스로 전송하면 컨테이너에서 4xx/5xx 오류가 발생하며, 이는 여기서 감지하는 것보다 진단하기 어렵습니다.

6.1 샘플링 매개변수 — 각 요청 전 필수 사용자 프롬프트

요청 본문을 구성하기 전에, vLLM 스타일의 샘플링 매개변수에 대해 사용자에게 명시적으로 입력 요청을 해야 합니다. 기본값을 자동으로 적용해서는 안 됩니다. 필드당 하나의 질문으로 구성된 구조화된 입력 요청을 사용하여, 다음을 수행해야 합니다:

  1. 적용 가능한 모든 필드를 해당 유형 및 기본값과 함께 나열해야 합니다.
  2. 사용자가 필드를 건너뛰거나 수락하여 해당 필드의 기본값을 적용할 수 있도록 해야 합니다. 값을 입력하는 것은 절대 필수 사항이 아닙니다.
  3. 모든 필드를 한 번에 수집해야 합니다.

프롬프트가 끝난 후, 사용자가 입력한 각 값을 그대로 적용하고, 건너뛴 필드에 대해서는 기본값을 대입해야 합니다. 임의로 값을 생성하거나 몰래 제한해서는 안 됩니다.

필드 목록, 기본값 및 아키텍처별 적용 가능성: references/request.yaml → chat_completions_request_body (기본 샘플링 필드: max_tokens, top_p, temperature) 및 network_arch_constraints. (아키텍처별 재정의 및 다음과 같은 추가 항목: guidance/num_steps/seed/negative_prompt for cosmos-predict2.5). 특정 필드가 활성 아키텍처에서 지원되지 않는 것으로 표시된 경우, 해당 필드에 대한 입력을 요청하지 말고 본문에 포함하지 마십시오.

6.2 요청 형식

다음과 같은 POST 를 {BASE_URL}/v1/chat/completions 로보내십시오. Content-Type: application/json 를 사용하여 최소 300초의 타임아웃을 설정하여 전송하십시오. 본문은 OpenAI 호환 형식(vLLM 채팅 완성)입니다. 전체 필드 스키마 및 콘텐츠 항목 형식(text / image_url / video_url)에 대해서는 references/request.yaml → chat_completions_request_body 에서 전체 필드 스키마 및 콘텐츠 항목 형식(text / image_url / video_url)을, code_examples 여기에서 바로 실행 가능한 Python 및 curl 샘플을 확인할 수 있습니다.

제약 사항: 첫 번째 사용자 메시지만 처리됩니다. 요청 본문에 비밀 값을 포함할 수 없습니다. 네트워크별 제약 사항이 적용됩니다(예: cosmos-rl의 경우 모든 요청에 이미지나 동영상이 포함되어야 하며, cosmos-rl은 data: URI를 거부함)은 references/request.yaml → network_arch_constraints.

6.3 응답 처리

HTTP 상태 의미 작업
200 성공 — choices[0].message.content 생성된 텍스트가 있습니다 결과 읽기
202 서버 초기화 중이거나 모델이 아직 로딩 중입니다 잠시 후 다시 시도하십시오
503 초기화 실패, 모델 로딩 실패 또는 모델이 아직 준비되지 않음 확인 error.type: model_not_ready → 다시 시도; initialization_error / model_load_error → 중단하고 로그 확인
400 JSON 본문이 없거나 비어 있습니다 요청 수정
500 추론 중 처리되지 않은 예외 발생 컨테이너 로그 확인

202 및 503 오류의 경우, 본문에는 다음이 포함됩니다 {"error": {"type": "", "message": ""}}. 참조: container_response_shapes 에서 references/request.yaml 에서 오류 유형 문자열을 참조하십시오.

GitHub에서 보기
---
name: tao-run-inference-service
description: Start, query, and stop a TAO inference microservice for a specific network architecture by delegating container execution to the appropriate platform skill.
license: Apache-2.0
---

# TAO Inference Microservice

## Instructions

**To start an inference service:**
1. Collect required inputs (Section 1) and resolve the container image (Section 2).
2. Build the job payload and inner command (Sections 3–4.1); use `references/code-templates.yaml` → `job_payload_builder`.
3. Read `skills/platform/<platform>/SKILL.md` and start the container (Section 4.2).
4. Write the service registry and poll readiness (Section 4.3); use `references/code-templates.yaml` → `registry_write.<platform>` and `readiness_check`.

**To send an inference request:**
1. Resolve which service receives the request per Section 6.0 (by `job_id`, by `network_arch`, or by explicit user choice when multiple services run — **never silently default to `"latest"` when more than one service exists**), then read the endpoint from `references/code-templates.yaml` → `request.registry_read` with the resolved `job_id`.
2. **Before building the request body, prompt the user for the vLLM-style sampling parameters (Section 6.1).** Present `max_tokens`, `top_p`, `temperature` (and any per-arch extras) with their defaults; let the user override or skip each one to accept the default. Never silently use defaults.
3. Build and send the body per Section 6.2; handle the response per Section 6.3.

**To stop a service:** Read `references/code-templates.yaml` → `stop.registry_read` to resolve the job_id, read `skills/platform/<platform>/SKILL.md`, then follow Section 5.

**Reference data** (schemas, mappings, valid values — no instructions):
- **`references/service.yaml`** — image mappings, valid `network_arch` names, job payload schema, env var names, secrets classification.
- **`references/request.yaml`** — endpoint definition, request field schema, response shapes, code examples.
- **`references/code-templates.yaml`** — Python templates for payload building, registry writes, readiness checks, and stop/request flows.

---

## Secrets rule (applies to every generated code block in this skill)

**Never ask the user to type a secret value into a prompt.** For every secret value:
1. Tell the user which environment variable to set (e.g. `export HF_TOKEN=...`).
2. Generate code that reads it with `os.environ["VAR_NAME"]` — never hard-code, interpolate, or prompt for the value.

**Secret env vars** (full list in `references/service.yaml` → `secrets_handling`):
`HF_TOKEN`, `WANDB_API_KEY`, `CLEARML_API_ACCESS_KEY`, `CLEARML_API_SECRET_KEY`, `TAO_API_KEY`, `TAO_USER_KEY`.

**Safe to collect in the prompt:** `network_arch`, `model_path`, `num_gpus`, prompt text, `WANDB_*` config URLs, `CLEARML_*_HOST` URLs.

---

## 1. What to collect from the user

| Input | Role |
|--------|------|
| **`network_arch`** | Chooses container image, the per-arch inner command shape (`references/service.yaml` → `container_commands.<network_arch>`), and `neural_network_name` in the job JSON when applicable. Must match a basename in `valid_network_arch_config_basenames` in `references/service.yaml` (e.g. `cosmos-rl`, `cosmos-predict2.5`). |
| **`model_path`** | The trained model checkpoint. Valid forms: `hf_model://<org>/<model>` (HuggingFace Hub — set `HF_TOKEN` for gated models) or a local container filesystem path. Cloud URIs (`s3://`, `gs://`, `az://`) are NOT supported — the inference service has no cloud-storage dependency. Always ask the user; never substitute a placeholder. See `references/service.yaml` → `model_path_protocols`. |
| **`platform`** | Compute platform: `local-docker`, `brev`, `slurm`, or `kubernetes`. |
| **`num_gpus`** | Defaults to **1**; minimum **1** for inference. |

---

## 2. Image resolution

Each `network_arch` has a sidecar config file named `{network_arch}.config.json`. Resolve the container image as follows:

1. Read `{network_arch}.config.json` and take `api_params.image` (e.g. `COSMOS_RL`). This is a key into `docker_image_defaults.mapping` in `references/service.yaml`.
2. Look up that key in the mapping. If the host env var `IMAGE_<KEY>` is set (e.g. `IMAGE_COSMOS_RL`), it overrides the mapped default.
3. The mapped value is normally a dotted key into the repo-root `versions.yaml` manifest (e.g. `tao_toolkit.cosmos_rl`). Resolve it to a concrete `nvcr.io/...` image URI by looking up `versions.yaml` → `images.<group>.<name>`. Absolute URIs pass through unchanged, so an `IMAGE_<KEY>` env-var override that contains a full URI still works. The Python helper for this lives in `references/code-templates.yaml`.
4. If the config file is missing or `api_params.image` is empty, fall back to the `COSMOS_RL` key.

The config file also has `spec_params.inference.model_path` which drives **folder vs file** path semantics: if the value contains the substring `folder`, the container treats the path as a directory.

---

## 3. Environment variables (no callbacks)

Set these in `env_payload` before encoding `env_json`. Do **not** set `TAO_LOGGING_SERVER_URL` or `TAO_ADMIN_KEY`.

**`TAO_EXECUTION_BACKEND`** — must match the platform:

| Platform | `TAO_EXECUTION_BACKEND` value |
|----------|-------------------------------|
| local-docker | `local-docker` |
| brev | `local-docker` |
| slurm | `slurm` |
| kubernetes | `local-k8s` |

**`CLOUD_BASED`** — always `"False"` for this skill (disables callback posting to `TAO_LOGGING_SERVER_URL`).

**GPU env vars** — only needed when the platform skill does not handle GPU injection automatically:
- Tegra / Jetson: `--runtime=nvidia` with `NVIDIA_DRIVER_CAPABILITIES=all` and `NVIDIA_VISIBLE_DEVICES=<ids>`.
- Standard x86 + nvidia-container-toolkit: use Docker `device_requests`. The platform skill handles this.

---

## 4. Executing across platforms

The job payload and inner command (Sections 1–3) are **platform-agnostic**. For each platform, read **`skills/platform/<name>/SKILL.md`** for preflight checks and credentials **before** generating any execution code.

### 4.1 Build the inner command (per arch)

The inner-command shape is **per `network_arch`** — there is no uniform template. Look up the per-arch entry in `references/service.yaml` → `container_commands.<network_arch>`; if not present, the arch is unsupported — stop and ask. Pick the matching sub-block in `references/code-templates.yaml` → `job_payload_builder.<network_arch>`. Prefix the command with `umask 0 &&` and keep it **identical across platforms** (local-docker, brev, slurm, kubernetes).

Common across arches:

- `job_id`: fresh `uuid.uuid4()` — becomes the container name and registry key.
- `image`: resolve per Section 2.
- Secrets (`access_key`, `secret_key`, `HF_TOKEN`, etc.) are read from env vars at runtime — never hard-code, never log or print.

Arch-specific notes (full details in `references/service.yaml` → `container_commands`):

- **`cosmos-rl`** — single `--job '<JOB_JSON>' --docker_env_vars '<ENV_JSON>'` blob; `json.dumps(...)` + `shlex.quote(...)`. `env_payload` carries `TAO_EXECUTION_BACKEND` (per Section 3 table), `TAO_API_JOB_ID`, `CLOUD_BASED=False`. The inference service has no cloud-storage dependency; `HF_TOKEN` is the only cred env var that ever applies (for gated HuggingFace models).
- **`cosmos-predict2.5`** — flag-style `cosmos_predict inference_microservice start ... --port 8080` (no `setup.` prefix; uses `tyro.conf.OmitArgPrefixes`). `--job`/`--docker_env_vars` are **not** accepted. Translate `model_path` to `--checkpoint-path` (local path) or `--model <registered_key>` (`hf_model://`); cloud URIs are rejected. The only cred env var that ever applies is `HF_TOKEN` for gated HuggingFace models. Per-request params (prompt, inference_type, num_output_frames, guidance, seed, num_steps, negative_prompt) go in the request body, not at startup. `TAO_EXECUTION_BACKEND`/`TAO_API_JOB_ID`/`CLOUD_BASED` are unused and may be omitted.

### 4.2 Delegate execution to the platform skill

Read **`skills/platform/<platform>/SKILL.md`** and follow it to start the container.

**Base parameters (all platforms):**

| Parameter | Value |
|-----------|-------|
| `image` | resolved container image (Section 2) |
| `command` | `inner` — the shell string built in Section 4.1 |
| `gpu_count` | `num_gpus` |
| `env_vars` | `env_payload` |
| job / container name | `job_id` — must equal the UUID from 4.1 so the registry can reference it |
| `host_port` *(local-docker, brev)* | host-side port to bind to container port 8080. Default `8080`, but **must be unique per concurrent service** — see the port-allocation rule below. |

**Platform-specific additional inputs:**

| Platform | Additional inputs |
|----------|------------------|
| **local-docker** | None beyond base |
| **brev** | `instance_id` (optional — reuse an existing instance); on multi-credential / multi-workspace accounts also `cloud_cred_id` and `workspace_group_id` for first-create — see `skills/platform/tao-run-on-brev/SKILL.md` |
| **slurm** | `partition` and `account` — check `SLURM_PARTITION`/`SLURM_ACCOUNT` env vars; ask user if unset |
| **kubernetes** | `namespace` (default: `default`); `image_pull_secret` (required for `nvcr.io` images) |

**Port binding (local-docker and brev):** use **direct docker run** (not DockerSDK) so that `-p <host_port>:8080` can be passed and the container name equals `job_id` exactly.

**Port allocation rule (local-docker and brev, REQUIRED for concurrent services):** Before starting a service, read the registry (`/tmp/tao-inf-ms-state.json`) and collect the set of `host_port` values from every existing entry on the same platform (and, for brev, the same `instance_id`). Pick the **lowest free port starting from 8080** that is not in that set — e.g. `host_port = next(p for p in range(8080, 8200) if p not in used_ports)`. The default `8080` only applies when no other service is running. This is what makes "start 3 services, each reachable at a distinct `host_url`" work; without it, services 2 and 3 fail with `bind: address already in use`. SLURM and kubernetes get distinct endpoints from their own platform mechanisms and do not need this step.

### 4.3 After start: service registry and endpoint

Write the service registry immediately after the platform confirms the container is running. The registry (`/tmp/tao-inf-ms-state.json`) is keyed by `job_id`; `"latest"` always points to the most recently started service.

See `references/code-templates.yaml` → `registry_write.<platform>` for the Python template.

| Platform | `host_url` | `platform_job_id` | Extra step before writing |
|----------|-----------|-------------------|--------------------------|
| **local-docker** | `http://localhost:{host_port}` | — | None |
| **brev** | `http://{brev_ip}:{host_port}` | — | `brev ls` → get instance IP (`localhost` is invalid on remote VM) |
| **slurm** | `http://localhost:{host_port}` | SLURM scheduler job ID | Wait until Running; SSH port-forward `localhost:{host_port}→{node}:8080` |
| **kubernetes** | `http://{external_ip}:8080` | k8s job name | `kubectl expose job … --type=LoadBalancer`; wait for external IP |

After writing the registry, print the job_id and URL:

```python
print(f"Inference service started.")
print(f"  Job ID : {job_id}")
print(f"  Arch   : {network_arch}")
print(f"  URL    : {state[job_id]['host_url']}/v1/chat/completions")
print(f"Use this Job ID to send requests or stop the service.")
```

Then poll for readiness — see `references/code-templates.yaml` → `readiness_check`. The container loads the model in the background; do not send requests before it returns 200.

---

## 5. Stopping the inference service

Ask the user for the `job_id` to stop. If they don't provide one, default to `state["latest"]` and confirm which job_id is being stopped. Read the registry using `references/code-templates.yaml` → `stop.registry_read`, then read **`skills/platform/<platform>/SKILL.md`** and use its cancellation / stop mechanism.

| Platform | Identifier to pass | Extra cleanup |
|----------|--------------------|---------------|
| **local-docker** | `job_id_to_stop` — container name | None |
| **brev** | `job_id_to_stop` — container name | None |
| **slurm** | `entry["platform_job_id"]` — SLURM job ID | `pkill -f "ssh.*-L.*{entry['host_port']}"` |
| **kubernetes** | `entry["platform_job_id"]` — k8s job name | `kubectl delete svc {entry["platform_job_id"]} -n <namespace>` |

where `entry = state[job_id_to_stop]`. After stopping, clean up the registry: `references/code-templates.yaml` → `stop.registry_cleanup`.

---

## 6. Sending inference requests

### 6.0 Resolve which service receives this request (REQUIRED)

Each request must be routed to the **specific** service that runs the matching model. Routing happens by `job_id` — the registry stores `network_arch` per entry, so you can resolve a target by arch when the user names a model instead of a `job_id`. Apply these rules in order:

1. **User provided an explicit `job_id`** → use it. Verify it exists in `state`.
2. **User named a `network_arch`** (e.g. "send this to the cosmos-rl service") → look up matching entries: `candidates = [j for j, e in state.items() if j != "latest" and isinstance(e, dict) and e["network_arch"] == arch]`.
   - Exactly one match → use it.
   - Multiple matches → **prompt the user** with the candidate `job_id`s and their `started_at`; do not auto-pick.
   - No match → stop and tell the user no service for that arch is running.
3. **No `job_id` and no `network_arch`** → count non-`"latest"` entries in `state`:
   - Exactly one running service → use it.
   - Two or more → **do not silently default to `state["latest"]`**. Prompt the user with the full list (`job_id`, `network_arch`, `host_url`) and require an explicit choice. The `"latest"` pointer is a convenience for single-service workflows, not a routing fallback when multiple services coexist.
   - Zero → stop and tell the user to start a service first.

After resolving, read the endpoint from the registry (`references/code-templates.yaml` → `request.registry_read`), passing the resolved `job_id` as `user_provided_job_id`. Confirm to the user: "Sending to job_id=… arch=… url=…". If the service may still be loading, poll readiness first (`references/code-templates.yaml` → `readiness_check`).

**Cross-check before sending:** if the user-supplied request body contains arch-specific fields (e.g. `guidance` / `num_steps` / `seed` / `negative_prompt` → cosmos-predict2.5; required `image_url`/`video_url` content items → cosmos-rl), verify they are consistent with `state[job_id]["network_arch"]`. On mismatch, stop and ask — sending a cosmos-predict2.5 body to a cosmos-rl service will fail at the container with a 4xx/5xx that is harder to diagnose than catching it here.

### 6.1 Sampling parameters — REQUIRED user prompt before each request

Before constructing the request body, you **MUST** explicitly prompt the user for the vLLM-style sampling parameters. Do **not** silently apply defaults. Use a structured prompt, one question per field, that:

1. Lists every applicable field with its **type** and **default value**.
2. Lets the user skip / accept any field to take that field's default — entering a value is never required.
3. Collects all fields in one round.

After the prompt, apply each user-entered value verbatim and substitute the default for any skipped field. Do not invent values or silently clamp.

**Field list, defaults, and per-arch applicability:** `references/request.yaml` → `chat_completions_request_body` (base sampling fields: `max_tokens`, `top_p`, `temperature`) and `network_arch_constraints.<network_arch>` (per-arch overrides and extras such as `guidance`/`num_steps`/`seed`/`negative_prompt` for `cosmos-predict2.5`). If a field is marked unsupported for the active arch, do **not** prompt for it and do **not** include it in the body.

### 6.2 Request format

Send a `POST` to `{BASE_URL}/v1/chat/completions` with `Content-Type: application/json` and a timeout of **at least 300 s**. The body is OpenAI-compatible (vLLM chat completions); see `references/request.yaml` → `chat_completions_request_body` for the full field schema and content-item shapes (text / image_url / video_url), and `code_examples` for ready-to-run Python and curl samples.

**Constraints:** only the first user message is processed. No secret values in request bodies. **Per-network constraints** (e.g. cosmos-rl requires every request to include an image or video; cosmos-rl rejects `data:` URIs) are in `references/request.yaml` → `network_arch_constraints`.

### 6.3 Response handling

| HTTP status | Meaning | Action |
|-------------|---------|--------|
| **200** | Success — `choices[0].message.content` has the generated text | Read result |
| **202** | Server still initializing or model still loading | Retry after a delay |
| **503** | Initialization failed, model load failed, **or model not yet ready** | Inspect `error.type`: `model_not_ready` → retry; `initialization_error` / `model_load_error` → give up and check logs |
| **400** | Missing or empty JSON body | Fix request |
| **500** | Unhandled exception during inference | Check container logs |

For 202 and 503, the body contains `{"error": {"type": "<error_type>", "message": "<reason>"}}`. See `container_response_shapes` in `references/request.yaml` for error type strings.

tao-run-inference-service 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/NVIDIA/skills/tree/main/skills/tao-run-inference-service # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 NVIDIA/skills

관련 스킬

klingai-upgrade-migration
업데이트 된 시간 2026년 7월 3일
Verification &amp; Quality Assurance
업데이트 된 시간 2026년 6월 29일
base44-cli
업데이트 된 시간 2026년 6월 29일
Railway CLI Management
업데이트 된 시간 2026년 7월 2일
OR