digital-health-clinical-asr-eval
NVIDIA/skills
선택한 NIM을 기준으로 임상 ASR 매니페스트를 채점하고, 5개 섹션으로 구성된 KER 순위표를 생성한 후, 평가 후 의사결정 트리를 통해 사용자를 안내합니다.
...모든 것을 확장하십시오임상 ASR 플라이휠 — 3단계 (평가)
⚠ 에이전트: 답변하기 전에 아래의 ‘중요 워크플로우 규칙’ 섹션을 반드시읽어보세요. 이 SKILL.md파일은 독립적으로 작동합니다.
evals/,references/,assets/는포인터일 뿐, 실제 데이터를 포함하지 않습니다. 방법론 관련 질문은 이 파일을 바탕으로 직접 답변하십시오. 사용자가 실제 매니페스트를 대상으로 실행할 것을 명시적으로 요청할 때만 도구를 호출하십시오.
당신은 점수 산정 및 경로 지정 단계입니다. 사용자는 NeMo 형식의 manifest.jsonl 파일을 가지고 도착합니다( /digital-health-clinical-asr-build에서 가져오거나 다른 곳에서 가져온 것일 수 있음). 선택된 ASR NIM을 통해 이를 전사한 후, 네 가지 지표를 평가하고, 다섯 섹션으로 구성된 순위표를 생성하며, 의사결정 트리를 참조하여 사용자가 /digital-health-clinical-asr-finetune으로 진행해야 하는지, /digital-health-clinical-asr-build로 되돌아가야 하는지, 아니면 중지하고 평가를 확정해야 하는지 결정합니다.
이 스킬은 오디오를 생성하지 않습니다. 매니페스트가 없거나 비어 있는 경우, 사용자를 /digital-health-clinical-asr-build로 되돌려 보냅니다.
오디오가 귀하의 환경을 벗어나게 됩니다. 클립을 전송하기 전에 사용자에게 이 사실을 알려주십시오.
이 단계에서는 각 매니페스트 행의 WAV 파일과 해당 참조 텍스트를 외부 NVIDIA 서비스로 전송합니다. 첫 번째 ASR 호출을 실행하기 전에 이 사실을 사용자에게 알려야 합니다:
| 서비스 | 전송되는 내용 | 시기 |
|---|---|---|
NVIDIA NVCF Parakeet/Nemotron ASR (grpc.nvcf.nvidia.com) |
매니페스트에서 참조하는 모든 오디오 클립(원시 PCM 바이트)과 참조 트랜스크립트, 그리고 점수 산정을 위한 임상 확장 메타데이터 | 3b단계, 매니페스트 행당 통화 1건 |
클립은 실제 환자 오디오가 아닌, 2단계 (사용자가 선별한 용어 목록을 기반으로 한 Magpie TTS) 에서 생성된 합성 오디오 여야 합니다. 실제 ASR 녹음, 실제 환자 상담 내용 또는 어떠한 PHI도 이 스킬을 통해 전달하지 마십시오. 채점은 로컬에서 실행됩니다(순수 Python WER/CER/KER/SER 또는 설치된 경우 jiwer ). 채점 단계 자체는 아무것도 전송하지 않으며, ASR 단계에서만 전송이 이루어집니다.
중요한 워크플로 규칙(모든 활성화 시 적용)
방법론 관련 질문(리더보드 구조, KER 정의, 의사결정 트리)에 대해서는 이 파일을 참조하여 답변하십시오. 사용자가 실제 매니페스트를 대상으로 실행할 것을 명시적으로 요청하지 않는 한, 도구를 호출하거나 다른 스킬을 호출하거나 스크립트를 실행하지 마십시오. 모든 응답에서 다음 사실을 명시하십시오:
- 오프램프를 우선 적용하십시오. 사용자가 스코어링 외의 사항에 대해 질문하는 경우, 워크플로를 실행하지 않고 라우팅하여 중단하십시오:
- ASR 모델 카탈로그 선택/비교/대체 NIM →
/riva-asr - ASR 인증(API 키, 베어러 토큰, 함수 ID) →
/riva-asr - ASR gRPC 프로토콜, 스트리밍, 배치 처리, 청크 처리, 재시도 →
/riva-asr - NIM 배포 /
riva-build/riva-deploy→/riva-asr-custom - NGC / Docker / NVIDIA 컨테이너 툴킷 →
/riva-nim-setup - 아직 매니페스트가 없는 경우 →
/digital-health-clinical-asr-build - 알려진 KER을 사용하여 지금 미세 조정하고자 함 →
/digital-health-clinical-asr-finetune
- ASR 모델 카탈로그 선택/비교/대체 NIM →
- 기본 ASR NIM은
nvidia/parakeet-tdt-0.6b-v2입니다(NVCF 함수 IDd3fe9151-442b-4204-a70d-5fcc597fd610, 오프라인 gRPC). 환경 변수 재정의:ASR_MODEL_NAME(리더보드 표시 이름),ASR_NVCF_FUNCTION_ID(다른 호스팅된 NIM으로 교체 — 예: Parakeet 백엔드에 오류가 발생한 경우 Whisper Large v3b702f636-…또는 미세 조정된 NIM으로 전환),ASR_ENDPOINT(자체 호스팅 gRPC; 우선 순위 적용). API 크레딧을 소비하기 전에 선택된 NIM과 확인된 function-id를 반환합니다. - ASR 전사 처리는 3b 단계에 내장되어 있습니다 (NVCF gRPC +
riva.client.ASRService.offline_recognize, 1단계와 동일한 인증 패턴). 프로토콜/인증에 대한 더 자세한 질문, 대체 NIM 카탈로그 또는 자체 호스팅 Riva NIM 구성에 대해서는/riva-asr을 참조하십시오. - KER가 핵심 지표입니다. 행별 검사: 플래그가 지정된
용어단어들은 정규화된 가설 내에서 순서대로, 연속적으로, 인접하게 나타나야 합니다.cefazolin → cefa zolin은오류입니다. 집계된 WER은 임상적으로 위험한 오류를 숨길 수 있으므로, 두 가지 모두 보고되며 KER이 필터 역할을 합니다. ipa_source별분할 값은 리더보드에서가장 유익한 단일 수치입니다.merriam-webster와magpie_g2p간의 차이는 SSML 오버라이드 파이프라인이 실제로 효과를 발휘하고 있음을 증명합니다. 사용자에게 이를 소리 내어 읽어주세요.- 특수 사례 라우팅.
merriam-webster행은 양호,magpie_g2p행은 불량 → 발음 범위 격차이지, 모델 격차가 아님./digital-health-clinical-asr-build단계 2d로 되돌아가십시오. 첫 번째 대응으로/digital-health-clinical-asr-finetune을권장하지 마십시오. - 5개 섹션으로 구성된 리더보드 순서. 헤드라인 (WER/CER/KER/SER) →
엔티티 카테고리별 KER →IPA 소스별 KER →노이즈 레벨별 KER → 용어별 KER (최악 순).IPA 소스별섹션은 필수이며, 이는 SSML 파이프라인이 제대로 작동한다는 증거입니다.
목적
임상 ASR 매니페스트를 평가하고, 5개 섹션으로 구성된 KER 리더보드를 생성하며, 평가 후 의사결정 트리를 통해 사용자를 라우팅합니다. 방법론 세부 사항(지표 정의, 정규화, 리더보드 순서, 특수 사례 라우팅)은 위의 ‘중요 워크플로 규칙’과 아래의 ‘지침’에 명시되어 있습니다.
이 스킬을 사용할 때
다음과 같은 사용자 문구에 대해 활성화합니다:
- "내 ASR 매니페스트 점수 매기기"
- "Parakeet TDT v2의 KER은 얼마인가요?"
- "사이클 N에 대해 평가를 실행해 주세요"
- "임상 벤치마크에서 두 ASR 모델을 비교해 줘"
- "리더보드 생성해 줘"
- "manifest.jsonl 파일이 있는데, 점수를 어떻게 매기나요?"
- "WER이 0.07인데 KER이 0.4인 이유는 무엇인가요?"
- "파인 튜닝을 해야 할까요?" (이는 평가 단계의 질문입니다. 평가 후 의사 결정 트리는 이 스킬에 포함되어 있습니다)
리터럴 키워드 비활성화 확인 — 사용자의 메시지에 authenticate, API key, bearer, function ID, gRPC, streaming, chunking, batching, transcription retry, riva-build, riva-deploy, NIM deploy, NGC, Docker, Container Toolkit 중 하나라도 포함되어 있거나, “어떤 ASR 모델이 가장 좋은가요” / “모델 비교” / “벤더별 차이점”과 같은 질문을 할 경우 — 스코어링 워크플로를 활성화하지 마십시오. 위의 중요 워크플로 규칙 #1을 적용하여 올바른 형제 스킬로 라우팅하고 중단하십시오. 이는 사용자가 키워드와 함께 "KER" 또는 "eval"을 언급한 경우에도 적용됩니다.
필수 조건
- 임상 확장 필드(
term,entity_category,ipa_source,voice_id,noise_level,context_type)가 포함된NeMo 형식의 매니페스트. 스키마는 빌드 스킬의references/manifest-schema.md에문서화되어 있습니다. NVIDIA_API_KEY가내보내져 있어야 합니다(1단계 전제 조건은 여전히 적용됨).nvidia-riva-client+soundfile이설치되어 있어야 합니다(1단계 필수 조건). 자체 호스팅 Riva NIM에 대한 자세한 내용은/riva-asr옵션 B를 참조하십시오.- 디스크에 오디오 파일이 실제로 존재해야 합니다. API 크레딧을 사용하기 전에 manifest-schema 참조 문서의 audio-existence 사전 점검을 실행하십시오.
지침
3a. ASR NIM 선택
기본값: NVCF gRPC(오프라인)를 통한 nvidia/parakeet-tdt-0.6b-v2, function-id d3fe9151-442b-4204-a70d-5fcc597fd610. NVIDIA의 현재 영어 ASR 권장 사항 — 카탈로그에서 가장 빠르고 저렴하며, NeMo의 기본 SFT 레시피에서 지원되므로 Stage 3 베이스라인과 Stage 4 미세 조정이 동일한 모델 패밀리를 사용합니다.
세 가지 런타임 환경 변수 재정의 옵션(리더보드 표시용ASR_MODEL_NAME, 다른 호스팅된 NIM으로 전환하기 위한 ASR_NVCF_FUNCTION_ID, 자체 호스팅 gRPC용 ASR_ENDPOINT )와 전체 대체 NIM 카탈로그(Parakeet TDT 1.1B, Parakeet CTC 1.1B, Whisper Large v3, Nemotron 스트리밍)이 제공되며, 각 항목의 함수 ID와 호출 형태에 대한 설명은 references/offline-asr-recipe.md에서 확인할 수 있습니다.
API 크레딧을 소모하기 전에 선택한 NIM, 확인된 함수 ID 및 환경 변수 재정의 사항을 사용자에게 알려주십시오. 호스팅된 Parakeet TDT v2의 200행 매니페스트는 비용이 저렴하지만, 1,000행 매니페스트에서 잘못된 모델로 실수로 실행하는 것은 그렇지 않습니다.
3b. 텍스트 변환
manifest.jsonl의 각 행에 대해 audio_filepath를 전사하고 per_sample.json을 작성합니다(행당 하나의 JSON 객체, JSONL 또는 JSON 배열 — 호출자가 선택):
{
"audio_filepath": "...",
"ref": "",
"hyp": "",
"term": "",
"entity_category": "",
"ipa_source": "",
"voice_id": "",
"noise_level": "",
"context_type": ""
}
레시피 (전체 Python 코드는 references/offline-asr-recipe.md에 있음): transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US")는 NVCF(또는 자체 호스팅 Riva의 경우 ASR_ENDPOINT로 설정된 경우)에 대한 오프라인 gRPC 스트림을 열고, 각 행마다 riva.client.ASRService.offline_recognize를 호출하며 — 임상 매니페스트의 문장은 30초 이하이므로 스트리밍이나 배치 처리가 필요하지 않음 — 위의 JSONL을 출력합니다. Stage 1 설정 스모크 테스트와 동일한 auth_for 형식을 사용합니다. 에이전트 하네스는 api_key를 명시적으로 전달하며, 레시피는 상단에서 세 가지 환경 변수 재정의(ASR_NVCF_FUNCTION_ID, ASR_MODEL_NAME, ASR_ENDPOINT)를 읽어오므로 감사 담당자가 모든 제어 항목을 한 곳에서 확인할 수 있습니다.
Whisper 대체 방식 (Parakeet의 NVCF 백엔드가 Triton으로부터 CUDA 불법 메모리 액세스 오류를 반환할 때) 및 자체 호스팅 Riva NIM (ASR_ENDPOINT=localhost:50051) 환경 변수 패턴: references/offline-asr-recipe.md의 §Whisper fallback, §Self-hosted Riva NIM 절을 참조하십시오.
복원력 관련 설정은 사용자에게 위임됩니다. NVCF가 배치 도중 RESOURCE_ENHAUSTED를 반환하면 루프가 해당 행에서 중단되며, 오류가 발생한 행부터 다시 실행됩니다. 스트리밍/배치 처리/백오프를 동반한 재시도는 본 문서의 범위를 벗어납니다 — /riva-asr 참조.
3c. 네 가지 지표 계산
각 행에 대해 다음을 계산합니다:
| 지표 | 측정 대상 | 이 지표를 유지하는 이유 |
|---|---|---|
| WER | 단어 오류율 (정규화 후 토큰에 대한 레벤슈타인 거리) | 업계 표준; 임상적 분석에는 다소 거친 도구 |
| CER | 문자 오류율 | 긴 복합명에서 아슬아슬하게 오류를 간신히 포착함 |
| KER ★ | 키워드 오류율 — 표시된 용어가 가설에 등장했는가(정규화, 연속 일치)? |
주요 임상 신호 |
| SER | 문장 오류율 (오류가 있으면 1, 완벽하면 0) | 합리성 한계; 의사가 경험하는 바 |
정규화 (네 가지 지표 모두를 계산하기 전에 참조 문장과 가설 문장 모두에 적용):
- 소문자로 변환.
- NFKD 정규화(스마트 따옴표 → ASCII 등으로 변환).
- 하이픈을 제외한 구두점 제거.
- 연속된 공백을 단일 공백으로 축소.
인라인 스코어링 레시피 — normalize / edit_distance / wer / cer / ker / ser (순수 Python, jiwer 의존성 없음): references/scoring-recipes.md 참조. 각 지표에 대해 (행별 점수의) 평균을 구하여 행별로 집계합니다.
엄격한 KER — 용어 단어들은 정규화된 가설에서 순서대로, 인접하게 나타나야 합니다. 이는 보수적인 방식입니다: cefazolin → cefa zolin은 오류로 간주됩니다. 이는 임상적으로 올바른 판단입니다 — 후속 약국 조회에서 철자가 틀린 토큰으로 인해 실패할 것이기 때문입니다.
KER는 주변의 오류를 감점하지 않습니다. 용어는 정확하지만 문장의 나머지 부분이 무의미한 내용인 행도 여전히 KER=0으로 점수가 매겨집니다. 해당 행의 WER은 더 광범위한 문제를 별도로 드러낼 것입니다.
3d. 세부 분석 + 리더보드
다음 순서대로 5개 섹션으로 구성된 마크다운 리더보드를 작성하십시오:
- 헤드라인 — 선택한 모델의 전체 WER, CER, KER, SER.
-
엔티티카테고리별 KER — 약물 vs 시술 vs 해부학 vs ... 이는 사용자가 실제 배포 시 가장 중요하게 여기는 부분입니다. -
ipa_source별KER — 리더보드에서 가장 유익한 단일 수치입니다.merriam-webster행과magpie_g2p행 간의 차이는 SSML 오버라이드 파이프라인이 실제로 효과를 발휘하고 있음을 증명합니다. 이 섹션을 사용자에게 소리 내어 읽어주세요. -
noise_level별KER — 임상 환경은 잡음이 많습니다.snr_5db행은‘clean’행보다 실제 상황에 더 가깝습니다. - 용어별 KER (최악 순) — 이것이 바로 4단계 미세 조정 대상입니다.
메리엄-웹스터 대 magpie_g2p 차이 해석이 포함된 대표적인 ipa_source 분할: references/scoring-recipes.md §Representative ipa_source split. 델타는 배포 상황을 알려줍니다. 사용자가 큰 격차를 확인하고 “미세 조정을 해야 할까요?”라고 묻는다면, 아직은 아닙니다. 사용자를 /digital-health-clinical-asr-build의 IPA QA 파이프라인(2d 단계)으로 다시 안내하십시오. 아래의 의사 결정 트리를 참조하십시오.
의사결정 트리 (평가 후)
우선순위 범주별 KER (대부분의 임상 워크플로에서는 약물 KER, 수술 워크플로에서는 시술 KER)을 확인하고 다음과 같이 안내합니다:
| 우선순위 범주별 KER | 권장 |
|---|---|
| > 0.3 | /digital-health-clinical-asr-finetune. 매니페스트는 이미 NeMo 형식을 지원합니다. 참고: 신뢰할 수 있는 미세 조정 신호를 얻으려면 행 수가 100개 이상이어야 합니다. 매니페스트가 이보다 작다면, 먼저 /digital-health-clinical-asr-build를 통해 규모를 늘리십시오. |
| 0.1 – 0.3 | 용어 목록을 확장 하거나 (새로운 도메인 용어를 추가하여 /digital-health-clinical-asr-build로 되돌아가면 — 일반적으로 튜닝보다 적은 비용으로 더 많은 오류를 발견할 수 있음) 파인튠을 수행하십시오. 첫 번째 평가에서는 용어 목록을 확장하고, 매니페스트를 이미 확장한 상태인 후속 평가에서는 튜닝을 수행하십시오. |
| < 0.1 | 강력한 베이스라인입니다. 아직 튜닝하지 마십시오 — 이미 포화 상태에 이른 지표를 대상으로 최적화를 시도하게 될 것입니다. 평가를 더욱 강화하십시오: 다양한 목소리, 소음 수준, 맥락, 적대적 용어를 추가하십시오. /digital-health-clinical-asr-build로 돌아가십시오. |
특수 사례 — merriam-webster 행은 점수가 좋지만 magpie_g2p 행은 나쁩니다. 이는 모델의 한계가 아니라 발음 힌트 커버리지의 부족 때문입니다. /digital-health-clinical-asr-finetune으로 가지 말고, /digital-health-clinical-asr-build의 2d 단계(IPA QA 검토)로 되돌아가세요. TTS 발음 격차를 바탕으로 파인 튜닝을 하면 모델이 자신의 실수를 잘못 인식하도록 학습하게 됩니다 — 이는 잘못된 해결책입니다.
예시
시나리오 A — 새로운 사이클 1 매니페스트에 대한 첫 번째 평가. 사용자: " term 및 entity_category 필드가 포함된 200개의 임상 오디오 행이 이미 있는 manifest.jsonl 파일이 있습니다. 어떻게 점수를 매기나요?" → 2단계를 완전히 건너뜁니다. audio-existence 사전 점검을 실행합니다. parakeet-tdt-0.6b-v2 (기본값) 를 선택하고, 선택 사항과 해결된 function-id를 반환합니다. 인라인으로 포함된 3b단계 레시피(transcribe_manifest(...))를 실행합니다. 네 가지 지표를 평가합니다. 다섯 섹션으로 구성된 리더보드를 생성합니다. 사용자에게ipa_source별 분할 결과를 읽어줍니다. 의약품 KER에 대해 의사결정 트리를 적용합니다.
시나리오 B — 혼합된 결과 해석. 사용자: "Eval 결과, merriam-webster 태그가 붙은 행에서는 KER이 0.05로 나오지만, magpie_g2p 태그가 붙은 행에서는 0.40입니다. 파인 튜닝을 해야 할까요?" → 아니요 — 이는 특수한 경우입니다. 모델은 정상이며, 발음 힌트가 롱테일 용어들을 포괄하지 못하고 있을 뿐입니다. 사용자를 /digital-health-clinical-asr-build 2d 단계로 되돌려 magpie_g2p 행을 검토하고, 검증된 IPA를 pronunciation_overrides.csv에 추가하도록 안내하십시오. 재구축 후 3단계를 다시 실행한 다음 4단계를 재검토하십시오.
생성된 아티팩트
per_sample.json— 모든 임상 확장 필드가 보존된 행별 전사 결과(ASRhyp가매니페스트의ref및 메타데이터와 연결됨)results.csv— 행별 WER/CER/KER/SER 점수leaderboard_cycle— 5개 섹션으로 구성된 마크다운 보고서.md
(파일 이름은 사용자가 직접 지정할 수 있으며, 위의 이름은 이 스킬의 나머지 부분에서 따르는 관례입니다.)
문제 해결
- "매니페스트를 찾을 수 없음" → 사용자가 2단계를 건너뛰었습니다.
/digital-health-clinical-asr-build로이동하거나$MANIFEST_PATH를확인하십시오. - 모든 행의 KER=1 →
ref와hyp간의 정규화 불일치. 양쪽 모두에 4단계 정규화 절차를 적용하십시오. - 모든 행의 KER=0이지만 WER이 높은 경우 → 매니페스트 정렬 오류(오디오 행 불일치)일 가능성이 높습니다. 몇 쌍
(참조, 가설)을 수동으로 무작위 확인하십시오. merriam-webster값이 낮고magpie_g2p값이 높음 → 발음 범위 차이./digital-health-clinical-asr-build단계 2d로 이동하십시오. 미세 조정은 하지 마십시오 — 모델에 문제는 없습니다.-
merriam-webster와magpie_g2p모두 높음 → 실제 모델 격차. 4단계가 올바른 경로입니다(매니페스트 ≥ 100행). 정제된행은 양호하나,snr_5db가급증 → 견고성 격차;/digital-health-clinical-asr-build를통해 노이즈 다양성 확대.- Riva-NIM과 오프라인 NeMo 결과가 상이함 → Riva 전처리 /
riva-build플래그./riva-asr-custom으로진행. - 대용량 매니페스트에서
RESOURCE_EXHAUSTED발생 → 30초 후 재시도; 누락된 행을 분할하여 재실행. 내장 백오프:/riva-asr. Auth.__init__()에서 'ssl_cert' 오류 발생/ Parakeet 함수 ID에서 CUDA 불법 메모리 액세스:references/offline-asr-recipe.md참조 (ssl_root_cert 이름 변경 + §Whisper 폴백).
기타 사항: 업스트림 소유자 확인. ASR 프로토콜 / NIM 배포 → /riva-asr. 점수 산정 → 여기.
제한 사항
- 기본적으로 영어 전용입니다. 토큰화 및 정규화는 라틴 문자 및 en-US 어휘집을 가정합니다.
- 엄격한 연속 KER은 보수적입니다.
cefa zolin과같은 근접 오류는 미일치로 간주됩니다. 이는 의도된 사항입니다 — 약국 검색은 근접 오류 시 실패합니다. “유연한” 매칭을 원하는 사용자는 음소 수준 편집 거리로 전환할 수 있으며, 이는 구성 조정 rather than 방법론의 확장입니다. - 평가 실행당 하나의 모델만 사용 가능합니다. 두 모델을 비교하려면 평가를 두 번 실행한 후, 두 개의
leaderboard_cycle파일을 비교해야 합니다(또는 레시피를 확장하여 다중 모델 행을 직접 작성할 수도 있습니다)..md - 호스팅 전용 경로를 가정합니다. 자체 호스팅 NIM도 작동하지만, 먼저
/riva-nim-setup을실행해야 합니다.
다음 단계
- 전방 모델(KER > 0.3, 매니페스트 ≥ 100행):
/digital-health-clinical-asr-finetune. - 빌드로 돌아가기(첫 번째 평가에서 KER 0.1–0.3, 또는
magpie_g2p격차):/digital-health-clinical-asr-build. - 중단(KER < 0.1): 평가가 포화 상태에 도달했습니다. 성공을 선언하기 전에 모델을 안정화하십시오.
- ASR 프로토콜 / 인증 / 스트리밍 / 자체 호스팅 NIM에 대한 자세한 내용은
/riva-asr을참조하십시오.
참고 문헌
references/offline-asr-recipe.md— 전체 3b단계 Python 레시피(transcribe_manifest,resolve_asr_config,build_asr_auth), 호출 형태 설명이 포함된 함수 ID 카탈로그, Whisper 폴백, 자체 호스팅 Riva NIM 설정references/scoring-recipes.md— 표준 4단계 정규화를 적용한 순수 파이썬 WER/CER/KER/SER 스코어링 함수
---
name: digital-health-clinical-asr-eval
description: Score a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
license: Apache-2.0
---
<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->
# Clinical ASR Flywheel — Stage 3 (Eval)
> **⚠ Agent: read the Critical Workflow Rules section below before answering.** This SKILL.md is self-contained — `evals/`, `references/`, and `assets/` are pointers, not load-bearing. Answer methodology questions from this file directly; only invoke tools when the user explicitly asks to execute against a real manifest.
You are the **score-and-route** stage. The user arrives with a NeMo-format `manifest.jsonl` (either from `/digital-health-clinical-asr-build` or carried in from elsewhere). You transcribe it via the chosen ASR NIM, score four metrics, produce a five-section leaderboard, and read the decision tree to decide whether the user should advance to `/digital-health-clinical-asr-finetune`, loop back to `/digital-health-clinical-asr-build`, or stop and harden the eval.
**This skill does not generate audio.** If the manifest is missing or empty, send the user back to `/digital-health-clinical-asr-build`.
## Audio leaves your environment — disclose this to the user before any clip is sent
This stage transmits each manifest row's WAV file plus its reference text to an external NVIDIA service. Surface this before invoking the first ASR call:
| Service | What gets sent | When |
|---|---|---|
| **NVIDIA NVCF Parakeet/Nemotron ASR** (`grpc.nvcf.nvidia.com`) | Every audio clip referenced by the manifest (raw PCM bytes), plus the reference transcript and the clinical-extension metadata for scoring | Step 3b, one call per manifest row |
The clips should be **synthetic audio generated by Stage 2** (Magpie TTS over a user-curated term list) — not real patient audio. **Do not pass real ASR recordings, real patient encounters, or any PHI through this skill.** Scoring then runs locally (pure-Python WER/CER/KER/SER, or `jiwer` if installed). The scoring step itself does not transmit anything; only the ASR step does.
## Critical workflow rules (apply on every activation)
For methodology questions (leaderboard structure, KER definition, decision tree), answer from this file. Don't invoke tools, call other skills, or run scripts unless the user explicitly asks to execute against a real manifest. Surface these facts in any response:
1. **Off-ramp first.** If the user is asking about something outside scoring, route and stop without running any workflow:
- ASR model-catalog selection / comparison / alternative NIMs → `/riva-asr`
- ASR auth (API keys, bearer tokens, function IDs) → `/riva-asr`
- ASR gRPC protocol, streaming, batching, chunking, retries → `/riva-asr`
- NIM deploy / `riva-build` / `riva-deploy` → `/riva-asr-custom`
- NGC / Docker / NVIDIA Container Toolkit → `/riva-nim-setup`
- No manifest yet → `/digital-health-clinical-asr-build`
- Wants to fine-tune now with a known KER → `/digital-health-clinical-asr-finetune`
2. **Default ASR NIM is `nvidia/parakeet-tdt-0.6b-v2`** (NVCF function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`, offline gRPC). Env-var overrides: `ASR_MODEL_NAME` (leaderboard display name), `ASR_NVCF_FUNCTION_ID` (swap to a different hosted NIM — e.g. Whisper Large v3 `b702f636-…` while the Parakeet backend is faulting, or a fine-tuned NIM), `ASR_ENDPOINT` (self-hosted gRPC; takes precedence). Echo the chosen NIM **and the resolved function-id** back before spending API credits.
3. **ASR transcription is inlined in Step 3b** (NVCF gRPC + `riva.client.ASRService.offline_recognize`, same auth pattern as Stage 1). For deeper protocol/auth questions, alternative NIM catalogs, or self-hosted Riva NIM configuration, defer to `/riva-asr`.
4. **KER is the headline.** Per-row check: the flagged `term` words must appear *in order, contiguous, adjacent* in the normalized hypothesis. `cefazolin → cefa zolin` is a miss. Aggregate WER hides clinically dangerous failures; both are reported, KER is the gate.
5. **The by-`ipa_source` split is the most informative single number** in the leaderboard. The `merriam-webster` vs `magpie_g2p` delta proves the SSML override pipeline is doing real work. Read it aloud to the user.
6. **Special-case routing.** `merriam-webster` rows good, `magpie_g2p` rows bad → pronunciation-coverage gap, **not** a model gap. Route back to `/digital-health-clinical-asr-build` Step 2d. **Do NOT recommend `/digital-health-clinical-asr-finetune`** as a first response.
7. **Five-section leaderboard order.** Headline (WER/CER/KER/SER) → KER by `entity_category` → KER by `ipa_source` → KER by `noise_level` → Per-term KER worst-first. The by-`ipa_source` section is mandatory; it is the proof the SSML pipeline works.
## Purpose
Score a clinical-ASR manifest, produce a five-section KER leaderboard, and route the user via the post-eval decision tree. Methodology details (metric definitions, normalization, leaderboard order, special-case routing) live in Critical Workflow Rules above and Instructions below.
## When to use this skill
Activate on user phrases like:
- "Score my ASR manifest"
- "What's the KER on Parakeet TDT v2?"
- "Run the eval on cycle-N"
- "Compare two ASR models on the clinical benchmark"
- "Generate the leaderboard"
- "I have a manifest.jsonl, how do I score it?"
- "Why is KER 0.4 when WER is 0.07?"
- "Should we fine-tune?" *(this is the eval-side question — the post-eval decision tree lives in this skill)*
**Literal-keyword non-activation check** — if the user's message contains any of `authenticate`, `API key`, `bearer`, `function ID`, `gRPC`, `streaming`, `chunking`, `batching`, `transcription retry`, `riva-build`, `riva-deploy`, `NIM deploy`, `NGC`, `Docker`, `Container Toolkit`, or asks "which ASR model is best" / "compare models" / "vendor differences" — **do NOT activate** the scoring workflow. Apply Critical Workflow Rule #1 above to route to the right sibling skill and stop. This applies even if the user mentions "KER" or "eval" alongside the keyword.
## Prerequisites
- **A NeMo-format manifest** with the clinical extension fields (`term`, `entity_category`, `ipa_source`, `voice_id`, `noise_level`, `context_type`). The schema is documented in the build skill's `references/manifest-schema.md`.
- **`NVIDIA_API_KEY`** exported (Stage 1 prerequisite still applies).
- **`nvidia-riva-client` + `soundfile`** installed (Stage 1 prerequisite). For self-hosted Riva NIM details, see `/riva-asr` Option B.
- **Audio files actually present on disk** — run the audio-existence pre-flight from the manifest-schema reference before spending API credits.
## Instructions
### 3a. Pick the ASR NIM
**Default**: `nvidia/parakeet-tdt-0.6b-v2` via NVCF gRPC (offline), function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`. NVIDIA's current English ASR recommendation — fastest/cheapest in the catalog, and supported in NeMo's stock SFT recipe so the Stage 3 baseline and a Stage 4 fine-tune ride the same model family.
Three runtime env-var override knobs (`ASR_MODEL_NAME` for leaderboard display, `ASR_NVCF_FUNCTION_ID` to swap to a different hosted NIM, `ASR_ENDPOINT` for self-hosted gRPC) plus the full alternate-NIM catalog (Parakeet TDT 1.1B, Parakeet CTC 1.1B, Whisper Large v3, Nemotron streaming) with function IDs and call-shape notes: `references/offline-asr-recipe.md`.
Echo the chosen NIM, the resolved function-id, and any env-var overrides to the user **before** spending API credits. A 200-row manifest on hosted Parakeet TDT v2 is cheap; an accidental run against the wrong model on a 1,000-row manifest is not.
### 3b. Transcribe
For each row in `manifest.jsonl`, transcribe `audio_filepath` and write `per_sample.json` (one JSON object per row, JSONL or a JSON array — caller's choice):
```json
{
"audio_filepath": "...",
"ref": "<row.text>",
"hyp": "<asr output>",
"term": "<row.term>",
"entity_category": "<row.entity_category>",
"ipa_source": "<row.ipa_source>",
"voice_id": "<row.voice_id>",
"noise_level": "<row.noise_level>",
"context_type": "<row.context_type>"
}
```
**Recipe** (full Python in `references/offline-asr-recipe.md`): `transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US")` opens an offline gRPC stream to NVCF (or to `ASR_ENDPOINT` if set for self-hosted Riva), calls `riva.client.ASRService.offline_recognize` per row — sentences in a clinical manifest are ≤ 30 s so no streaming/batching needed — and writes the JSONL above. Same `auth_for` shape as the Stage 1 setup smoke test. The agent harness passes `api_key` explicitly; the recipe reads the three env-var overrides (`ASR_NVCF_FUNCTION_ID`, `ASR_MODEL_NAME`, `ASR_ENDPOINT`) at the top so auditors see the knobs in one place.
**Whisper fallback** (when Parakeet's NVCF backend faults with `CUDA illegal-memory-access` from Triton) and **self-hosted Riva NIM** (`ASR_ENDPOINT=localhost:50051`) env-var patterns: see `references/offline-asr-recipe.md` (§Whisper fallback, §Self-hosted Riva NIM).
**Resilience knobs deferred to the user.** If NVCF returns `RESOURCE_EXHAUSTED` mid-batch, the loop raises on that row; re-run from the failing row. Streaming/batching/retry-with-backoff are out of scope — see `/riva-asr`.
### 3c. Score four metrics
For every row, compute:
| Metric | What it measures | Why we keep it |
|---|---|---|
| **WER** | Word error rate (Levenshtein on tokens, after normalization) | Industry standard; blunt instrument for clinical |
| **CER** | Character error rate | Catches near-misses on long compound names |
| **KER** ★ | Keyword error rate — did the flagged `term` appear in the hypothesis (normalized, **contiguous** match)? | **Headline clinical signal** |
| **SER** | Sentence error rate (1 if any wrong, 0 if perfect) | Sanity bound; what the doctor experiences |
**Normalization (apply to both `ref` and `hyp` before all four metrics):**
1. Lowercase.
2. NFKD-normalize (smart quotes → ASCII, etc.).
3. Strip punctuation **except hyphen**.
4. Collapse whitespace runs to a single space.
**Inline scoring recipes** — `normalize` / `edit_distance` / `wer` / `cer` / `ker` / `ser` (pure-Python, no `jiwer` dependency): see `references/scoring-recipes.md`. Aggregate across rows by taking `mean(per-row score)` for each metric.
**Strict KER** — term words must appear *in order, adjacent* in the normalized hypothesis. This is conservative: `cefazolin → cefa zolin` counts as a miss. That's the right call clinically — a downstream pharmacy lookup will fail on the misspelled token.
KER does **not** punish surrounding errors. A row where the term is correct and the rest of the sentence is garbage still scores KER=0; the WER on that row will surface the broader problem separately.
### 3d. Breakdowns + leaderboard
Write a five-section markdown leaderboard, **in this order**:
1. **Headline** — overall WER, CER, KER, SER for the chosen model.
2. **KER by `entity_category`** — drug vs procedure vs anatomy vs ... This is what the user actually cares about for deployment.
3. **KER by `ipa_source`** — **the most informative single number in the leaderboard.** The delta between `merriam-webster` and `magpie_g2p` rows is the proof the SSML override pipeline is doing real work. *Read this section aloud to the user.*
4. **KER by `noise_level`** — clinical environments are loud. `snr_5db` rows are closer to reality than `clean`.
5. **Per-term KER** (worst first) — these are your Stage 4 fine-tune targets.
A representative `ipa_source` split with the merriam-webster vs magpie_g2p delta interpretation: `references/scoring-recipes.md` §Representative ipa_source split. The delta tells the deployment story — if the user sees a wide gap and asks "should we fine-tune?", the answer is *not yet*; route them back to `/digital-health-clinical-asr-build`'s IPA QA pipeline (Stage 2d). See the decision tree below.
## Decision tree (after eval)
Read the **priority-category KER** (drug KER for most clinical workflows, procedure KER for surgical workflows) and route:
| KER on priority category | Recommend |
|---|---|
| **> 0.3** | `/digital-health-clinical-asr-finetune`. Manifest is already NeMo-format-ready. Note: rows ≥ 100 is the minimum for a believable fine-tune signal; if the manifest is smaller, grow it first via `/digital-health-clinical-asr-build`. |
| **0.1 – 0.3** | Either expand the term list (back to `/digital-health-clinical-asr-build` with new domain terms — usually surfaces more failures cheaper than tuning) **or** fine-tune. On a *first* eval, expand. On a *later* eval where you've already grown the manifest, tune. |
| **< 0.1** | Strong baseline. Don't tune yet — you'd be optimizing against a saturated metric. Push the eval harder: add voices, noise levels, contexts, adversarial terms. Loop back to `/digital-health-clinical-asr-build`. |
**Special case — `merriam-webster` rows score well but `magpie_g2p` rows are bad.** That's a pronunciation-hint coverage gap, **not a model gap**. Route back to `/digital-health-clinical-asr-build` Step 2d (IPA QA review), not to `/digital-health-clinical-asr-finetune`. Fine-tuning over a TTS-pronunciation gap teaches the model to mis-recognize the model's own mistakes — the wrong fix.
## Examples
**Scenario A — first eval on a fresh cycle-1 manifest.** User: *"I have `manifest.jsonl` with 200 clinical audio rows already, with `term` and `entity_category` fields. How do I score it?"* → Skip Stage 2 entirely. Run the audio-existence pre-flight. Pick `parakeet-tdt-0.6b-v2` (default) and echo the choice + resolved function-id. Run the inlined Step 3b recipe (`transcribe_manifest(...)`). Score the four metrics. Produce the five-section leaderboard. Read the by-`ipa_source` split to the user. Apply the decision tree against drug KER.
**Scenario B — interpreting a mixed result.** User: *"Eval shows KER 0.05 on rows tagged `merriam-webster` but 0.40 on rows tagged `magpie_g2p`. Should I fine-tune?"* → No — this is the special case. The model is fine; the pronunciation hints aren't covering the long-tail terms. Route the user back to `/digital-health-clinical-asr-build` Step 2d to audition the `magpie_g2p` rows and append verified IPA to `pronunciation_overrides.csv`. Re-run Stage 3 after the rebuild before reconsidering Stage 4.
## Artifacts produced
- `per_sample.json` — per-row transcription results with all clinical-extension fields preserved (the ASR `hyp` joined to the manifest's `ref` and metadata)
- `results.csv` — per-row WER/CER/KER/SER scores
- `leaderboard_cycle<N>.md` — five-section markdown report
(File names are user-chosen; the names above are conventions the rest of this skill assumes.)
## Troubleshooting
- **"No manifest found"** → user skipped Stage 2. Route to `/digital-health-clinical-asr-build` or confirm `$MANIFEST_PATH`.
- **All rows KER=1** → normalization mismatch between `ref` and `hyp`. Apply the four normalization steps to both sides.
- **All rows KER=0 but WER high** → likely misaligned manifest (audio row mismatch). Spot-check a few `(ref, hyp)` pairs by hand.
- **`merriam-webster` low, `magpie_g2p` high** → pronunciation-coverage gap. Route to `/digital-health-clinical-asr-build` Step 2d. **Don't fine-tune** — model isn't the problem.
- **Both `merriam-webster` and `magpie_g2p` high** → real model gap. Stage 4 is the right route (manifest ≥ 100 rows).
- **`clean` rows fine, `snr_5db` balloons** → robustness gap; expand noise diversity via `/digital-health-clinical-asr-build`.
- **Riva-NIM and offline NeMo results diverge** → Riva preprocessing / `riva-build` flags. Route to `/riva-asr-custom`.
- **`RESOURCE_EXHAUSTED` on large manifests** → retry after 30 s; slice + re-run dropped rows. Built-in backoff: `/riva-asr`.
- **`Auth.__init__() got 'ssl_cert'`** / **CUDA illegal-memory-access on Parakeet function ID**: see `references/offline-asr-recipe.md` (ssl_root_cert rename + §Whisper fallback).
Anything else: identify the upstream owner. ASR protocol / NIM deploy → `/riva-asr`. Scoring → here.
## Limitations
- **English-only by default.** Tokenization + normalization assume Latin script and en-US lexicon.
- **Strict-contiguous KER is conservative.** A near-miss like `cefa zolin` counts as a miss. That's intentional — pharmacy lookups fail on near-misses. Users wanting "soft" matching can switch to phoneme-level edit distance, which is a methodology extension, not a config tweak.
- **One model per eval run.** Comparing two models means running the eval twice and diffing the two `leaderboard_cycle<N>.md` files (or extending the recipe to write multi-model rows yourself).
- **Hosted-only paths assumed.** Self-hosted NIMs work but require `/riva-nim-setup` first.
## Next steps
- **Forward (KER > 0.3, manifest ≥ 100 rows):** `/digital-health-clinical-asr-finetune`.
- **Back to build (KER 0.1–0.3 on first eval, or `magpie_g2p` gap):** `/digital-health-clinical-asr-build`.
- **Stop (KER < 0.1):** the eval is saturated. Harden it before declaring victory.
- **Lateral** for ASR protocol / auth / streaming / self-hosted NIM details: `/riva-asr`.
## References
- [`references/offline-asr-recipe.md`](references/offline-asr-recipe.md) — full Step 3b Python recipe (`transcribe_manifest`, `resolve_asr_config`, `build_asr_auth`), function-ID catalog with call-shape notes, Whisper fallback, self-hosted Riva NIM setup
- [`references/scoring-recipes.md`](references/scoring-recipes.md) — pure-Python WER/CER/KER/SER scoring functions with the canonical 4-step normalization
digital-health-clinical-asr-eval 설치
스킬 파일을 다운로드한 후 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-eval # Copy SKILL.md to your .claude/skills/ directory
복사





집
