選項
首頁首頁 Skill 數據科學與機器學習 digital-health-clinical-asr-eval

digital-health-clinical-asr-eval

NVIDIA/skills NVIDIA/skills

針對選定的 NIM,對臨床 ASR 清單進行評分,產生包含五個區段的 KER 排行榜,並透過評估後的決策樹引導使用者。

...展開全部
1
更新時間 2026-09-28

臨床 ASR 飛輪 — 第 3 階段(評估)

⚠ 代理:回答前請先閱讀下方的「關鍵工作流程規則」章節。此 SKILL.md 檔案為自包含檔案 —evals/、references/ 及assets/僅為參考路徑,並非承載核心功能。請直接根據此檔案回答方法論相關問題;僅當使用者明確要求針對真實 manifest 執行時,才調用工具。

您負責「評分與路由」階段。使用者會攜帶一份 NeMo 格式的manifest.jsonl檔案(來源可能是/digital-health-clinical-asr-build,或從其他地方導入)。 您需透過選定的 ASR NIM 將其轉錄,對四項指標進行評分,生成五個區段的排行榜,並根據決策樹判斷使用者應繼續前往/digital-health-clinical-asr-finetune、迴圈回到/digital-health-clinical-asr-build,還是停止並鞏固評估結果。

此技能不會產生音訊。若清單檔案遺失或為空,請將使用者導回/digital-health-clinical-asr-build。

音訊將離開您的環境 — 在傳送任何片段之前,請向使用者說明此情況

此階段會將每個清單列對應的 WAV 檔案及其參考文字傳輸至外部 NVIDIA 服務。請在呼叫第一個 ASR 請求之前向使用者說明此事項:

服務 傳送內容 何時
NVIDIA NVCF Parakeet/Nemotron ASR(grpc.nvcf.nvidia.com) 清單中引用的每個音訊片段(原始 PCM 位元組),加上參考文字稿以及用於評分的臨床擴展元資料 步驟 3b,每份清單行對應一次呼叫

這些片段應為第二階段(基於使用者編纂術語表的 Magpie TTS)所生成的合成音訊——而非真實病患音訊。請勿透過此技能傳遞真實的 ASR 錄音、真實的患者診療紀錄,或任何受保護的健康資訊(PHI)。評分隨後在本地執行(純 Python 版的 WER/CER/KER/SER,或若已安裝則使用jiwer)。評分步驟本身不會傳輸任何資料;僅 ASR 步驟會進行傳輸。

關鍵工作流程規則(適用於每次觸發)

若涉及方法論問題(排行榜結構、KER 定義、決策樹),請從此檔案中提供答案。除非使用者明確要求針對真實清單執行,否則請勿調用工具、呼叫其他技能或執行腳本。請在任何回應中明確說明以下事項:

  1. 優先採用「Off-ramp」機制。若使用者詢問的內容超出評分範圍,請進行路由並停止執行,且勿執行任何工作流程:
    • ASR 模型目錄的選取/比較/替代 NIM →/riva-asr
    • ASR 授權(API 金鑰、承載者憑證、函式 ID)→/riva-asr
    • ASR gRPC 協定、串流、批次處理、分塊、重試 →/riva-asr
    • NIM 部署 /riva-build/riva-deploy→/riva-asr-custom
    • NGC / Docker / NVIDIA 容器工具包 →/riva-nim-setup
    • 尚未建立清單 →/digital-health-clinical-asr-build
    • 希望立即使用已知的 KER 進行微調 →/digital-health-clinical-asr-finetune
  2. 預設的 ASR NIM 為nvidia/parakeet-tdt-0.6b-v2(NVCF 函式 IDd3fe9151-442b-4204-a70d-5fcc597fd610,離線 gRPC)。 環境變數覆寫:ASR_MODEL_NAME(排行榜顯示名稱)、ASR_NVCF_FUNCTION_ID(切換至不同的託管 NIM — 例如 當 Parakeet 後端發生故障時切換至 Whisper Large v3b702f636-…,或切換至微調過的 NIM),ASR_ENDPOINT(自託管 gRPC;優先級最高)。在消耗 API 額度之前,先回傳所選的 NIM及解析出的 function-id。
  3. ASR 轉錄功能已內嵌於步驟 3b 中(NVCF gRPC +riva.client.ASRService.offline_recognize,與第一階段採用相同的認證模式)。 若有關於協定/認證的深入問題、替代 NIM 目錄,或自託管 Riva NIM 配置,請參閱/riva-asr。
  4. KER 是關鍵指標。每行檢查:標記的術語詞必須在正規化假設中依序出現、連續且相鄰。 cefazolin → cefa zolin即為錯誤。彙總 WER 會掩蓋臨床上危險的失敗;兩者均會通報,KER 才是門檻。
  5. 按 ipa_source劃分的分數是排行榜中最具參考價值的單一數值。Merriam-Webster與magpie_g2p之間的差異證明,SSML 覆寫管道確實發揮了實質作用。請將其大聲讀給使用者聽。
  6. 特殊情況路由。 merriam-webster行表現良好,magpie_g2p行表現不佳 → 這是發音覆蓋率差距,而非模型差距。請回溯至/digital-health-clinical-asr-build第 2d 步。切勿將/digital-health-clinical-asr-finetune作為首選解決方案。
  7. 五部分排行榜排序。標題 (WER/CER/KER/SER) → 按entity_category 分類的KER → 按ipa_source 來源分類的KER → 按noise_level 噪音等級分類的KER → 按單詞 KER 表現最差者優先。按 ipa_source 來源分類的部分為必備項目;這是 SSML 處理管線運作有效的證明。

目的

對臨床 ASR 清單進行評分,生成五部分組成的 KER 排行榜,並透過評估後決策樹引導使用者。方法論詳情(指標定義、正規化、排行榜排序、特殊情況路由)請參閱上方的「關鍵工作流程規則」及下方的「操作說明」。

何時使用此技能

當使用者說出以下類型的短語時觸發:

  • 「為我的 ASR 清單評分」
  • 「Parakeet TDT v2 的 KER 是多少?」
  • 「在第 N 循環執行評估」
  • 「在臨床基準測試上比較兩個 ASR 模型」
  • 「產生排行榜」
  • 「我有一個 manifest.jsonl 檔案,該如何對其進行評分?」
  • 「當 WER 為 0.07 時,為什麼 KER 會是 0.4?」
  • 「我們應該進行微調嗎?」(這是評估階段的問題——評估後的決策樹位於此技能中)

字面關鍵字未觸發檢查— 若使用者訊息包含以下任一詞彙:authenticate、API key、bearer、function ID、gRPC、streaming、chunking、batching、transcription retry、riva-build、riva-deploy、NIM deploy、NGC、Docker、Container Toolkit,或詢問「哪個 ASR 模型最好」/「比較模型」/「供應商差異」——請勿啟動評分工作流程。 請套用上述「關鍵工作流程規則 #1」,將請求路由至正確的同級技能並停止處理。即使使用者在關鍵字旁提及「KER」或「eval」,此規則仍適用。

先決條件

  • 一份包含臨床擴展欄位(term、entity_category、ipa_source、voice_id、noise_level、context_type)的 NeMo 格式清單。該架構說明詳見建置技能的references/manifest-schema.md 檔案。
  • 已匯出NVIDIA_API_KEY(第一階段的先決條件仍然適用)。
  • 已安裝nvidia-riva-client+soundfile(第一階段先決條件)。有關自架設 Riva NIM 的詳細資訊,請參閱/riva-asr方案 B。
  • 音訊檔案必須實際存在於磁碟上— 在消耗 API 額度之前,請先執行 manifest-schema 參考文件中的 audio-existence 預檢程序。

操作說明

3a. 選擇 ASR NIM

預設:透過 NVCF gRPC(離線)使用的nvidia/parakeet-tdt-0.6b-v2,函式識別碼d3fe9151-442b-4204-a70d-5fcc597fd610。 NVIDIA 目前針對英語的 ASR 建議選項 — 這是目錄中最快且最經濟的選項,並獲 NeMo 預設 SFT 配方支援,因此第 3 階段的基準模型與第 4 階段的微調模型皆採用同一模型家族。

三個執行時環境變數覆寫控制項(ASR_MODEL_NAME用於排行榜顯示、ASR_NVCF_FUNCTION_ID用於切換至不同的託管 NIM、ASR_ENDPOINT用於自託管 gRPC),外加完整的替代 NIM 目錄(Parakeet TDT 1.1B、 Parakeet CTC 1.1B、Whisper Large v3、Nemotron 串流)皆附有函式 ID 及呼叫結構說明:請參閱references/offline-asr-recipe.md。

在消耗 API 額度之前,應向使用者回報所選的 NIM、已解析的功能識別碼,以及任何環境變數覆寫設定。在託管的 Parakeet TDT v2 上執行 200 行的清單成本低廉;但若在 1,000 行的清單上誤用錯誤的模型進行執行,則代價高昂。

3b. 轉錄

針對manifest.jsonl 中的每一行,將audio_filepath進行轉錄,並寫入per_sample.json(每行一個 JSON 物件,格式可選 JSONL 或 JSON 陣列——由呼叫方決定):

{
  "audio_filepath": "...",
  "ref": "",
  "hyp": "",
  "term": "",
  "entity_category": "",
  "ipa_source": "",
  "voice_id": "",
  "noise_level": "",
  "context_type": ""
}

實作範例(完整的 Python 程式碼請參閱references/offline-asr-recipe.md):transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US")會開啟通往 NVCF 的離線 gRPC 串流(若設定為自架設的 Riva,則通往ASR_ENDPOINT),針對每一行呼叫riva.client.ASRService.offline_recognize函式 — 由於臨床清單中的句子長度均不超過 30 秒,因此無需進行串流或批次處理 — 並寫入上述 JSONL 格式資料。其 auth_for結構與第一階段設定的煙霧測試相同。 代理程式框架會明確傳遞api_key;該配方會在最上方讀取三項環境變數覆寫設定(ASR_NVCF_FUNCTION_ID、ASR_MODEL_NAME、ASR_ENDPOINT),以便審計人員能在一處查看所有可調整參數。

Whisper 備用方案(當 Parakeet 的 NVCF 後端因 Triton 引發的CUDA 非法記憶體存取錯誤而失效時)以及自託管 Riva NIM(ASR_ENDPOINT=localhost:50051)的環境變數設定模式: 請參閱references/offline-asr-recipe.md(§Whisper 備用方案、§自託管 Riva NIM)。

韌性調整參數交由使用者自行處理。若 NVCF 在批次處理過程中回傳RESOURCE_ENHAUSTED錯誤,迴圈將在該行拋出異常;並從發生錯誤的那一行重新執行。串流處理/批次處理/帶退避的重試機制不在本範疇內 — 請參閱/riva-asr。

3c. 計算四項指標

針對每一行,計算:

指標 衡量項目 保留此指標的原因
WER 單字錯誤率(對令牌進行萊文斯坦距離計算,經正規化後) 業界標準;用於臨床分析的粗糙工具
CER 字元錯誤率 可偵測長複合名稱中的「近誤」
KER★ 關鍵字錯誤率 — 被標記的術語是否出現在假設中(經正規化處理且連續匹配)? 主要臨床訊號
SER 句子錯誤率(若有錯誤則為 1,完美則為 0) 合理性上限;醫師的實際體驗

正規化(在計算所有四項指標前,同時對參考句與假設句進行處理):

  1. 轉為小寫。
  2. 進行 NFKD 標準化(智能引號 → ASCII 等)。
  3. 移除除連字號以外的所有標點符號。
  4. 將連續的空白符號壓縮為單一空格。

內嵌評分配方—normalize/edit_distance/wer/cer/ker/ser(純 Python 實現,不依賴jiwer):請參閱references/scoring-recipes.md。針對每項指標,透過計算各列評分的平均值(mean(per-row score))來進行跨列彙總。

嚴格的 KER— 術語單字必須按順序出現,且在正規化假設中緊鄰。此為保守做法:cefazolin → cefa zolin會被視為錯配。從臨床角度來看,這是正確的判斷 — 下游藥房查詢會因拼寫錯誤的詞元而失敗。

KER不會因周邊錯誤而扣分。若某行中的術語正確,但句子其餘部分全是垃圾內容,該行的 KER 仍為 0;該行的 WER 將另行揭示更廣泛的問題。

3d. 細項分析 + 排行榜

撰寫一份包含五個區段的 Markdown 排行榜,順序如下:

  1. 標題— 所選模型的整體 WER、CER、KER、SER。
  2. 按實體類別劃分的KER— 藥物 vs 醫療程序 vs 解剖學 vs ... 這正是使用者在部署時真正關心的重點。
  3. 按ipa_source分類的 KER—排行榜中最具參考價值的單一數值。 Merriam-Webster與magpie_g2p行之間的差異,正是 SSML 覆寫處理流程確實發揮作用的證明。請將此部分內容大聲朗讀給使用者聽。
  4. 依noise_level分類的 KER— 臨床環境的干擾聲很強。snr_5db行比「乾淨」的數據更貼近現實。
  5. 各術語的 KER(由最差者起)——這些就是您的第 4 階段微調目標。

一個具代表性的ipa_source分割數據,附有對 Merriam-Webster 與 magpie_g2p 差異的解讀:references/scoring-recipes.md§Representative ipa_source split。 差異值揭示了部署狀況 — 若使用者發現顯著差距並詢問「是否該進行微調?」,答案是「尚未需要」;請將其導回/digital-health-clinical-asr-build 的 IPA 品質保證管道(第 2d 階段)。請參閱下方的決策樹。

決策樹(評估後)

讀取優先級類別的KER(多數臨床工作流程為藥物 KER,外科工作流程為手術程序 KER),並進行路由:

優先類別的 KER 建議
> 0.3 /digital-health-clinical-asr-finetune。清單檔已符合 NeMo 格式規範。注意:行數 ≥ 100 是獲得可信微調訊號的最低門檻;若清單檔規模較小,請先透過/digital-health-clinical-asr-build 擴增其規模。
0.1 – 0.3 可選擇擴充術語清單(返回/digital-health-clinical-asr-build並加入新領域術語——通常比微調更能以較低成本揭露更多錯誤),或進行微調。首次評估時,請先擴充術語清單;若後續評估時已擴充過清單,則進行微調。
< 0.1 強勁的基準表現。暫勿進行微調——這等同於針對已達飽和的指標進行優化。應進一步強化評估:增加發音、噪音水準、語境及對抗性詞彙。回頭至/digital-health-clinical-asr-build。

特殊情況 —merriam-webster資料列的評分良好,但magpie_g2p資料列的評分很差。這是發音提示的覆蓋缺口,而非模型缺口。 請回溯至/digital-health-clinical-asr-build第 2d 步(國際音標 QA 審查),而非前往/digital-health-clinical-asr-finetune。針對 TTS 發音差距進行微調,會教導模型錯誤識別其自身的錯誤——這是一種錯誤的修正方式。

範例

情境 A — 首次在全新的週期 1 清單上進行評估。使用者:「我手邊有個manifest.jsonl檔案,裡面已有 200 筆臨床音訊資料,並包含term和entity_category欄位。該如何進行評分?」→ 完全跳過第 2 階段。執行音訊存在性預檢。 選擇parakeet-tdt-0.6b-v2(預設值),並回報該選擇及已解析的 function-id。執行內嵌的第 3b 步驟配方(transcribe_manifest(...))。 針對四項指標進行評分。產生五個區段的排行榜。向使用者讀取按 ipa_source 劃分的分組結果。針對藥物 KER 套用決策樹。

情境 B — 解讀混合結果。使用者:「評估顯示,標記為merriam-webster的資料列 KER 為 0.05,但標記為magpie_g2p 的資料列 KER 為 0.40。我該進行微調嗎?」→ 不需要 —— 這是特殊情況。模型表現良好;發音提示未能涵蓋長尾詞彙。 將使用者導回/digital-health-clinical-asr-build第 2d 步驟,以檢視magpie_g2p資料列,並將驗證過的 IPA 附加至pronunciation_overrides.csv。在重建完成後重新執行第 3 階段,再重新評估第 4 階段。

產出的成果

  • per_sample.json— 每行轉錄結果,完整保留所有臨床擴展欄位(ASR假名與清單的ref及元資料已關聯)
  • results.csv— 每行 WER/CER/KER/SER 評分
  • leaderboard_cycle.md— 五個部分組成的 Markdown 報告

(檔案名稱由使用者自行選擇;上述名稱僅為本技能後續操作所假設的慣例。)

疑難排解

  • 「未找到清單」→ 使用者跳過了第 2 階段。請導向/digital-health-clinical-asr-build或確認$MANIFEST_PATH。
  • 所有列的 KER=1→參考音軌與假設音軌之間的正規化不匹配。請對兩側均應用四個正規化步驟。
  • 所有資料列 KER=0 但 WER 偏高→ 可能是清單對齊錯誤(音訊資料列不匹配)。請手動抽查幾組(參考, 假設)配對。
  • merriam-webster值低、magpie_g2p值高→ 發音涵蓋範圍存在差距。請參閱/digital-health-clinical-asr-build第 2d 階段。切勿進行微調— 問題不在於模型。
  • Merriam-Webster與magpie_g2p皆高→ 真實模型差距。應走第 4 階段(對照表 ≥ 100 行)。
  • 乾淨的資料列沒問題,snr_5db值暴增→ 穩健性差距;透過/digital-health-clinical-asr-build 擴展雜訊多樣性。
  • Riva-NIM 與離線 NeMo 的結果出現分歧→ Riva 預處理 /riva-build參數。請改用/riva-asr-custom。
  • 大型清單出現RESOURCE_EXHAUSTED錯誤→ 30 秒後重試;切片並重新執行被捨棄的資料列。內建退避機制:/riva-asr。
  • Auth.__init__() 收到 'ssl_cert'/Parakeet 函式 ID 出現 CUDA 非法記憶體存取錯誤:參見references/offline-asr-recipe.md(ssl_root_cert 重命名 + §Whisper 備用方案)。

其他情況:識別上游所有者。ASR 協定 / NIM 部署 →/riva-asr。評分 → 此處。

限制

  • 預設僅支援英語。分詞 + 規範化假設使用拉丁字母及 en-US 詞彙表。
  • 嚴格連續的 KER 採用保守策略。類似「cefa zolin」這類近似錯配將被視為錯配。此設計屬刻意為之——藥品查詢系統在近似錯配時會失敗。若使用者希望採用「寬鬆」的匹配方式,可切換至音素層級的編輯距離,此為方法論的延伸,而非配置調整。
  • 每次評估執行僅限一個模型。若要比較兩個模型,需執行兩次評估,並比對兩個leaderboard_cycle.md檔案(或自行擴展配方以寫入多模型資料列)。
  • 預設僅支援託管模式。自行託管的 NIM 雖可運作,但需先執行/riva-nim-setup。

下一步

  • 前向(KER > 0.3,清單 ≥ 100 行): /digital-health-clinical-asr-finetune。
  • 返回建置階段(首次評估時 KER 介於 0.1–0.3 之間,或存在magpie_g2p差距): /digital-health-clinical-asr-build。
  • 停止(KER < 0.1):評估已達飽和狀態。在宣告成功前請先進行模型強化。
  • 關於ASR 協定/授權/串流/自架設 NIM 的相關資訊:/riva-asr。

參考文獻

  • references/offline-asr-recipe.md— 完整的第 3b 步驟 Python 實作範例(transcribe_manifest、resolve_asr_config、build_asr_auth)、附帶呼叫結構註解的函式 ID 目錄、Whisper 備用方案,以及自架設 Riva NIM 的設定
  • references/scoring-recipes.md— 採用標準 4 步驟正規化的純 Python WER/CER/KER/SER 評分函式
在 GitHub 上查看
---
name: digital-health-clinical-asr-eval
description: Score a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
license: Apache-2.0
---

<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->

# Clinical ASR Flywheel — Stage 3 (Eval)

> **⚠ Agent: read the Critical Workflow Rules section below before answering.** This SKILL.md is self-contained — `evals/`, `references/`, and `assets/` are pointers, not load-bearing. Answer methodology questions from this file directly; only invoke tools when the user explicitly asks to execute against a real manifest.

You are the **score-and-route** stage. The user arrives with a NeMo-format `manifest.jsonl` (either from `/digital-health-clinical-asr-build` or carried in from elsewhere). You transcribe it via the chosen ASR NIM, score four metrics, produce a five-section leaderboard, and read the decision tree to decide whether the user should advance to `/digital-health-clinical-asr-finetune`, loop back to `/digital-health-clinical-asr-build`, or stop and harden the eval.

**This skill does not generate audio.** If the manifest is missing or empty, send the user back to `/digital-health-clinical-asr-build`.

## Audio leaves your environment — disclose this to the user before any clip is sent

This stage transmits each manifest row's WAV file plus its reference text to an external NVIDIA service. Surface this before invoking the first ASR call:

| Service | What gets sent | When |
|---|---|---|
| **NVIDIA NVCF Parakeet/Nemotron ASR** (`grpc.nvcf.nvidia.com`) | Every audio clip referenced by the manifest (raw PCM bytes), plus the reference transcript and the clinical-extension metadata for scoring | Step 3b, one call per manifest row |

The clips should be **synthetic audio generated by Stage 2** (Magpie TTS over a user-curated term list) — not real patient audio. **Do not pass real ASR recordings, real patient encounters, or any PHI through this skill.** Scoring then runs locally (pure-Python WER/CER/KER/SER, or `jiwer` if installed). The scoring step itself does not transmit anything; only the ASR step does.

## Critical workflow rules (apply on every activation)

For methodology questions (leaderboard structure, KER definition, decision tree), answer from this file. Don't invoke tools, call other skills, or run scripts unless the user explicitly asks to execute against a real manifest. Surface these facts in any response:

1. **Off-ramp first.** If the user is asking about something outside scoring, route and stop without running any workflow:
   - ASR model-catalog selection / comparison / alternative NIMs → `/riva-asr`
   - ASR auth (API keys, bearer tokens, function IDs) → `/riva-asr`
   - ASR gRPC protocol, streaming, batching, chunking, retries → `/riva-asr`
   - NIM deploy / `riva-build` / `riva-deploy` → `/riva-asr-custom`
   - NGC / Docker / NVIDIA Container Toolkit → `/riva-nim-setup`
   - No manifest yet → `/digital-health-clinical-asr-build`
   - Wants to fine-tune now with a known KER → `/digital-health-clinical-asr-finetune`
2. **Default ASR NIM is `nvidia/parakeet-tdt-0.6b-v2`** (NVCF function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`, offline gRPC). Env-var overrides: `ASR_MODEL_NAME` (leaderboard display name), `ASR_NVCF_FUNCTION_ID` (swap to a different hosted NIM — e.g. Whisper Large v3 `b702f636-…` while the Parakeet backend is faulting, or a fine-tuned NIM), `ASR_ENDPOINT` (self-hosted gRPC; takes precedence). Echo the chosen NIM **and the resolved function-id** back before spending API credits.
3. **ASR transcription is inlined in Step 3b** (NVCF gRPC + `riva.client.ASRService.offline_recognize`, same auth pattern as Stage 1). For deeper protocol/auth questions, alternative NIM catalogs, or self-hosted Riva NIM configuration, defer to `/riva-asr`.
4. **KER is the headline.** Per-row check: the flagged `term` words must appear *in order, contiguous, adjacent* in the normalized hypothesis. `cefazolin → cefa zolin` is a miss. Aggregate WER hides clinically dangerous failures; both are reported, KER is the gate.
5. **The by-`ipa_source` split is the most informative single number** in the leaderboard. The `merriam-webster` vs `magpie_g2p` delta proves the SSML override pipeline is doing real work. Read it aloud to the user.
6. **Special-case routing.** `merriam-webster` rows good, `magpie_g2p` rows bad → pronunciation-coverage gap, **not** a model gap. Route back to `/digital-health-clinical-asr-build` Step 2d. **Do NOT recommend `/digital-health-clinical-asr-finetune`** as a first response.
7. **Five-section leaderboard order.** Headline (WER/CER/KER/SER) → KER by `entity_category` → KER by `ipa_source` → KER by `noise_level` → Per-term KER worst-first. The by-`ipa_source` section is mandatory; it is the proof the SSML pipeline works.

## Purpose

Score a clinical-ASR manifest, produce a five-section KER leaderboard, and route the user via the post-eval decision tree. Methodology details (metric definitions, normalization, leaderboard order, special-case routing) live in Critical Workflow Rules above and Instructions below.

## When to use this skill

Activate on user phrases like:

- "Score my ASR manifest"
- "What's the KER on Parakeet TDT v2?"
- "Run the eval on cycle-N"
- "Compare two ASR models on the clinical benchmark"
- "Generate the leaderboard"
- "I have a manifest.jsonl, how do I score it?"
- "Why is KER 0.4 when WER is 0.07?"
- "Should we fine-tune?" *(this is the eval-side question — the post-eval decision tree lives in this skill)*

**Literal-keyword non-activation check** — if the user's message contains any of `authenticate`, `API key`, `bearer`, `function ID`, `gRPC`, `streaming`, `chunking`, `batching`, `transcription retry`, `riva-build`, `riva-deploy`, `NIM deploy`, `NGC`, `Docker`, `Container Toolkit`, or asks "which ASR model is best" / "compare models" / "vendor differences" — **do NOT activate** the scoring workflow. Apply Critical Workflow Rule #1 above to route to the right sibling skill and stop. This applies even if the user mentions "KER" or "eval" alongside the keyword.

## Prerequisites

- **A NeMo-format manifest** with the clinical extension fields (`term`, `entity_category`, `ipa_source`, `voice_id`, `noise_level`, `context_type`). The schema is documented in the build skill's `references/manifest-schema.md`.
- **`NVIDIA_API_KEY`** exported (Stage 1 prerequisite still applies).
- **`nvidia-riva-client` + `soundfile`** installed (Stage 1 prerequisite). For self-hosted Riva NIM details, see `/riva-asr` Option B.
- **Audio files actually present on disk** — run the audio-existence pre-flight from the manifest-schema reference before spending API credits.

## Instructions

### 3a. Pick the ASR NIM

**Default**: `nvidia/parakeet-tdt-0.6b-v2` via NVCF gRPC (offline), function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`. NVIDIA's current English ASR recommendation — fastest/cheapest in the catalog, and supported in NeMo's stock SFT recipe so the Stage 3 baseline and a Stage 4 fine-tune ride the same model family.

Three runtime env-var override knobs (`ASR_MODEL_NAME` for leaderboard display, `ASR_NVCF_FUNCTION_ID` to swap to a different hosted NIM, `ASR_ENDPOINT` for self-hosted gRPC) plus the full alternate-NIM catalog (Parakeet TDT 1.1B, Parakeet CTC 1.1B, Whisper Large v3, Nemotron streaming) with function IDs and call-shape notes: `references/offline-asr-recipe.md`.

Echo the chosen NIM, the resolved function-id, and any env-var overrides to the user **before** spending API credits. A 200-row manifest on hosted Parakeet TDT v2 is cheap; an accidental run against the wrong model on a 1,000-row manifest is not.

### 3b. Transcribe

For each row in `manifest.jsonl`, transcribe `audio_filepath` and write `per_sample.json` (one JSON object per row, JSONL or a JSON array — caller's choice):

```json
{
  "audio_filepath": "...",
  "ref": "<row.text>",
  "hyp": "<asr output>",
  "term": "<row.term>",
  "entity_category": "<row.entity_category>",
  "ipa_source": "<row.ipa_source>",
  "voice_id": "<row.voice_id>",
  "noise_level": "<row.noise_level>",
  "context_type": "<row.context_type>"
}
```

**Recipe** (full Python in `references/offline-asr-recipe.md`): `transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US")` opens an offline gRPC stream to NVCF (or to `ASR_ENDPOINT` if set for self-hosted Riva), calls `riva.client.ASRService.offline_recognize` per row — sentences in a clinical manifest are ≤ 30 s so no streaming/batching needed — and writes the JSONL above. Same `auth_for` shape as the Stage 1 setup smoke test. The agent harness passes `api_key` explicitly; the recipe reads the three env-var overrides (`ASR_NVCF_FUNCTION_ID`, `ASR_MODEL_NAME`, `ASR_ENDPOINT`) at the top so auditors see the knobs in one place.

**Whisper fallback** (when Parakeet's NVCF backend faults with `CUDA illegal-memory-access` from Triton) and **self-hosted Riva NIM** (`ASR_ENDPOINT=localhost:50051`) env-var patterns: see `references/offline-asr-recipe.md` (§Whisper fallback, §Self-hosted Riva NIM).

**Resilience knobs deferred to the user.** If NVCF returns `RESOURCE_EXHAUSTED` mid-batch, the loop raises on that row; re-run from the failing row. Streaming/batching/retry-with-backoff are out of scope — see `/riva-asr`.

### 3c. Score four metrics

For every row, compute:

| Metric | What it measures | Why we keep it |
|---|---|---|
| **WER** | Word error rate (Levenshtein on tokens, after normalization) | Industry standard; blunt instrument for clinical |
| **CER** | Character error rate | Catches near-misses on long compound names |
| **KER** ★ | Keyword error rate — did the flagged `term` appear in the hypothesis (normalized, **contiguous** match)? | **Headline clinical signal** |
| **SER** | Sentence error rate (1 if any wrong, 0 if perfect) | Sanity bound; what the doctor experiences |

**Normalization (apply to both `ref` and `hyp` before all four metrics):**

1. Lowercase.
2. NFKD-normalize (smart quotes → ASCII, etc.).
3. Strip punctuation **except hyphen**.
4. Collapse whitespace runs to a single space.

**Inline scoring recipes** — `normalize` / `edit_distance` / `wer` / `cer` / `ker` / `ser` (pure-Python, no `jiwer` dependency): see `references/scoring-recipes.md`. Aggregate across rows by taking `mean(per-row score)` for each metric.

**Strict KER** — term words must appear *in order, adjacent* in the normalized hypothesis. This is conservative: `cefazolin → cefa zolin` counts as a miss. That's the right call clinically — a downstream pharmacy lookup will fail on the misspelled token.

KER does **not** punish surrounding errors. A row where the term is correct and the rest of the sentence is garbage still scores KER=0; the WER on that row will surface the broader problem separately.

### 3d. Breakdowns + leaderboard

Write a five-section markdown leaderboard, **in this order**:

1. **Headline** — overall WER, CER, KER, SER for the chosen model.
2. **KER by `entity_category`** — drug vs procedure vs anatomy vs ... This is what the user actually cares about for deployment.
3. **KER by `ipa_source`** — **the most informative single number in the leaderboard.** The delta between `merriam-webster` and `magpie_g2p` rows is the proof the SSML override pipeline is doing real work. *Read this section aloud to the user.*
4. **KER by `noise_level`** — clinical environments are loud. `snr_5db` rows are closer to reality than `clean`.
5. **Per-term KER** (worst first) — these are your Stage 4 fine-tune targets.

A representative `ipa_source` split with the merriam-webster vs magpie_g2p delta interpretation: `references/scoring-recipes.md` §Representative ipa_source split. The delta tells the deployment story — if the user sees a wide gap and asks "should we fine-tune?", the answer is *not yet*; route them back to `/digital-health-clinical-asr-build`'s IPA QA pipeline (Stage 2d). See the decision tree below.

## Decision tree (after eval)

Read the **priority-category KER** (drug KER for most clinical workflows, procedure KER for surgical workflows) and route:

| KER on priority category | Recommend |
|---|---|
| **> 0.3** | `/digital-health-clinical-asr-finetune`. Manifest is already NeMo-format-ready. Note: rows ≥ 100 is the minimum for a believable fine-tune signal; if the manifest is smaller, grow it first via `/digital-health-clinical-asr-build`. |
| **0.1 – 0.3** | Either expand the term list (back to `/digital-health-clinical-asr-build` with new domain terms — usually surfaces more failures cheaper than tuning) **or** fine-tune. On a *first* eval, expand. On a *later* eval where you've already grown the manifest, tune. |
| **< 0.1** | Strong baseline. Don't tune yet — you'd be optimizing against a saturated metric. Push the eval harder: add voices, noise levels, contexts, adversarial terms. Loop back to `/digital-health-clinical-asr-build`. |

**Special case — `merriam-webster` rows score well but `magpie_g2p` rows are bad.** That's a pronunciation-hint coverage gap, **not a model gap**. Route back to `/digital-health-clinical-asr-build` Step 2d (IPA QA review), not to `/digital-health-clinical-asr-finetune`. Fine-tuning over a TTS-pronunciation gap teaches the model to mis-recognize the model's own mistakes — the wrong fix.

## Examples

**Scenario A — first eval on a fresh cycle-1 manifest.** User: *"I have `manifest.jsonl` with 200 clinical audio rows already, with `term` and `entity_category` fields. How do I score it?"* → Skip Stage 2 entirely. Run the audio-existence pre-flight. Pick `parakeet-tdt-0.6b-v2` (default) and echo the choice + resolved function-id. Run the inlined Step 3b recipe (`transcribe_manifest(...)`). Score the four metrics. Produce the five-section leaderboard. Read the by-`ipa_source` split to the user. Apply the decision tree against drug KER.

**Scenario B — interpreting a mixed result.** User: *"Eval shows KER 0.05 on rows tagged `merriam-webster` but 0.40 on rows tagged `magpie_g2p`. Should I fine-tune?"* → No — this is the special case. The model is fine; the pronunciation hints aren't covering the long-tail terms. Route the user back to `/digital-health-clinical-asr-build` Step 2d to audition the `magpie_g2p` rows and append verified IPA to `pronunciation_overrides.csv`. Re-run Stage 3 after the rebuild before reconsidering Stage 4.

## Artifacts produced

- `per_sample.json` — per-row transcription results with all clinical-extension fields preserved (the ASR `hyp` joined to the manifest's `ref` and metadata)
- `results.csv` — per-row WER/CER/KER/SER scores
- `leaderboard_cycle<N>.md` — five-section markdown report

(File names are user-chosen; the names above are conventions the rest of this skill assumes.)

## Troubleshooting

- **"No manifest found"** → user skipped Stage 2. Route to `/digital-health-clinical-asr-build` or confirm `$MANIFEST_PATH`.
- **All rows KER=1** → normalization mismatch between `ref` and `hyp`. Apply the four normalization steps to both sides.
- **All rows KER=0 but WER high** → likely misaligned manifest (audio row mismatch). Spot-check a few `(ref, hyp)` pairs by hand.
- **`merriam-webster` low, `magpie_g2p` high** → pronunciation-coverage gap. Route to `/digital-health-clinical-asr-build` Step 2d. **Don't fine-tune** — model isn't the problem.
- **Both `merriam-webster` and `magpie_g2p` high** → real model gap. Stage 4 is the right route (manifest ≥ 100 rows).
- **`clean` rows fine, `snr_5db` balloons** → robustness gap; expand noise diversity via `/digital-health-clinical-asr-build`.
- **Riva-NIM and offline NeMo results diverge** → Riva preprocessing / `riva-build` flags. Route to `/riva-asr-custom`.
- **`RESOURCE_EXHAUSTED` on large manifests** → retry after 30 s; slice + re-run dropped rows. Built-in backoff: `/riva-asr`.
- **`Auth.__init__() got 'ssl_cert'`** / **CUDA illegal-memory-access on Parakeet function ID**: see `references/offline-asr-recipe.md` (ssl_root_cert rename + §Whisper fallback).

Anything else: identify the upstream owner. ASR protocol / NIM deploy → `/riva-asr`. Scoring → here.

## Limitations

- **English-only by default.** Tokenization + normalization assume Latin script and en-US lexicon.
- **Strict-contiguous KER is conservative.** A near-miss like `cefa zolin` counts as a miss. That's intentional — pharmacy lookups fail on near-misses. Users wanting "soft" matching can switch to phoneme-level edit distance, which is a methodology extension, not a config tweak.
- **One model per eval run.** Comparing two models means running the eval twice and diffing the two `leaderboard_cycle<N>.md` files (or extending the recipe to write multi-model rows yourself).
- **Hosted-only paths assumed.** Self-hosted NIMs work but require `/riva-nim-setup` first.

## Next steps

- **Forward (KER > 0.3, manifest ≥ 100 rows):** `/digital-health-clinical-asr-finetune`.
- **Back to build (KER 0.1–0.3 on first eval, or `magpie_g2p` gap):** `/digital-health-clinical-asr-build`.
- **Stop (KER < 0.1):** the eval is saturated. Harden it before declaring victory.
- **Lateral** for ASR protocol / auth / streaming / self-hosted NIM details: `/riva-asr`.

## References

- [`references/offline-asr-recipe.md`](references/offline-asr-recipe.md) — full Step 3b Python recipe (`transcribe_manifest`, `resolve_asr_config`, `build_asr_auth`), function-ID catalog with call-shape notes, Whisper fallback, self-hosted Riva NIM setup
- [`references/scoring-recipes.md`](references/scoring-recipes.md) — pure-Python WER/CER/KER/SER scoring functions with the canonical 4-step normalization


安裝 digital-health-clinical-asr-eval

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-eval # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 NVIDIA/skills

相關技能

web-search
更新時間 2026-06-29
webapp-testing
更新時間 2026-06-29
lark-base
更新時間 2026-07-05
agentmail
更新時間 2026-06-29
OR