digital-health-clinical-asr-eval
NVIDIA/skills
選択したNIMに対して臨床ASRマニフェストをスコアリングし、5つのセクションからなるKERリーダーボードを作成し、評価後の決定木を通じてユーザーを誘導する。
...すべて拡張します臨床ASRフライホイール — ステージ3(評価)
⚠ エージェント:回答する前に、以下の「重要なワークフロールール」のセクションをお読みください。この SKILL.mdは自己完結型です。
evals/、references/、assets/は参照先であり、実体ではありません。このファイルに関する方法論の質問には直接回答してください。ユーザーが実際のマニフェストに対して実行することを明示的に要求した場合にのみ、ツールを呼び出してください。
あなたは「スコア付けおよびルーティング」ステージを担当します。ユーザーは NeMo 形式のmanifest.jsonlファイル(/digital-health-clinical-asr-buildからのもの、または外部から持ち込まれたもの)を持って到着します。 あなたは、選択された ASR NIM を使用してこれを文字起こしし、4 つの指標を評価し、5 セクションからなるリーダーボードを作成し、決定木を参照して、ユーザーを/digital-health-clinical-asr-finetune に進めるか、/digital-health-clinical-asr-build にループバックさせるか、あるいは評価を確定して終了させるかを決定します。
このスキルは音声を生成しません。マニフェストが存在しないか空の場合、ユーザーを/digital-health-clinical-asr-build に戻します。
音声データはお客様の環境外へ送信されます。クリップを送信する前に、この旨をユーザーに明示してください
この段階では、各マニフェスト行の WAV ファイルとその参照テキストが、外部の NVIDIA サービスに送信されます。最初の ASR 呼び出しを実行する前に、この点をユーザーに明示してください:
| サービス | 送信される内容 | 送信タイミング |
|---|---|---|
NVIDIA NVCF Parakeet/Nemotron ASR(grpc.nvcf.nvidia.com) |
マニフェストで参照されるすべてのオーディオクリップ(生のPCMバイト)に加え、参照用トランスクリプトおよびスコアリング用の臨床拡張メタデータ | ステップ3b、マニフェストの各行につき1件の通話 |
クリップは、ステージ2(ユーザーが作成した用語リストに基づくMagpie TTS)によって生成された合成音声である必要があります。実際の患者の音声ではありません。実際の ASR 録音、実際の患者とのやり取り、または PHI をこのスキルを通じて送信しないでください。その後、スコアリングはローカルで実行されます(純粋な Python による WER/CER/KER/SER、またはインストールされている場合はjiwer)。スコアリングステップ自体では何も送信されず、ASR ステップでのみ送信が行われます。
重要なワークフロールール(すべてのアクティベーションに適用)
方法論に関する質問(リーダーボードの構造、KERの定義、決定木)については、このファイルに基づいて回答してください。ユーザーが実際のマニフェストに対して実行することを明示的に要求しない限り、ツールの起動、他のスキルの呼び出し、スクリプトの実行を行わないでください。どのような応答においても、以下の事実を明示してください:
- まずオフランプを優先します。ユーザーがスコアリング以外の事項について質問している場合は、ワークフローを実行せずにルーティングを行い、処理を停止してください:
- ASRモデルカタログの選択/比較/代替NIM →
/riva-asr - ASR認証(APIキー、ベアラートークン、関数ID) →
/riva-asr - ASRのgRPCプロトコル、ストリーミング、バッチ処理、チャンキング、再試行 →
/riva-asr - NIMのデプロイ/
riva-build/riva-deploy→/riva-asr-custom - NGC / Docker / NVIDIA Container Toolkit →
/riva-nim-setup - マニフェストはまだない →
/digital-health-clinical-asr-build - 既知のKERを使用して今すぐ微調整を行う →
/digital-health-clinical-asr-finetune
- ASRモデルカタログの選択/比較/代替NIM →
- デフォルトの ASR NIM は
nvidia/parakeet-tdt-0.6b-v2です(NVCF 関数 ID:d3fe9151-442b-4204-a70d-5fcc597fd610、オフライン gRPC)。 環境変数による上書き:ASR_MODEL_NAME(リーダーボードの表示名)、ASR_NVCF_FUNCTION_ID(別のホスト型 NIM に切り替える — 例: Parakeetバックエンドに障害が発生している間はWhisper Large v3b702f636-…、あるいは微調整済みのNIM)、ASR_ENDPOINT(セルフホスト型gRPC;優先されます)。APIクレジットを消費する前に、選択されたNIMと解決されたfunction-idをエコーバックします。 - ASR文字起こしはステップ3b(NVCF gRPC +
riva.client.ASRService.offline_recognize、ステージ1と同じ認証パターン)に組み込まれています。 プロトコルや認証に関するより詳細な質問、代替のNIMカタログ、またはセルフホスト型Riva NIMの設定については、/riva-asrを参照してください。 - KERが最重要指標です。行ごとのチェック:フラグが立てられた
用語は、正規化された仮説(normalized hypothesis)内で順序通り、連続して、隣接して出現している必要があります。cefazolin → cefa zolinはミスです。集計されたWERでは臨床的に危険な失敗が見逃されがちですが、両方が報告されており、KERがゲート役を果たします。 by-ipa_sourceによる分割は、リーダーボードの中で最も有益な単一の数値です。merriam-websterとmagpie_g2pの差は、SSMLオーバーライドパイプラインが実際に機能していることを証明しています。ユーザーに声に出して読み聞かせてください。- 特殊ケースのルーティング。「
merriam-webster」の行は良好、「magpie_g2p」の行は不良 → これは発音カバレッジのギャップであり、モデルのギャップではない。/digital-health-clinical-asr-buildのステップ 2d へ戻す。最初の対応として/digital-health-clinical-asr-finetune を推奨してはならない。 - 5セクション構成のリーダーボード順。見出し(WER/CER/KER/SER)→
エンティティカテゴリ別のKER →IPAソース別のKER →ノイズレベル別のKER → 用語ごとのKER(最悪のものから順)。IPAソース別のセクションは必須であり、これがSSMLパイプラインが機能していることの証明となる。
目的
臨床ASRマニフェストをスコアリングし、5セクションからなるKERリーダーボードを作成し、評価後の決定木を通じてユーザーをルーティングする。方法論の詳細(メトリクスの定義、正規化、リーダーボードの順序、特殊ケースのルーティング)は、上記の「重要なワークフロールール」および以下の「手順」に記載されている。
このスキルを使用するタイミング
次のようなユーザーのフレーズで起動します:
- 「私のASRマニフェストをスコアリングして」
- 「Parakeet TDT v2のKERは?」
- 「サイクルNで評価を実行して」
- 「臨床ベンチマークで2つのASRモデルを比較して」
- 「リーダーボードを生成して」
- 「manifest.jsonl があるのですが、どのようにスコアを算出すればよいですか?」
- 「WERが0.07なのに、なぜKERが0.4になるのですか?」
- 「ファインチューニングを行うべきか?」(これは評価側の質問です。評価後の意思決定ツリーはこのスキル内にあります)
リテラルキーワードによる非アクティベーションチェック— ユーザーのメッセージに「authenticate」、「API key」、「bearer」、「function ID」、「gRPC」、「streaming」、「chunking」、「batching」、「transcription retry」、「riva-build」、「riva-deploy」、「NIM deploy」、「NGC」、Docker、Container Toolkitのいずれかが含まれている場合、または「どのASRモデルが最適か」「モデルを比較」「ベンダー間の違い」といった質問がある場合 — スコアリングワークフローを起動しないでください。 上記の「重要なワークフロールール #1」を適用し、適切な兄弟スキルへルーティングして処理を停止してください。これは、ユーザーがキーワードとともに「KER」や「eval」に言及した場合でも適用されます。
前提条件
- 臨床拡張フィールド(
term、entity_category、ipa_source、voice_id、noise_level、context_type)を含むNeMo 形式のマニフェスト。スキーマは、スキル構築用のreferences/manifest-schema.mdに記載されています。 NVIDIA_API_KEY がエクスポートされていること(ステージ 1 の前提条件は引き続き適用されます)。nvidia-riva-clientおよびsoundfileがインストールされていること(ステージ1の前提条件)。セルフホスト型 Riva NIM の詳細については、/riva-asrオプション B を参照してください。- ディスク上に実際にオーディオファイルが存在すること— API クレジットを消費する前に、manifest-schema リファレンスに記載されている audio-existence プリフライトを実行してください。
手順
3a. ASR NIM を選択する
デフォルト:NVCF gRPC(オフライン)経由のnvidia/parakeet-tdt-0.6b-v2、function-idd3fe9151-442b-4204-a70d-5fcc597fd610。 NVIDIAが現在推奨する英語ASR — カタログ内で最速かつ最安であり、NeMoの標準SFTレシピでサポートされているため、ステージ3のベースラインとステージ4の微調整は同じモデルファミリーを使用します。
3つのランタイム環境変数オーバーライド設定(リーダーボード表示用のASR_MODEL_NAME、別のホスト型NIMに切り替えるためのASR_NVCF_FUNCTION_ID、ASR_ENDPOINT:セルフホスト型gRPC用)に加え、代替NIMの完全なカタログ(Parakeet TDT 1.1B、 Parakeet CTC 1.1B、Whisper Large v3、Nemotron ストリーミング)が、関数 ID および呼び出し形式に関する注釈とともに用意されています:references/offline-asr-recipe.md。
APIクレジットを消費する前に、選択されたNIM、解決された関数ID、および環境変数の上書き設定をユーザーに通知してください。ホスト型Parakeet TDT v2での200行のマニフェストは安価ですが、1,000行のマニフェストで誤って間違ったモデルに対して実行してしまうと、そうはいきません。
3b. 文字起こし
manifest.jsonlの各行について、audio_filepathを文字起こしし、per_sample.jsonを書き出します(1行につき1つのJSONオブジェクト、JSONLまたはJSON配列 — 呼び出し側の選択による):
{
"audio_filepath": "...",
"ref": "",
"hyp": "",
"term": "",
"entity_category": "",
"ipa_source": "",
"voice_id": "",
"noise_level": "",
"context_type": ""
}
レシピ(Pythonコード全文はreferences/offline-asr-recipe.md に記載):transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US") は、NVCF へのオフライン gRPC ストリーム(または、セルフホスト型 Riva で設定されている場合はASR_ENDPOINTへのストリーム)を開き、各行ごとにriva.client.ASRService.offline_recognizeを呼び出し、上記のJSONLを出力します。認証の形式は、ステージ1セットアップのスモークテストと同じです。 エージェント・ハーネスはapi_key を明示的に渡します。レシピは冒頭で 3 つの環境変数オーバーライド(ASR_NVCF_FUNCTION_ID、ASR_MODEL_NAME、ASR_ENDPOINT)を読み込むため、監査担当者は設定項目を一箇所で確認できます。
Whisperフォールバック(ParakeetのNVCFバックエンドが、TritonからのCUDA不正メモリアクセスにより障害が発生した場合)およびセルフホスト型Riva NIM(ASR_ENDPOINT=localhost:50051)の環境変数パターン:references/offline-asr-recipe.md(§Whisper フォールバック、§セルフホスト型 Riva NIM)を参照してください。
耐障害性に関する設定はユーザーに委ねられています。バッチ処理中にNVCFがRESOURCE_ENHAUSTEDを返した場合、ループはその行で例外を発生させ、失敗した行から再実行します。ストリーミング/バッチ処理/バックオフ付きリトライは本ドキュメントの範囲外です —/riva-asrを参照してください。
3c. 4つの指標を算出する
各行について、以下を計算します:
| メトリクス | 測定対象 | 記録する理由 |
|---|---|---|
| WER | 単語エラー率(正規化後のトークンに対するレベンシュタイン距離) | 業界標準;臨床用途には不十分な指標 |
| CER | 文字エラー率 | 長い複合名称におけるニアミスを検出 |
| KER★ | キーワード誤検出率 — フラグが立てられた用語が仮説に現れたか(正規化済み、連続一致)? |
主要な臨床シグナル |
| SER | 文のエラー率(誤りがあれば1、完璧なら0) | 妥当性の境界;医師が実際に経験すること |
正規化(4つの指標をすべて算出する前に、参照文と仮説文の両方に適用):
- 小文字化。
- NFKD正規化(スマートクォート→ASCIIなど)。
- ハイフン以外の句読点を削除。
- 連続する空白を1つのスペースにまとめる。
インラインのスコアリングレシピ—normalize/edit_distance/wer/cer/ker/ser(純粋なPython、jiwerへの依存なし):references/scoring-recipes.mdを参照。各メトリックについて、行ごとのスコアの平均(mean(per-row score))を算出し、行ごとに集計する。
厳格な KER— 用語は、正規化された仮説において順序通りに、かつ隣接して出現しなければならない。これは保守的な手法である:cefazolin → cefa zolin はミスとしてカウントされる。これは臨床的には正しい判断である — 下流の薬局検索では、スペルミスのあるトークンで検索に失敗するだろう。
KERは周辺の誤りを減点対象としません。用語が正しく、文の残りが無意味な文字列である行でも、KERは0と評価されます。その行のWERは、より広範な問題を別途浮き彫りにします。
3d. 内訳 + リーダーボード
以下の順序で、5つのセクションからなるMarkdown形式のリーダーボードを作成してください:
- 見出し— 選択したモデルの総合 WER、CER、KER、SER。
-
エンティティカテゴリ別のKER— 薬剤 vs 処置 vs 解剖学 vs … これは、ユーザーが実運用において実際に重視する点です。 -
ipa_source別の KER—リーダーボードの中で最も情報量の多い単一の数値。merriam-websterとmagpie_g2pの行間の差は、SSML オーバーライドパイプラインが実際に機能している証拠です。このセクションをユーザーに声に出して読み上げてください。 -
noise_level別の KER— 臨床環境はノイズが多いものです。snr_5dbの行は、クリーンなデータよりも現実に近い値を示しています。 - 用語ごとの KER(悪い順) — これらがステージ 4 の微調整ターゲットとなります。
merriam-webster と magpie_g2p の差の解釈を含む、代表的なipa_sourceスプリット:references/scoring-recipes.md§Representative ipa_source split。 デルタ値はデプロイメントの状況を物語ります。ユーザーが大きなギャップを認識し「ファインチューニングすべきか?」と尋ねた場合、答えは「まだ早い」です。ユーザーを/digital-health-clinical-asr-build の IPA QA パイプライン(ステージ 2d)に戻してください。以下の決定木を参照してください。
決定木(評価後)
優先度カテゴリのKER(ほとんどの臨床ワークフローでは薬剤KER、外科ワークフローでは処置KER)を確認し、以下のようにルーティングする:
| 優先度カテゴリ別のKER | 推奨 |
|---|---|
| > 0.3 | /digital-health-clinical-asr-finetune。マニフェストはすでにNeMo形式に対応しています。注:信頼できる微調整シグナルを得るには、行数が100以上である必要があります。マニフェストがこれより少ない場合は、まず/digital-health-clinical-asr-build を使用して行数を増やしてください。 |
| 0.1 – 0.3 | 用語リストを拡張するか(新しいドメイン用語を追加して/digital-health-clinical-asr-buildを実行する — 通常、チューニングよりも低コストでより多くの失敗を明らかにできる)、あるいはファインチューニングを行う。最初の評価では拡張を行い、マニフェストをすでに拡大している後の評価ではチューニングを行う。 |
| < 0.1 | 強力なベースラインです。まだチューニングは行わないでください。指標が飽和状態にある状態で最適化することになります。評価をさらに厳しくしてください:音声、ノイズレベル、文脈、敵対的用語を追加します。/digital-health-clinical-asr-build に戻って繰り返してください。 |
特例 —merriam-websterの行はスコアが高いが、magpie_g2pの行は悪い。これは発音ヒントのカバレッジ不足であり、モデルの問題ではない。/digital-health-clinical-asr-buildのステップ 2d(IPA QA レビュー)に戻り、/digital-health-clinical-asr-finetune には進まないでください。TTS の発音ギャップに対して微調整を行うと、モデルは自身の誤りを誤認識するよう学習してしまいます。これは誤った修正です。
例
シナリオA — 新しいサイクル1のマニフェストに対する最初の評価。ユーザー:「term およびentity_categoryフィールドを含む、200件の臨床音声レコードがすでに含まれたmanifest.jsonlがあります。どのようにスコア付けすればよいですか?」→ ステージ2を完全にスキップします。audio-existenceのプリフライトを実行します。parakeet-tdt-0.6b-v2(デフォルト)を選択し、その選択内容と解決済みのfunction-idをユーザーに通知する。インライン化されたステップ3bのレシピ(transcribe_manifest(...))を実行する。 4つのメトリクスに対してスコアを算出します。5セクションからなるリーダーボードを生成します。ipa_sourceごとの内訳をユーザーに読み上げます。医薬品KERに対して決定木を適用します。
シナリオ B — 混合結果の解釈。ユーザー:「評価結果によると、merriam-websterとタグ付けされた行ではKERが0.05ですが、magpie_g2pとタグ付けされた行では0.40です。微調整すべきでしょうか?」→ いいえ — これは特殊なケースです。モデルに問題はなく、発音のヒントがロングテール用語をカバーできていないだけです。 ユーザーを/digital-health-clinical-asr-buildのステップ 2d に戻し、magpie_g2p行を検証して、検証済みの IPA をpronunciation_overrides.csv に追加するよう指示する。再構築後にステージ 3 を再実行してから、ステージ 4 を再検討する。
生成された成果物
per_sample.json— すべての臨床拡張フィールドが保持された行ごとの転写結果(ASRのhypがマニフェストのrefおよびメタデータと結合されている)results.csv— 行ごとの WER/CER/KER/SER スコアleaderboard_cycle— 5つのセクションからなるMarkdown形式のレポート.md
(ファイル名はユーザーが自由に指定できます。上記の名前は、このスキルの残りの部分で想定されている規約です。)
トラブルシューティング
- 「マニフェストが見つかりません」→ ユーザーがステージ 2 をスキップしました。
/digital-health-clinical-asr-buildにリダイレクトするか、$MANIFEST_PATHを確認してください。 - すべての行で KER=1→
refとhyp間の正規化が一致していません。両側に 4 つの正規化手順を適用してください。 - すべての行で KER=0 だが WER が高い→ マニフェストの不整合(音声行の不一致)が考えられます。いくつかの
(参照、仮説)ペアを手作業で抜き打ち確認してください。 merriam-websterの値が低く、magpie_g2pの値が高い→ 発音のカバー範囲にギャップがある。/digital-health-clinical-asr-buildのステップ2dに進んでください。微調整は行わないでください— 問題はモデル側ではありません。-
merriam-websterとmagpie_g2pの両方が高い→ 実際のモデルギャップ。ステージ4が適切な手順(マニフェスト行数 ≥ 100行)。 クリーンな行は問題なし、snr_5dbが急増→ 頑健性のギャップ;/digital-health-clinical-asr-build経由でノイズの多様性を拡大。- Riva-NIMとオフラインNeMoの結果に乖離がある→ Rivaの前処理
/riva-buildのフラグ。/riva-asr-customへのルート。 - 大規模なマニフェストで
RESOURCE_EXHAUSTEDが発生 → 30 秒後に再試行;削除された行をスライスして再実行。組み込みのバックオフ:/riva-asr。 Auth.__init__() で 'ssl_cert' が取得された/Parakeet 関数 ID での CUDA による不正なメモリアクセス:references/offline-asr-recipe.mdを参照(ssl_root_cert の名称変更+§Whisper フォールバック)。
その他:上流の所有者を特定してください。ASRプロトコル/NIMデプロイ →/riva-asr。スコアリング → こちら。
制限事項
- デフォルトでは英語のみ対応。トークン化および正規化は、ラテン文字および en-US 辞書を前提としています。
- 厳密に連続したKERは保守的です。
cefa zolinのようなニアミスもミスとしてカウントされます。これは意図的な仕様です — 薬局での検索では、ニアミスでは検索に失敗します。「ソフト」なマッチングを希望するユーザーは、フォネムレベルの編集距離に切り替えることができます。これは設定の微調整ではなく、方法論の拡張です。 - 評価実行ごとに1つのモデルのみ使用可能です。2つのモデルを比較するには、評価を2回実行し、2つの
leaderboard_cycleファイルの差分を比較する必要があります(または、レシピを拡張してマルチモデル行を自分で記述することも可能です)。.md - ホスト型のみを想定しています。セルフホスト型のNIMも動作しますが、その場合はまず
/riva-nim-setupを実行する必要があります。
次のステップ
- フォワード(KER > 0.3、マニフェスト行数 ≥ 100行):
/digital-health-clinical-asr-finetune。 - ビルドに戻る(初回評価で KER が 0.1~0.3、または
magpie_g2pとのギャップがある場合):/digital-health-clinical-asr-build。 - 停止(KER < 0.1):評価は飽和状態です。成功を宣言する前に、モデルを堅牢化してください。
- ASRプロトコル、認証、ストリーミング、セルフホスト型NIMの詳細については、
/riva-asrを参照してください。
参考文献
references/offline-asr-recipe.md— ステップ 3b の完全な Python レシピ(transcribe_manifest、resolve_asr_config、build_asr_auth)、呼び出し形状の注釈付き関数 ID カタログ、Whisper フォールバック、セルフホスト型 Riva NIM のセットアップreferences/scoring-recipes.md— 標準的な4段階正規化を用いた、純粋なPythonによるWER/CER/KER/SERスコアリング関数
---
name: digital-health-clinical-asr-eval
description: Score a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
license: Apache-2.0
---
<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->
# Clinical ASR Flywheel — Stage 3 (Eval)
> **⚠ Agent: read the Critical Workflow Rules section below before answering.** This SKILL.md is self-contained — `evals/`, `references/`, and `assets/` are pointers, not load-bearing. Answer methodology questions from this file directly; only invoke tools when the user explicitly asks to execute against a real manifest.
You are the **score-and-route** stage. The user arrives with a NeMo-format `manifest.jsonl` (either from `/digital-health-clinical-asr-build` or carried in from elsewhere). You transcribe it via the chosen ASR NIM, score four metrics, produce a five-section leaderboard, and read the decision tree to decide whether the user should advance to `/digital-health-clinical-asr-finetune`, loop back to `/digital-health-clinical-asr-build`, or stop and harden the eval.
**This skill does not generate audio.** If the manifest is missing or empty, send the user back to `/digital-health-clinical-asr-build`.
## Audio leaves your environment — disclose this to the user before any clip is sent
This stage transmits each manifest row's WAV file plus its reference text to an external NVIDIA service. Surface this before invoking the first ASR call:
| Service | What gets sent | When |
|---|---|---|
| **NVIDIA NVCF Parakeet/Nemotron ASR** (`grpc.nvcf.nvidia.com`) | Every audio clip referenced by the manifest (raw PCM bytes), plus the reference transcript and the clinical-extension metadata for scoring | Step 3b, one call per manifest row |
The clips should be **synthetic audio generated by Stage 2** (Magpie TTS over a user-curated term list) — not real patient audio. **Do not pass real ASR recordings, real patient encounters, or any PHI through this skill.** Scoring then runs locally (pure-Python WER/CER/KER/SER, or `jiwer` if installed). The scoring step itself does not transmit anything; only the ASR step does.
## Critical workflow rules (apply on every activation)
For methodology questions (leaderboard structure, KER definition, decision tree), answer from this file. Don't invoke tools, call other skills, or run scripts unless the user explicitly asks to execute against a real manifest. Surface these facts in any response:
1. **Off-ramp first.** If the user is asking about something outside scoring, route and stop without running any workflow:
- ASR model-catalog selection / comparison / alternative NIMs → `/riva-asr`
- ASR auth (API keys, bearer tokens, function IDs) → `/riva-asr`
- ASR gRPC protocol, streaming, batching, chunking, retries → `/riva-asr`
- NIM deploy / `riva-build` / `riva-deploy` → `/riva-asr-custom`
- NGC / Docker / NVIDIA Container Toolkit → `/riva-nim-setup`
- No manifest yet → `/digital-health-clinical-asr-build`
- Wants to fine-tune now with a known KER → `/digital-health-clinical-asr-finetune`
2. **Default ASR NIM is `nvidia/parakeet-tdt-0.6b-v2`** (NVCF function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`, offline gRPC). Env-var overrides: `ASR_MODEL_NAME` (leaderboard display name), `ASR_NVCF_FUNCTION_ID` (swap to a different hosted NIM — e.g. Whisper Large v3 `b702f636-…` while the Parakeet backend is faulting, or a fine-tuned NIM), `ASR_ENDPOINT` (self-hosted gRPC; takes precedence). Echo the chosen NIM **and the resolved function-id** back before spending API credits.
3. **ASR transcription is inlined in Step 3b** (NVCF gRPC + `riva.client.ASRService.offline_recognize`, same auth pattern as Stage 1). For deeper protocol/auth questions, alternative NIM catalogs, or self-hosted Riva NIM configuration, defer to `/riva-asr`.
4. **KER is the headline.** Per-row check: the flagged `term` words must appear *in order, contiguous, adjacent* in the normalized hypothesis. `cefazolin → cefa zolin` is a miss. Aggregate WER hides clinically dangerous failures; both are reported, KER is the gate.
5. **The by-`ipa_source` split is the most informative single number** in the leaderboard. The `merriam-webster` vs `magpie_g2p` delta proves the SSML override pipeline is doing real work. Read it aloud to the user.
6. **Special-case routing.** `merriam-webster` rows good, `magpie_g2p` rows bad → pronunciation-coverage gap, **not** a model gap. Route back to `/digital-health-clinical-asr-build` Step 2d. **Do NOT recommend `/digital-health-clinical-asr-finetune`** as a first response.
7. **Five-section leaderboard order.** Headline (WER/CER/KER/SER) → KER by `entity_category` → KER by `ipa_source` → KER by `noise_level` → Per-term KER worst-first. The by-`ipa_source` section is mandatory; it is the proof the SSML pipeline works.
## Purpose
Score a clinical-ASR manifest, produce a five-section KER leaderboard, and route the user via the post-eval decision tree. Methodology details (metric definitions, normalization, leaderboard order, special-case routing) live in Critical Workflow Rules above and Instructions below.
## When to use this skill
Activate on user phrases like:
- "Score my ASR manifest"
- "What's the KER on Parakeet TDT v2?"
- "Run the eval on cycle-N"
- "Compare two ASR models on the clinical benchmark"
- "Generate the leaderboard"
- "I have a manifest.jsonl, how do I score it?"
- "Why is KER 0.4 when WER is 0.07?"
- "Should we fine-tune?" *(this is the eval-side question — the post-eval decision tree lives in this skill)*
**Literal-keyword non-activation check** — if the user's message contains any of `authenticate`, `API key`, `bearer`, `function ID`, `gRPC`, `streaming`, `chunking`, `batching`, `transcription retry`, `riva-build`, `riva-deploy`, `NIM deploy`, `NGC`, `Docker`, `Container Toolkit`, or asks "which ASR model is best" / "compare models" / "vendor differences" — **do NOT activate** the scoring workflow. Apply Critical Workflow Rule #1 above to route to the right sibling skill and stop. This applies even if the user mentions "KER" or "eval" alongside the keyword.
## Prerequisites
- **A NeMo-format manifest** with the clinical extension fields (`term`, `entity_category`, `ipa_source`, `voice_id`, `noise_level`, `context_type`). The schema is documented in the build skill's `references/manifest-schema.md`.
- **`NVIDIA_API_KEY`** exported (Stage 1 prerequisite still applies).
- **`nvidia-riva-client` + `soundfile`** installed (Stage 1 prerequisite). For self-hosted Riva NIM details, see `/riva-asr` Option B.
- **Audio files actually present on disk** — run the audio-existence pre-flight from the manifest-schema reference before spending API credits.
## Instructions
### 3a. Pick the ASR NIM
**Default**: `nvidia/parakeet-tdt-0.6b-v2` via NVCF gRPC (offline), function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`. NVIDIA's current English ASR recommendation — fastest/cheapest in the catalog, and supported in NeMo's stock SFT recipe so the Stage 3 baseline and a Stage 4 fine-tune ride the same model family.
Three runtime env-var override knobs (`ASR_MODEL_NAME` for leaderboard display, `ASR_NVCF_FUNCTION_ID` to swap to a different hosted NIM, `ASR_ENDPOINT` for self-hosted gRPC) plus the full alternate-NIM catalog (Parakeet TDT 1.1B, Parakeet CTC 1.1B, Whisper Large v3, Nemotron streaming) with function IDs and call-shape notes: `references/offline-asr-recipe.md`.
Echo the chosen NIM, the resolved function-id, and any env-var overrides to the user **before** spending API credits. A 200-row manifest on hosted Parakeet TDT v2 is cheap; an accidental run against the wrong model on a 1,000-row manifest is not.
### 3b. Transcribe
For each row in `manifest.jsonl`, transcribe `audio_filepath` and write `per_sample.json` (one JSON object per row, JSONL or a JSON array — caller's choice):
```json
{
"audio_filepath": "...",
"ref": "<row.text>",
"hyp": "<asr output>",
"term": "<row.term>",
"entity_category": "<row.entity_category>",
"ipa_source": "<row.ipa_source>",
"voice_id": "<row.voice_id>",
"noise_level": "<row.noise_level>",
"context_type": "<row.context_type>"
}
```
**Recipe** (full Python in `references/offline-asr-recipe.md`): `transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US")` opens an offline gRPC stream to NVCF (or to `ASR_ENDPOINT` if set for self-hosted Riva), calls `riva.client.ASRService.offline_recognize` per row — sentences in a clinical manifest are ≤ 30 s so no streaming/batching needed — and writes the JSONL above. Same `auth_for` shape as the Stage 1 setup smoke test. The agent harness passes `api_key` explicitly; the recipe reads the three env-var overrides (`ASR_NVCF_FUNCTION_ID`, `ASR_MODEL_NAME`, `ASR_ENDPOINT`) at the top so auditors see the knobs in one place.
**Whisper fallback** (when Parakeet's NVCF backend faults with `CUDA illegal-memory-access` from Triton) and **self-hosted Riva NIM** (`ASR_ENDPOINT=localhost:50051`) env-var patterns: see `references/offline-asr-recipe.md` (§Whisper fallback, §Self-hosted Riva NIM).
**Resilience knobs deferred to the user.** If NVCF returns `RESOURCE_EXHAUSTED` mid-batch, the loop raises on that row; re-run from the failing row. Streaming/batching/retry-with-backoff are out of scope — see `/riva-asr`.
### 3c. Score four metrics
For every row, compute:
| Metric | What it measures | Why we keep it |
|---|---|---|
| **WER** | Word error rate (Levenshtein on tokens, after normalization) | Industry standard; blunt instrument for clinical |
| **CER** | Character error rate | Catches near-misses on long compound names |
| **KER** ★ | Keyword error rate — did the flagged `term` appear in the hypothesis (normalized, **contiguous** match)? | **Headline clinical signal** |
| **SER** | Sentence error rate (1 if any wrong, 0 if perfect) | Sanity bound; what the doctor experiences |
**Normalization (apply to both `ref` and `hyp` before all four metrics):**
1. Lowercase.
2. NFKD-normalize (smart quotes → ASCII, etc.).
3. Strip punctuation **except hyphen**.
4. Collapse whitespace runs to a single space.
**Inline scoring recipes** — `normalize` / `edit_distance` / `wer` / `cer` / `ker` / `ser` (pure-Python, no `jiwer` dependency): see `references/scoring-recipes.md`. Aggregate across rows by taking `mean(per-row score)` for each metric.
**Strict KER** — term words must appear *in order, adjacent* in the normalized hypothesis. This is conservative: `cefazolin → cefa zolin` counts as a miss. That's the right call clinically — a downstream pharmacy lookup will fail on the misspelled token.
KER does **not** punish surrounding errors. A row where the term is correct and the rest of the sentence is garbage still scores KER=0; the WER on that row will surface the broader problem separately.
### 3d. Breakdowns + leaderboard
Write a five-section markdown leaderboard, **in this order**:
1. **Headline** — overall WER, CER, KER, SER for the chosen model.
2. **KER by `entity_category`** — drug vs procedure vs anatomy vs ... This is what the user actually cares about for deployment.
3. **KER by `ipa_source`** — **the most informative single number in the leaderboard.** The delta between `merriam-webster` and `magpie_g2p` rows is the proof the SSML override pipeline is doing real work. *Read this section aloud to the user.*
4. **KER by `noise_level`** — clinical environments are loud. `snr_5db` rows are closer to reality than `clean`.
5. **Per-term KER** (worst first) — these are your Stage 4 fine-tune targets.
A representative `ipa_source` split with the merriam-webster vs magpie_g2p delta interpretation: `references/scoring-recipes.md` §Representative ipa_source split. The delta tells the deployment story — if the user sees a wide gap and asks "should we fine-tune?", the answer is *not yet*; route them back to `/digital-health-clinical-asr-build`'s IPA QA pipeline (Stage 2d). See the decision tree below.
## Decision tree (after eval)
Read the **priority-category KER** (drug KER for most clinical workflows, procedure KER for surgical workflows) and route:
| KER on priority category | Recommend |
|---|---|
| **> 0.3** | `/digital-health-clinical-asr-finetune`. Manifest is already NeMo-format-ready. Note: rows ≥ 100 is the minimum for a believable fine-tune signal; if the manifest is smaller, grow it first via `/digital-health-clinical-asr-build`. |
| **0.1 – 0.3** | Either expand the term list (back to `/digital-health-clinical-asr-build` with new domain terms — usually surfaces more failures cheaper than tuning) **or** fine-tune. On a *first* eval, expand. On a *later* eval where you've already grown the manifest, tune. |
| **< 0.1** | Strong baseline. Don't tune yet — you'd be optimizing against a saturated metric. Push the eval harder: add voices, noise levels, contexts, adversarial terms. Loop back to `/digital-health-clinical-asr-build`. |
**Special case — `merriam-webster` rows score well but `magpie_g2p` rows are bad.** That's a pronunciation-hint coverage gap, **not a model gap**. Route back to `/digital-health-clinical-asr-build` Step 2d (IPA QA review), not to `/digital-health-clinical-asr-finetune`. Fine-tuning over a TTS-pronunciation gap teaches the model to mis-recognize the model's own mistakes — the wrong fix.
## Examples
**Scenario A — first eval on a fresh cycle-1 manifest.** User: *"I have `manifest.jsonl` with 200 clinical audio rows already, with `term` and `entity_category` fields. How do I score it?"* → Skip Stage 2 entirely. Run the audio-existence pre-flight. Pick `parakeet-tdt-0.6b-v2` (default) and echo the choice + resolved function-id. Run the inlined Step 3b recipe (`transcribe_manifest(...)`). Score the four metrics. Produce the five-section leaderboard. Read the by-`ipa_source` split to the user. Apply the decision tree against drug KER.
**Scenario B — interpreting a mixed result.** User: *"Eval shows KER 0.05 on rows tagged `merriam-webster` but 0.40 on rows tagged `magpie_g2p`. Should I fine-tune?"* → No — this is the special case. The model is fine; the pronunciation hints aren't covering the long-tail terms. Route the user back to `/digital-health-clinical-asr-build` Step 2d to audition the `magpie_g2p` rows and append verified IPA to `pronunciation_overrides.csv`. Re-run Stage 3 after the rebuild before reconsidering Stage 4.
## Artifacts produced
- `per_sample.json` — per-row transcription results with all clinical-extension fields preserved (the ASR `hyp` joined to the manifest's `ref` and metadata)
- `results.csv` — per-row WER/CER/KER/SER scores
- `leaderboard_cycle<N>.md` — five-section markdown report
(File names are user-chosen; the names above are conventions the rest of this skill assumes.)
## Troubleshooting
- **"No manifest found"** → user skipped Stage 2. Route to `/digital-health-clinical-asr-build` or confirm `$MANIFEST_PATH`.
- **All rows KER=1** → normalization mismatch between `ref` and `hyp`. Apply the four normalization steps to both sides.
- **All rows KER=0 but WER high** → likely misaligned manifest (audio row mismatch). Spot-check a few `(ref, hyp)` pairs by hand.
- **`merriam-webster` low, `magpie_g2p` high** → pronunciation-coverage gap. Route to `/digital-health-clinical-asr-build` Step 2d. **Don't fine-tune** — model isn't the problem.
- **Both `merriam-webster` and `magpie_g2p` high** → real model gap. Stage 4 is the right route (manifest ≥ 100 rows).
- **`clean` rows fine, `snr_5db` balloons** → robustness gap; expand noise diversity via `/digital-health-clinical-asr-build`.
- **Riva-NIM and offline NeMo results diverge** → Riva preprocessing / `riva-build` flags. Route to `/riva-asr-custom`.
- **`RESOURCE_EXHAUSTED` on large manifests** → retry after 30 s; slice + re-run dropped rows. Built-in backoff: `/riva-asr`.
- **`Auth.__init__() got 'ssl_cert'`** / **CUDA illegal-memory-access on Parakeet function ID**: see `references/offline-asr-recipe.md` (ssl_root_cert rename + §Whisper fallback).
Anything else: identify the upstream owner. ASR protocol / NIM deploy → `/riva-asr`. Scoring → here.
## Limitations
- **English-only by default.** Tokenization + normalization assume Latin script and en-US lexicon.
- **Strict-contiguous KER is conservative.** A near-miss like `cefa zolin` counts as a miss. That's intentional — pharmacy lookups fail on near-misses. Users wanting "soft" matching can switch to phoneme-level edit distance, which is a methodology extension, not a config tweak.
- **One model per eval run.** Comparing two models means running the eval twice and diffing the two `leaderboard_cycle<N>.md` files (or extending the recipe to write multi-model rows yourself).
- **Hosted-only paths assumed.** Self-hosted NIMs work but require `/riva-nim-setup` first.
## Next steps
- **Forward (KER > 0.3, manifest ≥ 100 rows):** `/digital-health-clinical-asr-finetune`.
- **Back to build (KER 0.1–0.3 on first eval, or `magpie_g2p` gap):** `/digital-health-clinical-asr-build`.
- **Stop (KER < 0.1):** the eval is saturated. Harden it before declaring victory.
- **Lateral** for ASR protocol / auth / streaming / self-hosted NIM details: `/riva-asr`.
## References
- [`references/offline-asr-recipe.md`](references/offline-asr-recipe.md) — full Step 3b Python recipe (`transcribe_manifest`, `resolve_asr_config`, `build_asr_auth`), function-ID catalog with call-shape notes, Whisper fallback, self-hosted Riva NIM setup
- [`references/scoring-recipes.md`](references/scoring-recipes.md) — pure-Python WER/CER/KER/SER scoring functions with the canonical 4-step normalization
digital-health-clinical-asr-evalをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/NVIDIA/skills/tree/main/skills/digital-health-clinical-asr-eval # Copy SKILL.md to your .claude/skills/ directory
コピー





家
