digital-health-clinical-asr-eval
NVIDIA/skills
根据选定的NIM对临床ASR清单进行评分,生成包含五个部分的KER排行榜,并通过评估后决策树引导用户。
...展开全部临床 ASR 飞轮 — 第 3 阶段(评估)
⚠ 代理:回答前请阅读下文的“关键工作流规则”部分。此 SKILL.md 文件是自包含
的——evals/、references/和assets/仅为引用,不承担实际功能。请直接根据此文件回答方法论问题;仅当用户明确要求针对真实清单执行操作时,才调用相关工具。
您负责评分与路由阶段。用户提交了一个 NeMo 格式的manifest.jsonl 文件(该文件可能来自/digital-health-clinical-asr-build,也可能从其他地方导入)。 您需通过选定的 ASR NIM 对该文件进行转录,计算四项指标的得分,生成包含五个部分的排行榜,并根据决策树判断用户应继续前往/digital-health-clinical-asr-finetune、回退至/digital-health-clinical-asr-build,还是停止并固化评估结果。
该技能不生成音频。如果清单文件缺失或为空,请将用户重定向回/digital-health-clinical-asr-build。
音频将离开您的环境——在发送任何音频片段之前,请向用户披露此信息
此阶段会将每个清单行对应的 WAV 文件及其参考文本传输至外部 NVIDIA 服务。请在调用首次 ASR 请求前向用户提示此信息:
| 服务 | 发送内容 | 何时 |
|---|---|---|
NVIDIA NVCF Parakeet/Nemotron ASR(grpc.nvcf.nvidia.com) |
清单中引用的每个音频片段(原始 PCM 字节),以及参考转录文本和用于评分所需的临床扩展元数据 | 步骤 3b,每个清单行对应一次调用 |
音频片段应为第 2 阶段生成的合成音频(基于用户精心编排的术语表通过 Magpie TTS生成的)——而非真实患者的音频。请勿通过此技能传输真实的 ASR 录音、真实的患者就诊记录或任何受保护健康信息(PHI)。随后评分将在本地运行(纯 Python 实现的 WER/CER/KER/SER,或已安装的jiwer)。评分步骤本身不会传输任何数据;仅 ASR 步骤会传输数据。
关键工作流规则(适用于每次激活)
对于方法论问题(排行榜结构、KER定义、决策树),请从本文件中给出答案。除非用户明确要求针对真实的清单执行操作,否则请勿调用工具、调用其他技能或运行脚本。在任何响应中均应明确说明以下事实:
- 优先采用“岔路”处理。如果用户询问的内容超出评分范围,请进行路由并停止,且不运行任何工作流:
- ASR 模型目录的选择/比较/替代 NIM →
/riva-asr - ASR 授权(API 密钥、承载令牌、函数 ID)→
/riva-asr - ASR gRPC 协议、流式处理、批处理、分块处理、重试 →
/riva-asr - NIM 部署 /
riva-build/riva-deploy→/riva-asr-custom - NGC / Docker / NVIDIA 容器工具包 →
/riva-nim-setup - 尚无清单 →
/digital-health-clinical-asr-build - 希望立即使用已知的 KER 进行微调 →
/digital-health-clinical-asr-finetune
- ASR 模型目录的选择/比较/替代 NIM →
- 默认 ASR NIM 为
nvidia/parakeet-tdt-0.6b-v2(NVCF 函数 IDd3fe9151-442b-4204-a70d-5fcc597fd610,离线 gRPC)。 环境变量覆盖:ASR_MODEL_NAME(排行榜显示名称)、ASR_NVCF_FUNCTION_ID(切换到另一个托管的 NIM — 例如 当 Parakeet 后端出现故障时切换至 Whisper Large v3b702f636-…,或切换至微调过的 NIM),ASR_ENDPOINT(自托管 gRPC;优先级最高)。在消耗 API 积分之前,先回显所选的 NIM和解析后的 function-id。 - ASR 转录功能内联于步骤 3b 中(NVCF gRPC +
riva.client.ASRService.offline_recognize,认证模式与第一阶段相同)。 有关更深入的协议/认证问题、替代 NIM 目录或自托管的 Riva NIM 配置,请参阅/riva-asr。 - KER 是关键指标。逐行检查:标记的
术语词必须按顺序、连续且相邻地出现在规范化假设中。cefazolin → cefa zolin即为错误。聚合 WER 会掩盖临床上的严重错误;两者均需报告,KER 是最终筛选标准。 按 ipa_source划分的数据是排行榜中最具参考价值的单一指标。Merriam-Webster与magpie_g2p之间的差异证明,SSML 覆盖管道确实发挥了作用。请向用户大声朗读结果。- 特殊情况处理。
merriam-webster行表现良好,magpie_g2p行表现不佳 → 这是发音覆盖范围的差距,而非模型差距。请回溯至/digital-health-clinical-asr-build第 2d 步。切勿将/digital-health-clinical-asr-finetune作为首选方案。 - 五部分排行榜排序。标题(WER/CER/KER/SER)→ 按
实体类别(entity_category)划分的 KER → 按IPA 来源(ipa_source)划分的 KER → 按噪声级别(noise_level)划分的 KER → 按术语划分的 KER(从最差开始)。按 IPA 来源(ipa_source)划分的部分是必备的;它是 SSML 管道正常工作的证明。
目的
对临床ASR清单进行评分,生成五部分的KER排行榜,并通过评估后决策树引导用户。方法论细节(指标定义、归一化、排行榜排序、特殊情况引导)详见上文的“关键工作流规则”及下文的“操作说明”。
何时使用此技能
在用户说出以下短语时触发:
- “为我的 ASR 清单评分”
- “Parakeet TDT v2 的 KER 是多少?”
- “对第N个周期运行评估”
- “在临床基准数据集上比较两个ASR模型”
- “生成排行榜”
- “我有一个 manifest.jsonl 文件,该如何对其进行评分?”
- “为什么WER是0.07时,KER却是0.4?”
- “是否需要进行微调?”(这是评估阶段的问题——评估后的决策树位于该技能中)
字面关键字非触发检查——如果用户消息中包含“authenticate”、“API key”、“bearer”、“function ID”、“gRPC”、“streaming”、“chunking”、“batching”、“transcription retry”、“riva-build”、“riva-deploy”、“NIM deploy”、“NGC”、Docker、Container Toolkit,或询问“哪个 ASR 模型最好”/“比较模型”/“供应商差异”——请勿激活评分工作流。 应用上述关键工作流规则 #1,将请求路由至正确的同级技能并终止处理。即使用户在提及关键词的同时还提到了“KER”或“eval”,此规则依然适用。
先决条件
- 一个包含临床扩展字段(
term、entity_category、ipa_source、voice_id、noise_level、context_type)的 NeMo 格式清单。该模式在构建技能的references/manifest-schema.md中有详细说明。 - 已导出
NVIDIA_API_KEY(第一阶段的先决条件仍然适用)。 - 已安装
nvidia-riva-client+soundfile(第一阶段先决条件)。有关自托管的 Riva NIM 的详细信息,请参阅/riva-asr方案 B。 - 确保音频文件实际存在于磁盘上——在消耗 API 积分之前,请运行 manifest-schema 参考文档中的 audio-existence 预检查。
操作指南
3a. 选择 ASR NIM
默认:通过 NVCF gRPC(离线)访问nvidia/parakeet-tdt-0.6b-v2,函数 ID 为d3fe9151-442b-4204-a70d-5fcc597fd610。 NVIDIA 当前推荐的英语 ASR 模型——目录中速度最快、成本最低,且在 NeMo 的标准 SFT 配方中受支持,因此第 3 阶段基线和第 4 阶段微调可基于同一模型家族。
三个运行时环境变量覆盖控件(ASR_MODEL_NAME用于排行榜显示,ASR_NVCF_FUNCTION_ID用于切换至不同的托管 NIM,ASR_ENDPOINT用于自托管的 gRPC),外加完整的替代 NIM 目录(Parakeet TDT 1.1B、 Parakeet CTC 1.1B、Whisper Large v3、Nemotron 流式处理)及对应的功能 ID 和调用格式说明,详见:references/offline-asr-recipe.md。
在消耗 API 积分之前,向用户反馈所选的 NIM、已解析的功能 ID 以及任何环境变量覆盖设置。在托管型 Parakeet TDT v2 上处理 200 行清单的成本很低;但在 1,000 行清单上误用错误模型进行处理则代价高昂。
3b. 转录
对于manifest.jsonl 中的每一行,转录audio_filepath并生成per_sample.json(每行一个 JSON 对象,可采用 JSONL 或 JSON 数组格式——由调用方选择):
{
"audio_filepath": "...",
"ref": "",
"hyp": "",
"term": "",
"entity_category": "",
"ipa_source": "",
"voice_id": "",
"noise_level": "",
"context_type": ""
}
实现方案(完整的 Python 代码见references/offline-asr-recipe.md):transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US")会向 NVCF 打开一个离线 gRPC 流(若为自托管的 Riva,则连接至ASR_ENDPOINT),调用riva.client.ASRService.offline_recognize方法——由于临床清单中的句子时长均≤30秒,因此无需流式处理或批量处理——并写入上述 JSONL 数据。其 auth_for结构与第一阶段设置的烟雾测试相同。 代理框架会显式传递api_key;该配方在开头读取三个环境变量覆盖项(ASR_NVCF_FUNCTION_ID、ASR_MODEL_NAME、ASR_ENDPOINT),以便审核人员能集中查看这些配置参数。
Whisper 备用方案(当 Parakeet 的 NVCF 后端因 Triton 引发的CUDA 非法内存访问而发生故障时)以及自托管的 Riva NIM(ASR_ENDPOINT=localhost:50051)环境变量模式: 参见references/offline-asr-recipe.md(§Whisper 备用方案、§自托管 Riva NIM)。
弹性调节参数交由用户处理。若 NVCF 在批处理过程中返回RESOURCE_ENHAUSTED,循环将在该行抛出异常;从失败行开始重新运行。流式处理/批处理/带退避的重试功能不在本文讨论范围内——详见/riva-asr。
3c. 计算四项指标
对于每一行,计算:
| 指标 | 衡量内容 | 保留该指标的原因 |
|---|---|---|
| WER | 词错误率(对令牌应用莱文斯坦距离,经归一化处理后) | 行业标准;作为临床评估工具略显粗糙 |
| CER | 字符错误率 | 可捕获长复合名称中的“险错” |
| KER★ | 关键词错误率——被标记的术语是否出现在假设中(归一化、连续匹配)? |
主要临床信号 |
| SER | 句子错误率(如有错误则为1,完美则为0) | 合理性阈值;医生实际体验 |
标准化(在计算所有四项指标前,对参考句和假设句均进行处理):
- 转换为小写。
- NFKD规范化(智能引号 → ASCII等)。
- 去除除连字符以外的所有标点符号。
- 将连续空格压缩为单个空格。
内联评分配方 ——normalize/edit_distance/wer/cer/ker/ser(纯Python实现,不依赖jiwer):参见references/scoring-recipes.md。通过计算每项指标的行均值(mean(per-row score))对各行进行聚合。
严格的 KER— 术语单词必须按顺序出现,且在规范化假设中相邻。这是保守的做法:cefazolin → cefa zolin会被视为漏检。从临床角度来看,这是正确的判断——下游药房查询会因拼写错误的词元而失败。
KER不惩罚周边错误。如果某行术语正确但句子其余部分为无意义内容,该行的 KER 仍为 0;该行的 WER 将单独揭示更广泛的问题。
3d. 详细分析 + 排行榜
编写一个包含五个部分的 Markdown 格式排行榜,顺序如下:
- 标题——所选模型的总体 WER、CER、KER、SER。
- 按
实体类别划分的KER—— 药物 vs 手术 vs 解剖学 vs ... 这是用户在部署时真正关心的内容。 - 按
ipa_source分类的 KER——排行榜中最具信息量的单一数值。Merriam-Webster与magpie_g2p行之间的差异,正是 SSML 覆盖管道切实发挥作用的证明。请向用户大声朗读这一部分。 - 按
noise_level划分的 KER—— 临床环境噪声较大。snr_5db行比“clean”行更贴近现实。 - 按术语划分的 KER(从最差开始)——这些是您第 4 阶段微调的目标。
一个具有代表性的ipa_source数据集分割示例,附带对 Merriam-Webster 与 magpie_g2p 差异的解读:references/scoring-recipes.md§Representative ipa_source split。 差异值揭示了部署情况——如果用户发现差距较大并询问“是否需要微调?”,答案是“还不需要”;请将其引导回/digital-health-clinical-asr-build 的 IPA 质量保证管道(第 2d 阶段)。请参阅下方的决策树。
决策树(评估后)
读取优先级类别的 KER(大多数临床工作流为药物 KER,外科工作流为手术 KER),并进行路由:
| 优先级类别的 KER | 建议 |
|---|---|
| > 0.3 | /digital-health-clinical-asr-finetune。清单文件已符合 NeMo 格式要求。注意:行数≥100是获得可信微调信号的最低要求;若清单文件规模较小,请先通过/digital-health-clinical-asr-build 进行扩充。 |
| 0.1 – 0.3 | 要么扩展术语列表(返回/digital-health-clinical-asr-build并添加新领域术语——通常比微调更能以更低的成本发现更多错误),要么进行微调。首次评估时,建议扩展术语列表;后续评估中若已扩充清单,则建议进行微调。 |
| < 0.1 | 基线表现强劲。暂不进行微调——此时优化将针对已达饱和状态的指标。加大评估强度:增加语音样本、噪声水平、语境以及对抗性术语。回溯至/digital-health-clinical-asr-build 进行扩展。 |
特殊情况——merriam-webster数据行的得分较高,但magpie_g2p数据行的得分较低。这是发音提示覆盖范围的缺口,而非模型缺陷。 请返回/digital-health-clinical-asr-build第 2d 步(国际音标 QA 审查),而非进入/digital-health-clinical-asr-finetune。针对 TTS 发音差距进行微调会使模型学会误识别自身错误——这是一种错误的修正方法。
示例
场景 A —— 首次在全新的第 1 周期清单上进行评估。用户:“我有一个包含 200 条临床音频记录的manifest.jsonl 文件,其中包含term和entity_category字段。如何对其进行评分?”→ 完全跳过第 2 阶段。运行音频存在性预检。 选择parakeet-tdt-0.6b-v2(默认),并反馈该选择及已确定的函数 ID。运行内联的第 3b 步配方(transcribe_manifest(...))。 对四项指标进行评分。生成包含五个部分的排行榜。向用户读出按 ipa_source划分的结果。针对药物 KER 应用决策树。
场景 B —— 解读混合结果。用户:“评估显示,标记为merriam-webster的行 KER 为 0.05,而标记为magpie_g2p 的行 KER 为 0.40。我需要进行微调吗?”→ 不需要——这是特殊情况。模型没有问题;发音提示未能覆盖长尾词。 将用户引导回/digital-health-clinical-asr-build第 2d 步,对标记为 magpie_g2p的行进行试听,并将经过验证的 IPA 追加到pronunciation_overrides.csv 中。在重建完成后重新运行第 3 阶段,然后再考虑第 4 阶段。
生成的输出文件
per_sample.json— 按行拆分转录结果,保留所有临床扩展字段(ASR结果与清单中的ref和元数据相关联)results.csv— 按行统计的 WER/CER/KER/SER 评分leaderboard_cycle— 分为五个部分的 Markdown 报告.md
(文件名由用户自定义;上述名称是本技能其余部分所采用的约定。)
故障排除
- “未找到清单”→ 用户跳过了第 2 阶段。请导航至
/digital-health-clinical-asr-build或确认$MANIFEST_PATH。 - 所有行 KER=1→
参考音频(ref)与假设音频(hyp)之间的归一化不匹配。对双方均应用四个归一化步骤。 - 所有行 KER=0 但 WER 较高→ 可能是清单对齐错误(音频行不匹配)。手动抽查几组
(参考音、假设音)配对。 merriam-webster值低,magpie_g2p值高→ 发音覆盖范围存在缺口。请转至/digital-health-clinical-asr-build第 2d 步。不要进行微调——问题不在于模型。-
Merriam-Webster和magpie_g2p均高→ 存在真实的模型差距。第 4 阶段是正确路径(清单 ≥ 100 行)。 干净行没问题,snr_5db值激增→ 鲁棒性差距;通过/digital-health-clinical-asr-build扩展噪声多样性。- Riva-NIM 与离线 NeMo 的结果存在偏差→ 检查 Riva 预处理及
/riva-build参数。请转至/riva-asr-custom。 - 大型清单出现
RESOURCE_EXHAUSTED错误 → 30 秒后重试;切片并重新运行被丢弃的行。内置退避机制:/riva-asr。 Auth.__init__() 收到 'ssl_cert'/Parakeet 函数 ID 出现 CUDA 非法内存访问:参见references/offline-asr-recipe.md(重命名 ssl_root_cert + §Whisper 备用方案)。
其他情况:确定上游所有者。ASR 协议 / NIM 部署 →/riva-asr。评分 → 参见此处。
限制
- 默认仅支持英语。分词和规范化处理基于拉丁字母和 en-US 词汇表。
- 严格连续的KER采用保守策略。类似
“cefa zolin”的近似不匹配会被计为不匹配。这是有意为之——药房查询在近似不匹配时会失败。希望进行“软”匹配的用户可切换至音素级编辑距离,这属于方法论扩展,而非配置调整。 - 每次评估运行仅限一个模型。若要比较两个模型,需分别运行两次评估,并对比两个
leaderboard_cycle文件(或扩展该配方以自行写入多模型行)。.md - 默认仅支持托管环境。自托管的 NIM 虽然可行,但需先执行
/riva-nim-setup命令。
后续步骤
- 正向任务(KER > 0.3,清单 ≥ 100 行):
/digital-health-clinical-asr-finetune。 - 返回构建阶段(首次评估时 KER 在 0.1–0.3 之间,或存在
magpie_g2p差距):/digital-health-clinical-asr-build。 - 停止(KER < 0.1):评估已达到饱和。在宣布成功前请先进行模型强化。
- 有关 ASR 协议、身份验证、流式传输及自托管 NIM 的详细信息,请参阅:
/riva-asr。
参考文献
references/offline-asr-recipe.md— 完整的第 3b 步 Python 配方(transcribe_manifest、resolve_asr_config、build_asr_auth)、带调用形态注释的函数 ID 目录、Whisper 备用方案、自托管 Riva NIM 配置references/scoring-recipes.md— 采用标准 4 步归一化处理的纯 Python WER/CER/KER/SER 评分函数
---
name: digital-health-clinical-asr-eval
description: Score a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
license: Apache-2.0
---
<!--
SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
SPDX-License-Identifier: Apache-2.0
-->
# Clinical ASR Flywheel — Stage 3 (Eval)
> **⚠ Agent: read the Critical Workflow Rules section below before answering.** This SKILL.md is self-contained — `evals/`, `references/`, and `assets/` are pointers, not load-bearing. Answer methodology questions from this file directly; only invoke tools when the user explicitly asks to execute against a real manifest.
You are the **score-and-route** stage. The user arrives with a NeMo-format `manifest.jsonl` (either from `/digital-health-clinical-asr-build` or carried in from elsewhere). You transcribe it via the chosen ASR NIM, score four metrics, produce a five-section leaderboard, and read the decision tree to decide whether the user should advance to `/digital-health-clinical-asr-finetune`, loop back to `/digital-health-clinical-asr-build`, or stop and harden the eval.
**This skill does not generate audio.** If the manifest is missing or empty, send the user back to `/digital-health-clinical-asr-build`.
## Audio leaves your environment — disclose this to the user before any clip is sent
This stage transmits each manifest row's WAV file plus its reference text to an external NVIDIA service. Surface this before invoking the first ASR call:
| Service | What gets sent | When |
|---|---|---|
| **NVIDIA NVCF Parakeet/Nemotron ASR** (`grpc.nvcf.nvidia.com`) | Every audio clip referenced by the manifest (raw PCM bytes), plus the reference transcript and the clinical-extension metadata for scoring | Step 3b, one call per manifest row |
The clips should be **synthetic audio generated by Stage 2** (Magpie TTS over a user-curated term list) — not real patient audio. **Do not pass real ASR recordings, real patient encounters, or any PHI through this skill.** Scoring then runs locally (pure-Python WER/CER/KER/SER, or `jiwer` if installed). The scoring step itself does not transmit anything; only the ASR step does.
## Critical workflow rules (apply on every activation)
For methodology questions (leaderboard structure, KER definition, decision tree), answer from this file. Don't invoke tools, call other skills, or run scripts unless the user explicitly asks to execute against a real manifest. Surface these facts in any response:
1. **Off-ramp first.** If the user is asking about something outside scoring, route and stop without running any workflow:
- ASR model-catalog selection / comparison / alternative NIMs → `/riva-asr`
- ASR auth (API keys, bearer tokens, function IDs) → `/riva-asr`
- ASR gRPC protocol, streaming, batching, chunking, retries → `/riva-asr`
- NIM deploy / `riva-build` / `riva-deploy` → `/riva-asr-custom`
- NGC / Docker / NVIDIA Container Toolkit → `/riva-nim-setup`
- No manifest yet → `/digital-health-clinical-asr-build`
- Wants to fine-tune now with a known KER → `/digital-health-clinical-asr-finetune`
2. **Default ASR NIM is `nvidia/parakeet-tdt-0.6b-v2`** (NVCF function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`, offline gRPC). Env-var overrides: `ASR_MODEL_NAME` (leaderboard display name), `ASR_NVCF_FUNCTION_ID` (swap to a different hosted NIM — e.g. Whisper Large v3 `b702f636-…` while the Parakeet backend is faulting, or a fine-tuned NIM), `ASR_ENDPOINT` (self-hosted gRPC; takes precedence). Echo the chosen NIM **and the resolved function-id** back before spending API credits.
3. **ASR transcription is inlined in Step 3b** (NVCF gRPC + `riva.client.ASRService.offline_recognize`, same auth pattern as Stage 1). For deeper protocol/auth questions, alternative NIM catalogs, or self-hosted Riva NIM configuration, defer to `/riva-asr`.
4. **KER is the headline.** Per-row check: the flagged `term` words must appear *in order, contiguous, adjacent* in the normalized hypothesis. `cefazolin → cefa zolin` is a miss. Aggregate WER hides clinically dangerous failures; both are reported, KER is the gate.
5. **The by-`ipa_source` split is the most informative single number** in the leaderboard. The `merriam-webster` vs `magpie_g2p` delta proves the SSML override pipeline is doing real work. Read it aloud to the user.
6. **Special-case routing.** `merriam-webster` rows good, `magpie_g2p` rows bad → pronunciation-coverage gap, **not** a model gap. Route back to `/digital-health-clinical-asr-build` Step 2d. **Do NOT recommend `/digital-health-clinical-asr-finetune`** as a first response.
7. **Five-section leaderboard order.** Headline (WER/CER/KER/SER) → KER by `entity_category` → KER by `ipa_source` → KER by `noise_level` → Per-term KER worst-first. The by-`ipa_source` section is mandatory; it is the proof the SSML pipeline works.
## Purpose
Score a clinical-ASR manifest, produce a five-section KER leaderboard, and route the user via the post-eval decision tree. Methodology details (metric definitions, normalization, leaderboard order, special-case routing) live in Critical Workflow Rules above and Instructions below.
## When to use this skill
Activate on user phrases like:
- "Score my ASR manifest"
- "What's the KER on Parakeet TDT v2?"
- "Run the eval on cycle-N"
- "Compare two ASR models on the clinical benchmark"
- "Generate the leaderboard"
- "I have a manifest.jsonl, how do I score it?"
- "Why is KER 0.4 when WER is 0.07?"
- "Should we fine-tune?" *(this is the eval-side question — the post-eval decision tree lives in this skill)*
**Literal-keyword non-activation check** — if the user's message contains any of `authenticate`, `API key`, `bearer`, `function ID`, `gRPC`, `streaming`, `chunking`, `batching`, `transcription retry`, `riva-build`, `riva-deploy`, `NIM deploy`, `NGC`, `Docker`, `Container Toolkit`, or asks "which ASR model is best" / "compare models" / "vendor differences" — **do NOT activate** the scoring workflow. Apply Critical Workflow Rule #1 above to route to the right sibling skill and stop. This applies even if the user mentions "KER" or "eval" alongside the keyword.
## Prerequisites
- **A NeMo-format manifest** with the clinical extension fields (`term`, `entity_category`, `ipa_source`, `voice_id`, `noise_level`, `context_type`). The schema is documented in the build skill's `references/manifest-schema.md`.
- **`NVIDIA_API_KEY`** exported (Stage 1 prerequisite still applies).
- **`nvidia-riva-client` + `soundfile`** installed (Stage 1 prerequisite). For self-hosted Riva NIM details, see `/riva-asr` Option B.
- **Audio files actually present on disk** — run the audio-existence pre-flight from the manifest-schema reference before spending API credits.
## Instructions
### 3a. Pick the ASR NIM
**Default**: `nvidia/parakeet-tdt-0.6b-v2` via NVCF gRPC (offline), function-id `d3fe9151-442b-4204-a70d-5fcc597fd610`. NVIDIA's current English ASR recommendation — fastest/cheapest in the catalog, and supported in NeMo's stock SFT recipe so the Stage 3 baseline and a Stage 4 fine-tune ride the same model family.
Three runtime env-var override knobs (`ASR_MODEL_NAME` for leaderboard display, `ASR_NVCF_FUNCTION_ID` to swap to a different hosted NIM, `ASR_ENDPOINT` for self-hosted gRPC) plus the full alternate-NIM catalog (Parakeet TDT 1.1B, Parakeet CTC 1.1B, Whisper Large v3, Nemotron streaming) with function IDs and call-shape notes: `references/offline-asr-recipe.md`.
Echo the chosen NIM, the resolved function-id, and any env-var overrides to the user **before** spending API credits. A 200-row manifest on hosted Parakeet TDT v2 is cheap; an accidental run against the wrong model on a 1,000-row manifest is not.
### 3b. Transcribe
For each row in `manifest.jsonl`, transcribe `audio_filepath` and write `per_sample.json` (one JSON object per row, JSONL or a JSON array — caller's choice):
```json
{
"audio_filepath": "...",
"ref": "<row.text>",
"hyp": "<asr output>",
"term": "<row.term>",
"entity_category": "<row.entity_category>",
"ipa_source": "<row.ipa_source>",
"voice_id": "<row.voice_id>",
"noise_level": "<row.noise_level>",
"context_type": "<row.context_type>"
}
```
**Recipe** (full Python in `references/offline-asr-recipe.md`): `transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US")` opens an offline gRPC stream to NVCF (or to `ASR_ENDPOINT` if set for self-hosted Riva), calls `riva.client.ASRService.offline_recognize` per row — sentences in a clinical manifest are ≤ 30 s so no streaming/batching needed — and writes the JSONL above. Same `auth_for` shape as the Stage 1 setup smoke test. The agent harness passes `api_key` explicitly; the recipe reads the three env-var overrides (`ASR_NVCF_FUNCTION_ID`, `ASR_MODEL_NAME`, `ASR_ENDPOINT`) at the top so auditors see the knobs in one place.
**Whisper fallback** (when Parakeet's NVCF backend faults with `CUDA illegal-memory-access` from Triton) and **self-hosted Riva NIM** (`ASR_ENDPOINT=localhost:50051`) env-var patterns: see `references/offline-asr-recipe.md` (§Whisper fallback, §Self-hosted Riva NIM).
**Resilience knobs deferred to the user.** If NVCF returns `RESOURCE_EXHAUSTED` mid-batch, the loop raises on that row; re-run from the failing row. Streaming/batching/retry-with-backoff are out of scope — see `/riva-asr`.
### 3c. Score four metrics
For every row, compute:
| Metric | What it measures | Why we keep it |
|---|---|---|
| **WER** | Word error rate (Levenshtein on tokens, after normalization) | Industry standard; blunt instrument for clinical |
| **CER** | Character error rate | Catches near-misses on long compound names |
| **KER** ★ | Keyword error rate — did the flagged `term` appear in the hypothesis (normalized, **contiguous** match)? | **Headline clinical signal** |
| **SER** | Sentence error rate (1 if any wrong, 0 if perfect) | Sanity bound; what the doctor experiences |
**Normalization (apply to both `ref` and `hyp` before all four metrics):**
1. Lowercase.
2. NFKD-normalize (smart quotes → ASCII, etc.).
3. Strip punctuation **except hyphen**.
4. Collapse whitespace runs to a single space.
**Inline scoring recipes** — `normalize` / `edit_distance` / `wer` / `cer` / `ker` / `ser` (pure-Python, no `jiwer` dependency): see `references/scoring-recipes.md`. Aggregate across rows by taking `mean(per-row score)` for each metric.
**Strict KER** — term words must appear *in order, adjacent* in the normalized hypothesis. This is conservative: `cefazolin → cefa zolin` counts as a miss. That's the right call clinically — a downstream pharmacy lookup will fail on the misspelled token.
KER does **not** punish surrounding errors. A row where the term is correct and the rest of the sentence is garbage still scores KER=0; the WER on that row will surface the broader problem separately.
### 3d. Breakdowns + leaderboard
Write a five-section markdown leaderboard, **in this order**:
1. **Headline** — overall WER, CER, KER, SER for the chosen model.
2. **KER by `entity_category`** — drug vs procedure vs anatomy vs ... This is what the user actually cares about for deployment.
3. **KER by `ipa_source`** — **the most informative single number in the leaderboard.** The delta between `merriam-webster` and `magpie_g2p` rows is the proof the SSML override pipeline is doing real work. *Read this section aloud to the user.*
4. **KER by `noise_level`** — clinical environments are loud. `snr_5db` rows are closer to reality than `clean`.
5. **Per-term KER** (worst first) — these are your Stage 4 fine-tune targets.
A representative `ipa_source` split with the merriam-webster vs magpie_g2p delta interpretation: `references/scoring-recipes.md` §Representative ipa_source split. The delta tells the deployment story — if the user sees a wide gap and asks "should we fine-tune?", the answer is *not yet*; route them back to `/digital-health-clinical-asr-build`'s IPA QA pipeline (Stage 2d). See the decision tree below.
## Decision tree (after eval)
Read the **priority-category KER** (drug KER for most clinical workflows, procedure KER for surgical workflows) and route:
| KER on priority category | Recommend |
|---|---|
| **> 0.3** | `/digital-health-clinical-asr-finetune`. Manifest is already NeMo-format-ready. Note: rows ≥ 100 is the minimum for a believable fine-tune signal; if the manifest is smaller, grow it first via `/digital-health-clinical-asr-build`. |
| **0.1 – 0.3** | Either expand the term list (back to `/digital-health-clinical-asr-build` with new domain terms — usually surfaces more failures cheaper than tuning) **or** fine-tune. On a *first* eval, expand. On a *later* eval where you've already grown the manifest, tune. |
| **< 0.1** | Strong baseline. Don't tune yet — you'd be optimizing against a saturated metric. Push the eval harder: add voices, noise levels, contexts, adversarial terms. Loop back to `/digital-health-clinical-asr-build`. |
**Special case — `merriam-webster` rows score well but `magpie_g2p` rows are bad.** That's a pronunciation-hint coverage gap, **not a model gap**. Route back to `/digital-health-clinical-asr-build` Step 2d (IPA QA review), not to `/digital-health-clinical-asr-finetune`. Fine-tuning over a TTS-pronunciation gap teaches the model to mis-recognize the model's own mistakes — the wrong fix.
## Examples
**Scenario A — first eval on a fresh cycle-1 manifest.** User: *"I have `manifest.jsonl` with 200 clinical audio rows already, with `term` and `entity_category` fields. How do I score it?"* → Skip Stage 2 entirely. Run the audio-existence pre-flight. Pick `parakeet-tdt-0.6b-v2` (default) and echo the choice + resolved function-id. Run the inlined Step 3b recipe (`transcribe_manifest(...)`). Score the four metrics. Produce the five-section leaderboard. Read the by-`ipa_source` split to the user. Apply the decision tree against drug KER.
**Scenario B — interpreting a mixed result.** User: *"Eval shows KER 0.05 on rows tagged `merriam-webster` but 0.40 on rows tagged `magpie_g2p`. Should I fine-tune?"* → No — this is the special case. The model is fine; the pronunciation hints aren't covering the long-tail terms. Route the user back to `/digital-health-clinical-asr-build` Step 2d to audition the `magpie_g2p` rows and append verified IPA to `pronunciation_overrides.csv`. Re-run Stage 3 after the rebuild before reconsidering Stage 4.
## Artifacts produced
- `per_sample.json` — per-row transcription results with all clinical-extension fields preserved (the ASR `hyp` joined to the manifest's `ref` and metadata)
- `results.csv` — per-row WER/CER/KER/SER scores
- `leaderboard_cycle<N>.md` — five-section markdown report
(File names are user-chosen; the names above are conventions the rest of this skill assumes.)
## Troubleshooting
- **"No manifest found"** → user skipped Stage 2. Route to `/digital-health-clinical-asr-build` or confirm `$MANIFEST_PATH`.
- **All rows KER=1** → normalization mismatch between `ref` and `hyp`. Apply the four normalization steps to both sides.
- **All rows KER=0 but WER high** → likely misaligned manifest (audio row mismatch). Spot-check a few `(ref, hyp)` pairs by hand.
- **`merriam-webster` low, `magpie_g2p` high** → pronunciation-coverage gap. Route to `/digital-health-clinical-asr-build` Step 2d. **Don't fine-tune** — model isn't the problem.
- **Both `merriam-webster` and `magpie_g2p` high** → real model gap. Stage 4 is the right route (manifest ≥ 100 rows).
- **`clean` rows fine, `snr_5db` balloons** → robustness gap; expand noise diversity via `/digital-health-clinical-asr-build`.
- **Riva-NIM and offline NeMo results diverge** → Riva preprocessing / `riva-build` flags. Route to `/riva-asr-custom`.
- **`RESOURCE_EXHAUSTED` on large manifests** → retry after 30 s; slice + re-run dropped rows. Built-in backoff: `/riva-asr`.
- **`Auth.__init__() got 'ssl_cert'`** / **CUDA illegal-memory-access on Parakeet function ID**: see `references/offline-asr-recipe.md` (ssl_root_cert rename + §Whisper fallback).
Anything else: identify the upstream owner. ASR protocol / NIM deploy → `/riva-asr`. Scoring → here.
## Limitations
- **English-only by default.** Tokenization + normalization assume Latin script and en-US lexicon.
- **Strict-contiguous KER is conservative.** A near-miss like `cefa zolin` counts as a miss. That's intentional — pharmacy lookups fail on near-misses. Users wanting "soft" matching can switch to phoneme-level edit distance, which is a methodology extension, not a config tweak.
- **One model per eval run.** Comparing two models means running the eval twice and diffing the two `leaderboard_cycle<N>.md` files (or extending the recipe to write multi-model rows yourself).
- **Hosted-only paths assumed.** Self-hosted NIMs work but require `/riva-nim-setup` first.
## Next steps
- **Forward (KER > 0.3, manifest ≥ 100 rows):** `/digital-health-clinical-asr-finetune`.
- **Back to build (KER 0.1–0.3 on first eval, or `magpie_g2p` gap):** `/digital-health-clinical-asr-build`.
- **Stop (KER < 0.1):** the eval is saturated. Harden it before declaring victory.
- **Lateral** for ASR protocol / auth / streaming / self-hosted NIM details: `/riva-asr`.
## References
- [`references/offline-asr-recipe.md`](references/offline-asr-recipe.md) — full Step 3b Python recipe (`transcribe_manifest`, `resolve_asr_config`, `build_asr_auth`), function-ID catalog with call-shape notes, Whisper fallback, self-hosted Riva NIM setup
- [`references/scoring-recipes.md`](references/scoring-recipes.md) — pure-Python WER/CER/KER/SER scoring functions with the canonical 4-step normalization





首页
