选项

mle-workflow

affaan-m/ECC affaan-m/ECC

利用数据契约、可重复的训练流程、可衡量的质量门槛、可部署的构建产物以及运维监控,将模型原型转化为生产级的机器学习系统。

...展开全部
0
更新时间 2026-10-01

机器学习工程工作流

运用此技能,将模型工作转化为具备清晰数据契约、可重复训练、可衡量质量门槛、可部署构建产物及运维监控的生产级机器学习系统。

何时启用

  • 规划或审查生产环境中的机器学习功能、模型更新、排序系统、推荐系统、分类器、嵌入工作流或预测管道时
  • 将笔记本代码转换为可重用的训练、评估、批量推理或在线推理管道
  • 设计模型推广标准、离线/在线评估、实验追踪或回滚路径
  • 调试由数据漂移、标签泄漏、过期特征、构建产物不匹配或训练与服务逻辑不一致导致的故障
  • 添加模型监控、金丝雀发布、影子流量或部署后质量检查

范围校准

仅使用适合您当前系统的“车道”。此技能适用于排序、搜索、推荐、分类、预测、嵌入、大型语言模型(LLM)工作流、异常检测和批量分析,但不应将单一架构强加于所有场景。

  • 不要假设每个模型都具备监督标签、在线服务、特征存储库、PyTorch、GPU、人工审核、A/B 测试或实时反馈。
  • 当数据契约、基线、评估脚本和回滚说明足以使变更可审核时,请勿添加繁重的 MLOps 机制。
  • 当项目缺乏标签、结果延迟、切片定义、生产流量或监控责任归属时,请明确说明这些假设。
  • 将示例视为可互换的框架。用项目特有的等效方案替换指标、服务模式、数据存储和部署机制。

相关技能

  • python-patterns 以及 python-testing 用于 Python 实现和 pytest 覆盖率
  • pytorch-patterns 用于深度学习模型、数据加载器、设备处理和训练循环
  • eval-harness 以及 ai-regression-testing 用于推广门和代理辅助回归检查
  • database-migrations, postgres-patterns,以及 clickhouse-io 用于数据存储和分析界面
  • deployment-patterns, docker-patterns,以及 security-review 用于服务、密钥、容器和生产环境加固

复用软件工程(SWE)工作流

请勿将机器学习工程(MLE)与软件工程割裂开来。大多数 ECC 软件工程工作流可直接应用于机器学习系统,且通常具有更严格的故障模式:

推荐的 minimal --with capability:machine-learning 安装方案会使核心代理接口与该技能并存。对于仅使用技能或受限于代理的测试框架,请将 skill:mle-workflow 与 agent:mle-reviewer ,前提是目标系统支持代理。

SWE 接口 MLE 的使用
product-capability / architecture-decision-records 将模型工作转化为明确的产品契约,并记录不可逆转的数据、模型和部署决策
repo-scan / codebase-onboarding / code-tour 在引入并行机器学习栈之前,先找出现有的训练、特征工程、服务、评估和监控路径
plan / feature-dev 将模型变更界定为产品功能,并包含数据、评估、服务和回滚阶段
tdd-workflow / python-testing 在实施前测试特征转换、拆分逻辑、指标计算、构建产物加载及推理架构
code-reviewer / mle-reviewer 审查代码质量,并关注机器学习特有的数据泄露、可重现性、推广及监控风险
build-fix / pr-test-analyzer 诊断持续集成(CI)故障、不稳定的评估、缺失的测试 fixture 以及特定于环境的模型或依赖项故障
quality-gate / test-coverage 要求针对转换、指标、推理契约、上线门槛及回滚行为提供自动化验证依据
eval-harness / verification-loop 将离线指标、切片检查、延迟预算和回滚演练转化为可重复的门控机制
ai-regression-testing 将每个生产环境中的缺陷都视为回归问题:缺失特征、过期标签、损坏构建产物、模式漂移或服务不匹配
api-design / backend-patterns 设计预测 API、批处理任务、幂等重训练端点和响应信封
database-migrations / postgres-patterns / clickhouse-io 版本标签、特征快照、预测日志、实验指标及漂移分析
deployment-patterns / docker-patterns 打包可重现的训练和服役镜像,并包含健康检查、资源限制和回滚机制
canary-watch / dashboard-builder 通过模型版本、切片、漂移、延迟、成本和延迟标签仪表盘,让部署健康状况一目了然
security-review / security-scan 检查模型构建产物、笔记本、提示词、数据集和日志中是否存在机密信息、个人身份信息(PII)、不安全的反序列化以及供应链风险
e2e-testing / browser-qa / accessibility 测试使用预测结果的关键产品流程,包括可解释性和备用 UI 状态
benchmark / performance-optimizer 衡量吞吐量、p95延迟、内存、GPU利用率以及每次预测或重新训练的成本
cost-aware-llm-pipeline / token-budget-advisor 根据质量、延迟和预算路由 LLM/嵌入式工作负载,而非默认使用最大模型
documentation-lookup / search-first 在编写代码前,验证当前用于模型服务、特征存储、向量数据库和评估工具的库的行为
git-workflow / github-ops / opensource-pipeline 将 MLE 变更打包提交审核,明确范围,排除生成的构建产物,并提供可重现的测试证据
strategic-compact / dmux-workflows 将耗时的机器学习工作拆分为并行任务:数据契约、评估框架、服务路径、监控和文档

十项 MLE 任务模拟

在规划或审查MLE工作时,将这些模拟用作覆盖率检查。完善的MLE工作流应将每项任务分解为明确的契约、可重用的软件工程接口、自动化证据以及可审查的交付成果。

ID 常见MLE任务 简化的ECC路径 所需输出 涵盖的管道通道
MLE-01 构建一个涵盖模糊预测、排序、推荐、分类、嵌入或预测能力的框架 product-capability, plan, architecture-decision-records, mle-workflow 迭代 明确关键关注方、决策负责人、成功指标、不可接受的错误、假设、约束条件以及首次实验 产品契约、利益相关方损失、风险、上线
MLE-02 定义指标目标、标签、数据源及容错预算 repo-scan, database-reviewer, database-migrations, postgres-patterns, clickhouse-io 包含实体粒度、标签时间、标签置信度、特征时间、特定时间点连接、拆分策略和数据集快照的数据与指标契约 数据契约、指标设计、数据泄漏、可重现性
MLE-03 在增加复杂度之前构建基线模型和评分路径 tdd-workflow, python-testing, python-patterns, code-reviewer 包含混淆矩阵、校准说明、延迟/成本估算、已知弱点,以及针对评分形态和确定性的测试的基线评分器 基线、评分、测试、服务一致性
MLE-04 基于对结果差异成因的假设生成特征 python-patterns, pytorch-patterns, docker-patterns, deployment-patterns 特征规划与转换模块,涵盖信号源、缺失值、异常值、相关性、泄漏检查以及训练与生产端等效性 特征管道、数据泄漏、训练、伪影
MLE-05 在权衡取舍下调整阈值、配置和模型复杂度 eval-harness, ai-regression-testing, quality-gate, test-coverage 阈值/配置报告,用于比较精度、召回率、F1 值、AUC、校准、分组切片、延迟、成本、复杂度以及可接受的错误类别 评估、阈值、模型推广、回归
MLE-06 运行错误分析,将错误转化为下一次实验 eval-harness, ai-regression-testing, mle-reviewer, silent-failure-hunter 针对误报、漏报、标签模糊、过时特征、信号缺失及包含经验教训的错误追踪的错误聚类报告 错误分析、错误追踪、迭代、回归
MLE-07 打包模型工件以供批量或在线推理使用 api-design, backend-patterns, security-review, security-scan 带版本控制的模型构建包,包含预处理、配置、依赖约束、模式验证、安全加载及个人身份信息(PII)安全日志 构建产物、安全性、推理契约
MLE-08 发布支持反馈捕获的在线服务或批量评分 api-design, backend-patterns, e2e-testing, browser-qa, accessibility 预测端点或批处理作业,包含响应封装、超时、批处理、备用方案、模型版本、置信度、反馈日志记录以及产品流程测试 服务、批量推理、备用方案、用户工作流
MLE-09 通过影子流量、金丝雀测试、A/B 测试或回滚部署模型 canary-watch, dashboard-builder, verification-loop, performance-optimizer 发布计划命名、流量拆分、仪表盘、p95 延迟、成本、质量防护机制、回滚构建产物及回滚触发条件 部署、金丝雀测试、回滚
MLE-10 发布后对生产模型进行运维、调试和刷新 silent-failure-hunter, dashboard-builder, mle-reviewer, doc-updater, github-ops 包含漂移检查、延迟标签健康状况、警报负责人、运行手册更新、重新训练标准及 PR 证据的观察日志和刷新计划 监控、事件响应、重新训练

迭代摘要

在触碰模型代码之前,将工作压缩为一个可供审查的成果。该成果应简短到足以放入 PR 描述中,且精确到足以让另一位工程师对其权衡方案提出质疑。

Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:

该摘要相当于 MLE 领域中一份严谨的软件工程师设计说明。它能防止团队对无人信任的指标进行优化,避免添加无法解决真实错误模式的功能,或发布无法回滚的复杂代码。

决策脑

每当任务存在模糊性、影响重大或涉及大量指标时,请使用以下循环:

  1. 从决策出发,而非模型。明确指出会改变下游行为的行动。
  2. 明确谁会受影响以及原因。不同的利益相关者会因误报、漏报、延迟、计算成本、不透明性或错失机会而付出不同的代价。
  3. 将模糊性转化为假设。思考:什么样的信号能区分不同结果?什么样的证据能推翻该假设?以及应该设定怎样的简单基准,使其难以被超越。
  4. 在开发定制系统之前,先研究现有技术或相关已知问题。
  5. 使用以下标准对选项进行评分: (probability, confidence) x (cost, severity, importance, impact).
  6. 考虑对抗性行为、激励机制、选择性披露、分布偏移以及反馈循环。
  7. 优先选择能减少最重要错误的最简单变更。简单并非懒惰;而是在保持迭代速度的同时,将重大失误降至最低的一种方式。
  8. 记录决策、证据、反论点以及下一步可逆操作。

指标与错误经济学

从失败成本而非惯例中选择指标:

  • 尽早使用混淆矩阵,以便团队能够讨论具体的假阳性与假阴性,而非抽象的准确率。
  • 当错误的阳性决策成本占主导地位时,应优先考虑精确度。
  • 当漏检的代价占主导地位时,应优先考虑召回率。
  • 仅当精度与召回率的权衡真正平衡且可解释时,才使用 F1 值。
  • 当排序质量比单一阈值更重要时,应使用AUC或排序指标。
  • 将延迟、吞吐量、内存和成本作为首要指标进行追踪,因为它们决定了模型复杂度的可行范围。
  • 在庆祝离线性能提升之前,应先与基线模型和当前生产环境中的模型进行对比。
  • 将现实世界的反馈信号视为带有偏差、延迟和覆盖缺口的延迟标签;未经分析时,切勿将其视为真实标签。

每次指标选择都应明确说明:它会降低哪种错误的成本、增加哪种错误发生的概率,以及由谁来承担该成本。

数据与特征假设

特征应源于分离理论:

  • 文本、分类型字段、数值历史数据、图关系、时效性、频率以及聚合数据均是潜在的信号类别,而非自动生成的特征。
  • 对于每个特征族,应说明其为何能区分结果,以及可能如何泄露未来信息。
  • 对于存在标签噪声的情况,应考虑裁决、标签置信度、软目标或置信度加权。
  • 对于类不平衡问题,应比较加权损失、重采样、阈值移动和校准决策规则。
  • 对于缺失值,需判断其缺失是具有信息价值、可插补,还是应予以忽略。
  • 对于异常值,需决定是进行裁剪、分桶、调查,还是将其保留为罕见但重要的信号。
  • 对于相关特征,需检查它们是否冗余、不稳定,或是不可获得的未来状态的代理变量。

除非错误分析表明基线模型的失败是由于额外信号或容量能够合理解决的原因所致,否则不要增加模型的复杂度。

误差分析循环

每次基线模型测试、训练运行、阈值调整或配置变更后:

  1. 将错误分为误报、漏报、未判定、低置信度案例和系统故障。
  2. 根据共同特征对错误进行聚类:语言、实体类型、来源、时间、地理位置、设备、稀疏性、时效性、特征新鲜度、标签来源或模型版本。
  3. 将模型错误与数据缺陷、标签歧义、产品歧义、监测缺失以及服务端不匹配区分开来。
  4. 将每个主要错误聚类追溯至以下四种改进措施之一:优化标签、优化特征、优化阈值/配置,或优化产品备用方案。
  5. 将每个重要错误保留为回归测试、评估切片、仪表盘面板或运行手册条目。
  6. 将下一次迭代设计为可证伪的实验,而非模糊的“改进模型”任务。

最强大的最大似然估计(MLE)循环并非“训练 → 指标 → 部署”,而是“错误 → 聚类 → 假设 → 实验 → 证据 → 更简化的系统”。

观察日志

在代码、拉取请求、实验报告或运行手册旁保留一份简洁的决策与证据记录:

Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:

利用该日志使模型工作具有累积性。目标是让每次迭代都能使下一次决策更轻松,而不仅仅是产出另一个成果。

核心工作流

1. 定义预测契约

在编写模型代码之前,先明确产品层面的协议:

  • 预测目标和决策负责人
  • 输入实体、输出模式、置信度/校准字段以及允许的延迟
  • 批处理、在线、流式或混合服务模式
  • 当模型、特征库或依赖项不可用时的备用行为
  • 针对高影响决策的人工审核或覆盖路径
  • 针对输入、预测和标签的隐私、保留及审计要求

请勿将“改进模型”作为要求。应将模型与可观察的产品行为及可衡量的验收标准挂钩。

2. 锁定数据契约

每个机器学习任务都需要明确的数据契约:

  • 实体粒度和主键
  • 标签定义、标签时间戳及标签可用性延迟
  • 特征时间戳、新鲜度服务水平协议(SLA)及特定时间点连接规则
  • 训练集、验证集、测试集和回测集的划分策略
  • 必填字段、允许为空的字段、取值范围、类别及单位
  • 不得进入训练成果或日志的个人身份信息(PII)或敏感字段
  • 用于可重现性的数据集版本或快照 ID

首先防范数据泄露。如果某特征在预测时不可用,或通过未来信息进行关联,则应将其移除或转移至仅用于分析的路径。

3. 构建可重现的管道

训练代码应可由其他工程师运行,且不依赖隐藏的笔记本状态:

  • 对所有超参数和路径使用类型化配置文件或数据类
  • 固定包和模型依赖项
  • 设置随机种子,并记录任何非确定性的 GPU 行为
  • 记录数据集版本、代码 SHA、配置哈希、指标以及构建产物 URI
  • 将预处理逻辑与模型构建产物一同保存,而非单独保存在笔记本中
  • 确保训练、评估和推理的转换操作保持共享,或由同一来源生成
  • 确保每个步骤均为幂等,以免重试操作导致构建产物或指标受损

优先使用不可变值和纯转换函数。在特征生成过程中,避免修改共享的数据框或全局配置。

import hashlib
from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class TrainingConfig:
    dataset_uri: str
    model_dir: Path
    seed: int
    learning_rate: float
    batch_size: int


def artifact_name(config: TrainingConfig, code_sha: str) -> str:
    config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
    config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
    return f"{code_sha[:12]}-{config_hash}"

4. 推广前进行评估

应在训练完成前声明推广标准:

  • 基线模型与当前生产模型的对比
  • 与产品行为对齐的核心指标
  • 针对延迟、校准、公平性切片、成本和误差集中度的防护指标
  • 针对重要用户群体、地区、设备、语言或数据源的切片指标
  • 当指标存在噪声时,应提供置信区间或重复运行的方差
  • 针对高影响模型,由人工审查失败案例
  • 明确的“禁止上线”阈值
PROMOTION_GATES = {
    "auc": ("min", 0.82),
    "calibration_error": ("max", 0.04),
    "p95_latency_ms": ("max", 80),
}


def assert_promotion_ready(metrics: dict[str, float]) -> None:
    missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
    if missing:
        raise ValueError(f"Model promotion metrics missing required gates: {missing}")

    failures = {
        name: value
        for name, (direction, threshold) in PROMOTION_GATES.items()
        for value in [metrics[name]]
        if (direction == "min" and value < threshold)
        or (direction == "max" and value > threshold)
    }
    if failures:
        raise ValueError(f"Model failed promotion gates: {failures}")

将离线指标用作门槛,而非保证。当模型改变产品行为时,在全面部署前应规划影子评估、金丝雀发布或 A/B 测试。

5. 打包以供部署

只有当服务契约可测试时,机器学习构建产物才算具备生产就绪性:

  • 模型构建包应包含版本号、训练数据引用、配置及预处理信息
  • 输入模式会拒绝无效、过时或超出范围的特征
  • 输出模式应包含模型版本,并在适用时包含置信度或解释字段
  • 服务路径设有超时、批处理、资源限制及备用行为
  • CPU/GPU 需求明确且经过测试
  • 预测日志避免包含个人身份信息(PII),并包含足够的标识符以供调试和标签关联
  • 集成测试涵盖缺失特征、过期特征、类型错误、空批次以及备用路径

切勿在未通过等价性测试验证的情况下,让仅用于训练的特征代码与服务端的特征代码产生差异。

6. 模型运维

模型监控需要同时关注系统指标和质量指标:

  • 可用性、错误率、超时率、队列深度以及 p50/p95/p99 延迟
  • 特征空值率、范围漂移、类别漂移和新鲜度漂移
  • 预测分布漂移和置信度分布漂移
  • 标签到达健康状况和延迟质量指标
  • 业务KPI防护阈值和回滚触发条件
  • 针对金丝雀发布和回滚的按版本仪表盘

每次部署都应制定回滚计划,其中需明确指定之前的构建产物、配置、数据依赖项以及流量切换机制。

审查清单

  • 预测契约明确且可测试
  • 数据契约定义了实体粒度、标签时间、特征时间以及快照/版本
  • 已根据预测时点的可用性核查了数据泄漏风险
  • 训练过程可基于代码、配置、数据版本和初始化种子进行重现
  • 指标与基线及当前生产模型进行对比
  • 针对高风险群体,已纳入切片指标和防护措施
  • 推广门控已实现自动化,并采用“失败即关闭”机制
  • 训练和服役转换过程共享或经过等价性测试
  • 模型构建产物包含版本、配置、数据集引用及预处理信息
  • 服务路径会验证输入,并具备超时、备用方案和回滚机制
  • 监控涵盖系统健康状况、特征漂移、预测漂移以及延迟标签
  • 敏感数据被排除在构建产物、日志、提示词和示例之外

反模式

  • 需要 Notebook 状态才能复现模型
  • 随机划分导致未来数据泄露到验证集或测试集中
  • 特征连接忽略了事件时间和标签的可用性
  • 离线指标有所提升,但关键切片却出现退步
  • 阈值在测试集上被反复调整
  • 训练预处理代码被手动复制到服务端代码中
  • 预测日志中缺少模型版本信息
  • 监控仅检查服务可用性,而不检查数据或预测质量
  • 回滚需要重新训练,而非切换到已知可用的构建版本

输出预期

使用此技能时,应返回具体的构建产物:数据契约、发布门、管道步骤、测试计划、部署计划或审查结果。应明确指出阻碍生产就绪的未知因素,而非用假设来填补空白。

在 GitHub 上查看
---
name: mle-workflow
description: Turn model work into a production ML system with data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.
---

# Machine Learning Engineering Workflow

Use this skill to turn model work into a production ML system with clear data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.

## When to Activate

- Planning or reviewing a production ML feature, model refresh, ranking system, recommender, classifier, embedding workflow, or forecasting pipeline
- Converting notebook code into a reusable training, evaluation, batch inference, or online inference pipeline
- Designing model promotion criteria, offline/online evals, experiment tracking, or rollback paths
- Debugging failures caused by data drift, label leakage, stale features, artifact mismatch, or inconsistent training and serving logic
- Adding model monitoring, canary rollout, shadow traffic, or post-deploy quality checks

## Scope Calibration

Use only the lanes that fit the system in front of you. This skill is useful for ranking, search, recommendations, classifiers, forecasting, embeddings, LLM workflows, anomaly detection, and batch analytics, but it should not force one architecture onto all of them.

- Do not assume every model has supervised labels, online serving, a feature store, PyTorch, GPUs, human review, A/B tests, or real-time feedback.
- Do not add heavyweight MLOps machinery when a data contract, baseline, eval script, and rollback note would make the change reviewable.
- Do make assumptions explicit when the project lacks labels, delayed outcomes, slice definitions, production traffic, or monitoring ownership.
- Treat examples as interchangeable scaffolds. Replace metrics, serving mode, data stores, and rollout mechanics with the project-native equivalents.

## Related Skills

- `python-patterns` and `python-testing` for Python implementation and pytest coverage
- `pytorch-patterns` for deep learning models, data loaders, device handling, and training loops
- `eval-harness` and `ai-regression-testing` for promotion gates and agent-assisted regression checks
- `database-migrations`, `postgres-patterns`, and `clickhouse-io` for data storage and analytics surfaces
- `deployment-patterns`, `docker-patterns`, and `security-review` for serving, secrets, containers, and production hardening

## Reuse the SWE Surface

Do not treat MLE as separate from software engineering. Most ECC SWE workflows apply directly to ML systems, often with stricter failure modes:

The recommended `minimal --with capability:machine-learning` install keeps the core agent surface available alongside this skill. For skill-only or agent-limited harnesses, pair `skill:mle-workflow` with `agent:mle-reviewer` where the target supports agents.

| SWE surface | MLE use |
|-------------|---------|
| `product-capability` / `architecture-decision-records` | Turn model work into explicit product contracts and record irreversible data, model, and rollout choices |
| `repo-scan` / `codebase-onboarding` / `code-tour` | Find existing training, feature, serving, eval, and monitoring paths before introducing a parallel ML stack |
| `plan` / `feature-dev` | Scope model changes as product capabilities with data, eval, serving, and rollback phases |
| `tdd-workflow` / `python-testing` | Test feature transforms, split logic, metric calculations, artifact loading, and inference schemas before implementation |
| `code-reviewer` / `mle-reviewer` | Review code quality plus ML-specific leakage, reproducibility, promotion, and monitoring risks |
| `build-fix` / `pr-test-analyzer` | Diagnose broken CI, flaky evals, missing fixtures, and environment-specific model or dependency failures |
| `quality-gate` / `test-coverage` | Require automated evidence for transforms, metrics, inference contracts, promotion gates, and rollback behavior |
| `eval-harness` / `verification-loop` | Turn offline metrics, slice checks, latency budgets, and rollback drills into repeatable gates |
| `ai-regression-testing` | Preserve every production bug as a regression: missing feature, stale label, bad artifact, schema drift, or serving mismatch |
| `api-design` / `backend-patterns` | Design prediction APIs, batch jobs, idempotent retraining endpoints, and response envelopes |
| `database-migrations` / `postgres-patterns` / `clickhouse-io` | Version labels, feature snapshots, prediction logs, experiment metrics, and drift analytics |
| `deployment-patterns` / `docker-patterns` | Package reproducible training and serving images with health checks, resource limits, and rollback |
| `canary-watch` / `dashboard-builder` | Make rollout health visible with model-version, slice, drift, latency, cost, and delayed-label dashboards |
| `security-review` / `security-scan` | Check model artifacts, notebooks, prompts, datasets, and logs for secrets, PII, unsafe deserialization, and supply-chain risk |
| `e2e-testing` / `browser-qa` / `accessibility` | Test critical product flows that consume predictions, including explainability and fallback UI states |
| `benchmark` / `performance-optimizer` | Measure throughput, p95 latency, memory, GPU utilization, and cost per prediction or retrain |
| `cost-aware-llm-pipeline` / `token-budget-advisor` | Route LLM/embedding workloads by quality, latency, and budget instead of defaulting to the largest model |
| `documentation-lookup` / `search-first` | Verify current library behavior for model serving, feature stores, vector DBs, and eval tooling before coding |
| `git-workflow` / `github-ops` / `opensource-pipeline` | Package MLE changes for review with crisp scope, generated artifacts excluded, and reproducible test evidence |
| `strategic-compact` / `dmux-workflows` | Split long ML work into parallel tracks: data contract, eval harness, serving path, monitoring, and docs |

## Ten MLE Task Simulations

Use these simulations as coverage checks when planning or reviewing MLE work. A strong MLE workflow should reduce each task to explicit contracts, reusable SWE surfaces, automated evidence, and a reviewable artifact.

| ID | Common MLE task | Streamlined ECC path | Required output | Pipeline lanes covered |
|----|-----------------|----------------------|-----------------|------------------------|
| MLE-01 | Frame an ambiguous prediction, ranking, recommender, classifier, embedding, or forecast capability | `product-capability`, `plan`, `architecture-decision-records`, `mle-workflow` | Iteration Compact naming who cares, decision owner, success metric, unacceptable mistakes, assumptions, constraints, and first experiment | product contract, stakeholder loss, risk, rollout |
| MLE-02 | Define metric goals, labels, data sources, and the mistake budget | `repo-scan`, `database-reviewer`, `database-migrations`, `postgres-patterns`, `clickhouse-io` | Data and metric contract with entity grain, label timing, label confidence, feature timing, point-in-time joins, split policy, and dataset snapshot | data contract, metric design, leakage, reproducibility |
| MLE-03 | Build a baseline model and scoring path before adding complexity | `tdd-workflow`, `python-testing`, `python-patterns`, `code-reviewer` | Baseline scorer with confusion matrix, calibration notes, latency/cost estimate, known weaknesses, and tests for score shape and determinism | baseline, scoring, testing, serving parity |
| MLE-04 | Generate features from hypotheses about what separates outcomes | `python-patterns`, `pytorch-patterns`, `docker-patterns`, `deployment-patterns` | Feature plan and transform module covering signal source, missing values, outliers, correlations, leakage checks, and train/serve equivalence | feature pipeline, leakage, training, artifacts |
| MLE-05 | Tune thresholds, configs, and model complexity under tradeoffs | `eval-harness`, `ai-regression-testing`, `quality-gate`, `test-coverage` | Threshold/config report comparing precision, recall, F1, AUC, calibration, group slices, latency, cost, complexity, and acceptable error classes | evaluation, threshold, promotion, regression |
| MLE-06 | Run error analysis and turn mistakes into the next experiment | `eval-harness`, `ai-regression-testing`, `mle-reviewer`, `silent-failure-hunter` | Error cluster report for false positives, false negatives, ambiguous labels, stale features, missing signals, and bug traces with lessons captured | error analysis, bug trace, iteration, regression |
| MLE-07 | Package a model artifact for batch or online inference | `api-design`, `backend-patterns`, `security-review`, `security-scan` | Versioned artifact bundle with preprocessing, config, dependency constraints, schema validation, safe loading, and PII-safe logs | artifact, security, inference contract |
| MLE-08 | Ship online serving or batch scoring with feedback capture | `api-design`, `backend-patterns`, `e2e-testing`, `browser-qa`, `accessibility` | Prediction endpoint or batch job with response envelope, timeout, batching, fallback, model version, confidence, feedback logging, and product-flow tests | serving, batch inference, fallback, user workflow |
| MLE-09 | Roll out a model with shadow traffic, canary, A/B test, or rollback | `canary-watch`, `dashboard-builder`, `verification-loop`, `performance-optimizer` | Rollout plan naming traffic split, dashboards, p95 latency, cost, quality guardrails, rollback artifact, and rollback trigger | deployment, canary, rollback |
| MLE-10 | Operate, debug, and refresh a production model after launch | `silent-failure-hunter`, `dashboard-builder`, `mle-reviewer`, `doc-updater`, `github-ops` | Observation ledger and refresh plan with drift checks, delayed-label health, alert owners, runbook updates, retrain criteria, and PR evidence | monitoring, incident response, retraining |

## Iteration Compact

Before touching model code, compress the work into one reviewable artifact. This should be short enough to fit in a PR description and precise enough that another engineer can challenge the tradeoffs.

```text
Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:
```

This compact is the MLE equivalent of a strong SWE design note. It keeps the team from optimizing a metric no one trusts, adding features that do not address the real error mode, or shipping complexity without a rollback.

## Decision Brain

Use this loop whenever the task is ambiguous, high-impact, or metric-heavy:

1. Start from the decision, not the model. Name the action that changes downstream behavior.
2. Name who cares and why. Different stakeholders pay different costs for false positives, false negatives, latency, compute spend, opacity, or missed opportunities.
3. Convert ambiguity into hypotheses. Ask what signal would separate outcomes, what evidence would disprove it, and what simple baseline should be hard to beat.
4. Research prior art or a nearby known problem before inventing a bespoke system.
5. Score choices with `(probability, confidence) x (cost, severity, importance, impact)`.
6. Consider adversarial behavior, incentives, selective disclosure, distribution shift, and feedback loops.
7. Prefer the simplest change that reduces the most important mistake. Simplicity is not laziness; it is a way to minimize blunders while preserving iteration speed.
8. Capture the decision, evidence, counterargument, and next reversible step.

## Metric and Mistake Economics

Choose metrics from failure costs, not habit:

- Use a confusion matrix early so the team can discuss concrete false positives and false negatives instead of abstract accuracy.
- Favor precision when the cost of an incorrect positive decision dominates.
- Favor recall when the cost of a missed positive dominates.
- Use F1 only when the precision/recall tradeoff is genuinely balanced and explainable.
- Use AUC or ranking metrics when ordering quality matters more than a single threshold.
- Track latency, throughput, memory, and cost as first-class metrics because they shape feasible model complexity.
- Compare against a baseline and the current production model before celebrating an offline gain.
- Treat real-world feedback signals as delayed labels with bias, lag, and coverage gaps; do not treat them as ground truth without analysis.

Every metric choice should state which mistake it makes cheaper, which mistake it makes more likely, and who absorbs that cost.

## Data and Feature Hypotheses

Features should come from a theory of separation:

- Text, categorical fields, numeric histories, graph relationships, recency, frequency, and aggregates are candidate signal families, not automatic features.
- For every feature family, state why it should separate outcomes and how it could leak future information.
- For noisy labels, consider adjudication, label confidence, soft targets, or confidence weighting.
- For class imbalance, compare weighted loss, resampling, threshold movement, and calibrated decision rules.
- For missing values, decide whether absence is informative, imputable, or a reason to abstain.
- For outliers, decide whether to clip, bucket, investigate, or preserve them as rare but important signal.
- For correlated features, check whether they are redundant, unstable, or proxies for unavailable future state.

Do not add model complexity until error analysis shows that the baseline is failing for a reason additional signal or capacity can plausibly fix.

## Error Analysis Loop

After each baseline, training run, threshold change, or config change:

1. Split mistakes into false positives, false negatives, abstentions, low-confidence cases, and system failures.
2. Cluster errors by shared traits: language, entity type, source, time, geography, device, sparsity, recency, feature freshness, label source, or model version.
3. Separate model mistakes from data bugs, label ambiguity, product ambiguity, instrumentation gaps, and serving mismatches.
4. Trace each major cluster to one of four moves: better labels, better features, better threshold/config, or better product fallback.
5. Preserve every important mistake as a regression test, eval slice, dashboard panel, or runbook entry.
6. Write the next iteration as a falsifiable experiment, not a vague "improve model" task.

The strongest MLE loop is not train -> metric -> ship. It is mistake -> cluster -> hypothesis -> experiment -> evidence -> simpler system.

## Observation Ledger

Keep a compact decision and evidence trail beside the code, PR, experiment report, or runbook:

```text
Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:
```

Use the ledger to make model work cumulative. The goal is for each iteration to make the next decision easier, not merely to produce another artifact.

## Core Workflow

### 1. Define the Prediction Contract

Capture the product-level contract before writing model code:

- Prediction target and decision owner
- Input entity, output schema, confidence/calibration fields, and allowed latency
- Batch, online, streaming, or hybrid serving mode
- Fallback behavior when the model, feature store, or dependency is unavailable
- Human review or override path for high-impact decisions
- Privacy, retention, and audit requirements for inputs, predictions, and labels

Do not accept "improve the model" as a requirement. Tie the model to an observable product behavior and a measurable acceptance gate.

### 2. Lock the Data Contract

Every ML task needs an explicit data contract:

- Entity grain and primary key
- Label definition, label timestamp, and label availability delay
- Feature timestamp, freshness SLA, and point-in-time join rules
- Train, validation, test, and backtest split policy
- Required columns, allowed nulls, ranges, categories, and units
- PII or sensitive fields that must not enter training artifacts or logs
- Dataset version or snapshot ID for reproducibility

Guard against leakage first. If a feature is not available at prediction time, or is joined using future information, remove it or move it to an analysis-only path.

### 3. Build a Reproducible Pipeline

Training code should be runnable by another engineer without hidden notebook state:

- Use typed config files or dataclasses for all hyperparameters and paths
- Pin package and model dependencies
- Set random seeds and document any nondeterministic GPU behavior
- Record dataset version, code SHA, config hash, metrics, and artifact URI
- Save preprocessing logic with the model artifact, not separately in a notebook
- Keep train, eval, and inference transformations shared or generated from one source
- Make every step idempotent so retries do not corrupt artifacts or metrics

Prefer immutable values and pure transformation functions. Avoid mutating shared data frames or global config during feature generation.

```python
import hashlib
from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class TrainingConfig:
    dataset_uri: str
    model_dir: Path
    seed: int
    learning_rate: float
    batch_size: int


def artifact_name(config: TrainingConfig, code_sha: str) -> str:
    config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
    config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
    return f"{code_sha[:12]}-{config_hash}"
```

### 4. Evaluate Before Promotion

Promotion criteria should be declared before training finishes:

- Baseline model and current production model comparison
- Primary metric aligned to product behavior
- Guardrail metrics for latency, calibration, fairness slices, cost, and error concentration
- Slice metrics for important cohorts, geographies, devices, languages, or data sources
- Confidence intervals or repeated-run variance when metrics are noisy
- Failure examples reviewed by a human for high-impact models
- Explicit "do not ship" thresholds

```python
PROMOTION_GATES = {
    "auc": ("min", 0.82),
    "calibration_error": ("max", 0.04),
    "p95_latency_ms": ("max", 80),
}


def assert_promotion_ready(metrics: dict[str, float]) -> None:
    missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
    if missing:
        raise ValueError(f"Model promotion metrics missing required gates: {missing}")

    failures = {
        name: value
        for name, (direction, threshold) in PROMOTION_GATES.items()
        for value in [metrics[name]]
        if (direction == "min" and value < threshold)
        or (direction == "max" and value > threshold)
    }
    if failures:
        raise ValueError(f"Model failed promotion gates: {failures}")
```

Use offline metrics as gates, not guarantees. When the model changes product behavior, plan shadow evaluation, canary rollout, or A/B testing before full rollout.

### 5. Package for Serving

An ML artifact is production-ready only when the serving contract is testable:

- Model artifact includes version, training data reference, config, and preprocessing
- Input schema rejects invalid, stale, or out-of-range features
- Output schema includes model version and confidence or explanation fields when useful
- Serving path has timeout, batching, resource limits, and fallback behavior
- CPU/GPU requirements are explicit and tested
- Prediction logs avoid PII and include enough identifiers for debugging and label joins
- Integration tests cover missing features, stale features, bad types, empty batches, and fallback path

Never let training-only feature code diverge from serving feature code without a test that proves equivalence.

### 6. Operate the Model

Model monitoring needs both system and quality signals:

- Availability, error rate, timeout rate, queue depth, and p50/p95/p99 latency
- Feature null rate, range drift, categorical drift, and freshness drift
- Prediction distribution drift and confidence distribution drift
- Label arrival health and delayed quality metrics
- Business KPI guardrails and rollback triggers
- Per-version dashboards for canaries and rollbacks

Every deployment should have a rollback plan that names the previous artifact, config, data dependency, and traffic-switch mechanism.

## Review Checklist

- [ ] Prediction contract is explicit and testable
- [ ] Data contract defines entity grain, label timing, feature timing, and snapshot/version
- [ ] Leakage risks were checked against prediction-time availability
- [ ] Training is reproducible from code, config, data version, and seed
- [ ] Metrics compare against baseline and current production model
- [ ] Slice metrics and guardrails are included for high-risk cohorts
- [ ] Promotion gates are automated and fail closed
- [ ] Training and serving transformations are shared or equivalence-tested
- [ ] Model artifact carries version, config, dataset reference, and preprocessing
- [ ] Serving path validates inputs and has timeout, fallback, and rollback behavior
- [ ] Monitoring covers system health, feature drift, prediction drift, and delayed labels
- [ ] Sensitive data is excluded from artifacts, logs, prompts, and examples

## Anti-Patterns

- Notebook state is required to reproduce the model
- Random split leaks future data into validation or test sets
- Feature joins ignore event time and label availability
- Offline metric improves while important slices regress
- Thresholds are tuned on the test set repeatedly
- Training preprocessing is copied manually into serving code
- Model version is missing from prediction logs
- Monitoring only checks service uptime, not data or prediction quality
- Rollback requires retraining instead of switching to a known-good artifact

## Output Expectations

When using this skill, return concrete artifacts: data contract, promotion gates, pipeline steps, test plan, deployment plan, or review findings. Call out unknowns that block production readiness instead of filling them with assumptions.

所有文件

1 个文件

安装 mle-workflow

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/affaan-m/ECC/tree/main/skills/mle-workflow # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 将自动检测并使用该技能
仓库 affaan-m/ECC

相关技能

web-search
更新时间 2026-06-29
webapp-testing
更新时间 2026-06-29
lark-base
更新时间 2026-07-05
agentmail
更新时间 2026-06-29
OR