選項

mle-workflow

affaan-m/ECC affaan-m/ECC

透過資料合約、可重複的訓練流程、可量化的品質門檻、可部署的建構產出,以及運作監控,將模型原型轉化為可投入生產的機器學習系統。

...展開全部
0
更新時間 2026-10-01

機器學習工程工作流程

運用這項技能,將模型開發成果轉化為具備清晰資料合約、可重複訓練流程、可量化品質門檻、可部署產出物以及營運監控功能的生產級機器學習系統。

何時啟用

  • 規劃或審查生產環境中的 ML 功能、模型更新、排名系統、推薦系統、分類器、嵌入式工作流程或預測管道時
  • 將筆記本程式碼轉換為可重複使用的訓練、評估、批次推論或線上推論管線
  • 設計模型推廣準則、離線/線上評估、實驗追蹤或回滾路徑
  • 除錯因資料漂移、標籤洩漏、過期特徵、產出物不匹配,或訓練與服務邏輯不一致所導致的故障
  • 新增模型監控、金絲雀部署、影子流量或部署後品質檢查

範圍校準

僅使用適合您眼前系統的「車道」。這項技能對於排名、搜尋、推薦、分類器、預測、嵌入、大型語言模型(LLM)工作流程、異常偵測及批次分析皆有所助益,但不應將單一架構強加於所有情境。

  • 切勿假設每個模型都具備監督式標籤、線上服務、特徵儲存庫、PyTorch、GPU、人工審查、A/B 測試或即時回饋。
  • 當透過資料合約、基準模型、評估腳本及回滾說明即可使變更可審查時,請勿添加繁瑣的 MLOps 機制。
  • 當專案缺乏標籤、結果延遲、切片定義、生產流量或監控責任歸屬時,請明確說明相關假設。
  • 將範例視為可互換的框架。將指標、部署模式、資料儲存庫及部署機制替換為該專案原生的等效方案。

相關技能

  • python-patterns 以及 python-testing 適用於 Python 實作及 pytest 覆蓋率
  • pytorch-patterns 針對深度學習模型、資料載入器、裝置處理及訓練迴圈
  • eval-harness 以及 ai-regression-testing 針對發布閘門與代理輔助回歸檢查
  • database-migrations, postgres-patterns,以及 clickhouse-io 用於資料儲存與分析介面
  • deployment-patterns, docker-patterns,以及 security-review 用於服務、機密資訊、容器及生產環境強化

重複利用 SWE 工作面

請勿將機器學習工程(MLE)視為與軟體工程截然不同的領域。大多數 ECC SWE 工作流程可直接套用至機器學習系統,且通常面臨更嚴格的故障模式:

建議的 minimal --with capability:machine-learning 安裝方式會讓核心代理程式介面與此技能並存。若採用僅含技能或受限於代理程式的測試框架,請將 skill:mle-workflow 與 agent:mle-reviewer ,前提是目標系統支援代理程式。

SWE 介面 MLE 用途
product-capability / architecture-decision-records 將模型開發工作轉化為明確的產品合約,並記錄不可逆轉的資料、模型及部署選擇
repo-scan / codebase-onboarding / code-tour 在導入並行機器學習架構之前,先找出現有的訓練、特徵、服務、評估及監控路徑
plan / feature-dev 將模型變更界定為產品功能,並包含資料、評估、部署及回滾階段
tdd-workflow / python-testing 在實作前測試特徵轉換、分組邏輯、指標計算、建構檔載入及推論架構
code-reviewer / mle-reviewer 審查程式碼品質,並評估機器學習特有的資料洩漏、可重現性、部署及監控風險
build-fix / pr-test-analyzer 診斷持續整合(CI)故障、評估結果不穩定、測試 fixture 缺失,以及特定環境下的模型或依賴項失敗
quality-gate / test-coverage 要求針對轉換、指標、推論合約、上線門檻及回滾行為提供自動化佐證
eval-harness / verification-loop 將離線指標、切片檢查、延遲預算及回滾演練轉化為可重複執行的門檻
ai-regression-testing 將每個生產環境中的錯誤視為回歸問題:缺失的特徵、過期的標籤、有問題的建構產出、架構漂移或服務端不匹配
api-design / backend-patterns 設計預測 API、批次工作、幂等重新訓練端點及回應封套
database-migrations / postgres-patterns / clickhouse-io 版本標籤、特徵快照、預測日誌、實驗指標及漂移分析
deployment-patterns / docker-patterns 將可重現的訓練與服務映像打包,並包含健康檢查、資源限制及回滾機制
canary-watch / dashboard-builder 透過模型版本、切片、漂移、延遲、成本及延遲標籤儀表板,讓部署狀態一目瞭然
security-review / security-scan 檢查模型產出、筆記本、提示詞、資料集及日誌,以偵測機密資訊、個人識別資訊(PII)、不安全的反序列化及供應鏈風險
e2e-testing / browser-qa / accessibility 測試使用預測結果的關鍵產品流程,包括可解釋性與備用 UI 狀態
benchmark / performance-optimizer 測量吞吐量、p95 延遲、記憶體、GPU 利用率,以及每次預測或重新訓練的成本
cost-aware-llm-pipeline / token-budget-advisor 根據品質、延遲和預算來路由大型語言模型(LLM)/嵌入式工作負載,而非預設使用最大型模型
documentation-lookup / search-first 在編寫程式碼前,驗證當前用於模型部署、特徵儲存庫、向量資料庫及評估工具的函式庫行為
git-workflow / github-ops / opensource-pipeline 將 MLE 變更打包以供審查,明確界定範圍、排除生成的副產品,並提供可重現的測試證據
strategic-compact / dmux-workflows 將冗長的 ML 工作拆分為並行軌跡:資料合約、評估框架、服務路徑、監控及文件

十項 MLE 任務模擬

在規劃或審查 MLE 工作時,將這些模擬用作覆蓋率檢查。完善的 MLE 工作流程應將每項任務簡化為明確的合約、可重複使用的軟體工程介面、自動化證據,以及可供審查的產出物。

ID 常見 MLE 任務 精簡的 ECC 路徑 所需輸出 涵蓋的管線通道
MLE-01 建構模糊的預測、排序、推薦、分類、嵌入或預報能力 product-capability, plan, architecture-decision-records, mle-workflow 迭代精簡命名:相關方、決策負責人、成功指標、不可接受的錯誤、假設、限制條件以及首次實驗 產品合約、利害關係人損失、風險、部署
MLE-02 定義指標目標、標籤、資料來源及錯誤預算 repo-scan, database-reviewer, database-migrations, postgres-patterns, clickhouse-io 包含實體粒度、標籤時間點、標籤信心度、特徵時間點、特定時間點聯結、分割政策及資料集快照的資料與指標合約 資料契約、指標設計、資料洩漏、可重現性
MLE-03 在增加複雜度之前,先建立基準模型與評分路徑 tdd-workflow, python-testing, python-patterns, code-reviewer 包含混淆矩陣、校準備註、延遲/成本估算、已知弱點,以及分數形狀與決定性測試的基準評分器 基準、評分、測試、服務一致性
MLE-04 根據「哪些因素導致結果差異」的假設來生成特徵 python-patterns, pytorch-patterns, docker-patterns, deployment-patterns 涵蓋訊號來源、缺失值、異常值、相關性、洩漏檢查,以及訓練與服務等效性的特徵規劃與轉換模組 特徵處理流程、資料洩漏、訓練、人工產物
MLE-05 在權衡取捨下調整閾值、設定及模型複雜度 eval-harness, ai-regression-testing, quality-gate, test-coverage 閾值/配置報告,比較精確度、召回率、F1 分數、AUC、校準度、群組切片、延遲、成本、複雜度及可接受的錯誤類別 評估、閾值、推廣、回歸
MLE-06 執行錯誤分析,並將錯誤轉化為下一個實驗 eval-harness, ai-regression-testing, mle-reviewer, silent-failure-hunter 針對假陽性、假陰性、標籤不明確、過時特徵、訊號遺漏及錯誤追蹤的錯誤叢集報告,並彙整相關經驗教訓 錯誤分析、錯誤追蹤、迭代、回歸
MLE-07 打包模型產出物以供批次或線上推論使用 api-design, backend-patterns, security-review, security-scan 帶有預處理、配置、依賴性約束、架構驗證、安全載入及個人識別資訊(PII)安全日誌的版本化產出包 建構成果、安全性、推論合約
MLE-08 部署具備回饋擷取功能的線上服務或批次評分 api-design, backend-patterns, e2e-testing, browser-qa, accessibility 具備回應封套、超時設定、批次處理、備用方案、模型版本、置信度、回饋記錄及產品流程測試的預測端點或批次工作 服務、批次推論、備用方案、使用者工作流程
MLE-09 透過影子流量、金絲雀測試、A/B 測試或回滾機制部署模型 canary-watch, dashboard-builder, verification-loop, performance-optimizer 部署計畫命名、流量分割、儀表板、p95 延遲、成本、品質防護措施、回滾構件及回滾觸發條件 部署、金絲雀測試、回滾
MLE-10 在模型上線後進行運作、除錯及刷新 silent-failure-hunter, dashboard-builder, mle-reviewer, doc-updater, github-ops 包含漂移檢查、延遲標籤健康狀態、警示負責人、操作手冊更新、重新訓練準則及 PR 證據的觀察日誌與更新計畫 監控、事件應變、重新訓練

迭代摘要

在觸碰模型程式碼之前,請將工作內容濃縮成一份可供審查的產出物。此產出物應簡短到足以放入拉取請求(PR)說明中,同時也應精確到足以讓其他工程師能針對其中的權衡取捨提出質疑。

Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:

此摘要相當於 MLE 領域中一份紮實的軟體工程師設計備忘錄。它能防止團隊去優化無人信賴的指標、新增無法解決真實錯誤模式的功能,或在缺乏回滾機制的情況下推出複雜功能。

決策大腦

每當任務存在模糊性、影響重大或涉及大量指標時,請使用此循環:

  1. 從決策出發,而非從模型出發。明確指出會改變下游行為的行動。
  2. 明確指出誰在乎以及原因。不同的利害關係人會因假陽性、假陰性、延遲、運算成本、不透明性或錯失機會而付出不同的代價。
  3. 將模糊性轉化為假設。詢問:什麼訊號能區分不同結果?什麼證據能推翻該假設?以及應設定何種難以超越的簡單基準?
  4. 在開發客製化系統之前,先研究現有技術或相關已知問題。
  5. 根據以下因素對選項進行評分: (probability, confidence) x (cost, severity, importance, impact).
  6. 考量對抗性行為、激勵機制、選擇性揭露、分佈偏移以及回饋迴路。
  7. 優先選擇能減少最重要錯誤的最簡單變更。簡潔並非懶惰;而是既能最小化重大失誤,又能維持迭代速度的一種方式。
  8. 記錄決策、證據、反駁論點以及下一個可逆步驟。

指標與錯誤經濟學

應根據失敗成本而非慣例來選擇指標:

  • 盡早使用混淆矩陣,讓團隊能針對具體的假陽性與假陰性進行討論,而非僅探討抽象的準確度。
  • 當錯誤的陽性決策所造成的成本佔主導地位時,應優先考量精確度。
  • 當漏檢陽性的成本佔主導地位時,應優先考慮召回率。
  • 僅在精確度與召回率的權衡真正達到平衡且可解釋時,才使用 F1 分數。
  • 當排序品質比單一閾值更重要時,應使用 AUC 或排序指標。
  • 將延遲、吞吐量、記憶體和成本視為一級指標進行追蹤,因為這些因素決定了模型複雜度的可行範圍。
  • 在慶祝離線提升之前,應先與基準模型及當前生產環境中的模型進行比較。
  • 將真實世界的回饋訊號視為具有偏誤、滯後及覆蓋缺口之延遲標籤;切勿在未經分析的情況下,將其視為真實標籤。

每項指標的選擇都應明確說明:它會使哪種錯誤的代價降低、會增加哪種錯誤發生的機率,以及由誰承擔該成本。

資料與特徵假設

特徵應源自分離理論:

  • 文字、分類欄位、數值歷史紀錄、圖形關係、近期性、頻率及聚合值,皆為候選訊號類別,而非自動生成的特徵。
  • 針對每個特徵類別,應說明其為何能區分結果,以及可能如何洩露未來資訊。
  • 針對標籤雜訊,應考慮裁決、標籤信心度、軟目標或信心加權等方法。
  • 針對類別不平衡問題,應比較加權損失、重採樣、閾值移動及校準決策規則。
  • 針對缺失值,需判斷其缺失是否具有資訊價值、可進行填補,抑或應因此棄用。
  • 針對異常值,需決定應進行裁剪、分桶、調查,抑或保留其作為罕見但重要的訊號。
  • 針對相關特徵,應檢視其是否屬於冗餘、不穩定,或是無法取得的未來狀態之代理變數。

除非錯誤分析顯示基準模型因某種原因而失效,且額外的訊號或能力可合理地解決該問題,否則切勿增加模型複雜度。

錯誤分析循環

在每次基線模型、訓練執行、閾值變更或配置變更後:

  1. 將錯誤分為假陽性、假陰性、未判定、低信心案例及系統故障。
  2. 依據共同特徵對錯誤進行聚類:語言、實體類型、來源、時間、地理位置、裝置、稀疏性、時效性、特徵新鮮度、標籤來源或模型版本。
  3. 將模型錯誤與資料錯誤、標籤歧義、產品歧義、監測缺口及服務端不匹配區分開來。
  4. 將每個主要群組追溯至四種改善措施之一:更精準的標籤、更優質的特徵、更佳的閾值/設定,或更完善的產品備用方案。
  5. 將每項重要錯誤保存為回歸測試、評估切片、儀表板面板或運維手冊條目。
  6. 將下一次迭代設計為可證偽的實驗,而非模糊的「改善模型」任務。

最強大的最大似然估計(MLE)迴圈並非「訓練 → 指標 → 部署」,而是「錯誤 → 群組 → 假設 → 實驗 → 證據 → 更簡潔的系統」。

觀察紀錄簿

在程式碼、拉取請求、實驗報告或運行手冊旁,保留簡潔的決策與證據軌跡:

Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:

利用此紀錄簿讓模型工作具有累積性。目標是讓每次迭代都能使下一個決策更為容易,而非僅僅產出另一項產出物。

核心工作流程

1. 定義預測合約

在撰寫模型程式碼之前,先釐清產品層級的契約:

  • 預測目標與決策負責人
  • 輸入實體、輸出模式、置信度/校準欄位,以及允許的延遲
  • 批次、線上、串流或混合式服務模式
  • 當模型、特徵儲存庫或依賴項不可用時的備用行為
  • 針對高影響力決策的人工審核或覆寫路徑
  • 針對輸入資料、預測結果及標籤的隱私權、保留期限與稽核要求

請勿將「改善模型」視為需求。應將模型與可觀察的產品行為及可量化的驗收標準掛鉤。

2. 鎖定資料合約

每個機器學習任務都需要明確的資料合約:

  • 實體粒度與主鍵
  • 標籤定義、標籤時間戳記及標籤可用性延遲
  • 特徵時間戳記、新鮮度服務水準協議(SLA)以及特定時間點的聯結規則
  • 訓練、驗證、測試及回測的分割政策
  • 必填欄位、允許為空的欄位、數值範圍、類別及單位
  • 不得進入訓練成果或日誌的個人識別資訊(PII)或敏感欄位
  • 用於確保可重現性的資料集版本或快照 ID

首要任務是防範資料外洩。若某個特徵在預測時不可用,或需使用未來資訊進行聯結,則應移除該特徵或將其移至「僅供分析」的路徑。

3. 建置可重現的處理流程

訓練程式碼應能由其他工程師執行,且不應包含隱藏的筆記本狀態:

  • 針對所有超參數與處理路徑,應使用類型化的配置檔案或資料類別
  • 固定套件與模型的依賴項
  • 設定隨機種子,並記錄任何非確定性的 GPU 行為
  • 記錄資料集版本、程式碼 SHA、設定檔雜湊值、指標以及建構產出 URI
  • 將預處理邏輯與模型產出檔案一併儲存,而非分開存放在筆記本中
  • 確保訓練、評估和推論的轉換保持共享,或由單一來源產生
  • 確保每個步驟皆為幺等,以免重試時損壞產出物或指標

優先採用不可變值與純轉換函數。在特徵生成過程中,避免修改共享的資料框或全域設定。

import hashlib
from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class TrainingConfig:
    dataset_uri: str
    model_dir: Path
    seed: int
    learning_rate: float
    batch_size: int


def artifact_name(config: TrainingConfig, code_sha: str) -> str:
    config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
    config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
    return f"{code_sha[:12]}-{config_hash}"

4. 推廣前進行評估

應在訓練結束前宣告部署標準:

  • 基準模型與當前生產模型的比較
  • 與產品行為對應的主要指標
  • 針對延遲、校準、公平性切片、成本及錯誤集中度的防護指標
  • 針對重要用戶群、地區、裝置、語言或資料來源的切片指標
  • 當指標存在噪訊時,應提供信賴區間或重複運行變異數
  • 針對高影響力模型,由人工審查失敗案例
  • 明確的「禁止上線」閾值
PROMOTION_GATES = {
    "auc": ("min", 0.82),
    "calibration_error": ("max", 0.04),
    "p95_latency_ms": ("max", 80),
}


def assert_promotion_ready(metrics: dict[str, float]) -> None:
    missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
    if missing:
        raise ValueError(f"Model promotion metrics missing required gates: {missing}")

    failures = {
        name: value
        for name, (direction, threshold) in PROMOTION_GATES.items()
        for value in [metrics[name]]
        if (direction == "min" and value < threshold)
        or (direction == "max" and value > threshold)
    }
    if failures:
        raise ValueError(f"Model failed promotion gates: {failures}")

將離線指標用作門檻,而非保證。當模型改變產品行為時,應在全面部署前規劃影子評估、金絲雀部署或 A/B 測試。

5. 部署準備

只有當服務合約可進行測試時,ML 建構產物才算具備生產就緒性:

  • 模型產出物應包含版本、訓練資料參考、配置及預處理資訊
  • 輸入資料結構會拒收無效、過期或超出範圍的特徵
  • 輸出模式應包含模型版本,並在適當時加入信心度或解釋欄位
  • 服務路徑具備超時設定、批次處理、資源限制及備用行為
  • CPU/GPU 需求已明確說明並經過測試
  • 預測日誌會避開個人識別資訊(PII),並包含足夠的識別碼以供除錯及標籤關聯
  • 整合測試涵蓋缺失特徵、過期特徵、類型錯誤、空批次及備用路徑

切勿讓僅用於訓練的特徵代碼與生產環境的特徵代碼產生分歧,除非有測試能證明兩者等效。

6. 模型運作

模型監控需同時涵蓋系統與品質指標:

  • 可用性、錯誤率、超時率、佇列深度,以及 p50/p95/p99 延遲
  • 特徵空值率、範圍漂移、類別漂移及新鮮度漂移
  • 預測分佈漂移與置信度分佈漂移
  • 標籤抵達狀態與延遲品質指標
  • 業務 KPI 警戒線與回滾觸發條件
  • 針對金絲雀測試與回滾的各版本儀表板

每次部署都應具備回滾計畫,其中須明確指定前一版本的構建產出、配置、資料依賴項以及流量切換機制。

審查清單

  • 預測合約明確且可測試
  • 資料合約定義實體粒度、標籤時間點、功能時間點以及快照/版本
  • 已針對預測時點的可用性檢查洩漏風險
  • 訓練過程可透過程式碼、設定、資料版本及初始化值進行重現
  • 指標需與基準模型及當前生產環境模型進行比較
  • 針對高風險群組,已納入切片指標與防護措施
  • 推廣閘門已自動化,並採用「失敗即關閉」機制
  • 訓練與服務端的轉換函式均已共享或通過等價性測試
  • 模型產出包含版本、配置、資料集參考及預處理資訊
  • 服務路徑會驗證輸入,並具備超時、備用方案及回滾機制
  • 監控範圍涵蓋系統健康狀態、特徵漂移、預測漂移及延遲標籤
  • 敏感資料不會包含在建構產出、日誌、提示詞及範例中

反模式

  • 重現模型需依賴筆記本狀態
  • 隨機分割會導致未來資料洩漏至驗證集或測試集中
  • 特徵聯結忽略事件時間與標籤可用性
  • 離線指標有所改善,但重要切片卻出現退步
  • 閾值反覆在測試集上進行微調
  • 訓練預處理步驟被手動複製到服務端程式碼中
  • 預測日誌中缺少模型版本資訊
  • 監控僅檢查服務正常運作時間,未檢查資料或預測品質
  • 回滾時需重新訓練,而非切換至已知正常的建置檔

輸出期望

使用此技能時,應回傳具體的產出成果:資料合約、發布閘門、管線步驟、測試計畫、部署計畫或審查結果。針對阻礙生產就緒的未知因素應明確指出,而非以假設來填補。

在 GitHub 上查看
---
name: mle-workflow
description: Turn model work into a production ML system with data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.
---

# Machine Learning Engineering Workflow

Use this skill to turn model work into a production ML system with clear data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.

## When to Activate

- Planning or reviewing a production ML feature, model refresh, ranking system, recommender, classifier, embedding workflow, or forecasting pipeline
- Converting notebook code into a reusable training, evaluation, batch inference, or online inference pipeline
- Designing model promotion criteria, offline/online evals, experiment tracking, or rollback paths
- Debugging failures caused by data drift, label leakage, stale features, artifact mismatch, or inconsistent training and serving logic
- Adding model monitoring, canary rollout, shadow traffic, or post-deploy quality checks

## Scope Calibration

Use only the lanes that fit the system in front of you. This skill is useful for ranking, search, recommendations, classifiers, forecasting, embeddings, LLM workflows, anomaly detection, and batch analytics, but it should not force one architecture onto all of them.

- Do not assume every model has supervised labels, online serving, a feature store, PyTorch, GPUs, human review, A/B tests, or real-time feedback.
- Do not add heavyweight MLOps machinery when a data contract, baseline, eval script, and rollback note would make the change reviewable.
- Do make assumptions explicit when the project lacks labels, delayed outcomes, slice definitions, production traffic, or monitoring ownership.
- Treat examples as interchangeable scaffolds. Replace metrics, serving mode, data stores, and rollout mechanics with the project-native equivalents.

## Related Skills

- `python-patterns` and `python-testing` for Python implementation and pytest coverage
- `pytorch-patterns` for deep learning models, data loaders, device handling, and training loops
- `eval-harness` and `ai-regression-testing` for promotion gates and agent-assisted regression checks
- `database-migrations`, `postgres-patterns`, and `clickhouse-io` for data storage and analytics surfaces
- `deployment-patterns`, `docker-patterns`, and `security-review` for serving, secrets, containers, and production hardening

## Reuse the SWE Surface

Do not treat MLE as separate from software engineering. Most ECC SWE workflows apply directly to ML systems, often with stricter failure modes:

The recommended `minimal --with capability:machine-learning` install keeps the core agent surface available alongside this skill. For skill-only or agent-limited harnesses, pair `skill:mle-workflow` with `agent:mle-reviewer` where the target supports agents.

| SWE surface | MLE use |
|-------------|---------|
| `product-capability` / `architecture-decision-records` | Turn model work into explicit product contracts and record irreversible data, model, and rollout choices |
| `repo-scan` / `codebase-onboarding` / `code-tour` | Find existing training, feature, serving, eval, and monitoring paths before introducing a parallel ML stack |
| `plan` / `feature-dev` | Scope model changes as product capabilities with data, eval, serving, and rollback phases |
| `tdd-workflow` / `python-testing` | Test feature transforms, split logic, metric calculations, artifact loading, and inference schemas before implementation |
| `code-reviewer` / `mle-reviewer` | Review code quality plus ML-specific leakage, reproducibility, promotion, and monitoring risks |
| `build-fix` / `pr-test-analyzer` | Diagnose broken CI, flaky evals, missing fixtures, and environment-specific model or dependency failures |
| `quality-gate` / `test-coverage` | Require automated evidence for transforms, metrics, inference contracts, promotion gates, and rollback behavior |
| `eval-harness` / `verification-loop` | Turn offline metrics, slice checks, latency budgets, and rollback drills into repeatable gates |
| `ai-regression-testing` | Preserve every production bug as a regression: missing feature, stale label, bad artifact, schema drift, or serving mismatch |
| `api-design` / `backend-patterns` | Design prediction APIs, batch jobs, idempotent retraining endpoints, and response envelopes |
| `database-migrations` / `postgres-patterns` / `clickhouse-io` | Version labels, feature snapshots, prediction logs, experiment metrics, and drift analytics |
| `deployment-patterns` / `docker-patterns` | Package reproducible training and serving images with health checks, resource limits, and rollback |
| `canary-watch` / `dashboard-builder` | Make rollout health visible with model-version, slice, drift, latency, cost, and delayed-label dashboards |
| `security-review` / `security-scan` | Check model artifacts, notebooks, prompts, datasets, and logs for secrets, PII, unsafe deserialization, and supply-chain risk |
| `e2e-testing` / `browser-qa` / `accessibility` | Test critical product flows that consume predictions, including explainability and fallback UI states |
| `benchmark` / `performance-optimizer` | Measure throughput, p95 latency, memory, GPU utilization, and cost per prediction or retrain |
| `cost-aware-llm-pipeline` / `token-budget-advisor` | Route LLM/embedding workloads by quality, latency, and budget instead of defaulting to the largest model |
| `documentation-lookup` / `search-first` | Verify current library behavior for model serving, feature stores, vector DBs, and eval tooling before coding |
| `git-workflow` / `github-ops` / `opensource-pipeline` | Package MLE changes for review with crisp scope, generated artifacts excluded, and reproducible test evidence |
| `strategic-compact` / `dmux-workflows` | Split long ML work into parallel tracks: data contract, eval harness, serving path, monitoring, and docs |

## Ten MLE Task Simulations

Use these simulations as coverage checks when planning or reviewing MLE work. A strong MLE workflow should reduce each task to explicit contracts, reusable SWE surfaces, automated evidence, and a reviewable artifact.

| ID | Common MLE task | Streamlined ECC path | Required output | Pipeline lanes covered |
|----|-----------------|----------------------|-----------------|------------------------|
| MLE-01 | Frame an ambiguous prediction, ranking, recommender, classifier, embedding, or forecast capability | `product-capability`, `plan`, `architecture-decision-records`, `mle-workflow` | Iteration Compact naming who cares, decision owner, success metric, unacceptable mistakes, assumptions, constraints, and first experiment | product contract, stakeholder loss, risk, rollout |
| MLE-02 | Define metric goals, labels, data sources, and the mistake budget | `repo-scan`, `database-reviewer`, `database-migrations`, `postgres-patterns`, `clickhouse-io` | Data and metric contract with entity grain, label timing, label confidence, feature timing, point-in-time joins, split policy, and dataset snapshot | data contract, metric design, leakage, reproducibility |
| MLE-03 | Build a baseline model and scoring path before adding complexity | `tdd-workflow`, `python-testing`, `python-patterns`, `code-reviewer` | Baseline scorer with confusion matrix, calibration notes, latency/cost estimate, known weaknesses, and tests for score shape and determinism | baseline, scoring, testing, serving parity |
| MLE-04 | Generate features from hypotheses about what separates outcomes | `python-patterns`, `pytorch-patterns`, `docker-patterns`, `deployment-patterns` | Feature plan and transform module covering signal source, missing values, outliers, correlations, leakage checks, and train/serve equivalence | feature pipeline, leakage, training, artifacts |
| MLE-05 | Tune thresholds, configs, and model complexity under tradeoffs | `eval-harness`, `ai-regression-testing`, `quality-gate`, `test-coverage` | Threshold/config report comparing precision, recall, F1, AUC, calibration, group slices, latency, cost, complexity, and acceptable error classes | evaluation, threshold, promotion, regression |
| MLE-06 | Run error analysis and turn mistakes into the next experiment | `eval-harness`, `ai-regression-testing`, `mle-reviewer`, `silent-failure-hunter` | Error cluster report for false positives, false negatives, ambiguous labels, stale features, missing signals, and bug traces with lessons captured | error analysis, bug trace, iteration, regression |
| MLE-07 | Package a model artifact for batch or online inference | `api-design`, `backend-patterns`, `security-review`, `security-scan` | Versioned artifact bundle with preprocessing, config, dependency constraints, schema validation, safe loading, and PII-safe logs | artifact, security, inference contract |
| MLE-08 | Ship online serving or batch scoring with feedback capture | `api-design`, `backend-patterns`, `e2e-testing`, `browser-qa`, `accessibility` | Prediction endpoint or batch job with response envelope, timeout, batching, fallback, model version, confidence, feedback logging, and product-flow tests | serving, batch inference, fallback, user workflow |
| MLE-09 | Roll out a model with shadow traffic, canary, A/B test, or rollback | `canary-watch`, `dashboard-builder`, `verification-loop`, `performance-optimizer` | Rollout plan naming traffic split, dashboards, p95 latency, cost, quality guardrails, rollback artifact, and rollback trigger | deployment, canary, rollback |
| MLE-10 | Operate, debug, and refresh a production model after launch | `silent-failure-hunter`, `dashboard-builder`, `mle-reviewer`, `doc-updater`, `github-ops` | Observation ledger and refresh plan with drift checks, delayed-label health, alert owners, runbook updates, retrain criteria, and PR evidence | monitoring, incident response, retraining |

## Iteration Compact

Before touching model code, compress the work into one reviewable artifact. This should be short enough to fit in a PR description and precise enough that another engineer can challenge the tradeoffs.

```text
Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:
```

This compact is the MLE equivalent of a strong SWE design note. It keeps the team from optimizing a metric no one trusts, adding features that do not address the real error mode, or shipping complexity without a rollback.

## Decision Brain

Use this loop whenever the task is ambiguous, high-impact, or metric-heavy:

1. Start from the decision, not the model. Name the action that changes downstream behavior.
2. Name who cares and why. Different stakeholders pay different costs for false positives, false negatives, latency, compute spend, opacity, or missed opportunities.
3. Convert ambiguity into hypotheses. Ask what signal would separate outcomes, what evidence would disprove it, and what simple baseline should be hard to beat.
4. Research prior art or a nearby known problem before inventing a bespoke system.
5. Score choices with `(probability, confidence) x (cost, severity, importance, impact)`.
6. Consider adversarial behavior, incentives, selective disclosure, distribution shift, and feedback loops.
7. Prefer the simplest change that reduces the most important mistake. Simplicity is not laziness; it is a way to minimize blunders while preserving iteration speed.
8. Capture the decision, evidence, counterargument, and next reversible step.

## Metric and Mistake Economics

Choose metrics from failure costs, not habit:

- Use a confusion matrix early so the team can discuss concrete false positives and false negatives instead of abstract accuracy.
- Favor precision when the cost of an incorrect positive decision dominates.
- Favor recall when the cost of a missed positive dominates.
- Use F1 only when the precision/recall tradeoff is genuinely balanced and explainable.
- Use AUC or ranking metrics when ordering quality matters more than a single threshold.
- Track latency, throughput, memory, and cost as first-class metrics because they shape feasible model complexity.
- Compare against a baseline and the current production model before celebrating an offline gain.
- Treat real-world feedback signals as delayed labels with bias, lag, and coverage gaps; do not treat them as ground truth without analysis.

Every metric choice should state which mistake it makes cheaper, which mistake it makes more likely, and who absorbs that cost.

## Data and Feature Hypotheses

Features should come from a theory of separation:

- Text, categorical fields, numeric histories, graph relationships, recency, frequency, and aggregates are candidate signal families, not automatic features.
- For every feature family, state why it should separate outcomes and how it could leak future information.
- For noisy labels, consider adjudication, label confidence, soft targets, or confidence weighting.
- For class imbalance, compare weighted loss, resampling, threshold movement, and calibrated decision rules.
- For missing values, decide whether absence is informative, imputable, or a reason to abstain.
- For outliers, decide whether to clip, bucket, investigate, or preserve them as rare but important signal.
- For correlated features, check whether they are redundant, unstable, or proxies for unavailable future state.

Do not add model complexity until error analysis shows that the baseline is failing for a reason additional signal or capacity can plausibly fix.

## Error Analysis Loop

After each baseline, training run, threshold change, or config change:

1. Split mistakes into false positives, false negatives, abstentions, low-confidence cases, and system failures.
2. Cluster errors by shared traits: language, entity type, source, time, geography, device, sparsity, recency, feature freshness, label source, or model version.
3. Separate model mistakes from data bugs, label ambiguity, product ambiguity, instrumentation gaps, and serving mismatches.
4. Trace each major cluster to one of four moves: better labels, better features, better threshold/config, or better product fallback.
5. Preserve every important mistake as a regression test, eval slice, dashboard panel, or runbook entry.
6. Write the next iteration as a falsifiable experiment, not a vague "improve model" task.

The strongest MLE loop is not train -> metric -> ship. It is mistake -> cluster -> hypothesis -> experiment -> evidence -> simpler system.

## Observation Ledger

Keep a compact decision and evidence trail beside the code, PR, experiment report, or runbook:

```text
Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:
```

Use the ledger to make model work cumulative. The goal is for each iteration to make the next decision easier, not merely to produce another artifact.

## Core Workflow

### 1. Define the Prediction Contract

Capture the product-level contract before writing model code:

- Prediction target and decision owner
- Input entity, output schema, confidence/calibration fields, and allowed latency
- Batch, online, streaming, or hybrid serving mode
- Fallback behavior when the model, feature store, or dependency is unavailable
- Human review or override path for high-impact decisions
- Privacy, retention, and audit requirements for inputs, predictions, and labels

Do not accept "improve the model" as a requirement. Tie the model to an observable product behavior and a measurable acceptance gate.

### 2. Lock the Data Contract

Every ML task needs an explicit data contract:

- Entity grain and primary key
- Label definition, label timestamp, and label availability delay
- Feature timestamp, freshness SLA, and point-in-time join rules
- Train, validation, test, and backtest split policy
- Required columns, allowed nulls, ranges, categories, and units
- PII or sensitive fields that must not enter training artifacts or logs
- Dataset version or snapshot ID for reproducibility

Guard against leakage first. If a feature is not available at prediction time, or is joined using future information, remove it or move it to an analysis-only path.

### 3. Build a Reproducible Pipeline

Training code should be runnable by another engineer without hidden notebook state:

- Use typed config files or dataclasses for all hyperparameters and paths
- Pin package and model dependencies
- Set random seeds and document any nondeterministic GPU behavior
- Record dataset version, code SHA, config hash, metrics, and artifact URI
- Save preprocessing logic with the model artifact, not separately in a notebook
- Keep train, eval, and inference transformations shared or generated from one source
- Make every step idempotent so retries do not corrupt artifacts or metrics

Prefer immutable values and pure transformation functions. Avoid mutating shared data frames or global config during feature generation.

```python
import hashlib
from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class TrainingConfig:
    dataset_uri: str
    model_dir: Path
    seed: int
    learning_rate: float
    batch_size: int


def artifact_name(config: TrainingConfig, code_sha: str) -> str:
    config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
    config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
    return f"{code_sha[:12]}-{config_hash}"
```

### 4. Evaluate Before Promotion

Promotion criteria should be declared before training finishes:

- Baseline model and current production model comparison
- Primary metric aligned to product behavior
- Guardrail metrics for latency, calibration, fairness slices, cost, and error concentration
- Slice metrics for important cohorts, geographies, devices, languages, or data sources
- Confidence intervals or repeated-run variance when metrics are noisy
- Failure examples reviewed by a human for high-impact models
- Explicit "do not ship" thresholds

```python
PROMOTION_GATES = {
    "auc": ("min", 0.82),
    "calibration_error": ("max", 0.04),
    "p95_latency_ms": ("max", 80),
}


def assert_promotion_ready(metrics: dict[str, float]) -> None:
    missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
    if missing:
        raise ValueError(f"Model promotion metrics missing required gates: {missing}")

    failures = {
        name: value
        for name, (direction, threshold) in PROMOTION_GATES.items()
        for value in [metrics[name]]
        if (direction == "min" and value < threshold)
        or (direction == "max" and value > threshold)
    }
    if failures:
        raise ValueError(f"Model failed promotion gates: {failures}")
```

Use offline metrics as gates, not guarantees. When the model changes product behavior, plan shadow evaluation, canary rollout, or A/B testing before full rollout.

### 5. Package for Serving

An ML artifact is production-ready only when the serving contract is testable:

- Model artifact includes version, training data reference, config, and preprocessing
- Input schema rejects invalid, stale, or out-of-range features
- Output schema includes model version and confidence or explanation fields when useful
- Serving path has timeout, batching, resource limits, and fallback behavior
- CPU/GPU requirements are explicit and tested
- Prediction logs avoid PII and include enough identifiers for debugging and label joins
- Integration tests cover missing features, stale features, bad types, empty batches, and fallback path

Never let training-only feature code diverge from serving feature code without a test that proves equivalence.

### 6. Operate the Model

Model monitoring needs both system and quality signals:

- Availability, error rate, timeout rate, queue depth, and p50/p95/p99 latency
- Feature null rate, range drift, categorical drift, and freshness drift
- Prediction distribution drift and confidence distribution drift
- Label arrival health and delayed quality metrics
- Business KPI guardrails and rollback triggers
- Per-version dashboards for canaries and rollbacks

Every deployment should have a rollback plan that names the previous artifact, config, data dependency, and traffic-switch mechanism.

## Review Checklist

- [ ] Prediction contract is explicit and testable
- [ ] Data contract defines entity grain, label timing, feature timing, and snapshot/version
- [ ] Leakage risks were checked against prediction-time availability
- [ ] Training is reproducible from code, config, data version, and seed
- [ ] Metrics compare against baseline and current production model
- [ ] Slice metrics and guardrails are included for high-risk cohorts
- [ ] Promotion gates are automated and fail closed
- [ ] Training and serving transformations are shared or equivalence-tested
- [ ] Model artifact carries version, config, dataset reference, and preprocessing
- [ ] Serving path validates inputs and has timeout, fallback, and rollback behavior
- [ ] Monitoring covers system health, feature drift, prediction drift, and delayed labels
- [ ] Sensitive data is excluded from artifacts, logs, prompts, and examples

## Anti-Patterns

- Notebook state is required to reproduce the model
- Random split leaks future data into validation or test sets
- Feature joins ignore event time and label availability
- Offline metric improves while important slices regress
- Thresholds are tuned on the test set repeatedly
- Training preprocessing is copied manually into serving code
- Model version is missing from prediction logs
- Monitoring only checks service uptime, not data or prediction quality
- Rollback requires retraining instead of switching to a known-good artifact

## Output Expectations

When using this skill, return concrete artifacts: data contract, promotion gates, pipeline steps, test plan, deployment plan, or review findings. Call out unknowns that block production readiness instead of filling them with assumptions.

所有檔案

1 個檔案

安裝 mle-workflow

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/affaan-m/ECC/tree/main/skills/mle-workflow # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 affaan-m/ECC

相關技能

web-search
更新時間 2026-06-29
webapp-testing
更新時間 2026-06-29
lark-base
更新時間 2026-07-05
agentmail
更新時間 2026-06-29
OR