オプション

mle-workflow

affaan-m/ECC affaan-m/ECC

データ契約、再現可能なトレーニング、測定可能な品質ゲート、デプロイ可能な成果物、および運用モニタリングを活用して、モデル開発を本番環境向けの機械学習システムへと発展させます。

...すべて拡張します
0
更新された時間 2026年10月1日

機械学習エンジニアリングのワークフロー

このスキルを活用して、モデル開発の成果を、明確なデータ契約、再現性のあるトレーニング、測定可能な品質ゲート、デプロイ可能な成果物、および運用モニタリングを備えた本番環境向けのMLシステムへと変換します。

活用すべき場面

  • 本番環境向けのML機能、モデルの更新、ランキングシステム、レコメンダー、分類器、埋め込みワークフロー、または予測パイプラインの計画またはレビュー時
  • ノートブックのコードを、再利用可能なトレーニング、評価、バッチ推論、またはオンライン推論パイプラインに変換する場合
  • モデルの本番移行基準、オフライン/オンライン評価、実験の追跡、またはロールバック経路の設計
  • データドリフト、ラベルリーク、古くなった特徴量、アーティファクトの不一致、またはトレーニングとサービングのロジックの不整合によって引き起こされる障害のデバッグ
  • モデルの監視、カナリア展開、シャドウトラフィック、またはデプロイ後の品質チェックの追加

スコープの調整

目の前のシステムに適したレーンのみを使用してください。このスキルは、ランキング、検索、レコメンデーション、分類、予測、埋め込み、LLMワークフロー、異常検知、バッチ分析に役立ちますが、それらすべてに単一のアーキテクチャを強要すべきではありません。

  • すべてのモデルが、教師ありラベル、オンラインサービング、特徴量ストア、PyTorch、GPU、人間によるレビュー、A/Bテスト、またはリアルタイムフィードバックを備えているとは仮定しないでください。
  • データ契約、ベースライン、評価スクリプト、ロールバックに関する注意事項があれば変更内容を検証できる場合、重厚なMLOpsの仕組みを追加しないでください。
  • プロジェクトにラベル、結果の遅延、スライスの定義、本番トラフィック、またはモニタリングの責任者が欠けている場合は、その前提条件を明示してください。
  • 例は互換性のある足場として扱ってください。指標、サービングモード、データストア、ロールアウトの仕組みは、プロジェクト固有の同等のものに置き換えてください。

関連スキル

  • python-patterns および python-testing Python実装およびpytestカバレッジ用
  • pytorch-patterns ディープラーニングモデル、データローダー、デバイス処理、およびトレーニングループについて
  • eval-harness および ai-regression-testing プロモーションゲートおよびエージェント支援型回帰チェック
  • database-migrations, postgres-patterns、および clickhouse-io データストレージおよび分析インターフェース
  • deployment-patterns, docker-patterns、および security-review サービング、シークレット、コンテナ、および本番環境の堅牢化

SWEの領域を再利用する

MLEをソフトウェアエンジニアリングとは切り離して扱わないでください。ECCのSWEワークフローの多くはMLシステムに直接適用可能ですが、多くの場合、より厳格な障害モードが伴います:

推奨される minimal --with capability:machine-learning インストールでは、このスキルと並行してコアエージェントのインターフェースも利用可能です。スキルのみ、またはエージェントが制限されたハーネスの場合は、 skill:mle-workflow を agent:mle-reviewer を組み合わせてください。

SWEインターフェース MLEの使用
product-capability / architecture-decision-records モデル開発の成果を明確なプロダクト契約に変換し、データ、モデル、ロールアウトに関する選択内容を不可逆的に記録する
repo-scan / codebase-onboarding / code-tour 並列MLスタックを導入する前に、既存のトレーニング、特徴量、サービング、評価、およびモニタリングのパスを特定する
plan / feature-dev モデル変更の範囲を、データ、評価、サービング、ロールバックの各フェーズを含む製品機能として定義する
tdd-workflow / python-testing 実装前に、特徴量変換、分割ロジック、メトリック計算、アーティファクトの読み込み、および推論スキーマをテストする
code-reviewer / mle-reviewer コードの品質に加え、ML特有のリーク、再現性、本番環境への移行、およびモニタリングに関するリスクをレビューする
build-fix / pr-test-analyzer CIの不具合、不安定な評価、フィクスチャの欠落、および環境固有のモデルや依存関係の障害を診断する
quality-gate / test-coverage 変換、メトリクス、推論契約、本番移行のゲート、およびロールバック動作について、自動化されたエビデンスを必須とする
eval-harness / verification-loop オフラインメトリクス、スライスチェック、レイテンシ予算、およびロールバック演習を、再現可能なゲートに変換する
ai-regression-testing 本番環境のバグをすべて回帰バグとして記録する:機能の欠落、ラベルの古さ、アーティファクトの不具合、スキーマのドリフト、サービングの不一致など
api-design / backend-patterns 予測API、バッチジョブ、冪等な再トレーニングエンドポイント、およびレスポンスエンベロープを設計する
database-migrations / postgres-patterns / clickhouse-io バージョンラベル、機能スナップショット、予測ログ、実験メトリクス、およびドリフト分析
deployment-patterns / docker-patterns 再現可能なトレーニングおよびサービングイメージを、ヘルスチェック、リソース制限、およびロールバック機能とともにパッケージ化する
canary-watch / dashboard-builder モデルバージョン、スライス、ドリフト、レイテンシ、コスト、遅延ラベルのダッシュボードを用いて、ロールアウトの健全性を可視化する
security-review / security-scan モデルアーティファクト、ノートブック、プロンプト、データセット、およびログについて、機密情報、PII、安全でない逆シリアライズ、およびサプライチェーンリスクの有無を確認する
e2e-testing / browser-qa / accessibility 説明可能性やフォールバック UI 状態を含め、予測を利用する重要なプロダクトフローをテストする
benchmark / performance-optimizer スループット、p95レイテンシ、メモリ、GPU使用率、および予測または再トレーニングあたりのコストを測定
cost-aware-llm-pipeline / token-budget-advisor LLM/エンベディングのワークロードを、デフォルトで最大のモデルに割り当てるのではなく、品質、レイテンシ、予算に基づいてルーティングする
documentation-lookup / search-first コーディング前に、モデルサービング、特徴量ストア、ベクトルDB、評価ツールに関する現在のライブラリの挙動を検証する
git-workflow / github-ops / opensource-pipeline MLEの変更内容を、明確な範囲を定義し、生成された成果物を除外し、再現可能なテスト証拠を添えてレビュー用にパッケージ化する
strategic-compact / dmux-workflows 長期的なML作業を、データ契約、評価ハーネス、サービングパス、モニタリング、ドキュメントという並行するトラックに分割する

10のMLEタスクシミュレーション

MLE作業の計画やレビューの際には、これらのシミュレーションをカバレッジチェックとして活用する。堅牢なMLEワークフローでは、各タスクを明確な契約、再利用可能なSWEインターフェース、自動化されたエビデンス、およびレビュー可能な成果物に還元すべきである。

ID 一般的なMLEタスク 合理化されたECCパス 必要な出力 対象となるパイプラインレーン
MLE-01 曖昧な予測、ランキング、レコメンデーション、分類、埋め込み、または予測機能の枠組みを構築する product-capability, plan, architecture-decision-records, mle-workflow 反復「誰が関与するか」、意思決定責任者、成功指標、許容できないミス、前提条件、制約、および最初の実験を簡潔に命名する 製品契約、ステークホルダーの損失、リスク、展開
MLE-02 指標目標、ラベル、データソース、および許容誤差を定義する repo-scan, database-reviewer, database-migrations, postgres-patterns, clickhouse-io エンティティ粒度、ラベルのタイミング、ラベルの信頼度、特徴量のタイミング、特定時点での結合、分割ポリシー、およびデータセットのスナップショットを含むデータおよび指標契約 データ契約、指標設計、リーク、再現性
MLE-03 複雑性を追加する前に、ベースラインモデルとスコアリングパスを構築する tdd-workflow, python-testing, python-patterns, code-reviewer 混同行列、キャリブレーションに関する注記、レイテンシ/コストの見積もり、既知の弱点、およびスコアの形状と決定性に関するテストを備えたベースライン・スコアラー ベースライン、スコアリング、テスト、サービングの整合性
MLE-04 結果を分ける要因に関する仮説から特徴量を生成する python-patterns, pytorch-patterns, docker-patterns, deployment-patterns シグナルソース、欠損値、外れ値、相関、リークチェック、およびトレーニング/サービングの等価性を網羅した特徴量計画および変換モジュール 特徴量パイプライン、リーク、トレーニング、アーティファクト
MLE-05 トレードオフを考慮したしきい値、設定、およびモデルの複雑度の調整 eval-harness, ai-regression-testing, quality-gate, test-coverage 精度、リコール、F1、AUC、キャリブレーション、グループスライス、レイテンシ、コスト、複雑度、および許容可能なエラークラスを比較する閾値/設定レポート 評価、閾値、プロモーション、回帰
MLE-06 エラー分析を実行し、ミスを次の実験に活かす eval-harness, ai-regression-testing, mle-reviewer, silent-failure-hunter 誤検知、検知漏れ、曖昧なラベル、陳腐化した特徴量、欠落した信号、およびバグトレースに関するエラークラスターレポート(得られた教訓を含む) エラー分析、バグトレース、反復、回帰
MLE-07 バッチまたはオンライン推論用のモデルアーティファクトをパッケージ化する api-design, backend-patterns, security-review, security-scan 前処理、設定、依存関係の制約、スキーマ検証、安全な読み込み、およびPIIを保護したログを含む、バージョン管理されたアーティファクトバンドル アーティファクト、セキュリティ、推論契約
MLE-08 フィードバックの取得機能を備えたオンラインサービングまたはバッチスコアリングの提供 api-design, backend-patterns, e2e-testing, browser-qa, accessibility レスポンスエンベロープ、タイムアウト、バッチ処理、フォールバック、モデルバージョン、信頼度、フィードバックロギング、およびプロダクトフローテストを備えた予測エンドポイントまたはバッチジョブ サービング、バッチ推論、フォールバック、ユーザーワークフロー
MLE-09 シャドウトラフィック、カナリアテスト、A/Bテスト、またはロールバックを用いたモデルの展開 canary-watch, dashboard-builder, verification-loop, performance-optimizer ロールアウト計画の命名、トラフィック分割、ダッシュボード、p95 レイテンシ、コスト、品質ガードレール、ロールバックアーティファクト、およびロールバックトリガー デプロイ、カナリア、ロールバック
MLE-10 リリース後の本番モデルの運用、デバッグ、およびリフレッシュ silent-failure-hunter, dashboard-builder, mle-reviewer, doc-updater, github-ops ドリフトチェック、遅延ラベルの健全性、アラート担当者、ランブックの更新、再学習基準、および PR の証拠を含む、観測台帳および更新計画 モニタリング、インシデント対応、再学習

反復の簡素化

モデルコードに手を加える前に、作業内容を1つのレビュー可能な成果物にまとめます。これはプルリクエストの説明欄に収まるほど簡潔であり、かつ他のエンジニアがそのトレードオフについて議論できるほど正確である必要があります。

Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:

このコンパクトは、MLEにおける優れたSWE設計ノートの役割を果たします。これにより、チームが誰も信頼していない指標を最適化したり、実際のエラーモードに対処しない機能を追加したり、ロールバック手段のない複雑な機能をリリースしたりすることを防ぎます。

意思決定の基盤

タスクが曖昧であったり、影響が大きかったり、メトリクスに依存していたりする場合は、常にこのループを使用してください:

  1. モデルではなく、意思決定から始めます。下流の挙動を変えるアクションを明確に定義します。
  2. 誰が関心を持ち、その理由は何かを明確にする。ステークホルダーによって、誤検知、検知漏れ、レイテンシ、計算コスト、不透明性、あるいは機会損失に対する負担は異なる。
  3. 曖昧さを仮説に変換する。「どのようなシグナルが結果を区別するか」「どのような証拠がそれを反証するか」「どのようなシンプルなベースラインが、上回るのが難しいものであるべきか」を問う。
  4. 特注のシステムを考案する前に、先行技術や類似の既知の問題を調査してください。
  5. 以下の観点から選択肢を評価する: (probability, confidence) x (cost, severity, importance, impact).
  6. 敵対的な行動、インセンティブ、選択的開示、分布のシフト、フィードバックループを考慮する。
  7. 最も重要なミスを減らす、最もシンプルな変更を優先する。シンプルさは怠慢ではなく、反復のスピードを維持しつつ、重大なミスを最小限に抑える方法である。
  8. 決定事項、証拠、反論、そして次に取り消せるステップを記録する。

指標とミス経済学

指標は習慣ではなく、失敗コストに基づいて選択する:

  • 早期に混同行列(コンフュージョンマトリックス)を活用し、チームが抽象的な精度ではなく、具体的な偽陽性や偽陰性について議論できるようにする。
  • 誤った肯定判定のコストが支配的である場合は、精度を優先する。
  • 見逃しによるコストが支配的である場合は、再現率を優先する。
  • F1スコアは、精度と再現率のトレードオフが真にバランスが取れており、説明可能な場合にのみ使用すること。
  • 単一の閾値よりも品質の順位付けが重要である場合は、AUCまたはランキング指標を使用する。
  • レイテンシ、スループット、メモリ、コストは、実現可能なモデルの複雑さを左右するため、第一級の指標として追跡する。
  • オフラインでの性能向上を喜ぶ前に、ベースラインおよび現在の本番環境モデルと比較すること。
  • 実世界のフィードバック信号は、バイアス、遅延、およびカバレッジのギャップを伴う遅延ラベルとして扱い、分析なしに「真値」として扱ってはならない。

指標を選択する際は、その指標によってどのミスがより安価になり、どのミスがより起こりやすくなり、そのコストを誰が負担するかを明示すべきです。

データと特徴量の仮説

特徴量は分離の理論に基づいて導出されるべきである:

  • テキスト、カテゴリ型フィールド、数値履歴、グラフ上の関係、最新性、頻度、および集計値は、候補となるシグナルのファミリーであり、自動生成された特徴量ではない。
  • 各特徴量ファミリーについて、なぜそれが結果を分離できるのか、またどのようにして将来的な情報が漏洩する可能性があるのかを明記すること。
  • ラベルにノイズが含まれる場合は、裁定、ラベルの信頼度、ソフトターゲット、または信頼度加重を検討する。
  • クラスの不均衡については、重み付き損失、リサンプリング、閾値移動、およびキャリブレーションされた決定ルールを比較検討する。
  • 欠損値については、その欠如が情報的であるか、補完可能か、あるいは判断を控えるべき理由であるかを決定する。
  • 外れ値については、クリップするか、バケット化するか、調査するか、あるいは稀ではあるが重要なシグナルとして保持するかを決定してください。
  • 相関のある特徴量については、それらが冗長か、不安定か、あるいは利用できない将来の状態の代用変数であるかを確認してください。

エラー分析により、追加のシグナルや処理能力によって合理的に修正できる理由によりベースラインが機能していないことが示されるまで、モデルの複雑さを増さないでください。

エラー分析ループ

各ベースライン、トレーニング実行、閾値変更、または設定変更の後に:

  1. 誤りを、偽陽性、偽陰性、判定保留、信頼度の低いケース、およびシステム障害に分類する。
  2. 言語、エンティティタイプ、ソース、時間、地理、デバイス、スパース性、最新性、特徴量の鮮度、ラベルソース、またはモデルバージョンといった共通の特徴に基づいてエラーをクラスタリングする。
  3. モデルの誤りを、データの不具合、ラベルの曖昧さ、製品の曖昧さ、計測の不備、およびサービングの不整合から区別する。
  4. 各主要なクラスターについて、次の4つの対策のいずれかに結びつける:ラベルの改善、特徴量の改善、閾値/設定の改善、または製品のフォールバックの改善。
  5. 重要なエラーはすべて、回帰テスト、評価スライス、ダッシュボードパネル、またはランブックのエントリとして保存する。
  6. 次の反復を、曖昧な「モデルの改善」タスクではなく、反証可能な実験として策定する。

最も強力なMLEループは、「トレーニング → メトリクス → デプロイ」ではありません。「エラー → クラスター → 仮説 → 実験 → 証拠 → よりシンプルなシステム」です。

観察記録

コード、プルリクエスト、実験レポート、またはランブックの横に、簡潔な意思決定と証拠の記録を残しておく:

Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:

この台帳を活用して、モデルの成果を累積させていく。目標は、単に新たな成果物を作り出すことではなく、各イテレーションが次の意思決定を容易にすることにある。

中核となるワークフロー

1. 予測契約を定義する

モデルコードを記述する前に、製品レベルの契約を明確にします:

  • 予測対象と意思決定責任者
  • 入力エンティティ、出力スキーマ、信頼度/キャリブレーションフィールド、および許容レイテンシ
  • バッチ、オンライン、ストリーミング、またはハイブリッドのサービングモード
  • モデル、フィーチャーストア、または依存関係が利用できない場合のフォールバック動作
  • 影響の大きい決定に対する人的レビューまたは上書きの経路
  • 入力、予測、およびラベルに関するプライバシー、保存期間、および監査要件

「モデルの改善」を要件として受け入れてはならない。モデルを、観察可能な製品の挙動および測定可能な受け入れ基準に結びつけること。

2. データ契約を確定する

すべての機械学習タスクには、明示的なデータ契約が必要です:

  • エンティティの粒度と主キー
  • ラベルの定義、ラベルのタイムスタンプ、およびラベルの利用可能までの遅延
  • 特徴量のタイムスタンプ、鮮度SLA、および特定時点での結合ルール
  • トレーニング、検証、テスト、およびバックテストの分割ポリシー
  • 必須列、NULL許容、範囲、カテゴリ、および単位
  • トレーニング用アーティファクトやログに含めてはならないPIIまたは機密フィールド
  • 再現性確保のためのデータセットバージョンまたはスナップショットID

まず情報漏洩の防止を最優先してください。予測時に利用できない特徴量、または将来の情報を使用して結合された特徴量がある場合は、それを削除するか、分析専用パスに移動してください。

3. 再現性のあるパイプラインを構築する

トレーニングコードは、ノートブックの隠れた状態に影響されることなく、他のエンジニアでも実行可能であるべきです:

  • すべてのハイパーパラメータおよびパスには、型指定された設定ファイルまたはデータクラスを使用する
  • パッケージおよびモデルの依存関係を固定する
  • ランダムシードを設定し、非決定論的なGPUの挙動をすべて文書化する
  • データセットのバージョン、コードのSHA、設定ファイルのハッシュ、メトリクス、およびアーティファクトのURIを記録する
  • 前処理ロジックは、ノートブックに個別に保存するのではなく、モデルアーティファクトと共に保存する
  • トレーニング、評価、推論の変換処理は共有するか、単一のソースから生成するようにする
  • 再試行によってアーティファクトやメトリクスが破損しないよう、すべてのステップを冪等にする

不変の値と純粋な変換関数を優先する。特徴量生成中に、共有データフレームやグローバル設定を変更することは避ける。

import hashlib
from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class TrainingConfig:
    dataset_uri: str
    model_dir: Path
    seed: int
    learning_rate: float
    batch_size: int


def artifact_name(config: TrainingConfig, code_sha: str) -> str:
    config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
    config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
    return f"{code_sha[:12]}-{config_hash}"

4. プロモーション前の評価

プロモーションの基準は、トレーニングが完了する前に宣言しておく必要があります:

  • ベースラインモデルと現在の本番モデルの比較
  • 製品の挙動に即した主要なメトリクス
  • レイテンシ、キャリブレーション、公平性のスライス、コスト、およびエラーの集中度に関するガードレール指標
  • 重要なコホート、地域、デバイス、言語、またはデータソースごとのスライス指標
  • 指標にノイズがある場合の信頼区間または反復実行によるばらつき
  • 影響の大きいモデルについては、人間による失敗事例のレビュー
  • 明示的な「リリース禁止」閾値
PROMOTION_GATES = {
    "auc": ("min", 0.82),
    "calibration_error": ("max", 0.04),
    "p95_latency_ms": ("max", 80),
}


def assert_promotion_ready(metrics: dict[str, float]) -> None:
    missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
    if missing:
        raise ValueError(f"Model promotion metrics missing required gates: {missing}")

    failures = {
        name: value
        for name, (direction, threshold) in PROMOTION_GATES.items()
        for value in [metrics[name]]
        if (direction == "min" and value < threshold)
        or (direction == "max" and value > threshold)
    }
    if failures:
        raise ValueError(f"Model failed promotion gates: {failures}")

オフライン指標は「保証」ではなく「ゲート」として使用してください。モデルが製品の挙動を変更する場合は、本格展開の前にシャドウ評価、カナリア展開、またはA/Bテストを計画してください。

5. サービング用パッケージ化

MLアーティファクトは、サービング契約がテスト可能である場合にのみ、本番環境での運用が可能となります:

  • モデルアーティファクトには、バージョン、トレーニングデータの参照情報、設定、前処理が含まれる
  • 入力スキーマは、無効、古くなった、または範囲外のフィーチャーを拒否する
  • 出力スキーマには、モデルバージョンに加え、必要に応じて信頼度や説明フィールドを含める
  • サービングパスには、タイムアウト、バッチ処理、リソース制限、およびフォールバック動作が設定されている
  • CPU/GPUの要件は明示されており、テスト済みです
  • 予測ログは個人を特定できる情報(PII)を排除し、デバッグおよびラベル結合に必要な識別子を十分に含んでいる
  • 統合テストでは、欠落した特徴量、古い特徴量、不正な型、空のバッチ、およびフォールバック処理を網羅している

同等性を証明するテストなしに、トレーニング専用機能コードとサービング用機能コードが乖離することのないようにしてください。

6. モデルの運用

モデルの監視には、システムおよび品質に関する指標の両方が必要です:

  • 可用性、エラー率、タイムアウト率、キューの深さ、およびp50/p95/p99レイテンシ
  • 特徴量のヌル率、範囲のドリフト、カテゴリのドリフト、および鮮度のドリフト
  • 予測分布のドリフトおよび信頼区間分布のドリフト
  • ラベル到着の健全性および遅延に関する品質指標
  • ビジネスKPIのガードレールおよびロールバックトリガー
  • カナリア展開およびロールバック用のバージョン別ダッシュボード

すべてのデプロイには、前回のアーティファクト、設定、データ依存関係、およびトラフィック切り替えの仕組みを明記したロールバック計画が必要です。

レビューチェックリスト

  • 予測契約は明示的かつ検証可能である
  • データ契約では、エンティティの粒度、ラベルのタイミング、機能のタイミング、およびスナップショット/バージョンが定義されている
  • 予測時点の可用性に対して、リークリスクが確認されている
  • トレーニングは、コード、設定、データバージョン、およびシードから再現可能である
  • メトリクスは、ベースラインおよび現在の本番モデルと比較されている
  • 高リスクのコホートに対しては、スライス指標とガードレールが組み込まれている
  • プロモーションゲートは自動化されており、失敗時はクローズされる
  • トレーニングおよびサービングの変換処理は共有されるか、同等性テストが実施される
  • モデルアーティファクトには、バージョン、設定、データセットの参照、および前処理情報が含まれる
  • サービングパスでは入力の検証が行われ、タイムアウト、フォールバック、ロールバックの処理が実装されている
  • モニタリングでは、システムの健全性、特徴量のドリフト、予測のドリフト、およびラベルの遅延をカバーしています
  • 機密データは、アーティファクト、ログ、プロンプト、および例から除外される

アンチパターン

  • モデルを再現するにはノートブックの状態が必要
  • ランダム分割により、将来のデータが検証セットまたはテストセットに漏洩する
  • 特徴量の結合処理において、イベント時刻やラベルの可用性が考慮されていない
  • 重要なスライスで性能が低下しているにもかかわらず、オフライン指標は改善している
  • 閾値がテストセット上で繰り返し調整される
  • トレーニングの前処理が、手動でサービングコードにコピーされている
  • 予測ログにモデルバージョンが記載されていない
  • モニタリングではサービスの稼働時間のみがチェックされ、データや予測の品質は確認されていない
  • ロールバックには、正常なことが確認済みのアーティファクトへの切り替えではなく、再学習が必要となる

期待される成果

このスキルを使用する際は、データ契約、プロモーションゲート、パイプラインのステップ、テスト計画、デプロイ計画、レビュー結果といった具体的な成果物を返すこと。本番環境への移行を妨げる未知の要素については、仮定で埋めるのではなく、明確に指摘すること。

GitHubで見る
---
name: mle-workflow
description: Turn model work into a production ML system with data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.
---

# Machine Learning Engineering Workflow

Use this skill to turn model work into a production ML system with clear data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.

## When to Activate

- Planning or reviewing a production ML feature, model refresh, ranking system, recommender, classifier, embedding workflow, or forecasting pipeline
- Converting notebook code into a reusable training, evaluation, batch inference, or online inference pipeline
- Designing model promotion criteria, offline/online evals, experiment tracking, or rollback paths
- Debugging failures caused by data drift, label leakage, stale features, artifact mismatch, or inconsistent training and serving logic
- Adding model monitoring, canary rollout, shadow traffic, or post-deploy quality checks

## Scope Calibration

Use only the lanes that fit the system in front of you. This skill is useful for ranking, search, recommendations, classifiers, forecasting, embeddings, LLM workflows, anomaly detection, and batch analytics, but it should not force one architecture onto all of them.

- Do not assume every model has supervised labels, online serving, a feature store, PyTorch, GPUs, human review, A/B tests, or real-time feedback.
- Do not add heavyweight MLOps machinery when a data contract, baseline, eval script, and rollback note would make the change reviewable.
- Do make assumptions explicit when the project lacks labels, delayed outcomes, slice definitions, production traffic, or monitoring ownership.
- Treat examples as interchangeable scaffolds. Replace metrics, serving mode, data stores, and rollout mechanics with the project-native equivalents.

## Related Skills

- `python-patterns` and `python-testing` for Python implementation and pytest coverage
- `pytorch-patterns` for deep learning models, data loaders, device handling, and training loops
- `eval-harness` and `ai-regression-testing` for promotion gates and agent-assisted regression checks
- `database-migrations`, `postgres-patterns`, and `clickhouse-io` for data storage and analytics surfaces
- `deployment-patterns`, `docker-patterns`, and `security-review` for serving, secrets, containers, and production hardening

## Reuse the SWE Surface

Do not treat MLE as separate from software engineering. Most ECC SWE workflows apply directly to ML systems, often with stricter failure modes:

The recommended `minimal --with capability:machine-learning` install keeps the core agent surface available alongside this skill. For skill-only or agent-limited harnesses, pair `skill:mle-workflow` with `agent:mle-reviewer` where the target supports agents.

| SWE surface | MLE use |
|-------------|---------|
| `product-capability` / `architecture-decision-records` | Turn model work into explicit product contracts and record irreversible data, model, and rollout choices |
| `repo-scan` / `codebase-onboarding` / `code-tour` | Find existing training, feature, serving, eval, and monitoring paths before introducing a parallel ML stack |
| `plan` / `feature-dev` | Scope model changes as product capabilities with data, eval, serving, and rollback phases |
| `tdd-workflow` / `python-testing` | Test feature transforms, split logic, metric calculations, artifact loading, and inference schemas before implementation |
| `code-reviewer` / `mle-reviewer` | Review code quality plus ML-specific leakage, reproducibility, promotion, and monitoring risks |
| `build-fix` / `pr-test-analyzer` | Diagnose broken CI, flaky evals, missing fixtures, and environment-specific model or dependency failures |
| `quality-gate` / `test-coverage` | Require automated evidence for transforms, metrics, inference contracts, promotion gates, and rollback behavior |
| `eval-harness` / `verification-loop` | Turn offline metrics, slice checks, latency budgets, and rollback drills into repeatable gates |
| `ai-regression-testing` | Preserve every production bug as a regression: missing feature, stale label, bad artifact, schema drift, or serving mismatch |
| `api-design` / `backend-patterns` | Design prediction APIs, batch jobs, idempotent retraining endpoints, and response envelopes |
| `database-migrations` / `postgres-patterns` / `clickhouse-io` | Version labels, feature snapshots, prediction logs, experiment metrics, and drift analytics |
| `deployment-patterns` / `docker-patterns` | Package reproducible training and serving images with health checks, resource limits, and rollback |
| `canary-watch` / `dashboard-builder` | Make rollout health visible with model-version, slice, drift, latency, cost, and delayed-label dashboards |
| `security-review` / `security-scan` | Check model artifacts, notebooks, prompts, datasets, and logs for secrets, PII, unsafe deserialization, and supply-chain risk |
| `e2e-testing` / `browser-qa` / `accessibility` | Test critical product flows that consume predictions, including explainability and fallback UI states |
| `benchmark` / `performance-optimizer` | Measure throughput, p95 latency, memory, GPU utilization, and cost per prediction or retrain |
| `cost-aware-llm-pipeline` / `token-budget-advisor` | Route LLM/embedding workloads by quality, latency, and budget instead of defaulting to the largest model |
| `documentation-lookup` / `search-first` | Verify current library behavior for model serving, feature stores, vector DBs, and eval tooling before coding |
| `git-workflow` / `github-ops` / `opensource-pipeline` | Package MLE changes for review with crisp scope, generated artifacts excluded, and reproducible test evidence |
| `strategic-compact` / `dmux-workflows` | Split long ML work into parallel tracks: data contract, eval harness, serving path, monitoring, and docs |

## Ten MLE Task Simulations

Use these simulations as coverage checks when planning or reviewing MLE work. A strong MLE workflow should reduce each task to explicit contracts, reusable SWE surfaces, automated evidence, and a reviewable artifact.

| ID | Common MLE task | Streamlined ECC path | Required output | Pipeline lanes covered |
|----|-----------------|----------------------|-----------------|------------------------|
| MLE-01 | Frame an ambiguous prediction, ranking, recommender, classifier, embedding, or forecast capability | `product-capability`, `plan`, `architecture-decision-records`, `mle-workflow` | Iteration Compact naming who cares, decision owner, success metric, unacceptable mistakes, assumptions, constraints, and first experiment | product contract, stakeholder loss, risk, rollout |
| MLE-02 | Define metric goals, labels, data sources, and the mistake budget | `repo-scan`, `database-reviewer`, `database-migrations`, `postgres-patterns`, `clickhouse-io` | Data and metric contract with entity grain, label timing, label confidence, feature timing, point-in-time joins, split policy, and dataset snapshot | data contract, metric design, leakage, reproducibility |
| MLE-03 | Build a baseline model and scoring path before adding complexity | `tdd-workflow`, `python-testing`, `python-patterns`, `code-reviewer` | Baseline scorer with confusion matrix, calibration notes, latency/cost estimate, known weaknesses, and tests for score shape and determinism | baseline, scoring, testing, serving parity |
| MLE-04 | Generate features from hypotheses about what separates outcomes | `python-patterns`, `pytorch-patterns`, `docker-patterns`, `deployment-patterns` | Feature plan and transform module covering signal source, missing values, outliers, correlations, leakage checks, and train/serve equivalence | feature pipeline, leakage, training, artifacts |
| MLE-05 | Tune thresholds, configs, and model complexity under tradeoffs | `eval-harness`, `ai-regression-testing`, `quality-gate`, `test-coverage` | Threshold/config report comparing precision, recall, F1, AUC, calibration, group slices, latency, cost, complexity, and acceptable error classes | evaluation, threshold, promotion, regression |
| MLE-06 | Run error analysis and turn mistakes into the next experiment | `eval-harness`, `ai-regression-testing`, `mle-reviewer`, `silent-failure-hunter` | Error cluster report for false positives, false negatives, ambiguous labels, stale features, missing signals, and bug traces with lessons captured | error analysis, bug trace, iteration, regression |
| MLE-07 | Package a model artifact for batch or online inference | `api-design`, `backend-patterns`, `security-review`, `security-scan` | Versioned artifact bundle with preprocessing, config, dependency constraints, schema validation, safe loading, and PII-safe logs | artifact, security, inference contract |
| MLE-08 | Ship online serving or batch scoring with feedback capture | `api-design`, `backend-patterns`, `e2e-testing`, `browser-qa`, `accessibility` | Prediction endpoint or batch job with response envelope, timeout, batching, fallback, model version, confidence, feedback logging, and product-flow tests | serving, batch inference, fallback, user workflow |
| MLE-09 | Roll out a model with shadow traffic, canary, A/B test, or rollback | `canary-watch`, `dashboard-builder`, `verification-loop`, `performance-optimizer` | Rollout plan naming traffic split, dashboards, p95 latency, cost, quality guardrails, rollback artifact, and rollback trigger | deployment, canary, rollback |
| MLE-10 | Operate, debug, and refresh a production model after launch | `silent-failure-hunter`, `dashboard-builder`, `mle-reviewer`, `doc-updater`, `github-ops` | Observation ledger and refresh plan with drift checks, delayed-label health, alert owners, runbook updates, retrain criteria, and PR evidence | monitoring, incident response, retraining |

## Iteration Compact

Before touching model code, compress the work into one reviewable artifact. This should be short enough to fit in a PR description and precise enough that another engineer can challenge the tradeoffs.

```text
Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:
```

This compact is the MLE equivalent of a strong SWE design note. It keeps the team from optimizing a metric no one trusts, adding features that do not address the real error mode, or shipping complexity without a rollback.

## Decision Brain

Use this loop whenever the task is ambiguous, high-impact, or metric-heavy:

1. Start from the decision, not the model. Name the action that changes downstream behavior.
2. Name who cares and why. Different stakeholders pay different costs for false positives, false negatives, latency, compute spend, opacity, or missed opportunities.
3. Convert ambiguity into hypotheses. Ask what signal would separate outcomes, what evidence would disprove it, and what simple baseline should be hard to beat.
4. Research prior art or a nearby known problem before inventing a bespoke system.
5. Score choices with `(probability, confidence) x (cost, severity, importance, impact)`.
6. Consider adversarial behavior, incentives, selective disclosure, distribution shift, and feedback loops.
7. Prefer the simplest change that reduces the most important mistake. Simplicity is not laziness; it is a way to minimize blunders while preserving iteration speed.
8. Capture the decision, evidence, counterargument, and next reversible step.

## Metric and Mistake Economics

Choose metrics from failure costs, not habit:

- Use a confusion matrix early so the team can discuss concrete false positives and false negatives instead of abstract accuracy.
- Favor precision when the cost of an incorrect positive decision dominates.
- Favor recall when the cost of a missed positive dominates.
- Use F1 only when the precision/recall tradeoff is genuinely balanced and explainable.
- Use AUC or ranking metrics when ordering quality matters more than a single threshold.
- Track latency, throughput, memory, and cost as first-class metrics because they shape feasible model complexity.
- Compare against a baseline and the current production model before celebrating an offline gain.
- Treat real-world feedback signals as delayed labels with bias, lag, and coverage gaps; do not treat them as ground truth without analysis.

Every metric choice should state which mistake it makes cheaper, which mistake it makes more likely, and who absorbs that cost.

## Data and Feature Hypotheses

Features should come from a theory of separation:

- Text, categorical fields, numeric histories, graph relationships, recency, frequency, and aggregates are candidate signal families, not automatic features.
- For every feature family, state why it should separate outcomes and how it could leak future information.
- For noisy labels, consider adjudication, label confidence, soft targets, or confidence weighting.
- For class imbalance, compare weighted loss, resampling, threshold movement, and calibrated decision rules.
- For missing values, decide whether absence is informative, imputable, or a reason to abstain.
- For outliers, decide whether to clip, bucket, investigate, or preserve them as rare but important signal.
- For correlated features, check whether they are redundant, unstable, or proxies for unavailable future state.

Do not add model complexity until error analysis shows that the baseline is failing for a reason additional signal or capacity can plausibly fix.

## Error Analysis Loop

After each baseline, training run, threshold change, or config change:

1. Split mistakes into false positives, false negatives, abstentions, low-confidence cases, and system failures.
2. Cluster errors by shared traits: language, entity type, source, time, geography, device, sparsity, recency, feature freshness, label source, or model version.
3. Separate model mistakes from data bugs, label ambiguity, product ambiguity, instrumentation gaps, and serving mismatches.
4. Trace each major cluster to one of four moves: better labels, better features, better threshold/config, or better product fallback.
5. Preserve every important mistake as a regression test, eval slice, dashboard panel, or runbook entry.
6. Write the next iteration as a falsifiable experiment, not a vague "improve model" task.

The strongest MLE loop is not train -> metric -> ship. It is mistake -> cluster -> hypothesis -> experiment -> evidence -> simpler system.

## Observation Ledger

Keep a compact decision and evidence trail beside the code, PR, experiment report, or runbook:

```text
Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:
```

Use the ledger to make model work cumulative. The goal is for each iteration to make the next decision easier, not merely to produce another artifact.

## Core Workflow

### 1. Define the Prediction Contract

Capture the product-level contract before writing model code:

- Prediction target and decision owner
- Input entity, output schema, confidence/calibration fields, and allowed latency
- Batch, online, streaming, or hybrid serving mode
- Fallback behavior when the model, feature store, or dependency is unavailable
- Human review or override path for high-impact decisions
- Privacy, retention, and audit requirements for inputs, predictions, and labels

Do not accept "improve the model" as a requirement. Tie the model to an observable product behavior and a measurable acceptance gate.

### 2. Lock the Data Contract

Every ML task needs an explicit data contract:

- Entity grain and primary key
- Label definition, label timestamp, and label availability delay
- Feature timestamp, freshness SLA, and point-in-time join rules
- Train, validation, test, and backtest split policy
- Required columns, allowed nulls, ranges, categories, and units
- PII or sensitive fields that must not enter training artifacts or logs
- Dataset version or snapshot ID for reproducibility

Guard against leakage first. If a feature is not available at prediction time, or is joined using future information, remove it or move it to an analysis-only path.

### 3. Build a Reproducible Pipeline

Training code should be runnable by another engineer without hidden notebook state:

- Use typed config files or dataclasses for all hyperparameters and paths
- Pin package and model dependencies
- Set random seeds and document any nondeterministic GPU behavior
- Record dataset version, code SHA, config hash, metrics, and artifact URI
- Save preprocessing logic with the model artifact, not separately in a notebook
- Keep train, eval, and inference transformations shared or generated from one source
- Make every step idempotent so retries do not corrupt artifacts or metrics

Prefer immutable values and pure transformation functions. Avoid mutating shared data frames or global config during feature generation.

```python
import hashlib
from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class TrainingConfig:
    dataset_uri: str
    model_dir: Path
    seed: int
    learning_rate: float
    batch_size: int


def artifact_name(config: TrainingConfig, code_sha: str) -> str:
    config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
    config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
    return f"{code_sha[:12]}-{config_hash}"
```

### 4. Evaluate Before Promotion

Promotion criteria should be declared before training finishes:

- Baseline model and current production model comparison
- Primary metric aligned to product behavior
- Guardrail metrics for latency, calibration, fairness slices, cost, and error concentration
- Slice metrics for important cohorts, geographies, devices, languages, or data sources
- Confidence intervals or repeated-run variance when metrics are noisy
- Failure examples reviewed by a human for high-impact models
- Explicit "do not ship" thresholds

```python
PROMOTION_GATES = {
    "auc": ("min", 0.82),
    "calibration_error": ("max", 0.04),
    "p95_latency_ms": ("max", 80),
}


def assert_promotion_ready(metrics: dict[str, float]) -> None:
    missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
    if missing:
        raise ValueError(f"Model promotion metrics missing required gates: {missing}")

    failures = {
        name: value
        for name, (direction, threshold) in PROMOTION_GATES.items()
        for value in [metrics[name]]
        if (direction == "min" and value < threshold)
        or (direction == "max" and value > threshold)
    }
    if failures:
        raise ValueError(f"Model failed promotion gates: {failures}")
```

Use offline metrics as gates, not guarantees. When the model changes product behavior, plan shadow evaluation, canary rollout, or A/B testing before full rollout.

### 5. Package for Serving

An ML artifact is production-ready only when the serving contract is testable:

- Model artifact includes version, training data reference, config, and preprocessing
- Input schema rejects invalid, stale, or out-of-range features
- Output schema includes model version and confidence or explanation fields when useful
- Serving path has timeout, batching, resource limits, and fallback behavior
- CPU/GPU requirements are explicit and tested
- Prediction logs avoid PII and include enough identifiers for debugging and label joins
- Integration tests cover missing features, stale features, bad types, empty batches, and fallback path

Never let training-only feature code diverge from serving feature code without a test that proves equivalence.

### 6. Operate the Model

Model monitoring needs both system and quality signals:

- Availability, error rate, timeout rate, queue depth, and p50/p95/p99 latency
- Feature null rate, range drift, categorical drift, and freshness drift
- Prediction distribution drift and confidence distribution drift
- Label arrival health and delayed quality metrics
- Business KPI guardrails and rollback triggers
- Per-version dashboards for canaries and rollbacks

Every deployment should have a rollback plan that names the previous artifact, config, data dependency, and traffic-switch mechanism.

## Review Checklist

- [ ] Prediction contract is explicit and testable
- [ ] Data contract defines entity grain, label timing, feature timing, and snapshot/version
- [ ] Leakage risks were checked against prediction-time availability
- [ ] Training is reproducible from code, config, data version, and seed
- [ ] Metrics compare against baseline and current production model
- [ ] Slice metrics and guardrails are included for high-risk cohorts
- [ ] Promotion gates are automated and fail closed
- [ ] Training and serving transformations are shared or equivalence-tested
- [ ] Model artifact carries version, config, dataset reference, and preprocessing
- [ ] Serving path validates inputs and has timeout, fallback, and rollback behavior
- [ ] Monitoring covers system health, feature drift, prediction drift, and delayed labels
- [ ] Sensitive data is excluded from artifacts, logs, prompts, and examples

## Anti-Patterns

- Notebook state is required to reproduce the model
- Random split leaks future data into validation or test sets
- Feature joins ignore event time and label availability
- Offline metric improves while important slices regress
- Thresholds are tuned on the test set repeatedly
- Training preprocessing is copied manually into serving code
- Model version is missing from prediction logs
- Monitoring only checks service uptime, not data or prediction quality
- Rollback requires retraining instead of switching to a known-good artifact

## Output Expectations

When using this skill, return concrete artifacts: data contract, promotion gates, pipeline steps, test plan, deployment plan, or review findings. Call out unknowns that block production readiness instead of filling them with assumptions.

すべてのファイル

1件のファイル

mle-workflowをインストール

スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。

ZIPをダウンロード

リポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。

git clone https://github.com/affaan-m/ECC/tree/main/skills/mle-workflow # Copy SKILL.md to your .claude/skills/ directory

コピー コピー
クイックセットアップ: スキルフォルダを .claude/skills/ にコピーしてください。 Claude が自動的にそのスキルを検出して使用します。
リポジトリ affaan-m/ECC

関連スキル

web-search
更新された時間 2026年6月29日
webapp-testing
更新された時間 2026年6月29日
lark-base
更新された時間 2026年7月5日
agentmail
更新された時間 2026年6月29日
OR