mle-workflow
affaan-m/ECC
데이터 계약, 반복 가능한 훈련, 측정 가능한 품질 게이트, 배포 가능한 아티팩트 및 운영 모니터링을 통해 모델 작업을 실제 ML 시스템으로 전환하십시오.
...모든 것을 확장하십시오머신러닝 엔지니어링 워크플로우
이 기술을 활용하여 모델 작업을 명확한 데이터 계약, 재현 가능한 훈련, 측정 가능한 품질 게이트, 배포 가능한 아티팩트 및 운영 모니터링을 갖춘 프로덕션 ML 시스템으로 전환하십시오.
활성화 시점
- 프로덕션 ML 기능, 모델 갱신, 순위 결정 시스템, 추천 시스템, 분류기, 임베딩 워크플로우 또는 예측 파이프라인을 계획하거나 검토할 때
- 노트북 코드를 재사용 가능한 훈련, 평가, 배치 추론 또는 온라인 추론 파이프라인으로 변환할 때
- 모델 적용 기준, 오프라인/온라인 평가, 실험 추적 또는 롤백 경로를 설계할 때
- 데이터 드리프트, 레이블 유출, 오래된 특징, 아티팩트 불일치, 또는 훈련 및 서빙 로직의 불일치로 인한 오류 디버깅
- 모델 모니터링, 카나리아 배포, 섀도우 트래픽 또는 배포 후 품질 점검을 추가할 때
범위 조정
현재 시스템에 적합한 레인만 사용하십시오. 이 기술은 랭킹, 검색, 추천, 분류기, 예측, 임베딩, LLM 워크플로우, 이상 탐지 및 배치 분석에 유용하지만, 모든 경우에 단일 아키텍처를 강제로 적용해서는 안 됩니다.
- 모든 모델이 지도 학습 라벨, 온라인 서빙, 특징 저장소, PyTorch, GPU, 수동 검토, A/B 테스트 또는 실시간 피드백을 갖추고 있다고 가정하지 마십시오.
- 데이터 계약, 베이스라인, 평가 스크립트, 롤백 메모만으로도 변경 사항을 검토할 수 있는 경우에는 무거운 MLOps 시스템을 추가하지 마십시오.
- 프로젝트에 레이블, 지연된 결과, 슬라이스 정의, 프로덕션 트래픽 또는 모니터링 책임자가 없는 경우에는 가정을 명확히 밝혀야 합니다.
- 예시는 서로 바꿔 쓸 수 있는 기본 틀로 간주하십시오. 지표, 서빙 모드, 데이터 저장소, 롤아웃 메커니즘은 해당 프로젝트에 적합한 것으로 대체하십시오.
관련 기술
python-patterns및python-testingPython 구현 및 pytest 커버리지pytorch-patterns딥러닝 모델, 데이터 로더, 디바이스 처리 및 훈련 루프eval-harness그리고ai-regression-testing프로모션 게이트 및 에이전트 지원 회귀 검사database-migrations,postgres-patterns, 그리고clickhouse-io데이터 저장 및 분석 인터페이스deployment-patterns,docker-patterns, 그리고security-review서비스, 시크릿, 컨테이너 및 프로덕션 환경 강화
SWE 영역 재사용
MLE를 소프트웨어 엔지니어링과 별개의 것으로 취급하지 마십시오. 대부분의 ECC SWE 워크플로는 ML 시스템에 직접 적용되며, 종종 더 엄격한 오류 모드를 수반합니다:
권장되는 minimal --with capability:machine-learning 설치 방식은 이 스킬과 함께 핵심 에이전트 표면을 계속 사용할 수 있게 합니다. 스킬 전용 또는 에이전트가 제한된 하네스의 경우, skill:mle-workflow 다음과 함께 사용하십시오 agent:mle-reviewer 대상 시스템이 에이전트를 지원하는 경우.
| SWE 인터페이스 | MLE 사용 |
|---|---|
product-capability / architecture-decision-records |
모델 작업을 명시적인 제품 계약으로 전환하고, 되돌릴 수 없는 데이터, 모델 및 롤아웃 선택 사항을 기록하십시오 |
repo-scan / codebase-onboarding / code-tour |
병렬 ML 스택을 도입하기 전에 기존의 훈련, 특징 추출, 서빙, 평가 및 모니터링 경로를 파악합니다 |
plan / feature-dev |
모델 변경 사항을 데이터, 평가, 서비스 제공 및 롤백 단계를 포함하는 제품 기능으로 범위를 정의 |
tdd-workflow / python-testing |
구현 전에 특징 변환, 분할 로직, 지표 계산, 아티팩트 로딩 및 추론 스키마를 테스트하십시오 |
code-reviewer / mle-reviewer |
코드 품질은 물론 ML 특유의 누수, 재현성, 배포 및 모니터링 위험을 검토하십시오 |
build-fix / pr-test-analyzer |
오작동하는 CI, 불안정한 평가, 누락된 고정 장치, 환경별 모델 또는 종속성 오류를 진단하십시오 |
quality-gate / test-coverage |
변환, 지표, 추론 계약, 배포 게이트 및 롤백 동작에 대한 자동화된 증거를 요구하십시오 |
eval-harness / verification-loop |
오프라인 메트릭, 슬라이스 검사, 지연 시간 예산 및 롤백 시뮬레이션을 반복 가능한 게이트로 전환 |
ai-regression-testing |
누락된 기능, 오래된 레이블, 불량 아티팩트, 스키마 드리프트, 서빙 불일치 등 모든 프로덕션 버그를 회귀 문제로 기록하십시오 |
api-design / backend-patterns |
예측 API, 배치 작업, 항등적 재훈련 엔드포인트 및 응답 엔벨로프 설계 |
database-migrations / postgres-patterns / clickhouse-io |
버전 라벨, 기능 스냅샷, 예측 로그, 실험 메트릭 및 드리프트 분석 |
deployment-patterns / docker-patterns |
재현 가능한 훈련 및 서빙 이미지를 상태 점검, 리소스 제한 및 롤백 기능과 함께 패키징 |
canary-watch / dashboard-builder |
모델 버전, 슬라이스, 드리프트, 지연 시간, 비용 및 지연된 레이블 대시보드를 통해 롤아웃 상태를 가시화하세요 |
security-review / security-scan |
모델 아티팩트, 노트북, 프롬프트, 데이터셋 및 로그에서 기밀 정보, 개인 식별 정보(PII), 안전하지 않은 역직렬화 및 공급망 위험을 점검하십시오 |
e2e-testing / browser-qa / accessibility |
설명 가능성 및 대체 UI 상태를 포함하여 예측을 활용하는 핵심 제품 흐름 테스트 |
benchmark / performance-optimizer |
처리량, p95 지연 시간, 메모리, GPU 사용률, 예측 또는 재훈련당 비용을 측정 |
cost-aware-llm-pipeline / token-budget-advisor |
기본적으로 가장 큰 모델을 사용하는 대신, 품질, 지연 시간 및 예산을 기준으로 LLM/임베딩 워크로드를 라우팅 |
documentation-lookup / search-first |
코딩 전에 모델 서빙, 피처 스토어, 벡터 DB 및 평가 도구에 대한 현재 라이브러리 동작을 확인하십시오 |
git-workflow / github-ops / opensource-pipeline |
명확한 범위, 생성된 아티팩트 제외, 재현 가능한 테스트 증거를 포함하여 검토를 위해 MLE 변경 사항을 패키징 |
strategic-compact / dmux-workflows |
긴 ML 작업을 데이터 계약, 평가 하네스, 서빙 경로, 모니터링, 문서 등 병렬 트랙으로 분할합니다. |
10가지 MLE 작업 시뮬레이션
MLE 작업을 계획하거나 검토할 때 이러한 시뮬레이션을 커버리지 점검 용도로 활용하십시오. 견고한 MLE 워크플로는 각 태스크를 명확한 계약, 재사용 가능한 SWE 인터페이스, 자동화된 증거, 검토 가능한 산출물로 정리해야 합니다.
| ID | 일반적인 MLE 작업 | 간소화된 ECC 경로 | 필수 산출물 | 포괄되는 파이프라인 레인 |
|---|---|---|---|---|
| MLE-01 | 모호한 예측, 순위 지정, 추천, 분류, 임베딩 또는 예측 기능을 정의하십시오 | product-capability, plan, architecture-decision-records, mle-workflow |
반복 관련자, 의사 결정권자, 성공 지표, 용납할 수 없는 실수, 가정, 제약 조건 및 첫 번째 실험을 간결하게 명명 | 제품 계약, 이해관계자 손실, 위험, 출시 |
| MLE-02 | 지표 목표, 레이블, 데이터 소스 및 오류 허용 한도를 정의합니다 | repo-scan, database-reviewer, database-migrations, postgres-patterns, clickhouse-io |
엔티티 단위, 레이블 타이밍, 레이블 신뢰도, 특징 타이밍, 특정 시점 조인, 분할 정책 및 데이터 세트 스냅샷을 포함한 데이터 및 지표 계약 | 데이터 계약, 지표 설계, 누수, 재현성 |
| MLE-03 | 복잡성을 추가하기 전에 기준 모델 및 스코어링 경로 구축 | tdd-workflow, python-testing, python-patterns, code-reviewer |
혼동 행렬, 보정 참고 사항, 지연 시간/비용 추정치, 알려진 약점, 점수 형태 및 결정성에 대한 테스트를 포함한 기준 스코어러 | 기준, 스코어링, 테스트, 서비스 패리티 |
| MLE-04 | 결과를 구분하는 요인에 대한 가설로부터 특징을 생성 | python-patterns, pytorch-patterns, docker-patterns, deployment-patterns |
신호 원천, 누락된 값, 이상치, 상관관계, 누수 점검, 훈련/서비스 동등성을 다루는 특징 계획 및 변환 모듈 | 특성 파이프라인, 누출, 훈련, 아티팩트 |
| MLE-05 | 상호 절충 관계를 고려하여 임계값, 구성 및 모델 복잡도 조정 | eval-harness, ai-regression-testing, quality-gate, test-coverage |
정밀도, 재현율, F1, AUC, 보정, 그룹 슬라이스, 지연 시간, 비용, 복잡도 및 허용 오차 클래스를 비교하는 임계값/구성 보고서 | 평가, 임계값, 승격, 회귀 |
| MLE-06 | 오류 분석을 수행하고 실수를 다음 실험으로 전환 | eval-harness, ai-regression-testing, mle-reviewer, silent-failure-hunter |
오탐지, 누락, 모호한 레이블, 유효 기간이 지난 특징, 누락된 신호 및 버그 추적에 대한 오류 클러스터 보고서와 함께 도출된 교훈 | 오류 분석, 버그 추적, 반복, 회귀 |
| MLE-07 | 배치 또는 온라인 추론을 위한 모델 아티팩트 패키징 | api-design, backend-patterns, security-review, security-scan |
전처리, 구성, 종속성 제약 조건, 스키마 유효성 검사, 안전한 로딩 및 PII 안전 로그가 포함된 버전 관리 아티팩트 번들 | 아티팩트, 보안, 추론 계약 |
| MLE-08 | 피드백 캡처 기능을 갖춘 온라인 서비스 또는 배치 스코어링 배포 | api-design, backend-patterns, e2e-testing, browser-qa, accessibility |
응답 엔벨로프, 타임아웃, 배치 처리, 폴백, 모델 버전, 신뢰도, 피드백 로깅 및 제품 흐름 테스트가 포함된 예측 엔드포인트 또는 배치 작업 | 서빙, 배치 추론, 폴백, 사용자 워크플로 |
| MLE-09 | 섀도우 트래픽, 카나리아 테스트, A/B 테스트 또는 롤백을 통해 모델 배포 | canary-watch, dashboard-builder, verification-loop, performance-optimizer |
롤아웃 계획 명명: 트래픽 분할, 대시보드, p95 지연 시간, 비용, 품질 가드레일, 롤백 아티팩트 및 롤백 트리거 | 배포, 카나리아, 롤백 |
| MLE-10 | 출시 후 프로덕션 모델 운영, 디버깅 및 갱신 | silent-failure-hunter, dashboard-builder, mle-reviewer, doc-updater, github-ops |
드리프트 점검, 지연 라벨 상태, 경보 담당자, 런북 업데이트, 재훈련 기준 및 PR 증거를 포함한 관찰 원장 및 갱신 계획 | 모니터링, 인시던트 대응, 재훈련 |
반복 작업 요약
모델 코드를 수정하기 전에, 작업을 검토 가능한 단일 아티팩트로 압축하십시오. 이 아티팩트는 PR 설명란에 들어갈 만큼 간결해야 하며, 다른 엔지니어가 그 타협점에 대해 의문을 제기할 수 있을 만큼 정확해야 합니다.
Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:
이 요약본은 MLE 분야에서 탄탄한 SWE 설계 노트에 해당합니다. 이를 통해 팀은 아무도 신뢰하지 않는 지표를 최적화하거나, 실제 오류 모드를 해결하지 못하는 기능을 추가하거나, 롤백 방안 없이 복잡한 기능을 출시하는 것을 방지할 수 있습니다.
의사결정 브레인
작업이 모호하거나, 영향력이 크거나, 지표에 지나치게 의존하는 경우 이 루프를 사용하십시오:
- 모델이 아닌 결정부터 시작하십시오. 하류 동작을 변경하는 조치를 명시하십시오.
- 누가 관심을 갖고 왜 그런지 명시하세요. 이해관계자마다 오탐, 누락, 지연, 컴퓨팅 비용, 불투명성, 또는 놓친 기회에 대해 각기 다른 대가를 치릅니다.
- 모호함을 가설로 전환하십시오. 어떤 신호가 결과를 구분해 줄지, 어떤 증거가 이를 반증할 수 있는지, 그리고 어떤 단순한 기준선이 깨기 어려워야 하는지 자문하십시오.
- 맞춤형 시스템을 개발하기 전에 선행 기술이나 유사한 알려진 문제를 조사하십시오.
- 다음 요소를 기준으로 선택 사항을 평가하십시오:
(probability, confidence) x (cost, severity, importance, impact). - 적대적 행동, 인센티브, 선택적 정보 공개, 분포 변화, 피드백 루프 등을 고려하십시오.
- 가장 중요한 실수를 줄여주는 가장 단순한 변경을 우선시하십시오. 단순함은 게으름이 아니라, 반복 속도를 유지하면서 실수를 최소화하는 방법입니다.
- 결정 사항, 증거, 반론, 그리고 다음 가역적 단계를 기록하십시오.
지표와 실수 경제학
습관이 아닌 실패 비용에 기반하여 지표를 선택하십시오:
- 팀이 추상적인 정확도 대신 구체적인 오탐(false positive)과 누락(false negative)에 대해 논의할 수 있도록 초기 단계부터 혼동 행렬을 활용하십시오.
- 잘못된 양성 판정 결정의 비용이 압도적으로 클 때는 정밀도를 우선시하십시오.
- 양성 판정을 놓치는 비용이 더 클 때는 재현율을 우선시하십시오.
- 정밀도와 재현율 간의 절충점이 진정으로 균형을 이루고 설명 가능한 경우에만 F1 점수를 사용하십시오.
- 단일 임계값보다 순위가 더 중요한 경우에는 AUC 또는 순위 지표를 사용하십시오.
- 지연 시간, 처리량, 메모리, 비용을 1순위 지표로 추적하십시오. 이는 모델의 실현 가능한 복잡성을 결정하기 때문입니다.
- 오프라인 성능 향상을 축하하기 전에 베이스라인 및 현재 운영 중인 모델과 비교하십시오.
- 실제 피드백 신호를 편향, 지연 및 커버리지 격차가 있는 지연된 레이블로 취급하십시오. 분석 없이 이를 그라운드 트루스로 간주해서는 안 됩니다.
모든 지표 선택 시, 어떤 실수의 비용을 줄이고, 어떤 실수의 발생 가능성을 높이며, 그 비용을 누가 부담하는지 명시해야 합니다.
데이터 및 특징 가설
특징은 분리 이론에서 비롯되어야 합니다:
- 텍스트, 범주형 필드, 수치 기록, 그래프 관계, 최근성, 빈도, 집계값 등은 후보 신호 군일 뿐, 자동 생성된 특징이 아닙니다.
- 각 특징 군에 대해, 왜 그것이 결과를 분리해야 하는지, 그리고 어떻게 미래 정보를 누설할 수 있는지 명시해야 합니다.
- 라벨에 노이즈가 있는 경우, 판정(adjudication), 라벨 신뢰도, 소프트 타겟(soft targets), 또는 신뢰도 가중치를 고려해야 합니다.
- 클래스 불균형의 경우, 가중 손실, 재표본 추출, 임계값 이동 및 보정된 의사결정 규칙을 비교해야 한다.
- 누락된 값의 경우, 해당 값의 부재가 정보적 의미를 지니는지, 추정 가능한지, 아니면 판단을 보류해야 할 이유인지 결정하십시오.
- 이상치에 대해서는, 이를 잘라낼지, 버킷으로 묶을지, 조사할지, 아니면 희귀하지만 중요한 신호로 보존할지 결정하십시오.
- 상관 관계가 있는 특징의 경우, 중복인지, 불안정한지, 아니면 이용 불가능한 미래 상태를 대리하는 지표인지 확인하십시오.
오류 분석 결과, 추가적인 신호나 용량이 타당하게 해결할 수 있는 이유로 인해 기준 모델이 실패하고 있음이 확인되기 전까지는 모델의 복잡성을 높이지 마십시오.
오류 분석 루프
각 기준 모델, 훈련 실행, 임계값 변경 또는 구성 변경 후에는:
- 오류를 오탐지, 누락, 판단 보류, 신뢰도 낮은 사례, 시스템 오류로 분류하십시오.
- 언어, 엔티티 유형, 출처, 시간, 지리적 위치, 기기, 희소성, 최신성, 특징의 최신도, 라벨 출처 또는 모델 버전과 같은 공통된 특성에 따라 오류를 클러스터링하십시오.
- 모델 오류를 데이터 버그, 라벨 모호성, 제품 모호성, 계측 누락, 서비스 불일치와 구분합니다.
- 각 주요 클러스터를 ‘라벨 개선’, ‘특징량 개선’, ‘임계값/구성 개선’ 또는 ‘제품 대체 방안 개선’이라는 네 가지 조치 중 하나로 추적합니다.
- 모든 중요한 오류를 회귀 테스트, 평가 슬라이스, 대시보드 패널 또는 런북 항목으로 보존하십시오.
- 다음 반복 작업을 모호한 “모델 개선” 과제가 아닌, 반증 가능한 실험으로 설계하십시오.
가장 강력한 MLE 루프는 ‘훈련 → 지표 → 배포’가 아닙니다. 바로 ‘오류 → 클러스터 → 가설 → 실험 → 증거 → 더 단순한 시스템’입니다.
관찰 기록부
코드, PR, 실험 보고서 또는 런북 옆에 간결한 의사결정 및 증거 기록을 유지하십시오:
Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:
이 레저를 활용하여 모델 작업의 누적 효과를 창출하세요. 목표는 단순히 또 다른 산출물을 생성하는 것이 아니라, 각 반복을 통해 다음 결정을 더 쉽게 내릴 수 있도록 하는 것입니다.
핵심 워크플로우
1. 예측 계약 정의
모델 코드를 작성하기 전에 제품 수준의 계약을 정리하세요:
- 예측 대상 및 의사결정 담당자
- 입력 엔티티, 출력 스키마, 신뢰도/보정 필드 및 허용 지연 시간
- 배치, 온라인, 스트리밍 또는 하이브리드 서빙 모드
- 모델, 피처 스토어 또는 종속성을 사용할 수 없을 때의 대체 처리 방식
- 영향이 큰 결정에 대한 수동 검토 또는 재정의 경로
- 입력 데이터, 예측 결과 및 레이블에 대한 개인정보 보호, 보존 및 감사 요구 사항
“모델 개선”을 요구 사항으로 받아들여서는 안 됩니다. 모델을 관찰 가능한 제품 동작 및 측정 가능한 승인 기준에 연계하십시오.
2. 데이터 계약 확정
모든 ML 작업에는 명시적인 데이터 계약이 필요합니다:
- 엔티티 단위 및 기본 키
- 라벨 정의, 라벨 타임스탬프 및 라벨 가용성 지연
- 특성 타임스탬프, 최신성 SLA, 특정 시점 조인 규칙
- 훈련, 검증, 테스트 및 백테스트 분할 정책
- 필수 열, 허용되는 NULL 값, 범위, 범주 및 단위
- 훈련 아티팩트나 로그에 포함되어서는 안 되는 PII 또는 민감한 필드
- 재현성을 위한 데이터셋 버전 또는 스냅샷 ID
먼저 정보 유출을 방지하십시오. 예측 시점에 특징을 사용할 수 없거나 미래 정보를 사용하여 조인된 경우, 해당 특징을 제거하거나 분석 전용 경로로 이동시키십시오.
3. 재현 가능한 파이프라인 구축
훈련 코드는 숨겨진 노트북 상태 없이도 다른 엔지니어가 실행할 수 있어야 합니다:
- 모든 하이퍼파라미터와 경로에는 타입이 지정된 구성 파일이나 데이터 클래스를 사용하십시오
- 패키지 및 모델 종속성을 고정하십시오
- 난수 시드를 설정하고 비결정적인 GPU 동작을 문서화하십시오
- 데이터셋 버전, 코드 SHA, 구성 해시, 메트릭 및 아티팩트 URI를 기록하십시오
- 전처리 로직은 노트북에 별도로 저장하지 말고 모델 아티팩트와 함께 저장하십시오
- 훈련, 평가 및 추론 변환을 공유하거나 단일 소스에서 생성하십시오
- 재시도 시 아티팩트나 메트릭이 손상되지 않도록 모든 단계를 이덱포텐트하게 만드십시오
불변 값과 순수 변환 함수를 우선적으로 사용하십시오. 특징 생성 과정에서 공유 데이터프레임이나 전역 구성을 변경하지 마십시오.
import hashlib
from dataclasses import dataclass
from pathlib import Path
@dataclass(frozen=True)
class TrainingConfig:
dataset_uri: str
model_dir: Path
seed: int
learning_rate: float
batch_size: int
def artifact_name(config: TrainingConfig, code_sha: str) -> str:
config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
return f"{code_sha[:12]}-{config_hash}"
4. 배포 전 평가
훈련이 완료되기 전에 배포 기준을 명시해야 합니다:
- 기준 모델과 현재 운영 중인 모델의 비교
- 제품 동작에 부합하는 주요 메트릭
- 지연 시간, 보정, 공정성 슬라이스, 비용 및 오류 집중도에 대한 안전 장치 지표
- 중요한 코호트, 지역, 기기, 언어 또는 데이터 소스에 대한 슬라이스 지표
- 지표에 잡음이 있을 경우 신뢰 구간 또는 반복 실행 분산
- 영향이 큰 모델의 경우, 사람이 직접 검토한 실패 사례
- 명시적인 “배포 금지” 임계값
PROMOTION_GATES = {
"auc": ("min", 0.82),
"calibration_error": ("max", 0.04),
"p95_latency_ms": ("max", 80),
}
def assert_promotion_ready(metrics: dict[str, float]) -> None:
missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
if missing:
raise ValueError(f"Model promotion metrics missing required gates: {missing}")
failures = {
name: value
for name, (direction, threshold) in PROMOTION_GATES.items()
for value in [metrics[name]]
if (direction == "min" and value < threshold)
or (direction == "max" and value > threshold)
}
if failures:
raise ValueError(f"Model failed promotion gates: {failures}")
오프라인 지표는 보장 수단이 아닌 진입 기준으로 활용하십시오. 모델이 제품 동작을 변경하는 경우, 전체 배포 전에 섀도우 평가, 카나리아 배포 또는 A/B 테스트를 계획하십시오.
5. 서빙을 위한 패키징
ML 아티팩트는 서빙 계약이 검증 가능할 때만 프로덕션에 배포할 준비가 된 것입니다:
- 모델 아티팩트에는 버전, 훈련 데이터 참조, 구성 및 전처리 정보가 포함됩니다
- 입력 스키마는 유효하지 않거나, 오래된, 또는 범위를 벗어난 특징을 거부해야 합니다
- 출력 스키마에는 모델 버전과, 필요한 경우 신뢰도 또는 설명 필드가 포함되어야 합니다
- 서빙 경로에는 타임아웃, 배치 처리, 리소스 제한 및 대체 동작이 설정되어 있습니다
- CPU/GPU 요구 사항이 명확하게 명시되어 있으며 테스트를 거쳤습니다
- 예측 로그는 개인 식별 정보(PII)를 포함하지 않으며, 디버깅 및 레이블 조인을 위해 충분한 식별자를 포함합니다
- 통합 테스트는 누락된 특징, 오래된 특징, 잘못된 유형, 빈 배치 및 대체 경로를 다룹니다
동등성을 입증하는 테스트 없이, 훈련 전용 특징 코드가 서빙용 특징 코드와 달라지지 않도록 해야 합니다.
6. 모델 운영
모델 모니터링에는 시스템 신호와 품질 신호가 모두 필요합니다:
- 가용성, 오류율, 타임아웃율, 큐 깊이, p50/p95/p99 지연 시간
- 특성 누락률, 범위 드리프트, 범주형 드리프트, 신선도 드리프트
- 예측 분포 드리프트 및 신뢰도 분포 드리프트
- 라벨 도착 상태 및 지연 품질 지표
- 비즈니스 KPI 안전 장치 및 롤백 트리거
- 캐나리 및 롤백을 위한 버전별 대시보드
모든 배포에는 이전 아티팩트, 구성, 데이터 종속성 및 트래픽 전환 메커니즘을 명시한 롤백 계획이 있어야 합니다.
검토 체크리스트
- 예측 계약이 명확하고 테스트 가능해야 합니다
- 데이터 계약은 엔티티 단위, 라벨 시기, 기능 시기, 스냅샷/버전을 정의해야 합니다
- 예측 시점의 가용성을 기준으로 유출 위험이 확인되었어야 합니다
- 코드, 구성, 데이터 버전 및 시드를 기반으로 훈련을 재현할 수 있습니다.
- 지표는 기준 모델 및 현재 운영 모델과 비교됩니다
- 고위험 코호트를 위해 슬라이스 메트릭과 가드레일이 포함됩니다
- 프로모션 게이트는 자동화되어 있으며, 실패 시 차단됩니다
- 훈련 및 서빙 변환은 공유되거나 동등성 검증을 거칩니다
- 모델 아티팩트에는 버전, 구성, 데이터셋 참조 및 전처리 정보가 포함됩니다
- 서빙 경로는 입력을 검증하며, 타임아웃, 폴백 및 롤백 동작을 지원합니다
- 모니터링은 시스템 상태, 특징 드리프트, 예측 드리프트 및 지연된 레이블을 포괄합니다
- 민감한 데이터는 아티팩트, 로그, 프롬프트 및 예시에서 제외됩니다
안티패턴
- 모델을 재현하려면 노트북 상태가 필요합니다
- 무작위 분할로 인해 향후 데이터가 검증 집합이나 테스트 집합으로 유출됩니다
- 특성 조인 시 이벤트 시간 및 레이블 가용성을 무시합니다
- 오프라인 지표는 개선되지만 중요한 슬라이스에서는 성능이 저하됩니다
- 임계값이 테스트 세트에서 반복적으로 조정된다
- 훈련 전처리 과정이 서비스 코드로 수동으로 복사됩니다
- 예측 로그에 모델 버전이 누락되어 있습니다
- 모니터링은 서비스 가동 시간만 확인하고, 데이터나 예측 품질은 확인하지 않습니다
- 롤백 시 정상으로 확인된 아티팩트로 전환하는 대신 재훈련이 필요합니다
기대 결과
이 스킬을 사용할 때는 데이터 계약, 프로모션 게이트, 파이프라인 단계, 테스트 계획, 배포 계획 또는 검토 결과와 같은 구체적인 아티팩트를 반환해야 합니다. 가정을 바탕으로 채워 넣기보다는, 프로덕션 준비 상태를 방해하는 불확실한 요소를 명확히 지적해야 합니다.
---
name: mle-workflow
description: Turn model work into a production ML system with data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.
---
# Machine Learning Engineering Workflow
Use this skill to turn model work into a production ML system with clear data contracts, repeatable training, measurable quality gates, deployable artifacts, and operational monitoring.
## When to Activate
- Planning or reviewing a production ML feature, model refresh, ranking system, recommender, classifier, embedding workflow, or forecasting pipeline
- Converting notebook code into a reusable training, evaluation, batch inference, or online inference pipeline
- Designing model promotion criteria, offline/online evals, experiment tracking, or rollback paths
- Debugging failures caused by data drift, label leakage, stale features, artifact mismatch, or inconsistent training and serving logic
- Adding model monitoring, canary rollout, shadow traffic, or post-deploy quality checks
## Scope Calibration
Use only the lanes that fit the system in front of you. This skill is useful for ranking, search, recommendations, classifiers, forecasting, embeddings, LLM workflows, anomaly detection, and batch analytics, but it should not force one architecture onto all of them.
- Do not assume every model has supervised labels, online serving, a feature store, PyTorch, GPUs, human review, A/B tests, or real-time feedback.
- Do not add heavyweight MLOps machinery when a data contract, baseline, eval script, and rollback note would make the change reviewable.
- Do make assumptions explicit when the project lacks labels, delayed outcomes, slice definitions, production traffic, or monitoring ownership.
- Treat examples as interchangeable scaffolds. Replace metrics, serving mode, data stores, and rollout mechanics with the project-native equivalents.
## Related Skills
- `python-patterns` and `python-testing` for Python implementation and pytest coverage
- `pytorch-patterns` for deep learning models, data loaders, device handling, and training loops
- `eval-harness` and `ai-regression-testing` for promotion gates and agent-assisted regression checks
- `database-migrations`, `postgres-patterns`, and `clickhouse-io` for data storage and analytics surfaces
- `deployment-patterns`, `docker-patterns`, and `security-review` for serving, secrets, containers, and production hardening
## Reuse the SWE Surface
Do not treat MLE as separate from software engineering. Most ECC SWE workflows apply directly to ML systems, often with stricter failure modes:
The recommended `minimal --with capability:machine-learning` install keeps the core agent surface available alongside this skill. For skill-only or agent-limited harnesses, pair `skill:mle-workflow` with `agent:mle-reviewer` where the target supports agents.
| SWE surface | MLE use |
|-------------|---------|
| `product-capability` / `architecture-decision-records` | Turn model work into explicit product contracts and record irreversible data, model, and rollout choices |
| `repo-scan` / `codebase-onboarding` / `code-tour` | Find existing training, feature, serving, eval, and monitoring paths before introducing a parallel ML stack |
| `plan` / `feature-dev` | Scope model changes as product capabilities with data, eval, serving, and rollback phases |
| `tdd-workflow` / `python-testing` | Test feature transforms, split logic, metric calculations, artifact loading, and inference schemas before implementation |
| `code-reviewer` / `mle-reviewer` | Review code quality plus ML-specific leakage, reproducibility, promotion, and monitoring risks |
| `build-fix` / `pr-test-analyzer` | Diagnose broken CI, flaky evals, missing fixtures, and environment-specific model or dependency failures |
| `quality-gate` / `test-coverage` | Require automated evidence for transforms, metrics, inference contracts, promotion gates, and rollback behavior |
| `eval-harness` / `verification-loop` | Turn offline metrics, slice checks, latency budgets, and rollback drills into repeatable gates |
| `ai-regression-testing` | Preserve every production bug as a regression: missing feature, stale label, bad artifact, schema drift, or serving mismatch |
| `api-design` / `backend-patterns` | Design prediction APIs, batch jobs, idempotent retraining endpoints, and response envelopes |
| `database-migrations` / `postgres-patterns` / `clickhouse-io` | Version labels, feature snapshots, prediction logs, experiment metrics, and drift analytics |
| `deployment-patterns` / `docker-patterns` | Package reproducible training and serving images with health checks, resource limits, and rollback |
| `canary-watch` / `dashboard-builder` | Make rollout health visible with model-version, slice, drift, latency, cost, and delayed-label dashboards |
| `security-review` / `security-scan` | Check model artifacts, notebooks, prompts, datasets, and logs for secrets, PII, unsafe deserialization, and supply-chain risk |
| `e2e-testing` / `browser-qa` / `accessibility` | Test critical product flows that consume predictions, including explainability and fallback UI states |
| `benchmark` / `performance-optimizer` | Measure throughput, p95 latency, memory, GPU utilization, and cost per prediction or retrain |
| `cost-aware-llm-pipeline` / `token-budget-advisor` | Route LLM/embedding workloads by quality, latency, and budget instead of defaulting to the largest model |
| `documentation-lookup` / `search-first` | Verify current library behavior for model serving, feature stores, vector DBs, and eval tooling before coding |
| `git-workflow` / `github-ops` / `opensource-pipeline` | Package MLE changes for review with crisp scope, generated artifacts excluded, and reproducible test evidence |
| `strategic-compact` / `dmux-workflows` | Split long ML work into parallel tracks: data contract, eval harness, serving path, monitoring, and docs |
## Ten MLE Task Simulations
Use these simulations as coverage checks when planning or reviewing MLE work. A strong MLE workflow should reduce each task to explicit contracts, reusable SWE surfaces, automated evidence, and a reviewable artifact.
| ID | Common MLE task | Streamlined ECC path | Required output | Pipeline lanes covered |
|----|-----------------|----------------------|-----------------|------------------------|
| MLE-01 | Frame an ambiguous prediction, ranking, recommender, classifier, embedding, or forecast capability | `product-capability`, `plan`, `architecture-decision-records`, `mle-workflow` | Iteration Compact naming who cares, decision owner, success metric, unacceptable mistakes, assumptions, constraints, and first experiment | product contract, stakeholder loss, risk, rollout |
| MLE-02 | Define metric goals, labels, data sources, and the mistake budget | `repo-scan`, `database-reviewer`, `database-migrations`, `postgres-patterns`, `clickhouse-io` | Data and metric contract with entity grain, label timing, label confidence, feature timing, point-in-time joins, split policy, and dataset snapshot | data contract, metric design, leakage, reproducibility |
| MLE-03 | Build a baseline model and scoring path before adding complexity | `tdd-workflow`, `python-testing`, `python-patterns`, `code-reviewer` | Baseline scorer with confusion matrix, calibration notes, latency/cost estimate, known weaknesses, and tests for score shape and determinism | baseline, scoring, testing, serving parity |
| MLE-04 | Generate features from hypotheses about what separates outcomes | `python-patterns`, `pytorch-patterns`, `docker-patterns`, `deployment-patterns` | Feature plan and transform module covering signal source, missing values, outliers, correlations, leakage checks, and train/serve equivalence | feature pipeline, leakage, training, artifacts |
| MLE-05 | Tune thresholds, configs, and model complexity under tradeoffs | `eval-harness`, `ai-regression-testing`, `quality-gate`, `test-coverage` | Threshold/config report comparing precision, recall, F1, AUC, calibration, group slices, latency, cost, complexity, and acceptable error classes | evaluation, threshold, promotion, regression |
| MLE-06 | Run error analysis and turn mistakes into the next experiment | `eval-harness`, `ai-regression-testing`, `mle-reviewer`, `silent-failure-hunter` | Error cluster report for false positives, false negatives, ambiguous labels, stale features, missing signals, and bug traces with lessons captured | error analysis, bug trace, iteration, regression |
| MLE-07 | Package a model artifact for batch or online inference | `api-design`, `backend-patterns`, `security-review`, `security-scan` | Versioned artifact bundle with preprocessing, config, dependency constraints, schema validation, safe loading, and PII-safe logs | artifact, security, inference contract |
| MLE-08 | Ship online serving or batch scoring with feedback capture | `api-design`, `backend-patterns`, `e2e-testing`, `browser-qa`, `accessibility` | Prediction endpoint or batch job with response envelope, timeout, batching, fallback, model version, confidence, feedback logging, and product-flow tests | serving, batch inference, fallback, user workflow |
| MLE-09 | Roll out a model with shadow traffic, canary, A/B test, or rollback | `canary-watch`, `dashboard-builder`, `verification-loop`, `performance-optimizer` | Rollout plan naming traffic split, dashboards, p95 latency, cost, quality guardrails, rollback artifact, and rollback trigger | deployment, canary, rollback |
| MLE-10 | Operate, debug, and refresh a production model after launch | `silent-failure-hunter`, `dashboard-builder`, `mle-reviewer`, `doc-updater`, `github-ops` | Observation ledger and refresh plan with drift checks, delayed-label health, alert owners, runbook updates, retrain criteria, and PR evidence | monitoring, incident response, retraining |
## Iteration Compact
Before touching model code, compress the work into one reviewable artifact. This should be short enough to fit in a PR description and precise enough that another engineer can challenge the tradeoffs.
```text
Goal:
Who cares:
Decision owner:
User or system action changed by the model:
Success metric:
Guardrail metrics:
Mistake budget:
Unacceptable mistakes:
Acceptable mistakes:
Assumptions:
Constraints:
Labels and data snapshot:
Baseline:
Candidate signals:
Threshold or config plan:
Eval slices:
Known risks:
Next experiment:
Rollback or fallback:
```
This compact is the MLE equivalent of a strong SWE design note. It keeps the team from optimizing a metric no one trusts, adding features that do not address the real error mode, or shipping complexity without a rollback.
## Decision Brain
Use this loop whenever the task is ambiguous, high-impact, or metric-heavy:
1. Start from the decision, not the model. Name the action that changes downstream behavior.
2. Name who cares and why. Different stakeholders pay different costs for false positives, false negatives, latency, compute spend, opacity, or missed opportunities.
3. Convert ambiguity into hypotheses. Ask what signal would separate outcomes, what evidence would disprove it, and what simple baseline should be hard to beat.
4. Research prior art or a nearby known problem before inventing a bespoke system.
5. Score choices with `(probability, confidence) x (cost, severity, importance, impact)`.
6. Consider adversarial behavior, incentives, selective disclosure, distribution shift, and feedback loops.
7. Prefer the simplest change that reduces the most important mistake. Simplicity is not laziness; it is a way to minimize blunders while preserving iteration speed.
8. Capture the decision, evidence, counterargument, and next reversible step.
## Metric and Mistake Economics
Choose metrics from failure costs, not habit:
- Use a confusion matrix early so the team can discuss concrete false positives and false negatives instead of abstract accuracy.
- Favor precision when the cost of an incorrect positive decision dominates.
- Favor recall when the cost of a missed positive dominates.
- Use F1 only when the precision/recall tradeoff is genuinely balanced and explainable.
- Use AUC or ranking metrics when ordering quality matters more than a single threshold.
- Track latency, throughput, memory, and cost as first-class metrics because they shape feasible model complexity.
- Compare against a baseline and the current production model before celebrating an offline gain.
- Treat real-world feedback signals as delayed labels with bias, lag, and coverage gaps; do not treat them as ground truth without analysis.
Every metric choice should state which mistake it makes cheaper, which mistake it makes more likely, and who absorbs that cost.
## Data and Feature Hypotheses
Features should come from a theory of separation:
- Text, categorical fields, numeric histories, graph relationships, recency, frequency, and aggregates are candidate signal families, not automatic features.
- For every feature family, state why it should separate outcomes and how it could leak future information.
- For noisy labels, consider adjudication, label confidence, soft targets, or confidence weighting.
- For class imbalance, compare weighted loss, resampling, threshold movement, and calibrated decision rules.
- For missing values, decide whether absence is informative, imputable, or a reason to abstain.
- For outliers, decide whether to clip, bucket, investigate, or preserve them as rare but important signal.
- For correlated features, check whether they are redundant, unstable, or proxies for unavailable future state.
Do not add model complexity until error analysis shows that the baseline is failing for a reason additional signal or capacity can plausibly fix.
## Error Analysis Loop
After each baseline, training run, threshold change, or config change:
1. Split mistakes into false positives, false negatives, abstentions, low-confidence cases, and system failures.
2. Cluster errors by shared traits: language, entity type, source, time, geography, device, sparsity, recency, feature freshness, label source, or model version.
3. Separate model mistakes from data bugs, label ambiguity, product ambiguity, instrumentation gaps, and serving mismatches.
4. Trace each major cluster to one of four moves: better labels, better features, better threshold/config, or better product fallback.
5. Preserve every important mistake as a regression test, eval slice, dashboard panel, or runbook entry.
6. Write the next iteration as a falsifiable experiment, not a vague "improve model" task.
The strongest MLE loop is not train -> metric -> ship. It is mistake -> cluster -> hypothesis -> experiment -> evidence -> simpler system.
## Observation Ledger
Keep a compact decision and evidence trail beside the code, PR, experiment report, or runbook:
```text
Iteration:
Change:
Why this mattered:
Metric movement:
Slice movement:
False positives:
False negatives:
Unexpected errors:
Decision:
Tradeoff accepted:
Lesson captured:
Regression added:
Debt created:
Next iteration:
```
Use the ledger to make model work cumulative. The goal is for each iteration to make the next decision easier, not merely to produce another artifact.
## Core Workflow
### 1. Define the Prediction Contract
Capture the product-level contract before writing model code:
- Prediction target and decision owner
- Input entity, output schema, confidence/calibration fields, and allowed latency
- Batch, online, streaming, or hybrid serving mode
- Fallback behavior when the model, feature store, or dependency is unavailable
- Human review or override path for high-impact decisions
- Privacy, retention, and audit requirements for inputs, predictions, and labels
Do not accept "improve the model" as a requirement. Tie the model to an observable product behavior and a measurable acceptance gate.
### 2. Lock the Data Contract
Every ML task needs an explicit data contract:
- Entity grain and primary key
- Label definition, label timestamp, and label availability delay
- Feature timestamp, freshness SLA, and point-in-time join rules
- Train, validation, test, and backtest split policy
- Required columns, allowed nulls, ranges, categories, and units
- PII or sensitive fields that must not enter training artifacts or logs
- Dataset version or snapshot ID for reproducibility
Guard against leakage first. If a feature is not available at prediction time, or is joined using future information, remove it or move it to an analysis-only path.
### 3. Build a Reproducible Pipeline
Training code should be runnable by another engineer without hidden notebook state:
- Use typed config files or dataclasses for all hyperparameters and paths
- Pin package and model dependencies
- Set random seeds and document any nondeterministic GPU behavior
- Record dataset version, code SHA, config hash, metrics, and artifact URI
- Save preprocessing logic with the model artifact, not separately in a notebook
- Keep train, eval, and inference transformations shared or generated from one source
- Make every step idempotent so retries do not corrupt artifacts or metrics
Prefer immutable values and pure transformation functions. Avoid mutating shared data frames or global config during feature generation.
```python
import hashlib
from dataclasses import dataclass
from pathlib import Path
@dataclass(frozen=True)
class TrainingConfig:
dataset_uri: str
model_dir: Path
seed: int
learning_rate: float
batch_size: int
def artifact_name(config: TrainingConfig, code_sha: str) -> str:
config_key = f"{config.dataset_uri}:{config.seed}:{config.learning_rate}:{config.batch_size}"
config_hash = hashlib.sha256(config_key.encode("utf-8")).hexdigest()[:12]
return f"{code_sha[:12]}-{config_hash}"
```
### 4. Evaluate Before Promotion
Promotion criteria should be declared before training finishes:
- Baseline model and current production model comparison
- Primary metric aligned to product behavior
- Guardrail metrics for latency, calibration, fairness slices, cost, and error concentration
- Slice metrics for important cohorts, geographies, devices, languages, or data sources
- Confidence intervals or repeated-run variance when metrics are noisy
- Failure examples reviewed by a human for high-impact models
- Explicit "do not ship" thresholds
```python
PROMOTION_GATES = {
"auc": ("min", 0.82),
"calibration_error": ("max", 0.04),
"p95_latency_ms": ("max", 80),
}
def assert_promotion_ready(metrics: dict[str, float]) -> None:
missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
if missing:
raise ValueError(f"Model promotion metrics missing required gates: {missing}")
failures = {
name: value
for name, (direction, threshold) in PROMOTION_GATES.items()
for value in [metrics[name]]
if (direction == "min" and value < threshold)
or (direction == "max" and value > threshold)
}
if failures:
raise ValueError(f"Model failed promotion gates: {failures}")
```
Use offline metrics as gates, not guarantees. When the model changes product behavior, plan shadow evaluation, canary rollout, or A/B testing before full rollout.
### 5. Package for Serving
An ML artifact is production-ready only when the serving contract is testable:
- Model artifact includes version, training data reference, config, and preprocessing
- Input schema rejects invalid, stale, or out-of-range features
- Output schema includes model version and confidence or explanation fields when useful
- Serving path has timeout, batching, resource limits, and fallback behavior
- CPU/GPU requirements are explicit and tested
- Prediction logs avoid PII and include enough identifiers for debugging and label joins
- Integration tests cover missing features, stale features, bad types, empty batches, and fallback path
Never let training-only feature code diverge from serving feature code without a test that proves equivalence.
### 6. Operate the Model
Model monitoring needs both system and quality signals:
- Availability, error rate, timeout rate, queue depth, and p50/p95/p99 latency
- Feature null rate, range drift, categorical drift, and freshness drift
- Prediction distribution drift and confidence distribution drift
- Label arrival health and delayed quality metrics
- Business KPI guardrails and rollback triggers
- Per-version dashboards for canaries and rollbacks
Every deployment should have a rollback plan that names the previous artifact, config, data dependency, and traffic-switch mechanism.
## Review Checklist
- [ ] Prediction contract is explicit and testable
- [ ] Data contract defines entity grain, label timing, feature timing, and snapshot/version
- [ ] Leakage risks were checked against prediction-time availability
- [ ] Training is reproducible from code, config, data version, and seed
- [ ] Metrics compare against baseline and current production model
- [ ] Slice metrics and guardrails are included for high-risk cohorts
- [ ] Promotion gates are automated and fail closed
- [ ] Training and serving transformations are shared or equivalence-tested
- [ ] Model artifact carries version, config, dataset reference, and preprocessing
- [ ] Serving path validates inputs and has timeout, fallback, and rollback behavior
- [ ] Monitoring covers system health, feature drift, prediction drift, and delayed labels
- [ ] Sensitive data is excluded from artifacts, logs, prompts, and examples
## Anti-Patterns
- Notebook state is required to reproduce the model
- Random split leaks future data into validation or test sets
- Feature joins ignore event time and label availability
- Offline metric improves while important slices regress
- Thresholds are tuned on the test set repeatedly
- Training preprocessing is copied manually into serving code
- Model version is missing from prediction logs
- Monitoring only checks service uptime, not data or prediction quality
- Rollback requires retraining instead of switching to a known-good artifact
## Output Expectations
When using this skill, return concrete artifacts: data contract, promotion gates, pipeline steps, test plan, deployment plan, or review findings. Call out unknowns that block production readiness instead of filling them with assumptions.
모든 파일
1개 파일mle-workflow 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/affaan-m/ECC/tree/main/skills/mle-workflow # Copy SKILL.md to your .claude/skills/ directory
복사





집
