skill-comply
affaan-m/ECC
通过生成不同提示严格程度的场景、运行智能体、对工具调用进行分类,并报告包含完整时间线的合规率,自动检测编码智能体是否遵循技能、规则或智能体定义。
...展开全部skill-comply: 自动化合规性评估
通过以下方式衡量编码代理是否实际遵循了技能、规则或代理定义:
- 从任意 .md 文件中自动生成预期行为序列(规格说明)
- 自动生成提示严格程度逐渐降低的场景(支持型 → 中立型 → 竞争型)
- 运行
claude -p并通过 stream-json 捕获工具调用轨迹 - 使用大型语言模型(而非正则表达式)将工具调用与规格步骤进行分类
- 确定性地检查时间顺序
- 生成包含规格说明、提示词和时间线的自包含报告
支持的目标
- 技能(
skills/*/SKILL.md):工作流技能,如 search-first、TDD 指南 - 规则(
rules/common/*.md):强制性规则,如 testing.md、security.md、git-workflow.md - 代理定义(
agents/*.md):代理是否在预期时间被调用(目前尚不支持内部工作流验证)
何时触发
- 用户运行
/skill-comply - 用户询问“这条规则是否真的被遵守了?”
- 在添加新规则/技能后,用于验证代理的合规性
- 作为质量维护的一部分定期执行
用法
# 完整运行
uv run python -m scripts.run ~/.claude/rules/common/testing.md
# 模拟运行(不产生费用,仅包含规范和场景)
uv run python -m scripts.run --dry-run ~/.claude/skills/search-first/SKILL.md
# 自定义模型
uv run python -m scripts.run --gen-model haiku --model sonnet
核心概念:提示语独立性
衡量即使提示词未明确支持,技能/规则是否仍被遵循。
报告内容
报告内容自成体系,包括:
- 预期行为序列(自动生成的规范)
- 场景提示(各严格度级别下提出的请求)
- 各场景的合规评分
- 带大语言模型(LLM)分类标签的工具调用时间线
高级功能(可选)
对于熟悉钩子的用户,报告还会针对合规性较低的步骤提供钩子提升建议。这仅供参考——其主要价值在于合规性可见性本身。
---
name: skill-comply
description: Automatically measures whether coding agents follow skills, rules, or agent definitions by generating scenarios at multiple prompt strictness levels, running agents, classifying tool calls, and reporting compliance rates with full timelines.
---
# skill-comply: Automated Compliance Measurement
Measures whether coding agents actually follow skills, rules, or agent definitions by:
1. Auto-generating expected behavioral sequences (specs) from any .md file
2. Auto-generating scenarios with decreasing prompt strictness (supportive → neutral → competing)
3. Running `claude -p` and capturing tool call traces via stream-json
4. Classifying tool calls against spec steps using LLM (not regex)
5. Checking temporal ordering deterministically
6. Generating self-contained reports with spec, prompts, and timelines
## Supported Targets
- **Skills** (`skills/*/SKILL.md`): Workflow skills like search-first, TDD guides
- **Rules** (`rules/common/*.md`): Mandatory rules like testing.md, security.md, git-workflow.md
- **Agent definitions** (`agents/*.md`): Whether an agent gets invoked when expected (internal workflow verification not yet supported)
## When to Activate
- User runs `/skill-comply <path>`
- User asks "is this rule actually being followed?"
- After adding new rules/skills, to verify agent compliance
- Periodically as part of quality maintenance
## Usage
```bash
# Full run
uv run python -m scripts.run ~/.claude/rules/common/testing.md
# Dry run (no cost, spec + scenarios only)
uv run python -m scripts.run --dry-run ~/.claude/skills/search-first/SKILL.md
# Custom models
uv run python -m scripts.run --gen-model haiku --model sonnet <path>
```
## Key Concept: Prompt Independence
Measures whether a skill/rule is followed even when the prompt doesn't explicitly support it.
## Report Contents
Reports are self-contained and include:
1. Expected behavioral sequence (auto-generated spec)
2. Scenario prompts (what was asked at each strictness level)
3. Compliance scores per scenario
4. Tool call timelines with LLM classification labels
### Advanced (optional)
For users familiar with hooks, reports also include hook promotion recommendations for steps with low compliance. This is informational — the main value is the compliance visibility itself.
所有文件
21 个文件安装 skill-comply
下载技能文件并将其解压到 .claude/skills/ 目录中。
下载ZIP克隆仓库并复制技能文件到您的项目中。
git clone https://github.com/affaan-m/ECC/tree/main/skills/skill-comply # Copy SKILL.md to your .claude/skills/ directory
复制





首页
