Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)
Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.
Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.
This commit is contained in:
@@ -0,0 +1 @@
|
||||
research:awaiting_approval
|
||||
@@ -0,0 +1,62 @@
|
||||
# SPEC — mde-interactive-enforcement
|
||||
|
||||
## Problem
|
||||
|
||||
`status.py:1432-1436` enforces that the *session* playing reviewer differs from the
|
||||
implementer (via `.state.implementer`), but there is no enforcement that the *model*
|
||||
playing bug-finder differs from adversarial-bug-finder, or that the referee model
|
||||
differs from the implementer model. Same-model conflict-of-interest yields
|
||||
rubber-stamping.
|
||||
|
||||
## Design
|
||||
|
||||
Full design in `design/framework/functional.md` §4 and `design/framework/technical.md` §§4-5, §12.
|
||||
|
||||
### Key decisions
|
||||
|
||||
- **Conflict matrix (locked)**: `code_review≠implement`; `bug_find≠implement`;
|
||||
`adversarial_bug_find≠implement+bug_find`; `referee≠implement+bug_find+adversarial_bug_find`;
|
||||
`loop-verify≠loop-implement`.
|
||||
- **Auto-assignment**: default model → next-available on conflict → user override
|
||||
via `--transition --model` or `loop.json roles.<role>.model`. Refuse only if no
|
||||
non-conflicting model exists.
|
||||
- **State**: `.state.models` per task recording which model filled which role.
|
||||
- **Audit**: new `model_divergence` category in `--audit`.
|
||||
- **Dashboard**: model badges on task cards.
|
||||
|
||||
## Scope
|
||||
|
||||
1. `scripts/status.py`:
|
||||
- `CONFLICT_MATRIX` constant (locked matrix from SPEC)
|
||||
- `_check_conflict(current_phase, current_model, next_phase, next_model)` helper
|
||||
- `cmd_transition`: `--model` arg + conflict check in multi-LLM mode
|
||||
- `cmd_claim`: `--model` arg + conflict check
|
||||
- `cmd_audit`: `model_divergence` category
|
||||
- `.state.models` writer (JSON: `{implementer, code_reviewer, bug_hunter, ...}`)
|
||||
2. `automaton/dashboard/html/dashboard.js` — model badge on task cards
|
||||
3. `automaton/dashboard/ui/app.py` — `.state.models` in `/api/tasks` response
|
||||
4. `tests/test_model_divergence.py` — conflict matrix, auto-assign, audit, badges
|
||||
|
||||
## Dependencies
|
||||
|
||||
- Depends on `mde-manifest-detection` (needs `_load_models_manifest()` and `_get_mode()`)
|
||||
|
||||
## Out of scope
|
||||
|
||||
- Loop model binding and `{model}` substitution (subtask `mde-loop-enforcement`)
|
||||
- Rule agents (FW-2, FW-3)
|
||||
|
||||
## Success criteria
|
||||
|
||||
1. Single-LLM mode: `--model` records but never refuses (advisory once if `advised: true`)
|
||||
2. Multi-LLM mode: `--transition --model` refuses on conflict (e.g., same model for implement + code_review)
|
||||
3. Multi-LLM mode: `--claim --model` refuses on conflict
|
||||
4. `--audit` flags `model_divergence` violations
|
||||
5. Dashboard shows model badges when `.state.models` exists
|
||||
6. `pytest tests/ -q` green
|
||||
|
||||
## References
|
||||
|
||||
- `design/framework/functional.md` §4 (Model-Divergence Enforcement)
|
||||
- `design/framework/technical.md` §§4-5 (status.py flags, conflict matrix), §12 (agent tab data flow)
|
||||
- `design/framework/README.md` (locked decisions F4, F5)
|
||||
Reference in New Issue
Block a user