CI / build (push) Has been cancelled
Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)
Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.
Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.
2.6 KiB
2.6 KiB
SPEC — mde-interactive-enforcement
Problem
status.py:1432-1436 enforces that the session playing reviewer differs from the
implementer (via .state.implementer), but there is no enforcement that the model
playing bug-finder differs from adversarial-bug-finder, or that the referee model
differs from the implementer model. Same-model conflict-of-interest yields
rubber-stamping.
Design
Full design in design/framework/functional.md §4 and design/framework/technical.md §§4-5, §12.
Key decisions
- Conflict matrix (locked):
code_review≠implement;bug_find≠implement;adversarial_bug_find≠implement+bug_find;referee≠implement+bug_find+adversarial_bug_find;loop-verify≠loop-implement. - Auto-assignment: default model → next-available on conflict → user override
via
--transition --modelorloop.json roles.<role>.model. Refuse only if no non-conflicting model exists. - State:
.state.modelsper task recording which model filled which role. - Audit: new
model_divergencecategory in--audit. - Dashboard: model badges on task cards.
Scope
scripts/status.py:CONFLICT_MATRIXconstant (locked matrix from SPEC)_check_conflict(current_phase, current_model, next_phase, next_model)helpercmd_transition:--modelarg + conflict check in multi-LLM modecmd_claim:--modelarg + conflict checkcmd_audit:model_divergencecategory.state.modelswriter (JSON:{implementer, code_reviewer, bug_hunter, ...})
automaton/dashboard/html/dashboard.js— model badge on task cardsautomaton/dashboard/ui/app.py—.state.modelsin/api/tasksresponsetests/test_model_divergence.py— conflict matrix, auto-assign, audit, badges
Dependencies
- Depends on
mde-manifest-detection(needs_load_models_manifest()and_get_mode())
Out of scope
- Loop model binding and
{model}substitution (subtaskmde-loop-enforcement) - Rule agents (FW-2, FW-3)
Success criteria
- Single-LLM mode:
--modelrecords but never refuses (advisory once ifadvised: true) - Multi-LLM mode:
--transition --modelrefuses on conflict (e.g., same model for implement + code_review) - Multi-LLM mode:
--claim --modelrefuses on conflict --auditflagsmodel_divergenceviolations- Dashboard shows model badges when
.state.modelsexists pytest tests/ -qgreen
References
design/framework/functional.md§4 (Model-Divergence Enforcement)design/framework/technical.md§§4-5 (status.py flags, conflict matrix), §12 (agent tab data flow)design/framework/README.md(locked decisions F4, F5)