Files
Lap Tran d325963644
CI / build (push) Has been cancelled
Flatten model-divergence subtasks into 3 independent tasks
Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)

Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.

Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.
2026-06-25 07:25:10 -04:00

2.6 KiB

SPEC — mde-interactive-enforcement

Problem

status.py:1432-1436 enforces that the session playing reviewer differs from the implementer (via .state.implementer), but there is no enforcement that the model playing bug-finder differs from adversarial-bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping.

Design

Full design in design/framework/functional.md §4 and design/framework/technical.md §§4-5, §12.

Key decisions

  • Conflict matrix (locked): code_review≠implement; bug_find≠implement; adversarial_bug_find≠implement+bug_find; referee≠implement+bug_find+adversarial_bug_find; loop-verify≠loop-implement.
  • Auto-assignment: default model → next-available on conflict → user override via --transition --model or loop.json roles.<role>.model. Refuse only if no non-conflicting model exists.
  • State: .state.models per task recording which model filled which role.
  • Audit: new model_divergence category in --audit.
  • Dashboard: model badges on task cards.

Scope

  1. scripts/status.py:
    • CONFLICT_MATRIX constant (locked matrix from SPEC)
    • _check_conflict(current_phase, current_model, next_phase, next_model) helper
    • cmd_transition: --model arg + conflict check in multi-LLM mode
    • cmd_claim: --model arg + conflict check
    • cmd_audit: model_divergence category
    • .state.models writer (JSON: {implementer, code_reviewer, bug_hunter, ...})
  2. automaton/dashboard/html/dashboard.js — model badge on task cards
  3. automaton/dashboard/ui/app.py — .state.models in /api/tasks response
  4. tests/test_model_divergence.py — conflict matrix, auto-assign, audit, badges

Dependencies

  • Depends on mde-manifest-detection (needs _load_models_manifest() and _get_mode())

Out of scope

  • Loop model binding and {model} substitution (subtask mde-loop-enforcement)
  • Rule agents (FW-2, FW-3)

Success criteria

  1. Single-LLM mode: --model records but never refuses (advisory once if advised: true)
  2. Multi-LLM mode: --transition --model refuses on conflict (e.g., same model for implement + code_review)
  3. Multi-LLM mode: --claim --model refuses on conflict
  4. --audit flags model_divergence violations
  5. Dashboard shows model badges when .state.models exists
  6. pytest tests/ -q green

References

  • design/framework/functional.md §4 (Model-Divergence Enforcement)
  • design/framework/technical.md §§4-5 (status.py flags, conflict matrix), §12 (agent tab data flow)
  • design/framework/README.md (locked decisions F4, F5)