Flatten model-divergence subtasks into 3 independent tasks
CI / build (push) Has been cancelled

Replaced parent+subtask structure with 3 standalone tasks:
- mde-manifest-detection (foundational: models.json schema, detect_models.py)
- mde-interactive-enforcement (conflict matrix, --model args, audit, badges)
- mde-loop-enforcement (loop.json per-role model, {model} substitution, check-gate)

Parent archived to tasks/complete/ for history. Each task has its own
SPEC.md and is at research:awaiting_approval.

Rationale: subtasks are independently actionable with their own lifecycle.
Parent container added complexity without benefit.
This commit is contained in:
Lap Tran
2026-06-25 07:25:10 -04:00
parent 7336db282d
commit d325963644
20 changed files with 170 additions and 100 deletions
@@ -0,0 +1 @@
complete
@@ -8,7 +8,7 @@ sequential subtasks with strict dependency ordering.
## Sub-tasks (3, sequential)
### subtask-1: `mde-manifest-detection`
### subtask-1: `mde-manifest-detection` (→ independent task)
**Scope:** `models.json` schema + loader + `scripts/detect_models.py` probe + install integration.
**Files touched:**
- `scripts/detect_models.py` (new — probe opencode.json providers + localhost endpoints 8080/11434/1234/8000)
+1
View File
@@ -0,0 +1 @@
research:awaiting_approval
+62
View File
@@ -0,0 +1,62 @@
# SPEC — mde-interactive-enforcement
## Problem
`status.py:1432-1436` enforces that the *session* playing reviewer differs from the
implementer (via `.state.implementer`), but there is no enforcement that the *model*
playing bug-finder differs from adversarial-bug-finder, or that the referee model
differs from the implementer model. Same-model conflict-of-interest yields
rubber-stamping.
## Design
Full design in `design/framework/functional.md` §4 and `design/framework/technical.md` §§4-5, §12.
### Key decisions
- **Conflict matrix (locked)**: `code_review≠implement`; `bug_find≠implement`;
`adversarial_bug_find≠implement+bug_find`; `referee≠implement+bug_find+adversarial_bug_find`;
`loop-verify≠loop-implement`.
- **Auto-assignment**: default model → next-available on conflict → user override
via `--transition --model` or `loop.json roles.<role>.model`. Refuse only if no
non-conflicting model exists.
- **State**: `.state.models` per task recording which model filled which role.
- **Audit**: new `model_divergence` category in `--audit`.
- **Dashboard**: model badges on task cards.
## Scope
1. `scripts/status.py`:
- `CONFLICT_MATRIX` constant (locked matrix from SPEC)
- `_check_conflict(current_phase, current_model, next_phase, next_model)` helper
- `cmd_transition`: `--model` arg + conflict check in multi-LLM mode
- `cmd_claim`: `--model` arg + conflict check
- `cmd_audit`: `model_divergence` category
- `.state.models` writer (JSON: `{implementer, code_reviewer, bug_hunter, ...}`)
2. `automaton/dashboard/html/dashboard.js` — model badge on task cards
3. `automaton/dashboard/ui/app.py` — `.state.models` in `/api/tasks` response
4. `tests/test_model_divergence.py` — conflict matrix, auto-assign, audit, badges
## Dependencies
- Depends on `mde-manifest-detection` (needs `_load_models_manifest()` and `_get_mode()`)
## Out of scope
- Loop model binding and `{model}` substitution (subtask `mde-loop-enforcement`)
- Rule agents (FW-2, FW-3)
## Success criteria
1. Single-LLM mode: `--model` records but never refuses (advisory once if `advised: true`)
2. Multi-LLM mode: `--transition --model` refuses on conflict (e.g., same model for implement + code_review)
3. Multi-LLM mode: `--claim --model` refuses on conflict
4. `--audit` flags `model_divergence` violations
5. Dashboard shows model badges when `.state.models` exists
6. `pytest tests/ -q` green
## References
- `design/framework/functional.md` §4 (Model-Divergence Enforcement)
- `design/framework/technical.md` §§4-5 (status.py flags, conflict matrix), §12 (agent tab data flow)
- `design/framework/README.md` (locked decisions F4, F5)
+1
View File
@@ -0,0 +1 @@
research:awaiting_approval
+56
View File
@@ -0,0 +1,56 @@
# SPEC — mde-loop-enforcement
## Problem
Loops have no model-divergence enforcement. A loop could use the same model for
both implementation and verification, defeating the purpose of the loop verifier
gate. Loop runners need per-role model binding and `{model}` substitution in
harness commands.
## Design
Full design in `design/framework/functional.md` §4 and `design/framework/technical.md` §§10, §13.
### Key decisions
- **Per-role model**: `loop.json` roles section with optional `model` field per role.
Defaults to `models.json` `default` model if not specified.
- **Substitution**: `{model}` placeholder in `loop.json` `harness.command` gets
replaced with the assigned model for that role.
- **Check-gate**: New model-divergence brake gate: loop-verify model must differ from
loop-implement model in multi-LLM mode. Single-LLM mode: advisory only.
## Scope
1. `scripts/loop-runner.py`:
- `_invoke_harness` (line 366-404): `{model}` placeholder substitution from `loop.json` role config
2. `scripts/status.py`:
- `--check-gate`: model-divergence brake gate (loop-verify ≠ loop-implement in multi-LLM mode)
- `cmd_create_loop` / `cmd_install_schedule`: validate `loop.json` per-role `model` fields
3. `templates/loops/`: update loop templates with `roles` schema example
4. `design/loops/technical.md`: document `{model}` substitution
5. `tests/test_model_divergence.py`: loop model binding, `{model}` substitution, check-gate halt
## Dependencies
- Depends on `mde-interactive-enforcement` (needs `CONFLICT_MATRIX` and `_check_conflict`)
## Out of scope
- Interactive `--transition --model` / `--claim --model` (already done in `mde-interactive-enforcement`)
- Rule agents (FW-2, FW-3)
- Agent tab redesign (FW-1)
## Success criteria
1. `loop.json` with `roles.implementer.model` → `_invoke_harness` substitutes `{model}` in harness command
2. `loop.json` without per-role `model` → defaults to `models.json` `default` model
3. `--check-gate` in multi-LLM mode halts if loop-verify model = loop-implement model
4. `--check-gate` in single-LLM mode does NOT halt (advisory only)
5. `pytest tests/ -q` green
## References
- `design/framework/functional.md` §4 (Model-Divergence Enforcement)
- `design/framework/technical.md` §§10 (harness invocation), §13 (loop integration)
- `design/loops/technical.md` (loop runner architecture)
+1
View File
@@ -0,0 +1 @@
research:awaiting_approval
+47
View File
@@ -0,0 +1,47 @@
# SPEC — mde-manifest-detection
## Problem
The framework has no model manifest system. Model-divergence enforcement (conflict
matrix, per-role model binding, loop enforcement) requires knowing which models are
available and whether the system is in single-LLM or multi-LLM mode.
## Design
Full design in `design/framework/functional.md` §4 and `design/framework/technical.md` §§2-3.
### Key decisions
- **Manifest**: `models.json` with `{default, advised, models:[{name, provider, context_window, location}]}`.
- **Mode detection**: 0-1 models → single-LLM (advisory once, then silent). 2+ → multi-LLM (hard block). Missing file → single-LLM (backward compatible).
- **Detection**: `scripts/detect_models.py` probes opencode.json providers + localhost endpoints (8080/11434/1234/8000).
## Scope
1. `scripts/detect_models.py` — probe opencode.json providers + localhost endpoints, emit candidate manifest as JSON to stdout
2. `scripts/status.py` — add `_load_models_manifest()` and `_get_mode()` helpers
3. `scripts/install.sh`, `update.sh`, `upgrade.sh` — call detect_models after vram_detect
4. `config.md` — `## Available Models` section template
5. `prompts/onboarding.md` — Step 2d model config check
6. `tests/test_model_divergence.py` — manifest loading, single vs multi mode, missing file
## Out of scope
- Conflict matrix enforcement (subtask `mde-interactive-enforcement`)
- Loop model binding (subtask `mde-loop-enforcement`)
- Rule agents (FW-2, FW-3) — separate backlog items
- Agent tab redesign (FW-1)
## Success criteria
1. `models.json` missing → `_get_mode()` returns `"single"`, all model commands are no-ops
2. `models.json` with 0-1 models → `"single"` mode
3. `models.json` with 2+ models → `"multi"` mode
4. `detect_models.py` probes localhost and prints candidate JSON
5. `pytest tests/ -q` green
## References
- `design/framework/functional.md` §4 (Model-Divergence Enforcement)
- `design/framework/technical.md` §§2-3 (manifest schema, state schemas)
- `design/framework/README.md` (locked decisions F4, F5)
@@ -1 +0,0 @@
decomposition:approved
@@ -1,34 +0,0 @@
# BRIEF — mde-interactive-enforcement
**Parent:** model-divergence-enforcement
**Subtask:** 2 of 3 (depends on mde-manifest-detection)
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 2
**Design ref:** `design/framework/technical.md` §§4-5, §12
## Objective
Add the conflict matrix, `--transition --model` / `--claim --model` enforcement,
`.state.models` tracking, audit category, and dashboard model badges. This is the
core enforcement layer for interactive (non-loop) workflows.
## Deliverables
1. `scripts/status.py`:
- `CONFLICT_MATRIX` constant (locked matrix from SPEC)
- `_check_conflict(current_phase, current_model, next_phase, next_model)` helper
- `cmd_transition`: `--model` arg + conflict check in multi-LLM mode
- `cmd_claim`: `--model` arg + conflict check
- `cmd_audit`: `model_divergence` category
- `.state.models` writer (JSON: `{implementer, code_reviewer, bug_hunter, ...}`)
2. `automaton/dashboard/html/dashboard.js` — model badge on task cards
3. `automaton/dashboard/ui/app.py` — `.state.models` in `/api/tasks` response
4. `tests/test_model_divergence.py` — conflict matrix, auto-assign, audit, badges
## Acceptance
- Single-LLM mode: `--model` records but never refuses (advisory once if `advised: true`)
- Multi-LLM mode: `--transition --model` refuses on conflict (e.g., same model for implement + code_review)
- Multi-LLM mode: `--claim --model` refuses on conflict
- `--audit` flags `model_divergence` violations
- Dashboard shows model badges when `.state.models` exists
- `pytest tests/ -q` green
@@ -1,31 +0,0 @@
# BRIEF — mde-loop-enforcement
**Parent:** model-divergence-enforcement
**Subtask:** 3 of 3 (depends on mde-interactive-enforcement)
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 3
**Design ref:** `design/framework/technical.md` §§10, §13
## Objective
Add loop-level model-divergence enforcement: `loop.json` per-role `model` field,
`{model}` substitution in `_invoke_harness`, and `--check-gate` model-divergence
brake gate.
## Deliverables
1. `scripts/loop-runner.py`:
- `_invoke_harness` (line 366-404): `{model}` placeholder substitution from `loop.json` role config
2. `scripts/status.py`:
- `--check-gate`: model-divergence brake gate (loop-verify ≠ loop-implement in multi-LLM mode)
- `cmd_create_loop` / `cmd_install_schedule`: validate `loop.json` per-role `model` fields
3. `templates/loops/`: update loop templates with `roles` schema example
4. `design/loops/technical.md`: document `{model}` substitution
5. `tests/test_model_divergence.py`: loop model binding, `{model}` substitution, check-gate halt
## Acceptance
- `loop.json` with `roles.implementer.model` → `_invoke_harness` substitutes `{model}` in harness command
- `loop.json` without per-role `model` → defaults to `models.json` `default` model
- `--check-gate` in multi-LLM mode halts if loop-verify model = loop-implement model
- `--check-gate` in single-LLM mode does NOT halt (advisory only)
- `pytest tests/ -q` green
@@ -1,30 +0,0 @@
# BRIEF — mde-manifest-detection
**Parent:** model-divergence-enforcement
**Subtask:** 1 of 3 (foundational, no deps)
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 1
**Design ref:** `design/framework/technical.md` §§2-3
## Objective
Create the model manifest system (`models.json` schema + loader) and model detection
script (`scripts/detect_models.py`). This is the foundation for all model-divergence
enforcement — subtasks 2 and 3 depend on it.
## Deliverables
1. `scripts/detect_models.py` — probes opencode.json providers + localhost endpoints
(8080/11434/1234/8000), emits candidate manifest as JSON to stdout
2. `scripts/status.py` — add `_load_models_manifest()` and `_get_mode()` helpers
3. `scripts/install.sh`, `update.sh`, `upgrade.sh` — call detect_models after vram_detect
4. `config.md` — `## Available Models` section template
5. `prompts/onboarding.md` — Step 2d model config check
6. `tests/test_model_divergence.py` — manifest loading, single vs multi mode, missing file
## Acceptance
- `models.json` missing → `_get_mode()` returns `"single"`, all model commands are no-ops
- `models.json` with 0-1 models → `"single"` mode
- `models.json` with 2+ models → `"multi"` mode
- `detect_models.py` probes localhost and prints candidate JSON
- `pytest tests/ -q` green