Bootstrap self-improvement loop, decompose model-divergence-enforcement
CI / build (push) Has been cancelled

Self-improvement loop:
- Created via --create-loop --from-template self-improvement
- Scheduled via launchd (3600s interval)
- State: running

model-divergence-enforcement task:
- Research approved, decomposed into 3 sequential subtasks:
  1. mde-manifest-detection (models.json + detect_models.py)
  2. mde-interactive-enforcement (conflict matrix + --model args + audit)
  3. mde-loop-enforcement (loop.json roles + {model} substitution + check-gate)
- Parent at decomposition:approved (stays active until subtasks complete)
- Each subtask has BRIEF.md with scope, deliverables, acceptance criteria
This commit is contained in:
Lap Tran
2026-06-25 07:15:25 -04:00
parent 715f6f9495
commit 7336db282d
17 changed files with 350 additions and 0 deletions
@@ -0,0 +1,34 @@
# BRIEF — mde-interactive-enforcement
**Parent:** model-divergence-enforcement
**Subtask:** 2 of 3 (depends on mde-manifest-detection)
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 2
**Design ref:** `design/framework/technical.md` §§4-5, §12
## Objective
Add the conflict matrix, `--transition --model` / `--claim --model` enforcement,
`.state.models` tracking, audit category, and dashboard model badges. This is the
core enforcement layer for interactive (non-loop) workflows.
## Deliverables
1. `scripts/status.py`:
- `CONFLICT_MATRIX` constant (locked matrix from SPEC)
- `_check_conflict(current_phase, current_model, next_phase, next_model)` helper
- `cmd_transition`: `--model` arg + conflict check in multi-LLM mode
- `cmd_claim`: `--model` arg + conflict check
- `cmd_audit`: `model_divergence` category
- `.state.models` writer (JSON: `{implementer, code_reviewer, bug_hunter, ...}`)
2. `automaton/dashboard/html/dashboard.js` — model badge on task cards
3. `automaton/dashboard/ui/app.py` — `.state.models` in `/api/tasks` response
4. `tests/test_model_divergence.py` — conflict matrix, auto-assign, audit, badges
## Acceptance
- Single-LLM mode: `--model` records but never refuses (advisory once if `advised: true`)
- Multi-LLM mode: `--transition --model` refuses on conflict (e.g., same model for implement + code_review)
- Multi-LLM mode: `--claim --model` refuses on conflict
- `--audit` flags `model_divergence` violations
- Dashboard shows model badges when `.state.models` exists
- `pytest tests/ -q` green
@@ -0,0 +1,31 @@
# BRIEF — mde-loop-enforcement
**Parent:** model-divergence-enforcement
**Subtask:** 3 of 3 (depends on mde-interactive-enforcement)
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 3
**Design ref:** `design/framework/technical.md` §§10, §13
## Objective
Add loop-level model-divergence enforcement: `loop.json` per-role `model` field,
`{model}` substitution in `_invoke_harness`, and `--check-gate` model-divergence
brake gate.
## Deliverables
1. `scripts/loop-runner.py`:
- `_invoke_harness` (line 366-404): `{model}` placeholder substitution from `loop.json` role config
2. `scripts/status.py`:
- `--check-gate`: model-divergence brake gate (loop-verify ≠ loop-implement in multi-LLM mode)
- `cmd_create_loop` / `cmd_install_schedule`: validate `loop.json` per-role `model` fields
3. `templates/loops/`: update loop templates with `roles` schema example
4. `design/loops/technical.md`: document `{model}` substitution
5. `tests/test_model_divergence.py`: loop model binding, `{model}` substitution, check-gate halt
## Acceptance
- `loop.json` with `roles.implementer.model` → `_invoke_harness` substitutes `{model}` in harness command
- `loop.json` without per-role `model` → defaults to `models.json` `default` model
- `--check-gate` in multi-LLM mode halts if loop-verify model = loop-implement model
- `--check-gate` in single-LLM mode does NOT halt (advisory only)
- `pytest tests/ -q` green
@@ -0,0 +1,30 @@
# BRIEF — mde-manifest-detection
**Parent:** model-divergence-enforcement
**Subtask:** 1 of 3 (foundational, no deps)
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 1
**Design ref:** `design/framework/technical.md` §§2-3
## Objective
Create the model manifest system (`models.json` schema + loader) and model detection
script (`scripts/detect_models.py`). This is the foundation for all model-divergence
enforcement — subtasks 2 and 3 depend on it.
## Deliverables
1. `scripts/detect_models.py` — probes opencode.json providers + localhost endpoints
(8080/11434/1234/8000), emits candidate manifest as JSON to stdout
2. `scripts/status.py` — add `_load_models_manifest()` and `_get_mode()` helpers
3. `scripts/install.sh`, `update.sh`, `upgrade.sh` — call detect_models after vram_detect
4. `config.md` — `## Available Models` section template
5. `prompts/onboarding.md` — Step 2d model config check
6. `tests/test_model_divergence.py` — manifest loading, single vs multi mode, missing file
## Acceptance
- `models.json` missing → `_get_mode()` returns `"single"`, all model commands are no-ops
- `models.json` with 0-1 models → `"single"` mode
- `models.json` with 2+ models → `"multi"` mode
- `detect_models.py` probes localhost and prints candidate JSON
- `pytest tests/ -q` green