Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7336db282d | ||
|
|
715f6f9495 | ||
|
|
4b3d92c9a2 |
@@ -111,3 +111,12 @@ When the agent receives a trigger command, it must:
|
|||||||
1. Read the corresponding template file.
|
1. Read the corresponding template file.
|
||||||
2. Replace all `{placeholders}` with the actual project values.
|
2. Replace all `{placeholders}` with the actual project values.
|
||||||
3. Execute the rendered prompt.
|
3. Execute the rendered prompt.
|
||||||
|
|
||||||
|
## Backlog
|
||||||
|
|
||||||
|
Design backlogs are the framework's outstanding-work store when no active tasks exist. Each design area has its own `BACKLOG.md`:
|
||||||
|
|
||||||
|
- `~/.automaton/design/loops/BACKLOG.md` — loop engineering v1.1+ and deferred items.
|
||||||
|
- `~/.automaton/design/framework/BACKLOG.md` — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement).
|
||||||
|
|
||||||
|
The self-improvement loop consumes these automatically when configured with `work_source.kind = "backlog"` and `work_source.area` set to the relevant design area (`"loops"` or `"framework"`). For manual work, read the topmost `- [ ]` item and create a task via `status.py --create-task`.
|
||||||
|
|||||||
@@ -47,6 +47,9 @@ Automaton is a **contract-based, state-enforced workflow framework** for LLM age
|
|||||||
│ ├── config.py
|
│ ├── config.py
|
||||||
│ ├── core/
|
│ ├── core/
|
||||||
│ └── ui/
|
│ └── ui/
|
||||||
|
├── design/ # Design docs for major features
|
||||||
|
│ ├── loops/ # Loop engineering system (v1 locked)
|
||||||
|
│ └── framework/ # Framework agent features (rule agents, Agent tab, model-divergence)
|
||||||
├── tests/ # pytest suite
|
├── tests/ # pytest suite
|
||||||
└── tasks/ # Framework development tasks
|
└── tasks/ # Framework development tasks
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -2,6 +2,15 @@
|
|||||||
|
|
||||||
## [unreleased]
|
## [unreleased]
|
||||||
|
|
||||||
|
### Added — framework agent features design docs
|
||||||
|
|
||||||
|
- **New `design/framework/`** directory: design index, functional design, technical design, and backlog for three framework-level agent features:
|
||||||
|
- **Model-Divergence Enforcement** — `models.json` manifest, single vs multi-LLM mode detection, conflict matrix (`code_review≠implement`, `bug_find≠implement`, `adversarial_bug_find≠implement+bug_find`, `referee≠implement+bug_find+adversarial_bug_find`, `loop-verify≠loop-implement`), `--transition --model`, `.state.models`, `loop.json` per-role model binding + `{model}` substitution, `--audit` model_divergence category, dashboard badges. Shipped as a manual task (3 subtasks), not a backlog item.
|
||||||
|
- **Rule Agents** — Rule Proposer (daily scheduled standalone agent, scans completed tasks' `BUG_REPORT.md`/`ADVERSARIAL_BUG_REPORT.md`/`VERDICT.md`, proposes rules to `RULE_PROPOSALS.md`) and Rule Reviewer (monthly scheduled standalone agent, consolidates `.rules.md` → `RULE_REVIEW.md`). Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers. Conflict-of-interest: Reviewer must differ from Proposer; both must differ from tasks they review. Enforcement deferred until model-divergence ships.
|
||||||
|
- **Agent Tab Redesign** — replace 4 fake `AGENT_TYPE_META` types with two sections: Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) + Scheduled Jobs (real `job.kind`: cleanup, loop, rule-scan, rule-review). New `/api/phase-roles` endpoint.
|
||||||
|
- **`design/framework/BACKLOG.md`**: 3 v1 items (`agent-tab-real-roles`, `rule-proposer-agent`, `rule-reviewer-agent`) + 3 deferred items. Consumable by the self-improvement loop via `work_source.area = "framework"`.
|
||||||
|
- **Cross-references**: `design/loops/README.md` and `design/loops/BACKLOG.md` now point to the sibling `design/framework/` area. `AGENTS.md` repo layout updated.
|
||||||
|
|
||||||
### Added — cross-loop task claim (task `add-claim-loop-task`)
|
### Added — cross-loop task claim (task `add-claim-loop-task`)
|
||||||
|
|
||||||
- **`scripts/status.py`**: New `--claim-loop-task <name> --task <taskname> [--project P]` command. Exit 0 = claimed (or already self-claimed, idempotent). Exit 2 = already claimed by another running/paused loop (`task_already_claimed:{other}` on stderr) or untracked loop (`loop_untracked`). Uses `_loop_lock` to serialize writes; cross-loop scan is advisory (self-healing on next tick).
|
- **`scripts/status.py`**: New `--claim-loop-task <name> --task <taskname> [--project P]` command. Exit 0 = claimed (or already self-claimed, idempotent). Exit 2 = already claimed by another running/paused loop (`task_already_claimed:{other}` on stderr) or untracked loop (`loop_untracked`). Uses `_loop_lock` to serialize writes; cross-loop scan is advisory (self-healing on next tick).
|
||||||
|
|||||||
@@ -250,6 +250,7 @@ The runner resolves prompt files from `loop.json` `roles.*.prompt` (e.g. `loop-i
|
|||||||
| `blast_radius.file_scope` | List of paths the loop may edit |
|
| `blast_radius.file_scope` | List of paths the loop may edit |
|
||||||
| `blast_radius.use_worktree` | If true, tick runs in a per-loop git worktree |
|
| `blast_radius.use_worktree` | If true, tick runs in a per-loop git worktree |
|
||||||
| `work_source.kind` | `single`, `audit`, or `backlog` |
|
| `work_source.kind` | `single`, `audit`, or `backlog` |
|
||||||
|
| `work_source.area` | Design area for `backlog` kind (default `"loops"`; `"framework"` reads `design/framework/BACKLOG.md`) |
|
||||||
| `roles.implement.prompt` | Prompt file for Implement role |
|
| `roles.implement.prompt` | Prompt file for Implement role |
|
||||||
| `roles.verify.prompt` | Prompt file for Verify role |
|
| `roles.verify.prompt` | Prompt file for Verify role |
|
||||||
| `roles.orchestrate.prompt` | Prompt file for Orchestrate role |
|
| `roles.orchestrate.prompt` | Prompt file for Orchestrate role |
|
||||||
|
|||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# Framework Backlog
|
||||||
|
|
||||||
|
Status: **v1 draft 2026-06-25**. These items are captured for the self-improvement loop (or manual pickup) once the self-improvement loop is running and the design docs are in place.
|
||||||
|
|
||||||
|
This backlog mirrors the pattern of `design/loops/BACKLOG.md` but for framework-level agent features. The self-improvement loop's `work_source.area` can be set to `"framework"` to pull from here.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## v1 (Priority)
|
||||||
|
|
||||||
|
| ID | Item | Description | Dependencies |
|
||||||
|
|---|---|---|---|
|
||||||
|
| FW-1 | **agent-tab-real-roles** | Redesign the dashboard Agent tab: replace 4 fake `AGENT_TYPE_META` types with two sections — (1) Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) showing active/inactive based on current task phases; (2) Scheduled Jobs (cleanup, loop-tick, rule-scan, rule-review) using real `job.kind` from `/api/scheduled`. New `/api/phase-roles` endpoint. | None |
|
||||||
|
| FW-2 | **rule-proposer-agent** | Daily scheduled standalone agent that scans completed tasks since last run, reads their `BUG_REPORT.md` / `ADVERSARIAL_BUG_REPORT.md` / `VERDICT.md`, proposes new rules with concrete examples to `RULE_PROPOSALS.md`. Uses direct harness invocation (reuses `loop-runner._invoke_harness`). State in `.state.rule-scan`. Requires different LLM than the tasks it reviews (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding) |
|
||||||
|
| FW-3 | **rule-reviewer-agent** | Monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, missing examples — writes `RULE_REVIEW.md`. Uses direct harness invocation. State in `.state.rule-review`. Must use different LLM than Rule Proposer (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding), `rule-proposer-agent` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Deferred / Tier 2 (Post-v1)
|
||||||
|
|
||||||
|
| ID | Item | Description | Notes |
|
||||||
|
|---|---|---|---|
|
||||||
|
| FW-4 | **model-divergence-enforcement** | Manifest (`models.json`), mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model + `{model}` substitution, audit category, dashboard badges. This is a **manual task**, not a backlog item — tracked separately. | See separate task `model-divergence-enforcement` |
|
||||||
|
| FW-5 | **dashboard-agent-tab-v2** | Enhance Agent tab with: role details (click → task list), schedule management (enable/disable from UI), last-run timestamps, per-agent logs. | Requires FW-1 |
|
||||||
|
| FW-6 | **rule-enforcement-gate** | Pre-edit hook (`--can-edit --rule <name>`) that enforces rules from `.rules.md` before edits. Separate from rule agents (which only propose/review). | Requires FW-2, FW-3 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- **Dependency on `model-divergence-enforcement`**: FW-2 and FW-3 are designed with model-binding in their configs (per-role `model` in their schedule config), but hard-block enforcement is deferred until the model-divergence feature ships. The design docs note the dependency; implementation tasks in the backlog carry a "depends on" annotation.
|
||||||
|
- **Self-improvement loop**: Once the self-improvement loop is running with `work_source.area = "framework"`, it will pick up items from this backlog automatically. The loop template `templates/loops/self-improvement/loop.json` already supports `work_source.area`.
|
||||||
|
- **FW-4 is NOT in this backlog** — it's a manual parent task created via `status.py --create-task model-divergence-enforcement` and decomposed into subtasks.
|
||||||
@@ -0,0 +1,67 @@
|
|||||||
|
# Framework Agent Features — Design Index
|
||||||
|
|
||||||
|
Status: **v1 draft 2026-06-25**. Design for three framework-level agent features: rule agents (proposer + reviewer), Agent tab redesign, and model-divergence enforcement.
|
||||||
|
|
||||||
|
## What this is
|
||||||
|
|
||||||
|
The Automaton framework's agent-feature layer: scheduled standalone agents for rule maintenance (Rule Proposer, Rule Reviewer), a redesigned dashboard Agent tab showing real phase roles instead of fake types, and a model-divergence enforcement system that prevents conflict-of-interest by ensuring different LLM roles use different models.
|
||||||
|
|
||||||
|
Built on top of the existing `status.py` phase machine and loop infrastructure — no second enforcement surface.
|
||||||
|
|
||||||
|
## Why
|
||||||
|
|
||||||
|
- `.rules.md` is maintained manually; failure patterns from completed tasks are not systematically captured.
|
||||||
|
- The dashboard Agent tab shows 4 made-up agent types (`completed_task_archiver`, `single`, `audit`, `backlog`) — not the real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator).
|
||||||
|
- Conflict-of-interest: two sessions with the same LLM playing bug finder and adversarial bug finder (or implementer and reviewer) don't yield useful results. The framework needs computational enforcement of model divergence.
|
||||||
|
|
||||||
|
## Documents
|
||||||
|
|
||||||
|
- [`functional.md`](functional.md) — what v1 does, roles, triggers, outputs, schedules, conflict-of-interest, success criteria. **Read this first.**
|
||||||
|
- [`technical.md`](technical.md) — the implementation contract: file map, state schemas, `status.py` flags, runner flows, dashboard API changes, test coverage. **Read this if you're implementing.**
|
||||||
|
- [`BACKLOG.md`](BACKLOG.md) — v1 work queue (3 items) and deferred items. The self-improvement loop's work queue when `work_source.area = "framework"`.
|
||||||
|
|
||||||
|
## v1 scope (draft)
|
||||||
|
|
||||||
|
Three feature areas, each with a clear boundary:
|
||||||
|
|
||||||
|
1. **Model-Divergence Enforcement** — Foundational layer. `models.json` manifest, mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, audit category, dashboard badges. Enables the other two features.
|
||||||
|
|
||||||
|
2. **Rule Agents** — Two standalone scheduled agents (NOT task lifecycle phases):
|
||||||
|
- **Rule Proposer**: Daily scan of completed tasks → `RULE_PROPOSALS.md` with concrete examples.
|
||||||
|
- **Rule Reviewer**: Monthly consolidation of `.rules.md` → `RULE_REVIEW.md`.
|
||||||
|
- Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers (mirror `--install-cleanup-schedule`).
|
||||||
|
|
||||||
|
3. **Agent Tab Redesign** — Dashboard `/agent` view:
|
||||||
|
- **Phase Roles section**: 6 roles from `.agent.md`, active/inactive based on current task phases.
|
||||||
|
- **Scheduled Jobs section**: Real job kinds from `/api/scheduled` (cleanup, loop, rule-scan, rule-review), replacing 4 fake `AGENT_TYPE_META` types.
|
||||||
|
- New `/api/phase-roles` endpoint.
|
||||||
|
|
||||||
|
## Locked decisions (referenced as `(Fn)` in functional/technical)
|
||||||
|
|
||||||
|
| ID | Decision |
|
||||||
|
|---|---|
|
||||||
|
| F1 | Rule agents are **standalone scheduled agents**, NOT task lifecycle phases. They don't block task completion. |
|
||||||
|
| F2 | Rule Proposer runs **daily**; Rule Reviewer runs **monthly**. OS-native schedulers (launchd/cron/schtasks). |
|
||||||
|
| F3 | Rule agents use **direct harness invocation** (reuse `loop-runner._invoke_harness`), not loop infrastructure. |
|
||||||
|
| F4 | Conflict-of-interest: Rule Reviewer **must use different LLM** than Rule Proposer. Rule agents must use different LLM than tasks they review. Enforcement **deferred** until model-divergence ships (F5). |
|
||||||
|
| F5 | Model-divergence enforcement is a **foundational layer** shipped first (manual task, not backlog). Rule agents designed with model-binding config but hard-block enforced later. |
|
||||||
|
| F6 | Agent tab has **two sections**: Phase Roles (from `.agent.md` + task state) + Scheduled Jobs (from `/api/scheduled`). |
|
||||||
|
| F7 | Phase Roles section needs **new `/api/phase-roles` endpoint** mapping active tasks to their phase roles. |
|
||||||
|
| F8 | Scheduled Jobs section uses **real `job.kind`** (cleanup, loop, rule-scan, rule-review) — remove `AGENT_TYPE_META` fake types. |
|
||||||
|
| F9 | State files: `.state.rule-scan` (last scanned task, timestamp, proposed rules), `.state.rule-review` (last run timestamp). |
|
||||||
|
| F10 | New `status.py` flags: `--rule-scan`, `--install-rule-scan-schedule`, `--rule-review`, `--install-rule-review-schedule` (mirror `--cleanup-done` / `--install-cleanup-schedule`). |
|
||||||
|
|
||||||
|
## Relationship to loops
|
||||||
|
|
||||||
|
- `design/loops/` is the loop engineering system (state-enforced unattended work).
|
||||||
|
- `design/framework/` is framework-level agent features (rule maintenance, observability, model governance).
|
||||||
|
- The self-improvement loop (`templates/loops/self-improvement/`) can drive `design/framework/` work by setting `work_source.area = "framework"` — no code change needed (see `loop-runner.py:_find_work_backlog`).
|
||||||
|
- Model-divergence enforcement (F5) is a prerequisite for hard-blocking conflict-of-interest in rule agents and loops.
|
||||||
|
|
||||||
|
## Post-v1
|
||||||
|
|
||||||
|
After v1 lands and the self-improvement loop is running:
|
||||||
|
|
||||||
|
- Rule enforcement gate (`--can-edit --rule <name>`) — separate feature.
|
||||||
|
- Dashboard Agent tab v2: role details, schedule management, per-agent logs.
|
||||||
|
- Rule agents gain comprehension-debt tracking (`last-read-sha` per rule file).
|
||||||
@@ -0,0 +1,335 @@
|
|||||||
|
# Framework Agent Features — Functional Design
|
||||||
|
|
||||||
|
Status: v1 (draft 2026-06-25). Supersedes any prior informal discussions of rule agents or Agent tab redesign.
|
||||||
|
|
||||||
|
Audience: framework maintainers (currently: the human and one AI assistant). After handoff the self-improvement loop is also an audience — designs must be legible to a fresh-context LLM verifier.
|
||||||
|
|
||||||
|
## 1. Problem
|
||||||
|
|
||||||
|
Three gaps in the framework's agent layer:
|
||||||
|
|
||||||
|
1. **`.rules.md` is maintained manually.** Failure patterns from completed tasks (`BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md`) are not systematically captured into rules. The Self-Improvement section of `.rules.md` (lines 33-36) says "Add one rule per observed failure mode with a concrete example" and "Consolidate contradictions monthly" — but no agent does this. It relies on the human or a session agent remembering.
|
||||||
|
|
||||||
|
2. **The dashboard Agent tab shows fake agent types.** `AGENT_TYPE_META` in `dashboard.js:330-334` defines 4 types (`completed_task_archiver`, `single`, `audit`, `backlog`) that are loop work-source kinds, not real agent roles. The 6 real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) are not surfaced.
|
||||||
|
|
||||||
|
3. **No conflict-of-interest enforcement on model divergence.** `status.py:1432-1436` enforces that the reviewer session differs from the implementer session (via `.state.implementer`), but there is no enforcement that the *model* playing bug-finder differs from the model playing adversarial bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping.
|
||||||
|
|
||||||
|
## 2. Goals
|
||||||
|
|
||||||
|
v1 — **the framework maintains its own rules, surfaces real agent roles, and computationally enforces model divergence**:
|
||||||
|
|
||||||
|
1. **Rule Proposer**: a daily scheduled standalone agent that scans completed tasks since its last run, reads their failure artifacts, and proposes new rules with concrete examples to `RULE_PROPOSALS.md`.
|
||||||
|
2. **Rule Reviewer**: a monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, rules missing examples — and writes `RULE_REVIEW.md`.
|
||||||
|
3. **Agent Tab Redesign**: replace 4 fake `AGENT_TYPE_META` types with two sections — Phase Roles (6 roles from `.agent.md`, active/inactive based on current task phases) and Scheduled Jobs (real `job.kind` from `/api/scheduled`).
|
||||||
|
4. **Model-Divergence Enforcement**: a `models.json` manifest, mode detection (single vs multi-LLM), a conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, an audit category, and dashboard badges.
|
||||||
|
5. **Conflict-of-interest for rule agents**: Rule Reviewer must use a different LLM than Rule Proposer. Both must use different LLMs than the tasks they review. Enforcement is *designed now* but *hard-blocked only after model-divergence ships* (F4, F5).
|
||||||
|
|
||||||
|
## 3. Non-Goals (v1)
|
||||||
|
|
||||||
|
- **Rule enforcement during task execution.** Rule agents only *propose* and *review* rules; they do not block edits. A future `--can-edit --rule <name>` gate (FW-6) is separate.
|
||||||
|
- **Auto-approving rules.** A human always reviews `RULE_PROPOSALS.md` and `RULE_REVIEW.md` before rules are merged into `.rules.md`. No auto-merge path in v1.
|
||||||
|
- **Rule agents as task lifecycle phases.** Rule agents are *standalone scheduled agents* (F1). They do not block task completion. They are not phases in the state machine.
|
||||||
|
- **A new agent harness.** Rule agents use direct harness invocation (reusing `loop-runner._invoke_harness`). No new runtime.
|
||||||
|
- **Network-fetched dependencies.** New code is Python stdlib only. No new pip installs.
|
||||||
|
- **Model capability inspection.** The framework never inspects model capability, provider, or size (D8). It only tracks *which* model fills *which* role and enforces the conflict matrix.
|
||||||
|
|
||||||
|
## 4. Model-Divergence Enforcement (foundational layer)
|
||||||
|
|
||||||
|
Shipped first as a manual task (`model-divergence-enforcement`), not a backlog item. Enables conflict-of-interest hard-blocking for rule agents and loops.
|
||||||
|
|
||||||
|
### 4.1 Manifest: `models.json`
|
||||||
|
|
||||||
|
A new file at `~/.automaton/models.json` (or `{project}/.automaton/models.json`):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"default": "glm-4.6",
|
||||||
|
"advised": true,
|
||||||
|
"models": [
|
||||||
|
{"name": "glm-4.6", "provider": "opencode", "context_window": 131072, "location": "remote"},
|
||||||
|
{"name": "qwen3-coder", "provider": "opencode", "context_window": 131072, "location": "remote"},
|
||||||
|
{"name": "llama-3.3-70b", "provider": "localhost", "context_window": 32768, "location": "http://localhost:8080"}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `default`: the model used when no role-specific binding is set.
|
||||||
|
- `advised`: if `true`, the framework prints a one-time advisory in single-LLM mode recommending a second model for conflict-of-interest roles, then goes silent.
|
||||||
|
- `models[]`: the roster. `location` is `"remote"` or a localhost URL for probing.
|
||||||
|
|
||||||
|
### 4.2 Mode Detection
|
||||||
|
|
||||||
|
- **0-1 models** in `models.json` (or file missing) → **single-LLM mode**. Advisory once (if `advised: true`), then silent. No hard blocks.
|
||||||
|
- **2+ models** → **multi-LLM mode**. Hard-block on conflict-matrix violations. Auto-assign next-available non-conflicting model on conflict; refuse only if no non-conflicting model exists.
|
||||||
|
- **Missing file** → single-LLM mode (backward compatible). Existing behavior preserved.
|
||||||
|
|
||||||
|
### 4.3 Conflict Matrix (locked)
|
||||||
|
|
||||||
|
| Role | Must differ from |
|
||||||
|
|---|---|
|
||||||
|
| `code_review` | `implement` |
|
||||||
|
| `bug_find` | `implement` |
|
||||||
|
| `adversarial_bug_find` | `implement`, `bug_find` |
|
||||||
|
| `referee` | `implement`, `bug_find`, `adversarial_bug_find` |
|
||||||
|
| `loop-verify` | `loop-implement` |
|
||||||
|
|
||||||
|
`doc_review`, `code_review`, and `bug_find` are independent of each other (not conflicts). Only `bug_find` ↔ `adversarial_bug_find` conflicts (they are adversary pairs).
|
||||||
|
|
||||||
|
### 4.4 Auto-Assignment (multi-LLM mode)
|
||||||
|
|
||||||
|
1. Default model → assigned to `implement` (and `loop-implement`).
|
||||||
|
2. On conflict, pick the next-available model from `models[]` that does not conflict.
|
||||||
|
3. User override: `loop.json` `roles.<role>.model` or `status.py --transition --model <name>`.
|
||||||
|
4. Refuse only if no non-conflicting model exists.
|
||||||
|
|
||||||
|
### 4.5 State: `.state.models`
|
||||||
|
|
||||||
|
Each task gets `{task}/.state.models` recording which model filled which role:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"implement": "glm-4.6", "code_review": "qwen3-coder", "bug_find": "qwen3-coder", "adversarial_bug_find": "llama-3.3-70b", "referee": "llama-3.3-70b"}
|
||||||
|
```
|
||||||
|
|
||||||
|
`status.py --transition --model <name>` records the model for the role being transitioned into. `--claim` in multi-LLM mode checks the conflict matrix against `.state.models` and refuses on violation.
|
||||||
|
|
||||||
|
### 4.6 Loop Integration
|
||||||
|
|
||||||
|
`loop.json` gains per-role `model` and `harness.command` with `{model}` substitution:
|
||||||
|
|
||||||
|
```json
|
||||||
|
"roles": {
|
||||||
|
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
|
||||||
|
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
|
||||||
|
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`loop-runner.py:_invoke_harness` substitutes `{model}` into the harness command. `--check-gate` enforces `loop-verify` ≠ `loop-implement` model in multi-LLM mode.
|
||||||
|
|
||||||
|
### 4.7 Audit + Dashboard
|
||||||
|
|
||||||
|
- `status.py --audit` gains a `model_divergence` category: flags tasks where `.state.models` violates the conflict matrix.
|
||||||
|
- Dashboard task cards show model badges (one per role filled).
|
||||||
|
|
||||||
|
## 5. Rule Proposer
|
||||||
|
|
||||||
|
A **standalone scheduled agent** (F1) that proposes new rules from completed-task failure patterns.
|
||||||
|
|
||||||
|
### 5.1 Trigger
|
||||||
|
|
||||||
|
Daily, via OS-native scheduler (mirrors `--install-cleanup-schedule`). `status.py --install-rule-scan-schedule [--interval 86400]` installs the schedule unit. Manual: `status.py --rule-scan`.
|
||||||
|
|
||||||
|
### 5.2 Inputs
|
||||||
|
|
||||||
|
- `.state.rule-scan`: state file tracking the last scanned task and timestamp.
|
||||||
|
- Completed tasks (in `tasks/complete/` or tasks with `.state` phase `complete`) that were completed since the last scan.
|
||||||
|
- For each such task: `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md` (if present).
|
||||||
|
- Current `.rules.md` (to deduplicate against existing rules).
|
||||||
|
|
||||||
|
### 5.3 Output
|
||||||
|
|
||||||
|
`RULE_PROPOSALS.md` (at `~/.automaton/RULE_PROPOSALS.md` in framework mode, or `{project}/.automaton/RULE_PROPOSALS.md` in project mode). Format:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# Rule Proposals — {date}
|
||||||
|
|
||||||
|
## Proposed Rule: {title}
|
||||||
|
**Source**: tasks/{task-name}/VERDICT.md
|
||||||
|
**Pattern**: {one-line description of the failure mode}
|
||||||
|
**Example**:
|
||||||
|
{concrete code/config snippet from the task}
|
||||||
|
**Proposed rule text**:
|
||||||
|
{the rule as it would appear in .rules.md}
|
||||||
|
|
||||||
|
---
|
||||||
|
```
|
||||||
|
|
||||||
|
Proposals are *appended* per run. A human reviews and merges accepted rules into `.rules.md`. No auto-merge (Non-Goal).
|
||||||
|
|
||||||
|
### 5.4 LLM Session
|
||||||
|
|
||||||
|
Direct harness invocation (F3): `status.py --rule-scan` reuses `loop-runner._invoke_harness` to spawn one LLM session with a prompt that includes the failure artifacts and current rules. The session proposes rules in the `RULE_PROPOSALS.md` format.
|
||||||
|
|
||||||
|
### 5.5 Conflict-of-Interest
|
||||||
|
|
||||||
|
The Rule Proposer's LLM must differ from the implementer + bug-hunter + adversarial-bug-hunter models of the tasks it scans. This prevents the model that made the bug from proposing the rule about its own bug.
|
||||||
|
|
||||||
|
**Enforcement deferred** (F4): until model-divergence ships, the Rule Proposer runs with the default model. The design includes a `model` field in its schedule config; hard-blocking activates once `models.json` exists and multi-LLM mode is detected.
|
||||||
|
|
||||||
|
### 5.6 State: `.state.rule-scan`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"last_scan_at": "2026-06-25T10:00:00Z", "last_scanned_task": "fix-context-sizing", "proposals_count": 3}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 6. Rule Reviewer
|
||||||
|
|
||||||
|
A **standalone scheduled agent** (F1) that consolidates `.rules.md` periodically.
|
||||||
|
|
||||||
|
### 6.1 Trigger
|
||||||
|
|
||||||
|
Monthly, via OS-native scheduler. `status.py --install-rule-review-schedule [--interval 2592000]` installs the schedule unit. Manual: `status.py --rule-review`.
|
||||||
|
|
||||||
|
### 6.2 Inputs
|
||||||
|
|
||||||
|
- `.state.rule-review`: state file tracking the last run timestamp.
|
||||||
|
- Current `.rules.md` (full file).
|
||||||
|
- Recent `RULE_PROPOSALS.md` entries (since last review).
|
||||||
|
- Recent completed-task summaries (last 30 days) for context on stale rules.
|
||||||
|
|
||||||
|
### 6.3 Output
|
||||||
|
|
||||||
|
`RULE_REVIEW.md` (at `~/.automaton/RULE_REVIEW.md` or project equivalent). Format:
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
# Rule Review — {date}
|
||||||
|
|
||||||
|
## Contradictions Found
|
||||||
|
- Rule A ("...") contradicts Rule B ("..."). Suggested resolution: {merge/drop/keep A}.
|
||||||
|
|
||||||
|
## Stale Rules (no observed instance in last 30 days)
|
||||||
|
- Rule C ("..."). Suggested action: drop or annotate as low-priority.
|
||||||
|
|
||||||
|
## Rules Missing Examples
|
||||||
|
- Rule D ("..."). Suggested example: {from a recent task}.
|
||||||
|
|
||||||
|
## Merge Candidates
|
||||||
|
- Rules E and F overlap. Suggested merged text: {...}.
|
||||||
|
|
||||||
|
---
|
||||||
|
```
|
||||||
|
|
||||||
|
A human reviews and applies accepted changes to `.rules.md`. No auto-apply (Non-Goal).
|
||||||
|
|
||||||
|
### 6.4 LLM Session
|
||||||
|
|
||||||
|
Direct harness invocation (F3), same as Rule Proposer. One LLM session with a prompt that includes the full `.rules.md` and recent proposals.
|
||||||
|
|
||||||
|
### 6.5 Conflict-of-Interest
|
||||||
|
|
||||||
|
The Rule Reviewer's LLM **must differ from the Rule Proposer's LLM** (F4). The Proposer proposes (bias toward adding); the Reviewer consolidates (bias toward pruning). Same model = self-review = rubber-stamping.
|
||||||
|
|
||||||
|
**Enforcement deferred** until model-divergence ships.
|
||||||
|
|
||||||
|
### 6.6 State: `.state.rule-review`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"last_review_at": "2026-06-25T10:00:00Z", "contradictions_found": 2, "stale_rules": 5, "merges_suggested": 1}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 7. Agent Tab Redesign
|
||||||
|
|
||||||
|
### 7.1 Current State (to be replaced)
|
||||||
|
|
||||||
|
`dashboard.js:330-334` defines `AGENT_TYPE_META` with 4 fake types:
|
||||||
|
- `completed_task_archiver` — actually the cleanup scheduled job.
|
||||||
|
- `single` — actually a loop with `work_source.kind = "single"`.
|
||||||
|
- `audit` — actually a loop with `work_source.kind = "audit"`.
|
||||||
|
- `backlog` — actually a loop with `work_source.kind = "backlog"`.
|
||||||
|
|
||||||
|
These are loop work-source kinds, not agent roles. They conflate two different concepts.
|
||||||
|
|
||||||
|
### 7.2 New Design: Two Sections
|
||||||
|
|
||||||
|
**Section 1 — Phase Roles**
|
||||||
|
|
||||||
|
Shows the 6 phase roles from `.agent.md` Agent Configuration:
|
||||||
|
|
||||||
|
| Role ID | Phases | Active when |
|
||||||
|
|---|---|---|
|
||||||
|
| `researcher` | research, decomposition, design, test_design | A task is in one of these phases |
|
||||||
|
| `implementer` | implement | A task is in `implement` phase |
|
||||||
|
| `code-reviewer` | code_review | A task is in `code_review` phase |
|
||||||
|
| `bug-hunter` | bug_find, adversarial_bug_find | A task is in one of these phases |
|
||||||
|
| `referee` | referee | A task is in `referee` phase |
|
||||||
|
| `orchestrator` | new, complete, human_intervention | A task is in one of these phases |
|
||||||
|
|
||||||
|
Each role card shows: role icon, role label, status (Active/Idle — based on whether any task is in that role's phases), and the count of tasks in that role's phases. Clicking a role filters the task list to tasks in that role's phases.
|
||||||
|
|
||||||
|
**Data source**: new `/api/phase-roles` endpoint. Returns:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"roles": [
|
||||||
|
{"id": "researcher", "label": "Researcher", "icon": "🔬", "phases": ["research", "decomposition", "design", "test_design"], "active_tasks": 2, "status": "active"},
|
||||||
|
{"id": "implementer", "label": "Implementer", "icon": "⚙️", "phases": ["implement"], "active_tasks": 1, "status": "active"},
|
||||||
|
...
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Section 2 — Scheduled Jobs**
|
||||||
|
|
||||||
|
Shows real scheduled jobs from `/api/scheduled`, using `job.kind` (not fake agent types):
|
||||||
|
|
||||||
|
| Job Kind | Icon | Label | Source |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `cleanup` | 🧹 | Cleanup Archiver | `com.automaton.cleanup` |
|
||||||
|
| `loop` | 🔄 | Loop: {name} | `com.automaton.loop.{name}` |
|
||||||
|
| `rule-scan` | 📝 | Rule Proposer | `com.automaton.rule-scan` (new) |
|
||||||
|
| `rule-review` | 📋 | Rule Reviewer | `com.automaton.rule-review` (new) |
|
||||||
|
|
||||||
|
Each job card shows: job label, status (Enabled/Disabled/Misconfigured), next-run interval, runtime state (for loops: iteration count, halt status; for rule agents: last-scan/review timestamp). `AGENT_TYPE_META` is removed entirely; rendering uses `job.kind` directly.
|
||||||
|
|
||||||
|
### 7.3 Self-Documenting Names
|
||||||
|
|
||||||
|
Per `.rules.md` "Self-Documenting UI Names" (lines 53-64), all schedule unit names and stub filenames are self-documenting:
|
||||||
|
- `com.automaton.rule-scan` (launchd label)
|
||||||
|
- `automaton-rule-scan.sh` (stub filename)
|
||||||
|
- `com.automaton.rule-review` / `automaton-rule-review.sh`
|
||||||
|
|
||||||
|
## 8. Schedules
|
||||||
|
|
||||||
|
Rule agents use OS-native schedulers, mirroring the `--install-cleanup-schedule` pattern (`status.py:2188-2260`):
|
||||||
|
|
||||||
|
| Agent | Flag | Default Interval | Launchd Label | Stub |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Rule Proposer | `--install-rule-scan-schedule` | 86400s (daily) | `com.automaton.rule-scan` | `automaton-rule-scan.sh` |
|
||||||
|
| Rule Reviewer | `--install-rule-review-schedule` | 2592000s (monthly) | `com.automaton.rule-review` | `automaton-rule-review.sh` |
|
||||||
|
|
||||||
|
Platform dispatch via `platform.system()`:
|
||||||
|
- **Darwin**: `~/Library/LaunchAgents/com.automaton.rule-scan.plist` with `StartInterval`.
|
||||||
|
- **Linux**: crontab line via `_install_cron_block_generic`.
|
||||||
|
- **Windows**: `schtasks /create /tn "AutomatonRuleScan" ...`.
|
||||||
|
|
||||||
|
Stub scripts are 3-line bash/bat files that call `python3 status.py --rule-scan` (or `--rule-review`).
|
||||||
|
|
||||||
|
## 9. Conflict-of-Interest
|
||||||
|
|
||||||
|
### 9.1 Dependency Chain
|
||||||
|
|
||||||
|
```
|
||||||
|
model-divergence-enforcement (shipped first, manual task)
|
||||||
|
↓ enables hard-block
|
||||||
|
rule-proposer-agent (FW-2)
|
||||||
|
↓ conflict-of-interest
|
||||||
|
rule-reviewer-agent (FW-3) — must differ from Proposer
|
||||||
|
```
|
||||||
|
|
||||||
|
### 9.2 Design Now, Enforce Later (F4, F5)
|
||||||
|
|
||||||
|
- Rule agents are *designed* with `model` fields in their schedule config.
|
||||||
|
- The design docs specify the conflict-of-interest rules.
|
||||||
|
- Hard-block enforcement *activates* when `models.json` exists and multi-LLM mode is detected.
|
||||||
|
- Until then, rule agents run with the default model (single-LLM mode, advisory only).
|
||||||
|
|
||||||
|
### 9.3 Why Different Models
|
||||||
|
|
||||||
|
- **Proposer vs Reviewer**: Proposer has a bias toward *adding* rules (more is better). Reviewer has a bias toward *pruning* (less is better). Same model = self-review = the proposer's rules never get pruned.
|
||||||
|
- **Rule agent vs scanned tasks**: The model that introduced a bug should not propose the rule about its own bug — it has a blind spot for that failure mode.
|
||||||
|
|
||||||
|
## 10. Success Criteria for v1
|
||||||
|
|
||||||
|
1. **Model-divergence**: `--transition --model` records the model in `.state.models`; `--claim` refuses conflict-matrix violations in multi-LLM mode; `--audit` flags violations; dashboard shows model badges.
|
||||||
|
2. **Rule Proposer**: `status.py --rule-scan` reads completed tasks since last scan, proposes rules to `RULE_PROPOSALS.md`, updates `.state.rule-scan`. Test with a seeded completed task containing a `VERDICT.md`.
|
||||||
|
3. **Rule Reviewer**: `status.py --rule-review` reads `.rules.md` + recent proposals, writes `RULE_REVIEW.md`, updates `.state.rule-review`. Test with a seeded `.rules.md` containing a contradiction.
|
||||||
|
4. **Agent Tab**: `/api/phase-roles` returns 6 roles with active-task counts; dashboard renders Phase Roles + Scheduled Jobs sections; `AGENT_TYPE_META` is removed; `job.kind` drives rendering.
|
||||||
|
5. **Schedules**: `--install-rule-scan-schedule` and `--install-rule-review-schedule` install OS-native units with self-documenting names.
|
||||||
|
6. **Tests**: `pytest tests/ -v` is green; new tests cover state schemas, scan flows, dashboard API, and conflict-matrix enforcement.
|
||||||
|
7. **No regression**: pre-existing test suite passes unchanged.
|
||||||
|
|
||||||
|
## 11. Locked Decision Index
|
||||||
|
|
||||||
|
All decisions referenced by `(Fn)` above are recorded in the v1 design conversation (this session). They are non-negotiable for v1 implementation. Changes require a design doc update and a new `[unreleased]` changelog entry.
|
||||||
|
|
||||||
|
See `README.md` § "Locked decisions" for the full table.
|
||||||
@@ -0,0 +1,536 @@
|
|||||||
|
# Framework Agent Features — Technical Design
|
||||||
|
|
||||||
|
Companion to `functional.md`. This file is the implementation contract: every line here is what the implementation tasks build. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update.
|
||||||
|
|
||||||
|
## 1. File Map (what v1 adds)
|
||||||
|
|
||||||
|
```
|
||||||
|
~/.automaton/
|
||||||
|
├── models.json # NEW — model manifest (see §2)
|
||||||
|
├── scripts/
|
||||||
|
│ ├── detect_models.py # NEW — probes opencode.json + localhost endpoints
|
||||||
|
│ └── status.py # EXTENDED — new flags (see §4)
|
||||||
|
├── prompts/
|
||||||
|
│ ├── rule-proposer.md # NEW — Rule Proposer session prompt
|
||||||
|
│ ├── rule-reviewer.md # NEW — Rule Reviewer session prompt
|
||||||
|
│ └── onboarding.md # EXTENDED — Step 2e (backlog check)
|
||||||
|
├── automaton/
|
||||||
|
│ └── dashboard/
|
||||||
|
│ ├── html/dashboard.js # EXTENDED — remove AGENT_TYPE_META, two-section render
|
||||||
|
│ └── ui/app.py # EXTENDED — /api/phase-roles endpoint
|
||||||
|
├── .automaton/ # (framework self-hosting: this is ~/.automaton/.automaton/)
|
||||||
|
│ ├── .state.rule-scan # NEW — Rule Proposer state (see §3)
|
||||||
|
│ ├── .state.rule-review # NEW — Rule Reviewer state (see §3)
|
||||||
|
│ ├── RULE_PROPOSALS.md # NEW — Rule Proposer output (append-per-run)
|
||||||
|
│ ├── RULE_REVIEW.md # NEW — Rule Reviewer output (append-per-run)
|
||||||
|
│ └── automaton-rule-scan.sh # NEW — generated by --install-rule-scan-schedule
|
||||||
|
│ automaton-rule-review.sh # NEW — generated by --install-rule-review-schedule
|
||||||
|
└── tests/
|
||||||
|
├── test_model_divergence.py # NEW — manifest, conflict matrix, --transition --model
|
||||||
|
├── test_rule_agents.py # NEW — scan flows, state files, output schemas
|
||||||
|
└── test_dashboard_phase_roles.py # NEW — /api/phase-roles, two-section render
|
||||||
|
```
|
||||||
|
|
||||||
|
Per-project paths mirror the loop convention: `{project}/.automaton/.state.rule-scan`, `{project}/.automaton/RULE_PROPOSALS.md`, etc. For framework self-hosting, the project is `~/.automaton/` itself.
|
||||||
|
|
||||||
|
## 2. `models.json` Schema
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"default": "glm-4.6",
|
||||||
|
"advised": true,
|
||||||
|
"models": [
|
||||||
|
{
|
||||||
|
"name": "glm-4.6",
|
||||||
|
"provider": "opencode",
|
||||||
|
"context_window": 131072,
|
||||||
|
"location": "remote"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "qwen3-coder",
|
||||||
|
"provider": "opencode",
|
||||||
|
"context_window": 131072,
|
||||||
|
"location": "remote"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "llama-3.3-70b",
|
||||||
|
"provider": "localhost",
|
||||||
|
"context_window": 32768,
|
||||||
|
"location": "http://localhost:8080"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `default`: model name used when no role-specific binding exists. Must be present in `models[]`.
|
||||||
|
- `advised`: bool. If `true`, single-LLM mode prints a one-time advisory recommending a second model, then goes silent.
|
||||||
|
- `models[]`: roster. `name` is the unique key. `provider` is informational. `context_window` is informational (framework never inspects capability, D8). `location` is `"remote"` or a localhost URL (for `detect_models.py` probing).
|
||||||
|
- **Missing file** → single-LLM mode (backward compatible). All model-divergence commands are no-ops.
|
||||||
|
- **0-1 models** → single-LLM mode. Advisory once if `advised: true`.
|
||||||
|
- **2+ models** → multi-LLM mode. Hard-block on conflict matrix.
|
||||||
|
|
||||||
|
### 2.1 `detect_models.py`
|
||||||
|
|
||||||
|
```
|
||||||
|
python3 scripts/detect_models.py [--json]
|
||||||
|
```
|
||||||
|
|
||||||
|
1. Parse `opencode.json` (or `opencode.jsonc`) for provider+model entries.
|
||||||
|
2. Probe localhost endpoints: `http://localhost:8080/v1/models`, `http://localhost:11434/api/tags` (Ollama), `http://localhost:1234/v1/models` (LM Studio), `http://localhost:8000/v1/models` (vLLM).
|
||||||
|
3. Merge results, emit a candidate `models.json` to stdout (or write if `--json` not set).
|
||||||
|
4. Used by `install.sh` / `update.sh` / `upgrade.sh` to bootstrap or refresh `models.json`.
|
||||||
|
|
||||||
|
## 3. State Schemas
|
||||||
|
|
||||||
|
### 3.1 `.state.rule-scan`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"last_scan_at": "2026-06-25T10:00:00Z",
|
||||||
|
"last_scanned_task": "fix-context-sizing",
|
||||||
|
"proposals_count": 3,
|
||||||
|
"scanned_tasks_count": 12
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `last_scanned_task`: the most recent task name scanned. Next scan starts after this task (alphabetical or mtime order).
|
||||||
|
- `proposals_count`: cumulative count of proposals written to `RULE_PROPOSALS.md`.
|
||||||
|
- Stored at `{project}/.automaton/.state.rule-scan`. Missing file → first run scans all completed tasks.
|
||||||
|
|
||||||
|
### 3.2 `.state.rule-review`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"last_review_at": "2026-06-25T10:00:00Z",
|
||||||
|
"contradictions_found": 2,
|
||||||
|
"stale_rules": 5,
|
||||||
|
"merges_suggested": 1
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- Stored at `{project}/.automaton/.state.rule-review`. Missing file → first run reviews all rules.
|
||||||
|
|
||||||
|
### 3.3 `.state.models` (per-task)
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"implement": "glm-4.6",
|
||||||
|
"code_review": "qwen3-coder",
|
||||||
|
"bug_find": "qwen3-coder",
|
||||||
|
"adversarial_bug_find": "llama-3.3-70b",
|
||||||
|
"referee": "llama-3.3-70b",
|
||||||
|
"doc_review": null
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- Stored at `{task}/.state.models`. One file per task.
|
||||||
|
- Written by `--transition --model <name>` when entering a phase.
|
||||||
|
- Read by `--claim` (conflict-matrix check) and `--audit` (violation detection).
|
||||||
|
- Roles not yet filled are `null` or absent.
|
||||||
|
|
||||||
|
## 4. `status.py` New Flags
|
||||||
|
|
||||||
|
All model-divergence and rule-agent commands route through `status.py` — no second enforcement surface.
|
||||||
|
|
||||||
|
```
|
||||||
|
# Model-divergence
|
||||||
|
status.py --transition <phase> --task <t> [--model <name>] Records model in .state.models; checks conflict matrix
|
||||||
|
status.py --claim --task <t> --agent <a> [--model <name>] Refuses if model conflicts with filled roles (multi-LLM mode)
|
||||||
|
status.py --audit EXTENDED — +model_divergence category
|
||||||
|
status.py --can-edit [...] UNCHANGED
|
||||||
|
|
||||||
|
# Rule agents
|
||||||
|
status.py --rule-scan [--project <p>] [--dry-run] Scan completed tasks, propose rules to RULE_PROPOSALS.md
|
||||||
|
status.py --install-rule-scan-schedule [--interval S] Install OS-native unit for --rule-scan (default daily)
|
||||||
|
status.py --rule-review [--project <p>] [--dry-run] Consolidate .rules.md, write RULE_REVIEW.md
|
||||||
|
status.py --install-rule-review-schedule [--interval S] Install OS-native unit for --rule-review (default monthly)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4.1 `--transition --model` flow
|
||||||
|
|
||||||
|
1. Load `models.json`. If missing or single-LLM mode → record model (advisory), no conflict check.
|
||||||
|
2. If multi-LLM mode: load `.state.models` for the task. Check the role being entered against the conflict matrix (§5).
|
||||||
|
3. If `--model` not provided: auto-assign next-available non-conflicting model from `models[]`. Refuse if none available.
|
||||||
|
4. If `--model` provided: verify it's in `models[]`. Check conflict matrix. Refuse on violation.
|
||||||
|
5. Write `role: model` to `.state.models`. Transition the phase.
|
||||||
|
|
||||||
|
### 4.2 `--claim --model` flow
|
||||||
|
|
||||||
|
1. Load `models.json`. If single-LLM mode → existing claim logic, no model check.
|
||||||
|
2. If multi-LLM mode: load `.state.models`. Determine the role for the phase being claimed. Check conflict matrix against already-filled roles.
|
||||||
|
3. Refuse if the claiming agent's model conflicts. Error message names the conflicting role and model.
|
||||||
|
|
||||||
|
### 4.3 `--audit` extension
|
||||||
|
|
||||||
|
New audit category `model_divergence`:
|
||||||
|
- For each task with `.state.models`: check all filled roles against the conflict matrix.
|
||||||
|
- Flag violations as `severity: high` (conflict-of-interest is a correctness issue, not a style issue).
|
||||||
|
- Output format mirrors existing audit categories.
|
||||||
|
|
||||||
|
## 5. Conflict Matrix (implementation)
|
||||||
|
|
||||||
|
```python
|
||||||
|
CONFLICT_MATRIX = {
|
||||||
|
"code_review": {"implement"},
|
||||||
|
"bug_find": {"implement"},
|
||||||
|
"adversarial_bug_find": {"implement", "bug_find"},
|
||||||
|
"referee": {"implement", "bug_find", "adversarial_bug_find"},
|
||||||
|
"loop-verify": {"loop-implement"},
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- Key = role being entered. Value = set of roles that must have a different model.
|
||||||
|
- `doc_review`, `code_review`, `bug_find` are NOT in conflict with each other (only `bug_find` ↔ `adversarial_bug_find` conflicts).
|
||||||
|
- Check function: `def _check_conflict(state_models: dict, role: str, model: str, matrix: dict) -> Optional[str]` — returns the conflicting role name or `None`.
|
||||||
|
|
||||||
|
## 6. Rule Proposer Flow (`--rule-scan`)
|
||||||
|
|
||||||
|
```
|
||||||
|
1. Load .state.rule-scan (or init if missing).
|
||||||
|
2. Find completed tasks since last_scanned_task:
|
||||||
|
- Scan tasks/complete/ and tasks with .state phase=complete
|
||||||
|
- Filter by mtime > last_scan_at (or all if first run)
|
||||||
|
- Sort by mtime ascending
|
||||||
|
3. For each task:
|
||||||
|
a. Read BUG_REPORT.md, ADVERSARIAL_BUG_REPORT.md, VERDICT.md (skip if none exist)
|
||||||
|
b. Read current .rules.md (for dedup context — capped at 4k tokens)
|
||||||
|
c. Build proposer prompt (see §7)
|
||||||
|
d. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_proposer_model)
|
||||||
|
e. Parse LLM output for proposed rules (expect RULE_PROPOSALS.md format)
|
||||||
|
f. Append proposals to RULE_PROPOSALS.md
|
||||||
|
g. Update .state.rule-scan (last_scanned_task, proposals_count)
|
||||||
|
4. Write final .state.rule-scan with last_scan_at = now.
|
||||||
|
```
|
||||||
|
|
||||||
|
- `--dry-run`: list tasks that would be scanned, do not invoke harness.
|
||||||
|
- `--project`: scope to a project (default: framework dir).
|
||||||
|
- Errors during a single task scan do not abort the run; the scan continues to the next task and logs the error.
|
||||||
|
|
||||||
|
## 7. Rule Proposer Prompt Shape (`rule-proposer.md`)
|
||||||
|
|
||||||
|
```
|
||||||
|
# Rule Proposer — {date}
|
||||||
|
|
||||||
|
You are scanning completed tasks for failure patterns that should become rules.
|
||||||
|
|
||||||
|
## Current rules (read-only, for dedup)
|
||||||
|
{current_rules} # .rules.md content, capped at 4k tokens
|
||||||
|
|
||||||
|
## Task failure artifacts
|
||||||
|
{bug_report} # BUG_REPORT.md content, capped at 2k tokens
|
||||||
|
{adversarial_report} # ADVERSARIAL_BUG_REPORT.md, capped at 2k tokens
|
||||||
|
{verdict} # VERDICT.md, capped at 2k tokens
|
||||||
|
|
||||||
|
## What to do
|
||||||
|
For each distinct failure pattern you observe:
|
||||||
|
1. Check if a rule already exists in .rules.md that covers it. If so, skip.
|
||||||
|
2. If no existing rule covers it, propose a new rule with:
|
||||||
|
- A concrete example from the task artifacts
|
||||||
|
- The proposed rule text as it would appear in .rules.md
|
||||||
|
|
||||||
|
## Output (strict markdown, no JSON)
|
||||||
|
## Proposed Rule: {title}
|
||||||
|
**Source**: tasks/{task-name}/VERDICT.md
|
||||||
|
**Pattern**: {one-line description}
|
||||||
|
**Example**:
|
||||||
|
{concrete snippet}
|
||||||
|
**Proposed rule text**:
|
||||||
|
{rule text}
|
||||||
|
|
||||||
|
---
|
||||||
|
```
|
||||||
|
|
||||||
|
No `{model}` token in the prompt — the model is selected by the caller and passed to `_invoke_harness`.
|
||||||
|
|
||||||
|
## 8. Rule Reviewer Flow (`--rule-review`)
|
||||||
|
|
||||||
|
```
|
||||||
|
1. Load .state.rule-review (or init if missing).
|
||||||
|
2. Read .rules.md (full file).
|
||||||
|
3. Read recent RULE_PROPOSALS.md entries (since last_review_at).
|
||||||
|
4. Read recent completed-task summaries (last 30 days) for staleness context.
|
||||||
|
5. Build reviewer prompt (see §9).
|
||||||
|
6. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_reviewer_model).
|
||||||
|
7. Parse LLM output for review sections (contradictions, stale, missing examples, merges).
|
||||||
|
8. Append to RULE_REVIEW.md.
|
||||||
|
9. Update .state.rule-review.
|
||||||
|
```
|
||||||
|
|
||||||
|
- `--dry-run`: report what would be reviewed, do not invoke harness.
|
||||||
|
|
||||||
|
## 9. Rule Reviewer Prompt Shape (`rule-reviewer.md`)
|
||||||
|
|
||||||
|
```
|
||||||
|
# Rule Reviewer — {date}
|
||||||
|
|
||||||
|
You are consolidating .rules.md for contradictions, staleness, and missing examples.
|
||||||
|
|
||||||
|
## Current rules (full)
|
||||||
|
{rules_content} # .rules.md, full file
|
||||||
|
|
||||||
|
## Recent proposals (since last review)
|
||||||
|
{recent_proposals} # RULE_PROPOSALS.md entries since last_review_at
|
||||||
|
|
||||||
|
## Recent completed tasks (last 30 days, for staleness context)
|
||||||
|
{task_summaries} # one-line per task: name + phase + completion date
|
||||||
|
|
||||||
|
## What to check
|
||||||
|
1. Contradictions: rules that conflict with each other.
|
||||||
|
2. Stale rules: no observed instance in last 30 days.
|
||||||
|
3. Rules missing examples: any rule without a concrete example.
|
||||||
|
4. Merge candidates: overlapping rules that could be consolidated.
|
||||||
|
|
||||||
|
## Output (strict markdown, no JSON)
|
||||||
|
## Contradictions Found
|
||||||
|
- ...
|
||||||
|
## Stale Rules
|
||||||
|
- ...
|
||||||
|
## Rules Missing Examples
|
||||||
|
- ...
|
||||||
|
## Merge Candidates
|
||||||
|
- ...
|
||||||
|
```
|
||||||
|
|
||||||
|
## 10. Harness Invocation (direct, not loop)
|
||||||
|
|
||||||
|
Rule agents reuse `loop-runner._invoke_harness` directly — they are NOT loops. The function signature (from `loop-runner.py:366-404`):
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _invoke_harness(harness_command: str, prompt_content: str, cwd: str, env: dict = None) -> str:
|
||||||
|
```
|
||||||
|
|
||||||
|
`status.py --rule-scan` calls this as:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from loop_runner import _invoke_harness
|
||||||
|
output = _invoke_harness(
|
||||||
|
harness_command=rule_harness_command, # from schedule config or default
|
||||||
|
prompt_content=resolved_prompt, # rule-proposer.md with tokens substituted
|
||||||
|
cwd=str(project_dir),
|
||||||
|
env={"AUTOMATON_RULE_ROLE": "proposer"}
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
`{model}` substitution: if the harness command contains `{model}`, it's replaced with the rule agent's configured model. Until model-divergence ships, this is the default model.
|
||||||
|
|
||||||
|
### 10.1 Schedule Config for Rule Agents
|
||||||
|
|
||||||
|
Rule agents do not use `loop.json`. Their config is embedded in the schedule stub:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
#!/usr/bin/env bash
|
||||||
|
cd "<project_root>"
|
||||||
|
python3 "<framework>/scripts/status.py" --rule-scan --model <name>
|
||||||
|
```
|
||||||
|
|
||||||
|
The `--model` flag is optional and ignored in single-LLM mode. In multi-LLM mode it sets the rule agent's model (subject to conflict-of-interest checks once enforced).
|
||||||
|
|
||||||
|
## 11. Scheduler Unit Generation
|
||||||
|
|
||||||
|
Mirrors `cmd_install_cleanup_schedule` (`status.py:2188-2260`) exactly:
|
||||||
|
|
||||||
|
### 11.1 `--install-rule-scan-schedule`
|
||||||
|
|
||||||
|
```python
|
||||||
|
def cmd_install_rule_scan_schedule(args) -> int:
|
||||||
|
interval = args.interval if args.interval else 86400 # daily
|
||||||
|
# 1. Write stub: automaton-rule-scan.sh
|
||||||
|
# 2. Platform dispatch:
|
||||||
|
# Darwin → ~/Library/LaunchAgents/com.automaton.rule-scan.plist
|
||||||
|
# Linux → crontab block via _install_cron_block_generic
|
||||||
|
# Windows → schtasks /create /tn "AutomatonRuleScan"
|
||||||
|
```
|
||||||
|
|
||||||
|
### 11.2 `--install-rule-review-schedule`
|
||||||
|
|
||||||
|
```python
|
||||||
|
def cmd_install_rule_review_schedule(args) -> int:
|
||||||
|
interval = args.interval if args.interval else 2592000 # monthly
|
||||||
|
# Same pattern, labels: com.automaton.rule-review / AutomatonRuleReview
|
||||||
|
```
|
||||||
|
|
||||||
|
### 11.3 `_list_scheduled_jobs` extension
|
||||||
|
|
||||||
|
`_list_scheduled_jobs` (`status.py:2282`) gains recognition for new labels:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _launchd_label_kind(label: str) -> tuple[str, str]:
|
||||||
|
if label.startswith("com.automaton.loop."):
|
||||||
|
return ("loop", label[len("com.automaton.loop."):])
|
||||||
|
if label in ("com.automaton.cleanup",):
|
||||||
|
return ("cleanup", "")
|
||||||
|
if label in ("com.automaton.rule-scan",):
|
||||||
|
return ("rule-scan", "")
|
||||||
|
if label in ("com.automaton.rule-review",):
|
||||||
|
return ("rule-review", "")
|
||||||
|
...
|
||||||
|
```
|
||||||
|
|
||||||
|
This makes rule-scan and rule-review jobs appear in `/api/scheduled` with their real `kind`, which the Agent tab renders directly.
|
||||||
|
|
||||||
|
## 12. Agent Tab Data Flow
|
||||||
|
|
||||||
|
### 12.1 New endpoint: `/api/phase-roles`
|
||||||
|
|
||||||
|
`app.py` gains a handler:
|
||||||
|
|
||||||
|
```python
|
||||||
|
elif self.path == "/api/phase-roles":
|
||||||
|
self._serve_phase_roles()
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _serve_phase_roles(self):
|
||||||
|
# 1. Parse .agent.md Agent Configuration for role definitions
|
||||||
|
# 2. Load all tasks via status.py module
|
||||||
|
# 3. For each role, count tasks in that role's phases
|
||||||
|
# 4. Return JSON:
|
||||||
|
{
|
||||||
|
"roles": [
|
||||||
|
{"id": "researcher", "label": "Researcher", "icon": "🔬",
|
||||||
|
"phases": ["research", "decomposition", "design", "test_design"],
|
||||||
|
"active_tasks": 2, "status": "active"},
|
||||||
|
...
|
||||||
|
],
|
||||||
|
"available": True
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Role icons (self-documenting, per `.rules.md` Self-Documenting UI Names):
|
||||||
|
|
||||||
|
| Role | Icon |
|
||||||
|
|---|---|
|
||||||
|
| researcher | 🔬 |
|
||||||
|
| implementer | ⚙️ |
|
||||||
|
| code-reviewer | 👁️ |
|
||||||
|
| bug-hunter | 🐛 |
|
||||||
|
| referee | ⚖️ |
|
||||||
|
| orchestrator | 🎯 |
|
||||||
|
|
||||||
|
### 12.2 `dashboard.js` changes
|
||||||
|
|
||||||
|
**Remove**: `AGENT_TYPE_META` (lines 330-334), `AGENT_TYPES` (337), `AGENT_TYPE_META_FALLBACK` (338), `_resolveAgentType` (351-354).
|
||||||
|
|
||||||
|
**Replace `renderAgentTab`** with a two-section render:
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
async function renderAgentTab() {
|
||||||
|
const panel = document.getElementById('agent-panel');
|
||||||
|
panel.innerHTML = '<div class="bg-loading">Loading…</div>';
|
||||||
|
|
||||||
|
const [rolesRes, schedRes] = await Promise.all([
|
||||||
|
fetch('/api/phase-roles').then(r => r.json()).catch(() => ({roles: [], available: false})),
|
||||||
|
fetchSchedule(),
|
||||||
|
]);
|
||||||
|
|
||||||
|
// Section 1: Phase Roles
|
||||||
|
const rolesHtml = rolesRes.available ? renderPhaseRoles(rolesRes.roles)
|
||||||
|
: '<div class="bg-empty">Phase roles require .agent.md Agent Configuration.</div>';
|
||||||
|
|
||||||
|
// Section 2: Scheduled Jobs
|
||||||
|
const jobsHtml = renderScheduledJobs(schedRes.jobs || []);
|
||||||
|
|
||||||
|
panel.innerHTML = `
|
||||||
|
<div class="agent-section">
|
||||||
|
<h3>Phase Roles</h3>
|
||||||
|
<div class="bg-grid">${rolesHtml}</div>
|
||||||
|
</div>
|
||||||
|
<div class="agent-section">
|
||||||
|
<h3>Scheduled Jobs</h3>
|
||||||
|
<div class="bg-grid">${jobsHtml}</div>
|
||||||
|
</div>`;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**`renderScheduledJobs`** uses `job.kind` directly (no fake type resolution):
|
||||||
|
|
||||||
|
```javascript
|
||||||
|
const JOB_META = {
|
||||||
|
cleanup: { icon: '🧹', label: 'Cleanup Archiver' },
|
||||||
|
loop: { icon: '🔄', label: (j) => `Loop: ${j.name}` },
|
||||||
|
'rule-scan': { icon: '📝', label: 'Rule Proposer' },
|
||||||
|
'rule-review':{ icon: '📋', label: 'Rule Reviewer' },
|
||||||
|
};
|
||||||
|
```
|
||||||
|
|
||||||
|
## 13. Loop Integration (model-divergence)
|
||||||
|
|
||||||
|
### 13.1 `loop.json` per-role model
|
||||||
|
|
||||||
|
```json
|
||||||
|
"roles": {
|
||||||
|
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
|
||||||
|
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
|
||||||
|
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `model` is optional. If absent, uses `models.json` `default`.
|
||||||
|
- `loop-verify` model is checked against `loop-implement` model in `--check-gate` (multi-LLM mode).
|
||||||
|
|
||||||
|
### 13.2 `{model}` substitution in `_invoke_harness`
|
||||||
|
|
||||||
|
`loop-runner.py:366-404` `_invoke_harness` gains `{model}` token substitution:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _invoke_harness(harness_command, prompt_content, cwd, env=None, model=None):
|
||||||
|
if model and "{model}" in harness_command:
|
||||||
|
harness_command = harness_command.replace("{model}", model)
|
||||||
|
...
|
||||||
|
```
|
||||||
|
|
||||||
|
The caller passes `model` from the role config. If the harness command has no `{model}` token, the model is informational only (the harness picks its own).
|
||||||
|
|
||||||
|
### 13.3 `--check-gate` model-divergence check
|
||||||
|
|
||||||
|
In multi-LLM mode, `--check-gate` adds:
|
||||||
|
- Load `loop.json` roles. Compare `verify.model` vs `implement.model`.
|
||||||
|
- If same model and multi-LLM mode → halt as `model_conflict` (new halt reason, or reuse `human_intervention` with a descriptive message).
|
||||||
|
|
||||||
|
## 14. Test Coverage
|
||||||
|
|
||||||
|
### 14.1 `test_model_divergence.py`
|
||||||
|
|
||||||
|
- `test_models_json_missing_single_llm_mode` — no file → advisory, no blocks.
|
||||||
|
- `test_single_model_advisory_once` — 1 model, `advised: true` → advisory printed once, then silent.
|
||||||
|
- `test_multi_llm_conflict_matrix` — 2+ models, `--transition --model` records, `--claim` refuses conflict.
|
||||||
|
- `test_auto_assign_next_available` — no `--model` flag → auto-assigns non-conflicting model.
|
||||||
|
- `test_auto_assign_exhausted` — all models conflict → refuse.
|
||||||
|
- `test_audit_model_divergence` — `--audit` flags conflict-matrix violations.
|
||||||
|
- `test_loop_verify_neq_implement` — `--check-gate` halts on same model in multi-LLM mode.
|
||||||
|
|
||||||
|
### 14.2 `test_rule_agents.py`
|
||||||
|
|
||||||
|
- `test_rule_scan_finds_completed_tasks` — seeded completed task with VERDICT.md → proposal written.
|
||||||
|
- `test_rule_scan_state_tracking` — `.state.rule-scan` updated with last_scanned_task + count.
|
||||||
|
- `test_rule_scan_dedup` — existing rule in `.rules.md` → not re-proposed.
|
||||||
|
- `test_rule_scan_dry_run` — no harness invocation, lists candidates.
|
||||||
|
- `test_rule_review_finds_contradictions` — seeded `.rules.md` with contradiction → review written.
|
||||||
|
- `test_rule_review_state_tracking` — `.state.rule-review` updated.
|
||||||
|
- `test_install_rule_scan_schedule` — stub + plist created with correct labels.
|
||||||
|
- `test_install_rule_review_schedule` — stub + plist created with correct labels.
|
||||||
|
|
||||||
|
### 14.3 `test_dashboard_phase_roles.py`
|
||||||
|
|
||||||
|
- `test_api_phase_roles` — `/api/phase-roles` returns 6 roles with correct phases.
|
||||||
|
- `test_phase_roles_active_count` — tasks in phases → correct active_tasks count.
|
||||||
|
- `test_scheduled_jobs_new_kinds` — rule-scan and rule-review jobs appear with correct `kind`.
|
||||||
|
- `test_agent_type_meta_removed` — `AGENT_TYPE_META` no longer in dashboard.js (grep test).
|
||||||
|
|
||||||
|
## 15. Rollout (3 sequential tasks for model-divergence)
|
||||||
|
|
||||||
|
The model-divergence-enforcement parent task decomposes into 3 subtasks:
|
||||||
|
|
||||||
|
1. **manifest+detection**: `models.json` schema, `detect_models.py`, `install.sh`/`update.sh`/`upgrade.sh` integration, `config.md` section, onboarding Step 2d.
|
||||||
|
2. **interactive enforcement+audit**: `.state.models`, `--transition --model`, `--claim --model` conflict check, `--audit` model_divergence category, dashboard badges.
|
||||||
|
3. **loop enforcement+dashboard**: `loop.json` per-role model, `{model}` substitution, `--check-gate` model check, loop dashboard badges.
|
||||||
|
|
||||||
|
Rule agents (FW-2, FW-3) and Agent tab (FW-1) are backlog items, picked up after model-divergence ships (for FW-2/FW-3) or independently (for FW-1).
|
||||||
|
|
||||||
|
## 16. Locked Decision Index
|
||||||
|
|
||||||
|
All decisions referenced by `(Fn)` are in `README.md` § "Locked decisions". Implementation must conform. Deviations require a design doc update + `[unreleased]` CHANGELOG entry.
|
||||||
@@ -6,6 +6,8 @@ Work queue for the self-improvement loop after task 7 lands. Items not assigned
|
|||||||
|
|
||||||
A loop configured with `work_source.kind = "backlog"` reads this file, picks the topmost `[ ]` item, drafts an implementation, transitions through phases, hands off to a human reviewer. Mark items `[x]` when complete; move items to `DONE.md` (created later) on closure.
|
A loop configured with `work_source.kind = "backlog"` reads this file, picks the topmost `[ ]` item, drafts an implementation, transitions through phases, hands off to a human reviewer. Mark items `[x]` when complete; move items to `DONE.md` (created later) on closure.
|
||||||
|
|
||||||
|
**Sibling backlog**: `design/framework/BACKLOG.md` covers framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). A loop with `work_source.area = "framework"` reads that file instead of this one.
|
||||||
|
|
||||||
## v1.1 — framework manages its own docs
|
## v1.1 — framework manages its own docs
|
||||||
|
|
||||||
- [ ] **design-update-loop-template** — `templates/loops/design-update/` that keeps `design/<area>/*.md` in sync with the code it documents. Triggered by `last-read-sha` drift detection.
|
- [ ] **design-update-loop-template** — `templates/loops/design-update/` that keeps `design/<area>/*.md` in sync with the code it documents. Triggered by `last-read-sha` drift detection.
|
||||||
|
|||||||
@@ -15,6 +15,7 @@ Per-session manual driving doesn't scale against the framework's growing backlog
|
|||||||
- [`functional.md`](functional.md) — what v1 does, roles, the five deaths, blast radius, schedules, success criteria. **Read this first.**
|
- [`functional.md`](functional.md) — what v1 does, roles, the five deaths, blast radius, schedules, success criteria. **Read this first.**
|
||||||
- [`technical.md`](technical.md) — the implementation contract: file map, `.state.loop` schema, `status.py` flags, gate checks, runner flow, test coverage. **Read this if you're implementing.**
|
- [`technical.md`](technical.md) — the implementation contract: file map, `.state.loop` schema, `status.py` flags, gate checks, runner flow, test coverage. **Read this if you're implementing.**
|
||||||
- [`BACKLOG.md`](BACKLOG.md) — v1.1 and deferred items (Scope 2 design-update loop, Scope 3 self-designing, parallel mode, dashboard panel). The self-improvement loop's work queue.
|
- [`BACKLOG.md`](BACKLOG.md) — v1.1 and deferred items (Scope 2 design-update loop, Scope 3 self-designing, parallel mode, dashboard panel). The self-improvement loop's work queue.
|
||||||
|
- **Sibling design**: [`../framework/`](../framework/) — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). Has its own `BACKLOG.md` consumable via `work_source.area = "framework"`.
|
||||||
|
|
||||||
## v1 scope (locked)
|
## v1 scope (locked)
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,14 @@
|
|||||||
|
{
|
||||||
|
"current_task": null,
|
||||||
|
"halt_reason": null,
|
||||||
|
"iteration_count": 0,
|
||||||
|
"last_tick_at": null,
|
||||||
|
"last_verdict": null,
|
||||||
|
"name": "self-improvement",
|
||||||
|
"resumed_count": 0,
|
||||||
|
"schema_version": 1,
|
||||||
|
"score_history": [],
|
||||||
|
"status": "running",
|
||||||
|
"worktree_branch": null,
|
||||||
|
"worktree_path": null
|
||||||
|
}
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
cd "/Users/laptran/.automaton"
|
||||||
|
python3 "/Users/laptran/.automaton/scripts/loop-runner.py" --mode tick --loop "self-improvement"
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
{
|
||||||
|
"name": "self-improvement",
|
||||||
|
"description": "Ticks against status.py --audit on the framework's own repo",
|
||||||
|
"schedule": {
|
||||||
|
"interval_seconds": 3600
|
||||||
|
},
|
||||||
|
"brakes": {
|
||||||
|
"max_iterations": 10,
|
||||||
|
"max_budget_usd": null,
|
||||||
|
"score_plateau_window": 3
|
||||||
|
},
|
||||||
|
"blast_radius": {
|
||||||
|
"file_scope": ["scripts/", "prompts/", "tests/", "design/"],
|
||||||
|
"base_branch": "main",
|
||||||
|
"use_worktree": true
|
||||||
|
},
|
||||||
|
"work_source": {
|
||||||
|
"kind": "audit",
|
||||||
|
"project": "~/.automaton/"
|
||||||
|
},
|
||||||
|
"acceptance_criteria": [
|
||||||
|
"Audit findings resolved (no outstanding Cat-1/Cat-2/Cat-4 violations on the resolved task)",
|
||||||
|
"All R-numbers from the task SPEC.md implemented",
|
||||||
|
"Tests pass with no regressions",
|
||||||
|
"Pipeline driven to complete"
|
||||||
|
],
|
||||||
|
"roles": {
|
||||||
|
"implement": {"prompt": "loop-implement.md"},
|
||||||
|
"verify": {"prompt": "loop-verifier.md"},
|
||||||
|
"orchestrate": {"prompt": "loop-orchestrate.md"}
|
||||||
|
},
|
||||||
|
"outputs": {
|
||||||
|
"retention": 20
|
||||||
|
},
|
||||||
|
"harness": {
|
||||||
|
"command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -53,4 +53,15 @@ Picked from BUG_REPORTs of the loop tasks and from `design/loops/BACKLOG.md` def
|
|||||||
|
|
||||||
## NEXT
|
## NEXT
|
||||||
|
|
||||||
Resume task 2 (`add-state-loop-lock`): SPEC is already written, state is `research:awaiting_approval`. Approve it, transition to implement, write the `_loop_lock` helper + wrap callsites in `status.py` and `loop-runner.py`, add `tests/test_state_loop_lock.py` (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order.
|
Resume task 2 (`add-state-loop-lock`): SPEC is already written, state is `research:awaiting_approval`. Approve it, transition to implement, write the `_loop_lock` helper + wrap callsites in `status.py` and `loop-runner.py`, add `tests/test_state_loop_lock.py` (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order.
|
||||||
|
|
||||||
|
## Session 2026-06-25 — Framework agent features design
|
||||||
|
|
||||||
|
- Created `design/framework/` with `README.md`, `functional.md`, `technical.md`, `BACKLOG.md` — design for three framework-level agent features: model-divergence enforcement, rule agents (Proposer + Reviewer), Agent tab redesign.
|
||||||
|
- **Model-divergence enforcement** is a manual task (`model-divergence-enforcement`), decomposed into 3 subtasks: (1) manifest+detection, (2) interactive enforcement+audit, (3) loop enforcement+dashboard. Not a backlog item.
|
||||||
|
- **Rule agents** (FW-2, FW-3) are backlog items depending on model-divergence shipping first (conflict-of-interest LLM binding). Rule Proposer runs daily, Rule Reviewer runs monthly. Both use direct harness invocation (reuse `loop-runner._invoke_harness`), not loop infrastructure.
|
||||||
|
- **Agent tab redesign** (FW-1) is a backlog item with no dependencies. Replaces 4 fake `AGENT_TYPE_META` types with Phase Roles (6 roles from `.agent.md`) + Scheduled Jobs (real `job.kind`).
|
||||||
|
- Decision: rule agents use **direct harness invocation** (not loops, not standalone status.py commands).
|
||||||
|
- Decision: model-divergence conflict-of-interest is **designed now, enforced later** — rule agents carry `model` fields in config but hard-blocking activates only when `models.json` exists and multi-LLM mode is detected.
|
||||||
|
- Cross-references updated: `AGENTS.md` repo layout, `README.md` loop config table, `design/loops/README.md`, `design/loops/BACKLOG.md`, `CHANGELOG.md`, `.onboarding.md` (Backlog section), `prompts/onboarding.md` (Step 2d).
|
||||||
|
- Next: switch on self-improvement loop (`--create-loop self-improvement --from-template self-improvement` + `--install-schedule`), then create `model-divergence-enforcement` parent task and decompose into 3 subtasks.
|
||||||
@@ -75,6 +75,13 @@ Install the automaton pre-commit hook to block commits when no task is in an edi
|
|||||||
4. Verify the hook: `python ~/.automaton/scripts/status.py --can-edit --project {project}` should return exit code 1 (DENIED) since no tasks exist yet.
|
4. Verify the hook: `python ~/.automaton/scripts/status.py --can-edit --project {project}` should return exit code 1 (DENIED) since no tasks exist yet.
|
||||||
5. Note in the onboarding report whether the hook was installed.
|
5. Note in the onboarding report whether the hook was installed.
|
||||||
|
|
||||||
|
### Step 2d: Backlog Check
|
||||||
|
|
||||||
|
Check for outstanding design work that may need attention:
|
||||||
|
1. Read `~/.automaton/design/loops/BACKLOG.md` for loop engineering work queue items.
|
||||||
|
2. Read `~/.automaton/design/framework/BACKLOG.md` for framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement).
|
||||||
|
3. Note in the onboarding report whether there are unchecked `- [ ]` items in either backlog that the project owner may want to pick up manually or via the self-improvement loop (`work_source.area = "loops"` or `"framework"`).
|
||||||
|
|
||||||
## Output
|
## Output
|
||||||
|
|
||||||
Create or update the following inside {project}/.automaton/:
|
Create or update the following inside {project}/.automaton/:
|
||||||
|
|||||||
@@ -1,2 +1,2 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7
|
python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/private/var/folders/f5/yv0dzbnx47x3yp8sc_2519gh0000gn/T/pytest-of-laptran/pytest-98/test_uninstall_via_disabled_re0"
|
||||||
|
|||||||
+59
-17
@@ -24,15 +24,17 @@ AUTOMATON_DIR = Path.home() / ".automaton"
|
|||||||
|
|
||||||
PHASE_PRIORITY = {
|
PHASE_PRIORITY = {
|
||||||
"referee": 12,
|
"referee": 12,
|
||||||
"doc_review": 10,
|
"doc_review": 11,
|
||||||
"adversarial_bug_find": 8,
|
"adversarial_bug_find": 10,
|
||||||
"bug_find": 6,
|
"bug_find": 9,
|
||||||
"implement": 5,
|
"code_review": 8,
|
||||||
"test_design": 4,
|
"implement": 7,
|
||||||
"design": 3,
|
"test_design": 6,
|
||||||
"decomposition": 2,
|
"design": 5,
|
||||||
"research": 1,
|
"decomposition": 4,
|
||||||
"new": 0,
|
"research": 3,
|
||||||
|
"new": 2,
|
||||||
|
"human_intervention": 1,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
@@ -60,13 +62,28 @@ def _read_state(task_path: Path) -> Optional[str]:
|
|||||||
|
|
||||||
|
|
||||||
def _all_tasks(project_dir: Path) -> list[tuple[str, Path]]:
|
def _all_tasks(project_dir: Path) -> list[tuple[str, Path]]:
|
||||||
tasks_dir = project_dir / "tasks"
|
"""Scan all tasks (including subtasks) for a project.
|
||||||
if not tasks_dir.is_dir():
|
|
||||||
|
Mirrors ``status.py:_all_task_dirs``: skips ``tasks/complete/``,
|
||||||
|
recurses into ``subtasks/``, uses ``.automaton/tasks`` for non-framework
|
||||||
|
projects.
|
||||||
|
"""
|
||||||
|
if project_dir == AUTOMATON_DIR:
|
||||||
|
base = AUTOMATON_DIR / "tasks"
|
||||||
|
else:
|
||||||
|
base = project_dir / ".automaton" / "tasks"
|
||||||
|
if not base.is_dir():
|
||||||
return []
|
return []
|
||||||
result = []
|
result = []
|
||||||
for subdir in sorted(tasks_dir.iterdir()):
|
for entry in sorted(base.iterdir()):
|
||||||
if subdir.is_dir():
|
if not entry.is_dir() or entry.name.startswith(".") or entry.name == "complete":
|
||||||
result.append((subdir.name, subdir))
|
continue
|
||||||
|
result.append((entry.name, entry))
|
||||||
|
subtasks = entry / "subtasks"
|
||||||
|
if subtasks.exists():
|
||||||
|
for sub in sorted(subtasks.iterdir()):
|
||||||
|
if sub.is_dir() and not sub.name.startswith("."):
|
||||||
|
result.append((f"{entry.name}/{sub.name}", sub))
|
||||||
return result
|
return result
|
||||||
|
|
||||||
|
|
||||||
@@ -87,11 +104,16 @@ def scan_all_tasks(project_dir: Path) -> list[dict]:
|
|||||||
|
|
||||||
|
|
||||||
def is_terminal(task: dict) -> bool:
|
def is_terminal(task: dict) -> bool:
|
||||||
"""Check if a task is in a terminal state."""
|
"""Check if a task is in a terminal state.
|
||||||
|
|
||||||
|
Only ``complete`` is terminal. ``human_intervention`` has legal
|
||||||
|
transitions (→ referee, → complete) so the autopilot can still
|
||||||
|
suggest next steps for it.
|
||||||
|
"""
|
||||||
phase = task.get("phase")
|
phase = task.get("phase")
|
||||||
if phase is None:
|
if phase is None:
|
||||||
return False
|
return False
|
||||||
return phase in ("complete", "human_intervention")
|
return phase == "complete"
|
||||||
|
|
||||||
|
|
||||||
def needs_user_input(task: dict) -> bool:
|
def needs_user_input(task: dict) -> bool:
|
||||||
@@ -101,6 +123,8 @@ def needs_user_input(task: dict) -> bool:
|
|||||||
return False
|
return False
|
||||||
if phase.endswith(":awaiting_approval"):
|
if phase.endswith(":awaiting_approval"):
|
||||||
return True
|
return True
|
||||||
|
if phase == "human_intervention":
|
||||||
|
return True
|
||||||
task_path = Path(task["path"])
|
task_path = Path(task["path"])
|
||||||
verdict_file = task_path / "VERDICT.md"
|
verdict_file = task_path / "VERDICT.md"
|
||||||
if verdict_file.exists():
|
if verdict_file.exists():
|
||||||
@@ -246,6 +270,13 @@ def cmd_drive(args):
|
|||||||
print(
|
print(
|
||||||
f" python ~/.automaton/scripts/status.py --approve --task {t['name']} --project {project_dir}"
|
f" python ~/.automaton/scripts/status.py --approve --task {t['name']} --project {project_dir}"
|
||||||
)
|
)
|
||||||
|
elif t["phase"] == "human_intervention":
|
||||||
|
print(
|
||||||
|
f" Review {t['name']} — transition to referee or complete:"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f" python ~/.automaton/scripts/status.py --transition referee --task {t['name']} --project {project_dir}"
|
||||||
|
)
|
||||||
else:
|
else:
|
||||||
print(f" Review {t['name']}/VERDICT.md and take action")
|
print(f" Review {t['name']}/VERDICT.md and take action")
|
||||||
return 0
|
return 0
|
||||||
@@ -313,8 +344,19 @@ def cmd_drive(args):
|
|||||||
elif base == "implement":
|
elif base == "implement":
|
||||||
print("→ Write implementation, generate IMPLEMENTATION.md, then transition:")
|
print("→ Write implementation, generate IMPLEMENTATION.md, then transition:")
|
||||||
print(
|
print(
|
||||||
f" python ~/.automaton/scripts/status.py --transition bug_find --task {task['name']} --project {project_dir}"
|
f" python ~/.automaton/scripts/status.py --transition code_review --task {task['name']} --project {project_dir}"
|
||||||
)
|
)
|
||||||
|
elif base == "code_review":
|
||||||
|
if phase == "code_review":
|
||||||
|
print("→ Generate CODE_REVIEW.md, then transition to awaiting_approval:")
|
||||||
|
print(
|
||||||
|
f" python ~/.automaton/scripts/status.py --transition code_review:awaiting_approval --task {task['name']} --project {project_dir}"
|
||||||
|
)
|
||||||
|
elif phase == "code_review:approved":
|
||||||
|
print("→ Transition to bug_find:")
|
||||||
|
print(
|
||||||
|
f" python ~/.automaton/scripts/status.py --transition bug_find --task {task['name']} --project {project_dir}"
|
||||||
|
)
|
||||||
elif base == "bug_find":
|
elif base == "bug_find":
|
||||||
print("→ Generate BUG_REPORT.md, then transition:")
|
print("→ Generate BUG_REPORT.md, then transition:")
|
||||||
print(
|
print(
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
decomposition:approved
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
research:approved|2026-06-25T11:11:51.910546+00:00|user
|
||||||
|
decomposition:approved|2026-06-25T11:12:40.642310+00:00|user
|
||||||
@@ -0,0 +1,113 @@
|
|||||||
|
# DECOMPOSITION — model-divergence-enforcement
|
||||||
|
|
||||||
|
## Method
|
||||||
|
|
||||||
|
Decompose by **dependency layer**, not by file. Each subtask builds on the previous
|
||||||
|
one's foundation. The SPEC (`tasks/model-divergence-enforcement/SPEC.md`) defines 3
|
||||||
|
sequential subtasks with strict dependency ordering.
|
||||||
|
|
||||||
|
## Sub-tasks (3, sequential)
|
||||||
|
|
||||||
|
### subtask-1: `mde-manifest-detection`
|
||||||
|
**Scope:** `models.json` schema + loader + `scripts/detect_models.py` probe + install integration.
|
||||||
|
**Files touched:**
|
||||||
|
- `scripts/detect_models.py` (new — probe opencode.json providers + localhost endpoints 8080/11434/1234/8000)
|
||||||
|
- `scripts/status.py` (add `_load_models_manifest()` helper, `_get_mode()` — single vs multi-LLM)
|
||||||
|
- `scripts/install.sh` (call `detect_models.py` after `vram_detect.py`)
|
||||||
|
- `scripts/update.sh` (same)
|
||||||
|
- `scripts/upgrade.sh` (same)
|
||||||
|
- `config.md` (add `## Available Models` section template)
|
||||||
|
- `prompts/onboarding.md` (Step 2d — model config check)
|
||||||
|
- `tests/test_model_divergence.py` (new — manifest loading, single vs multi mode, missing file backward compat)
|
||||||
|
**Not touched:** `status.py --transition`, `--claim`, `--audit`, `loop-runner.py`, dashboard code.
|
||||||
|
**Acceptance:**
|
||||||
|
1. `models.json` missing → `_get_mode()` returns `"single"`, all model commands are no-ops.
|
||||||
|
2. `models.json` with 0-1 models → `_get_mode()` returns `"single"`.
|
||||||
|
3. `models.json` with 2+ models → `_get_mode()` returns `"multi"`.
|
||||||
|
4. `detect_models.py` probes localhost endpoints and prints a candidate manifest (JSON to stdout).
|
||||||
|
5. `pytest tests/test_model_divergence.py -v` green.
|
||||||
|
6. `pytest tests/ -q` green (no regressions).
|
||||||
|
**Peak context estimate:** ~6k tokens (new script + status.py helper + tests).
|
||||||
|
**Run order:** first. Foundational — subtasks 2 and 3 depend on this.
|
||||||
|
|
||||||
|
### subtask-2: `mde-interactive-enforcement`
|
||||||
|
**Scope:** `.state.models` schema + conflict matrix + `--transition --model` / `--claim --model` enforcement + audit category + dashboard badges.
|
||||||
|
**Depends on:** subtask-1 (needs `_load_models_manifest()` and `_get_mode()`).
|
||||||
|
**Files touched:**
|
||||||
|
- `scripts/status.py`:
|
||||||
|
- `CONFLICT_MATRIX` constant (locked matrix from SPEC §15)
|
||||||
|
- `_check_conflict(current_phase, current_model, next_phase, next_model)` helper
|
||||||
|
- `cmd_transition`: add `--model` arg; in multi-LLM mode, check conflict matrix before allowing transition
|
||||||
|
- `cmd_claim`: add `--model` arg; refuse if model conflicts with existing `.state.models` entry
|
||||||
|
- `cmd_audit`: add `model_divergence` category (scan `.state.models` for violations)
|
||||||
|
- `.state.models` writer (JSON: `{implementer: "model-name", code_reviewer: "model-name", ...}`)
|
||||||
|
- `automaton/dashboard/html/dashboard.js`: model badge on task cards (read from `.state.models`)
|
||||||
|
- `automaton/dashboard/ui/app.py`: include `.state.models` in `/api/tasks` response
|
||||||
|
- `tests/test_model_divergence.py`: conflict matrix tests, auto-assign tests, audit category tests, dashboard badge tests
|
||||||
|
**Not touched:** `loop-runner.py`, `loop.json` schema, `--check-gate`.
|
||||||
|
**Acceptance:**
|
||||||
|
1. Single-LLM mode: `--transition --model <name>` records model but never refuses. Advisory printed once if `advised: true`.
|
||||||
|
2. Multi-LLM mode: `--transition --model <name>` refuses if `<name>` conflicts with `.state.models` for a conflicting phase.
|
||||||
|
3. Multi-LLM mode: `--claim --model <name>` refuses on conflict.
|
||||||
|
4. `--audit` flags `model_divergence` violations (e.g., same model in implementer + code_reviewer).
|
||||||
|
5. Dashboard task cards show model badges when `.state.models` exists.
|
||||||
|
6. `pytest tests/test_model_divergence.py -v` green.
|
||||||
|
7. `pytest tests/ -q` green (no regressions).
|
||||||
|
**Peak context estimate:** ~8k tokens (status.py surgery + dashboard + tests).
|
||||||
|
**Run order:** second, AFTER subtask-1.
|
||||||
|
|
||||||
|
### subtask-3: `mde-loop-enforcement`
|
||||||
|
**Scope:** `loop.json` per-role model field + `{model}` substitution in loop-runner + `--check-gate` model-divergence brake.
|
||||||
|
**Depends on:** subtask-2 (needs `CONFLICT_MATRIX` and `_check_conflict`).
|
||||||
|
**Files touched:**
|
||||||
|
- `scripts/loop-runner.py`:
|
||||||
|
- `_invoke_harness` (line 366-404): add `{model}` placeholder substitution from `loop.json` role config
|
||||||
|
- `_find_work_backlog`: no change (already supports `work_source.area`)
|
||||||
|
- `scripts/status.py`:
|
||||||
|
- `--check-gate`: add model-divergence brake gate (loop-verify model ≠ loop-implement model in multi-LLM mode)
|
||||||
|
- `cmd_install_schedule` / `cmd_create_loop`: validate `loop.json` per-role `model` fields against `models.json`
|
||||||
|
- `templates/loops/`: update loop templates with `roles` schema example
|
||||||
|
- `design/loops/technical.md`: document `{model}` substitution
|
||||||
|
- `tests/test_model_divergence.py`: loop model binding tests, `{model}` substitution tests, check-gate halt tests
|
||||||
|
**Not touched:** interactive `--transition` / `--claim` (already done in subtask-2), dashboard badges (already done in subtask-2).
|
||||||
|
**Acceptance:**
|
||||||
|
1. `loop.json` with `roles.implementer.model: "llama-3.3-70b"` → `_invoke_harness` substitutes `{model}` in harness command.
|
||||||
|
2. `loop.json` without per-role `model` → defaults to `models.json` `default` model.
|
||||||
|
3. `--check-gate` in multi-LLM mode halts loop if loop-verify model = loop-implement model.
|
||||||
|
4. `--check-gate` in single-LLM mode does NOT halt (advisory only).
|
||||||
|
5. `pytest tests/test_model_divergence.py -v` green.
|
||||||
|
6. `pytest tests/ -q` green (no regressions).
|
||||||
|
**Peak context estimate:** ~6k tokens (loop-runner + check-gate + tests).
|
||||||
|
**Run order:** third, AFTER subtask-2.
|
||||||
|
|
||||||
|
## Dependency graph
|
||||||
|
|
||||||
|
```
|
||||||
|
subtask-1 (manifest+detection)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
subtask-2 (interactive enforcement+audit)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
subtask-3 (loop enforcement+dashboard)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
parent model-divergence-enforcement → complete
|
||||||
|
```
|
||||||
|
|
||||||
|
Parent is complete only when ALL three subtasks pass their acceptance criteria AND
|
||||||
|
`pytest tests/ -q` is green.
|
||||||
|
|
||||||
|
## Parent non-goals
|
||||||
|
|
||||||
|
- No rule agents (FW-2, FW-3) — separate backlog items that *consume* this feature.
|
||||||
|
- No agent tab redesign (FW-1) — separate backlog item, no dependency on this task.
|
||||||
|
- No model capability inspection — framework never inspects capability/size/provider (decision D8).
|
||||||
|
- No `detect_models.py` auto-writing `models.json` — detection is advisory; user confirms the manifest.
|
||||||
|
|
||||||
|
## Fallback
|
||||||
|
|
||||||
|
If any subtask hits a blocker (e.g., `status.py` surgery is too large for context budget),
|
||||||
|
it must report back to the Orchestrator via `human_intervention` rather than skipping
|
||||||
|
enforcement logic. Partial enforcement is worse than no enforcement — it creates a
|
||||||
|
false sense of security.
|
||||||
@@ -0,0 +1,81 @@
|
|||||||
|
# SPEC — model-divergence-enforcement
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The framework has no computational enforcement of model divergence for conflict-of-interest roles. `status.py:1432-1436` enforces that the *session* playing reviewer differs from the implementer (via `.state.implementer`), but there is no enforcement that the *model* playing bug-finder differs from adversarial-bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping.
|
||||||
|
|
||||||
|
## Design (locked)
|
||||||
|
|
||||||
|
Full design is in `design/framework/functional.md` §4 and `design/framework/technical.md` §§2-5,13. This SPEC references those docs and does not repeat them.
|
||||||
|
|
||||||
|
### Key decisions (from design session 2026-06-25)
|
||||||
|
|
||||||
|
- **Manifest**: `models.json` with `{default, advised, models:[{name, provider, context_window, location}]}`.
|
||||||
|
- **Mode detection**: 0-1 models → single-LLM (advisory once, then silent). 2+ → multi-LLM (hard block). Missing file → single-LLM (backward compatible).
|
||||||
|
- **Conflict matrix (locked)**: `code_review≠implement`; `bug_find≠implement`; `adversarial_bug_find≠implement+bug_find`; `referee≠implement+bug_find+adversarial_bug_find`; `loop-verify≠loop-implement`.
|
||||||
|
- **Auto-assignment**: default model → next-available on conflict → user override via `--transition --model` or `loop.json roles.<role>.model`. Refuse only if no non-conflicting model exists.
|
||||||
|
- **State**: `.state.models` per task recording which model filled which role.
|
||||||
|
- **Loop integration**: `loop.json` per-role `model` field + `{model}` substitution in `_invoke_harness`.
|
||||||
|
- **Audit**: new `model_divergence` category in `--audit`.
|
||||||
|
- **Dashboard**: model badges on task cards.
|
||||||
|
- **Detection**: `scripts/detect_models.py` probes opencode.json + localhost endpoints (8080/11434/1234/8000).
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This is a **parent task**. It decomposes into 3 sequential subtasks:
|
||||||
|
|
||||||
|
### Subtask 1: manifest+detection
|
||||||
|
- `models.json` schema + loader
|
||||||
|
- `scripts/detect_models.py` (probe opencode.json + localhost endpoints)
|
||||||
|
- `install.sh` / `update.sh` / `upgrade.sh` integration (run detect_models after VRAM detection)
|
||||||
|
- `config.md` `## Available Models` section
|
||||||
|
- `prompts/onboarding.md` Step 2d (model config)
|
||||||
|
- Tests: `test_model_divergence.py` (manifest loading, single vs multi mode detection, missing file backward compat)
|
||||||
|
|
||||||
|
### Subtask 2: interactive enforcement+audit
|
||||||
|
- `.state.models` schema + writer
|
||||||
|
- `status.py --transition --model <name>` (records model, checks conflict matrix in multi-LLM mode)
|
||||||
|
- `status.py --claim --model <name>` (refuses on conflict)
|
||||||
|
- `CONFLICT_MATRIX` constant + `_check_conflict` helper in `status.py`
|
||||||
|
- `status.py --audit` `model_divergence` category
|
||||||
|
- Dashboard model badges on task cards
|
||||||
|
- Tests: `test_model_divergence.py` (conflict matrix, auto-assign, audit category, dashboard badges)
|
||||||
|
|
||||||
|
### Subtask 3: loop enforcement+dashboard
|
||||||
|
- `loop.json` per-role `model` field (optional, defaults to `models.json default`)
|
||||||
|
- `loop-runner.py _invoke_harness` `{model}` substitution
|
||||||
|
- `status.py --check-gate` model-divergence check (loop-verify ≠ loop-implement in multi-LLM mode)
|
||||||
|
- Loop dashboard badges (model per role)
|
||||||
|
- Tests: `test_model_divergence.py` (loop model binding, {model} substitution, check-gate halt)
|
||||||
|
|
||||||
|
## Dependencies
|
||||||
|
|
||||||
|
- Subtask 1 → no deps (foundational)
|
||||||
|
- Subtask 2 → depends on subtask 1 (needs manifest loader)
|
||||||
|
- Subtask 3 → depends on subtask 2 (needs .state.models + conflict matrix)
|
||||||
|
|
||||||
|
## Out of scope
|
||||||
|
|
||||||
|
- Rule agents (FW-2, FW-3) — separate backlog items that *consume* this feature
|
||||||
|
- Agent tab redesign (FW-1) — separate backlog item, no dependency on this task
|
||||||
|
- Model capability inspection (D8: framework never inspects capability/size/provider)
|
||||||
|
- `detect_models.py` auto-writing `models.json` without user confirmation (detection is advisory; user confirms the manifest)
|
||||||
|
|
||||||
|
## Success criteria
|
||||||
|
|
||||||
|
1. `models.json` missing → single-LLM mode, all model commands are no-ops, existing behavior unchanged.
|
||||||
|
2. `models.json` with 1 model → single-LLM mode, advisory printed once if `advised: true`, then silent.
|
||||||
|
3. `models.json` with 2+ models → multi-LLM mode, conflict matrix enforced on `--transition --model` and `--claim --model`.
|
||||||
|
4. `--audit` flags conflict-matrix violations as `model_divergence` category.
|
||||||
|
5. `loop.json` per-role model binding works; `--check-gate` halts on loop-verify = loop-implement in multi-LLM mode.
|
||||||
|
6. `detect_models.py` probes opencode.json + localhost endpoints and emits a candidate manifest.
|
||||||
|
7. `pytest tests/ -v` green; no regressions.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `design/framework/functional.md` §4 (Model-Divergence Enforcement)
|
||||||
|
- `design/framework/technical.md` §§2-5 (manifest, state, flags, conflict matrix), §13 (loop integration)
|
||||||
|
- `design/framework/README.md` (locked decisions F4, F5)
|
||||||
|
- `.agent.md` Agent Configuration (6 phase roles)
|
||||||
|
- `scripts/status.py:1432-1436` (existing session-level reviewer≠implementer enforcement)
|
||||||
|
- `scripts/loop-runner.py:366-404` (`_invoke_harness`, needs `{model}` substitution)
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
new
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
# BRIEF — mde-interactive-enforcement
|
||||||
|
|
||||||
|
**Parent:** model-divergence-enforcement
|
||||||
|
**Subtask:** 2 of 3 (depends on mde-manifest-detection)
|
||||||
|
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 2
|
||||||
|
**Design ref:** `design/framework/technical.md` §§4-5, §12
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
|
||||||
|
Add the conflict matrix, `--transition --model` / `--claim --model` enforcement,
|
||||||
|
`.state.models` tracking, audit category, and dashboard model badges. This is the
|
||||||
|
core enforcement layer for interactive (non-loop) workflows.
|
||||||
|
|
||||||
|
## Deliverables
|
||||||
|
|
||||||
|
1. `scripts/status.py`:
|
||||||
|
- `CONFLICT_MATRIX` constant (locked matrix from SPEC)
|
||||||
|
- `_check_conflict(current_phase, current_model, next_phase, next_model)` helper
|
||||||
|
- `cmd_transition`: `--model` arg + conflict check in multi-LLM mode
|
||||||
|
- `cmd_claim`: `--model` arg + conflict check
|
||||||
|
- `cmd_audit`: `model_divergence` category
|
||||||
|
- `.state.models` writer (JSON: `{implementer, code_reviewer, bug_hunter, ...}`)
|
||||||
|
2. `automaton/dashboard/html/dashboard.js` — model badge on task cards
|
||||||
|
3. `automaton/dashboard/ui/app.py` — `.state.models` in `/api/tasks` response
|
||||||
|
4. `tests/test_model_divergence.py` — conflict matrix, auto-assign, audit, badges
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
|
||||||
|
- Single-LLM mode: `--model` records but never refuses (advisory once if `advised: true`)
|
||||||
|
- Multi-LLM mode: `--transition --model` refuses on conflict (e.g., same model for implement + code_review)
|
||||||
|
- Multi-LLM mode: `--claim --model` refuses on conflict
|
||||||
|
- `--audit` flags `model_divergence` violations
|
||||||
|
- Dashboard shows model badges when `.state.models` exists
|
||||||
|
- `pytest tests/ -q` green
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
new
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# BRIEF — mde-loop-enforcement
|
||||||
|
|
||||||
|
**Parent:** model-divergence-enforcement
|
||||||
|
**Subtask:** 3 of 3 (depends on mde-interactive-enforcement)
|
||||||
|
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 3
|
||||||
|
**Design ref:** `design/framework/technical.md` §§10, §13
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
|
||||||
|
Add loop-level model-divergence enforcement: `loop.json` per-role `model` field,
|
||||||
|
`{model}` substitution in `_invoke_harness`, and `--check-gate` model-divergence
|
||||||
|
brake gate.
|
||||||
|
|
||||||
|
## Deliverables
|
||||||
|
|
||||||
|
1. `scripts/loop-runner.py`:
|
||||||
|
- `_invoke_harness` (line 366-404): `{model}` placeholder substitution from `loop.json` role config
|
||||||
|
2. `scripts/status.py`:
|
||||||
|
- `--check-gate`: model-divergence brake gate (loop-verify ≠ loop-implement in multi-LLM mode)
|
||||||
|
- `cmd_create_loop` / `cmd_install_schedule`: validate `loop.json` per-role `model` fields
|
||||||
|
3. `templates/loops/`: update loop templates with `roles` schema example
|
||||||
|
4. `design/loops/technical.md`: document `{model}` substitution
|
||||||
|
5. `tests/test_model_divergence.py`: loop model binding, `{model}` substitution, check-gate halt
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
|
||||||
|
- `loop.json` with `roles.implementer.model` → `_invoke_harness` substitutes `{model}` in harness command
|
||||||
|
- `loop.json` without per-role `model` → defaults to `models.json` `default` model
|
||||||
|
- `--check-gate` in multi-LLM mode halts if loop-verify model = loop-implement model
|
||||||
|
- `--check-gate` in single-LLM mode does NOT halt (advisory only)
|
||||||
|
- `pytest tests/ -q` green
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
new
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
# BRIEF — mde-manifest-detection
|
||||||
|
|
||||||
|
**Parent:** model-divergence-enforcement
|
||||||
|
**Subtask:** 1 of 3 (foundational, no deps)
|
||||||
|
**Spec ref:** `tasks/model-divergence-enforcement/SPEC.md` §Subtask 1
|
||||||
|
**Design ref:** `design/framework/technical.md` §§2-3
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
|
||||||
|
Create the model manifest system (`models.json` schema + loader) and model detection
|
||||||
|
script (`scripts/detect_models.py`). This is the foundation for all model-divergence
|
||||||
|
enforcement — subtasks 2 and 3 depend on it.
|
||||||
|
|
||||||
|
## Deliverables
|
||||||
|
|
||||||
|
1. `scripts/detect_models.py` — probes opencode.json providers + localhost endpoints
|
||||||
|
(8080/11434/1234/8000), emits candidate manifest as JSON to stdout
|
||||||
|
2. `scripts/status.py` — add `_load_models_manifest()` and `_get_mode()` helpers
|
||||||
|
3. `scripts/install.sh`, `update.sh`, `upgrade.sh` — call detect_models after vram_detect
|
||||||
|
4. `config.md` — `## Available Models` section template
|
||||||
|
5. `prompts/onboarding.md` — Step 2d model config check
|
||||||
|
6. `tests/test_model_divergence.py` — manifest loading, single vs multi mode, missing file
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
|
||||||
|
- `models.json` missing → `_get_mode()` returns `"single"`, all model commands are no-ops
|
||||||
|
- `models.json` with 0-1 models → `"single"` mode
|
||||||
|
- `models.json` with 2+ models → `"multi"` mode
|
||||||
|
- `detect_models.py` probes localhost and prints candidate JSON
|
||||||
|
- `pytest tests/ -q` green
|
||||||
@@ -0,0 +1,316 @@
|
|||||||
|
"""Tests for scripts/autopilot.py — functional correctness review.
|
||||||
|
|
||||||
|
Covers bugs found during 2026-06-25 review:
|
||||||
|
1. PHASE_PRIORITY missing code_review
|
||||||
|
2. cmd_drive suggested implement→bug_find (illegal; should be implement→code_review)
|
||||||
|
3. cmd_drive had no handler for code_review / code_review:approved
|
||||||
|
4. _all_tasks didn't skip tasks/complete/ archive dir
|
||||||
|
5. _all_tasks didn't recurse into subtasks/
|
||||||
|
6. _all_tasks used wrong path for non-framework projects
|
||||||
|
7. is_terminal treated human_intervention as terminal (dead-code handler)
|
||||||
|
8. needs_user_input didn't flag human_intervention as blocked
|
||||||
|
"""
|
||||||
|
|
||||||
|
import importlib.util
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Optional
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
_AP_PATH = Path.home() / ".automaton" / "scripts" / "autopilot.py"
|
||||||
|
_spec = importlib.util.spec_from_file_location("autopilot_mod", _AP_PATH)
|
||||||
|
ap = importlib.util.module_from_spec(_spec)
|
||||||
|
_spec.loader.exec_module(ap)
|
||||||
|
|
||||||
|
|
||||||
|
def _make_task(project: Path, name: str, phase: str = "new") -> Path:
|
||||||
|
"""Create a task dir + .state (no subprocess)."""
|
||||||
|
if project == ap.AUTOMATON_DIR:
|
||||||
|
tp = project / "tasks" / name
|
||||||
|
else:
|
||||||
|
tp = project / ".automaton" / "tasks" / name
|
||||||
|
tp.mkdir(parents=True, exist_ok=True)
|
||||||
|
(tp / ".state").write_text(phase + "\n")
|
||||||
|
return tp
|
||||||
|
|
||||||
|
|
||||||
|
def _make_subtask(project: Path, parent: str, name: str, phase: str = "new") -> Path:
|
||||||
|
if project == ap.AUTOMATON_DIR:
|
||||||
|
sp = project / "tasks" / parent / "subtasks" / name
|
||||||
|
else:
|
||||||
|
sp = project / ".automaton" / "tasks" / parent / "subtasks" / name
|
||||||
|
sp.mkdir(parents=True, exist_ok=True)
|
||||||
|
(sp / ".state").write_text(phase + "\n")
|
||||||
|
return sp
|
||||||
|
|
||||||
|
|
||||||
|
def _write_verdict(task_path: Path, first_line: str) -> None:
|
||||||
|
(task_path / "VERDICT.md").write_text(first_line + "\nrest of verdict\n")
|
||||||
|
|
||||||
|
|
||||||
|
class _Args:
|
||||||
|
"""Minimal args object for cmd_* functions."""
|
||||||
|
def __init__(self, **kwargs):
|
||||||
|
self.project = kwargs.get("project")
|
||||||
|
self.max_iterations = kwargs.get("max_iterations")
|
||||||
|
self.delay = kwargs.get("delay")
|
||||||
|
self.threshold = kwargs.get("threshold")
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def tmp_project(tmp_path):
|
||||||
|
(tmp_path / ".automaton" / "tasks").mkdir(parents=True)
|
||||||
|
return tmp_path
|
||||||
|
|
||||||
|
|
||||||
|
class TestPhasePriority:
|
||||||
|
"""Bug 1: code_review was missing from PHASE_PRIORITY."""
|
||||||
|
|
||||||
|
def test_code_review_in_priority(self):
|
||||||
|
assert "code_review" in ap.PHASE_PRIORITY
|
||||||
|
|
||||||
|
def test_code_review_priority_between_implement_and_bug_find(self):
|
||||||
|
assert ap.PHASE_PRIORITY["code_review"] > ap.PHASE_PRIORITY["implement"]
|
||||||
|
assert ap.PHASE_PRIORITY["code_review"] < ap.PHASE_PRIORITY["bug_find"]
|
||||||
|
|
||||||
|
def test_human_intervention_in_priority(self):
|
||||||
|
assert "human_intervention" in ap.PHASE_PRIORITY
|
||||||
|
|
||||||
|
def test_priorities_monotonically_increase_along_pipeline(self):
|
||||||
|
pipeline = [
|
||||||
|
"new", "research", "decomposition", "design", "test_design",
|
||||||
|
"implement", "code_review", "bug_find", "adversarial_bug_find",
|
||||||
|
"doc_review", "referee",
|
||||||
|
]
|
||||||
|
for i in range(len(pipeline) - 1):
|
||||||
|
assert ap.PHASE_PRIORITY[pipeline[i]] < ap.PHASE_PRIORITY[pipeline[i + 1]], \
|
||||||
|
f"{pipeline[i]} ({ap.PHASE_PRIORITY[pipeline[i]]}) should be < {pipeline[i+1]} ({ap.PHASE_PRIORITY[pipeline[i+1]]})"
|
||||||
|
|
||||||
|
|
||||||
|
class TestCmdDriveCodeReview:
|
||||||
|
"""Bugs 2+3: implement→bug_find was illegal; code_review handlers missing."""
|
||||||
|
|
||||||
|
def test_implement_suggests_code_review_not_bug_find(self, tmp_project, capsys):
|
||||||
|
_make_task(tmp_project, "task-a", "implement")
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
ap.cmd_drive(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert "--transition code_review" in out
|
||||||
|
assert "--transition bug_find" not in out
|
||||||
|
|
||||||
|
def test_code_review_suggests_awaiting_approval(self, tmp_project, capsys):
|
||||||
|
_make_task(tmp_project, "task-cr", "code_review")
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
ap.cmd_drive(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert "--transition code_review:awaiting_approval" in out
|
||||||
|
|
||||||
|
def test_code_review_approved_suggests_bug_find(self, tmp_project, capsys):
|
||||||
|
_make_task(tmp_project, "task-cr2", "code_review:approved")
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
ap.cmd_drive(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert "--transition bug_find" in out
|
||||||
|
|
||||||
|
def test_implement_task_not_stalled(self, tmp_project, capsys):
|
||||||
|
"""Ensure cmd_drive produces a transition command for implement phase."""
|
||||||
|
_make_task(tmp_project, "task-impl", "implement")
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
rc = ap.cmd_drive(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert rc == 0
|
||||||
|
assert "status.py --transition" in out
|
||||||
|
|
||||||
|
|
||||||
|
class TestAllTasksSkipComplete:
|
||||||
|
"""Bug 4: _all_tasks included tasks/complete/ as a task."""
|
||||||
|
|
||||||
|
def test_complete_dir_excluded(self, tmp_project):
|
||||||
|
_make_task(tmp_project, "active-task", "implement")
|
||||||
|
complete_dir = tmp_project / ".automaton" / "tasks" / "complete"
|
||||||
|
complete_dir.mkdir(parents=True)
|
||||||
|
(complete_dir / "old-task").mkdir()
|
||||||
|
(complete_dir / "old-task" / ".state").write_text("complete\n")
|
||||||
|
|
||||||
|
tasks = ap._all_tasks(tmp_project)
|
||||||
|
names = [t[0] for t in tasks]
|
||||||
|
assert "active-task" in names
|
||||||
|
assert "complete" not in names
|
||||||
|
|
||||||
|
|
||||||
|
class TestAllTasksSubtasks:
|
||||||
|
"""Bug 5: _all_tasks didn't recurse into subtasks/."""
|
||||||
|
|
||||||
|
def test_subtask_found(self, tmp_project):
|
||||||
|
_make_task(tmp_project, "parent-task", "decomposition:approved")
|
||||||
|
_make_subtask(tmp_project, "parent-task", "sub-a", "implement")
|
||||||
|
_make_subtask(tmp_project, "parent-task", "sub-b", "new")
|
||||||
|
|
||||||
|
tasks = ap._all_tasks(tmp_project)
|
||||||
|
names = [t[0] for t in tasks]
|
||||||
|
assert "parent-task" in names
|
||||||
|
assert "parent-task/sub-a" in names
|
||||||
|
assert "parent-task/sub-b" in names
|
||||||
|
|
||||||
|
def test_subtask_scan_all_tasks(self, tmp_project):
|
||||||
|
_make_task(tmp_project, "parent-task", "decomposition:approved")
|
||||||
|
_make_subtask(tmp_project, "parent-task", "sub-a", "implement")
|
||||||
|
|
||||||
|
scanned = ap.scan_all_tasks(tmp_project)
|
||||||
|
names = [t["name"] for t in scanned]
|
||||||
|
assert "parent-task" in names
|
||||||
|
assert "parent-task/sub-a" in names
|
||||||
|
|
||||||
|
|
||||||
|
class TestAllTasksProjectPath:
|
||||||
|
"""Bug 6: _all_tasks used project_dir/tasks instead of project_dir/.automaton/tasks."""
|
||||||
|
|
||||||
|
def test_non_framework_project_path(self, tmp_project):
|
||||||
|
_make_task(tmp_project, "proj-task", "implement")
|
||||||
|
tasks = ap._all_tasks(tmp_project)
|
||||||
|
names = [t[0] for t in tasks]
|
||||||
|
assert "proj-task" in names
|
||||||
|
|
||||||
|
def test_wrong_path_returns_empty(self, tmp_path):
|
||||||
|
"""If .automaton/tasks doesn't exist, return empty (not crash)."""
|
||||||
|
tasks = ap._all_tasks(tmp_path)
|
||||||
|
assert tasks == []
|
||||||
|
|
||||||
|
|
||||||
|
class TestIsTerminal:
|
||||||
|
"""Bug 7: human_intervention was treated as terminal."""
|
||||||
|
|
||||||
|
def test_complete_is_terminal(self):
|
||||||
|
assert ap.is_terminal({"phase": "complete"}) is True
|
||||||
|
|
||||||
|
def test_human_intervention_not_terminal(self):
|
||||||
|
assert ap.is_terminal({"phase": "human_intervention"}) is False
|
||||||
|
|
||||||
|
def test_none_phase_not_terminal(self):
|
||||||
|
assert ap.is_terminal({"phase": None}) is False
|
||||||
|
|
||||||
|
def test_new_not_terminal(self):
|
||||||
|
assert ap.is_terminal({"phase": "new"}) is False
|
||||||
|
|
||||||
|
|
||||||
|
class TestNeedsUserInput:
|
||||||
|
"""Bug 8: human_intervention wasn't flagged as needing user input."""
|
||||||
|
|
||||||
|
def test_awaiting_approval_needs_input(self, tmp_project):
|
||||||
|
tp = _make_task(tmp_project, "task-r", "research:awaiting_approval")
|
||||||
|
assert ap.needs_user_input({"phase": "research:awaiting_approval", "path": str(tp)}) is True
|
||||||
|
|
||||||
|
def test_human_intervention_needs_input(self, tmp_project):
|
||||||
|
tp = _make_task(tmp_project, "task-hi", "human_intervention")
|
||||||
|
assert ap.needs_user_input({"phase": "human_intervention", "path": str(tp)}) is True
|
||||||
|
|
||||||
|
def test_implement_no_input(self, tmp_project):
|
||||||
|
tp = _make_task(tmp_project, "task-i", "implement")
|
||||||
|
assert ap.needs_user_input({"phase": "implement", "path": str(tp)}) is False
|
||||||
|
|
||||||
|
def test_verdict_fail_needs_input(self, tmp_project):
|
||||||
|
tp = _make_task(tmp_project, "task-v", "referee")
|
||||||
|
_write_verdict(tp, "FAIL: bugs found")
|
||||||
|
assert ap.needs_user_input({"phase": "referee", "path": str(tp)}) is True
|
||||||
|
|
||||||
|
def test_verdict_pass_no_input(self, tmp_project):
|
||||||
|
tp = _make_task(tmp_project, "task-vp", "referee")
|
||||||
|
_write_verdict(tp, "PASS: all good")
|
||||||
|
assert ap.needs_user_input({"phase": "referee", "path": str(tp)}) is False
|
||||||
|
|
||||||
|
|
||||||
|
class TestSortByAdvancement:
|
||||||
|
"""Verify sort orders code_review correctly after the fix."""
|
||||||
|
|
||||||
|
def test_code_review_more_advanced_than_implement(self):
|
||||||
|
tasks = [
|
||||||
|
{"name": "impl-task", "base_phase": "implement"},
|
||||||
|
{"name": "cr-task", "base_phase": "code_review"},
|
||||||
|
]
|
||||||
|
sorted_tasks = ap.sort_by_advancement(tasks)
|
||||||
|
assert sorted_tasks[0]["name"] == "cr-task"
|
||||||
|
|
||||||
|
def test_bug_find_more_advanced_than_code_review(self):
|
||||||
|
tasks = [
|
||||||
|
{"name": "cr-task", "base_phase": "code_review"},
|
||||||
|
{"name": "bf-task", "base_phase": "bug_find"},
|
||||||
|
]
|
||||||
|
sorted_tasks = ap.sort_by_advancement(tasks)
|
||||||
|
assert sorted_tasks[0]["name"] == "bf-task"
|
||||||
|
|
||||||
|
|
||||||
|
class TestCmdDriveHumanIntervention:
|
||||||
|
"""Verify human_intervention is now drivable (not dead code)."""
|
||||||
|
|
||||||
|
def test_human_intervention_shows_in_blocked(self, tmp_project, capsys):
|
||||||
|
_make_task(tmp_project, "task-hi", "human_intervention")
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
ap.cmd_drive(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert "ORCHESTRATION_BLOCKED" in out
|
||||||
|
assert "task-hi" in out
|
||||||
|
assert "--transition referee" in out
|
||||||
|
|
||||||
|
def test_human_intervention_not_in_unblocked(self, tmp_project, capsys):
|
||||||
|
"""human_intervention should be blocked, not selected as NEXT_TASK."""
|
||||||
|
_make_task(tmp_project, "task-hi", "human_intervention")
|
||||||
|
_make_task(tmp_project, "task-active", "implement")
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
ap.cmd_drive(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert "NEXT_TASK: task-active" in out
|
||||||
|
assert "NEXT_TASK: task-hi" not in out
|
||||||
|
|
||||||
|
|
||||||
|
class TestCmdDriveSummary:
|
||||||
|
"""Integration: cmd_summary with mixed states."""
|
||||||
|
|
||||||
|
def test_summary_counts(self, tmp_project, capsys):
|
||||||
|
_make_task(tmp_project, "t-new", "new")
|
||||||
|
_make_task(tmp_project, "t-impl", "implement")
|
||||||
|
_make_task(tmp_project, "t-blocked", "research:awaiting_approval")
|
||||||
|
_make_task(tmp_project, "t-hi", "human_intervention")
|
||||||
|
complete_dir = tmp_project / ".automaton" / "tasks" / "complete"
|
||||||
|
complete_dir.mkdir()
|
||||||
|
(complete_dir / "t-done").mkdir()
|
||||||
|
(complete_dir / "t-done" / ".state").write_text("complete\n")
|
||||||
|
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
ap.cmd_summary(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert "3" in out # 3 non-terminal (new, implement, blocked, hi... wait, hi is now non-terminal too)
|
||||||
|
assert "READY TO DRIVE" in out
|
||||||
|
assert "AWAITING USER" in out
|
||||||
|
assert "t-done" not in out # complete dir should not appear
|
||||||
|
assert "complete" not in out.split("READY")[0] # no "complete" as a task name
|
||||||
|
|
||||||
|
|
||||||
|
class TestCmdDriveAllPhases:
|
||||||
|
"""Every phase produces a transition command — no silent stalls."""
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("phase,expected_fragment", [
|
||||||
|
("new", "--transition research"),
|
||||||
|
("research", "--transition research:awaiting_approval"),
|
||||||
|
("research:approved", "--transition decomposition"),
|
||||||
|
("decomposition", "--transition decomposition:awaiting_approval"),
|
||||||
|
("decomposition:approved", "--transition complete"),
|
||||||
|
("design", "--transition design:awaiting_approval"),
|
||||||
|
("design:approved", "--transition test_design"),
|
||||||
|
("test_design", "--transition test_design:awaiting_approval"),
|
||||||
|
("test_design:approved", "--transition implement"),
|
||||||
|
("implement", "--transition code_review"),
|
||||||
|
("code_review", "--transition code_review:awaiting_approval"),
|
||||||
|
("code_review:approved", "--transition bug_find"),
|
||||||
|
("bug_find", "--transition adversarial_bug_find"),
|
||||||
|
("adversarial_bug_find", "--transition doc_review"),
|
||||||
|
("doc_review", "--transition referee"),
|
||||||
|
("referee", "--transition complete"),
|
||||||
|
])
|
||||||
|
def test_phase_produces_transition(self, tmp_project, capsys, phase, expected_fragment):
|
||||||
|
_make_task(tmp_project, "test-task", phase)
|
||||||
|
args = _Args(project=str(tmp_project))
|
||||||
|
rc = ap.cmd_drive(args)
|
||||||
|
out = capsys.readouterr().out
|
||||||
|
assert rc == 0
|
||||||
|
assert expected_fragment in out, f"Phase '{phase}' should suggest '{expected_fragment}', got:\n{out}"
|
||||||
Reference in New Issue
Block a user