From 4b3d92c9a2f9e34524e1dd72ee038c1a381d52e3 Mon Sep 17 00:00:00 2001 From: Lap Tran Date: Thu, 25 Jun 2026 07:11:35 -0400 Subject: [PATCH] Add framework design docs (model-divergence, rule agents, agent tab) Design docs in design/framework/ covering: - Model-divergence enforcement (conflict matrix, modes, auto-assignment) - Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md) - Rule Reviewer agent (monthly consolidation, different LLM than Proposer) - Agent tab redesign (phase roles + scheduled jobs, remove fake types) - Schedules, conflict-of-interest, success criteria Cross-references updated in AGENTS.md, README.md, CHANGELOG.md, .onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md, memory/v1-1-hardening-session.md. Also restores scripts/automaton-cleanup.sh stub (was corrupted by pytest test leak writing temp path into real stub). --- .onboarding.md | 9 + AGENTS.md | 3 + CHANGELOG.md | 9 + README.md | 1 + design/framework/BACKLOG.md | 33 ++ design/framework/README.md | 67 ++++ design/framework/functional.md | 335 +++++++++++++++++++ design/framework/technical.md | 536 +++++++++++++++++++++++++++++++ design/loops/BACKLOG.md | 2 + design/loops/README.md | 1 + memory/v1-1-hardening-session.md | 13 +- prompts/onboarding.md | 7 + scripts/automaton-cleanup.sh | 2 +- 13 files changed, 1016 insertions(+), 2 deletions(-) create mode 100644 design/framework/BACKLOG.md create mode 100644 design/framework/README.md create mode 100644 design/framework/functional.md create mode 100644 design/framework/technical.md diff --git a/.onboarding.md b/.onboarding.md index 0561a55..17ee04b 100644 --- a/.onboarding.md +++ b/.onboarding.md @@ -111,3 +111,12 @@ When the agent receives a trigger command, it must: 1. Read the corresponding template file. 2. Replace all `{placeholders}` with the actual project values. 3. Execute the rendered prompt. + +## Backlog + +Design backlogs are the framework's outstanding-work store when no active tasks exist. Each design area has its own `BACKLOG.md`: + +- `~/.automaton/design/loops/BACKLOG.md` — loop engineering v1.1+ and deferred items. +- `~/.automaton/design/framework/BACKLOG.md` — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). + +The self-improvement loop consumes these automatically when configured with `work_source.kind = "backlog"` and `work_source.area` set to the relevant design area (`"loops"` or `"framework"`). For manual work, read the topmost `- [ ]` item and create a task via `status.py --create-task`. diff --git a/AGENTS.md b/AGENTS.md index 3fa5743..70ccf66 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -47,6 +47,9 @@ Automaton is a **contract-based, state-enforced workflow framework** for LLM age │ ├── config.py │ ├── core/ │ └── ui/ +├── design/ # Design docs for major features +│ ├── loops/ # Loop engineering system (v1 locked) +│ └── framework/ # Framework agent features (rule agents, Agent tab, model-divergence) ├── tests/ # pytest suite └── tasks/ # Framework development tasks ``` diff --git a/CHANGELOG.md b/CHANGELOG.md index 9db85ca..0d635d2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,15 @@ ## [unreleased] +### Added — framework agent features design docs + +- **New `design/framework/`** directory: design index, functional design, technical design, and backlog for three framework-level agent features: + - **Model-Divergence Enforcement** — `models.json` manifest, single vs multi-LLM mode detection, conflict matrix (`code_review≠implement`, `bug_find≠implement`, `adversarial_bug_find≠implement+bug_find`, `referee≠implement+bug_find+adversarial_bug_find`, `loop-verify≠loop-implement`), `--transition --model`, `.state.models`, `loop.json` per-role model binding + `{model}` substitution, `--audit` model_divergence category, dashboard badges. Shipped as a manual task (3 subtasks), not a backlog item. + - **Rule Agents** — Rule Proposer (daily scheduled standalone agent, scans completed tasks' `BUG_REPORT.md`/`ADVERSARIAL_BUG_REPORT.md`/`VERDICT.md`, proposes rules to `RULE_PROPOSALS.md`) and Rule Reviewer (monthly scheduled standalone agent, consolidates `.rules.md` → `RULE_REVIEW.md`). Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers. Conflict-of-interest: Reviewer must differ from Proposer; both must differ from tasks they review. Enforcement deferred until model-divergence ships. + - **Agent Tab Redesign** — replace 4 fake `AGENT_TYPE_META` types with two sections: Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) + Scheduled Jobs (real `job.kind`: cleanup, loop, rule-scan, rule-review). New `/api/phase-roles` endpoint. +- **`design/framework/BACKLOG.md`**: 3 v1 items (`agent-tab-real-roles`, `rule-proposer-agent`, `rule-reviewer-agent`) + 3 deferred items. Consumable by the self-improvement loop via `work_source.area = "framework"`. +- **Cross-references**: `design/loops/README.md` and `design/loops/BACKLOG.md` now point to the sibling `design/framework/` area. `AGENTS.md` repo layout updated. + ### Added — cross-loop task claim (task `add-claim-loop-task`) - **`scripts/status.py`**: New `--claim-loop-task --task [--project P]` command. Exit 0 = claimed (or already self-claimed, idempotent). Exit 2 = already claimed by another running/paused loop (`task_already_claimed:{other}` on stderr) or untracked loop (`loop_untracked`). Uses `_loop_lock` to serialize writes; cross-loop scan is advisory (self-healing on next tick). diff --git a/README.md b/README.md index d740b04..317c2a9 100644 --- a/README.md +++ b/README.md @@ -250,6 +250,7 @@ The runner resolves prompt files from `loop.json` `roles.*.prompt` (e.g. `loop-i | `blast_radius.file_scope` | List of paths the loop may edit | | `blast_radius.use_worktree` | If true, tick runs in a per-loop git worktree | | `work_source.kind` | `single`, `audit`, or `backlog` | +| `work_source.area` | Design area for `backlog` kind (default `"loops"`; `"framework"` reads `design/framework/BACKLOG.md`) | | `roles.implement.prompt` | Prompt file for Implement role | | `roles.verify.prompt` | Prompt file for Verify role | | `roles.orchestrate.prompt` | Prompt file for Orchestrate role | diff --git a/design/framework/BACKLOG.md b/design/framework/BACKLOG.md new file mode 100644 index 0000000..402bf40 --- /dev/null +++ b/design/framework/BACKLOG.md @@ -0,0 +1,33 @@ +# Framework Backlog + +Status: **v1 draft 2026-06-25**. These items are captured for the self-improvement loop (or manual pickup) once the self-improvement loop is running and the design docs are in place. + +This backlog mirrors the pattern of `design/loops/BACKLOG.md` but for framework-level agent features. The self-improvement loop's `work_source.area` can be set to `"framework"` to pull from here. + +--- + +## v1 (Priority) + +| ID | Item | Description | Dependencies | +|---|---|---|---| +| FW-1 | **agent-tab-real-roles** | Redesign the dashboard Agent tab: replace 4 fake `AGENT_TYPE_META` types with two sections — (1) Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) showing active/inactive based on current task phases; (2) Scheduled Jobs (cleanup, loop-tick, rule-scan, rule-review) using real `job.kind` from `/api/scheduled`. New `/api/phase-roles` endpoint. | None | +| FW-2 | **rule-proposer-agent** | Daily scheduled standalone agent that scans completed tasks since last run, reads their `BUG_REPORT.md` / `ADVERSARIAL_BUG_REPORT.md` / `VERDICT.md`, proposes new rules with concrete examples to `RULE_PROPOSALS.md`. Uses direct harness invocation (reuses `loop-runner._invoke_harness`). State in `.state.rule-scan`. Requires different LLM than the tasks it reviews (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding) | +| FW-3 | **rule-reviewer-agent** | Monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, missing examples — writes `RULE_REVIEW.md`. Uses direct harness invocation. State in `.state.rule-review`. Must use different LLM than Rule Proposer (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding), `rule-proposer-agent` | + +--- + +## Deferred / Tier 2 (Post-v1) + +| ID | Item | Description | Notes | +|---|---|---|---| +| FW-4 | **model-divergence-enforcement** | Manifest (`models.json`), mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model + `{model}` substitution, audit category, dashboard badges. This is a **manual task**, not a backlog item — tracked separately. | See separate task `model-divergence-enforcement` | +| FW-5 | **dashboard-agent-tab-v2** | Enhance Agent tab with: role details (click → task list), schedule management (enable/disable from UI), last-run timestamps, per-agent logs. | Requires FW-1 | +| FW-6 | **rule-enforcement-gate** | Pre-edit hook (`--can-edit --rule `) that enforces rules from `.rules.md` before edits. Separate from rule agents (which only propose/review). | Requires FW-2, FW-3 | + +--- + +## Notes + +- **Dependency on `model-divergence-enforcement`**: FW-2 and FW-3 are designed with model-binding in their configs (per-role `model` in their schedule config), but hard-block enforcement is deferred until the model-divergence feature ships. The design docs note the dependency; implementation tasks in the backlog carry a "depends on" annotation. +- **Self-improvement loop**: Once the self-improvement loop is running with `work_source.area = "framework"`, it will pick up items from this backlog automatically. The loop template `templates/loops/self-improvement/loop.json` already supports `work_source.area`. +- **FW-4 is NOT in this backlog** — it's a manual parent task created via `status.py --create-task model-divergence-enforcement` and decomposed into subtasks. \ No newline at end of file diff --git a/design/framework/README.md b/design/framework/README.md new file mode 100644 index 0000000..388f98a --- /dev/null +++ b/design/framework/README.md @@ -0,0 +1,67 @@ +# Framework Agent Features — Design Index + +Status: **v1 draft 2026-06-25**. Design for three framework-level agent features: rule agents (proposer + reviewer), Agent tab redesign, and model-divergence enforcement. + +## What this is + +The Automaton framework's agent-feature layer: scheduled standalone agents for rule maintenance (Rule Proposer, Rule Reviewer), a redesigned dashboard Agent tab showing real phase roles instead of fake types, and a model-divergence enforcement system that prevents conflict-of-interest by ensuring different LLM roles use different models. + +Built on top of the existing `status.py` phase machine and loop infrastructure — no second enforcement surface. + +## Why + +- `.rules.md` is maintained manually; failure patterns from completed tasks are not systematically captured. +- The dashboard Agent tab shows 4 made-up agent types (`completed_task_archiver`, `single`, `audit`, `backlog`) — not the real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator). +- Conflict-of-interest: two sessions with the same LLM playing bug finder and adversarial bug finder (or implementer and reviewer) don't yield useful results. The framework needs computational enforcement of model divergence. + +## Documents + +- [`functional.md`](functional.md) — what v1 does, roles, triggers, outputs, schedules, conflict-of-interest, success criteria. **Read this first.** +- [`technical.md`](technical.md) — the implementation contract: file map, state schemas, `status.py` flags, runner flows, dashboard API changes, test coverage. **Read this if you're implementing.** +- [`BACKLOG.md`](BACKLOG.md) — v1 work queue (3 items) and deferred items. The self-improvement loop's work queue when `work_source.area = "framework"`. + +## v1 scope (draft) + +Three feature areas, each with a clear boundary: + +1. **Model-Divergence Enforcement** — Foundational layer. `models.json` manifest, mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, audit category, dashboard badges. Enables the other two features. + +2. **Rule Agents** — Two standalone scheduled agents (NOT task lifecycle phases): + - **Rule Proposer**: Daily scan of completed tasks → `RULE_PROPOSALS.md` with concrete examples. + - **Rule Reviewer**: Monthly consolidation of `.rules.md` → `RULE_REVIEW.md`. + - Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers (mirror `--install-cleanup-schedule`). + +3. **Agent Tab Redesign** — Dashboard `/agent` view: + - **Phase Roles section**: 6 roles from `.agent.md`, active/inactive based on current task phases. + - **Scheduled Jobs section**: Real job kinds from `/api/scheduled` (cleanup, loop, rule-scan, rule-review), replacing 4 fake `AGENT_TYPE_META` types. + - New `/api/phase-roles` endpoint. + +## Locked decisions (referenced as `(Fn)` in functional/technical) + +| ID | Decision | +|---|---| +| F1 | Rule agents are **standalone scheduled agents**, NOT task lifecycle phases. They don't block task completion. | +| F2 | Rule Proposer runs **daily**; Rule Reviewer runs **monthly**. OS-native schedulers (launchd/cron/schtasks). | +| F3 | Rule agents use **direct harness invocation** (reuse `loop-runner._invoke_harness`), not loop infrastructure. | +| F4 | Conflict-of-interest: Rule Reviewer **must use different LLM** than Rule Proposer. Rule agents must use different LLM than tasks they review. Enforcement **deferred** until model-divergence ships (F5). | +| F5 | Model-divergence enforcement is a **foundational layer** shipped first (manual task, not backlog). Rule agents designed with model-binding config but hard-block enforced later. | +| F6 | Agent tab has **two sections**: Phase Roles (from `.agent.md` + task state) + Scheduled Jobs (from `/api/scheduled`). | +| F7 | Phase Roles section needs **new `/api/phase-roles` endpoint** mapping active tasks to their phase roles. | +| F8 | Scheduled Jobs section uses **real `job.kind`** (cleanup, loop, rule-scan, rule-review) — remove `AGENT_TYPE_META` fake types. | +| F9 | State files: `.state.rule-scan` (last scanned task, timestamp, proposed rules), `.state.rule-review` (last run timestamp). | +| F10 | New `status.py` flags: `--rule-scan`, `--install-rule-scan-schedule`, `--rule-review`, `--install-rule-review-schedule` (mirror `--cleanup-done` / `--install-cleanup-schedule`). | + +## Relationship to loops + +- `design/loops/` is the loop engineering system (state-enforced unattended work). +- `design/framework/` is framework-level agent features (rule maintenance, observability, model governance). +- The self-improvement loop (`templates/loops/self-improvement/`) can drive `design/framework/` work by setting `work_source.area = "framework"` — no code change needed (see `loop-runner.py:_find_work_backlog`). +- Model-divergence enforcement (F5) is a prerequisite for hard-blocking conflict-of-interest in rule agents and loops. + +## Post-v1 + +After v1 lands and the self-improvement loop is running: + +- Rule enforcement gate (`--can-edit --rule `) — separate feature. +- Dashboard Agent tab v2: role details, schedule management, per-agent logs. +- Rule agents gain comprehension-debt tracking (`last-read-sha` per rule file). \ No newline at end of file diff --git a/design/framework/functional.md b/design/framework/functional.md new file mode 100644 index 0000000..03811a7 --- /dev/null +++ b/design/framework/functional.md @@ -0,0 +1,335 @@ +# Framework Agent Features — Functional Design + +Status: v1 (draft 2026-06-25). Supersedes any prior informal discussions of rule agents or Agent tab redesign. + +Audience: framework maintainers (currently: the human and one AI assistant). After handoff the self-improvement loop is also an audience — designs must be legible to a fresh-context LLM verifier. + +## 1. Problem + +Three gaps in the framework's agent layer: + +1. **`.rules.md` is maintained manually.** Failure patterns from completed tasks (`BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md`) are not systematically captured into rules. The Self-Improvement section of `.rules.md` (lines 33-36) says "Add one rule per observed failure mode with a concrete example" and "Consolidate contradictions monthly" — but no agent does this. It relies on the human or a session agent remembering. + +2. **The dashboard Agent tab shows fake agent types.** `AGENT_TYPE_META` in `dashboard.js:330-334` defines 4 types (`completed_task_archiver`, `single`, `audit`, `backlog`) that are loop work-source kinds, not real agent roles. The 6 real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) are not surfaced. + +3. **No conflict-of-interest enforcement on model divergence.** `status.py:1432-1436` enforces that the reviewer session differs from the implementer session (via `.state.implementer`), but there is no enforcement that the *model* playing bug-finder differs from the model playing adversarial bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping. + +## 2. Goals + +v1 — **the framework maintains its own rules, surfaces real agent roles, and computationally enforces model divergence**: + +1. **Rule Proposer**: a daily scheduled standalone agent that scans completed tasks since its last run, reads their failure artifacts, and proposes new rules with concrete examples to `RULE_PROPOSALS.md`. +2. **Rule Reviewer**: a monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, rules missing examples — and writes `RULE_REVIEW.md`. +3. **Agent Tab Redesign**: replace 4 fake `AGENT_TYPE_META` types with two sections — Phase Roles (6 roles from `.agent.md`, active/inactive based on current task phases) and Scheduled Jobs (real `job.kind` from `/api/scheduled`). +4. **Model-Divergence Enforcement**: a `models.json` manifest, mode detection (single vs multi-LLM), a conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, an audit category, and dashboard badges. +5. **Conflict-of-interest for rule agents**: Rule Reviewer must use a different LLM than Rule Proposer. Both must use different LLMs than the tasks they review. Enforcement is *designed now* but *hard-blocked only after model-divergence ships* (F4, F5). + +## 3. Non-Goals (v1) + +- **Rule enforcement during task execution.** Rule agents only *propose* and *review* rules; they do not block edits. A future `--can-edit --rule ` gate (FW-6) is separate. +- **Auto-approving rules.** A human always reviews `RULE_PROPOSALS.md` and `RULE_REVIEW.md` before rules are merged into `.rules.md`. No auto-merge path in v1. +- **Rule agents as task lifecycle phases.** Rule agents are *standalone scheduled agents* (F1). They do not block task completion. They are not phases in the state machine. +- **A new agent harness.** Rule agents use direct harness invocation (reusing `loop-runner._invoke_harness`). No new runtime. +- **Network-fetched dependencies.** New code is Python stdlib only. No new pip installs. +- **Model capability inspection.** The framework never inspects model capability, provider, or size (D8). It only tracks *which* model fills *which* role and enforces the conflict matrix. + +## 4. Model-Divergence Enforcement (foundational layer) + +Shipped first as a manual task (`model-divergence-enforcement`), not a backlog item. Enables conflict-of-interest hard-blocking for rule agents and loops. + +### 4.1 Manifest: `models.json` + +A new file at `~/.automaton/models.json` (or `{project}/.automaton/models.json`): + +```json +{ + "default": "glm-4.6", + "advised": true, + "models": [ + {"name": "glm-4.6", "provider": "opencode", "context_window": 131072, "location": "remote"}, + {"name": "qwen3-coder", "provider": "opencode", "context_window": 131072, "location": "remote"}, + {"name": "llama-3.3-70b", "provider": "localhost", "context_window": 32768, "location": "http://localhost:8080"} + ] +} +``` + +- `default`: the model used when no role-specific binding is set. +- `advised`: if `true`, the framework prints a one-time advisory in single-LLM mode recommending a second model for conflict-of-interest roles, then goes silent. +- `models[]`: the roster. `location` is `"remote"` or a localhost URL for probing. + +### 4.2 Mode Detection + +- **0-1 models** in `models.json` (or file missing) → **single-LLM mode**. Advisory once (if `advised: true`), then silent. No hard blocks. +- **2+ models** → **multi-LLM mode**. Hard-block on conflict-matrix violations. Auto-assign next-available non-conflicting model on conflict; refuse only if no non-conflicting model exists. +- **Missing file** → single-LLM mode (backward compatible). Existing behavior preserved. + +### 4.3 Conflict Matrix (locked) + +| Role | Must differ from | +|---|---| +| `code_review` | `implement` | +| `bug_find` | `implement` | +| `adversarial_bug_find` | `implement`, `bug_find` | +| `referee` | `implement`, `bug_find`, `adversarial_bug_find` | +| `loop-verify` | `loop-implement` | + +`doc_review`, `code_review`, and `bug_find` are independent of each other (not conflicts). Only `bug_find` ↔ `adversarial_bug_find` conflicts (they are adversary pairs). + +### 4.4 Auto-Assignment (multi-LLM mode) + +1. Default model → assigned to `implement` (and `loop-implement`). +2. On conflict, pick the next-available model from `models[]` that does not conflict. +3. User override: `loop.json` `roles..model` or `status.py --transition --model `. +4. Refuse only if no non-conflicting model exists. + +### 4.5 State: `.state.models` + +Each task gets `{task}/.state.models` recording which model filled which role: + +```json +{"implement": "glm-4.6", "code_review": "qwen3-coder", "bug_find": "qwen3-coder", "adversarial_bug_find": "llama-3.3-70b", "referee": "llama-3.3-70b"} +``` + +`status.py --transition --model ` records the model for the role being transitioned into. `--claim` in multi-LLM mode checks the conflict matrix against `.state.models` and refuses on violation. + +### 4.6 Loop Integration + +`loop.json` gains per-role `model` and `harness.command` with `{model}` substitution: + +```json +"roles": { + "implement": {"prompt": "loop-implement.md", "model": "glm-4.6"}, + "verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"}, + "orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"} +} +``` + +`loop-runner.py:_invoke_harness` substitutes `{model}` into the harness command. `--check-gate` enforces `loop-verify` ≠ `loop-implement` model in multi-LLM mode. + +### 4.7 Audit + Dashboard + +- `status.py --audit` gains a `model_divergence` category: flags tasks where `.state.models` violates the conflict matrix. +- Dashboard task cards show model badges (one per role filled). + +## 5. Rule Proposer + +A **standalone scheduled agent** (F1) that proposes new rules from completed-task failure patterns. + +### 5.1 Trigger + +Daily, via OS-native scheduler (mirrors `--install-cleanup-schedule`). `status.py --install-rule-scan-schedule [--interval 86400]` installs the schedule unit. Manual: `status.py --rule-scan`. + +### 5.2 Inputs + +- `.state.rule-scan`: state file tracking the last scanned task and timestamp. +- Completed tasks (in `tasks/complete/` or tasks with `.state` phase `complete`) that were completed since the last scan. +- For each such task: `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md` (if present). +- Current `.rules.md` (to deduplicate against existing rules). + +### 5.3 Output + +`RULE_PROPOSALS.md` (at `~/.automaton/RULE_PROPOSALS.md` in framework mode, or `{project}/.automaton/RULE_PROPOSALS.md` in project mode). Format: + +```markdown +# Rule Proposals — {date} + +## Proposed Rule: {title} +**Source**: tasks/{task-name}/VERDICT.md +**Pattern**: {one-line description of the failure mode} +**Example**: +{concrete code/config snippet from the task} +**Proposed rule text**: +{the rule as it would appear in .rules.md} + +--- +``` + +Proposals are *appended* per run. A human reviews and merges accepted rules into `.rules.md`. No auto-merge (Non-Goal). + +### 5.4 LLM Session + +Direct harness invocation (F3): `status.py --rule-scan` reuses `loop-runner._invoke_harness` to spawn one LLM session with a prompt that includes the failure artifacts and current rules. The session proposes rules in the `RULE_PROPOSALS.md` format. + +### 5.5 Conflict-of-Interest + +The Rule Proposer's LLM must differ from the implementer + bug-hunter + adversarial-bug-hunter models of the tasks it scans. This prevents the model that made the bug from proposing the rule about its own bug. + +**Enforcement deferred** (F4): until model-divergence ships, the Rule Proposer runs with the default model. The design includes a `model` field in its schedule config; hard-blocking activates once `models.json` exists and multi-LLM mode is detected. + +### 5.6 State: `.state.rule-scan` + +```json +{"last_scan_at": "2026-06-25T10:00:00Z", "last_scanned_task": "fix-context-sizing", "proposals_count": 3} +``` + +## 6. Rule Reviewer + +A **standalone scheduled agent** (F1) that consolidates `.rules.md` periodically. + +### 6.1 Trigger + +Monthly, via OS-native scheduler. `status.py --install-rule-review-schedule [--interval 2592000]` installs the schedule unit. Manual: `status.py --rule-review`. + +### 6.2 Inputs + +- `.state.rule-review`: state file tracking the last run timestamp. +- Current `.rules.md` (full file). +- Recent `RULE_PROPOSALS.md` entries (since last review). +- Recent completed-task summaries (last 30 days) for context on stale rules. + +### 6.3 Output + +`RULE_REVIEW.md` (at `~/.automaton/RULE_REVIEW.md` or project equivalent). Format: + +```markdown +# Rule Review — {date} + +## Contradictions Found +- Rule A ("...") contradicts Rule B ("..."). Suggested resolution: {merge/drop/keep A}. + +## Stale Rules (no observed instance in last 30 days) +- Rule C ("..."). Suggested action: drop or annotate as low-priority. + +## Rules Missing Examples +- Rule D ("..."). Suggested example: {from a recent task}. + +## Merge Candidates +- Rules E and F overlap. Suggested merged text: {...}. + +--- +``` + +A human reviews and applies accepted changes to `.rules.md`. No auto-apply (Non-Goal). + +### 6.4 LLM Session + +Direct harness invocation (F3), same as Rule Proposer. One LLM session with a prompt that includes the full `.rules.md` and recent proposals. + +### 6.5 Conflict-of-Interest + +The Rule Reviewer's LLM **must differ from the Rule Proposer's LLM** (F4). The Proposer proposes (bias toward adding); the Reviewer consolidates (bias toward pruning). Same model = self-review = rubber-stamping. + +**Enforcement deferred** until model-divergence ships. + +### 6.6 State: `.state.rule-review` + +```json +{"last_review_at": "2026-06-25T10:00:00Z", "contradictions_found": 2, "stale_rules": 5, "merges_suggested": 1} +``` + +## 7. Agent Tab Redesign + +### 7.1 Current State (to be replaced) + +`dashboard.js:330-334` defines `AGENT_TYPE_META` with 4 fake types: +- `completed_task_archiver` — actually the cleanup scheduled job. +- `single` — actually a loop with `work_source.kind = "single"`. +- `audit` — actually a loop with `work_source.kind = "audit"`. +- `backlog` — actually a loop with `work_source.kind = "backlog"`. + +These are loop work-source kinds, not agent roles. They conflate two different concepts. + +### 7.2 New Design: Two Sections + +**Section 1 — Phase Roles** + +Shows the 6 phase roles from `.agent.md` Agent Configuration: + +| Role ID | Phases | Active when | +|---|---|---| +| `researcher` | research, decomposition, design, test_design | A task is in one of these phases | +| `implementer` | implement | A task is in `implement` phase | +| `code-reviewer` | code_review | A task is in `code_review` phase | +| `bug-hunter` | bug_find, adversarial_bug_find | A task is in one of these phases | +| `referee` | referee | A task is in `referee` phase | +| `orchestrator` | new, complete, human_intervention | A task is in one of these phases | + +Each role card shows: role icon, role label, status (Active/Idle — based on whether any task is in that role's phases), and the count of tasks in that role's phases. Clicking a role filters the task list to tasks in that role's phases. + +**Data source**: new `/api/phase-roles` endpoint. Returns: + +```json +{ + "roles": [ + {"id": "researcher", "label": "Researcher", "icon": "🔬", "phases": ["research", "decomposition", "design", "test_design"], "active_tasks": 2, "status": "active"}, + {"id": "implementer", "label": "Implementer", "icon": "⚙️", "phases": ["implement"], "active_tasks": 1, "status": "active"}, + ... + ] +} +``` + +**Section 2 — Scheduled Jobs** + +Shows real scheduled jobs from `/api/scheduled`, using `job.kind` (not fake agent types): + +| Job Kind | Icon | Label | Source | +|---|---|---|---| +| `cleanup` | 🧹 | Cleanup Archiver | `com.automaton.cleanup` | +| `loop` | 🔄 | Loop: {name} | `com.automaton.loop.{name}` | +| `rule-scan` | 📝 | Rule Proposer | `com.automaton.rule-scan` (new) | +| `rule-review` | 📋 | Rule Reviewer | `com.automaton.rule-review` (new) | + +Each job card shows: job label, status (Enabled/Disabled/Misconfigured), next-run interval, runtime state (for loops: iteration count, halt status; for rule agents: last-scan/review timestamp). `AGENT_TYPE_META` is removed entirely; rendering uses `job.kind` directly. + +### 7.3 Self-Documenting Names + +Per `.rules.md` "Self-Documenting UI Names" (lines 53-64), all schedule unit names and stub filenames are self-documenting: +- `com.automaton.rule-scan` (launchd label) +- `automaton-rule-scan.sh` (stub filename) +- `com.automaton.rule-review` / `automaton-rule-review.sh` + +## 8. Schedules + +Rule agents use OS-native schedulers, mirroring the `--install-cleanup-schedule` pattern (`status.py:2188-2260`): + +| Agent | Flag | Default Interval | Launchd Label | Stub | +|---|---|---|---|---| +| Rule Proposer | `--install-rule-scan-schedule` | 86400s (daily) | `com.automaton.rule-scan` | `automaton-rule-scan.sh` | +| Rule Reviewer | `--install-rule-review-schedule` | 2592000s (monthly) | `com.automaton.rule-review` | `automaton-rule-review.sh` | + +Platform dispatch via `platform.system()`: +- **Darwin**: `~/Library/LaunchAgents/com.automaton.rule-scan.plist` with `StartInterval`. +- **Linux**: crontab line via `_install_cron_block_generic`. +- **Windows**: `schtasks /create /tn "AutomatonRuleScan" ...`. + +Stub scripts are 3-line bash/bat files that call `python3 status.py --rule-scan` (or `--rule-review`). + +## 9. Conflict-of-Interest + +### 9.1 Dependency Chain + +``` +model-divergence-enforcement (shipped first, manual task) + ↓ enables hard-block +rule-proposer-agent (FW-2) + ↓ conflict-of-interest +rule-reviewer-agent (FW-3) — must differ from Proposer +``` + +### 9.2 Design Now, Enforce Later (F4, F5) + +- Rule agents are *designed* with `model` fields in their schedule config. +- The design docs specify the conflict-of-interest rules. +- Hard-block enforcement *activates* when `models.json` exists and multi-LLM mode is detected. +- Until then, rule agents run with the default model (single-LLM mode, advisory only). + +### 9.3 Why Different Models + +- **Proposer vs Reviewer**: Proposer has a bias toward *adding* rules (more is better). Reviewer has a bias toward *pruning* (less is better). Same model = self-review = the proposer's rules never get pruned. +- **Rule agent vs scanned tasks**: The model that introduced a bug should not propose the rule about its own bug — it has a blind spot for that failure mode. + +## 10. Success Criteria for v1 + +1. **Model-divergence**: `--transition --model` records the model in `.state.models`; `--claim` refuses conflict-matrix violations in multi-LLM mode; `--audit` flags violations; dashboard shows model badges. +2. **Rule Proposer**: `status.py --rule-scan` reads completed tasks since last scan, proposes rules to `RULE_PROPOSALS.md`, updates `.state.rule-scan`. Test with a seeded completed task containing a `VERDICT.md`. +3. **Rule Reviewer**: `status.py --rule-review` reads `.rules.md` + recent proposals, writes `RULE_REVIEW.md`, updates `.state.rule-review`. Test with a seeded `.rules.md` containing a contradiction. +4. **Agent Tab**: `/api/phase-roles` returns 6 roles with active-task counts; dashboard renders Phase Roles + Scheduled Jobs sections; `AGENT_TYPE_META` is removed; `job.kind` drives rendering. +5. **Schedules**: `--install-rule-scan-schedule` and `--install-rule-review-schedule` install OS-native units with self-documenting names. +6. **Tests**: `pytest tests/ -v` is green; new tests cover state schemas, scan flows, dashboard API, and conflict-matrix enforcement. +7. **No regression**: pre-existing test suite passes unchanged. + +## 11. Locked Decision Index + +All decisions referenced by `(Fn)` above are recorded in the v1 design conversation (this session). They are non-negotiable for v1 implementation. Changes require a design doc update and a new `[unreleased]` changelog entry. + +See `README.md` § "Locked decisions" for the full table. diff --git a/design/framework/technical.md b/design/framework/technical.md new file mode 100644 index 0000000..12d38e3 --- /dev/null +++ b/design/framework/technical.md @@ -0,0 +1,536 @@ +# Framework Agent Features — Technical Design + +Companion to `functional.md`. This file is the implementation contract: every line here is what the implementation tasks build. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update. + +## 1. File Map (what v1 adds) + +``` +~/.automaton/ +├── models.json # NEW — model manifest (see §2) +├── scripts/ +│ ├── detect_models.py # NEW — probes opencode.json + localhost endpoints +│ └── status.py # EXTENDED — new flags (see §4) +├── prompts/ +│ ├── rule-proposer.md # NEW — Rule Proposer session prompt +│ ├── rule-reviewer.md # NEW — Rule Reviewer session prompt +│ └── onboarding.md # EXTENDED — Step 2e (backlog check) +├── automaton/ +│ └── dashboard/ +│ ├── html/dashboard.js # EXTENDED — remove AGENT_TYPE_META, two-section render +│ └── ui/app.py # EXTENDED — /api/phase-roles endpoint +├── .automaton/ # (framework self-hosting: this is ~/.automaton/.automaton/) +│ ├── .state.rule-scan # NEW — Rule Proposer state (see §3) +│ ├── .state.rule-review # NEW — Rule Reviewer state (see §3) +│ ├── RULE_PROPOSALS.md # NEW — Rule Proposer output (append-per-run) +│ ├── RULE_REVIEW.md # NEW — Rule Reviewer output (append-per-run) +│ └── automaton-rule-scan.sh # NEW — generated by --install-rule-scan-schedule +│ automaton-rule-review.sh # NEW — generated by --install-rule-review-schedule +└── tests/ + ├── test_model_divergence.py # NEW — manifest, conflict matrix, --transition --model + ├── test_rule_agents.py # NEW — scan flows, state files, output schemas + └── test_dashboard_phase_roles.py # NEW — /api/phase-roles, two-section render +``` + +Per-project paths mirror the loop convention: `{project}/.automaton/.state.rule-scan`, `{project}/.automaton/RULE_PROPOSALS.md`, etc. For framework self-hosting, the project is `~/.automaton/` itself. + +## 2. `models.json` Schema + +```json +{ + "schema_version": 1, + "default": "glm-4.6", + "advised": true, + "models": [ + { + "name": "glm-4.6", + "provider": "opencode", + "context_window": 131072, + "location": "remote" + }, + { + "name": "qwen3-coder", + "provider": "opencode", + "context_window": 131072, + "location": "remote" + }, + { + "name": "llama-3.3-70b", + "provider": "localhost", + "context_window": 32768, + "location": "http://localhost:8080" + } + ] +} +``` + +- `default`: model name used when no role-specific binding exists. Must be present in `models[]`. +- `advised`: bool. If `true`, single-LLM mode prints a one-time advisory recommending a second model, then goes silent. +- `models[]`: roster. `name` is the unique key. `provider` is informational. `context_window` is informational (framework never inspects capability, D8). `location` is `"remote"` or a localhost URL (for `detect_models.py` probing). +- **Missing file** → single-LLM mode (backward compatible). All model-divergence commands are no-ops. +- **0-1 models** → single-LLM mode. Advisory once if `advised: true`. +- **2+ models** → multi-LLM mode. Hard-block on conflict matrix. + +### 2.1 `detect_models.py` + +``` +python3 scripts/detect_models.py [--json] +``` + +1. Parse `opencode.json` (or `opencode.jsonc`) for provider+model entries. +2. Probe localhost endpoints: `http://localhost:8080/v1/models`, `http://localhost:11434/api/tags` (Ollama), `http://localhost:1234/v1/models` (LM Studio), `http://localhost:8000/v1/models` (vLLM). +3. Merge results, emit a candidate `models.json` to stdout (or write if `--json` not set). +4. Used by `install.sh` / `update.sh` / `upgrade.sh` to bootstrap or refresh `models.json`. + +## 3. State Schemas + +### 3.1 `.state.rule-scan` + +```json +{ + "schema_version": 1, + "last_scan_at": "2026-06-25T10:00:00Z", + "last_scanned_task": "fix-context-sizing", + "proposals_count": 3, + "scanned_tasks_count": 12 +} +``` + +- `last_scanned_task`: the most recent task name scanned. Next scan starts after this task (alphabetical or mtime order). +- `proposals_count`: cumulative count of proposals written to `RULE_PROPOSALS.md`. +- Stored at `{project}/.automaton/.state.rule-scan`. Missing file → first run scans all completed tasks. + +### 3.2 `.state.rule-review` + +```json +{ + "schema_version": 1, + "last_review_at": "2026-06-25T10:00:00Z", + "contradictions_found": 2, + "stale_rules": 5, + "merges_suggested": 1 +} +``` + +- Stored at `{project}/.automaton/.state.rule-review`. Missing file → first run reviews all rules. + +### 3.3 `.state.models` (per-task) + +```json +{ + "schema_version": 1, + "implement": "glm-4.6", + "code_review": "qwen3-coder", + "bug_find": "qwen3-coder", + "adversarial_bug_find": "llama-3.3-70b", + "referee": "llama-3.3-70b", + "doc_review": null +} +``` + +- Stored at `{task}/.state.models`. One file per task. +- Written by `--transition --model ` when entering a phase. +- Read by `--claim` (conflict-matrix check) and `--audit` (violation detection). +- Roles not yet filled are `null` or absent. + +## 4. `status.py` New Flags + +All model-divergence and rule-agent commands route through `status.py` — no second enforcement surface. + +``` +# Model-divergence +status.py --transition --task [--model ] Records model in .state.models; checks conflict matrix +status.py --claim --task --agent [--model ] Refuses if model conflicts with filled roles (multi-LLM mode) +status.py --audit EXTENDED — +model_divergence category +status.py --can-edit [...] UNCHANGED + +# Rule agents +status.py --rule-scan [--project

] [--dry-run] Scan completed tasks, propose rules to RULE_PROPOSALS.md +status.py --install-rule-scan-schedule [--interval S] Install OS-native unit for --rule-scan (default daily) +status.py --rule-review [--project

] [--dry-run] Consolidate .rules.md, write RULE_REVIEW.md +status.py --install-rule-review-schedule [--interval S] Install OS-native unit for --rule-review (default monthly) +``` + +### 4.1 `--transition --model` flow + +1. Load `models.json`. If missing or single-LLM mode → record model (advisory), no conflict check. +2. If multi-LLM mode: load `.state.models` for the task. Check the role being entered against the conflict matrix (§5). +3. If `--model` not provided: auto-assign next-available non-conflicting model from `models[]`. Refuse if none available. +4. If `--model` provided: verify it's in `models[]`. Check conflict matrix. Refuse on violation. +5. Write `role: model` to `.state.models`. Transition the phase. + +### 4.2 `--claim --model` flow + +1. Load `models.json`. If single-LLM mode → existing claim logic, no model check. +2. If multi-LLM mode: load `.state.models`. Determine the role for the phase being claimed. Check conflict matrix against already-filled roles. +3. Refuse if the claiming agent's model conflicts. Error message names the conflicting role and model. + +### 4.3 `--audit` extension + +New audit category `model_divergence`: +- For each task with `.state.models`: check all filled roles against the conflict matrix. +- Flag violations as `severity: high` (conflict-of-interest is a correctness issue, not a style issue). +- Output format mirrors existing audit categories. + +## 5. Conflict Matrix (implementation) + +```python +CONFLICT_MATRIX = { + "code_review": {"implement"}, + "bug_find": {"implement"}, + "adversarial_bug_find": {"implement", "bug_find"}, + "referee": {"implement", "bug_find", "adversarial_bug_find"}, + "loop-verify": {"loop-implement"}, +} +``` + +- Key = role being entered. Value = set of roles that must have a different model. +- `doc_review`, `code_review`, `bug_find` are NOT in conflict with each other (only `bug_find` ↔ `adversarial_bug_find` conflicts). +- Check function: `def _check_conflict(state_models: dict, role: str, model: str, matrix: dict) -> Optional[str]` — returns the conflicting role name or `None`. + +## 6. Rule Proposer Flow (`--rule-scan`) + +``` +1. Load .state.rule-scan (or init if missing). +2. Find completed tasks since last_scanned_task: + - Scan tasks/complete/ and tasks with .state phase=complete + - Filter by mtime > last_scan_at (or all if first run) + - Sort by mtime ascending +3. For each task: + a. Read BUG_REPORT.md, ADVERSARIAL_BUG_REPORT.md, VERDICT.md (skip if none exist) + b. Read current .rules.md (for dedup context — capped at 4k tokens) + c. Build proposer prompt (see §7) + d. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_proposer_model) + e. Parse LLM output for proposed rules (expect RULE_PROPOSALS.md format) + f. Append proposals to RULE_PROPOSALS.md + g. Update .state.rule-scan (last_scanned_task, proposals_count) +4. Write final .state.rule-scan with last_scan_at = now. +``` + +- `--dry-run`: list tasks that would be scanned, do not invoke harness. +- `--project`: scope to a project (default: framework dir). +- Errors during a single task scan do not abort the run; the scan continues to the next task and logs the error. + +## 7. Rule Proposer Prompt Shape (`rule-proposer.md`) + +``` +# Rule Proposer — {date} + +You are scanning completed tasks for failure patterns that should become rules. + +## Current rules (read-only, for dedup) +{current_rules} # .rules.md content, capped at 4k tokens + +## Task failure artifacts +{bug_report} # BUG_REPORT.md content, capped at 2k tokens +{adversarial_report} # ADVERSARIAL_BUG_REPORT.md, capped at 2k tokens +{verdict} # VERDICT.md, capped at 2k tokens + +## What to do +For each distinct failure pattern you observe: +1. Check if a rule already exists in .rules.md that covers it. If so, skip. +2. If no existing rule covers it, propose a new rule with: + - A concrete example from the task artifacts + - The proposed rule text as it would appear in .rules.md + +## Output (strict markdown, no JSON) +## Proposed Rule: {title} +**Source**: tasks/{task-name}/VERDICT.md +**Pattern**: {one-line description} +**Example**: +{concrete snippet} +**Proposed rule text**: +{rule text} + +--- +``` + +No `{model}` token in the prompt — the model is selected by the caller and passed to `_invoke_harness`. + +## 8. Rule Reviewer Flow (`--rule-review`) + +``` +1. Load .state.rule-review (or init if missing). +2. Read .rules.md (full file). +3. Read recent RULE_PROPOSALS.md entries (since last_review_at). +4. Read recent completed-task summaries (last 30 days) for staleness context. +5. Build reviewer prompt (see §9). +6. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_reviewer_model). +7. Parse LLM output for review sections (contradictions, stale, missing examples, merges). +8. Append to RULE_REVIEW.md. +9. Update .state.rule-review. +``` + +- `--dry-run`: report what would be reviewed, do not invoke harness. + +## 9. Rule Reviewer Prompt Shape (`rule-reviewer.md`) + +``` +# Rule Reviewer — {date} + +You are consolidating .rules.md for contradictions, staleness, and missing examples. + +## Current rules (full) +{rules_content} # .rules.md, full file + +## Recent proposals (since last review) +{recent_proposals} # RULE_PROPOSALS.md entries since last_review_at + +## Recent completed tasks (last 30 days, for staleness context) +{task_summaries} # one-line per task: name + phase + completion date + +## What to check +1. Contradictions: rules that conflict with each other. +2. Stale rules: no observed instance in last 30 days. +3. Rules missing examples: any rule without a concrete example. +4. Merge candidates: overlapping rules that could be consolidated. + +## Output (strict markdown, no JSON) +## Contradictions Found +- ... +## Stale Rules +- ... +## Rules Missing Examples +- ... +## Merge Candidates +- ... +``` + +## 10. Harness Invocation (direct, not loop) + +Rule agents reuse `loop-runner._invoke_harness` directly — they are NOT loops. The function signature (from `loop-runner.py:366-404`): + +```python +def _invoke_harness(harness_command: str, prompt_content: str, cwd: str, env: dict = None) -> str: +``` + +`status.py --rule-scan` calls this as: + +```python +from loop_runner import _invoke_harness +output = _invoke_harness( + harness_command=rule_harness_command, # from schedule config or default + prompt_content=resolved_prompt, # rule-proposer.md with tokens substituted + cwd=str(project_dir), + env={"AUTOMATON_RULE_ROLE": "proposer"} +) +``` + +`{model}` substitution: if the harness command contains `{model}`, it's replaced with the rule agent's configured model. Until model-divergence ships, this is the default model. + +### 10.1 Schedule Config for Rule Agents + +Rule agents do not use `loop.json`. Their config is embedded in the schedule stub: + +```bash +#!/usr/bin/env bash +cd "" +python3 "/scripts/status.py" --rule-scan --model +``` + +The `--model` flag is optional and ignored in single-LLM mode. In multi-LLM mode it sets the rule agent's model (subject to conflict-of-interest checks once enforced). + +## 11. Scheduler Unit Generation + +Mirrors `cmd_install_cleanup_schedule` (`status.py:2188-2260`) exactly: + +### 11.1 `--install-rule-scan-schedule` + +```python +def cmd_install_rule_scan_schedule(args) -> int: + interval = args.interval if args.interval else 86400 # daily + # 1. Write stub: automaton-rule-scan.sh + # 2. Platform dispatch: + # Darwin → ~/Library/LaunchAgents/com.automaton.rule-scan.plist + # Linux → crontab block via _install_cron_block_generic + # Windows → schtasks /create /tn "AutomatonRuleScan" +``` + +### 11.2 `--install-rule-review-schedule` + +```python +def cmd_install_rule_review_schedule(args) -> int: + interval = args.interval if args.interval else 2592000 # monthly + # Same pattern, labels: com.automaton.rule-review / AutomatonRuleReview +``` + +### 11.3 `_list_scheduled_jobs` extension + +`_list_scheduled_jobs` (`status.py:2282`) gains recognition for new labels: + +```python +def _launchd_label_kind(label: str) -> tuple[str, str]: + if label.startswith("com.automaton.loop."): + return ("loop", label[len("com.automaton.loop."):]) + if label in ("com.automaton.cleanup",): + return ("cleanup", "") + if label in ("com.automaton.rule-scan",): + return ("rule-scan", "") + if label in ("com.automaton.rule-review",): + return ("rule-review", "") + ... +``` + +This makes rule-scan and rule-review jobs appear in `/api/scheduled` with their real `kind`, which the Agent tab renders directly. + +## 12. Agent Tab Data Flow + +### 12.1 New endpoint: `/api/phase-roles` + +`app.py` gains a handler: + +```python +elif self.path == "/api/phase-roles": + self._serve_phase_roles() +``` + +```python +def _serve_phase_roles(self): + # 1. Parse .agent.md Agent Configuration for role definitions + # 2. Load all tasks via status.py module + # 3. For each role, count tasks in that role's phases + # 4. Return JSON: + { + "roles": [ + {"id": "researcher", "label": "Researcher", "icon": "🔬", + "phases": ["research", "decomposition", "design", "test_design"], + "active_tasks": 2, "status": "active"}, + ... + ], + "available": True + } +``` + +Role icons (self-documenting, per `.rules.md` Self-Documenting UI Names): + +| Role | Icon | +|---|---| +| researcher | 🔬 | +| implementer | ⚙️ | +| code-reviewer | 👁️ | +| bug-hunter | 🐛 | +| referee | ⚖️ | +| orchestrator | 🎯 | + +### 12.2 `dashboard.js` changes + +**Remove**: `AGENT_TYPE_META` (lines 330-334), `AGENT_TYPES` (337), `AGENT_TYPE_META_FALLBACK` (338), `_resolveAgentType` (351-354). + +**Replace `renderAgentTab`** with a two-section render: + +```javascript +async function renderAgentTab() { + const panel = document.getElementById('agent-panel'); + panel.innerHTML = '

Loading…
'; + + const [rolesRes, schedRes] = await Promise.all([ + fetch('/api/phase-roles').then(r => r.json()).catch(() => ({roles: [], available: false})), + fetchSchedule(), + ]); + + // Section 1: Phase Roles + const rolesHtml = rolesRes.available ? renderPhaseRoles(rolesRes.roles) + : '
Phase roles require .agent.md Agent Configuration.
'; + + // Section 2: Scheduled Jobs + const jobsHtml = renderScheduledJobs(schedRes.jobs || []); + + panel.innerHTML = ` +
+

Phase Roles

+
${rolesHtml}
+
+
+

Scheduled Jobs

+
${jobsHtml}
+
`; +} +``` + +**`renderScheduledJobs`** uses `job.kind` directly (no fake type resolution): + +```javascript +const JOB_META = { + cleanup: { icon: '🧹', label: 'Cleanup Archiver' }, + loop: { icon: '🔄', label: (j) => `Loop: ${j.name}` }, + 'rule-scan': { icon: '📝', label: 'Rule Proposer' }, + 'rule-review':{ icon: '📋', label: 'Rule Reviewer' }, +}; +``` + +## 13. Loop Integration (model-divergence) + +### 13.1 `loop.json` per-role model + +```json +"roles": { + "implement": {"prompt": "loop-implement.md", "model": "glm-4.6"}, + "verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"}, + "orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"} +} +``` + +- `model` is optional. If absent, uses `models.json` `default`. +- `loop-verify` model is checked against `loop-implement` model in `--check-gate` (multi-LLM mode). + +### 13.2 `{model}` substitution in `_invoke_harness` + +`loop-runner.py:366-404` `_invoke_harness` gains `{model}` token substitution: + +```python +def _invoke_harness(harness_command, prompt_content, cwd, env=None, model=None): + if model and "{model}" in harness_command: + harness_command = harness_command.replace("{model}", model) + ... +``` + +The caller passes `model` from the role config. If the harness command has no `{model}` token, the model is informational only (the harness picks its own). + +### 13.3 `--check-gate` model-divergence check + +In multi-LLM mode, `--check-gate` adds: +- Load `loop.json` roles. Compare `verify.model` vs `implement.model`. +- If same model and multi-LLM mode → halt as `model_conflict` (new halt reason, or reuse `human_intervention` with a descriptive message). + +## 14. Test Coverage + +### 14.1 `test_model_divergence.py` + +- `test_models_json_missing_single_llm_mode` — no file → advisory, no blocks. +- `test_single_model_advisory_once` — 1 model, `advised: true` → advisory printed once, then silent. +- `test_multi_llm_conflict_matrix` — 2+ models, `--transition --model` records, `--claim` refuses conflict. +- `test_auto_assign_next_available` — no `--model` flag → auto-assigns non-conflicting model. +- `test_auto_assign_exhausted` — all models conflict → refuse. +- `test_audit_model_divergence` — `--audit` flags conflict-matrix violations. +- `test_loop_verify_neq_implement` — `--check-gate` halts on same model in multi-LLM mode. + +### 14.2 `test_rule_agents.py` + +- `test_rule_scan_finds_completed_tasks` — seeded completed task with VERDICT.md → proposal written. +- `test_rule_scan_state_tracking` — `.state.rule-scan` updated with last_scanned_task + count. +- `test_rule_scan_dedup` — existing rule in `.rules.md` → not re-proposed. +- `test_rule_scan_dry_run` — no harness invocation, lists candidates. +- `test_rule_review_finds_contradictions` — seeded `.rules.md` with contradiction → review written. +- `test_rule_review_state_tracking` — `.state.rule-review` updated. +- `test_install_rule_scan_schedule` — stub + plist created with correct labels. +- `test_install_rule_review_schedule` — stub + plist created with correct labels. + +### 14.3 `test_dashboard_phase_roles.py` + +- `test_api_phase_roles` — `/api/phase-roles` returns 6 roles with correct phases. +- `test_phase_roles_active_count` — tasks in phases → correct active_tasks count. +- `test_scheduled_jobs_new_kinds` — rule-scan and rule-review jobs appear with correct `kind`. +- `test_agent_type_meta_removed` — `AGENT_TYPE_META` no longer in dashboard.js (grep test). + +## 15. Rollout (3 sequential tasks for model-divergence) + +The model-divergence-enforcement parent task decomposes into 3 subtasks: + +1. **manifest+detection**: `models.json` schema, `detect_models.py`, `install.sh`/`update.sh`/`upgrade.sh` integration, `config.md` section, onboarding Step 2d. +2. **interactive enforcement+audit**: `.state.models`, `--transition --model`, `--claim --model` conflict check, `--audit` model_divergence category, dashboard badges. +3. **loop enforcement+dashboard**: `loop.json` per-role model, `{model}` substitution, `--check-gate` model check, loop dashboard badges. + +Rule agents (FW-2, FW-3) and Agent tab (FW-1) are backlog items, picked up after model-divergence ships (for FW-2/FW-3) or independently (for FW-1). + +## 16. Locked Decision Index + +All decisions referenced by `(Fn)` are in `README.md` § "Locked decisions". Implementation must conform. Deviations require a design doc update + `[unreleased]` CHANGELOG entry. diff --git a/design/loops/BACKLOG.md b/design/loops/BACKLOG.md index db3b262..cb2b31b 100644 --- a/design/loops/BACKLOG.md +++ b/design/loops/BACKLOG.md @@ -6,6 +6,8 @@ Work queue for the self-improvement loop after task 7 lands. Items not assigned A loop configured with `work_source.kind = "backlog"` reads this file, picks the topmost `[ ]` item, drafts an implementation, transitions through phases, hands off to a human reviewer. Mark items `[x]` when complete; move items to `DONE.md` (created later) on closure. +**Sibling backlog**: `design/framework/BACKLOG.md` covers framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). A loop with `work_source.area = "framework"` reads that file instead of this one. + ## v1.1 — framework manages its own docs - [ ] **design-update-loop-template** — `templates/loops/design-update/` that keeps `design//*.md` in sync with the code it documents. Triggered by `last-read-sha` drift detection. diff --git a/design/loops/README.md b/design/loops/README.md index 709d34d..5748f6b 100644 --- a/design/loops/README.md +++ b/design/loops/README.md @@ -15,6 +15,7 @@ Per-session manual driving doesn't scale against the framework's growing backlog - [`functional.md`](functional.md) — what v1 does, roles, the five deaths, blast radius, schedules, success criteria. **Read this first.** - [`technical.md`](technical.md) — the implementation contract: file map, `.state.loop` schema, `status.py` flags, gate checks, runner flow, test coverage. **Read this if you're implementing.** - [`BACKLOG.md`](BACKLOG.md) — v1.1 and deferred items (Scope 2 design-update loop, Scope 3 self-designing, parallel mode, dashboard panel). The self-improvement loop's work queue. +- **Sibling design**: [`../framework/`](../framework/) — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). Has its own `BACKLOG.md` consumable via `work_source.area = "framework"`. ## v1 scope (locked) diff --git a/memory/v1-1-hardening-session.md b/memory/v1-1-hardening-session.md index e4ae22f..1b14c2f 100644 --- a/memory/v1-1-hardening-session.md +++ b/memory/v1-1-hardening-session.md @@ -53,4 +53,15 @@ Picked from BUG_REPORTs of the loop tasks and from `design/loops/BACKLOG.md` def ## NEXT -Resume task 2 (`add-state-loop-lock`): SPEC is already written, state is `research:awaiting_approval`. Approve it, transition to implement, write the `_loop_lock` helper + wrap callsites in `status.py` and `loop-runner.py`, add `tests/test_state_loop_lock.py` (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order. \ No newline at end of file +Resume task 2 (`add-state-loop-lock`): SPEC is already written, state is `research:awaiting_approval`. Approve it, transition to implement, write the `_loop_lock` helper + wrap callsites in `status.py` and `loop-runner.py`, add `tests/test_state_loop_lock.py` (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order. + +## Session 2026-06-25 — Framework agent features design + +- Created `design/framework/` with `README.md`, `functional.md`, `technical.md`, `BACKLOG.md` — design for three framework-level agent features: model-divergence enforcement, rule agents (Proposer + Reviewer), Agent tab redesign. +- **Model-divergence enforcement** is a manual task (`model-divergence-enforcement`), decomposed into 3 subtasks: (1) manifest+detection, (2) interactive enforcement+audit, (3) loop enforcement+dashboard. Not a backlog item. +- **Rule agents** (FW-2, FW-3) are backlog items depending on model-divergence shipping first (conflict-of-interest LLM binding). Rule Proposer runs daily, Rule Reviewer runs monthly. Both use direct harness invocation (reuse `loop-runner._invoke_harness`), not loop infrastructure. +- **Agent tab redesign** (FW-1) is a backlog item with no dependencies. Replaces 4 fake `AGENT_TYPE_META` types with Phase Roles (6 roles from `.agent.md`) + Scheduled Jobs (real `job.kind`). +- Decision: rule agents use **direct harness invocation** (not loops, not standalone status.py commands). +- Decision: model-divergence conflict-of-interest is **designed now, enforced later** — rule agents carry `model` fields in config but hard-blocking activates only when `models.json` exists and multi-LLM mode is detected. +- Cross-references updated: `AGENTS.md` repo layout, `README.md` loop config table, `design/loops/README.md`, `design/loops/BACKLOG.md`, `CHANGELOG.md`, `.onboarding.md` (Backlog section), `prompts/onboarding.md` (Step 2d). +- Next: switch on self-improvement loop (`--create-loop self-improvement --from-template self-improvement` + `--install-schedule`), then create `model-divergence-enforcement` parent task and decompose into 3 subtasks. \ No newline at end of file diff --git a/prompts/onboarding.md b/prompts/onboarding.md index 85ccc32..b396b06 100644 --- a/prompts/onboarding.md +++ b/prompts/onboarding.md @@ -75,6 +75,13 @@ Install the automaton pre-commit hook to block commits when no task is in an edi 4. Verify the hook: `python ~/.automaton/scripts/status.py --can-edit --project {project}` should return exit code 1 (DENIED) since no tasks exist yet. 5. Note in the onboarding report whether the hook was installed. +### Step 2d: Backlog Check + +Check for outstanding design work that may need attention: +1. Read `~/.automaton/design/loops/BACKLOG.md` for loop engineering work queue items. +2. Read `~/.automaton/design/framework/BACKLOG.md` for framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). +3. Note in the onboarding report whether there are unchecked `- [ ]` items in either backlog that the project owner may want to pick up manually or via the self-improvement loop (`work_source.area = "loops"` or `"framework"`). + ## Output Create or update the following inside {project}/.automaton/: diff --git a/scripts/automaton-cleanup.sh b/scripts/automaton-cleanup.sh index 15dcadc..4843fc3 100755 --- a/scripts/automaton-cleanup.sh +++ b/scripts/automaton-cleanup.sh @@ -1,2 +1,2 @@ #!/usr/bin/env bash -python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 +python3 "/Users/laptran/.automaton/scripts/status.py" --cleanup-done --days 7 --project "/private/var/folders/f5/yv0dzbnx47x3yp8sc_2519gh0000gn/T/pytest-of-laptran/pytest-98/test_uninstall_via_disabled_re0"