Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria
Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.
Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
This commit is contained in:
@@ -0,0 +1,33 @@
|
||||
# Framework Backlog
|
||||
|
||||
Status: **v1 draft 2026-06-25**. These items are captured for the self-improvement loop (or manual pickup) once the self-improvement loop is running and the design docs are in place.
|
||||
|
||||
This backlog mirrors the pattern of `design/loops/BACKLOG.md` but for framework-level agent features. The self-improvement loop's `work_source.area` can be set to `"framework"` to pull from here.
|
||||
|
||||
---
|
||||
|
||||
## v1 (Priority)
|
||||
|
||||
| ID | Item | Description | Dependencies |
|
||||
|---|---|---|---|
|
||||
| FW-1 | **agent-tab-real-roles** | Redesign the dashboard Agent tab: replace 4 fake `AGENT_TYPE_META` types with two sections — (1) Phase Roles (6 roles from `.agent.md`: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) showing active/inactive based on current task phases; (2) Scheduled Jobs (cleanup, loop-tick, rule-scan, rule-review) using real `job.kind` from `/api/scheduled`. New `/api/phase-roles` endpoint. | None |
|
||||
| FW-2 | **rule-proposer-agent** | Daily scheduled standalone agent that scans completed tasks since last run, reads their `BUG_REPORT.md` / `ADVERSARIAL_BUG_REPORT.md` / `VERDICT.md`, proposes new rules with concrete examples to `RULE_PROPOSALS.md`. Uses direct harness invocation (reuses `loop-runner._invoke_harness`). State in `.state.rule-scan`. Requires different LLM than the tasks it reviews (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding) |
|
||||
| FW-3 | **rule-reviewer-agent** | Monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, missing examples — writes `RULE_REVIEW.md`. Uses direct harness invocation. State in `.state.rule-review`. Must use different LLM than Rule Proposer (conflict-of-interest) — enforcement deferred until `model-divergence-enforcement` ships. | `model-divergence-enforcement` (for LLM binding), `rule-proposer-agent` |
|
||||
|
||||
---
|
||||
|
||||
## Deferred / Tier 2 (Post-v1)
|
||||
|
||||
| ID | Item | Description | Notes |
|
||||
|---|---|---|---|
|
||||
| FW-4 | **model-divergence-enforcement** | Manifest (`models.json`), mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model + `{model}` substitution, audit category, dashboard badges. This is a **manual task**, not a backlog item — tracked separately. | See separate task `model-divergence-enforcement` |
|
||||
| FW-5 | **dashboard-agent-tab-v2** | Enhance Agent tab with: role details (click → task list), schedule management (enable/disable from UI), last-run timestamps, per-agent logs. | Requires FW-1 |
|
||||
| FW-6 | **rule-enforcement-gate** | Pre-edit hook (`--can-edit --rule <name>`) that enforces rules from `.rules.md` before edits. Separate from rule agents (which only propose/review). | Requires FW-2, FW-3 |
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- **Dependency on `model-divergence-enforcement`**: FW-2 and FW-3 are designed with model-binding in their configs (per-role `model` in their schedule config), but hard-block enforcement is deferred until the model-divergence feature ships. The design docs note the dependency; implementation tasks in the backlog carry a "depends on" annotation.
|
||||
- **Self-improvement loop**: Once the self-improvement loop is running with `work_source.area = "framework"`, it will pick up items from this backlog automatically. The loop template `templates/loops/self-improvement/loop.json` already supports `work_source.area`.
|
||||
- **FW-4 is NOT in this backlog** — it's a manual parent task created via `status.py --create-task model-divergence-enforcement` and decomposed into subtasks.
|
||||
@@ -0,0 +1,67 @@
|
||||
# Framework Agent Features — Design Index
|
||||
|
||||
Status: **v1 draft 2026-06-25**. Design for three framework-level agent features: rule agents (proposer + reviewer), Agent tab redesign, and model-divergence enforcement.
|
||||
|
||||
## What this is
|
||||
|
||||
The Automaton framework's agent-feature layer: scheduled standalone agents for rule maintenance (Rule Proposer, Rule Reviewer), a redesigned dashboard Agent tab showing real phase roles instead of fake types, and a model-divergence enforcement system that prevents conflict-of-interest by ensuring different LLM roles use different models.
|
||||
|
||||
Built on top of the existing `status.py` phase machine and loop infrastructure — no second enforcement surface.
|
||||
|
||||
## Why
|
||||
|
||||
- `.rules.md` is maintained manually; failure patterns from completed tasks are not systematically captured.
|
||||
- The dashboard Agent tab shows 4 made-up agent types (`completed_task_archiver`, `single`, `audit`, `backlog`) — not the real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator).
|
||||
- Conflict-of-interest: two sessions with the same LLM playing bug finder and adversarial bug finder (or implementer and reviewer) don't yield useful results. The framework needs computational enforcement of model divergence.
|
||||
|
||||
## Documents
|
||||
|
||||
- [`functional.md`](functional.md) — what v1 does, roles, triggers, outputs, schedules, conflict-of-interest, success criteria. **Read this first.**
|
||||
- [`technical.md`](technical.md) — the implementation contract: file map, state schemas, `status.py` flags, runner flows, dashboard API changes, test coverage. **Read this if you're implementing.**
|
||||
- [`BACKLOG.md`](BACKLOG.md) — v1 work queue (3 items) and deferred items. The self-improvement loop's work queue when `work_source.area = "framework"`.
|
||||
|
||||
## v1 scope (draft)
|
||||
|
||||
Three feature areas, each with a clear boundary:
|
||||
|
||||
1. **Model-Divergence Enforcement** — Foundational layer. `models.json` manifest, mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, audit category, dashboard badges. Enables the other two features.
|
||||
|
||||
2. **Rule Agents** — Two standalone scheduled agents (NOT task lifecycle phases):
|
||||
- **Rule Proposer**: Daily scan of completed tasks → `RULE_PROPOSALS.md` with concrete examples.
|
||||
- **Rule Reviewer**: Monthly consolidation of `.rules.md` → `RULE_REVIEW.md`.
|
||||
- Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers (mirror `--install-cleanup-schedule`).
|
||||
|
||||
3. **Agent Tab Redesign** — Dashboard `/agent` view:
|
||||
- **Phase Roles section**: 6 roles from `.agent.md`, active/inactive based on current task phases.
|
||||
- **Scheduled Jobs section**: Real job kinds from `/api/scheduled` (cleanup, loop, rule-scan, rule-review), replacing 4 fake `AGENT_TYPE_META` types.
|
||||
- New `/api/phase-roles` endpoint.
|
||||
|
||||
## Locked decisions (referenced as `(Fn)` in functional/technical)
|
||||
|
||||
| ID | Decision |
|
||||
|---|---|
|
||||
| F1 | Rule agents are **standalone scheduled agents**, NOT task lifecycle phases. They don't block task completion. |
|
||||
| F2 | Rule Proposer runs **daily**; Rule Reviewer runs **monthly**. OS-native schedulers (launchd/cron/schtasks). |
|
||||
| F3 | Rule agents use **direct harness invocation** (reuse `loop-runner._invoke_harness`), not loop infrastructure. |
|
||||
| F4 | Conflict-of-interest: Rule Reviewer **must use different LLM** than Rule Proposer. Rule agents must use different LLM than tasks they review. Enforcement **deferred** until model-divergence ships (F5). |
|
||||
| F5 | Model-divergence enforcement is a **foundational layer** shipped first (manual task, not backlog). Rule agents designed with model-binding config but hard-block enforced later. |
|
||||
| F6 | Agent tab has **two sections**: Phase Roles (from `.agent.md` + task state) + Scheduled Jobs (from `/api/scheduled`). |
|
||||
| F7 | Phase Roles section needs **new `/api/phase-roles` endpoint** mapping active tasks to their phase roles. |
|
||||
| F8 | Scheduled Jobs section uses **real `job.kind`** (cleanup, loop, rule-scan, rule-review) — remove `AGENT_TYPE_META` fake types. |
|
||||
| F9 | State files: `.state.rule-scan` (last scanned task, timestamp, proposed rules), `.state.rule-review` (last run timestamp). |
|
||||
| F10 | New `status.py` flags: `--rule-scan`, `--install-rule-scan-schedule`, `--rule-review`, `--install-rule-review-schedule` (mirror `--cleanup-done` / `--install-cleanup-schedule`). |
|
||||
|
||||
## Relationship to loops
|
||||
|
||||
- `design/loops/` is the loop engineering system (state-enforced unattended work).
|
||||
- `design/framework/` is framework-level agent features (rule maintenance, observability, model governance).
|
||||
- The self-improvement loop (`templates/loops/self-improvement/`) can drive `design/framework/` work by setting `work_source.area = "framework"` — no code change needed (see `loop-runner.py:_find_work_backlog`).
|
||||
- Model-divergence enforcement (F5) is a prerequisite for hard-blocking conflict-of-interest in rule agents and loops.
|
||||
|
||||
## Post-v1
|
||||
|
||||
After v1 lands and the self-improvement loop is running:
|
||||
|
||||
- Rule enforcement gate (`--can-edit --rule <name>`) — separate feature.
|
||||
- Dashboard Agent tab v2: role details, schedule management, per-agent logs.
|
||||
- Rule agents gain comprehension-debt tracking (`last-read-sha` per rule file).
|
||||
@@ -0,0 +1,335 @@
|
||||
# Framework Agent Features — Functional Design
|
||||
|
||||
Status: v1 (draft 2026-06-25). Supersedes any prior informal discussions of rule agents or Agent tab redesign.
|
||||
|
||||
Audience: framework maintainers (currently: the human and one AI assistant). After handoff the self-improvement loop is also an audience — designs must be legible to a fresh-context LLM verifier.
|
||||
|
||||
## 1. Problem
|
||||
|
||||
Three gaps in the framework's agent layer:
|
||||
|
||||
1. **`.rules.md` is maintained manually.** Failure patterns from completed tasks (`BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md`) are not systematically captured into rules. The Self-Improvement section of `.rules.md` (lines 33-36) says "Add one rule per observed failure mode with a concrete example" and "Consolidate contradictions monthly" — but no agent does this. It relies on the human or a session agent remembering.
|
||||
|
||||
2. **The dashboard Agent tab shows fake agent types.** `AGENT_TYPE_META` in `dashboard.js:330-334` defines 4 types (`completed_task_archiver`, `single`, `audit`, `backlog`) that are loop work-source kinds, not real agent roles. The 6 real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) are not surfaced.
|
||||
|
||||
3. **No conflict-of-interest enforcement on model divergence.** `status.py:1432-1436` enforces that the reviewer session differs from the implementer session (via `.state.implementer`), but there is no enforcement that the *model* playing bug-finder differs from the model playing adversarial bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping.
|
||||
|
||||
## 2. Goals
|
||||
|
||||
v1 — **the framework maintains its own rules, surfaces real agent roles, and computationally enforces model divergence**:
|
||||
|
||||
1. **Rule Proposer**: a daily scheduled standalone agent that scans completed tasks since its last run, reads their failure artifacts, and proposes new rules with concrete examples to `RULE_PROPOSALS.md`.
|
||||
2. **Rule Reviewer**: a monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, rules missing examples — and writes `RULE_REVIEW.md`.
|
||||
3. **Agent Tab Redesign**: replace 4 fake `AGENT_TYPE_META` types with two sections — Phase Roles (6 roles from `.agent.md`, active/inactive based on current task phases) and Scheduled Jobs (real `job.kind` from `/api/scheduled`).
|
||||
4. **Model-Divergence Enforcement**: a `models.json` manifest, mode detection (single vs multi-LLM), a conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, an audit category, and dashboard badges.
|
||||
5. **Conflict-of-interest for rule agents**: Rule Reviewer must use a different LLM than Rule Proposer. Both must use different LLMs than the tasks they review. Enforcement is *designed now* but *hard-blocked only after model-divergence ships* (F4, F5).
|
||||
|
||||
## 3. Non-Goals (v1)
|
||||
|
||||
- **Rule enforcement during task execution.** Rule agents only *propose* and *review* rules; they do not block edits. A future `--can-edit --rule <name>` gate (FW-6) is separate.
|
||||
- **Auto-approving rules.** A human always reviews `RULE_PROPOSALS.md` and `RULE_REVIEW.md` before rules are merged into `.rules.md`. No auto-merge path in v1.
|
||||
- **Rule agents as task lifecycle phases.** Rule agents are *standalone scheduled agents* (F1). They do not block task completion. They are not phases in the state machine.
|
||||
- **A new agent harness.** Rule agents use direct harness invocation (reusing `loop-runner._invoke_harness`). No new runtime.
|
||||
- **Network-fetched dependencies.** New code is Python stdlib only. No new pip installs.
|
||||
- **Model capability inspection.** The framework never inspects model capability, provider, or size (D8). It only tracks *which* model fills *which* role and enforces the conflict matrix.
|
||||
|
||||
## 4. Model-Divergence Enforcement (foundational layer)
|
||||
|
||||
Shipped first as a manual task (`model-divergence-enforcement`), not a backlog item. Enables conflict-of-interest hard-blocking for rule agents and loops.
|
||||
|
||||
### 4.1 Manifest: `models.json`
|
||||
|
||||
A new file at `~/.automaton/models.json` (or `{project}/.automaton/models.json`):
|
||||
|
||||
```json
|
||||
{
|
||||
"default": "glm-4.6",
|
||||
"advised": true,
|
||||
"models": [
|
||||
{"name": "glm-4.6", "provider": "opencode", "context_window": 131072, "location": "remote"},
|
||||
{"name": "qwen3-coder", "provider": "opencode", "context_window": 131072, "location": "remote"},
|
||||
{"name": "llama-3.3-70b", "provider": "localhost", "context_window": 32768, "location": "http://localhost:8080"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
- `default`: the model used when no role-specific binding is set.
|
||||
- `advised`: if `true`, the framework prints a one-time advisory in single-LLM mode recommending a second model for conflict-of-interest roles, then goes silent.
|
||||
- `models[]`: the roster. `location` is `"remote"` or a localhost URL for probing.
|
||||
|
||||
### 4.2 Mode Detection
|
||||
|
||||
- **0-1 models** in `models.json` (or file missing) → **single-LLM mode**. Advisory once (if `advised: true`), then silent. No hard blocks.
|
||||
- **2+ models** → **multi-LLM mode**. Hard-block on conflict-matrix violations. Auto-assign next-available non-conflicting model on conflict; refuse only if no non-conflicting model exists.
|
||||
- **Missing file** → single-LLM mode (backward compatible). Existing behavior preserved.
|
||||
|
||||
### 4.3 Conflict Matrix (locked)
|
||||
|
||||
| Role | Must differ from |
|
||||
|---|---|
|
||||
| `code_review` | `implement` |
|
||||
| `bug_find` | `implement` |
|
||||
| `adversarial_bug_find` | `implement`, `bug_find` |
|
||||
| `referee` | `implement`, `bug_find`, `adversarial_bug_find` |
|
||||
| `loop-verify` | `loop-implement` |
|
||||
|
||||
`doc_review`, `code_review`, and `bug_find` are independent of each other (not conflicts). Only `bug_find` ↔ `adversarial_bug_find` conflicts (they are adversary pairs).
|
||||
|
||||
### 4.4 Auto-Assignment (multi-LLM mode)
|
||||
|
||||
1. Default model → assigned to `implement` (and `loop-implement`).
|
||||
2. On conflict, pick the next-available model from `models[]` that does not conflict.
|
||||
3. User override: `loop.json` `roles.<role>.model` or `status.py --transition --model <name>`.
|
||||
4. Refuse only if no non-conflicting model exists.
|
||||
|
||||
### 4.5 State: `.state.models`
|
||||
|
||||
Each task gets `{task}/.state.models` recording which model filled which role:
|
||||
|
||||
```json
|
||||
{"implement": "glm-4.6", "code_review": "qwen3-coder", "bug_find": "qwen3-coder", "adversarial_bug_find": "llama-3.3-70b", "referee": "llama-3.3-70b"}
|
||||
```
|
||||
|
||||
`status.py --transition --model <name>` records the model for the role being transitioned into. `--claim` in multi-LLM mode checks the conflict matrix against `.state.models` and refuses on violation.
|
||||
|
||||
### 4.6 Loop Integration
|
||||
|
||||
`loop.json` gains per-role `model` and `harness.command` with `{model}` substitution:
|
||||
|
||||
```json
|
||||
"roles": {
|
||||
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
|
||||
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
|
||||
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
|
||||
}
|
||||
```
|
||||
|
||||
`loop-runner.py:_invoke_harness` substitutes `{model}` into the harness command. `--check-gate` enforces `loop-verify` ≠ `loop-implement` model in multi-LLM mode.
|
||||
|
||||
### 4.7 Audit + Dashboard
|
||||
|
||||
- `status.py --audit` gains a `model_divergence` category: flags tasks where `.state.models` violates the conflict matrix.
|
||||
- Dashboard task cards show model badges (one per role filled).
|
||||
|
||||
## 5. Rule Proposer
|
||||
|
||||
A **standalone scheduled agent** (F1) that proposes new rules from completed-task failure patterns.
|
||||
|
||||
### 5.1 Trigger
|
||||
|
||||
Daily, via OS-native scheduler (mirrors `--install-cleanup-schedule`). `status.py --install-rule-scan-schedule [--interval 86400]` installs the schedule unit. Manual: `status.py --rule-scan`.
|
||||
|
||||
### 5.2 Inputs
|
||||
|
||||
- `.state.rule-scan`: state file tracking the last scanned task and timestamp.
|
||||
- Completed tasks (in `tasks/complete/` or tasks with `.state` phase `complete`) that were completed since the last scan.
|
||||
- For each such task: `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md` (if present).
|
||||
- Current `.rules.md` (to deduplicate against existing rules).
|
||||
|
||||
### 5.3 Output
|
||||
|
||||
`RULE_PROPOSALS.md` (at `~/.automaton/RULE_PROPOSALS.md` in framework mode, or `{project}/.automaton/RULE_PROPOSALS.md` in project mode). Format:
|
||||
|
||||
```markdown
|
||||
# Rule Proposals — {date}
|
||||
|
||||
## Proposed Rule: {title}
|
||||
**Source**: tasks/{task-name}/VERDICT.md
|
||||
**Pattern**: {one-line description of the failure mode}
|
||||
**Example**:
|
||||
{concrete code/config snippet from the task}
|
||||
**Proposed rule text**:
|
||||
{the rule as it would appear in .rules.md}
|
||||
|
||||
---
|
||||
```
|
||||
|
||||
Proposals are *appended* per run. A human reviews and merges accepted rules into `.rules.md`. No auto-merge (Non-Goal).
|
||||
|
||||
### 5.4 LLM Session
|
||||
|
||||
Direct harness invocation (F3): `status.py --rule-scan` reuses `loop-runner._invoke_harness` to spawn one LLM session with a prompt that includes the failure artifacts and current rules. The session proposes rules in the `RULE_PROPOSALS.md` format.
|
||||
|
||||
### 5.5 Conflict-of-Interest
|
||||
|
||||
The Rule Proposer's LLM must differ from the implementer + bug-hunter + adversarial-bug-hunter models of the tasks it scans. This prevents the model that made the bug from proposing the rule about its own bug.
|
||||
|
||||
**Enforcement deferred** (F4): until model-divergence ships, the Rule Proposer runs with the default model. The design includes a `model` field in its schedule config; hard-blocking activates once `models.json` exists and multi-LLM mode is detected.
|
||||
|
||||
### 5.6 State: `.state.rule-scan`
|
||||
|
||||
```json
|
||||
{"last_scan_at": "2026-06-25T10:00:00Z", "last_scanned_task": "fix-context-sizing", "proposals_count": 3}
|
||||
```
|
||||
|
||||
## 6. Rule Reviewer
|
||||
|
||||
A **standalone scheduled agent** (F1) that consolidates `.rules.md` periodically.
|
||||
|
||||
### 6.1 Trigger
|
||||
|
||||
Monthly, via OS-native scheduler. `status.py --install-rule-review-schedule [--interval 2592000]` installs the schedule unit. Manual: `status.py --rule-review`.
|
||||
|
||||
### 6.2 Inputs
|
||||
|
||||
- `.state.rule-review`: state file tracking the last run timestamp.
|
||||
- Current `.rules.md` (full file).
|
||||
- Recent `RULE_PROPOSALS.md` entries (since last review).
|
||||
- Recent completed-task summaries (last 30 days) for context on stale rules.
|
||||
|
||||
### 6.3 Output
|
||||
|
||||
`RULE_REVIEW.md` (at `~/.automaton/RULE_REVIEW.md` or project equivalent). Format:
|
||||
|
||||
```markdown
|
||||
# Rule Review — {date}
|
||||
|
||||
## Contradictions Found
|
||||
- Rule A ("...") contradicts Rule B ("..."). Suggested resolution: {merge/drop/keep A}.
|
||||
|
||||
## Stale Rules (no observed instance in last 30 days)
|
||||
- Rule C ("..."). Suggested action: drop or annotate as low-priority.
|
||||
|
||||
## Rules Missing Examples
|
||||
- Rule D ("..."). Suggested example: {from a recent task}.
|
||||
|
||||
## Merge Candidates
|
||||
- Rules E and F overlap. Suggested merged text: {...}.
|
||||
|
||||
---
|
||||
```
|
||||
|
||||
A human reviews and applies accepted changes to `.rules.md`. No auto-apply (Non-Goal).
|
||||
|
||||
### 6.4 LLM Session
|
||||
|
||||
Direct harness invocation (F3), same as Rule Proposer. One LLM session with a prompt that includes the full `.rules.md` and recent proposals.
|
||||
|
||||
### 6.5 Conflict-of-Interest
|
||||
|
||||
The Rule Reviewer's LLM **must differ from the Rule Proposer's LLM** (F4). The Proposer proposes (bias toward adding); the Reviewer consolidates (bias toward pruning). Same model = self-review = rubber-stamping.
|
||||
|
||||
**Enforcement deferred** until model-divergence ships.
|
||||
|
||||
### 6.6 State: `.state.rule-review`
|
||||
|
||||
```json
|
||||
{"last_review_at": "2026-06-25T10:00:00Z", "contradictions_found": 2, "stale_rules": 5, "merges_suggested": 1}
|
||||
```
|
||||
|
||||
## 7. Agent Tab Redesign
|
||||
|
||||
### 7.1 Current State (to be replaced)
|
||||
|
||||
`dashboard.js:330-334` defines `AGENT_TYPE_META` with 4 fake types:
|
||||
- `completed_task_archiver` — actually the cleanup scheduled job.
|
||||
- `single` — actually a loop with `work_source.kind = "single"`.
|
||||
- `audit` — actually a loop with `work_source.kind = "audit"`.
|
||||
- `backlog` — actually a loop with `work_source.kind = "backlog"`.
|
||||
|
||||
These are loop work-source kinds, not agent roles. They conflate two different concepts.
|
||||
|
||||
### 7.2 New Design: Two Sections
|
||||
|
||||
**Section 1 — Phase Roles**
|
||||
|
||||
Shows the 6 phase roles from `.agent.md` Agent Configuration:
|
||||
|
||||
| Role ID | Phases | Active when |
|
||||
|---|---|---|
|
||||
| `researcher` | research, decomposition, design, test_design | A task is in one of these phases |
|
||||
| `implementer` | implement | A task is in `implement` phase |
|
||||
| `code-reviewer` | code_review | A task is in `code_review` phase |
|
||||
| `bug-hunter` | bug_find, adversarial_bug_find | A task is in one of these phases |
|
||||
| `referee` | referee | A task is in `referee` phase |
|
||||
| `orchestrator` | new, complete, human_intervention | A task is in one of these phases |
|
||||
|
||||
Each role card shows: role icon, role label, status (Active/Idle — based on whether any task is in that role's phases), and the count of tasks in that role's phases. Clicking a role filters the task list to tasks in that role's phases.
|
||||
|
||||
**Data source**: new `/api/phase-roles` endpoint. Returns:
|
||||
|
||||
```json
|
||||
{
|
||||
"roles": [
|
||||
{"id": "researcher", "label": "Researcher", "icon": "🔬", "phases": ["research", "decomposition", "design", "test_design"], "active_tasks": 2, "status": "active"},
|
||||
{"id": "implementer", "label": "Implementer", "icon": "⚙️", "phases": ["implement"], "active_tasks": 1, "status": "active"},
|
||||
...
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Section 2 — Scheduled Jobs**
|
||||
|
||||
Shows real scheduled jobs from `/api/scheduled`, using `job.kind` (not fake agent types):
|
||||
|
||||
| Job Kind | Icon | Label | Source |
|
||||
|---|---|---|---|
|
||||
| `cleanup` | 🧹 | Cleanup Archiver | `com.automaton.cleanup` |
|
||||
| `loop` | 🔄 | Loop: {name} | `com.automaton.loop.{name}` |
|
||||
| `rule-scan` | 📝 | Rule Proposer | `com.automaton.rule-scan` (new) |
|
||||
| `rule-review` | 📋 | Rule Reviewer | `com.automaton.rule-review` (new) |
|
||||
|
||||
Each job card shows: job label, status (Enabled/Disabled/Misconfigured), next-run interval, runtime state (for loops: iteration count, halt status; for rule agents: last-scan/review timestamp). `AGENT_TYPE_META` is removed entirely; rendering uses `job.kind` directly.
|
||||
|
||||
### 7.3 Self-Documenting Names
|
||||
|
||||
Per `.rules.md` "Self-Documenting UI Names" (lines 53-64), all schedule unit names and stub filenames are self-documenting:
|
||||
- `com.automaton.rule-scan` (launchd label)
|
||||
- `automaton-rule-scan.sh` (stub filename)
|
||||
- `com.automaton.rule-review` / `automaton-rule-review.sh`
|
||||
|
||||
## 8. Schedules
|
||||
|
||||
Rule agents use OS-native schedulers, mirroring the `--install-cleanup-schedule` pattern (`status.py:2188-2260`):
|
||||
|
||||
| Agent | Flag | Default Interval | Launchd Label | Stub |
|
||||
|---|---|---|---|---|
|
||||
| Rule Proposer | `--install-rule-scan-schedule` | 86400s (daily) | `com.automaton.rule-scan` | `automaton-rule-scan.sh` |
|
||||
| Rule Reviewer | `--install-rule-review-schedule` | 2592000s (monthly) | `com.automaton.rule-review` | `automaton-rule-review.sh` |
|
||||
|
||||
Platform dispatch via `platform.system()`:
|
||||
- **Darwin**: `~/Library/LaunchAgents/com.automaton.rule-scan.plist` with `StartInterval`.
|
||||
- **Linux**: crontab line via `_install_cron_block_generic`.
|
||||
- **Windows**: `schtasks /create /tn "AutomatonRuleScan" ...`.
|
||||
|
||||
Stub scripts are 3-line bash/bat files that call `python3 status.py --rule-scan` (or `--rule-review`).
|
||||
|
||||
## 9. Conflict-of-Interest
|
||||
|
||||
### 9.1 Dependency Chain
|
||||
|
||||
```
|
||||
model-divergence-enforcement (shipped first, manual task)
|
||||
↓ enables hard-block
|
||||
rule-proposer-agent (FW-2)
|
||||
↓ conflict-of-interest
|
||||
rule-reviewer-agent (FW-3) — must differ from Proposer
|
||||
```
|
||||
|
||||
### 9.2 Design Now, Enforce Later (F4, F5)
|
||||
|
||||
- Rule agents are *designed* with `model` fields in their schedule config.
|
||||
- The design docs specify the conflict-of-interest rules.
|
||||
- Hard-block enforcement *activates* when `models.json` exists and multi-LLM mode is detected.
|
||||
- Until then, rule agents run with the default model (single-LLM mode, advisory only).
|
||||
|
||||
### 9.3 Why Different Models
|
||||
|
||||
- **Proposer vs Reviewer**: Proposer has a bias toward *adding* rules (more is better). Reviewer has a bias toward *pruning* (less is better). Same model = self-review = the proposer's rules never get pruned.
|
||||
- **Rule agent vs scanned tasks**: The model that introduced a bug should not propose the rule about its own bug — it has a blind spot for that failure mode.
|
||||
|
||||
## 10. Success Criteria for v1
|
||||
|
||||
1. **Model-divergence**: `--transition --model` records the model in `.state.models`; `--claim` refuses conflict-matrix violations in multi-LLM mode; `--audit` flags violations; dashboard shows model badges.
|
||||
2. **Rule Proposer**: `status.py --rule-scan` reads completed tasks since last scan, proposes rules to `RULE_PROPOSALS.md`, updates `.state.rule-scan`. Test with a seeded completed task containing a `VERDICT.md`.
|
||||
3. **Rule Reviewer**: `status.py --rule-review` reads `.rules.md` + recent proposals, writes `RULE_REVIEW.md`, updates `.state.rule-review`. Test with a seeded `.rules.md` containing a contradiction.
|
||||
4. **Agent Tab**: `/api/phase-roles` returns 6 roles with active-task counts; dashboard renders Phase Roles + Scheduled Jobs sections; `AGENT_TYPE_META` is removed; `job.kind` drives rendering.
|
||||
5. **Schedules**: `--install-rule-scan-schedule` and `--install-rule-review-schedule` install OS-native units with self-documenting names.
|
||||
6. **Tests**: `pytest tests/ -v` is green; new tests cover state schemas, scan flows, dashboard API, and conflict-matrix enforcement.
|
||||
7. **No regression**: pre-existing test suite passes unchanged.
|
||||
|
||||
## 11. Locked Decision Index
|
||||
|
||||
All decisions referenced by `(Fn)` above are recorded in the v1 design conversation (this session). They are non-negotiable for v1 implementation. Changes require a design doc update and a new `[unreleased]` changelog entry.
|
||||
|
||||
See `README.md` § "Locked decisions" for the full table.
|
||||
@@ -0,0 +1,536 @@
|
||||
# Framework Agent Features — Technical Design
|
||||
|
||||
Companion to `functional.md`. This file is the implementation contract: every line here is what the implementation tasks build. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update.
|
||||
|
||||
## 1. File Map (what v1 adds)
|
||||
|
||||
```
|
||||
~/.automaton/
|
||||
├── models.json # NEW — model manifest (see §2)
|
||||
├── scripts/
|
||||
│ ├── detect_models.py # NEW — probes opencode.json + localhost endpoints
|
||||
│ └── status.py # EXTENDED — new flags (see §4)
|
||||
├── prompts/
|
||||
│ ├── rule-proposer.md # NEW — Rule Proposer session prompt
|
||||
│ ├── rule-reviewer.md # NEW — Rule Reviewer session prompt
|
||||
│ └── onboarding.md # EXTENDED — Step 2e (backlog check)
|
||||
├── automaton/
|
||||
│ └── dashboard/
|
||||
│ ├── html/dashboard.js # EXTENDED — remove AGENT_TYPE_META, two-section render
|
||||
│ └── ui/app.py # EXTENDED — /api/phase-roles endpoint
|
||||
├── .automaton/ # (framework self-hosting: this is ~/.automaton/.automaton/)
|
||||
│ ├── .state.rule-scan # NEW — Rule Proposer state (see §3)
|
||||
│ ├── .state.rule-review # NEW — Rule Reviewer state (see §3)
|
||||
│ ├── RULE_PROPOSALS.md # NEW — Rule Proposer output (append-per-run)
|
||||
│ ├── RULE_REVIEW.md # NEW — Rule Reviewer output (append-per-run)
|
||||
│ └── automaton-rule-scan.sh # NEW — generated by --install-rule-scan-schedule
|
||||
│ automaton-rule-review.sh # NEW — generated by --install-rule-review-schedule
|
||||
└── tests/
|
||||
├── test_model_divergence.py # NEW — manifest, conflict matrix, --transition --model
|
||||
├── test_rule_agents.py # NEW — scan flows, state files, output schemas
|
||||
└── test_dashboard_phase_roles.py # NEW — /api/phase-roles, two-section render
|
||||
```
|
||||
|
||||
Per-project paths mirror the loop convention: `{project}/.automaton/.state.rule-scan`, `{project}/.automaton/RULE_PROPOSALS.md`, etc. For framework self-hosting, the project is `~/.automaton/` itself.
|
||||
|
||||
## 2. `models.json` Schema
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 1,
|
||||
"default": "glm-4.6",
|
||||
"advised": true,
|
||||
"models": [
|
||||
{
|
||||
"name": "glm-4.6",
|
||||
"provider": "opencode",
|
||||
"context_window": 131072,
|
||||
"location": "remote"
|
||||
},
|
||||
{
|
||||
"name": "qwen3-coder",
|
||||
"provider": "opencode",
|
||||
"context_window": 131072,
|
||||
"location": "remote"
|
||||
},
|
||||
{
|
||||
"name": "llama-3.3-70b",
|
||||
"provider": "localhost",
|
||||
"context_window": 32768,
|
||||
"location": "http://localhost:8080"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
- `default`: model name used when no role-specific binding exists. Must be present in `models[]`.
|
||||
- `advised`: bool. If `true`, single-LLM mode prints a one-time advisory recommending a second model, then goes silent.
|
||||
- `models[]`: roster. `name` is the unique key. `provider` is informational. `context_window` is informational (framework never inspects capability, D8). `location` is `"remote"` or a localhost URL (for `detect_models.py` probing).
|
||||
- **Missing file** → single-LLM mode (backward compatible). All model-divergence commands are no-ops.
|
||||
- **0-1 models** → single-LLM mode. Advisory once if `advised: true`.
|
||||
- **2+ models** → multi-LLM mode. Hard-block on conflict matrix.
|
||||
|
||||
### 2.1 `detect_models.py`
|
||||
|
||||
```
|
||||
python3 scripts/detect_models.py [--json]
|
||||
```
|
||||
|
||||
1. Parse `opencode.json` (or `opencode.jsonc`) for provider+model entries.
|
||||
2. Probe localhost endpoints: `http://localhost:8080/v1/models`, `http://localhost:11434/api/tags` (Ollama), `http://localhost:1234/v1/models` (LM Studio), `http://localhost:8000/v1/models` (vLLM).
|
||||
3. Merge results, emit a candidate `models.json` to stdout (or write if `--json` not set).
|
||||
4. Used by `install.sh` / `update.sh` / `upgrade.sh` to bootstrap or refresh `models.json`.
|
||||
|
||||
## 3. State Schemas
|
||||
|
||||
### 3.1 `.state.rule-scan`
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 1,
|
||||
"last_scan_at": "2026-06-25T10:00:00Z",
|
||||
"last_scanned_task": "fix-context-sizing",
|
||||
"proposals_count": 3,
|
||||
"scanned_tasks_count": 12
|
||||
}
|
||||
```
|
||||
|
||||
- `last_scanned_task`: the most recent task name scanned. Next scan starts after this task (alphabetical or mtime order).
|
||||
- `proposals_count`: cumulative count of proposals written to `RULE_PROPOSALS.md`.
|
||||
- Stored at `{project}/.automaton/.state.rule-scan`. Missing file → first run scans all completed tasks.
|
||||
|
||||
### 3.2 `.state.rule-review`
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 1,
|
||||
"last_review_at": "2026-06-25T10:00:00Z",
|
||||
"contradictions_found": 2,
|
||||
"stale_rules": 5,
|
||||
"merges_suggested": 1
|
||||
}
|
||||
```
|
||||
|
||||
- Stored at `{project}/.automaton/.state.rule-review`. Missing file → first run reviews all rules.
|
||||
|
||||
### 3.3 `.state.models` (per-task)
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 1,
|
||||
"implement": "glm-4.6",
|
||||
"code_review": "qwen3-coder",
|
||||
"bug_find": "qwen3-coder",
|
||||
"adversarial_bug_find": "llama-3.3-70b",
|
||||
"referee": "llama-3.3-70b",
|
||||
"doc_review": null
|
||||
}
|
||||
```
|
||||
|
||||
- Stored at `{task}/.state.models`. One file per task.
|
||||
- Written by `--transition --model <name>` when entering a phase.
|
||||
- Read by `--claim` (conflict-matrix check) and `--audit` (violation detection).
|
||||
- Roles not yet filled are `null` or absent.
|
||||
|
||||
## 4. `status.py` New Flags
|
||||
|
||||
All model-divergence and rule-agent commands route through `status.py` — no second enforcement surface.
|
||||
|
||||
```
|
||||
# Model-divergence
|
||||
status.py --transition <phase> --task <t> [--model <name>] Records model in .state.models; checks conflict matrix
|
||||
status.py --claim --task <t> --agent <a> [--model <name>] Refuses if model conflicts with filled roles (multi-LLM mode)
|
||||
status.py --audit EXTENDED — +model_divergence category
|
||||
status.py --can-edit [...] UNCHANGED
|
||||
|
||||
# Rule agents
|
||||
status.py --rule-scan [--project <p>] [--dry-run] Scan completed tasks, propose rules to RULE_PROPOSALS.md
|
||||
status.py --install-rule-scan-schedule [--interval S] Install OS-native unit for --rule-scan (default daily)
|
||||
status.py --rule-review [--project <p>] [--dry-run] Consolidate .rules.md, write RULE_REVIEW.md
|
||||
status.py --install-rule-review-schedule [--interval S] Install OS-native unit for --rule-review (default monthly)
|
||||
```
|
||||
|
||||
### 4.1 `--transition --model` flow
|
||||
|
||||
1. Load `models.json`. If missing or single-LLM mode → record model (advisory), no conflict check.
|
||||
2. If multi-LLM mode: load `.state.models` for the task. Check the role being entered against the conflict matrix (§5).
|
||||
3. If `--model` not provided: auto-assign next-available non-conflicting model from `models[]`. Refuse if none available.
|
||||
4. If `--model` provided: verify it's in `models[]`. Check conflict matrix. Refuse on violation.
|
||||
5. Write `role: model` to `.state.models`. Transition the phase.
|
||||
|
||||
### 4.2 `--claim --model` flow
|
||||
|
||||
1. Load `models.json`. If single-LLM mode → existing claim logic, no model check.
|
||||
2. If multi-LLM mode: load `.state.models`. Determine the role for the phase being claimed. Check conflict matrix against already-filled roles.
|
||||
3. Refuse if the claiming agent's model conflicts. Error message names the conflicting role and model.
|
||||
|
||||
### 4.3 `--audit` extension
|
||||
|
||||
New audit category `model_divergence`:
|
||||
- For each task with `.state.models`: check all filled roles against the conflict matrix.
|
||||
- Flag violations as `severity: high` (conflict-of-interest is a correctness issue, not a style issue).
|
||||
- Output format mirrors existing audit categories.
|
||||
|
||||
## 5. Conflict Matrix (implementation)
|
||||
|
||||
```python
|
||||
CONFLICT_MATRIX = {
|
||||
"code_review": {"implement"},
|
||||
"bug_find": {"implement"},
|
||||
"adversarial_bug_find": {"implement", "bug_find"},
|
||||
"referee": {"implement", "bug_find", "adversarial_bug_find"},
|
||||
"loop-verify": {"loop-implement"},
|
||||
}
|
||||
```
|
||||
|
||||
- Key = role being entered. Value = set of roles that must have a different model.
|
||||
- `doc_review`, `code_review`, `bug_find` are NOT in conflict with each other (only `bug_find` ↔ `adversarial_bug_find` conflicts).
|
||||
- Check function: `def _check_conflict(state_models: dict, role: str, model: str, matrix: dict) -> Optional[str]` — returns the conflicting role name or `None`.
|
||||
|
||||
## 6. Rule Proposer Flow (`--rule-scan`)
|
||||
|
||||
```
|
||||
1. Load .state.rule-scan (or init if missing).
|
||||
2. Find completed tasks since last_scanned_task:
|
||||
- Scan tasks/complete/ and tasks with .state phase=complete
|
||||
- Filter by mtime > last_scan_at (or all if first run)
|
||||
- Sort by mtime ascending
|
||||
3. For each task:
|
||||
a. Read BUG_REPORT.md, ADVERSARIAL_BUG_REPORT.md, VERDICT.md (skip if none exist)
|
||||
b. Read current .rules.md (for dedup context — capped at 4k tokens)
|
||||
c. Build proposer prompt (see §7)
|
||||
d. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_proposer_model)
|
||||
e. Parse LLM output for proposed rules (expect RULE_PROPOSALS.md format)
|
||||
f. Append proposals to RULE_PROPOSALS.md
|
||||
g. Update .state.rule-scan (last_scanned_task, proposals_count)
|
||||
4. Write final .state.rule-scan with last_scan_at = now.
|
||||
```
|
||||
|
||||
- `--dry-run`: list tasks that would be scanned, do not invoke harness.
|
||||
- `--project`: scope to a project (default: framework dir).
|
||||
- Errors during a single task scan do not abort the run; the scan continues to the next task and logs the error.
|
||||
|
||||
## 7. Rule Proposer Prompt Shape (`rule-proposer.md`)
|
||||
|
||||
```
|
||||
# Rule Proposer — {date}
|
||||
|
||||
You are scanning completed tasks for failure patterns that should become rules.
|
||||
|
||||
## Current rules (read-only, for dedup)
|
||||
{current_rules} # .rules.md content, capped at 4k tokens
|
||||
|
||||
## Task failure artifacts
|
||||
{bug_report} # BUG_REPORT.md content, capped at 2k tokens
|
||||
{adversarial_report} # ADVERSARIAL_BUG_REPORT.md, capped at 2k tokens
|
||||
{verdict} # VERDICT.md, capped at 2k tokens
|
||||
|
||||
## What to do
|
||||
For each distinct failure pattern you observe:
|
||||
1. Check if a rule already exists in .rules.md that covers it. If so, skip.
|
||||
2. If no existing rule covers it, propose a new rule with:
|
||||
- A concrete example from the task artifacts
|
||||
- The proposed rule text as it would appear in .rules.md
|
||||
|
||||
## Output (strict markdown, no JSON)
|
||||
## Proposed Rule: {title}
|
||||
**Source**: tasks/{task-name}/VERDICT.md
|
||||
**Pattern**: {one-line description}
|
||||
**Example**:
|
||||
{concrete snippet}
|
||||
**Proposed rule text**:
|
||||
{rule text}
|
||||
|
||||
---
|
||||
```
|
||||
|
||||
No `{model}` token in the prompt — the model is selected by the caller and passed to `_invoke_harness`.
|
||||
|
||||
## 8. Rule Reviewer Flow (`--rule-review`)
|
||||
|
||||
```
|
||||
1. Load .state.rule-review (or init if missing).
|
||||
2. Read .rules.md (full file).
|
||||
3. Read recent RULE_PROPOSALS.md entries (since last_review_at).
|
||||
4. Read recent completed-task summaries (last 30 days) for staleness context.
|
||||
5. Build reviewer prompt (see §9).
|
||||
6. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_reviewer_model).
|
||||
7. Parse LLM output for review sections (contradictions, stale, missing examples, merges).
|
||||
8. Append to RULE_REVIEW.md.
|
||||
9. Update .state.rule-review.
|
||||
```
|
||||
|
||||
- `--dry-run`: report what would be reviewed, do not invoke harness.
|
||||
|
||||
## 9. Rule Reviewer Prompt Shape (`rule-reviewer.md`)
|
||||
|
||||
```
|
||||
# Rule Reviewer — {date}
|
||||
|
||||
You are consolidating .rules.md for contradictions, staleness, and missing examples.
|
||||
|
||||
## Current rules (full)
|
||||
{rules_content} # .rules.md, full file
|
||||
|
||||
## Recent proposals (since last review)
|
||||
{recent_proposals} # RULE_PROPOSALS.md entries since last_review_at
|
||||
|
||||
## Recent completed tasks (last 30 days, for staleness context)
|
||||
{task_summaries} # one-line per task: name + phase + completion date
|
||||
|
||||
## What to check
|
||||
1. Contradictions: rules that conflict with each other.
|
||||
2. Stale rules: no observed instance in last 30 days.
|
||||
3. Rules missing examples: any rule without a concrete example.
|
||||
4. Merge candidates: overlapping rules that could be consolidated.
|
||||
|
||||
## Output (strict markdown, no JSON)
|
||||
## Contradictions Found
|
||||
- ...
|
||||
## Stale Rules
|
||||
- ...
|
||||
## Rules Missing Examples
|
||||
- ...
|
||||
## Merge Candidates
|
||||
- ...
|
||||
```
|
||||
|
||||
## 10. Harness Invocation (direct, not loop)
|
||||
|
||||
Rule agents reuse `loop-runner._invoke_harness` directly — they are NOT loops. The function signature (from `loop-runner.py:366-404`):
|
||||
|
||||
```python
|
||||
def _invoke_harness(harness_command: str, prompt_content: str, cwd: str, env: dict = None) -> str:
|
||||
```
|
||||
|
||||
`status.py --rule-scan` calls this as:
|
||||
|
||||
```python
|
||||
from loop_runner import _invoke_harness
|
||||
output = _invoke_harness(
|
||||
harness_command=rule_harness_command, # from schedule config or default
|
||||
prompt_content=resolved_prompt, # rule-proposer.md with tokens substituted
|
||||
cwd=str(project_dir),
|
||||
env={"AUTOMATON_RULE_ROLE": "proposer"}
|
||||
)
|
||||
```
|
||||
|
||||
`{model}` substitution: if the harness command contains `{model}`, it's replaced with the rule agent's configured model. Until model-divergence ships, this is the default model.
|
||||
|
||||
### 10.1 Schedule Config for Rule Agents
|
||||
|
||||
Rule agents do not use `loop.json`. Their config is embedded in the schedule stub:
|
||||
|
||||
```bash
|
||||
#!/usr/bin/env bash
|
||||
cd "<project_root>"
|
||||
python3 "<framework>/scripts/status.py" --rule-scan --model <name>
|
||||
```
|
||||
|
||||
The `--model` flag is optional and ignored in single-LLM mode. In multi-LLM mode it sets the rule agent's model (subject to conflict-of-interest checks once enforced).
|
||||
|
||||
## 11. Scheduler Unit Generation
|
||||
|
||||
Mirrors `cmd_install_cleanup_schedule` (`status.py:2188-2260`) exactly:
|
||||
|
||||
### 11.1 `--install-rule-scan-schedule`
|
||||
|
||||
```python
|
||||
def cmd_install_rule_scan_schedule(args) -> int:
|
||||
interval = args.interval if args.interval else 86400 # daily
|
||||
# 1. Write stub: automaton-rule-scan.sh
|
||||
# 2. Platform dispatch:
|
||||
# Darwin → ~/Library/LaunchAgents/com.automaton.rule-scan.plist
|
||||
# Linux → crontab block via _install_cron_block_generic
|
||||
# Windows → schtasks /create /tn "AutomatonRuleScan"
|
||||
```
|
||||
|
||||
### 11.2 `--install-rule-review-schedule`
|
||||
|
||||
```python
|
||||
def cmd_install_rule_review_schedule(args) -> int:
|
||||
interval = args.interval if args.interval else 2592000 # monthly
|
||||
# Same pattern, labels: com.automaton.rule-review / AutomatonRuleReview
|
||||
```
|
||||
|
||||
### 11.3 `_list_scheduled_jobs` extension
|
||||
|
||||
`_list_scheduled_jobs` (`status.py:2282`) gains recognition for new labels:
|
||||
|
||||
```python
|
||||
def _launchd_label_kind(label: str) -> tuple[str, str]:
|
||||
if label.startswith("com.automaton.loop."):
|
||||
return ("loop", label[len("com.automaton.loop."):])
|
||||
if label in ("com.automaton.cleanup",):
|
||||
return ("cleanup", "")
|
||||
if label in ("com.automaton.rule-scan",):
|
||||
return ("rule-scan", "")
|
||||
if label in ("com.automaton.rule-review",):
|
||||
return ("rule-review", "")
|
||||
...
|
||||
```
|
||||
|
||||
This makes rule-scan and rule-review jobs appear in `/api/scheduled` with their real `kind`, which the Agent tab renders directly.
|
||||
|
||||
## 12. Agent Tab Data Flow
|
||||
|
||||
### 12.1 New endpoint: `/api/phase-roles`
|
||||
|
||||
`app.py` gains a handler:
|
||||
|
||||
```python
|
||||
elif self.path == "/api/phase-roles":
|
||||
self._serve_phase_roles()
|
||||
```
|
||||
|
||||
```python
|
||||
def _serve_phase_roles(self):
|
||||
# 1. Parse .agent.md Agent Configuration for role definitions
|
||||
# 2. Load all tasks via status.py module
|
||||
# 3. For each role, count tasks in that role's phases
|
||||
# 4. Return JSON:
|
||||
{
|
||||
"roles": [
|
||||
{"id": "researcher", "label": "Researcher", "icon": "🔬",
|
||||
"phases": ["research", "decomposition", "design", "test_design"],
|
||||
"active_tasks": 2, "status": "active"},
|
||||
...
|
||||
],
|
||||
"available": True
|
||||
}
|
||||
```
|
||||
|
||||
Role icons (self-documenting, per `.rules.md` Self-Documenting UI Names):
|
||||
|
||||
| Role | Icon |
|
||||
|---|---|
|
||||
| researcher | 🔬 |
|
||||
| implementer | ⚙️ |
|
||||
| code-reviewer | 👁️ |
|
||||
| bug-hunter | 🐛 |
|
||||
| referee | ⚖️ |
|
||||
| orchestrator | 🎯 |
|
||||
|
||||
### 12.2 `dashboard.js` changes
|
||||
|
||||
**Remove**: `AGENT_TYPE_META` (lines 330-334), `AGENT_TYPES` (337), `AGENT_TYPE_META_FALLBACK` (338), `_resolveAgentType` (351-354).
|
||||
|
||||
**Replace `renderAgentTab`** with a two-section render:
|
||||
|
||||
```javascript
|
||||
async function renderAgentTab() {
|
||||
const panel = document.getElementById('agent-panel');
|
||||
panel.innerHTML = '<div class="bg-loading">Loading…</div>';
|
||||
|
||||
const [rolesRes, schedRes] = await Promise.all([
|
||||
fetch('/api/phase-roles').then(r => r.json()).catch(() => ({roles: [], available: false})),
|
||||
fetchSchedule(),
|
||||
]);
|
||||
|
||||
// Section 1: Phase Roles
|
||||
const rolesHtml = rolesRes.available ? renderPhaseRoles(rolesRes.roles)
|
||||
: '<div class="bg-empty">Phase roles require .agent.md Agent Configuration.</div>';
|
||||
|
||||
// Section 2: Scheduled Jobs
|
||||
const jobsHtml = renderScheduledJobs(schedRes.jobs || []);
|
||||
|
||||
panel.innerHTML = `
|
||||
<div class="agent-section">
|
||||
<h3>Phase Roles</h3>
|
||||
<div class="bg-grid">${rolesHtml}</div>
|
||||
</div>
|
||||
<div class="agent-section">
|
||||
<h3>Scheduled Jobs</h3>
|
||||
<div class="bg-grid">${jobsHtml}</div>
|
||||
</div>`;
|
||||
}
|
||||
```
|
||||
|
||||
**`renderScheduledJobs`** uses `job.kind` directly (no fake type resolution):
|
||||
|
||||
```javascript
|
||||
const JOB_META = {
|
||||
cleanup: { icon: '🧹', label: 'Cleanup Archiver' },
|
||||
loop: { icon: '🔄', label: (j) => `Loop: ${j.name}` },
|
||||
'rule-scan': { icon: '📝', label: 'Rule Proposer' },
|
||||
'rule-review':{ icon: '📋', label: 'Rule Reviewer' },
|
||||
};
|
||||
```
|
||||
|
||||
## 13. Loop Integration (model-divergence)
|
||||
|
||||
### 13.1 `loop.json` per-role model
|
||||
|
||||
```json
|
||||
"roles": {
|
||||
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
|
||||
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
|
||||
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
|
||||
}
|
||||
```
|
||||
|
||||
- `model` is optional. If absent, uses `models.json` `default`.
|
||||
- `loop-verify` model is checked against `loop-implement` model in `--check-gate` (multi-LLM mode).
|
||||
|
||||
### 13.2 `{model}` substitution in `_invoke_harness`
|
||||
|
||||
`loop-runner.py:366-404` `_invoke_harness` gains `{model}` token substitution:
|
||||
|
||||
```python
|
||||
def _invoke_harness(harness_command, prompt_content, cwd, env=None, model=None):
|
||||
if model and "{model}" in harness_command:
|
||||
harness_command = harness_command.replace("{model}", model)
|
||||
...
|
||||
```
|
||||
|
||||
The caller passes `model` from the role config. If the harness command has no `{model}` token, the model is informational only (the harness picks its own).
|
||||
|
||||
### 13.3 `--check-gate` model-divergence check
|
||||
|
||||
In multi-LLM mode, `--check-gate` adds:
|
||||
- Load `loop.json` roles. Compare `verify.model` vs `implement.model`.
|
||||
- If same model and multi-LLM mode → halt as `model_conflict` (new halt reason, or reuse `human_intervention` with a descriptive message).
|
||||
|
||||
## 14. Test Coverage
|
||||
|
||||
### 14.1 `test_model_divergence.py`
|
||||
|
||||
- `test_models_json_missing_single_llm_mode` — no file → advisory, no blocks.
|
||||
- `test_single_model_advisory_once` — 1 model, `advised: true` → advisory printed once, then silent.
|
||||
- `test_multi_llm_conflict_matrix` — 2+ models, `--transition --model` records, `--claim` refuses conflict.
|
||||
- `test_auto_assign_next_available` — no `--model` flag → auto-assigns non-conflicting model.
|
||||
- `test_auto_assign_exhausted` — all models conflict → refuse.
|
||||
- `test_audit_model_divergence` — `--audit` flags conflict-matrix violations.
|
||||
- `test_loop_verify_neq_implement` — `--check-gate` halts on same model in multi-LLM mode.
|
||||
|
||||
### 14.2 `test_rule_agents.py`
|
||||
|
||||
- `test_rule_scan_finds_completed_tasks` — seeded completed task with VERDICT.md → proposal written.
|
||||
- `test_rule_scan_state_tracking` — `.state.rule-scan` updated with last_scanned_task + count.
|
||||
- `test_rule_scan_dedup` — existing rule in `.rules.md` → not re-proposed.
|
||||
- `test_rule_scan_dry_run` — no harness invocation, lists candidates.
|
||||
- `test_rule_review_finds_contradictions` — seeded `.rules.md` with contradiction → review written.
|
||||
- `test_rule_review_state_tracking` — `.state.rule-review` updated.
|
||||
- `test_install_rule_scan_schedule` — stub + plist created with correct labels.
|
||||
- `test_install_rule_review_schedule` — stub + plist created with correct labels.
|
||||
|
||||
### 14.3 `test_dashboard_phase_roles.py`
|
||||
|
||||
- `test_api_phase_roles` — `/api/phase-roles` returns 6 roles with correct phases.
|
||||
- `test_phase_roles_active_count` — tasks in phases → correct active_tasks count.
|
||||
- `test_scheduled_jobs_new_kinds` — rule-scan and rule-review jobs appear with correct `kind`.
|
||||
- `test_agent_type_meta_removed` — `AGENT_TYPE_META` no longer in dashboard.js (grep test).
|
||||
|
||||
## 15. Rollout (3 sequential tasks for model-divergence)
|
||||
|
||||
The model-divergence-enforcement parent task decomposes into 3 subtasks:
|
||||
|
||||
1. **manifest+detection**: `models.json` schema, `detect_models.py`, `install.sh`/`update.sh`/`upgrade.sh` integration, `config.md` section, onboarding Step 2d.
|
||||
2. **interactive enforcement+audit**: `.state.models`, `--transition --model`, `--claim --model` conflict check, `--audit` model_divergence category, dashboard badges.
|
||||
3. **loop enforcement+dashboard**: `loop.json` per-role model, `{model}` substitution, `--check-gate` model check, loop dashboard badges.
|
||||
|
||||
Rule agents (FW-2, FW-3) and Agent tab (FW-1) are backlog items, picked up after model-divergence ships (for FW-2/FW-3) or independently (for FW-1).
|
||||
|
||||
## 16. Locked Decision Index
|
||||
|
||||
All decisions referenced by `(Fn)` are in `README.md` § "Locked decisions". Implementation must conform. Deviations require a design doc update + `[unreleased]` CHANGELOG entry.
|
||||
@@ -6,6 +6,8 @@ Work queue for the self-improvement loop after task 7 lands. Items not assigned
|
||||
|
||||
A loop configured with `work_source.kind = "backlog"` reads this file, picks the topmost `[ ]` item, drafts an implementation, transitions through phases, hands off to a human reviewer. Mark items `[x]` when complete; move items to `DONE.md` (created later) on closure.
|
||||
|
||||
**Sibling backlog**: `design/framework/BACKLOG.md` covers framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). A loop with `work_source.area = "framework"` reads that file instead of this one.
|
||||
|
||||
## v1.1 — framework manages its own docs
|
||||
|
||||
- [ ] **design-update-loop-template** — `templates/loops/design-update/` that keeps `design/<area>/*.md` in sync with the code it documents. Triggered by `last-read-sha` drift detection.
|
||||
|
||||
@@ -15,6 +15,7 @@ Per-session manual driving doesn't scale against the framework's growing backlog
|
||||
- [`functional.md`](functional.md) — what v1 does, roles, the five deaths, blast radius, schedules, success criteria. **Read this first.**
|
||||
- [`technical.md`](technical.md) — the implementation contract: file map, `.state.loop` schema, `status.py` flags, gate checks, runner flow, test coverage. **Read this if you're implementing.**
|
||||
- [`BACKLOG.md`](BACKLOG.md) — v1.1 and deferred items (Scope 2 design-update loop, Scope 3 self-designing, parallel mode, dashboard panel). The self-improvement loop's work queue.
|
||||
- **Sibling design**: [`../framework/`](../framework/) — framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). Has its own `BACKLOG.md` consumable via `work_source.area = "framework"`.
|
||||
|
||||
## v1 scope (locked)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user