336 lines
17 KiB
Markdown
336 lines
17 KiB
Markdown
# Framework Agent Features — Functional Design
|
|||
|
|
|
||
|
|
Status: v1 (draft 2026-06-25). Supersedes any prior informal discussions of rule agents or Agent tab redesign.
|
||
|
|
|
||
|
|
Audience: framework maintainers (currently: the human and one AI assistant). After handoff the self-improvement loop is also an audience — designs must be legible to a fresh-context LLM verifier.
|
||
|
|
|
||
|
|
## 1. Problem
|
||
|
|
|
||
|
|
Three gaps in the framework's agent layer:
|
||
|
|
|
||
|
|
1. **`.rules.md` is maintained manually.** Failure patterns from completed tasks (`BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md`) are not systematically captured into rules. The Self-Improvement section of `.rules.md` (lines 33-36) says "Add one rule per observed failure mode with a concrete example" and "Consolidate contradictions monthly" — but no agent does this. It relies on the human or a session agent remembering.
|
||
|
|
|
||
|
|
2. **The dashboard Agent tab shows fake agent types.** `AGENT_TYPE_META` in `dashboard.js:330-334` defines 4 types (`completed_task_archiver`, `single`, `audit`, `backlog`) that are loop work-source kinds, not real agent roles. The 6 real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) are not surfaced.
|
||
|
|
|
||
|
|
3. **No conflict-of-interest enforcement on model divergence.** `status.py:1432-1436` enforces that the reviewer session differs from the implementer session (via `.state.implementer`), but there is no enforcement that the *model* playing bug-finder differs from the model playing adversarial bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping.
|
||
|
|
|
||
|
|
## 2. Goals
|
||
|
|
|
||
|
|
v1 — **the framework maintains its own rules, surfaces real agent roles, and computationally enforces model divergence**:
|
||
|
|
|
||
|
|
1. **Rule Proposer**: a daily scheduled standalone agent that scans completed tasks since its last run, reads their failure artifacts, and proposes new rules with concrete examples to `RULE_PROPOSALS.md`.
|
||
|
|
2. **Rule Reviewer**: a monthly scheduled standalone agent that consolidates `.rules.md` — finds contradictions, stale rules, rules missing examples — and writes `RULE_REVIEW.md`.
|
||
|
|
3. **Agent Tab Redesign**: replace 4 fake `AGENT_TYPE_META` types with two sections — Phase Roles (6 roles from `.agent.md`, active/inactive based on current task phases) and Scheduled Jobs (real `job.kind` from `/api/scheduled`).
|
||
|
|
4. **Model-Divergence Enforcement**: a `models.json` manifest, mode detection (single vs multi-LLM), a conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, an audit category, and dashboard badges.
|
||
|
|
5. **Conflict-of-interest for rule agents**: Rule Reviewer must use a different LLM than Rule Proposer. Both must use different LLMs than the tasks they review. Enforcement is *designed now* but *hard-blocked only after model-divergence ships* (F4, F5).
|
||
|
|
|
||
|
|
## 3. Non-Goals (v1)
|
||
|
|
|
||
|
|
- **Rule enforcement during task execution.** Rule agents only *propose* and *review* rules; they do not block edits. A future `--can-edit --rule <name>` gate (FW-6) is separate.
|
||
|
|
- **Auto-approving rules.** A human always reviews `RULE_PROPOSALS.md` and `RULE_REVIEW.md` before rules are merged into `.rules.md`. No auto-merge path in v1.
|
||
|
|
- **Rule agents as task lifecycle phases.** Rule agents are *standalone scheduled agents* (F1). They do not block task completion. They are not phases in the state machine.
|
||
|
|
- **A new agent harness.** Rule agents use direct harness invocation (reusing `loop-runner._invoke_harness`). No new runtime.
|
||
|
|
- **Network-fetched dependencies.** New code is Python stdlib only. No new pip installs.
|
||
|
|
- **Model capability inspection.** The framework never inspects model capability, provider, or size (D8). It only tracks *which* model fills *which* role and enforces the conflict matrix.
|
||
|
|
|
||
|
|
## 4. Model-Divergence Enforcement (foundational layer)
|
||
|
|
|
||
|
|
Shipped first as a manual task (`model-divergence-enforcement`), not a backlog item. Enables conflict-of-interest hard-blocking for rule agents and loops.
|
||
|
|
|
||
|
|
### 4.1 Manifest: `models.json`
|
||
|
|
|
||
|
|
A new file at `~/.automaton/models.json` (or `{project}/.automaton/models.json`):
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"default": "glm-4.6",
|
||
|
|
"advised": true,
|
||
|
|
"models": [
|
||
|
|
{"name": "glm-4.6", "provider": "opencode", "context_window": 131072, "location": "remote"},
|
||
|
|
{"name": "qwen3-coder", "provider": "opencode", "context_window": 131072, "location": "remote"},
|
||
|
|
{"name": "llama-3.3-70b", "provider": "localhost", "context_window": 32768, "location": "http://localhost:8080"}
|
||
|
|
]
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
- `default`: the model used when no role-specific binding is set.
|
||
|
|
- `advised`: if `true`, the framework prints a one-time advisory in single-LLM mode recommending a second model for conflict-of-interest roles, then goes silent.
|
||
|
|
- `models[]`: the roster. `location` is `"remote"` or a localhost URL for probing.
|
||
|
|
|
||
|
|
### 4.2 Mode Detection
|
||
|
|
|
||
|
|
- **0-1 models** in `models.json` (or file missing) → **single-LLM mode**. Advisory once (if `advised: true`), then silent. No hard blocks.
|
||
|
|
- **2+ models** → **multi-LLM mode**. Hard-block on conflict-matrix violations. Auto-assign next-available non-conflicting model on conflict; refuse only if no non-conflicting model exists.
|
||
|
|
- **Missing file** → single-LLM mode (backward compatible). Existing behavior preserved.
|
||
|
|
|
||
|
|
### 4.3 Conflict Matrix (locked)
|
||
|
|
|
||
|
|
| Role | Must differ from |
|
||
|
|
|---|---|
|
||
|
|
| `code_review` | `implement` |
|
||
|
|
| `bug_find` | `implement` |
|
||
|
|
| `adversarial_bug_find` | `implement`, `bug_find` |
|
||
|
|
| `referee` | `implement`, `bug_find`, `adversarial_bug_find` |
|
||
|
|
| `loop-verify` | `loop-implement` |
|
||
|
|
|
||
|
|
`doc_review`, `code_review`, and `bug_find` are independent of each other (not conflicts). Only `bug_find` ↔ `adversarial_bug_find` conflicts (they are adversary pairs).
|
||
|
|
|
||
|
|
### 4.4 Auto-Assignment (multi-LLM mode)
|
||
|
|
|
||
|
|
1. Default model → assigned to `implement` (and `loop-implement`).
|
||
|
|
2. On conflict, pick the next-available model from `models[]` that does not conflict.
|
||
|
|
3. User override: `loop.json` `roles.<role>.model` or `status.py --transition --model <name>`.
|
||
|
|
4. Refuse only if no non-conflicting model exists.
|
||
|
|
|
||
|
|
### 4.5 State: `.state.models`
|
||
|
|
|
||
|
|
Each task gets `{task}/.state.models` recording which model filled which role:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{"implement": "glm-4.6", "code_review": "qwen3-coder", "bug_find": "qwen3-coder", "adversarial_bug_find": "llama-3.3-70b", "referee": "llama-3.3-70b"}
|
||
|
|
```
|
||
|
|
|
||
|
|
`status.py --transition --model <name>` records the model for the role being transitioned into. `--claim` in multi-LLM mode checks the conflict matrix against `.state.models` and refuses on violation.
|
||
|
|
|
||
|
|
### 4.6 Loop Integration
|
||
|
|
|
||
|
|
`loop.json` gains per-role `model` and `harness.command` with `{model}` substitution:
|
||
|
|
|
||
|
|
```json
|
||
|
|
"roles": {
|
||
|
|
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
|
||
|
|
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
|
||
|
|
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
`loop-runner.py:_invoke_harness` substitutes `{model}` into the harness command. `--check-gate` enforces `loop-verify` ≠ `loop-implement` model in multi-LLM mode.
|
||
|
|
|
||
|
|
### 4.7 Audit + Dashboard
|
||
|
|
|
||
|
|
- `status.py --audit` gains a `model_divergence` category: flags tasks where `.state.models` violates the conflict matrix.
|
||
|
|
- Dashboard task cards show model badges (one per role filled).
|
||
|
|
|
||
|
|
## 5. Rule Proposer
|
||
|
|
|
||
|
|
A **standalone scheduled agent** (F1) that proposes new rules from completed-task failure patterns.
|
||
|
|
|
||
|
|
### 5.1 Trigger
|
||
|
|
|
||
|
|
Daily, via OS-native scheduler (mirrors `--install-cleanup-schedule`). `status.py --install-rule-scan-schedule [--interval 86400]` installs the schedule unit. Manual: `status.py --rule-scan`.
|
||
|
|
|
||
|
|
### 5.2 Inputs
|
||
|
|
|
||
|
|
- `.state.rule-scan`: state file tracking the last scanned task and timestamp.
|
||
|
|
- Completed tasks (in `tasks/complete/` or tasks with `.state` phase `complete`) that were completed since the last scan.
|
||
|
|
- For each such task: `BUG_REPORT.md`, `ADVERSARIAL_BUG_REPORT.md`, `VERDICT.md` (if present).
|
||
|
|
- Current `.rules.md` (to deduplicate against existing rules).
|
||
|
|
|
||
|
|
### 5.3 Output
|
||
|
|
|
||
|
|
`RULE_PROPOSALS.md` (at `~/.automaton/RULE_PROPOSALS.md` in framework mode, or `{project}/.automaton/RULE_PROPOSALS.md` in project mode). Format:
|
||
|
|
|
||
|
|
```markdown
|
||
|
|
# Rule Proposals — {date}
|
||
|
|
|
||
|
|
## Proposed Rule: {title}
|
||
|
|
**Source**: tasks/{task-name}/VERDICT.md
|
||
|
|
**Pattern**: {one-line description of the failure mode}
|
||
|
|
**Example**:
|
||
|
|
{concrete code/config snippet from the task}
|
||
|
|
**Proposed rule text**:
|
||
|
|
{the rule as it would appear in .rules.md}
|
||
|
|
|
||
|
|
---
|
||
|
|
```
|
||
|
|
|
||
|
|
Proposals are *appended* per run. A human reviews and merges accepted rules into `.rules.md`. No auto-merge (Non-Goal).
|
||
|
|
|
||
|
|
### 5.4 LLM Session
|
||
|
|
|
||
|
|
Direct harness invocation (F3): `status.py --rule-scan` reuses `loop-runner._invoke_harness` to spawn one LLM session with a prompt that includes the failure artifacts and current rules. The session proposes rules in the `RULE_PROPOSALS.md` format.
|
||
|
|
|
||
|
|
### 5.5 Conflict-of-Interest
|
||
|
|
|
||
|
|
The Rule Proposer's LLM must differ from the implementer + bug-hunter + adversarial-bug-hunter models of the tasks it scans. This prevents the model that made the bug from proposing the rule about its own bug.
|
||
|
|
|
||
|
|
**Enforcement deferred** (F4): until model-divergence ships, the Rule Proposer runs with the default model. The design includes a `model` field in its schedule config; hard-blocking activates once `models.json` exists and multi-LLM mode is detected.
|
||
|
|
|
||
|
|
### 5.6 State: `.state.rule-scan`
|
||
|
|
|
||
|
|
```json
|
||
|
|
{"last_scan_at": "2026-06-25T10:00:00Z", "last_scanned_task": "fix-context-sizing", "proposals_count": 3}
|
||
|
|
```
|
||
|
|
|
||
|
|
## 6. Rule Reviewer
|
||
|
|
|
||
|
|
A **standalone scheduled agent** (F1) that consolidates `.rules.md` periodically.
|
||
|
|
|
||
|
|
### 6.1 Trigger
|
||
|
|
|
||
|
|
Monthly, via OS-native scheduler. `status.py --install-rule-review-schedule [--interval 2592000]` installs the schedule unit. Manual: `status.py --rule-review`.
|
||
|
|
|
||
|
|
### 6.2 Inputs
|
||
|
|
|
||
|
|
- `.state.rule-review`: state file tracking the last run timestamp.
|
||
|
|
- Current `.rules.md` (full file).
|
||
|
|
- Recent `RULE_PROPOSALS.md` entries (since last review).
|
||
|
|
- Recent completed-task summaries (last 30 days) for context on stale rules.
|
||
|
|
|
||
|
|
### 6.3 Output
|
||
|
|
|
||
|
|
`RULE_REVIEW.md` (at `~/.automaton/RULE_REVIEW.md` or project equivalent). Format:
|
||
|
|
|
||
|
|
```markdown
|
||
|
|
# Rule Review — {date}
|
||
|
|
|
||
|
|
## Contradictions Found
|
||
|
|
- Rule A ("...") contradicts Rule B ("..."). Suggested resolution: {merge/drop/keep A}.
|
||
|
|
|
||
|
|
## Stale Rules (no observed instance in last 30 days)
|
||
|
|
- Rule C ("..."). Suggested action: drop or annotate as low-priority.
|
||
|
|
|
||
|
|
## Rules Missing Examples
|
||
|
|
- Rule D ("..."). Suggested example: {from a recent task}.
|
||
|
|
|
||
|
|
## Merge Candidates
|
||
|
|
- Rules E and F overlap. Suggested merged text: {...}.
|
||
|
|
|
||
|
|
---
|
||
|
|
```
|
||
|
|
|
||
|
|
A human reviews and applies accepted changes to `.rules.md`. No auto-apply (Non-Goal).
|
||
|
|
|
||
|
|
### 6.4 LLM Session
|
||
|
|
|
||
|
|
Direct harness invocation (F3), same as Rule Proposer. One LLM session with a prompt that includes the full `.rules.md` and recent proposals.
|
||
|
|
|
||
|
|
### 6.5 Conflict-of-Interest
|
||
|
|
|
||
|
|
The Rule Reviewer's LLM **must differ from the Rule Proposer's LLM** (F4). The Proposer proposes (bias toward adding); the Reviewer consolidates (bias toward pruning). Same model = self-review = rubber-stamping.
|
||
|
|
|
||
|
|
**Enforcement deferred** until model-divergence ships.
|
||
|
|
|
||
|
|
### 6.6 State: `.state.rule-review`
|
||
|
|
|
||
|
|
```json
|
||
|
|
{"last_review_at": "2026-06-25T10:00:00Z", "contradictions_found": 2, "stale_rules": 5, "merges_suggested": 1}
|
||
|
|
```
|
||
|
|
|
||
|
|
## 7. Agent Tab Redesign
|
||
|
|
|
||
|
|
### 7.1 Current State (to be replaced)
|
||
|
|
|
||
|
|
`dashboard.js:330-334` defines `AGENT_TYPE_META` with 4 fake types:
|
||
|
|
- `completed_task_archiver` — actually the cleanup scheduled job.
|
||
|
|
- `single` — actually a loop with `work_source.kind = "single"`.
|
||
|
|
- `audit` — actually a loop with `work_source.kind = "audit"`.
|
||
|
|
- `backlog` — actually a loop with `work_source.kind = "backlog"`.
|
||
|
|
|
||
|
|
These are loop work-source kinds, not agent roles. They conflate two different concepts.
|
||
|
|
|
||
|
|
### 7.2 New Design: Two Sections
|
||
|
|
|
||
|
|
**Section 1 — Phase Roles**
|
||
|
|
|
||
|
|
Shows the 6 phase roles from `.agent.md` Agent Configuration:
|
||
|
|
|
||
|
|
| Role ID | Phases | Active when |
|
||
|
|
|---|---|---|
|
||
|
|
| `researcher` | research, decomposition, design, test_design | A task is in one of these phases |
|
||
|
|
| `implementer` | implement | A task is in `implement` phase |
|
||
|
|
| `code-reviewer` | code_review | A task is in `code_review` phase |
|
||
|
|
| `bug-hunter` | bug_find, adversarial_bug_find | A task is in one of these phases |
|
||
|
|
| `referee` | referee | A task is in `referee` phase |
|
||
|
|
| `orchestrator` | new, complete, human_intervention | A task is in one of these phases |
|
||
|
|
|
||
|
|
Each role card shows: role icon, role label, status (Active/Idle — based on whether any task is in that role's phases), and the count of tasks in that role's phases. Clicking a role filters the task list to tasks in that role's phases.
|
||
|
|
|
||
|
|
**Data source**: new `/api/phase-roles` endpoint. Returns:
|
||
|
|
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"roles": [
|
||
|
|
{"id": "researcher", "label": "Researcher", "icon": "🔬", "phases": ["research", "decomposition", "design", "test_design"], "active_tasks": 2, "status": "active"},
|
||
|
|
{"id": "implementer", "label": "Implementer", "icon": "⚙️", "phases": ["implement"], "active_tasks": 1, "status": "active"},
|
||
|
|
...
|
||
|
|
]
|
||
|
|
}
|
||
|
|
```
|
||
|
|
|
||
|
|
**Section 2 — Scheduled Jobs**
|
||
|
|
|
||
|
|
Shows real scheduled jobs from `/api/scheduled`, using `job.kind` (not fake agent types):
|
||
|
|
|
||
|
|
| Job Kind | Icon | Label | Source |
|
||
|
|
|---|---|---|---|
|
||
|
|
| `cleanup` | 🧹 | Cleanup Archiver | `com.automaton.cleanup` |
|
||
|
|
| `loop` | 🔄 | Loop: {name} | `com.automaton.loop.{name}` |
|
||
|
|
| `rule-scan` | 📝 | Rule Proposer | `com.automaton.rule-scan` (new) |
|
||
|
|
| `rule-review` | 📋 | Rule Reviewer | `com.automaton.rule-review` (new) |
|
||
|
|
|
||
|
|
Each job card shows: job label, status (Enabled/Disabled/Misconfigured), next-run interval, runtime state (for loops: iteration count, halt status; for rule agents: last-scan/review timestamp). `AGENT_TYPE_META` is removed entirely; rendering uses `job.kind` directly.
|
||
|
|
|
||
|
|
### 7.3 Self-Documenting Names
|
||
|
|
|
||
|
|
Per `.rules.md` "Self-Documenting UI Names" (lines 53-64), all schedule unit names and stub filenames are self-documenting:
|
||
|
|
- `com.automaton.rule-scan` (launchd label)
|
||
|
|
- `automaton-rule-scan.sh` (stub filename)
|
||
|
|
- `com.automaton.rule-review` / `automaton-rule-review.sh`
|
||
|
|
|
||
|
|
## 8. Schedules
|
||
|
|
|
||
|
|
Rule agents use OS-native schedulers, mirroring the `--install-cleanup-schedule` pattern (`status.py:2188-2260`):
|
||
|
|
|
||
|
|
| Agent | Flag | Default Interval | Launchd Label | Stub |
|
||
|
|
|---|---|---|---|---|
|
||
|
|
| Rule Proposer | `--install-rule-scan-schedule` | 86400s (daily) | `com.automaton.rule-scan` | `automaton-rule-scan.sh` |
|
||
|
|
| Rule Reviewer | `--install-rule-review-schedule` | 2592000s (monthly) | `com.automaton.rule-review` | `automaton-rule-review.sh` |
|
||
|
|
|
||
|
|
Platform dispatch via `platform.system()`:
|
||
|
|
- **Darwin**: `~/Library/LaunchAgents/com.automaton.rule-scan.plist` with `StartInterval`.
|
||
|
|
- **Linux**: crontab line via `_install_cron_block_generic`.
|
||
|
|
- **Windows**: `schtasks /create /tn "AutomatonRuleScan" ...`.
|
||
|
|
|
||
|
|
Stub scripts are 3-line bash/bat files that call `python3 status.py --rule-scan` (or `--rule-review`).
|
||
|
|
|
||
|
|
## 9. Conflict-of-Interest
|
||
|
|
|
||
|
|
### 9.1 Dependency Chain
|
||
|
|
|
||
|
|
```
|
||
|
|
model-divergence-enforcement (shipped first, manual task)
|
||
|
|
↓ enables hard-block
|
||
|
|
rule-proposer-agent (FW-2)
|
||
|
|
↓ conflict-of-interest
|
||
|
|
rule-reviewer-agent (FW-3) — must differ from Proposer
|
||
|
|
```
|
||
|
|
|
||
|
|
### 9.2 Design Now, Enforce Later (F4, F5)
|
||
|
|
|
||
|
|
- Rule agents are *designed* with `model` fields in their schedule config.
|
||
|
|
- The design docs specify the conflict-of-interest rules.
|
||
|
|
- Hard-block enforcement *activates* when `models.json` exists and multi-LLM mode is detected.
|
||
|
|
- Until then, rule agents run with the default model (single-LLM mode, advisory only).
|
||
|
|
|
||
|
|
### 9.3 Why Different Models
|
||
|
|
|
||
|
|
- **Proposer vs Reviewer**: Proposer has a bias toward *adding* rules (more is better). Reviewer has a bias toward *pruning* (less is better). Same model = self-review = the proposer's rules never get pruned.
|
||
|
|
- **Rule agent vs scanned tasks**: The model that introduced a bug should not propose the rule about its own bug — it has a blind spot for that failure mode.
|
||
|
|
|
||
|
|
## 10. Success Criteria for v1
|
||
|
|
|
||
|
|
1. **Model-divergence**: `--transition --model` records the model in `.state.models`; `--claim` refuses conflict-matrix violations in multi-LLM mode; `--audit` flags violations; dashboard shows model badges.
|
||
|
|
2. **Rule Proposer**: `status.py --rule-scan` reads completed tasks since last scan, proposes rules to `RULE_PROPOSALS.md`, updates `.state.rule-scan`. Test with a seeded completed task containing a `VERDICT.md`.
|
||
|
|
3. **Rule Reviewer**: `status.py --rule-review` reads `.rules.md` + recent proposals, writes `RULE_REVIEW.md`, updates `.state.rule-review`. Test with a seeded `.rules.md` containing a contradiction.
|
||
|
|
4. **Agent Tab**: `/api/phase-roles` returns 6 roles with active-task counts; dashboard renders Phase Roles + Scheduled Jobs sections; `AGENT_TYPE_META` is removed; `job.kind` drives rendering.
|
||
|
|
5. **Schedules**: `--install-rule-scan-schedule` and `--install-rule-review-schedule` install OS-native units with self-documenting names.
|
||
|
|
6. **Tests**: `pytest tests/ -v` is green; new tests cover state schemas, scan flows, dashboard API, and conflict-matrix enforcement.
|
||
|
|
7. **No regression**: pre-existing test suite passes unchanged.
|
||
|
|
|
||
|
|
## 11. Locked Decision Index
|
||
|
|
|
||
|
|
All decisions referenced by `(Fn)` above are recorded in the v1 design conversation (this session). They are non-negotiable for v1 implementation. Changes require a design doc update and a new `[unreleased]` changelog entry.
|
||
|
|
|
||
|
|
See `README.md` § "Locked decisions" for the full table.
|