Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria
Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.
Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
17 KiB
Framework Agent Features — Functional Design
Status: v1 (draft 2026-06-25). Supersedes any prior informal discussions of rule agents or Agent tab redesign.
Audience: framework maintainers (currently: the human and one AI assistant). After handoff the self-improvement loop is also an audience — designs must be legible to a fresh-context LLM verifier.
1. Problem
Three gaps in the framework's agent layer:
-
.rules.mdis maintained manually. Failure patterns from completed tasks (BUG_REPORT.md,ADVERSARIAL_BUG_REPORT.md,VERDICT.md) are not systematically captured into rules. The Self-Improvement section of.rules.md(lines 33-36) says "Add one rule per observed failure mode with a concrete example" and "Consolidate contradictions monthly" — but no agent does this. It relies on the human or a session agent remembering. -
The dashboard Agent tab shows fake agent types.
AGENT_TYPE_METAindashboard.js:330-334defines 4 types (completed_task_archiver,single,audit,backlog) that are loop work-source kinds, not real agent roles. The 6 real phase roles from.agent.md(researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) are not surfaced. -
No conflict-of-interest enforcement on model divergence.
status.py:1432-1436enforces that the reviewer session differs from the implementer session (via.state.implementer), but there is no enforcement that the model playing bug-finder differs from the model playing adversarial bug-finder, or that the referee model differs from the implementer model. Same-model conflict-of-interest yields rubber-stamping.
2. Goals
v1 — the framework maintains its own rules, surfaces real agent roles, and computationally enforces model divergence:
- Rule Proposer: a daily scheduled standalone agent that scans completed tasks since its last run, reads their failure artifacts, and proposes new rules with concrete examples to
RULE_PROPOSALS.md. - Rule Reviewer: a monthly scheduled standalone agent that consolidates
.rules.md— finds contradictions, stale rules, rules missing examples — and writesRULE_REVIEW.md. - Agent Tab Redesign: replace 4 fake
AGENT_TYPE_METAtypes with two sections — Phase Roles (6 roles from.agent.md, active/inactive based on current task phases) and Scheduled Jobs (realjob.kindfrom/api/scheduled). - Model-Divergence Enforcement: a
models.jsonmanifest, mode detection (single vs multi-LLM), a conflict matrix,--transition --model,.state.models,loop.jsonper-role model binding,{model}substitution in harness commands, an audit category, and dashboard badges. - Conflict-of-interest for rule agents: Rule Reviewer must use a different LLM than Rule Proposer. Both must use different LLMs than the tasks they review. Enforcement is designed now but hard-blocked only after model-divergence ships (F4, F5).
3. Non-Goals (v1)
- Rule enforcement during task execution. Rule agents only propose and review rules; they do not block edits. A future
--can-edit --rule <name>gate (FW-6) is separate. - Auto-approving rules. A human always reviews
RULE_PROPOSALS.mdandRULE_REVIEW.mdbefore rules are merged into.rules.md. No auto-merge path in v1. - Rule agents as task lifecycle phases. Rule agents are standalone scheduled agents (F1). They do not block task completion. They are not phases in the state machine.
- A new agent harness. Rule agents use direct harness invocation (reusing
loop-runner._invoke_harness). No new runtime. - Network-fetched dependencies. New code is Python stdlib only. No new pip installs.
- Model capability inspection. The framework never inspects model capability, provider, or size (D8). It only tracks which model fills which role and enforces the conflict matrix.
4. Model-Divergence Enforcement (foundational layer)
Shipped first as a manual task (model-divergence-enforcement), not a backlog item. Enables conflict-of-interest hard-blocking for rule agents and loops.
4.1 Manifest: models.json
A new file at ~/.automaton/models.json (or {project}/.automaton/models.json):
{
"default": "glm-4.6",
"advised": true,
"models": [
{"name": "glm-4.6", "provider": "opencode", "context_window": 131072, "location": "remote"},
{"name": "qwen3-coder", "provider": "opencode", "context_window": 131072, "location": "remote"},
{"name": "llama-3.3-70b", "provider": "localhost", "context_window": 32768, "location": "http://localhost:8080"}
]
}
default: the model used when no role-specific binding is set.advised: iftrue, the framework prints a one-time advisory in single-LLM mode recommending a second model for conflict-of-interest roles, then goes silent.models[]: the roster.locationis"remote"or a localhost URL for probing.
4.2 Mode Detection
- 0-1 models in
models.json(or file missing) → single-LLM mode. Advisory once (ifadvised: true), then silent. No hard blocks. - 2+ models → multi-LLM mode. Hard-block on conflict-matrix violations. Auto-assign next-available non-conflicting model on conflict; refuse only if no non-conflicting model exists.
- Missing file → single-LLM mode (backward compatible). Existing behavior preserved.
4.3 Conflict Matrix (locked)
| Role | Must differ from |
|---|---|
code_review |
implement |
bug_find |
implement |
adversarial_bug_find |
implement, bug_find |
referee |
implement, bug_find, adversarial_bug_find |
loop-verify |
loop-implement |
doc_review, code_review, and bug_find are independent of each other (not conflicts). Only bug_find ↔ adversarial_bug_find conflicts (they are adversary pairs).
4.4 Auto-Assignment (multi-LLM mode)
- Default model → assigned to
implement(andloop-implement). - On conflict, pick the next-available model from
models[]that does not conflict. - User override:
loop.jsonroles.<role>.modelorstatus.py --transition --model <name>. - Refuse only if no non-conflicting model exists.
4.5 State: .state.models
Each task gets {task}/.state.models recording which model filled which role:
{"implement": "glm-4.6", "code_review": "qwen3-coder", "bug_find": "qwen3-coder", "adversarial_bug_find": "llama-3.3-70b", "referee": "llama-3.3-70b"}
status.py --transition --model <name> records the model for the role being transitioned into. --claim in multi-LLM mode checks the conflict matrix against .state.models and refuses on violation.
4.6 Loop Integration
loop.json gains per-role model and harness.command with {model} substitution:
"roles": {
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
}
loop-runner.py:_invoke_harness substitutes {model} into the harness command. --check-gate enforces loop-verify ≠ loop-implement model in multi-LLM mode.
4.7 Audit + Dashboard
status.py --auditgains amodel_divergencecategory: flags tasks where.state.modelsviolates the conflict matrix.- Dashboard task cards show model badges (one per role filled).
5. Rule Proposer
A standalone scheduled agent (F1) that proposes new rules from completed-task failure patterns.
5.1 Trigger
Daily, via OS-native scheduler (mirrors --install-cleanup-schedule). status.py --install-rule-scan-schedule [--interval 86400] installs the schedule unit. Manual: status.py --rule-scan.
5.2 Inputs
.state.rule-scan: state file tracking the last scanned task and timestamp.- Completed tasks (in
tasks/complete/or tasks with.statephasecomplete) that were completed since the last scan. - For each such task:
BUG_REPORT.md,ADVERSARIAL_BUG_REPORT.md,VERDICT.md(if present). - Current
.rules.md(to deduplicate against existing rules).
5.3 Output
RULE_PROPOSALS.md (at ~/.automaton/RULE_PROPOSALS.md in framework mode, or {project}/.automaton/RULE_PROPOSALS.md in project mode). Format:
# Rule Proposals — {date}
## Proposed Rule: {title}
**Source**: tasks/{task-name}/VERDICT.md
**Pattern**: {one-line description of the failure mode}
**Example**:
{concrete code/config snippet from the task}
**Proposed rule text**:
{the rule as it would appear in .rules.md}
---
Proposals are appended per run. A human reviews and merges accepted rules into .rules.md. No auto-merge (Non-Goal).
5.4 LLM Session
Direct harness invocation (F3): status.py --rule-scan reuses loop-runner._invoke_harness to spawn one LLM session with a prompt that includes the failure artifacts and current rules. The session proposes rules in the RULE_PROPOSALS.md format.
5.5 Conflict-of-Interest
The Rule Proposer's LLM must differ from the implementer + bug-hunter + adversarial-bug-hunter models of the tasks it scans. This prevents the model that made the bug from proposing the rule about its own bug.
Enforcement deferred (F4): until model-divergence ships, the Rule Proposer runs with the default model. The design includes a model field in its schedule config; hard-blocking activates once models.json exists and multi-LLM mode is detected.
5.6 State: .state.rule-scan
{"last_scan_at": "2026-06-25T10:00:00Z", "last_scanned_task": "fix-context-sizing", "proposals_count": 3}
6. Rule Reviewer
A standalone scheduled agent (F1) that consolidates .rules.md periodically.
6.1 Trigger
Monthly, via OS-native scheduler. status.py --install-rule-review-schedule [--interval 2592000] installs the schedule unit. Manual: status.py --rule-review.
6.2 Inputs
.state.rule-review: state file tracking the last run timestamp.- Current
.rules.md(full file). - Recent
RULE_PROPOSALS.mdentries (since last review). - Recent completed-task summaries (last 30 days) for context on stale rules.
6.3 Output
RULE_REVIEW.md (at ~/.automaton/RULE_REVIEW.md or project equivalent). Format:
# Rule Review — {date}
## Contradictions Found
- Rule A ("...") contradicts Rule B ("..."). Suggested resolution: {merge/drop/keep A}.
## Stale Rules (no observed instance in last 30 days)
- Rule C ("..."). Suggested action: drop or annotate as low-priority.
## Rules Missing Examples
- Rule D ("..."). Suggested example: {from a recent task}.
## Merge Candidates
- Rules E and F overlap. Suggested merged text: {...}.
---
A human reviews and applies accepted changes to .rules.md. No auto-apply (Non-Goal).
6.4 LLM Session
Direct harness invocation (F3), same as Rule Proposer. One LLM session with a prompt that includes the full .rules.md and recent proposals.
6.5 Conflict-of-Interest
The Rule Reviewer's LLM must differ from the Rule Proposer's LLM (F4). The Proposer proposes (bias toward adding); the Reviewer consolidates (bias toward pruning). Same model = self-review = rubber-stamping.
Enforcement deferred until model-divergence ships.
6.6 State: .state.rule-review
{"last_review_at": "2026-06-25T10:00:00Z", "contradictions_found": 2, "stale_rules": 5, "merges_suggested": 1}
7. Agent Tab Redesign
7.1 Current State (to be replaced)
dashboard.js:330-334 defines AGENT_TYPE_META with 4 fake types:
completed_task_archiver— actually the cleanup scheduled job.single— actually a loop withwork_source.kind = "single".audit— actually a loop withwork_source.kind = "audit".backlog— actually a loop withwork_source.kind = "backlog".
These are loop work-source kinds, not agent roles. They conflate two different concepts.
7.2 New Design: Two Sections
Section 1 — Phase Roles
Shows the 6 phase roles from .agent.md Agent Configuration:
| Role ID | Phases | Active when |
|---|---|---|
researcher |
research, decomposition, design, test_design | A task is in one of these phases |
implementer |
implement | A task is in implement phase |
code-reviewer |
code_review | A task is in code_review phase |
bug-hunter |
bug_find, adversarial_bug_find | A task is in one of these phases |
referee |
referee | A task is in referee phase |
orchestrator |
new, complete, human_intervention | A task is in one of these phases |
Each role card shows: role icon, role label, status (Active/Idle — based on whether any task is in that role's phases), and the count of tasks in that role's phases. Clicking a role filters the task list to tasks in that role's phases.
Data source: new /api/phase-roles endpoint. Returns:
{
"roles": [
{"id": "researcher", "label": "Researcher", "icon": "🔬", "phases": ["research", "decomposition", "design", "test_design"], "active_tasks": 2, "status": "active"},
{"id": "implementer", "label": "Implementer", "icon": "⚙️", "phases": ["implement"], "active_tasks": 1, "status": "active"},
...
]
}
Section 2 — Scheduled Jobs
Shows real scheduled jobs from /api/scheduled, using job.kind (not fake agent types):
| Job Kind | Icon | Label | Source |
|---|---|---|---|
cleanup |
🧹 | Cleanup Archiver | com.automaton.cleanup |
loop |
🔄 | Loop: {name} | com.automaton.loop.{name} |
rule-scan |
📝 | Rule Proposer | com.automaton.rule-scan (new) |
rule-review |
📋 | Rule Reviewer | com.automaton.rule-review (new) |
Each job card shows: job label, status (Enabled/Disabled/Misconfigured), next-run interval, runtime state (for loops: iteration count, halt status; for rule agents: last-scan/review timestamp). AGENT_TYPE_META is removed entirely; rendering uses job.kind directly.
7.3 Self-Documenting Names
Per .rules.md "Self-Documenting UI Names" (lines 53-64), all schedule unit names and stub filenames are self-documenting:
com.automaton.rule-scan(launchd label)automaton-rule-scan.sh(stub filename)com.automaton.rule-review/automaton-rule-review.sh
8. Schedules
Rule agents use OS-native schedulers, mirroring the --install-cleanup-schedule pattern (status.py:2188-2260):
| Agent | Flag | Default Interval | Launchd Label | Stub |
|---|---|---|---|---|
| Rule Proposer | --install-rule-scan-schedule |
86400s (daily) | com.automaton.rule-scan |
automaton-rule-scan.sh |
| Rule Reviewer | --install-rule-review-schedule |
2592000s (monthly) | com.automaton.rule-review |
automaton-rule-review.sh |
Platform dispatch via platform.system():
- Darwin:
~/Library/LaunchAgents/com.automaton.rule-scan.plistwithStartInterval. - Linux: crontab line via
_install_cron_block_generic. - Windows:
schtasks /create /tn "AutomatonRuleScan" ....
Stub scripts are 3-line bash/bat files that call python3 status.py --rule-scan (or --rule-review).
9. Conflict-of-Interest
9.1 Dependency Chain
model-divergence-enforcement (shipped first, manual task)
↓ enables hard-block
rule-proposer-agent (FW-2)
↓ conflict-of-interest
rule-reviewer-agent (FW-3) — must differ from Proposer
9.2 Design Now, Enforce Later (F4, F5)
- Rule agents are designed with
modelfields in their schedule config. - The design docs specify the conflict-of-interest rules.
- Hard-block enforcement activates when
models.jsonexists and multi-LLM mode is detected. - Until then, rule agents run with the default model (single-LLM mode, advisory only).
9.3 Why Different Models
- Proposer vs Reviewer: Proposer has a bias toward adding rules (more is better). Reviewer has a bias toward pruning (less is better). Same model = self-review = the proposer's rules never get pruned.
- Rule agent vs scanned tasks: The model that introduced a bug should not propose the rule about its own bug — it has a blind spot for that failure mode.
10. Success Criteria for v1
- Model-divergence:
--transition --modelrecords the model in.state.models;--claimrefuses conflict-matrix violations in multi-LLM mode;--auditflags violations; dashboard shows model badges. - Rule Proposer:
status.py --rule-scanreads completed tasks since last scan, proposes rules toRULE_PROPOSALS.md, updates.state.rule-scan. Test with a seeded completed task containing aVERDICT.md. - Rule Reviewer:
status.py --rule-reviewreads.rules.md+ recent proposals, writesRULE_REVIEW.md, updates.state.rule-review. Test with a seeded.rules.mdcontaining a contradiction. - Agent Tab:
/api/phase-rolesreturns 6 roles with active-task counts; dashboard renders Phase Roles + Scheduled Jobs sections;AGENT_TYPE_METAis removed;job.kinddrives rendering. - Schedules:
--install-rule-scan-scheduleand--install-rule-review-scheduleinstall OS-native units with self-documenting names. - Tests:
pytest tests/ -vis green; new tests cover state schemas, scan flows, dashboard API, and conflict-matrix enforcement. - No regression: pre-existing test suite passes unchanged.
11. Locked Decision Index
All decisions referenced by (Fn) above are recorded in the v1 design conversation (this session). They are non-negotiable for v1 implementation. Changes require a design doc update and a new [unreleased] changelog entry.
See README.md § "Locked decisions" for the full table.