Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria
Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.
Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
5.3 KiB
Framework Agent Features — Design Index
Status: v1 draft 2026-06-25. Design for three framework-level agent features: rule agents (proposer + reviewer), Agent tab redesign, and model-divergence enforcement.
What this is
The Automaton framework's agent-feature layer: scheduled standalone agents for rule maintenance (Rule Proposer, Rule Reviewer), a redesigned dashboard Agent tab showing real phase roles instead of fake types, and a model-divergence enforcement system that prevents conflict-of-interest by ensuring different LLM roles use different models.
Built on top of the existing status.py phase machine and loop infrastructure — no second enforcement surface.
Why
.rules.mdis maintained manually; failure patterns from completed tasks are not systematically captured.- The dashboard Agent tab shows 4 made-up agent types (
completed_task_archiver,single,audit,backlog) — not the real phase roles from.agent.md(researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator). - Conflict-of-interest: two sessions with the same LLM playing bug finder and adversarial bug finder (or implementer and reviewer) don't yield useful results. The framework needs computational enforcement of model divergence.
Documents
functional.md— what v1 does, roles, triggers, outputs, schedules, conflict-of-interest, success criteria. Read this first.technical.md— the implementation contract: file map, state schemas,status.pyflags, runner flows, dashboard API changes, test coverage. Read this if you're implementing.BACKLOG.md— v1 work queue (3 items) and deferred items. The self-improvement loop's work queue whenwork_source.area = "framework".
v1 scope (draft)
Three feature areas, each with a clear boundary:
-
Model-Divergence Enforcement — Foundational layer.
models.jsonmanifest, mode detection (single vs multi-LLM), conflict matrix,--transition --model,.state.models,loop.jsonper-role model binding,{model}substitution in harness commands, audit category, dashboard badges. Enables the other two features. -
Rule Agents — Two standalone scheduled agents (NOT task lifecycle phases):
- Rule Proposer: Daily scan of completed tasks →
RULE_PROPOSALS.mdwith concrete examples. - Rule Reviewer: Monthly consolidation of
.rules.md→RULE_REVIEW.md. - Both use direct harness invocation (reuse
loop-runner._invoke_harness), state files (.state.rule-scan,.state.rule-review), OS-native schedulers (mirror--install-cleanup-schedule).
- Rule Proposer: Daily scan of completed tasks →
-
Agent Tab Redesign — Dashboard
/agentview:- Phase Roles section: 6 roles from
.agent.md, active/inactive based on current task phases. - Scheduled Jobs section: Real job kinds from
/api/scheduled(cleanup, loop, rule-scan, rule-review), replacing 4 fakeAGENT_TYPE_METAtypes. - New
/api/phase-rolesendpoint.
- Phase Roles section: 6 roles from
Locked decisions (referenced as (Fn) in functional/technical)
| ID | Decision |
|---|---|
| F1 | Rule agents are standalone scheduled agents, NOT task lifecycle phases. They don't block task completion. |
| F2 | Rule Proposer runs daily; Rule Reviewer runs monthly. OS-native schedulers (launchd/cron/schtasks). |
| F3 | Rule agents use direct harness invocation (reuse loop-runner._invoke_harness), not loop infrastructure. |
| F4 | Conflict-of-interest: Rule Reviewer must use different LLM than Rule Proposer. Rule agents must use different LLM than tasks they review. Enforcement deferred until model-divergence ships (F5). |
| F5 | Model-divergence enforcement is a foundational layer shipped first (manual task, not backlog). Rule agents designed with model-binding config but hard-block enforced later. |
| F6 | Agent tab has two sections: Phase Roles (from .agent.md + task state) + Scheduled Jobs (from /api/scheduled). |
| F7 | Phase Roles section needs new /api/phase-roles endpoint mapping active tasks to their phase roles. |
| F8 | Scheduled Jobs section uses real job.kind (cleanup, loop, rule-scan, rule-review) — remove AGENT_TYPE_META fake types. |
| F9 | State files: .state.rule-scan (last scanned task, timestamp, proposed rules), .state.rule-review (last run timestamp). |
| F10 | New status.py flags: --rule-scan, --install-rule-scan-schedule, --rule-review, --install-rule-review-schedule (mirror --cleanup-done / --install-cleanup-schedule). |
Relationship to loops
design/loops/is the loop engineering system (state-enforced unattended work).design/framework/is framework-level agent features (rule maintenance, observability, model governance).- The self-improvement loop (
templates/loops/self-improvement/) can drivedesign/framework/work by settingwork_source.area = "framework"— no code change needed (seeloop-runner.py:_find_work_backlog). - Model-divergence enforcement (F5) is a prerequisite for hard-blocking conflict-of-interest in rule agents and loops.
Post-v1
After v1 lands and the self-improvement loop is running:
- Rule enforcement gate (
--can-edit --rule <name>) — separate feature. - Dashboard Agent tab v2: role details, schedule management, per-agent logs.
- Rule agents gain comprehension-debt tracking (
last-read-shaper rule file).