Files
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00

5.3 KiB

Framework Agent Features — Design Index

Status: v1 draft 2026-06-25. Design for three framework-level agent features: rule agents (proposer + reviewer), Agent tab redesign, and model-divergence enforcement.

What this is

The Automaton framework's agent-feature layer: scheduled standalone agents for rule maintenance (Rule Proposer, Rule Reviewer), a redesigned dashboard Agent tab showing real phase roles instead of fake types, and a model-divergence enforcement system that prevents conflict-of-interest by ensuring different LLM roles use different models.

Built on top of the existing status.py phase machine and loop infrastructure — no second enforcement surface.

Why

  • .rules.md is maintained manually; failure patterns from completed tasks are not systematically captured.
  • The dashboard Agent tab shows 4 made-up agent types (completed_task_archiver, single, audit, backlog) — not the real phase roles from .agent.md (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator).
  • Conflict-of-interest: two sessions with the same LLM playing bug finder and adversarial bug finder (or implementer and reviewer) don't yield useful results. The framework needs computational enforcement of model divergence.

Documents

  • functional.md — what v1 does, roles, triggers, outputs, schedules, conflict-of-interest, success criteria. Read this first.
  • technical.md — the implementation contract: file map, state schemas, status.py flags, runner flows, dashboard API changes, test coverage. Read this if you're implementing.
  • BACKLOG.md — v1 work queue (3 items) and deferred items. The self-improvement loop's work queue when work_source.area = "framework".

v1 scope (draft)

Three feature areas, each with a clear boundary:

  1. Model-Divergence Enforcement — Foundational layer. models.json manifest, mode detection (single vs multi-LLM), conflict matrix, --transition --model, .state.models, loop.json per-role model binding, {model} substitution in harness commands, audit category, dashboard badges. Enables the other two features.

  2. Rule Agents — Two standalone scheduled agents (NOT task lifecycle phases):

    • Rule Proposer: Daily scan of completed tasks → RULE_PROPOSALS.md with concrete examples.
    • Rule Reviewer: Monthly consolidation of .rules.md → RULE_REVIEW.md.
    • Both use direct harness invocation (reuse loop-runner._invoke_harness), state files (.state.rule-scan, .state.rule-review), OS-native schedulers (mirror --install-cleanup-schedule).
  3. Agent Tab Redesign — Dashboard /agent view:

    • Phase Roles section: 6 roles from .agent.md, active/inactive based on current task phases.
    • Scheduled Jobs section: Real job kinds from /api/scheduled (cleanup, loop, rule-scan, rule-review), replacing 4 fake AGENT_TYPE_META types.
    • New /api/phase-roles endpoint.

Locked decisions (referenced as (Fn) in functional/technical)

ID Decision
F1 Rule agents are standalone scheduled agents, NOT task lifecycle phases. They don't block task completion.
F2 Rule Proposer runs daily; Rule Reviewer runs monthly. OS-native schedulers (launchd/cron/schtasks).
F3 Rule agents use direct harness invocation (reuse loop-runner._invoke_harness), not loop infrastructure.
F4 Conflict-of-interest: Rule Reviewer must use different LLM than Rule Proposer. Rule agents must use different LLM than tasks they review. Enforcement deferred until model-divergence ships (F5).
F5 Model-divergence enforcement is a foundational layer shipped first (manual task, not backlog). Rule agents designed with model-binding config but hard-block enforced later.
F6 Agent tab has two sections: Phase Roles (from .agent.md + task state) + Scheduled Jobs (from /api/scheduled).
F7 Phase Roles section needs new /api/phase-roles endpoint mapping active tasks to their phase roles.
F8 Scheduled Jobs section uses real job.kind (cleanup, loop, rule-scan, rule-review) — remove AGENT_TYPE_META fake types.
F9 State files: .state.rule-scan (last scanned task, timestamp, proposed rules), .state.rule-review (last run timestamp).
F10 New status.py flags: --rule-scan, --install-rule-scan-schedule, --rule-review, --install-rule-review-schedule (mirror --cleanup-done / --install-cleanup-schedule).

Relationship to loops

  • design/loops/ is the loop engineering system (state-enforced unattended work).
  • design/framework/ is framework-level agent features (rule maintenance, observability, model governance).
  • The self-improvement loop (templates/loops/self-improvement/) can drive design/framework/ work by setting work_source.area = "framework" — no code change needed (see loop-runner.py:_find_work_backlog).
  • Model-divergence enforcement (F5) is a prerequisite for hard-blocking conflict-of-interest in rule agents and loops.

Post-v1

After v1 lands and the self-improvement loop is running:

  • Rule enforcement gate (--can-edit --rule <name>) — separate feature.
  • Dashboard Agent tab v2: role details, schedule management, per-agent logs.
  • Rule agents gain comprehension-debt tracking (last-read-sha per rule file).