Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria
Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.
Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
This commit is contained in:
@@ -0,0 +1,67 @@
|
||||
# Framework Agent Features — Design Index
|
||||
|
||||
Status: **v1 draft 2026-06-25**. Design for three framework-level agent features: rule agents (proposer + reviewer), Agent tab redesign, and model-divergence enforcement.
|
||||
|
||||
## What this is
|
||||
|
||||
The Automaton framework's agent-feature layer: scheduled standalone agents for rule maintenance (Rule Proposer, Rule Reviewer), a redesigned dashboard Agent tab showing real phase roles instead of fake types, and a model-divergence enforcement system that prevents conflict-of-interest by ensuring different LLM roles use different models.
|
||||
|
||||
Built on top of the existing `status.py` phase machine and loop infrastructure — no second enforcement surface.
|
||||
|
||||
## Why
|
||||
|
||||
- `.rules.md` is maintained manually; failure patterns from completed tasks are not systematically captured.
|
||||
- The dashboard Agent tab shows 4 made-up agent types (`completed_task_archiver`, `single`, `audit`, `backlog`) — not the real phase roles from `.agent.md` (researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator).
|
||||
- Conflict-of-interest: two sessions with the same LLM playing bug finder and adversarial bug finder (or implementer and reviewer) don't yield useful results. The framework needs computational enforcement of model divergence.
|
||||
|
||||
## Documents
|
||||
|
||||
- [`functional.md`](functional.md) — what v1 does, roles, triggers, outputs, schedules, conflict-of-interest, success criteria. **Read this first.**
|
||||
- [`technical.md`](technical.md) — the implementation contract: file map, state schemas, `status.py` flags, runner flows, dashboard API changes, test coverage. **Read this if you're implementing.**
|
||||
- [`BACKLOG.md`](BACKLOG.md) — v1 work queue (3 items) and deferred items. The self-improvement loop's work queue when `work_source.area = "framework"`.
|
||||
|
||||
## v1 scope (draft)
|
||||
|
||||
Three feature areas, each with a clear boundary:
|
||||
|
||||
1. **Model-Divergence Enforcement** — Foundational layer. `models.json` manifest, mode detection (single vs multi-LLM), conflict matrix, `--transition --model`, `.state.models`, `loop.json` per-role model binding, `{model}` substitution in harness commands, audit category, dashboard badges. Enables the other two features.
|
||||
|
||||
2. **Rule Agents** — Two standalone scheduled agents (NOT task lifecycle phases):
|
||||
- **Rule Proposer**: Daily scan of completed tasks → `RULE_PROPOSALS.md` with concrete examples.
|
||||
- **Rule Reviewer**: Monthly consolidation of `.rules.md` → `RULE_REVIEW.md`.
|
||||
- Both use direct harness invocation (reuse `loop-runner._invoke_harness`), state files (`.state.rule-scan`, `.state.rule-review`), OS-native schedulers (mirror `--install-cleanup-schedule`).
|
||||
|
||||
3. **Agent Tab Redesign** — Dashboard `/agent` view:
|
||||
- **Phase Roles section**: 6 roles from `.agent.md`, active/inactive based on current task phases.
|
||||
- **Scheduled Jobs section**: Real job kinds from `/api/scheduled` (cleanup, loop, rule-scan, rule-review), replacing 4 fake `AGENT_TYPE_META` types.
|
||||
- New `/api/phase-roles` endpoint.
|
||||
|
||||
## Locked decisions (referenced as `(Fn)` in functional/technical)
|
||||
|
||||
| ID | Decision |
|
||||
|---|---|
|
||||
| F1 | Rule agents are **standalone scheduled agents**, NOT task lifecycle phases. They don't block task completion. |
|
||||
| F2 | Rule Proposer runs **daily**; Rule Reviewer runs **monthly**. OS-native schedulers (launchd/cron/schtasks). |
|
||||
| F3 | Rule agents use **direct harness invocation** (reuse `loop-runner._invoke_harness`), not loop infrastructure. |
|
||||
| F4 | Conflict-of-interest: Rule Reviewer **must use different LLM** than Rule Proposer. Rule agents must use different LLM than tasks they review. Enforcement **deferred** until model-divergence ships (F5). |
|
||||
| F5 | Model-divergence enforcement is a **foundational layer** shipped first (manual task, not backlog). Rule agents designed with model-binding config but hard-block enforced later. |
|
||||
| F6 | Agent tab has **two sections**: Phase Roles (from `.agent.md` + task state) + Scheduled Jobs (from `/api/scheduled`). |
|
||||
| F7 | Phase Roles section needs **new `/api/phase-roles` endpoint** mapping active tasks to their phase roles. |
|
||||
| F8 | Scheduled Jobs section uses **real `job.kind`** (cleanup, loop, rule-scan, rule-review) — remove `AGENT_TYPE_META` fake types. |
|
||||
| F9 | State files: `.state.rule-scan` (last scanned task, timestamp, proposed rules), `.state.rule-review` (last run timestamp). |
|
||||
| F10 | New `status.py` flags: `--rule-scan`, `--install-rule-scan-schedule`, `--rule-review`, `--install-rule-review-schedule` (mirror `--cleanup-done` / `--install-cleanup-schedule`). |
|
||||
|
||||
## Relationship to loops
|
||||
|
||||
- `design/loops/` is the loop engineering system (state-enforced unattended work).
|
||||
- `design/framework/` is framework-level agent features (rule maintenance, observability, model governance).
|
||||
- The self-improvement loop (`templates/loops/self-improvement/`) can drive `design/framework/` work by setting `work_source.area = "framework"` — no code change needed (see `loop-runner.py:_find_work_backlog`).
|
||||
- Model-divergence enforcement (F5) is a prerequisite for hard-blocking conflict-of-interest in rule agents and loops.
|
||||
|
||||
## Post-v1
|
||||
|
||||
After v1 lands and the self-improvement loop is running:
|
||||
|
||||
- Rule enforcement gate (`--can-edit --rule <name>`) — separate feature.
|
||||
- Dashboard Agent tab v2: role details, schedule management, per-agent logs.
|
||||
- Rule agents gain comprehension-debt tracking (`last-read-sha` per rule file).
|
||||
Reference in New Issue
Block a user