Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria
Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.
Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
3.5 KiB
3.5 KiB
Framework Backlog
Status: v1 draft 2026-06-25. These items are captured for the self-improvement loop (or manual pickup) once the self-improvement loop is running and the design docs are in place.
This backlog mirrors the pattern of design/loops/BACKLOG.md but for framework-level agent features. The self-improvement loop's work_source.area can be set to "framework" to pull from here.
v1 (Priority)
| ID | Item | Description | Dependencies |
|---|---|---|---|
| FW-1 | agent-tab-real-roles | Redesign the dashboard Agent tab: replace 4 fake AGENT_TYPE_META types with two sections — (1) Phase Roles (6 roles from .agent.md: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) showing active/inactive based on current task phases; (2) Scheduled Jobs (cleanup, loop-tick, rule-scan, rule-review) using real job.kind from /api/scheduled. New /api/phase-roles endpoint. |
None |
| FW-2 | rule-proposer-agent | Daily scheduled standalone agent that scans completed tasks since last run, reads their BUG_REPORT.md / ADVERSARIAL_BUG_REPORT.md / VERDICT.md, proposes new rules with concrete examples to RULE_PROPOSALS.md. Uses direct harness invocation (reuses loop-runner._invoke_harness). State in .state.rule-scan. Requires different LLM than the tasks it reviews (conflict-of-interest) — enforcement deferred until model-divergence-enforcement ships. |
model-divergence-enforcement (for LLM binding) |
| FW-3 | rule-reviewer-agent | Monthly scheduled standalone agent that consolidates .rules.md — finds contradictions, stale rules, missing examples — writes RULE_REVIEW.md. Uses direct harness invocation. State in .state.rule-review. Must use different LLM than Rule Proposer (conflict-of-interest) — enforcement deferred until model-divergence-enforcement ships. |
model-divergence-enforcement (for LLM binding), rule-proposer-agent |
Deferred / Tier 2 (Post-v1)
| ID | Item | Description | Notes |
|---|---|---|---|
| FW-4 | model-divergence-enforcement | Manifest (models.json), mode detection (single vs multi-LLM), conflict matrix, --transition --model, .state.models, loop.json per-role model + {model} substitution, audit category, dashboard badges. This is a manual task, not a backlog item — tracked separately. |
See separate task model-divergence-enforcement |
| FW-5 | dashboard-agent-tab-v2 | Enhance Agent tab with: role details (click → task list), schedule management (enable/disable from UI), last-run timestamps, per-agent logs. | Requires FW-1 |
| FW-6 | rule-enforcement-gate | Pre-edit hook (--can-edit --rule <name>) that enforces rules from .rules.md before edits. Separate from rule agents (which only propose/review). |
Requires FW-2, FW-3 |
Notes
- Dependency on
model-divergence-enforcement: FW-2 and FW-3 are designed with model-binding in their configs (per-rolemodelin their schedule config), but hard-block enforcement is deferred until the model-divergence feature ships. The design docs note the dependency; implementation tasks in the backlog carry a "depends on" annotation. - Self-improvement loop: Once the self-improvement loop is running with
work_source.area = "framework", it will pick up items from this backlog automatically. The loop templatetemplates/loops/self-improvement/loop.jsonalready supportswork_source.area. - FW-4 is NOT in this backlog — it's a manual parent task created via
status.py --create-task model-divergence-enforcementand decomposed into subtasks.