Files
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00

3.5 KiB

Framework Backlog

Status: v1 draft 2026-06-25. These items are captured for the self-improvement loop (or manual pickup) once the self-improvement loop is running and the design docs are in place.

This backlog mirrors the pattern of design/loops/BACKLOG.md but for framework-level agent features. The self-improvement loop's work_source.area can be set to "framework" to pull from here.


v1 (Priority)

ID Item Description Dependencies
FW-1 agent-tab-real-roles Redesign the dashboard Agent tab: replace 4 fake AGENT_TYPE_META types with two sections — (1) Phase Roles (6 roles from .agent.md: researcher, implementer, code-reviewer, bug-hunter, referee, orchestrator) showing active/inactive based on current task phases; (2) Scheduled Jobs (cleanup, loop-tick, rule-scan, rule-review) using real job.kind from /api/scheduled. New /api/phase-roles endpoint. None
FW-2 rule-proposer-agent Daily scheduled standalone agent that scans completed tasks since last run, reads their BUG_REPORT.md / ADVERSARIAL_BUG_REPORT.md / VERDICT.md, proposes new rules with concrete examples to RULE_PROPOSALS.md. Uses direct harness invocation (reuses loop-runner._invoke_harness). State in .state.rule-scan. Requires different LLM than the tasks it reviews (conflict-of-interest) — enforcement deferred until model-divergence-enforcement ships. model-divergence-enforcement (for LLM binding)
FW-3 rule-reviewer-agent Monthly scheduled standalone agent that consolidates .rules.md — finds contradictions, stale rules, missing examples — writes RULE_REVIEW.md. Uses direct harness invocation. State in .state.rule-review. Must use different LLM than Rule Proposer (conflict-of-interest) — enforcement deferred until model-divergence-enforcement ships. model-divergence-enforcement (for LLM binding), rule-proposer-agent

Deferred / Tier 2 (Post-v1)

ID Item Description Notes
FW-4 model-divergence-enforcement Manifest (models.json), mode detection (single vs multi-LLM), conflict matrix, --transition --model, .state.models, loop.json per-role model + {model} substitution, audit category, dashboard badges. This is a manual task, not a backlog item — tracked separately. See separate task model-divergence-enforcement
FW-5 dashboard-agent-tab-v2 Enhance Agent tab with: role details (click → task list), schedule management (enable/disable from UI), last-run timestamps, per-agent logs. Requires FW-1
FW-6 rule-enforcement-gate Pre-edit hook (--can-edit --rule <name>) that enforces rules from .rules.md before edits. Separate from rule agents (which only propose/review). Requires FW-2, FW-3

Notes

  • Dependency on model-divergence-enforcement: FW-2 and FW-3 are designed with model-binding in their configs (per-role model in their schedule config), but hard-block enforcement is deferred until the model-divergence feature ships. The design docs note the dependency; implementation tasks in the backlog carry a "depends on" annotation.
  • Self-improvement loop: Once the self-improvement loop is running with work_source.area = "framework", it will pick up items from this backlog automatically. The loop template templates/loops/self-improvement/loop.json already supports work_source.area.
  • FW-4 is NOT in this backlog — it's a manual parent task created via status.py --create-task model-divergence-enforcement and decomposed into subtasks.