Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria
Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.
Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
3.3 KiB
3.3 KiB
Loop Engineering — Backlog
Work queue for the self-improvement loop after task 7 lands. Items not assigned to v1 implementation; they're picked up by loops in priority order.
How loops consume this
A loop configured with work_source.kind = "backlog" reads this file, picks the topmost [ ] item, drafts an implementation, transitions through phases, hands off to a human reviewer. Mark items [x] when complete; move items to DONE.md (created later) on closure.
Sibling backlog: design/framework/BACKLOG.md covers framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). A loop with work_source.area = "framework" reads that file instead of this one.
v1.1 — framework manages its own docs
- design-update-loop-template —
templates/loops/design-update/that keepsdesign/<area>/*.mdin sync with the code it documents. Triggered bylast-read-shadrift detection. - dashboard-loops-panel — new view in
automaton/dashboard/html/showing loop statuses, halt reasons, iteration counts, and a four-deaths audit table. - test_design_drift —
tests/test_design_drift.pythat fails whendesign/docs and code diverge beyondlast-read-sha. - loop-integration-contract —
contracts/loop-integration.mddocumenting the harness-side contract (mirrors the existingharness-integration.md). - last-read-sha-tracking —
design/<area>/.last-read-shaper file; dashboard flags drift when file mtime > sha. - context-sizing-tier-2 — first real loop workstream. Items seeded from
design/context-sizing/BACKLOG.md(sibling design, written after task 7).
Deferred (no v1.1 commitment)
- parallel-mode-default — flip
--mode parallelto default-on (currently opt-in per D6). Blocked on production observation of single-mode loops. - scope-3-self-designing — framework drafts its own designs from audit patterns, not just runs pre-authored ones. Blocked on Scope 2 proving loop discipline.
- auto-approve-relax — allow specific low-risk halt categories (
iterations_exhaustedon a green-scoring run) to auto-resume. Blocked on D4 staying firm in v1. - harness-adapter-spec — formal adapter contract for non-opencode harnesses (aider, Cline, Cursor, Copilot). v1 works with any harness via the generic
harness.commandinloop.json; a spec layer is a later cleanup. - multi-budget-currency —
max_budget_usdbecomesmax_budgetwith a configurable unit (tokens, seconds, USD). Blocked on remote-only informational usage holding up in practice. - compaction-auto-trigger —
prompts/compaction.mdauto-fires when tick context approaches the tier budget. Currently manual/advisory.
Tier 3 (optimizations, never required for v1.1)
- loop-concurrency-limit — cap concurrent loops per project when parallel mode lands.
- verifier-caching — cache verdicts for identical (task, artifact sha) pairs to avoid re-grading on no-op tick retries.
- schedule-coalescing — multiple loops with the same interval share a single wake event to reduce idle overhead.
- worktree-gc — garbage-collect stale worktree branches past
--max-worktree-age. - observability-hook — emit OTel spans for tick phases. Optional; depends on someone running an observability stack.