Files
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00

3.3 KiB

Loop Engineering — Backlog

Work queue for the self-improvement loop after task 7 lands. Items not assigned to v1 implementation; they're picked up by loops in priority order.

How loops consume this

A loop configured with work_source.kind = "backlog" reads this file, picks the topmost [ ] item, drafts an implementation, transitions through phases, hands off to a human reviewer. Mark items [x] when complete; move items to DONE.md (created later) on closure.

Sibling backlog: design/framework/BACKLOG.md covers framework-level agent features (rule agents, Agent tab redesign, model-divergence enforcement). A loop with work_source.area = "framework" reads that file instead of this one.

v1.1 — framework manages its own docs

  • design-update-loop-template — templates/loops/design-update/ that keeps design/<area>/*.md in sync with the code it documents. Triggered by last-read-sha drift detection.
  • dashboard-loops-panel — new view in automaton/dashboard/html/ showing loop statuses, halt reasons, iteration counts, and a four-deaths audit table.
  • test_design_drift — tests/test_design_drift.py that fails when design/ docs and code diverge beyond last-read-sha.
  • loop-integration-contract — contracts/loop-integration.md documenting the harness-side contract (mirrors the existing harness-integration.md).
  • last-read-sha-tracking — design/<area>/.last-read-sha per file; dashboard flags drift when file mtime > sha.
  • context-sizing-tier-2 — first real loop workstream. Items seeded from design/context-sizing/BACKLOG.md (sibling design, written after task 7).

Deferred (no v1.1 commitment)

  • parallel-mode-default — flip --mode parallel to default-on (currently opt-in per D6). Blocked on production observation of single-mode loops.
  • scope-3-self-designing — framework drafts its own designs from audit patterns, not just runs pre-authored ones. Blocked on Scope 2 proving loop discipline.
  • auto-approve-relax — allow specific low-risk halt categories (iterations_exhausted on a green-scoring run) to auto-resume. Blocked on D4 staying firm in v1.
  • harness-adapter-spec — formal adapter contract for non-opencode harnesses (aider, Cline, Cursor, Copilot). v1 works with any harness via the generic harness.command in loop.json; a spec layer is a later cleanup.
  • multi-budget-currency — max_budget_usd becomes max_budget with a configurable unit (tokens, seconds, USD). Blocked on remote-only informational usage holding up in practice.
  • compaction-auto-trigger — prompts/compaction.md auto-fires when tick context approaches the tier budget. Currently manual/advisory.

Tier 3 (optimizations, never required for v1.1)

  • loop-concurrency-limit — cap concurrent loops per project when parallel mode lands.
  • verifier-caching — cache verdicts for identical (task, artifact sha) pairs to avoid re-grading on no-op tick retries.
  • schedule-coalescing — multiple loops with the same interval share a single wake event to reduce idle overhead.
  • worktree-gc — garbage-collect stale worktree branches past --max-worktree-age.
  • observability-hook — emit OTel spans for tick phases. Optional; depends on someone running an observability stack.