Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria
Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.
Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
7.4 KiB
v1.1 hardening session — in progress
Loop v1 is fully shipped (9/9 bootstrap tasks done; see loop-v1-session-state.md). v1.1 hardening work is underway.
v1.1 plan (7 hardening tasks, priority order)
Picked from BUG_REPORTs of the loop tasks and from design/loops/BACKLOG.md deferred items.
- fix-harness-command-template — DONE (driven to complete; in
tasks/complete/) - add-state-loop-lock — SPEC written, in
research:awaiting_approval(paused to do task 1 first; resume next) - harden-parse-verdict —
parse_verdictscore clamp to[0,1]+passstring coercion ("true"/"false"strings;bool("false")isTruebug). Source:add-loop-runner/BUG_REPORT.mdO6. - add-outputs-retention —
loop.jsonoutputs.retentionfield; GC last N tick dirs ( Source:add-loop-runner/BUG_REPORT.mdO5.) - parametrize-base-branch —
loop.jsonblast_radius.base_branchreplaces hardcodedmainin_gate_worktree_drift. (Source:add-status-brakes/BUG_REPORT.mdO3.) - linux-schedule-parity —
_enable_scheduleLinux parity with Darwin/Windows. (Source:add-status-brakes/BUG_REPORT.mdO2/O4.) - add-claim-loop-task — atomic
current_taskownership before runner touches task. (Source:add-status-brakes/ADVERSARIAL_BUG_REPORT.mdA2.)
Task 1 details: fix-harness-command-template
Root cause: v1 default harness.command was ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"] — opencode run has no --prompt-file or --cwd flags. Only ever exercised via mocked-subprocess unit tests; never invoked live. A real --mode tick would fail on first invocation.
Fix: new {prompt_content} substitution token (single argv element under subprocess.run list mode; safe for any prompt text including quotes/special chars). New default: ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]. {prompt} (file path) and {cwd} retained for backwards compat.
Harness-agnostic contract: preserved and extended. Runner core has zero harness awareness (D8 intact).Pi Dev / aider / Cursor / Copilot / Cline / generic shell wrapper all work via loop.json harness.command override. The new {prompt_content} token makes the framework MORE harness-agnostic (covers harnesses that want a message arg, not a file path).
Per-loop model override: default does NOT hardcode --model; inherits from opencode config. Users route ticks to a specific local LLM (e.g. Qwen3.6-27B on local-mlx) by overriding harness.command in loop.json:
"harness": {"command": ["opencode", "run", "--model", "local-mlx/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-FP4", "--dir", "{cwd}", "{prompt_content}"]}
Inline bug found + fixed (A5): Path(resolved_prompt).read_text() could raise UnicodeDecodeError (subclass of ValueError, not OSError) on non-default-encoding prompt files. Broadened except to (OSError, UnicodeDecodeError); falls back to empty prompt content rather than crashing mid-tick.
Test scaffolding updates: _make_loop helpers in tests/test_loop_runner.py, tests/test_blast_radius.py, tests/test_goal_mode.py, tests/test_loop_templates.py now write loop-local prompt stubs (role-marker content prompt: <ref>) when the framework prompt at ~/.automaton/prompts/<ref> does NOT already exist. Preserves substring-matcher strategy for default-command tests while not overriding real framework prompts in test_loop_templates (which need {current_task} etc. substitution tokens). Custom-{prompt}-command tests in test_goal_mode.py updated matcher substrings from test-impl/test-verify/test-orch to implement-prompt/verify-prompt/orchestrate-prompt (temp file path contains those markers).
Tests: tests/test_harness_command.py (NEW) — 7 tests including the Pi Dev-shaped command test proving the substitution mechanism is harness-agnostic.
Test count: 440 passed (was 433; +7 new).
Custom opencode config for this machine
glm-5.2(headroom-routed, this session's model):opencode-go/glm-5.2gemma4-26b(llama-server on :8080):local-llm/gemma4-26b- Qwen3.6-27B (mlx-vlm on :8000):
local-mlx/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-FP4- mlx server confirmed running and serving (plus an MTP drafter variant)
- User wanted to see Qwen in action via the loop runner tied-role mechanism — that's what motivated doing task 1 first.
Workflow reminders (still in effect)
- Pipeline staging pattern is NOT used in v1.1 tasks (the
.staging/convention was mid-loop-v1 only; v1.1 tasks write artifacts directly to the task folder). - Pipeline:
new -> research -> research:awaiting_approval -> approve -> research:approved -> implement -> code_review -> code_review:awaiting_approval -> approve -> code_review:approved -> bug_find -> adversarial_bug_find -> doc_review -> referee -> complete. (Skip decomposition/design/test_test_design for simpler v1.1 hardening tasks.) --transition completemoves the task dir totasks/complete/automatically (per the relocate feature).- Test command:
python3 -m pytest tests/ -q. Lint:python3 -m py_compile <file>. No new pip deps; stdlib only. - Backlog items deferred past v1.1 are in
design/loops/BACKLOG.md(parallel-mode-default, scope-3-self-designing, auto-approve-relax, harness-adapter-spec, multi-budget-currency, compaction-auto-trigger, plus Tier 3 optimizations).
NEXT
Resume task 2 (add-state-loop-lock): SPEC is already written, state is research:awaiting_approval. Approve it, transition to implement, write the _loop_lock helper + wrap callsites in status.py and loop-runner.py, add tests/test_state_loop_lock.py (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order.
Session 2026-06-25 — Framework agent features design
- Created
design/framework/withREADME.md,functional.md,technical.md,BACKLOG.md— design for three framework-level agent features: model-divergence enforcement, rule agents (Proposer + Reviewer), Agent tab redesign. - Model-divergence enforcement is a manual task (
model-divergence-enforcement), decomposed into 3 subtasks: (1) manifest+detection, (2) interactive enforcement+audit, (3) loop enforcement+dashboard. Not a backlog item. - Rule agents (FW-2, FW-3) are backlog items depending on model-divergence shipping first (conflict-of-interest LLM binding). Rule Proposer runs daily, Rule Reviewer runs monthly. Both use direct harness invocation (reuse
loop-runner._invoke_harness), not loop infrastructure. - Agent tab redesign (FW-1) is a backlog item with no dependencies. Replaces 4 fake
AGENT_TYPE_METAtypes with Phase Roles (6 roles from.agent.md) + Scheduled Jobs (realjob.kind). - Decision: rule agents use direct harness invocation (not loops, not standalone status.py commands).
- Decision: model-divergence conflict-of-interest is designed now, enforced later — rule agents carry
modelfields in config but hard-blocking activates only whenmodels.jsonexists and multi-LLM mode is detected. - Cross-references updated:
AGENTS.mdrepo layout,README.mdloop config table,design/loops/README.md,design/loops/BACKLOG.md,CHANGELOG.md,.onboarding.md(Backlog section),prompts/onboarding.md(Step 2d). - Next: switch on self-improvement loop (
--create-loop self-improvement --from-template self-improvement+--install-schedule), then createmodel-divergence-enforcementparent task and decompose into 3 subtasks.