Files
automaton/memory/v1-1-hardening-session.md
T
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00

7.4 KiB

v1.1 hardening session — in progress

Loop v1 is fully shipped (9/9 bootstrap tasks done; see loop-v1-session-state.md). v1.1 hardening work is underway.

v1.1 plan (7 hardening tasks, priority order)

Picked from BUG_REPORTs of the loop tasks and from design/loops/BACKLOG.md deferred items.

  1. fix-harness-command-template — DONE (driven to complete; in tasks/complete/)
  2. add-state-loop-lock — SPEC written, in research:awaiting_approval (paused to do task 1 first; resume next)
  3. harden-parse-verdict — parse_verdict score clamp to [0,1] + pass string coercion ("true"/"false" strings; bool("false") is True bug). Source: add-loop-runner/BUG_REPORT.md O6.
  4. add-outputs-retention — loop.json outputs.retention field; GC last N tick dirs ( Source: add-loop-runner/BUG_REPORT.md O5.)
  5. parametrize-base-branch — loop.json blast_radius.base_branch replaces hardcoded main in _gate_worktree_drift. (Source: add-status-brakes/BUG_REPORT.md O3.)
  6. linux-schedule-parity — _enable_schedule Linux parity with Darwin/Windows. (Source: add-status-brakes/BUG_REPORT.md O2/O4.)
  7. add-claim-loop-task — atomic current_task ownership before runner touches task. (Source: add-status-brakes/ADVERSARIAL_BUG_REPORT.md A2.)

Task 1 details: fix-harness-command-template

Root cause: v1 default harness.command was ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"] — opencode run has no --prompt-file or --cwd flags. Only ever exercised via mocked-subprocess unit tests; never invoked live. A real --mode tick would fail on first invocation.

Fix: new {prompt_content} substitution token (single argv element under subprocess.run list mode; safe for any prompt text including quotes/special chars). New default: ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]. {prompt} (file path) and {cwd} retained for backwards compat.

Harness-agnostic contract: preserved and extended. Runner core has zero harness awareness (D8 intact).Pi Dev / aider / Cursor / Copilot / Cline / generic shell wrapper all work via loop.json harness.command override. The new {prompt_content} token makes the framework MORE harness-agnostic (covers harnesses that want a message arg, not a file path).

Per-loop model override: default does NOT hardcode --model; inherits from opencode config. Users route ticks to a specific local LLM (e.g. Qwen3.6-27B on local-mlx) by overriding harness.command in loop.json:

"harness": {"command": ["opencode", "run", "--model", "local-mlx/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-FP4", "--dir", "{cwd}", "{prompt_content}"]}

Inline bug found + fixed (A5): Path(resolved_prompt).read_text() could raise UnicodeDecodeError (subclass of ValueError, not OSError) on non-default-encoding prompt files. Broadened except to (OSError, UnicodeDecodeError); falls back to empty prompt content rather than crashing mid-tick.

Test scaffolding updates: _make_loop helpers in tests/test_loop_runner.py, tests/test_blast_radius.py, tests/test_goal_mode.py, tests/test_loop_templates.py now write loop-local prompt stubs (role-marker content prompt: <ref>) when the framework prompt at ~/.automaton/prompts/<ref> does NOT already exist. Preserves substring-matcher strategy for default-command tests while not overriding real framework prompts in test_loop_templates (which need {current_task} etc. substitution tokens). Custom-{prompt}-command tests in test_goal_mode.py updated matcher substrings from test-impl/test-verify/test-orch to implement-prompt/verify-prompt/orchestrate-prompt (temp file path contains those markers).

Tests: tests/test_harness_command.py (NEW) — 7 tests including the Pi Dev-shaped command test proving the substitution mechanism is harness-agnostic.

Test count: 440 passed (was 433; +7 new).

Custom opencode config for this machine

  • glm-5.2 (headroom-routed, this session's model): opencode-go/glm-5.2
  • gemma4-26b (llama-server on :8080): local-llm/gemma4-26b
  • Qwen3.6-27B (mlx-vlm on :8000): local-mlx/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-FP4
    • mlx server confirmed running and serving (plus an MTP drafter variant)
    • User wanted to see Qwen in action via the loop runner tied-role mechanism — that's what motivated doing task 1 first.

Workflow reminders (still in effect)

  • Pipeline staging pattern is NOT used in v1.1 tasks (the .staging/ convention was mid-loop-v1 only; v1.1 tasks write artifacts directly to the task folder).
  • Pipeline: new -> research -> research:awaiting_approval -> approve -> research:approved -> implement -> code_review -> code_review:awaiting_approval -> approve -> code_review:approved -> bug_find -> adversarial_bug_find -> doc_review -> referee -> complete. (Skip decomposition/design/test_test_design for simpler v1.1 hardening tasks.)
  • --transition complete moves the task dir to tasks/complete/ automatically (per the relocate feature).
  • Test command: python3 -m pytest tests/ -q. Lint: python3 -m py_compile <file>. No new pip deps; stdlib only.
  • Backlog items deferred past v1.1 are in design/loops/BACKLOG.md (parallel-mode-default, scope-3-self-designing, auto-approve-relax, harness-adapter-spec, multi-budget-currency, compaction-auto-trigger, plus Tier 3 optimizations).

NEXT

Resume task 2 (add-state-loop-lock): SPEC is already written, state is research:awaiting_approval. Approve it, transition to implement, write the _loop_lock helper + wrap callsites in status.py and loop-runner.py, add tests/test_state_loop_lock.py (7 tests per the SPEC), drive to complete. Then proceed to tasks 3-7 in order.

Session 2026-06-25 — Framework agent features design

  • Created design/framework/ with README.md, functional.md, technical.md, BACKLOG.md — design for three framework-level agent features: model-divergence enforcement, rule agents (Proposer + Reviewer), Agent tab redesign.
  • Model-divergence enforcement is a manual task (model-divergence-enforcement), decomposed into 3 subtasks: (1) manifest+detection, (2) interactive enforcement+audit, (3) loop enforcement+dashboard. Not a backlog item.
  • Rule agents (FW-2, FW-3) are backlog items depending on model-divergence shipping first (conflict-of-interest LLM binding). Rule Proposer runs daily, Rule Reviewer runs monthly. Both use direct harness invocation (reuse loop-runner._invoke_harness), not loop infrastructure.
  • Agent tab redesign (FW-1) is a backlog item with no dependencies. Replaces 4 fake AGENT_TYPE_META types with Phase Roles (6 roles from .agent.md) + Scheduled Jobs (real job.kind).
  • Decision: rule agents use direct harness invocation (not loops, not standalone status.py commands).
  • Decision: model-divergence conflict-of-interest is designed now, enforced later — rule agents carry model fields in config but hard-blocking activates only when models.json exists and multi-LLM mode is detected.
  • Cross-references updated: AGENTS.md repo layout, README.md loop config table, design/loops/README.md, design/loops/BACKLOG.md, CHANGELOG.md, .onboarding.md (Backlog section), prompts/onboarding.md (Step 2d).
  • Next: switch on self-improvement loop (--create-loop self-improvement --from-template self-improvement + --install-schedule), then create model-divergence-enforcement parent task and decompose into 3 subtasks.