Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

12 KiB
Raw Permalink Blame History

Fix Harness Command Template

The loop runner's default harness.command uses opencode run --prompt-file {prompt} --cwd {cwd}, but opencode run has no --prompt-file flag and no --cwd flag. The actual flags are --dir (cwd equivalent) and the message passed as a positional. The v1 runner has only been exercised via unit tests with a mocked subprocess (fake_run stub matches on --prompt-file), so the bug was never caught. A real --mode tick invocation against a live harness fails immediately.

This is a v1.1 correctness fix, not a feature. Without it, the entire loop runtime is non-functional out of the box.

Goal

Make the default harness.command in loop-runner.py actually invokable. Introduce a {prompt_content} substitution token that carries the resolved prompt file's text as a single argv element (safe under subprocess.run list mode — no shell parsing). Switch the default to use --dir and the positional message.

Root cause

loop-runner.py:324,328:

command = ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]

opencode run --help confirms available flags: --dir, --model, -f/--file, --format, --agent. No --prompt-file. No --cwd. The command would exit with a usage error on first real invocation.

The unit tests (tests/test_loop_runner.py) mock subprocess.run via a fake_run fixture that matches on --prompt-file as a generic stub rule (fake_run.add_simple("--prompt-file", "")). The mock never validates that the flag exists in the real opencode CLI.

Requirements

R1. New substitution token: {prompt_content}

In _invoke_harness (loop-runner.py), after resolving the prompt to a temp file path via _resolve_prompt, read the file's text content and substitute a new {prompt_content} token with it. The content becomes a single argv element in the final command list. Since subprocess.run is invoked with a list (no shell=True), no quoting/escaping is needed — the full prompt text is passed as one argv element regardless of content.

{prompt} (file path) remains available as a separate token for users who prefer to pass the file via -f attachment or a custom harness that reads files.

R2. New default harness command

Replace both fallback paths (loop-runner.py:324 for harness_cfg is None and loop-runner.py:328 for empty command in config) with:

command = ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]

This passes:

  • --dir {cwd} — the working directory for the spawned opencode process.
  • {prompt_content} — the full prompt text as the positional message argument.

The spawned opencode run process receives the prompt as its message, runs non-interactively, produces stdout, and exits. The runner captures stdout as before.

R3. Optional --model in the default

The default command does NOT hardcode a --model flag. The spawned opencode run inherits the model from the project/user config (opencode.json). Users who want a different model per loop (e.g. local Qwen for implement, subscription model for verify) override harness.command in their loop.json:

"harness": {
  "command": ["opencode", "run", "--model", "local-mlx/AEON-7/Qwen3.6-27B-...", "--dir", "{cwd}", "{prompt_content}"]
}

Per-role model override (if needed later) is a separate feature; out of scope for this fix.

R4. Per-role harness command override

The current code reads a single harness.command from loop.json and applies it to all three roles. The roles.<role>.harness override pattern is not added in this task — it's a feature, not a fix. The single harness.command applies to all roles. If a user wants per-role models, they can use different harness.command entries only after we add per-role override (future task). For now, one command for all roles.

R5. Update test stubs

The fake_run fixture in tests/test_loop_runner.py matches on --prompt-file as a generic stub rule. After the fix, the default command no longer contains --prompt-file. Update:

  • fake_run.add_simple("--prompt-file", "") → fake_run.add_simple("--dir", "") or a more generic matcher that catches the default opencode run shape. The stub should match on "opencode" as the binary name, or on --dir as a flag.
  • Any test assertions that check for --prompt-file in invocations → update to check for --dir and the prompt content positional.
  • The custom-command test (TestHarnessSubstitution.test_custom_command_with_output_token) uses --cwd in the custom command — that's the user's custom command, not the default, so it stays as-is (users can use whatever flags their harness supports).

R6. Update design doc

design/loops/technical.md §7 (lines 211, 218, 222, 262) references the old default ["opencode", "run", "--prompt-file", "{prompt}", "--cwd", "{cwd}"]. Update to the new default and document the {prompt_content} token alongside the existing {prompt}, {cwd}, {output}, {artifact} tokens.

R7. No breaking change to custom harness commands

Users with existing loop.json files that set a custom harness.command using {prompt} (file path) and {cwd} tokens continue to work. The {prompt} and {cwd} tokens are still populated by the substitution mapping. Only the default (when no harness.command is set) changes.

Harness agnosticism

The framework's harness contract (design/loops/functional.md §13, design/loops/technical.md §8, contracts/harness-integration.md):

  • The shape is generic: the runner substitutes tokens into whatever harness.command the user configures in loop.json. The runner core has zero knowledge of which harness is invoked.
  • The default is opencode-specific by design: the framework dogfoods opencode (D24). Users override harness.command for any other harness.
  • No harness/model inspection (D8): the framework never inspects harness type, model capability, size, or provider. The harness.command string is opaque to the runner; it just substitutes tokens and invokes.
  • Concrete adapters out of scope for v1 (BACKLOG.md: harness-adapter-spec deferred). The generic harness.command covers all harnesses that can (a) run a session against a given prompt and (b) write the resulting artifact to stdout.

This fix preserves and extends that contract:

  • Preserves: {prompt} (file path), {cwd}, {output}, {artifact} tokens still work; custom commands using them are unchanged (R7).
  • Extends: new {prompt_content} token (R1) carries the resolved prompt's text as a single argv element, enabling harnesses that prefer a message argument over a file path. This makes the framework more harness-agnostic than v1, not less.
  • No new harness awareness: the runner core still does not know which harness is invoked. The opencode-specific shape lives only in the default command string, which is overridable.

Examples — harness.command overrides in loop.json

// opencode (DEFAULT — no override needed; shown for clarity)
"harness": {"command": ["opencode", "run", "--dir", "{cwd}", "{prompt_content}"]}

// Pi Dev — pi binary; adjust flags to match `pi run --help`
"harness": {"command": ["pi", "run", "--cwd", "{cwd}", "{prompt_content}"]}

// Pi Dev — alternative shape if pi prefers a prompt file
"harness": {"command": ["pi", "run", "--cwd", "{cwd}", "--prompt-file", "{prompt}"]}

// aider — message argument, no file
"harness": {"command": ["aider", "--message", "{prompt_content}", "--yes"]}

// aider — alternative using a prompt file
"harness": {"command": ["aider", "--message-file", "{prompt}", "--yes"]}

// Cursor / Copilot / Cline — depends on each tool's CLI; same override pattern
"harness": {"command": ["cursor", "--cwd", "{cwd}", "--prompt-file", "{prompt}"]}

// Generic — any tool that reads prompt from stdin via a shell wrapper
"harness": {"command": ["sh", "-c", "cat {prompt} | my-tool --cwd {cwd}"]}

The Pi Dev examples are illustrative — the actual pi run flags depend on Pi Dev's CLI, which the user confirms against pi run --help on their machine. The point is that any harness can be wired in via this override; the runner does not care.

What this fix does NOT change about harness agnosticism

  • The runner core remains harness-agnostic (token substitution only).
  • D8 (no model/provider inspection) is preserved.
  • The contracts/harness-integration.md enforcement matrix (pre-edit/pre-commit/pre-push hooks, prompt rules per harness) is unaffected — this fix is about the loop tick harness invocation, not the pre-edit guard layer.
  • The plugins/automaton-guard-pi/ plugin (Pi Dev pre-edit guard) is unaffected.

Non-goals

  • No per-role harness command override (R4 explains why).
  • No per-role model selection (needs R4 first).
  • No opencode run --format json integration for machine-readable harness output (future; the verifier parses stdout as before).
  • No change to _resolve_prompt (temp file creation stays; the file is still created because {prompt} token users need the path and the runner needs a stable artifact path for the tick output dir).
  • No Pi Dev CLI probing or auto-detection — the user configures harness.command for their Pi Dev invocation; the framework does not detect or special-case Pi Dev.

Test plan (tests/test_harness_command.py — new, or extend tests/test_loop_runner.py)

  1. test_default_command_uses_dir_not_cwd: invoke _invoke_harness with harness_cfg=None; assert the final argv contains --dir and does NOT contain --cwd or --prompt-file.

  2. test_default_command_passes_prompt_content: invoke _invoke_harness with harness_cfg=None and a prompt file containing "hello world"; assert the final argv contains "hello world" as a positional element (not as a file path).

  3. test_prompt_content_handles_special_chars: prompt file contains "hello 'world' with $vars and \"quotes\""; assert the content appears as a single argv element (no shell expansion, no splitting).

  4. test_prompt_token_still_available: custom command ["cat", "{prompt}"] still receives the temp file path (backwards compat).

  5. test_custom_command_with_cwd_still_works: custom command using {cwd} still gets cwd substituted (backwards compat).

  6. test_empty_command_falls_back_to_new_default: harness_cfg={"command": []} falls back to the new default (not the old one).

  7. test_tick_with_new_default_completes: end-to-end tick test using the new default; fake_run stub matches opencode binary and returns canned stdout for each role. Assert tick completes with verdict and iteration increment.

  8. test_pi_shaped_command_substitutes_correctly: configure harness.command as ["pi", "run", "--cwd", "{cwd}", "{prompt_content}"] (Pi Dev example from the Harness agnosticism section). Invoke _invoke_harness with a prompt file containing "implement the lock". Assert the final argv is ["pi", "run", "--cwd", "<path>", "implement the lock"] — proving the substitution mechanism works for a non-opencode harness with no runner changes. The pi binary is never actually invoked (mocked via fake_run); this test validates token substitution, not pi's CLI.

Update existing tests:

  • test_tick_pass: change fake_run.add_simple("--prompt-file", "") to match the new default shape.
  • Any other test that stubs the harness via --prompt-file.

D-items

  • D-H1: {prompt_content} is a single argv element, not shell-expanded. Safe under subprocess.run list mode.
  • D-H2: {prompt} (file path) remains for backwards compat and file-attachment use cases.
  • D-H3: default does not hardcode --model; inherits from opencode config.
  • D-H4: no per-role override in this task (single harness.command for all roles).

Risks

  • Argv length: very large prompts (>128KB) could hit OS argv limits. Prompts in this framework are typically 2–10KB. Acceptable; document the limit in the helper docstring.
  • Test mock drift: the fake_run fixture now mocks a different default shape. If opencode's CLI flags change again in the future, the mock won't catch it. Mitigation: a separate smoke test that shells out to opencode run --help and asserts --dir exists (skip if opencode not on PATH). Add as an optional test marked @pytest.mark.skipif(not shutil.which("opencode")).

Verification

  • python3 -m py_compile scripts/loop-runner.py
  • python3 -m pytest tests/test_loop_runner.py -v
  • python3 -m pytest tests/ -q (full suite must remain green; 433 baseline)
  • Manual smoke (if opencode on PATH): opencode run --dir /tmp "echo hello" — confirm non-interactive execution produces stdout and exits.