Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

15 KiB

SPEC: add-claim-loop-task

Problem

tasks/add-status-brakes/ADVERSARIAL_BUG_REPORT.md A2:

if the loop never current_task-claimed the task, _loop_owning_task returns None and the transition proceeds. The agent can edit a task that isn't claimed by any loop. That is correct behavior (humans and ad-hoc agents can still work), but it means a hostile agent could race the loop runner to claim a task. Mitigation: loop runner should call a --claim-loop-task (not in v1) or set current_task atomically before transitioning.

The runner's cmd_tick currently sets state["current_task"] = current_task (line 714) inside _loop_lock but without cross-loop visibility. Two loops could both claim the same task through concurrent audit work-source dispatch — each loop's _loop_lock is per-loop, so they DON'T serialize across loops. A race scenario:

  1. Loop-alpha audit finds task fix-X → sets state["current_task"] = "fix-X" → writes
  2. Loop-beta audit also finds task fix-X → re-reads state (after loop-alpha wrote) → ALSO sets state["current_task"] = "fix-X" → writes
  3. Both loops now own the same task → both may transition it → state corruption

Task 2's _loop_lock closed same-loop races. This task closes cross-loop races by adding a claim command that checks no OTHER loop already claims the task before setting current_task.

Goal

Add an atomic claim operation that a loop runner calls BEFORE adopting a candidate task. The claim operation:

  1. Scans ALL loops to verify no OTHER loop owns the task
  2. Acquires the claiming loop's per-loop lock
  3. Sets state["current_task"] = task_name
  4. Writes .state.loop

If claim fails (another loop owns the task), the tick skips and picks a different candidate next iteration.

Non-goals

  • Claim timeout / expiry. Task 7 is a write-once-claimed, release-on-complete model. No lease.
  • Forced unclaim. Only the owning loop releases (current_task cleared when task transitions to complete inside the tick flow). Manual escape: --claim-loop-task with --force (separate task; backlog).
  • Claim for non-loop workflows (ad-hoc agents). Human agents still work unconstrained; claim is only checked inside the loop runner, not in --can-edit or --transition (those check _loop_owning_task which returns None for unclaimed tasks — correct, because humans intended to work unclaimed).

Key design

New command: --claim-loop-task <name> --task <taskname>

Invoked by the loop runner. Operates inside the per-loop _loop_lock (same as cmd_tick). Steps:

  1. Scan all loops via _all_loop_dirs() (or the runner provides project_dir).
  2. For each loop whose .state.loop.status == "running" (or "paused"), check if current_task == taskname.
  3. If any OTHER loop (not self) owns the task → return exit 2 with message "Task already claimed by {loop_name}". Stderr only, no state mutation.
  4. If self already owns the task → return exit 0, no-op, success (idempotent re-claim).
  5. If no loop owns the task → set state["current_task"] = taskname, _write_state_loop(...), return exit 0.

Runs inside _loop_lock(loop_path) to serialize concurrent --claim-loop-task against the same loop.

Runner integration

cmd_tick currently does this inside the _loop_lock block:

current_task, skip_reason = _find_work(state, cfg, loop_path, project_dir)
if current_task is None: ...
state["current_task"] = current_task

Replaced by:

current_task, skip_reason = _find_work(state, cfg, loop_path, project_dir)
if current_task is None: ...
claim_ok = _claim_task(loop_path, current_task, state, cfg, project_dir)
if not claim_ok:
    skip_reason = "task_claimed_by_other_loop"
    ... (skip, don't halt)
state["current_task"] = current_task  # still set for downstream tokens

OR as a subprocess call:

claim_rc = _gate(["--claim-loop-task", loop_name, "--task", current_task, ...])
if claim_rc != 0:  skip

Subprocess approach is simpler (reuses status.py as the authority), but it adds another subprocess per tick. In-process approach (new helper) avoids subprocess overhead. Decision: in-process helper _claim_task(loop_path, task_name, state, cfg, project_dir) — since it runs inside the _loop_lock already (wrapping cmd_tick), no new lock needed. The cross-loop scan is un-locked but idempotent (the per-loop lock serializes writes; the scan is a read-only advisory — race window reopens after the scan releases the loop-owning lock, BUT the scan is done inside the claiming loop's OWN lock, and the subsequent state write is atomic. If two loops race to claim the same task, the second loop's lock blocks until the first's _write_state_loop completes; when it re-acquires, its re-read sees the first loop's current_task set and aborts.)

Wait — that's the key insight: with _loop_lock wrapping both the scan and the write, the scan is performed inside the lock. But the scan iterates OTHER loops' .state.loop files — those are NOT locked by the claiming loop's lock. Between the scan (reading other loops' state) and the write, another loop could claim the task. So the subprocess approach that acquires the TARGET task's loop lock would be ideal, but that introduces lock ordering issues.

Simpler: rely on the runner's existing _loop_lock. The claim runs inside the locking loop's lock. The cross-loop scan is advisory: if it finds another loop claiming the task, it refuses. If it finds no one else, it sets current_task. If two loops race, the second loop's lock blocks the write until the first releases, then the second loop re-reads _read_state_loop (which now shows the first loop's current_task). The second loop will detect the conflict on the NEXT iteration (when _find_work re-picks the task, and current_task is already claimed by the first loop in state). The tick simply skips.

This is acceptable: the race window is one tick (_find_work → re-read under lock → re-check). A stale claim on loop 2 is self-healing on the next tick. No corruption.

Better: after cross-loop scan succeeds AND before writing, re-read ALL loops' state under the lock (the scan is done while holding the lock; the re-read captures any concurrent claim from another loop). But this still can't atomically lock all loops.

Final design: use subprocess approach. The runner spawns status.py --claim-loop-task <name> --task <taskname> --project <p>. Inside status.py, cmd_claim_loop_task:

  1. Opens <self_loop_path>/.state.lock and acquires flock.
  2. Re-reads self .state.loop.
  3. Scans all loops (reads each .state.loop without their locks — race possible but self-healing as described above).
  4. If other loop owns it → exit 2 with message.
  5. If self owns it → exit 0.
  6. If nobody owns it → sets state["current_task"] = taskname, _write_state_loop(...), exit 0.

The lock prevents another concurrent --claim-loop-task on the same loop. The cross-loop scan is advisory but the "re-read under self-lock" captures any concurrent write to self's own state.

Release

When does a task get un-claimed? Currently the runner never clears current_task. The task's phase advances to complete via the orchestrator, but current_task stays in .state.loop.

For v1.1, the orchestrator clears current_task when the task reaches complete. The orchestrator's loop-orchestrate.md prompt already says "the orchestrator calls exactly one status.py call (transition, approve, or escalate)". We extend: if the orchestrator transitions the task to a terminal phase (complete or human_intervention), the runner detects this post-orch via state-re-read and clears current_task. Implementation: after the orchestrate subprocess, the runner re-reads the task's phase; if complete or human_intervention, set state["current_task"] = None before the step-10 write.

Requirements

R1 — --claim-loop-task subprocess command

status.py accepts --claim-loop-task <name> --task <taskname> [--project P]. Exit codes:

  • 0 = claimed (or already self-claimed, idempotent)
  • 2 = already claimed by another loop, or untracked loop, or missing task/name Stderr messages:
  • OK or already_self_claimed → exit 0
  • task_already_claimed:{other_loop_name} → exit 2
  • loop_untracked → exit 2

R2 — Runner calls claim before _find_work

In cmd_tick, inside _loop_lock:

  1. After _find_work returns a task candidate (and before setting state["current_task"])
  2. Call _claim_task (subprocess invocation of status.py --claim-loop-task ...)
  3. If exit 0 → proceed (claim is self-no-op if already owned; or new claim registered)
  4. If exit 2 → skip tick with SKIP task_claimed_by_other_loop (do NOT halt; the gate already passed; this is a transient race). The next tick will re-try.

R3 — Release on terminal phase

After step 9 (orchestrate subprocess), before step 10 (_write_state_loop), the runner re-reads the task's .state file. If the phase is complete or human_intervention, set state["current_task"] = None. Write to .state.loop normally.

R4 — Cross-loop ownership check

--claim-loop-task scans all loops via _all_loop_dirs(project) and reads each .state.loop's current_task. If any OTHER loop (name ≠ self) has status == "running" (or "paused") and current_task == taskname, the claim is refused.

Self-ownership check: if self has current_task == taskname, return success (exit 0) without re-writing state (idempotent).

R5 — No race breakage

The cross-loop scan is advisory (not cross-lock). Best-effort: the _loop_lock on the claiming loop serializes writes to self's state. If two loops race, the second's --claim-loop-task blocks on the first's lock; after the first releases, the second re-reads self state and re-scans — seeing the first's current_task → refuses. The second loop's tick skips. Self-healing on next tick.

R6 — No new pip deps

subprocess, json, pathlib, argparse — all stdlib.

Test plan

Tests in tests/test_claim_loop_task.py (NEW). Use tmp_path for loop dirs.

  1. Claim succeeds (no one owns): create 2 loop dirs, .state.loop with current_task: null. Invoke cmd_claim_loop_task for loop1 task fix-X. Assert exit 0. Assert loop1's .state.loop.current_task == "fix-X".
  2. Claim refuses (other loop owns): set loop2's .state.loop.current_task = "fix-X". Claim loop1 for fix-X. Assert exit 2 with task_already_claimed:loop2. Assert loop1's .state.loop.current_task unchanged (null or whatever).
  3. Claim idempotent (self owns): set loop1's current_task = "fix-X". Claim loop1 for same task. Assert exit 0. Assert no state re-written (check mtime unchanged).
  4. Claim on untracked loop: no .state.loop file. Assert exit 2.
  5. Missing task arg: invoke cmd_claim_loop_task without --task. Assert error message + exit 2.
  6. Release on complete: in runner flow, after orchestrate, mock task .state as complete. Assert state["current_task"] = None.
  7. Release on human_intervention: same as R6 but phase human_intervention. Assert current_task = None.
  8. Release does NOT fire on implement phase: task in implement, assert current_task stays as-is.
  9. Cross-loop self-healing race: create two loops, set up race condition (loop2's state shows current_task = "fix-X" but the .state.loop file was written by a concurrent thread). Claim loop1 → refuses. Then remove loop2's claim, re-claim loop1 → succeeds.
  10. Claim on paused loop allowed: loop is paused but state["status"] == "paused"; claim should succeed (paused loop still owns its current_task).
  11. Runner integration: mock --claim-loop-task subprocess in cmd_tick; assert tick skips when exit 2, proceeds when exit 0.
  12. Runner release integration: mock .state file as complete; assert state["current_task"] cleared after step 9.

Decisions

  • D-C1: Claim is a status.py subprocess, not an in-process helper. Keeps status.py as the single authority for loop state. Avoids duplicating _all_loop_dirs / _read_state_loop scanning logic into the runner.
  • D-C2: Cross-loop scan is advisory (no cross-loop lock). Self-healing on next tick. Acceptable for v1.1: the race window is one tick, and the tick simply skips — no state corruption.
  • D-C3: paused loops retain their current_task claim. A resumed loop resumes work without re-claiming. Consistent with "paused = temporary stop, not release".
  • D-C4: halted loops' claim persists. Operator must --approve --loop to resume; the task remains claimed. No stealth unclaim on halt.
  • D-C5: Release on terminal phase (complete/human_intervention) is the runner's responsibility, not the orchestrator's. The orchestrator just calls --transition. The runner re-reads the task state after the orchestrator subprocess and clears current_task if terminal. This avoids coupling the orchestrator prompt to the current_task lifecycle.
  • D-C6: Runner clears current_task in the same _write_state_loop call that writes iteration_count++. Atomic: if writing fails, the next tick retries the orchestrate step (idempotent).
  • D-C7: _find_work still returns state.get("current_task"). The claim command SETS current_task, and the release flow CLEARS it. _find_work itself doesn't change.

Runner flow changes (cmd_tick, inside _loop_lock)

  7. parse verdict (unchanged)
  8. cap score_history (unchanged)
  9. spawn Orchestrate (unchanged)
  9.5 re-read task state; if terminal → current_task = None   ← NEW (R3)
  10. advance state (unchanged: iteration_count++ + write)

And for the claim path (steps 3-5):

  3. find_work (unchanged — returns candidate task or None)
  3.5 if candidate is not None AND candidate ≠ state.get("current_task"):
        claim_ok = _claim_subprocess(name, candidate, project_dir)  ← NEW (R1-R2)
        if not claim_ok:
            skip tick "task_claimed_by_other_loop"
  4. ensure worktree (unchanged)

Files touched

  • scripts/status.py — add cmd_claim_loop_task(args); add --claim-loop-task arg; add _claim_loop_task_impl(...) (the scanning logic).
  • scripts/loop-runner.py — in cmd_tick step 3-3.5: subprocess claim; step 9.5: release.
  • CHANGELOG.md — new entry under [unreleased].
  • design/loops/technical.md §7 — update tick-flow table for steps 3.5 (claim) and 9.5 (release).
  • design/loops/functional.md — add claim semantics to the loop lifecycle.
  • tests/test_claim_loop_task.py (NEW) — 12 tests per plan above.

Out of scope

  • --force flag to override another loop's claim (separate task; backlog).
  • Claim-then-stale detection (loop halts while claiming a task; the task stays claimed forever). Future: --audit could flag loops that are halted/non-existing while current_task is set.
  • --release-loop-task subcommand (release is automatic via terminal phase; operator escape is --claim-loop-task --force or manual current_task = None edit).
  • Claim status in --loop-list output. Future UX improvement.

Pipeline plan

research → research:awaiting_approval → research:approved → implement → code_review → code_review:awaiting_approval → code_review:approved → bug_find → adversarial_bug_find → doc_review → referee → complete.