Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

1.5 KiB

Adversarial Bug Report: Multi-Agent Support

Deep Review

The multi-agent system is well-designed for file-system-based coordination. Single-agent mode has no overhead. Claim/release uses atomic writes. Work discovery correctly prioritizes tasks closer to completion.

Potential Issues

  1. Agent identity is self-reported: --agent is a command-line flag with no authentication. Any agent can claim to be any agent-id. In a trusted environment (single machine, same user), this is fine. In adversarial or distributed scenarios, this would need cryptographic signing.

  2. Lock file race on NFS/Linux: The atomic rename pattern (.state.lock.tmp → .state.lock) is atomic on local filesystems but may not be atomic on NFS. The spec explicitly scopes this out ("file-based locks are sufficient for local agent coordination").

  3. Expired lock window: Between lock expiry and overclaiming, there's a window where two agents could both see an expired lock and both try to claim. The atomic write pattern means only one wins, but the loser gets an error rather than a graceful retry message.

  4. No lock inheritance on sub-task creation: When the coordinator creates a sub-task via --create-task, the sub-task is unclaimed by default. The coordinator must explicitly claim it on behalf of an agent. This is correct behavior but could be surprising.

Verdict: PASS — the self-reported identity is a known design choice (trusted environment), not a security vulnerability in the intended threat model.