Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

1.6 KiB

Bug Report: project-scoping-enforcement

Summary

Critical process violation: the agent performed all implementation work before creating a task, completely bypassing the framework's workflow enforcement.

Bugs Found

Bug 1: No framework self-enforcement prevents untasked work (Critical)

  • Severity: Critical
  • Location: Agent behavior, not code
  • Description: The agent identified 7 scoping issues, then directly implemented all fixes across 20+ files without first creating a task through status.py --create-task. The task was only created after all work was done, as a retrospective documentation exercise.
  • Reproduction: Any agent session where the user asks for work to be done. Nothing prevents the agent from editing files directly.
  • Suggested Fix: This is a behavioral fix, not a code fix. The agent should always create a task first for any non-trivial work, then implement within that task's phase constraints.

Bug 2: Process gap — no automated check that edits have a corresponding task

  • Severity: Medium
  • Location: Framework enforcement model
  • Description: status.py --can-edit only checks if a task is in the right phase for code edits. But it doesn't verify that the files being edited are within that task's scope. An agent can create task "foo" for project A, then edit files in project B without any task at all.
  • Suggested Fix: Future enhancement — --can-edit could optionally check that the files being modified are relevant to the task's SPEC.md or DESIGN.md scope.

Score

+10 (Bug 1 is a process violation worth documenting; Bug 2 is a future enhancement)