Files
automaton/tasks/status-script/BUG_REPORT.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

27 lines
1.5 KiB
Markdown

# Bug Report: Status Script
## Methodology
Reviewed the 980-line status.py implementation and 25 passing tests. Verified all commands, state transitions, validation logic, and error handling.
## Acceptance Criteria
| # | Criterion | Result |
|---|-----------|--------|
| 1 | `--task` shows current phase, allowed/forbidden actions | ✅ |
| 2 | `--transition` validates and writes `.state` | ✅ |
| 3 | `--transition` refuses past `:awaiting_approval` | ✅ |
| 4 | `--approve` transitions awaiting → approved | ✅ |
| 5 | `--approve` refuses if not awaiting approval | ✅ |
| 6 | `--create-task` creates folder with `.state` = `new` | ✅ |
| 7 | `--list` shows all tasks with phase and approval status | ✅ |
| 8 | `--validate-folder` checks out-of-order artifacts | ✅ |
| 9 | `--audit` runs all check categories | ⚠️ See finding 1 |
| 10 | `--claim`/`--release`/`--next-available`/`--available` | ✅ |
| 11 | `--can-edit`/`--scope-check`/`--same-session` | ✅ |
| 12 | Atomic writes (tmp+rename) | ✅ |
| 13 | `.state.approvals` log maintained | ✅ |
| 14 | 25 tests passing | ✅ |
## Findings
1. **Minor**: `--audit` Category 3 (git modification check) is stubbed — it checks whether the project is a git repo and reports "Skipped: not a git repository" or falls back to a simplified check. Full git-log-based timestamp comparison is not implemented. This is documented as part of the spec's "graceful skip" clause but the implementation is simpler than the spec describes.
## Verdict: PASS