Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test

- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
This commit is contained in:
Lap Tran
2026-06-26 10:05:18 -04:00
parent fe43b9e1fc
commit bc7daf8590
666 changed files with 15994 additions and 69 deletions
+28
View File
@@ -0,0 +1,28 @@
# BUG_REPORT: add-goal-mode
Probed goal-mode work sources, token substitution, and audit --json against edge cases.
## Bugs found
None blocking. Informational observations below.
## Observations (non-blocking)
### O1 -- `_find_work_audit` create-task subprocess is fire-and-forget
When a violation has no `task` field, the runner calls `status.py --create-task <slug>` with `timeout=15` and swallows all exceptions. If the create-task fails (e.g. disk full, permission error), the runner returns the slug anyway. The next tick's `--audit` will see the same violation (still no task dir) and try again. Self-healing on next tick. Accepted for v1.
### O2 -- `_find_work_backlog` bold-marker regex is strict
The regex `\*\*([A-Za-z0-9._-]+)\*\*` requires the bold text to be a valid slug (alphanumerics, dots, hyphens, underscores only). A backlog item like `- [ ] **fix user auth**` would fail the regex and fall back to `_slugify("fix user auth")` -> `fix-user-auth`. This is correct behavior but worth noting: the bold marker is a convention, not a requirement. Accepted.
### O3 -- `--audit --json` violations lack `resolved: true` entries
The audit collector only emits unresolved violations (those with actual defects). Resolved violations are not included in the JSON output. This is correct for the runner's use case (it filters on `not v.get("resolved", False)` anyway), but a consumer expecting a full audit history would need the human-readable `--audit` output instead. Accepted.
### O4 -- Token substitution tests require custom harness command
The R4 tests (`test_task_brief_substituted_from_research`, `test_acceptance_criteria_substituted_from_loop_json_list`, `test_next_hint_substituted_from_last_verdict`) use a custom `harness.command` that includes the token placeholders. The default harness command (`opencode run --prompt-file {prompt} --cwd {cwd}`) does not contain `{task_brief}` etc., so the tokens are only useful when a loop config explicitly adds them to its harness command. This is by design (SPEC R4: "tokens absent from the prompt stay literal"). The actual prompt files (task 6) will need to either reference these tokens or the harness command will need to pass them as CLI args. Accepted.
### O5 -- `_truncate_tokens` marker length can exceed budget by 1
The marker is ` ...[truncated]` (14 chars with leading space). The code does `text[:char_budget - len(_TRUNCATE_MARKER)]` + marker. If `char_budget` is smaller than `len(_TRUNCTATE_MARKER)`, the slice goes negative and Python returns the whole string (not empty). For `max_tokens=1` (budget=4), the result would be the full text + marker. This only happens with absurdly small token budgets (the real caps are 1000-4000). Not blocking. Noted for v1.1 hardening: clamp `char_budget` to `len(marker) + 1` minimum.
## Verdict
PASS -- no blocker bugs. All observations are accepted trade-offs or v1.1 hardening items.