- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
2.6 KiB
2.6 KiB
Implementation: State File Enforcement
Changes Made
1. .state file format (implemented in scripts/status.py)
.statefile contains a single phase name (e.g.,research,research:awaiting_approval,implement)- Written atomically via
.state.tmp→.staterename .state.approvalsappend-only log records all approvals with timestamp and approver.state.lockfor multi-agent claiming (optional, only in multi-agent mode)- Non-artifact metadata files (
.state,.state.approvals,.state.lock,.state.tmp,VRAM_CONFIG.md,PARENT_SPEC.md,REVIEW.md) are excluded from artifact checks
2. State transitions with approval sub-states
- Approval-gated phases: research, decomposition, design, test_design now have
:awaiting_approval→:approvedsub-states - Non-approval phases: implement, bug_find, adversarial_bug_find, doc_review, referee have no sub-states
status.py --transitionrefuses transitions past:awaiting_approvalwithout--approvestatus.py --approvetransitions:awaiting_approval→:approvedand records approval in.state.approvals
3. Backward compatibility
- If
.statedoesn't exist,status.pyinfers phase from artifacts and writes.state upgrade.shbootstraps.statefor all existing tasks- Phase prompts work with or without
.state(warns if missing)
4. Forbidden artifacts per phase (implemented in status.py --validate-folder)
- Each phase has a defined set of artifacts that must NOT exist (artifacts from future phases)
--validate-folderchecks and reports violations--transitionrefuses to proceed if forbidden artifacts exist
5. Task creation gate (implemented in status.py --create-task)
--create-taskcreates task folder with.state=newand empty.state.approvals- Validates kebab-case task names
- Refuses if task already exists
--auditCategory 4 flags manually created task folders
6. Orchestrator and workflow updates
prompts/workflow.mdrewritten:.stateis canonical, approval sub-states documented,status.pycommands referencedprompts/orchestrate.mdreduced from 493 to 143 lines, referencesworkflow.mdandsubtask_management.md- All phase prompts include
.stateprecondition check
Files Modified
scripts/status.py(new, 980 lines)prompts/workflow.md(rewritten)prompts/orchestrate.md(rewritten, 143 lines)prompts/subtask_management.md(new, extracted from orchestrate.md)tests/test_status.py(new, 25 tests)scripts/upgrade.sh(new)
Test Results
- 183 tests passing (including 25 new status.py tests)
- Python compilation clean
- Shell script syntax clean