- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
4.3 KiB
Framework Self-Consistency Tests
Goal
Add automated tests that enforce the framework's own rules — verifying prompt consistency, stop condition presence, canonical paths, and state machine integrity. These tests lock in the fixes from prior tasks and prevent regression.
Requirements
R1. Add tests/test_framework_self_consistency.py
Create a new test file that performs compile-time checks on the framework itself:
A. All delivery prompts have a stop condition block
- Assert every
.mdfile inprompts/that produces a deliverable artifact contains## Stop Condition (MANDATORY)or a documented equivalent - Exclude
orchestrate.md(not a delivery prompt),compaction.md(usesCOMPACTION_COMPLETE),adversarial_bug_find.md(now hasCONTRACT_MET),workflow.md(reference, not prompt)
B. No hardcoded repository URLs in prompts/contracts/templates
- Assert that prompt files don't contain the Gitea URL (
10.37.0.86:3003) or any other hardcoded repo URL install.shat line 13 is the only allowed location (the install script legitimately needs it)
C. .rules.md contains all mandatory rule sections
- Assert
.rules.mdmentions: Task-Driven Development, VRAM, Changelog, Session Discipline, Scope Confinement, Artifact Integrity
D. .rules.md self-improvement rule has concrete examples
- Assert the Self-Improvement section references at least one real failure mode
E. Canonical task path used in all prompts
- Assert all
{project}/.automaton/tasks/{task-name}/paths match the canonical format - Assert zero instances of
{project}/tasks/(the deprecated location) — including concrete task names like{project}/tasks/onboarding/
F. pyproject.toml has no stale extras
- Assert
pyproject.tomldoes not referenceinotify
G. Dashboard CSS theme variables are complete
- Assert both
:rootand[data-theme="light"]sections contain the same set of CSS variable names - This prevents the common bug where a variable is added to one theme but not the other
R2. Add a minimal JS logic test
dashboard.js has 470 lines of untested UI logic. At minimum, test the pure functions:
getTaskDisplayGroup()— review-based group advancementgetFilteredTasks()— filter/sort behaviorSTATE_ICONSmap completeness (matchesTaskStatevalues)
This can be done in Python by parsing the JS file and extracting the function logic, or by adding a small Node.js test with jsdom.
R3. Add verdict parsing regression test
Create a dedicated test file tests/test_parsing.py (or extend test_task.py) with:
- PASS verdict mentioning FAIL → DONE (not BLOCKED) — the regression test for the critical bug
- PASS verdict mentioning NEEDS_REVIEW → DONE (not BLOCKED)
- Structured verdict with
## Status: PASS→ DONE - Structured verdict with
## Status: FAIL→ BLOCKED - Unstructured verdict with just "FAIL" → BLOCKED (fallback behavior)
- Verdict with no status line → RESEARCH (since no SPEC either) or BACKLOG
- IMPLEMENTATION.md alone → BUG_FIND (state machine alignment)
- Empty VERDICT.md → BLOCKED
R4. CI configuration validation
Add a test that parses .gitea/workflows/ci.yml and asserts:
- It runs
py_compileon all Python source directories - It runs
pytest - It runs
bash -non shell scripts
This catches the case where a new directory is added but CI isn't updated.
Acceptance Criteria
python -m pytest tests/test_framework_self_consistency.py -vpasses- All delivery prompts have stop condition blocks (tested by R1.A)
- No hardcoded URLs in prompts (tested by R1.B)
.rules.mdcontains all mandatory sections (tested by R1.C)- Zero deprecated
{project}/tasks/paths in prompts (tested by R1.E) pyproject.tomlhas noinotifyreference (tested by R1.F)tests/test_parsing.pyincludes all regression cases from R3- All existing tests still pass (72/72 minimum)
- CI workflow correctly includes all framework source directories (tested by R4)
Non-Goals
- Not adding a linter/formatter (ruff/black) — framework policy doesn't require one
- Not adding mypy type checking
- Not changing the soft-enforcement philosophy — these tests verify prompts and docs, not runtime behavior
- Not testing the dashboard server integration (too heavy for unit tests)