Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

3.2 KiB

SPEC: Comprehensive Framework Self-Consistency Audit

Motivation

Several gaps were found where the framework doesn't apply its own principles to itself:

  • No task-driven enforcement: Framework prescribes task-driven development for projects but has no rules that enforce it for framework changes
  • No resource check before task scoping: VRAM detection exists but nothing ensures tasks are sized to fit system context limits
  • No changelog/release notes: VERDICT.md exists per-task but no aggregate change history
  • No migration path: Framework evolved but existing projects have no cleanup process

These aren't random bugs — they follow a pattern. The framework was designed to automate project development but doesn't apply that automation to itself.

Goal

Systematically audit the entire framework to find every place where it fails to follow its own design principles. Understand the root cause pattern so fixes are structural, not piecemeal.

Method

Step 1: Extract All Design Principles

Read every file in the framework and extract explicit and implicit design principles:

  • system-prompt.md — agent instructions
  • .agent.md — routing rules
  • .rules.md — project rules
  • prompts/orchestrate.md — orchestrator behavior
  • prompts/onboarding.md — project initialization
  • prompts/workflow.md — workflow state machine
  • config.md — configuration rules
  • README.md — documented principles
  • scripts/*.sh — automation scripts
  • automaton/dashboard/ — dashboard design
  • references/*.md — reference docs

Step 2: Self-Consistency Check

For each principle, ask: "Does the framework apply this to itself?"

Principle Applied to projects? Applied to framework? Gap?
Task-driven development Yes (onboarding.md) No YES
VRAM-aware task sizing Yes (config.md) No YES
Layered filesystem Yes (orchestrate.md) N/A (framework is the base layer) ?
Changelog/release notes Not documented No YES
... (find all)

Step 3: Categorize Gaps

For each gap, identify which category it falls into:

  1. Self-reference gap: Framework doesn't apply its rule to itself
  2. Missing rule: Principle exists in one place but isn't codified where agents read it
  3. Enforcement gap: Rule exists but nothing checks compliance
  4. Lifecycle gap: Feature exists (task completion) but follow-up step is missing (changelog, migration)

Step 4: Prioritize Fixes

Rank gaps by impact and effort. Flag which should be in the existing 4 tasks vs which need new tasks.

Acceptance Criteria

  • All design principles extracted and documented
  • All gaps identified with root cause category
  • Gaps prioritized with impact/effort estimate
  • Existing 4 tasks validated or adjusted based on findings
  • New tasks created for any gaps not already covered

Output

The audit produces tasks/framework-audit/RESEARCH.md containing:

  1. Complete principle inventory
  2. Gap analysis with root cause categories
  3. Prioritized action items
  4. Recommended task structure

Context

  • 46GB RAM, 16-core AMD CPU, no active GPU driver
  • Target context: 16k tokens, 25% headroom, 12k peak per sub-task
  • Framework location: ~/.automaton/