- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
5.6 KiB
5.6 KiB
Framework Self-Consistency Audit
Principle Inventory
Extracted from all framework files. Each principle is a rule the framework prescribes for project work.
| # | Principle | Source | Applied to Framework? |
|---|---|---|---|
| P1 | Task-driven development: All changes go through tasks (SPEC → phases → VERDICT) | onboarding.md, workflow.md, .rules.md |
❌ No rule enforces this for framework itself |
| P2 | VRAM-aware task sizing: Tasks must fit system context limits; check before scoping | config.md, orchestrate.md:21-101 |
❌ Never checked when creating framework tasks |
| P3 | Layered filesystem: Project overrides global, read project first then fallback to global | orchestrate.md:5-19, README.md:177-202 |
⚠️ Broken design — project-first read encourages full copies |
| P4 | Minimal project footprint: Projects should only have .agent.md + .rules.md |
onboarding.md:42, README.md:188 |
⚠️ Violated by P3's project-first read order |
| P5 | No manual task creation: Orchestrator creates task folders, never the user | workflow.md:22 |
❌ No rule forbids manual mkdir tasks/ |
| P6 | Agent reads rules at startup: Must read .agent.md + .rules.md before working |
system-prompt.md, session-starter.md |
⚠️ Doesn't read global .rules.md, only project's |
| P7 | Stop condition enforcement: "CONTRACT_MET" prevents early termination | references/stop-hook-pattern.md, various prompts |
✅ Phase-level prompts have stop conditions |
| P8 | Self-improving rules: .rules.md is a living document, add rules per failure mode |
.rules.md:3-5 |
❌ No rules were added for any of these gaps |
| P9 | Customization via extension, not copy: Override additively, never duplicate | Implicit from design intent | ❌ Not codified anywhere |
| P10 | Changelog/release notes: Changes should be recorded for release notes | Not stated anywhere | ❌ No changelog exists |
| P11 | One-time setup, then task flow: Onboarding is one-time, normal task flow after | onboarding.md:95 |
✅ Framework itself doesn't need onboarding |
Gap Analysis
Category Definitions
- Self-reference gap: Framework doesn't apply rule to itself
- Missing rule: Principle isn't codified where agents can read it
- Enforcement gap: Rule exists but nothing checks compliance
- Lifecycle gap: Feature exists but follow-up step is missing
- Design flaw: Architecture encourages violation of own principles
Gap Details
| # | Principle Violated | Category | Description | Covered By |
|---|---|---|---|---|
| G1 | P1 (Task-driven) | Self-reference + Missing rule | No rule says "create task before editing framework files" | framework-self-enforcement |
| G2 | P2 (VRAM check) | Self-reference + Missing rule | No rule says "check config.md VRAM before scoping tasks" | framework-self-enforcement (new rule) |
| G3 | P3/P4 (Layered filesystem) | Design flaw | orchestrate.md reads prompts/contracts/scripts from project first, encouraging full copies instead of minimal footprint |
additive-extension-model |
| G4 | P5 (No manual task creation) | Missing rule + Enforcement | No rule forbids mkdir tasks/ — tasks should be created by Orchestrator |
framework-self-enforcement (new rule) |
| G5 | P6 (Agent reads rules) | Missing rule | system-prompt.md doesn't instruct agent to read global .rules.md |
framework-self-enforcement |
| G6 | P8 (Self-improving rules) | Enforcement | Gaps G1-G9 exist but no rules were added to prevent recurrence | framework-self-enforcement |
| G7 | P9 (Extension model) | Missing rule | No documentation or rule says "additive overrides only, never copy" | additive-extension-model |
| G8 | P10 (Changelog) | Lifecycle | VERDICT.md written per-task but no aggregate changelog | changelog |
| G9 | None (missing feature) | Missing feature | No project migration path for existing projects with stale copies | project-migration |
| G10 | None (missing feature) | Missing feature | No dashboard-based task review/approval workflow | dashboard-task-review |
Impact/Effort Matrix
High Impact
│
│ G3 (design flaw) G1 (self-ref)
│ G2 (VRAM check) G5 (rules)
│ G7 (extension doc)
│
│ G10 (review UI) G4 (manual mkdir)
│ G8 (changelog) G6 (living rules)
│ G9 (migration)
│
└─────────────────────────────→
Low Effort High Effort
Task Structure Validation
Existing tasks vs. gaps covered
| Task | Gaps Covered |
|---|---|
additive-extension-model |
G3, G7 |
framework-self-enforcement |
G1, G2, G4, G5, G6 |
changelog |
G8 |
project-migration |
G9 |
New tasks needed
| Task | Gap | Reason for separate task |
|---|---|---|
dashboard-task-review |
G10 | UI feature, not a rule change. Separate from framework-self-enforcement which is about rules/docs only. |
Merged into framework-self-enforcement
G2, G4, G6 are all rule additions to .rules.md — they fit naturally in that single task alongside G1 and G5. No need to split further.
Recommendations
- Keep existing 4 tasks as-is — each covers its gaps cleanly
- Add
dashboard-task-reviewas a new task (G10 — user requested feature) - Expand
framework-self-enforcementspec to include VRAM check rule (G2), no-manual-creation rule (G4), and living-rules reminder (G6) - Mark
framework-auditas complete once RESEARCH.md is written and tasks are validated