Files
automaton/tasks/framework-audit/RESEARCH.md
T
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

90 lines
5.6 KiB
Markdown

# Framework Self-Consistency Audit
## Principle Inventory
Extracted from all framework files. Each principle is a rule the framework prescribes for project work.
| # | Principle | Source | Applied to Framework? |
|---|-----------|--------|-----------------------|
| P1 | **Task-driven development**: All changes go through tasks (SPEC → phases → VERDICT) | `onboarding.md`, `workflow.md`, `.rules.md` | ❌ No rule enforces this for framework itself |
| P2 | **VRAM-aware task sizing**: Tasks must fit system context limits; check before scoping | `config.md`, `orchestrate.md:21-101` | ❌ Never checked when creating framework tasks |
| P3 | **Layered filesystem**: Project overrides global, read project first then fallback to global | `orchestrate.md:5-19`, `README.md:177-202` | ⚠️ Broken design — project-first read encourages full copies |
| P4 | **Minimal project footprint**: Projects should only have `.agent.md` + `.rules.md` | `onboarding.md:42`, `README.md:188` | ⚠️ Violated by P3's project-first read order |
| P5 | **No manual task creation**: Orchestrator creates task folders, never the user | `workflow.md:22` | ❌ No rule forbids manual `mkdir tasks/` |
| P6 | **Agent reads rules at startup**: Must read `.agent.md` + `.rules.md` before working | `system-prompt.md`, `session-starter.md` | ⚠️ Doesn't read global `.rules.md`, only project's |
| P7 | **Stop condition enforcement**: "CONTRACT_MET" prevents early termination | `references/stop-hook-pattern.md`, various prompts | ✅ Phase-level prompts have stop conditions |
| P8 | **Self-improving rules**: `.rules.md` is a living document, add rules per failure mode | `.rules.md:3-5` | ❌ No rules were added for any of these gaps |
| P9 | **Customization via extension, not copy**: Override additively, never duplicate | Implicit from design intent | ❌ Not codified anywhere |
| P10 | **Changelog/release notes**: Changes should be recorded for release notes | Not stated anywhere | ❌ No changelog exists |
| P11 | **One-time setup, then task flow**: Onboarding is one-time, normal task flow after | `onboarding.md:95` | ✅ Framework itself doesn't need onboarding |
## Gap Analysis
### Category Definitions
- **Self-reference gap**: Framework doesn't apply rule to itself
- **Missing rule**: Principle isn't codified where agents can read it
- **Enforcement gap**: Rule exists but nothing checks compliance
- **Lifecycle gap**: Feature exists but follow-up step is missing
- **Design flaw**: Architecture encourages violation of own principles
### Gap Details
| # | Principle Violated | Category | Description | Covered By |
|---|--------------------|----------|-------------|------------|
| G1 | P1 (Task-driven) | Self-reference + Missing rule | No rule says "create task before editing framework files" | `framework-self-enforcement` |
| G2 | P2 (VRAM check) | Self-reference + Missing rule | No rule says "check config.md VRAM before scoping tasks" | `framework-self-enforcement` (new rule) |
| G3 | P3/P4 (Layered filesystem) | Design flaw | `orchestrate.md` reads prompts/contracts/scripts from project first, encouraging full copies instead of minimal footprint | `additive-extension-model` |
| G4 | P5 (No manual task creation) | Missing rule + Enforcement | No rule forbids `mkdir tasks/` — tasks should be created by Orchestrator | `framework-self-enforcement` (new rule) |
| G5 | P6 (Agent reads rules) | Missing rule | `system-prompt.md` doesn't instruct agent to read global `.rules.md` | `framework-self-enforcement` |
| G6 | P8 (Self-improving rules) | Enforcement | Gaps G1-G9 exist but no rules were added to prevent recurrence | `framework-self-enforcement` |
| G7 | P9 (Extension model) | Missing rule | No documentation or rule says "additive overrides only, never copy" | `additive-extension-model` |
| G8 | P10 (Changelog) | Lifecycle | VERDICT.md written per-task but no aggregate changelog | `changelog` |
| G9 | None (missing feature) | Missing feature | No project migration path for existing projects with stale copies | `project-migration` |
| G10 | None (missing feature) | Missing feature | No dashboard-based task review/approval workflow | `dashboard-task-review` |
## Impact/Effort Matrix
```
High Impact
│
│ G3 (design flaw) G1 (self-ref)
│ G2 (VRAM check) G5 (rules)
│ G7 (extension doc)
│
│ G10 (review UI) G4 (manual mkdir)
│ G8 (changelog) G6 (living rules)
│ G9 (migration)
│
└─────────────────────────────→
Low Effort High Effort
```
## Task Structure Validation
### Existing tasks vs. gaps covered
| Task | Gaps Covered |
|------|-------------|
| `additive-extension-model` | G3, G7 |
| `framework-self-enforcement` | G1, G2, G4, G5, G6 |
| `changelog` | G8 |
| `project-migration` | G9 |
### New tasks needed
| Task | Gap | Reason for separate task |
|------|-----|-------------------------|
| `dashboard-task-review` | G10 | UI feature, not a rule change. Separate from `framework-self-enforcement` which is about rules/docs only. |
### Merged into `framework-self-enforcement`
G2, G4, G6 are all rule additions to `.rules.md` — they fit naturally in that single task alongside G1 and G5. No need to split further.
## Recommendations
1. **Keep existing 4 tasks as-is** — each covers its gaps cleanly
2. **Add `dashboard-task-review`** as a new task (G10 — user requested feature)
3. **Expand `framework-self-enforcement` spec** to include VRAM check rule (G2), no-manual-creation rule (G4), and living-rules reminder (G6)
4. **Mark `framework-audit` as complete** once RESEARCH.md is written and tasks are validated