Archive completed tasks, add cleanup commands, self-documenting dashboard UI
CI / build (push) Has been cancelled

- Archive 79 completed framework-dev tasks from tasks/ -> tasks/complete/
- status.py: add --cleanup-done and --install-cleanup-schedule commands
- Add scripts/automaton-cleanup.sh for periodic task archiving
- Dashboard: rename 'Background' tab -> 'Agent', 'Cleanup' agent -> 'Completed Task Archiver', remove redundant group headers and pill badges, dim inactive agent placeholders
- .rules.md: add Self-Documenting UI Names rule
- New tests: test_cleanup_done.py, expanded test_app.py and test_task.py
This commit is contained in:
Lap Tran
2026-06-24 22:43:33 -04:00
parent e13513faaa
commit 4a2301b077
572 changed files with 856 additions and 101 deletions
-1
View File
@@ -1 +0,0 @@
complete
-18
View File
@@ -1,18 +0,0 @@
# Bug Report: Framework Self-Consistency Audit
## Methodology
Reviewed RESEARCH.md for completeness against SPEC.
## Acceptance Criteria
| # | Criterion | Result |
|---|-----------|--------|
| 1 | All principles extracted | ✅ (11 principles) |
| 2 | All gaps identified with root cause | ✅ (10 gaps, G1-G10) |
| 3 | Gaps prioritized | ✅ (impact/effort matrix) |
| 4 | Tasks validated | ✅ (existing + new task created) |
| 5 | New tasks for uncovered gaps | ✅ (dashboard-task-review) |
## Findings
None.
## Verdict: PASS
-12
View File
@@ -1,12 +0,0 @@
# Doc Review: Framework Self-Consistency Audit
## Documents Checked
| Doc | Status |
|-----|--------|
| RESEARCH.md | ✅ Comprehensive audit |
| CHANGELOG.md | ✅ Entry added |
## Findings
None.
## Verdict: PASS
-89
View File
@@ -1,89 +0,0 @@
# Framework Self-Consistency Audit
## Principle Inventory
Extracted from all framework files. Each principle is a rule the framework prescribes for project work.
| # | Principle | Source | Applied to Framework? |
|---|-----------|--------|-----------------------|
| P1 | **Task-driven development**: All changes go through tasks (SPEC → phases → VERDICT) | `onboarding.md`, `workflow.md`, `.rules.md` | ❌ No rule enforces this for framework itself |
| P2 | **VRAM-aware task sizing**: Tasks must fit system context limits; check before scoping | `config.md`, `orchestrate.md:21-101` | ❌ Never checked when creating framework tasks |
| P3 | **Layered filesystem**: Project overrides global, read project first then fallback to global | `orchestrate.md:5-19`, `README.md:177-202` | ⚠️ Broken design — project-first read encourages full copies |
| P4 | **Minimal project footprint**: Projects should only have `.agent.md` + `.rules.md` | `onboarding.md:42`, `README.md:188` | ⚠️ Violated by P3's project-first read order |
| P5 | **No manual task creation**: Orchestrator creates task folders, never the user | `workflow.md:22` | ❌ No rule forbids manual `mkdir tasks/` |
| P6 | **Agent reads rules at startup**: Must read `.agent.md` + `.rules.md` before working | `system-prompt.md`, `session-starter.md` | ⚠️ Doesn't read global `.rules.md`, only project's |
| P7 | **Stop condition enforcement**: "CONTRACT_MET" prevents early termination | `references/stop-hook-pattern.md`, various prompts | ✅ Phase-level prompts have stop conditions |
| P8 | **Self-improving rules**: `.rules.md` is a living document, add rules per failure mode | `.rules.md:3-5` | ❌ No rules were added for any of these gaps |
| P9 | **Customization via extension, not copy**: Override additively, never duplicate | Implicit from design intent | ❌ Not codified anywhere |
| P10 | **Changelog/release notes**: Changes should be recorded for release notes | Not stated anywhere | ❌ No changelog exists |
| P11 | **One-time setup, then task flow**: Onboarding is one-time, normal task flow after | `onboarding.md:95` | ✅ Framework itself doesn't need onboarding |
## Gap Analysis
### Category Definitions
- **Self-reference gap**: Framework doesn't apply rule to itself
- **Missing rule**: Principle isn't codified where agents can read it
- **Enforcement gap**: Rule exists but nothing checks compliance
- **Lifecycle gap**: Feature exists but follow-up step is missing
- **Design flaw**: Architecture encourages violation of own principles
### Gap Details
| # | Principle Violated | Category | Description | Covered By |
|---|--------------------|----------|-------------|------------|
| G1 | P1 (Task-driven) | Self-reference + Missing rule | No rule says "create task before editing framework files" | `framework-self-enforcement` |
| G2 | P2 (VRAM check) | Self-reference + Missing rule | No rule says "check config.md VRAM before scoping tasks" | `framework-self-enforcement` (new rule) |
| G3 | P3/P4 (Layered filesystem) | Design flaw | `orchestrate.md` reads prompts/contracts/scripts from project first, encouraging full copies instead of minimal footprint | `additive-extension-model` |
| G4 | P5 (No manual task creation) | Missing rule + Enforcement | No rule forbids `mkdir tasks/` — tasks should be created by Orchestrator | `framework-self-enforcement` (new rule) |
| G5 | P6 (Agent reads rules) | Missing rule | `system-prompt.md` doesn't instruct agent to read global `.rules.md` | `framework-self-enforcement` |
| G6 | P8 (Self-improving rules) | Enforcement | Gaps G1-G9 exist but no rules were added to prevent recurrence | `framework-self-enforcement` |
| G7 | P9 (Extension model) | Missing rule | No documentation or rule says "additive overrides only, never copy" | `additive-extension-model` |
| G8 | P10 (Changelog) | Lifecycle | VERDICT.md written per-task but no aggregate changelog | `changelog` |
| G9 | None (missing feature) | Missing feature | No project migration path for existing projects with stale copies | `project-migration` |
| G10 | None (missing feature) | Missing feature | No dashboard-based task review/approval workflow | `dashboard-task-review` |
## Impact/Effort Matrix
```
High Impact
│
│ G3 (design flaw) G1 (self-ref)
│ G2 (VRAM check) G5 (rules)
│ G7 (extension doc)
│
│ G10 (review UI) G4 (manual mkdir)
│ G8 (changelog) G6 (living rules)
│ G9 (migration)
│
└─────────────────────────────→
Low Effort High Effort
```
## Task Structure Validation
### Existing tasks vs. gaps covered
| Task | Gaps Covered |
|------|-------------|
| `additive-extension-model` | G3, G7 |
| `framework-self-enforcement` | G1, G2, G4, G5, G6 |
| `changelog` | G8 |
| `project-migration` | G9 |
### New tasks needed
| Task | Gap | Reason for separate task |
|------|-----|-------------------------|
| `dashboard-task-review` | G10 | UI feature, not a rule change. Separate from `framework-self-enforcement` which is about rules/docs only. |
### Merged into `framework-self-enforcement`
G2, G4, G6 are all rule additions to `.rules.md` — they fit naturally in that single task alongside G1 and G5. No need to split further.
## Recommendations
1. **Keep existing 4 tasks as-is** — each covers its gaps cleanly
2. **Add `dashboard-task-review`** as a new task (G10 — user requested feature)
3. **Expand `framework-self-enforcement` spec** to include VRAM check rule (G2), no-manual-creation rule (G4), and living-rules reminder (G6)
4. **Mark `framework-audit` as complete** once RESEARCH.md is written and tasks are validated
-3
View File
@@ -1,3 +0,0 @@
# Review
- **Status**: approved
- **Timestamp**: 2026-06-13T18:04:41.366656
-80
View File
@@ -1,80 +0,0 @@
# SPEC: Comprehensive Framework Self-Consistency Audit
## Motivation
Several gaps were found where the framework doesn't apply its own principles to itself:
- **No task-driven enforcement**: Framework prescribes task-driven development for projects but has no rules that enforce it for framework changes
- **No resource check before task scoping**: VRAM detection exists but nothing ensures tasks are sized to fit system context limits
- **No changelog/release notes**: VERDICT.md exists per-task but no aggregate change history
- **No migration path**: Framework evolved but existing projects have no cleanup process
These aren't random bugs — they follow a pattern. The framework was designed to automate project development but doesn't apply that automation to **itself**.
## Goal
Systematically audit the entire framework to find every place where it fails to follow its own design principles. Understand the **root cause pattern** so fixes are structural, not piecemeal.
## Method
### Step 1: Extract All Design Principles
Read every file in the framework and extract explicit and implicit design principles:
- `system-prompt.md` — agent instructions
- `.agent.md` — routing rules
- `.rules.md` — project rules
- `prompts/orchestrate.md` — orchestrator behavior
- `prompts/onboarding.md` — project initialization
- `prompts/workflow.md` — workflow state machine
- `config.md` — configuration rules
- `README.md` — documented principles
- `scripts/*.sh` — automation scripts
- `automaton/dashboard/` — dashboard design
- `references/*.md` — reference docs
### Step 2: Self-Consistency Check
For each principle, ask: "Does the framework apply this to itself?"
| Principle | Applied to projects? | Applied to framework? | Gap? |
|---|---|---|---|
| Task-driven development | Yes (onboarding.md) | No | YES |
| VRAM-aware task sizing | Yes (config.md) | No | YES |
| Layered filesystem | Yes (orchestrate.md) | N/A (framework is the base layer) | ? |
| Changelog/release notes | Not documented | No | YES |
| ... (find all) | | | |
### Step 3: Categorize Gaps
For each gap, identify which category it falls into:
1. **Self-reference gap**: Framework doesn't apply its rule to itself
2. **Missing rule**: Principle exists in one place but isn't codified where agents read it
3. **Enforcement gap**: Rule exists but nothing checks compliance
4. **Lifecycle gap**: Feature exists (task completion) but follow-up step is missing (changelog, migration)
### Step 4: Prioritize Fixes
Rank gaps by impact and effort. Flag which should be in the existing 4 tasks vs which need new tasks.
## Acceptance Criteria
- [ ] All design principles extracted and documented
- [ ] All gaps identified with root cause category
- [ ] Gaps prioritized with impact/effort estimate
- [ ] Existing 4 tasks validated or adjusted based on findings
- [ ] New tasks created for any gaps not already covered
## Output
The audit produces `tasks/framework-audit/RESEARCH.md` containing:
1. Complete principle inventory
2. Gap analysis with root cause categories
3. Prioritized action items
4. Recommended task structure
## Context
- 46GB RAM, 16-core AMD CPU, no active GPU driver
- Target context: 16k tokens, 25% headroom, 12k peak per sub-task
- Framework location: `~/.automaton/`
-16
View File
@@ -1,16 +0,0 @@
# VERDICT: Framework Self-Consistency Audit
## Status: PASS
## Summary
Performed comprehensive audit of the framework against its own design principles. Produced RESEARCH.md with 11 principles, 10 gaps, and 5 tasks.
## Phase Results
| Phase | Result |
|-------|--------|
| Research | ✅ PASS — RESEARCH.md produced |
| Bug Find | ✅ PASS |
| Doc Review | ✅ PASS |
## Final Verdict
**PASS** — Audit is comprehensive. All identified gaps were covered by existing or newly created tasks.