From 46ceed0122285348935a5dfa2385828268c1c3f1 Mon Sep 17 00:00:00 2001 From: laptran Date: Sat, 13 Jun 2026 12:16:01 -0400 Subject: [PATCH] Add Blocked dashboard column, framework audit, and 5 new task specs --- automaton/dashboard/core/board.py | 5 +- automaton/dashboard/html/dashboard.js | 3 +- tasks/additive-extension-model/SPEC.md | 69 ++++++++++++++++++ tasks/changelog/SPEC.md | 33 +++++++++ tasks/dashboard-task-review/SPEC.md | 77 ++++++++++++++++++++ tasks/framework-audit/RESEARCH.md | 89 ++++++++++++++++++++++++ tasks/framework-audit/SPEC.md | 80 +++++++++++++++++++++ tasks/framework-self-enforcement/SPEC.md | 69 ++++++++++++++++++ tasks/project-migration/SPEC.md | 34 +++++++++ 9 files changed, 457 insertions(+), 2 deletions(-) create mode 100644 tasks/additive-extension-model/SPEC.md create mode 100644 tasks/changelog/SPEC.md create mode 100644 tasks/dashboard-task-review/SPEC.md create mode 100644 tasks/framework-audit/RESEARCH.md create mode 100644 tasks/framework-audit/SPEC.md create mode 100644 tasks/framework-self-enforcement/SPEC.md create mode 100644 tasks/project-migration/SPEC.md diff --git a/automaton/dashboard/core/board.py b/automaton/dashboard/core/board.py index ef57fe2..fb1f1f6 100644 --- a/automaton/dashboard/core/board.py +++ b/automaton/dashboard/core/board.py @@ -20,8 +20,11 @@ class KanbanBoard: "Verification": [ TaskState.BUG_FIND, TaskState.ADV_BUG_FIND, TaskState.DOC_REVIEW, TaskState.REFEREE ], + "Blocked": [ + TaskState.BLOCKED + ], "Resolution": [ - TaskState.DONE, TaskState.BLOCKED + TaskState.DONE ], } diff --git a/automaton/dashboard/html/dashboard.js b/automaton/dashboard/html/dashboard.js index 8e242de..22ac2ba 100644 --- a/automaton/dashboard/html/dashboard.js +++ b/automaton/dashboard/html/dashboard.js @@ -29,7 +29,8 @@ const PHASE_GROUPS = [ { id: 'design', label: 'Design', color: '#26c6da', states: ['design', 'test_design'] }, { id: 'implementation', label: 'Implementation', color: '#66bb6a', states: ['implement'] }, { id: 'verification', label: 'Verification', color: '#ffa726', states: ['bug_find', 'adv_bug_find', 'doc_review', 'referee'] }, - { id: 'resolution', label: 'Resolution', color: '#66bb6a', states: ['done', 'blocked'] }, + { id: 'blocked', label: 'Blocked', color: '#ef5350', states: ['blocked'] }, + { id: 'resolution', label: 'Resolution', color: '#66bb6a', states: ['done'] }, ]; // State icon map diff --git a/tasks/additive-extension-model/SPEC.md b/tasks/additive-extension-model/SPEC.md new file mode 100644 index 0000000..7bce01a --- /dev/null +++ b/tasks/additive-extension-model/SPEC.md @@ -0,0 +1,69 @@ +# SPEC: Additive Extension Model for Project Upgrades + +## Overview + +Replace the current diff/merge upgrade process with a simpler additive extension model. Projects should never copy framework files. Instead, they provide overrides via `.agent.md`, `.rules.md`, and an optional `extensions/` directory. Updating the framework becomes a simple `git pull` with no project-level file comparison. + +## Motivation + +The current design has a design-vs-reality gap: + +| Design Intent | Reality | +|---|---| +| Projects only have `.agent.md` + `.rules.md` | Projects have full copies of framework files | +| Additive overrides only | Diff/merge required on upgrade | +| Simple `git pull` update | Complex file-by-file comparison | + +The root cause: `orchestrate.md` reads prompts/contracts/scripts from the **project first**, then falls back to global. This encourages copying files into the project, which breaks the clean separation. + +## Required Changes + +### 1. `prompts/orchestrate.md` + +Remove the project-first fallback for prompts, contracts, and scripts. The Orchestrator should: +- Always read base prompts/contracts/scripts from `~/.automaton/` (global) +- Check `{project}/.automaton/extensions/` for additive extensions (not replacements) +- Specific extension files to check: + - `{project}/.automaton/extensions/prompts/*.md` - loaded after the corresponding global prompt + - `{project}/.automaton/extensions/contracts/*.md` - loaded after global contracts + - `{project}/.automaton/extensions/scripts/*.sh` - loaded before global scripts (to allow pre-processing) +- The read order for `.agent.md` stays layered (project override is correct for routing) + +### 2. `prompts/onboarding.md` + +- Remove the "Project Upgrade" section (lines 99-155) that performs diff/merge +- Simplify to: if `.agent.md` or `.rules.md` are missing, create minimal defaults +- Add documentation for the `extensions/` directory pattern +- Remove any instructions that copy framework files into the project + +### 3. `scripts/update.sh` + +- Simplify: remove the reset hard HEAD step. Just `git pull` with a clean working tree check. +- Ensure it only touches `~/.automaton/`, never project directories + +### 4. `README.md` + +- Update the "Upgrading existing projects" section to describe the new additive model +- Document the `extensions/` directory pattern + +## Acceptance Criteria + +- [ ] `prompts/orchestrate.md` reads prompts/contracts/scripts from `~/.automaton/` first, not from project +- [ ] `prompts/orchestrate.md` checks `{project}/.automaton/extensions/` for additive extensions +- [ ] `prompts/onboarding.md` no longer has diff/merge upgrade logic +- [ ] `prompts/onboarding.md` creates minimal `.agent.md` and `.rules.md` if missing +- [ ] `prompts/onboarding.md` documents the `extensions/` directory +- [ ] `scripts/update.sh` does a simple `git pull` without resetting local changes +- [ ] `README.md` describes the new upgrade model +- [ ] No existing prompt/contract/script behavior is broken (regression check) + +## Non-Goals + +- Moving the dashboard (`automaton/dashboard/`) - it already reads from `~/.automaton/` at runtime +- Changing how `config.md` is read (already always global per `orchestrate.md` line 15) +- Changing the task directory structure + +## Notes + +- The `invest-copilot` project has a full copy of the framework in its `.automaton/` - this task should include a migration path to clean it up +- The extension model should be documented in `references/extensions.md` as well diff --git a/tasks/changelog/SPEC.md b/tasks/changelog/SPEC.md new file mode 100644 index 0000000..d5d8e96 --- /dev/null +++ b/tasks/changelog/SPEC.md @@ -0,0 +1,33 @@ +# SPEC: Changelog and Release Notes Process + +## Overview + +There's no changelog or release note process. Completed tasks produce a VERDICT.md but there's no aggregate record of what changed. This makes it hard to generate release notes or see the project's history at a glance. + +## Changes + +### 1. `CHANGELOG.md` at `~/.automaton/` + +Root-level changelog file with entries appended on task completion: + +```markdown +# Changelog + +## [unreleased] + +### Added +- description (#task-name) + +### Fixed +- description (#task-name) +``` + +### 2. Update `.rules.md` + +Add a rule: "When a task reaches Resolution (VERDICT.md written), append an entry to CHANGELOG.md." + +## Acceptance Criteria + +- [ ] `~/.automaton/CHANGELOG.md` exists with [unreleased] header +- [ ] `.rules.md` mentions changelog updates as part of task completion +- [ ] Format is clean enough to use as release notes source diff --git a/tasks/dashboard-task-review/SPEC.md b/tasks/dashboard-task-review/SPEC.md new file mode 100644 index 0000000..94e2de0 --- /dev/null +++ b/tasks/dashboard-task-review/SPEC.md @@ -0,0 +1,77 @@ +# SPEC: Dashboard Task Review and Approval + +## Overview + +The dashboard currently shows tasks grouped by phase but has no workflow for reviewing and approving tasks before they proceed to the next phase. A user should be able to review a task's artifacts (SPEC.md, DESIGN.md, IMPLEMENTATION.md, etc.) and approve or reject it directly from the dashboard. + +## Motivation + +The framework has phases that require user sign-off (Research → SPEC.md review, Design → DESIGN.md review, etc.) but this sign-off happens via agent interaction, not through the dashboard. Adding review/approval to the dashboard provides: + +1. **Asynchronous review** — approve or flag tasks without an active agent session +2. **Audit trail** — who approved what and when +3. **Blocked task management** — reject a task to move it to Blocked column +4. **Self-service** — approve multiple tasks at a glance + +## Requirements + +### 1. Review State Per Task + +Each task can have a review status: + +| Status | Meaning | +|--------|---------| +| `pending` | Awaiting review (default for new artifacts) | +| `approved` | Reviewer approved, task can proceed | +| `changes_requested` | Reviewer wants changes, task moves to Blocked | +| `not_needed` | No review needed (e.g., automated phases) | + +### 2. Review Data Storage + +Store review state in the task directory: +- `REVIEW.md` — review metadata (status, reviewer, timestamp, comments) +- Or embed in existing artifact files if simpler + +### 3. Dashboard UI + +- Each task card shows review status badge (🟡 pending, ✅ approved, ❌ changes requested) +- Clicking a task opens a detail panel with: + - Artifact preview (SPEC.md, DESIGN.md, etc.) + - Approve / Request Changes buttons + - Comment box +- Phase columns show a review filter toggle (show all / show pending only) +- Stats view includes review metrics (pending approvals count) + +### 4. Integration with Phase Flow + +- A task in "Research" phase with a new SPEC.md starts as `pending` review +- When approved, the task proceeds to next phase +- When "changes requested", the task moves to Blocked column with a note +- Review status is checked by the Orchestrator before auto-executing next phase + +### 5. API Endpoints + +- `GET /api/tasks/{name}/review` — get review status +- `POST /api/tasks/{name}/review` — submit review (approve/changes_requested + comment) +- `GET /api/review-summary` — aggregate review metrics for all tasks + +## Acceptance Criteria + +- [ ] Tasks have review state stored in REVIEW.md +- [ ] Dashboard shows review status badge on task cards +- [ ] Detail panel has Approve / Request Changes buttons +- [ ] Filter for pending reviews only +- [ ] Stats shows pending approval count +- [ ] API serves review data and accepts review submissions +- [ ] Orchestrator checks review status before auto-executing + +## Out of Scope + +- Multi-user review workflow (single user for now) +- Email/notification system for pending reviews +- Role-based permissions + +## Notes + +- This builds on the existing Blocked column (when changes_requested, task goes to Blocked) +- Review state is simple — approve or changes_requested, no multi-level approval diff --git a/tasks/framework-audit/RESEARCH.md b/tasks/framework-audit/RESEARCH.md new file mode 100644 index 0000000..e06ae3b --- /dev/null +++ b/tasks/framework-audit/RESEARCH.md @@ -0,0 +1,89 @@ +# Framework Self-Consistency Audit + +## Principle Inventory + +Extracted from all framework files. Each principle is a rule the framework prescribes for project work. + +| # | Principle | Source | Applied to Framework? | +|---|-----------|--------|-----------------------| +| P1 | **Task-driven development**: All changes go through tasks (SPEC → phases → VERDICT) | `onboarding.md`, `workflow.md`, `.rules.md` | ❌ No rule enforces this for framework itself | +| P2 | **VRAM-aware task sizing**: Tasks must fit system context limits; check before scoping | `config.md`, `orchestrate.md:21-101` | ❌ Never checked when creating framework tasks | +| P3 | **Layered filesystem**: Project overrides global, read project first then fallback to global | `orchestrate.md:5-19`, `README.md:177-202` | ⚠️ Broken design — project-first read encourages full copies | +| P4 | **Minimal project footprint**: Projects should only have `.agent.md` + `.rules.md` | `onboarding.md:42`, `README.md:188` | ⚠️ Violated by P3's project-first read order | +| P5 | **No manual task creation**: Orchestrator creates task folders, never the user | `workflow.md:22` | ❌ No rule forbids manual `mkdir tasks/` | +| P6 | **Agent reads rules at startup**: Must read `.agent.md` + `.rules.md` before working | `system-prompt.md`, `session-starter.md` | ⚠️ Doesn't read global `.rules.md`, only project's | +| P7 | **Stop condition enforcement**: "CONTRACT_MET" prevents early termination | `references/stop-hook-pattern.md`, various prompts | ✅ Phase-level prompts have stop conditions | +| P8 | **Self-improving rules**: `.rules.md` is a living document, add rules per failure mode | `.rules.md:3-5` | ❌ No rules were added for any of these gaps | +| P9 | **Customization via extension, not copy**: Override additively, never duplicate | Implicit from design intent | ❌ Not codified anywhere | +| P10 | **Changelog/release notes**: Changes should be recorded for release notes | Not stated anywhere | ❌ No changelog exists | +| P11 | **One-time setup, then task flow**: Onboarding is one-time, normal task flow after | `onboarding.md:95` | ✅ Framework itself doesn't need onboarding | + +## Gap Analysis + +### Category Definitions +- **Self-reference gap**: Framework doesn't apply rule to itself +- **Missing rule**: Principle isn't codified where agents can read it +- **Enforcement gap**: Rule exists but nothing checks compliance +- **Lifecycle gap**: Feature exists but follow-up step is missing +- **Design flaw**: Architecture encourages violation of own principles + +### Gap Details + +| # | Principle Violated | Category | Description | Covered By | +|---|--------------------|----------|-------------|------------| +| G1 | P1 (Task-driven) | Self-reference + Missing rule | No rule says "create task before editing framework files" | `framework-self-enforcement` | +| G2 | P2 (VRAM check) | Self-reference + Missing rule | No rule says "check config.md VRAM before scoping tasks" | `framework-self-enforcement` (new rule) | +| G3 | P3/P4 (Layered filesystem) | Design flaw | `orchestrate.md` reads prompts/contracts/scripts from project first, encouraging full copies instead of minimal footprint | `additive-extension-model` | +| G4 | P5 (No manual task creation) | Missing rule + Enforcement | No rule forbids `mkdir tasks/` — tasks should be created by Orchestrator | `framework-self-enforcement` (new rule) | +| G5 | P6 (Agent reads rules) | Missing rule | `system-prompt.md` doesn't instruct agent to read global `.rules.md` | `framework-self-enforcement` | +| G6 | P8 (Self-improving rules) | Enforcement | Gaps G1-G9 exist but no rules were added to prevent recurrence | `framework-self-enforcement` | +| G7 | P9 (Extension model) | Missing rule | No documentation or rule says "additive overrides only, never copy" | `additive-extension-model` | +| G8 | P10 (Changelog) | Lifecycle | VERDICT.md written per-task but no aggregate changelog | `changelog` | +| G9 | None (missing feature) | Missing feature | No project migration path for existing projects with stale copies | `project-migration` | +| G10 | None (missing feature) | Missing feature | No dashboard-based task review/approval workflow | `dashboard-task-review` | + +## Impact/Effort Matrix + +``` +High Impact + │ + │ G3 (design flaw) G1 (self-ref) + │ G2 (VRAM check) G5 (rules) + │ G7 (extension doc) + │ + │ G10 (review UI) G4 (manual mkdir) + │ G8 (changelog) G6 (living rules) + │ G9 (migration) + │ + └─────────────────────────────→ + Low Effort High Effort + +``` + +## Task Structure Validation + +### Existing tasks vs. gaps covered + +| Task | Gaps Covered | +|------|-------------| +| `additive-extension-model` | G3, G7 | +| `framework-self-enforcement` | G1, G2, G4, G5, G6 | +| `changelog` | G8 | +| `project-migration` | G9 | + +### New tasks needed + +| Task | Gap | Reason for separate task | +|------|-----|-------------------------| +| `dashboard-task-review` | G10 | UI feature, not a rule change. Separate from `framework-self-enforcement` which is about rules/docs only. | + +### Merged into `framework-self-enforcement` + +G2, G4, G6 are all rule additions to `.rules.md` — they fit naturally in that single task alongside G1 and G5. No need to split further. + +## Recommendations + +1. **Keep existing 4 tasks as-is** — each covers its gaps cleanly +2. **Add `dashboard-task-review`** as a new task (G10 — user requested feature) +3. **Expand `framework-self-enforcement` spec** to include VRAM check rule (G2), no-manual-creation rule (G4), and living-rules reminder (G6) +4. **Mark `framework-audit` as complete** once RESEARCH.md is written and tasks are validated diff --git a/tasks/framework-audit/SPEC.md b/tasks/framework-audit/SPEC.md new file mode 100644 index 0000000..b8359da --- /dev/null +++ b/tasks/framework-audit/SPEC.md @@ -0,0 +1,80 @@ +# SPEC: Comprehensive Framework Self-Consistency Audit + +## Motivation + +Several gaps were found where the framework doesn't apply its own principles to itself: + +- **No task-driven enforcement**: Framework prescribes task-driven development for projects but has no rules that enforce it for framework changes +- **No resource check before task scoping**: VRAM detection exists but nothing ensures tasks are sized to fit system context limits +- **No changelog/release notes**: VERDICT.md exists per-task but no aggregate change history +- **No migration path**: Framework evolved but existing projects have no cleanup process + +These aren't random bugs — they follow a pattern. The framework was designed to automate project development but doesn't apply that automation to **itself**. + +## Goal + +Systematically audit the entire framework to find every place where it fails to follow its own design principles. Understand the **root cause pattern** so fixes are structural, not piecemeal. + +## Method + +### Step 1: Extract All Design Principles + +Read every file in the framework and extract explicit and implicit design principles: +- `system-prompt.md` — agent instructions +- `.agent.md` — routing rules +- `.rules.md` — project rules +- `prompts/orchestrate.md` — orchestrator behavior +- `prompts/onboarding.md` — project initialization +- `prompts/workflow.md` — workflow state machine +- `config.md` — configuration rules +- `README.md` — documented principles +- `scripts/*.sh` — automation scripts +- `automaton/dashboard/` — dashboard design +- `references/*.md` — reference docs + +### Step 2: Self-Consistency Check + +For each principle, ask: "Does the framework apply this to itself?" + +| Principle | Applied to projects? | Applied to framework? | Gap? | +|---|---|---|---| +| Task-driven development | Yes (onboarding.md) | No | YES | +| VRAM-aware task sizing | Yes (config.md) | No | YES | +| Layered filesystem | Yes (orchestrate.md) | N/A (framework is the base layer) | ? | +| Changelog/release notes | Not documented | No | YES | +| ... (find all) | | | | + +### Step 3: Categorize Gaps + +For each gap, identify which category it falls into: + +1. **Self-reference gap**: Framework doesn't apply its rule to itself +2. **Missing rule**: Principle exists in one place but isn't codified where agents read it +3. **Enforcement gap**: Rule exists but nothing checks compliance +4. **Lifecycle gap**: Feature exists (task completion) but follow-up step is missing (changelog, migration) + +### Step 4: Prioritize Fixes + +Rank gaps by impact and effort. Flag which should be in the existing 4 tasks vs which need new tasks. + +## Acceptance Criteria + +- [ ] All design principles extracted and documented +- [ ] All gaps identified with root cause category +- [ ] Gaps prioritized with impact/effort estimate +- [ ] Existing 4 tasks validated or adjusted based on findings +- [ ] New tasks created for any gaps not already covered + +## Output + +The audit produces `tasks/framework-audit/RESEARCH.md` containing: +1. Complete principle inventory +2. Gap analysis with root cause categories +3. Prioritized action items +4. Recommended task structure + +## Context + +- 46GB RAM, 16-core AMD CPU, no active GPU driver +- Target context: 16k tokens, 25% headroom, 12k peak per sub-task +- Framework location: `~/.automaton/` diff --git a/tasks/framework-self-enforcement/SPEC.md b/tasks/framework-self-enforcement/SPEC.md new file mode 100644 index 0000000..31a098e --- /dev/null +++ b/tasks/framework-self-enforcement/SPEC.md @@ -0,0 +1,69 @@ +# SPEC: Enforce Task-Driven Development via Framework Rules + +## Overview + +The automaton framework has a task-driven workflow (SPEC → Design → Implementation → Verification → Resolution) but nothing tells the AI agent to follow it. The agent makes ad-hoc edits instead of creating proper tasks first. + +## Root Cause + +The agent reads `system-prompt.md`, `.agent.md`, and `.rules.md` at startup, but none instruct it to use the task-driven process. The agent treats framework modifications as "normal coding." + +## Framework Audit Context + +This task covers 4 gaps identified in `framework-audit/RESEARCH.md`: + +| Gap | Principle | Rule to Add | +|-----|-----------|-------------| +| G1 | Task-driven development | Create task before editing files | +| G2 | VRAM-aware task sizing | Check config.md VRAM before scoping tasks | +| G4 | No manual task creation | Never mkdir tasks/ — use Orchestrator | +| G5 | Agent reads global rules | system-prompt.md must read global .rules.md | +| G6 | Self-improving rules | Add rules per failure mode observed | + +## Changes + +### 1. `~/.automaton/.rules.md` + +Replace the template with concrete rules: + +```markdown +# .rules.md + +## Task-Driven Development +- All changes must go through a task in tasks/{name}/ with SPEC.md → phases → VERDICT.md +- Never edit files directly without a corresponding task +- Never create task directories manually (mkdir tasks/) — use the Orchestrator + +## VRAM-Aware Task Sizing +- Before creating or scoping a task, check ~/.automaton/config.md for VRAM limits +- Verify the task fits within max peak context (default: 12k tokens) +- If no config exists, default to 8k with 25% headroom + +## Self-Improvement +- Add one rule per observed failure mode with a concrete example +- Consolidate contradictions monthly. Remove stale rules. +- No rule without a real example of the problem it prevents +``` + +### 2. `system-prompt.md` + +After existing instructions, add: + +``` +4. Read ~/.automaton/.rules.md (global framework rules) +``` + +## Acceptance Criteria + +- [ ] `~/.automaton/.rules.md` has Task-Driven Development, VRAM-Aware Sizing, and Self-Improvement sections +- [ ] `system-prompt.md` instructs agent to read global `.rules.md` +- [ ] Agent creates tasks before making changes +- [ ] Agent checks VRAM limits before scoping tasks +- [ ] Agent never creates task directories manually + +## Out of Scope + +- Documentation/changelog process (separate task) +- Project migration/cleanup (separate task) +- Onboarding changes (separate task) +- Dashboard task review UI (separate task) diff --git a/tasks/project-migration/SPEC.md b/tasks/project-migration/SPEC.md new file mode 100644 index 0000000..812a72d --- /dev/null +++ b/tasks/project-migration/SPEC.md @@ -0,0 +1,34 @@ +# SPEC: Project Migration Script for Additive Extension Model + +## Overview + +After switching to the additive extension model, existing projects have stale copies of framework files in their `{project}/.automaton/` directories. These need to be cleaned up. + +## Context + +Under the old model, framework files (prompts, contracts, scripts) were copied into the project's `.automaton/`. Under the new model, the project should only contain `.agent.md`, `.rules.md`, and optionally an `extensions/` directory. Everything else comes from `~/.automaton/`. + +## Changes + +### 1. `scripts/migrate-project.sh` + +Script that takes a project path and: +1. Scans `{project}/.automaton/` for files that now live in `~/.automaton/` +2. For each matching file: + - **Identical to global** → delete (framework provides them) + - **Different from global** → move to `.automaton/extensions/` +3. Preserves `.agent.md` and `.rules.md` as-is +4. Reports what was deleted, moved, and kept + +### 2. `prompts/onboarding.md` + +Add a section at the top that checks if migration is needed: +- If project has prompt/contract/script copies in `.automaton/`, offer to run migration + +## Acceptance Criteria + +- [ ] `scripts/migrate-project.sh` exists and is executable +- [ ] Script correctly identifies identical vs customized files +- [ ] Customized files are moved to `extensions/`, not deleted +- [ ] `.agent.md` and `.rules.md` are never touched +- [ ] Onboarding detects stale projects and offers migration