81 lines
3.2 KiB
Markdown
81 lines
3.2 KiB
Markdown
# SPEC: Comprehensive Framework Self-Consistency Audit
|
|||
|
|
|
||
|
|
## Motivation
|
||
|
|
|
||
|
|
Several gaps were found where the framework doesn't apply its own principles to itself:
|
||
|
|
|
||
|
|
- **No task-driven enforcement**: Framework prescribes task-driven development for projects but has no rules that enforce it for framework changes
|
||
|
|
- **No resource check before task scoping**: VRAM detection exists but nothing ensures tasks are sized to fit system context limits
|
||
|
|
- **No changelog/release notes**: VERDICT.md exists per-task but no aggregate change history
|
||
|
|
- **No migration path**: Framework evolved but existing projects have no cleanup process
|
||
|
|
|
||
|
|
These aren't random bugs — they follow a pattern. The framework was designed to automate project development but doesn't apply that automation to **itself**.
|
||
|
|
|
||
|
|
## Goal
|
||
|
|
|
||
|
|
Systematically audit the entire framework to find every place where it fails to follow its own design principles. Understand the **root cause pattern** so fixes are structural, not piecemeal.
|
||
|
|
|
||
|
|
## Method
|
||
|
|
|
||
|
|
### Step 1: Extract All Design Principles
|
||
|
|
|
||
|
|
Read every file in the framework and extract explicit and implicit design principles:
|
||
|
|
- `system-prompt.md` — agent instructions
|
||
|
|
- `.agent.md` — routing rules
|
||
|
|
- `.rules.md` — project rules
|
||
|
|
- `prompts/orchestrate.md` — orchestrator behavior
|
||
|
|
- `prompts/onboarding.md` — project initialization
|
||
|
|
- `prompts/workflow.md` — workflow state machine
|
||
|
|
- `config.md` — configuration rules
|
||
|
|
- `README.md` — documented principles
|
||
|
|
- `scripts/*.sh` — automation scripts
|
||
|
|
- `automaton/dashboard/` — dashboard design
|
||
|
|
- `references/*.md` — reference docs
|
||
|
|
|
||
|
|
### Step 2: Self-Consistency Check
|
||
|
|
|
||
|
|
For each principle, ask: "Does the framework apply this to itself?"
|
||
|
|
|
||
|
|
| Principle | Applied to projects? | Applied to framework? | Gap? |
|
||
|
|
|---|---|---|---|
|
||
|
|
| Task-driven development | Yes (onboarding.md) | No | YES |
|
||
|
|
| VRAM-aware task sizing | Yes (config.md) | No | YES |
|
||
|
|
| Layered filesystem | Yes (orchestrate.md) | N/A (framework is the base layer) | ? |
|
||
|
|
| Changelog/release notes | Not documented | No | YES |
|
||
|
|
| ... (find all) | | | |
|
||
|
|
|
||
|
|
### Step 3: Categorize Gaps
|
||
|
|
|
||
|
|
For each gap, identify which category it falls into:
|
||
|
|
|
||
|
|
1. **Self-reference gap**: Framework doesn't apply its rule to itself
|
||
|
|
2. **Missing rule**: Principle exists in one place but isn't codified where agents read it
|
||
|
|
3. **Enforcement gap**: Rule exists but nothing checks compliance
|
||
|
|
4. **Lifecycle gap**: Feature exists (task completion) but follow-up step is missing (changelog, migration)
|
||
|
|
|
||
|
|
### Step 4: Prioritize Fixes
|
||
|
|
|
||
|
|
Rank gaps by impact and effort. Flag which should be in the existing 4 tasks vs which need new tasks.
|
||
|
|
|
||
|
|
## Acceptance Criteria
|
||
|
|
|
||
|
|
- [ ] All design principles extracted and documented
|
||
|
|
- [ ] All gaps identified with root cause category
|
||
|
|
- [ ] Gaps prioritized with impact/effort estimate
|
||
|
|
- [ ] Existing 4 tasks validated or adjusted based on findings
|
||
|
|
- [ ] New tasks created for any gaps not already covered
|
||
|
|
|
||
|
|
## Output
|
||
|
|
|
||
|
|
The audit produces `tasks/framework-audit/RESEARCH.md` containing:
|
||
|
|
1. Complete principle inventory
|
||
|
|
2. Gap analysis with root cause categories
|
||
|
|
3. Prioritized action items
|
||
|
|
4. Recommended task structure
|
||
|
|
|
||
|
|
## Context
|
||
|
|
|
||
|
|
- 46GB RAM, 16-core AMD CPU, no active GPU driver
|
||
|
|
- Target context: 16k tokens, 25% headroom, 12k peak per sub-task
|
||
|
|
- Framework location: `~/.automaton/`
|