- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top level (all <7 days old per the cleanup policy; premature bulk archive was fixed). - **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds the board via innerHTML every 2s, destroying each column-body's scrollTop. Now snapshots column-body scrollTop + board.scrollLeft + view.scrollTop before rebuild and restores after (matched by PHASE_GROUPS index). - **Dashboard UI additions** (pre-existing unstaged work): approval section cards, transition buttons, inline artifact editor (textarea for writing missing SPEC/VERDICT/etc from the detail modal). - **Bind ornith as Implement model** — config.md: Model explicit to omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive autopilot already used ornith via opencode default; now explicit. - **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg pointing at a pytest temp dir (test isolation leak). Rewired to point at ~/.automaton. - **Fix plist-isolation test** — test asserted host plist doesn't exist, but a real install creates it. Now snapshots mtime before run, asserts unchanged after (only a write during the test counts as bleed). - **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests: board renders tasks, column scroll survives auto-refresh tick. Verified the test fails without the scroll fix (scrollTop resets to 0). Skipped via importorskip when playwright is absent (main CI stays green). - **Clarify SI loop scope in README** — new-project onboarding section documents the framework-scoped self-improvement loop and options (leave/pause/create project loop). - **CHANGELOG** documents all changes including the known model-divergence gap (mde tasks marked complete but per-role model binding was never implemented).
3.1 KiB
Framework Configuration
This file contains global framework settings that apply across all projects.
VRAM Configuration
Settings for task decomposition based on available VRAM.
- Auto-detect: Yes # Detect GPU VRAM, RAM, and model context window automatically
- Target context: 16k tokens # Override auto-detect if needed
- Headroom: 25% # Leave headroom for code, context, and reasoning
- Max peak context per sub-task: 12k tokens # Max context for any single sub-task
Auto-detection
When Auto-detect: Yes, the framework probes your system to detect:
- GPU VRAM (via
nvidia-smiorlspci) - System RAM (via
/proc/meminfoorsysctl) - Model context window (via API config or model name lookup)
- Framework overhead (by reading all loaded prompt files)
To disable auto-detection and use manual values:
## VRAM Configuration
- **Auto-detect**: No
- **Target context**: 8k
- **Headroom**: 30%
- **Max peak context per sub-task**: 5.6k
Model Configuration
Settings for the LLM model being used.
- Model: omlx/Ornith-1.0-35B-4bit-mlx # Local LLM (opencode provider); used as the Implement role
- Override context window: 32768 # Matches opencode.json limit.context for ornith
Auto-detection
When Model: auto, the framework detects the model name from:
.agent.mdin the project (if specified there)- API config files (
.env,config.yaml,config.json, etc.) - Model name lookup by name (e.g., gpt-4o → 128k, claude-3-5-sonnet → 200k)
To disable auto-detection and use manual values:
## Model Configuration
- **Model**: gpt-4o
- **Override context window**: 128k
System Requirements
Requirements for the environment the framework runs in.
- nvidia-smi: Required if NVIDIA GPU (for VRAM detection)
- lspci: Fallback if NVIDIA GPU not available (for AMD GPU VRAM detection)
- /proc/meminfo: Required for RAM detection (Linux)
- sysctl: Fallback for RAM detection (macOS)
Loop Role Models
Loop ticks run three session roles. Roles are sessions, not models — a single model can fill multiple roles. Configure each loop's role-to-prompt binding in its loop.json; this section documents the framework's expectations only.
- Implement: — produces the artifact for this tick. Bound to
prompts/loop-implement.mdby default. - Verify: — grades the artifact and emits the JSON verdict
{pass, score, reasons, next_hint}. Bound toprompts/loop-verifier.md. The framework never inspects this role's model (D8); only its session. - Orchestrate: — applies the verdict, calls exactly one
status.pyoperation per tick, enforces brakes. Bound toprompts/loop-orchestrate.md.
Conflict-of-interest rule (D12): Verify: and Implement: must never be the same session. When two distinct sessions are infeasible (single-session harness), the runner falls back to session-only divergence — still safe.
Role context tiers are set per-loop in loop.json, not globally. The 16k floor (D13) applies regardless of tier.
Framework Version
- Version: 2.0
- State enforcement: enabled (
.statefile +status.py)