Files
Lap Tran bc7daf8590 Restore archived tasks, fix dashboard scroll-reset, bind ornith, add Playwright smoke test
- **Restore 82 completed tasks** from tasks/complete/ back to tasks/ top
  level (all <7 days old per the cleanup policy; premature bulk archive
  was fixed).
- **Dashboard: fix scroll-reset on auto-refresh** — renderBoard rebuilds
  the board via innerHTML every 2s, destroying each column-body's
  scrollTop. Now snapshots column-body scrollTop + board.scrollLeft +
  view.scrollTop before rebuild and restores after (matched by
  PHASE_GROUPS index).
- **Dashboard UI additions** (pre-existing unstaged work): approval
  section cards, transition buttons, inline artifact editor (textarea for
  writing missing SPEC/VERDICT/etc from the detail modal).
- **Bind ornith as Implement model** — config.md: Model explicit to
  omlx/Ornith-1.0-35B-4bit-mlx, context window 32768. Interactive
  autopilot already used ornith via opencode default; now explicit.
- **Fix cleanup stub** — automaton-cleanup.sh had a stale --project arg
  pointing at a pytest temp dir (test isolation leak). Rewired to point
  at ~/.automaton.
- **Fix plist-isolation test** — test asserted host plist doesn't exist,
  but a real install creates it. Now snapshots mtime before run, asserts
  unchanged after (only a write during the test counts as bleed).
- **New Playwright smoke test** (tests/test_dashboard_ui.py) — 2 tests:
  board renders tasks, column scroll survives auto-refresh tick.
  Verified the test fails without the scroll fix (scrollTop resets to 0).
  Skipped via importorskip when playwright is absent (main CI stays
  green).
- **Clarify SI loop scope in README** — new-project onboarding section
  documents the framework-scoped self-improvement loop and options
  (leave/pause/create project loop).
- **CHANGELOG** documents all changes including the known model-divergence
  gap (mde tasks marked complete but per-role model binding was never
  implemented).
2026-06-26 10:05:18 -04:00

11 KiB

Fix Context Sizing

Tier 1 context-sizing fixes — the foundational layer that add-status-brakes and downstream loop tasks consume. Per design/loops/technical.md §11 and the locked Tier 1 list.

Goal

Make the framework's context-budget reporting honest, single-headroom-applied, and machine-readable with a hard floor. Today vram_detect.py lies: it reports fake 8k/6k defaults when the actual budget is zero, silently applies headroom twice, and never refuses to run on an unknown model. The loop runner (task add-loop-runner) needs accurate, authoritative numbers — its 16k floor check (D13) is meaningless against fabricated defaults.

Requirements

R1. Remove double-headroom application in recommend_context

scripts/vram_detect.py:618-656 applies headroom twice:

  • Line 642 / 644 / 648: applies (100 - headroom_pct) // 100 while constructing recommended_kb from vram_context_kb / model_context_kb / ram_context_kb.
  • Line 654: applies (100 - headroom_pct) // 100 again when deriving max_peak_kb = net_kb * (100 - headroom_pct) // 100.

Net effect: max_peak_kb is discounted by headroom_pct twice, so a 25% headroom becomes a 44% reduction.

Fix: Restructure recommend_context so headroom is applied exactly once. Build recommended_kb as the raw budget (no * (100 - headroom_pct) // 100 at lines 642, 644, 648), subtract overhead, then apply headroom once to derive max_peak_kb:

def recommend_context(...) -> tuple[int, int, int]:
    headroom_pct = int(config.get("headroom_pct", DEFAULT_HEADROOM_PCT))

    if not config.get("auto_detect", True):
        target_kb = int(config.get("target_context_kb", 0))
        max_peak_kb = int(config.get("max_peak_kb", 0))
        if target_kb > 0:
            if max_peak_kb == 0 and headroom_pct > 0:
                max_peak_kb = target_kb * (100 - headroom_pct) // 100
            return headroom_pct, target_kb, max_peak_kb

    recommended_kb = 0
    if gpu_vram_gb >= 4:
        recommended_kb = gpu_vram_gb * 2000
    elif model_context_kb > 0:
        recommended_kb = model_context_kb
    else:
        recommended_kb = ram_gb * 750

    net_kb = recommended_kb - overhead_tokens
    max_peak_kb = net_kb * (100 - headroom_pct) // 100
    return headroom_pct, net_kb, max_peak_kb

Manual-override branch unchanged (it already applies headroom once via max_peak_kb = target_kb * (100 - headroom_pct) // 100).

R2. Stop lying about zero/negative budgets

scripts/vram_detect.py:651:

net_kb = max(0, recommended_kb - overhead_tokens)

and :698-699:

recommended_k = recommended_kb // 1000 if recommended_kb > 0 else 8
max_peak_k = max_peak_kb // 1000 if max_peak_kb > 0 else 6

The max(0, ...) silently clamps an actually-negative budget to zero, and the else 8 / else 6 report fabricated 8k/6k numbers when the true budget is zero or unknown. Any downstream consumer — the dashboard, decompose.md, the future loop runner — reads 8k and proceeds as if it's safe.

Fix:

  1. Drop the max(0, ...) clamp. Keep net_kb as the true arithmetic value (may be negative or zero). Already-floored callers (e.g. the dashboard) can compute max(0, ...) themselves; vram_detect.py returns the honest number.
  2. Drop the else 8 / else 6 fallbacks. Report the real quotient even when zero.
  3. Add a warning line to stdout when net_kb <= 0 or model_context_kb == 0 (see R3 for the loop refuse). For non-loop CLI invocations this is a human-readable warning, not an error exit.
recommended_k = recommended_kb // 1000
max_peak_k = max_peak_kb // 1000
if recommended_kb <= 0:
    print("WARNING: recommended context budget is zero or negative; "
          "no usable context headroom for the configured system.")

R3. Refuse unknown models (model_context_kb: 0) in loop mode

Today detect_model_context returns 0 on unknown models and recommend_context silently falls through to the VRAM/RAM branches. A loop tick with an unknown model could still proceed against an arbitrarily-deranged budget.

Fix: Add a --loop-mode flag to vram_detect.py's CLI. When set:

  • model_context_kb == 0 is a hard error → print "ERROR: model context window is unknown in --loop-mode. Set 'Override context window' in config.md or pass --model." and exit 2.
  • net_kb < 16000 is a hard error → print "ERROR: available context ({}k) below 16k floor in --loop-mode (D13)." and exit 2.

The flag is optional. Non-loop callers (the dashboard, manual invocations) keep current behavior — only loops opt into the strict check. The future loop-runner.py will invoke vram_detect.py --loop-mode --json and expect either a 0 exit with a {...} JSON payload, or a 2 exit with a refuse message.

User's explicit Override context window in config.md (see _parse_config_model) is authoritative per D13: if a user has set an override, detect_model_context returns that override directly and the model_context_kb == 0 refuse never fires. The flow already honors this — no special code needed.

R4. Expose available_context_kb in JSON output

The loop runner needs a single authoritative figure for its per-tick budget. Today it would have to derive it from recommended_kb - framework_overhead_tokens itself, duplicating math.

Fix: Add available_context_kb to the JSON output block in main():

output = {
    "gpu_vram_gb": gpu_vram_gb,
    "ram_gb": ram_gb,
    "model_context_kb": model_context_kb,
    "framework_overhead_tokens": overhead_tokens,
    "recommended_kb": recommended_kb,           # net of overhead, before headroom
    "recommended_k": recommended_k,
    "headroom": headroom_pct / 100.0,
    "max_peak_context_kb": max_peak_kb,         # per-subtask peak (loop worktrees consume this)
    "available_context_kb": max_peak_kb,        # alias consumed by loop-runner.py; explicit field
    "loop_mode_eligible": max_peak_kb >= 16000, # boolean: passes the 16k floor check
}

available_context_kb = max_peak_kb (post-R1 value, headroom applied exactly once). Two field names for the same number so both human-readable names and the runner's contract field are stable.

R5. Add ## Loop Role Models section to config.md

config.md today only documents VRAM settings. The loop system needs an explicit place for users to declare which model/session plays each role. Per design/loops/functional.md §5, three roles exist: Implement:, Verify:, Orchestrate:. The framework never inspects the model of each role (D8/D12) — it only needs to know which harness session to invoke per role, which is a harness-level concern that loop.json already handles via roles.{implement,verify,orchestrate}.prompt. So config.md should document the expectation, not encode it.

Fix: Append a new ## Loop Role Models section to ~/.automaton/config.md:

## Loop Role Models

Loop ticks run three session roles. Roles are *sessions*, not models — a single model can fill multiple roles. Configure each loop's role-to-prompt binding in its `loop.json`; this section documents the framework's expectations only.

- **Implement:** — produces the artifact for this tick. Bound to `prompts/loop-implement.md` by default.
- **Verify:** — grades the artifact and emits the JSON verdict `{pass, score, reasons, next_hint}`. Bound to `prompts/loop-verifier.md`. The framework never inspects this role's model (D8); only its session.
- **Orchestrate:** — applies the verdict, calls exactly one `status.py` operation per tick, enforces brakes. Bound to `prompts/loop-orchestrate.md`.

Conflict-of-interest rule (D12): `Verify:` and `Implement:` must never be the same *session*. When two distinct sessions are infeasible (single-session harness), the runner falls back to session-only divergence — still safe.

Role context tiers are set per-loop in `loop.json`, not globally. The 16k floor (D13) applies regardless of tier.

R6. Add 4k tier to decompose.md and tighten the table

prompts/decompose.md (around :82-84 per the design audit) has a context budget table that omits the small-context 4k tier that a single implement role might fit in when overhead + task brief is small. Per the design audit it also states the 16k floor.

Fix: Open prompts/decompose.md, find the existing context budget table (search for 4k or context near the cited lines), add a row for the 4k tier and an explicit "≤ 16k: refuse" line above the table. Exact edits to be confirmed by reading the file at implementation time — this requirement locks the intent, not the diff.

If the existing table already covers 4k, this requirement is satisfied without edits; otherwise it is added. The 16k floor is the only hard refuse — 4k is a per-subtask peak recommendation, not a floor.

Acceptance Criteria

  • recommend_context returns max_peak_kb with headroom applied exactly once (verified by reading the function body — no inner * (100 - headroom_pct) // 100 at the three budget-construction sites).
  • vram_detect.py's JSON output no longer reports 8 / 6 for recommended_k / max_peak_k when the underlying budget is zero. The actual quotients (including 0) are emitted.
  • vram_detect.py --loop-mode exits 2 with the refuse message when the computed available context is < 16000 tokens OR model_context_kb == 0.
  • Without --loop-mode, the script preserves prior non-zero behavior on unknown / zero budgets (only a warning is added; no exit code change).
  • JSON output includes available_context_kb and loop_mode_eligible fields.
  • config.md includes the ## Loop Role Models section verbatim (text may be condensed, intent preserved).
  • prompts/decompose.md either acknowledges an existing 4k tier in its table or grows a 4k tier row, plus a ≤ 16k: refuse line.
  • New tests in tests/ (Python) cover: double-headroom removed (regression test), --loop-mode refuse paths, JSON field presence, decompose.md tier presence.
  • Pre-existing framework tests stay green: python3 -m pytest tests/ -v.

Non-Goals

  • Per-tick context budget enforcement inside vram_detect.py — that lives in loop-runner.py (task add-loop-runner). This task only exposes the numbers.
  • Removing model_context_kb == 0 fallback-to-VRAM in non-loop mode — that behavior is preserved for human CLI calls.
  • Changing how Override context window is parsed (it's already authoritative per D13; this task just relies on it).
  • Touching loop-verifier.md prompt contents — that's task add-loop-templates-onboarding (task 6). This task only adds a ## Loop Role Models reference section to config.md.

Dependencies

None. This is the first task in the bootstrap queue; downstream brakes and runner depend on it.

Out of Scope (handled by Tier 2 design/context-sizing/)

Per D17: comprehensive context-sizing cleanup (last-read-sha, drift detection, decompose.md full rework, dashboard "model context" panel) is a sibling design driven by the first loop workstream after task 7 lands. This task limits itself to the six Tier 1 items above.