Files
Lap Tran 4b3d92c9a2 Add framework design docs (model-divergence, rule agents, agent tab)
Design docs in design/framework/ covering:
- Model-divergence enforcement (conflict matrix, modes, auto-assignment)
- Rule Proposer agent (daily scan, proposes rules to RULE_PROPOSALS.md)
- Rule Reviewer agent (monthly consolidation, different LLM than Proposer)
- Agent tab redesign (phase roles + scheduled jobs, remove fake types)
- Schedules, conflict-of-interest, success criteria

Cross-references updated in AGENTS.md, README.md, CHANGELOG.md,
.onboarding.md, prompts/onboarding.md, design/loops/{README,BACKLOG}.md,
memory/v1-1-hardening-session.md.

Also restores scripts/automaton-cleanup.sh stub (was corrupted by
pytest test leak writing temp path into real stub).
2026-06-25 07:11:35 -04:00

21 KiB

Framework Agent Features — Technical Design

Companion to functional.md. This file is the implementation contract: every line here is what the implementation tasks build. Deviations require a [unreleased] CHANGELOG entry and a design doc update.

1. File Map (what v1 adds)

~/.automaton/
├── models.json                            # NEW — model manifest (see §2)
├── scripts/
│   ├── detect_models.py                   # NEW — probes opencode.json + localhost endpoints
│   └── status.py                          # EXTENDED — new flags (see §4)
├── prompts/
│   ├── rule-proposer.md                   # NEW — Rule Proposer session prompt
│   ├── rule-reviewer.md                   # NEW — Rule Reviewer session prompt
│   └── onboarding.md                      # EXTENDED — Step 2e (backlog check)
├── automaton/
│   └── dashboard/
│       ├── html/dashboard.js              # EXTENDED — remove AGENT_TYPE_META, two-section render
│       └── ui/app.py                      # EXTENDED — /api/phase-roles endpoint
├── .automaton/                            # (framework self-hosting: this is ~/.automaton/.automaton/)
│   ├── .state.rule-scan                   # NEW — Rule Proposer state (see §3)
│   ├── .state.rule-review                 # NEW — Rule Reviewer state (see §3)
│   ├── RULE_PROPOSALS.md                  # NEW — Rule Proposer output (append-per-run)
│   ├── RULE_REVIEW.md                     # NEW — Rule Reviewer output (append-per-run)
│   └── automaton-rule-scan.sh             # NEW — generated by --install-rule-scan-schedule
│       automaton-rule-review.sh           # NEW — generated by --install-rule-review-schedule
└── tests/
    ├── test_model_divergence.py           # NEW — manifest, conflict matrix, --transition --model
    ├── test_rule_agents.py                # NEW — scan flows, state files, output schemas
    └── test_dashboard_phase_roles.py      # NEW — /api/phase-roles, two-section render

Per-project paths mirror the loop convention: {project}/.automaton/.state.rule-scan, {project}/.automaton/RULE_PROPOSALS.md, etc. For framework self-hosting, the project is ~/.automaton/ itself.

2. models.json Schema

{
  "schema_version": 1,
  "default": "glm-4.6",
  "advised": true,
  "models": [
    {
      "name": "glm-4.6",
      "provider": "opencode",
      "context_window": 131072,
      "location": "remote"
    },
    {
      "name": "qwen3-coder",
      "provider": "opencode",
      "context_window": 131072,
      "location": "remote"
    },
    {
      "name": "llama-3.3-70b",
      "provider": "localhost",
      "context_window": 32768,
      "location": "http://localhost:8080"
    }
  ]
}
  • default: model name used when no role-specific binding exists. Must be present in models[].
  • advised: bool. If true, single-LLM mode prints a one-time advisory recommending a second model, then goes silent.
  • models[]: roster. name is the unique key. provider is informational. context_window is informational (framework never inspects capability, D8). location is "remote" or a localhost URL (for detect_models.py probing).
  • Missing file → single-LLM mode (backward compatible). All model-divergence commands are no-ops.
  • 0-1 models → single-LLM mode. Advisory once if advised: true.
  • 2+ models → multi-LLM mode. Hard-block on conflict matrix.

2.1 detect_models.py

python3 scripts/detect_models.py [--json]
  1. Parse opencode.json (or opencode.jsonc) for provider+model entries.
  2. Probe localhost endpoints: http://localhost:8080/v1/models, http://localhost:11434/api/tags (Ollama), http://localhost:1234/v1/models (LM Studio), http://localhost:8000/v1/models (vLLM).
  3. Merge results, emit a candidate models.json to stdout (or write if --json not set).
  4. Used by install.sh / update.sh / upgrade.sh to bootstrap or refresh models.json.

3. State Schemas

3.1 .state.rule-scan

{
  "schema_version": 1,
  "last_scan_at": "2026-06-25T10:00:00Z",
  "last_scanned_task": "fix-context-sizing",
  "proposals_count": 3,
  "scanned_tasks_count": 12
}
  • last_scanned_task: the most recent task name scanned. Next scan starts after this task (alphabetical or mtime order).
  • proposals_count: cumulative count of proposals written to RULE_PROPOSALS.md.
  • Stored at {project}/.automaton/.state.rule-scan. Missing file → first run scans all completed tasks.

3.2 .state.rule-review

{
  "schema_version": 1,
  "last_review_at": "2026-06-25T10:00:00Z",
  "contradictions_found": 2,
  "stale_rules": 5,
  "merges_suggested": 1
}
  • Stored at {project}/.automaton/.state.rule-review. Missing file → first run reviews all rules.

3.3 .state.models (per-task)

{
  "schema_version": 1,
  "implement": "glm-4.6",
  "code_review": "qwen3-coder",
  "bug_find": "qwen3-coder",
  "adversarial_bug_find": "llama-3.3-70b",
  "referee": "llama-3.3-70b",
  "doc_review": null
}
  • Stored at {task}/.state.models. One file per task.
  • Written by --transition --model <name> when entering a phase.
  • Read by --claim (conflict-matrix check) and --audit (violation detection).
  • Roles not yet filled are null or absent.

4. status.py New Flags

All model-divergence and rule-agent commands route through status.py — no second enforcement surface.

# Model-divergence
status.py --transition <phase> --task <t> [--model <name>]   Records model in .state.models; checks conflict matrix
status.py --claim --task <t> --agent <a> [--model <name>]    Refuses if model conflicts with filled roles (multi-LLM mode)
status.py --audit                                            EXTENDED — +model_divergence category
status.py --can-edit [...]                                   UNCHANGED

# Rule agents
status.py --rule-scan [--project <p>] [--dry-run]            Scan completed tasks, propose rules to RULE_PROPOSALS.md
status.py --install-rule-scan-schedule [--interval S]        Install OS-native unit for --rule-scan (default daily)
status.py --rule-review [--project <p>] [--dry-run]          Consolidate .rules.md, write RULE_REVIEW.md
status.py --install-rule-review-schedule [--interval S]      Install OS-native unit for --rule-review (default monthly)

4.1 --transition --model flow

  1. Load models.json. If missing or single-LLM mode → record model (advisory), no conflict check.
  2. If multi-LLM mode: load .state.models for the task. Check the role being entered against the conflict matrix (§5).
  3. If --model not provided: auto-assign next-available non-conflicting model from models[]. Refuse if none available.
  4. If --model provided: verify it's in models[]. Check conflict matrix. Refuse on violation.
  5. Write role: model to .state.models. Transition the phase.

4.2 --claim --model flow

  1. Load models.json. If single-LLM mode → existing claim logic, no model check.
  2. If multi-LLM mode: load .state.models. Determine the role for the phase being claimed. Check conflict matrix against already-filled roles.
  3. Refuse if the claiming agent's model conflicts. Error message names the conflicting role and model.

4.3 --audit extension

New audit category model_divergence:

  • For each task with .state.models: check all filled roles against the conflict matrix.
  • Flag violations as severity: high (conflict-of-interest is a correctness issue, not a style issue).
  • Output format mirrors existing audit categories.

5. Conflict Matrix (implementation)

CONFLICT_MATRIX = {
    "code_review":             {"implement"},
    "bug_find":                {"implement"},
    "adversarial_bug_find":    {"implement", "bug_find"},
    "referee":                 {"implement", "bug_find", "adversarial_bug_find"},
    "loop-verify":             {"loop-implement"},
}
  • Key = role being entered. Value = set of roles that must have a different model.
  • doc_review, code_review, bug_find are NOT in conflict with each other (only bug_find ↔ adversarial_bug_find conflicts).
  • Check function: def _check_conflict(state_models: dict, role: str, model: str, matrix: dict) -> Optional[str] — returns the conflicting role name or None.

6. Rule Proposer Flow (--rule-scan)

1. Load .state.rule-scan (or init if missing).
2. Find completed tasks since last_scanned_task:
   - Scan tasks/complete/ and tasks with .state phase=complete
   - Filter by mtime > last_scan_at (or all if first run)
   - Sort by mtime ascending
3. For each task:
   a. Read BUG_REPORT.md, ADVERSARIAL_BUG_REPORT.md, VERDICT.md (skip if none exist)
   b. Read current .rules.md (for dedup context — capped at 4k tokens)
   c. Build proposer prompt (see §7)
   d. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_proposer_model)
   e. Parse LLM output for proposed rules (expect RULE_PROPOSALS.md format)
   f. Append proposals to RULE_PROPOSALS.md
   g. Update .state.rule-scan (last_scanned_task, proposals_count)
4. Write final .state.rule-scan with last_scan_at = now.
  • --dry-run: list tasks that would be scanned, do not invoke harness.
  • --project: scope to a project (default: framework dir).
  • Errors during a single task scan do not abort the run; the scan continues to the next task and logs the error.

7. Rule Proposer Prompt Shape (rule-proposer.md)

# Rule Proposer — {date}

You are scanning completed tasks for failure patterns that should become rules.

## Current rules (read-only, for dedup)
{current_rules}        # .rules.md content, capped at 4k tokens

## Task failure artifacts
{bug_report}           # BUG_REPORT.md content, capped at 2k tokens
{adversarial_report}   # ADVERSARIAL_BUG_REPORT.md, capped at 2k tokens
{verdict}              # VERDICT.md, capped at 2k tokens

## What to do
For each distinct failure pattern you observe:
1. Check if a rule already exists in .rules.md that covers it. If so, skip.
2. If no existing rule covers it, propose a new rule with:
   - A concrete example from the task artifacts
   - The proposed rule text as it would appear in .rules.md

## Output (strict markdown, no JSON)
## Proposed Rule: {title}
**Source**: tasks/{task-name}/VERDICT.md
**Pattern**: {one-line description}
**Example**:
{concrete snippet}
**Proposed rule text**:
{rule text}

---

No {model} token in the prompt — the model is selected by the caller and passed to _invoke_harness.

8. Rule Reviewer Flow (--rule-review)

1. Load .state.rule-review (or init if missing).
2. Read .rules.md (full file).
3. Read recent RULE_PROPOSALS.md entries (since last_review_at).
4. Read recent completed-task summaries (last 30 days) for staleness context.
5. Build reviewer prompt (see §9).
6. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_reviewer_model).
7. Parse LLM output for review sections (contradictions, stale, missing examples, merges).
8. Append to RULE_REVIEW.md.
9. Update .state.rule-review.
  • --dry-run: report what would be reviewed, do not invoke harness.

9. Rule Reviewer Prompt Shape (rule-reviewer.md)

# Rule Reviewer — {date}

You are consolidating .rules.md for contradictions, staleness, and missing examples.

## Current rules (full)
{rules_content}        # .rules.md, full file

## Recent proposals (since last review)
{recent_proposals}     # RULE_PROPOSALS.md entries since last_review_at

## Recent completed tasks (last 30 days, for staleness context)
{task_summaries}       # one-line per task: name + phase + completion date

## What to check
1. Contradictions: rules that conflict with each other.
2. Stale rules: no observed instance in last 30 days.
3. Rules missing examples: any rule without a concrete example.
4. Merge candidates: overlapping rules that could be consolidated.

## Output (strict markdown, no JSON)
## Contradictions Found
- ...
## Stale Rules
- ...
## Rules Missing Examples
- ...
## Merge Candidates
- ...

10. Harness Invocation (direct, not loop)

Rule agents reuse loop-runner._invoke_harness directly — they are NOT loops. The function signature (from loop-runner.py:366-404):

def _invoke_harness(harness_command: str, prompt_content: str, cwd: str, env: dict = None) -> str:

status.py --rule-scan calls this as:

from loop_runner import _invoke_harness
output = _invoke_harness(
    harness_command=rule_harness_command,  # from schedule config or default
    prompt_content=resolved_prompt,        # rule-proposer.md with tokens substituted
    cwd=str(project_dir),
    env={"AUTOMATON_RULE_ROLE": "proposer"}
)

{model} substitution: if the harness command contains {model}, it's replaced with the rule agent's configured model. Until model-divergence ships, this is the default model.

10.1 Schedule Config for Rule Agents

Rule agents do not use loop.json. Their config is embedded in the schedule stub:

#!/usr/bin/env bash
cd "<project_root>"
python3 "<framework>/scripts/status.py" --rule-scan --model <name>

The --model flag is optional and ignored in single-LLM mode. In multi-LLM mode it sets the rule agent's model (subject to conflict-of-interest checks once enforced).

11. Scheduler Unit Generation

Mirrors cmd_install_cleanup_schedule (status.py:2188-2260) exactly:

11.1 --install-rule-scan-schedule

def cmd_install_rule_scan_schedule(args) -> int:
    interval = args.interval if args.interval else 86400  # daily
    # 1. Write stub: automaton-rule-scan.sh
    # 2. Platform dispatch:
    #    Darwin  → ~/Library/LaunchAgents/com.automaton.rule-scan.plist
    #    Linux   → crontab block via _install_cron_block_generic
    #    Windows → schtasks /create /tn "AutomatonRuleScan"

11.2 --install-rule-review-schedule

def cmd_install_rule_review_schedule(args) -> int:
    interval = args.interval if args.interval else 2592000  # monthly
    # Same pattern, labels: com.automaton.rule-review / AutomatonRuleReview

11.3 _list_scheduled_jobs extension

_list_scheduled_jobs (status.py:2282) gains recognition for new labels:

def _launchd_label_kind(label: str) -> tuple[str, str]:
    if label.startswith("com.automaton.loop."):
        return ("loop", label[len("com.automaton.loop."):])
    if label in ("com.automaton.cleanup",):
        return ("cleanup", "")
    if label in ("com.automaton.rule-scan",):
        return ("rule-scan", "")
    if label in ("com.automaton.rule-review",):
        return ("rule-review", "")
    ...

This makes rule-scan and rule-review jobs appear in /api/scheduled with their real kind, which the Agent tab renders directly.

12. Agent Tab Data Flow

12.1 New endpoint: /api/phase-roles

app.py gains a handler:

elif self.path == "/api/phase-roles":
    self._serve_phase_roles()
def _serve_phase_roles(self):
    # 1. Parse .agent.md Agent Configuration for role definitions
    # 2. Load all tasks via status.py module
    # 3. For each role, count tasks in that role's phases
    # 4. Return JSON:
    {
      "roles": [
        {"id": "researcher", "label": "Researcher", "icon": "🔬",
         "phases": ["research", "decomposition", "design", "test_design"],
         "active_tasks": 2, "status": "active"},
        ...
      ],
      "available": True
    }

Role icons (self-documenting, per .rules.md Self-Documenting UI Names):

Role Icon
researcher 🔬
implementer ⚙️
code-reviewer 👁️
bug-hunter 🐛
referee ⚖️
orchestrator 🎯

12.2 dashboard.js changes

Remove: AGENT_TYPE_META (lines 330-334), AGENT_TYPES (337), AGENT_TYPE_META_FALLBACK (338), _resolveAgentType (351-354).

Replace renderAgentTab with a two-section render:

async function renderAgentTab() {
    const panel = document.getElementById('agent-panel');
    panel.innerHTML = '<div class="bg-loading">Loading…</div>';

    const [rolesRes, schedRes] = await Promise.all([
        fetch('/api/phase-roles').then(r => r.json()).catch(() => ({roles: [], available: false})),
        fetchSchedule(),
    ]);

    // Section 1: Phase Roles
    const rolesHtml = rolesRes.available ? renderPhaseRoles(rolesRes.roles)
        : '<div class="bg-empty">Phase roles require .agent.md Agent Configuration.</div>';

    // Section 2: Scheduled Jobs
    const jobsHtml = renderScheduledJobs(schedRes.jobs || []);

    panel.innerHTML = `
        <div class="agent-section">
            <h3>Phase Roles</h3>
            <div class="bg-grid">${rolesHtml}</div>
        </div>
        <div class="agent-section">
            <h3>Scheduled Jobs</h3>
            <div class="bg-grid">${jobsHtml}</div>
        </div>`;
}

renderScheduledJobs uses job.kind directly (no fake type resolution):

const JOB_META = {
    cleanup:     { icon: '🧹', label: 'Cleanup Archiver' },
    loop:        { icon: '🔄', label: (j) => `Loop: ${j.name}` },
    'rule-scan': { icon: '📝', label: 'Rule Proposer' },
    'rule-review':{ icon: '📋', label: 'Rule Reviewer' },
};

13. Loop Integration (model-divergence)

13.1 loop.json per-role model

"roles": {
  "implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
  "verify":    {"prompt": "loop-verifier.md",  "model": "qwen3-coder"},
  "orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
}
  • model is optional. If absent, uses models.json default.
  • loop-verify model is checked against loop-implement model in --check-gate (multi-LLM mode).

13.2 {model} substitution in _invoke_harness

loop-runner.py:366-404 _invoke_harness gains {model} token substitution:

def _invoke_harness(harness_command, prompt_content, cwd, env=None, model=None):
    if model and "{model}" in harness_command:
        harness_command = harness_command.replace("{model}", model)
    ...

The caller passes model from the role config. If the harness command has no {model} token, the model is informational only (the harness picks its own).

13.3 --check-gate model-divergence check

In multi-LLM mode, --check-gate adds:

  • Load loop.json roles. Compare verify.model vs implement.model.
  • If same model and multi-LLM mode → halt as model_conflict (new halt reason, or reuse human_intervention with a descriptive message).

14. Test Coverage

14.1 test_model_divergence.py

  • test_models_json_missing_single_llm_mode — no file → advisory, no blocks.
  • test_single_model_advisory_once — 1 model, advised: true → advisory printed once, then silent.
  • test_multi_llm_conflict_matrix — 2+ models, --transition --model records, --claim refuses conflict.
  • test_auto_assign_next_available — no --model flag → auto-assigns non-conflicting model.
  • test_auto_assign_exhausted — all models conflict → refuse.
  • test_audit_model_divergence — --audit flags conflict-matrix violations.
  • test_loop_verify_neq_implement — --check-gate halts on same model in multi-LLM mode.

14.2 test_rule_agents.py

  • test_rule_scan_finds_completed_tasks — seeded completed task with VERDICT.md → proposal written.
  • test_rule_scan_state_tracking — .state.rule-scan updated with last_scanned_task + count.
  • test_rule_scan_dedup — existing rule in .rules.md → not re-proposed.
  • test_rule_scan_dry_run — no harness invocation, lists candidates.
  • test_rule_review_finds_contradictions — seeded .rules.md with contradiction → review written.
  • test_rule_review_state_tracking — .state.rule-review updated.
  • test_install_rule_scan_schedule — stub + plist created with correct labels.
  • test_install_rule_review_schedule — stub + plist created with correct labels.

14.3 test_dashboard_phase_roles.py

  • test_api_phase_roles — /api/phase-roles returns 6 roles with correct phases.
  • test_phase_roles_active_count — tasks in phases → correct active_tasks count.
  • test_scheduled_jobs_new_kinds — rule-scan and rule-review jobs appear with correct kind.
  • test_agent_type_meta_removed — AGENT_TYPE_META no longer in dashboard.js (grep test).

15. Rollout (3 sequential tasks for model-divergence)

The model-divergence-enforcement parent task decomposes into 3 subtasks:

  1. manifest+detection: models.json schema, detect_models.py, install.sh/update.sh/upgrade.sh integration, config.md section, onboarding Step 2d.
  2. interactive enforcement+audit: .state.models, --transition --model, --claim --model conflict check, --audit model_divergence category, dashboard badges.
  3. loop enforcement+dashboard: loop.json per-role model, {model} substitution, --check-gate model check, loop dashboard badges.

Rule agents (FW-2, FW-3) and Agent tab (FW-1) are backlog items, picked up after model-divergence ships (for FW-2/FW-3) or independently (for FW-1).

16. Locked Decision Index

All decisions referenced by (Fn) are in README.md § "Locked decisions". Implementation must conform. Deviations require a design doc update + [unreleased] CHANGELOG entry.