Files

537 lines
21 KiB
Markdown
Raw Permalink Normal View History

# Framework Agent Features — Technical Design
Companion to `functional.md`. This file is the implementation contract: every line here is what the implementation tasks build. Deviations require a `[unreleased]` CHANGELOG entry and a design doc update.
## 1. File Map (what v1 adds)
```
~/.automaton/
├── models.json # NEW — model manifest (see §2)
├── scripts/
│ ├── detect_models.py # NEW — probes opencode.json + localhost endpoints
│ └── status.py # EXTENDED — new flags (see §4)
├── prompts/
│ ├── rule-proposer.md # NEW — Rule Proposer session prompt
│ ├── rule-reviewer.md # NEW — Rule Reviewer session prompt
│ └── onboarding.md # EXTENDED — Step 2e (backlog check)
├── automaton/
│ └── dashboard/
│ ├── html/dashboard.js # EXTENDED — remove AGENT_TYPE_META, two-section render
│ └── ui/app.py # EXTENDED — /api/phase-roles endpoint
├── .automaton/ # (framework self-hosting: this is ~/.automaton/.automaton/)
│ ├── .state.rule-scan # NEW — Rule Proposer state (see §3)
│ ├── .state.rule-review # NEW — Rule Reviewer state (see §3)
│ ├── RULE_PROPOSALS.md # NEW — Rule Proposer output (append-per-run)
│ ├── RULE_REVIEW.md # NEW — Rule Reviewer output (append-per-run)
│ └── automaton-rule-scan.sh # NEW — generated by --install-rule-scan-schedule
│ automaton-rule-review.sh # NEW — generated by --install-rule-review-schedule
└── tests/
├── test_model_divergence.py # NEW — manifest, conflict matrix, --transition --model
├── test_rule_agents.py # NEW — scan flows, state files, output schemas
└── test_dashboard_phase_roles.py # NEW — /api/phase-roles, two-section render
```
Per-project paths mirror the loop convention: `{project}/.automaton/.state.rule-scan`, `{project}/.automaton/RULE_PROPOSALS.md`, etc. For framework self-hosting, the project is `~/.automaton/` itself.
## 2. `models.json` Schema
```json
{
"schema_version": 1,
"default": "glm-4.6",
"advised": true,
"models": [
{
"name": "glm-4.6",
"provider": "opencode",
"context_window": 131072,
"location": "remote"
},
{
"name": "qwen3-coder",
"provider": "opencode",
"context_window": 131072,
"location": "remote"
},
{
"name": "llama-3.3-70b",
"provider": "localhost",
"context_window": 32768,
"location": "http://localhost:8080"
}
]
}
```
- `default`: model name used when no role-specific binding exists. Must be present in `models[]`.
- `advised`: bool. If `true`, single-LLM mode prints a one-time advisory recommending a second model, then goes silent.
- `models[]`: roster. `name` is the unique key. `provider` is informational. `context_window` is informational (framework never inspects capability, D8). `location` is `"remote"` or a localhost URL (for `detect_models.py` probing).
- **Missing file** → single-LLM mode (backward compatible). All model-divergence commands are no-ops.
- **0-1 models** → single-LLM mode. Advisory once if `advised: true`.
- **2+ models** → multi-LLM mode. Hard-block on conflict matrix.
### 2.1 `detect_models.py`
```
python3 scripts/detect_models.py [--json]
```
1. Parse `opencode.json` (or `opencode.jsonc`) for provider+model entries.
2. Probe localhost endpoints: `http://localhost:8080/v1/models`, `http://localhost:11434/api/tags` (Ollama), `http://localhost:1234/v1/models` (LM Studio), `http://localhost:8000/v1/models` (vLLM).
3. Merge results, emit a candidate `models.json` to stdout (or write if `--json` not set).
4. Used by `install.sh` / `update.sh` / `upgrade.sh` to bootstrap or refresh `models.json`.
## 3. State Schemas
### 3.1 `.state.rule-scan`
```json
{
"schema_version": 1,
"last_scan_at": "2026-06-25T10:00:00Z",
"last_scanned_task": "fix-context-sizing",
"proposals_count": 3,
"scanned_tasks_count": 12
}
```
- `last_scanned_task`: the most recent task name scanned. Next scan starts after this task (alphabetical or mtime order).
- `proposals_count`: cumulative count of proposals written to `RULE_PROPOSALS.md`.
- Stored at `{project}/.automaton/.state.rule-scan`. Missing file → first run scans all completed tasks.
### 3.2 `.state.rule-review`
```json
{
"schema_version": 1,
"last_review_at": "2026-06-25T10:00:00Z",
"contradictions_found": 2,
"stale_rules": 5,
"merges_suggested": 1
}
```
- Stored at `{project}/.automaton/.state.rule-review`. Missing file → first run reviews all rules.
### 3.3 `.state.models` (per-task)
```json
{
"schema_version": 1,
"implement": "glm-4.6",
"code_review": "qwen3-coder",
"bug_find": "qwen3-coder",
"adversarial_bug_find": "llama-3.3-70b",
"referee": "llama-3.3-70b",
"doc_review": null
}
```
- Stored at `{task}/.state.models`. One file per task.
- Written by `--transition --model <name>` when entering a phase.
- Read by `--claim` (conflict-matrix check) and `--audit` (violation detection).
- Roles not yet filled are `null` or absent.
## 4. `status.py` New Flags
All model-divergence and rule-agent commands route through `status.py` — no second enforcement surface.
```
# Model-divergence
status.py --transition <phase> --task <t> [--model <name>] Records model in .state.models; checks conflict matrix
status.py --claim --task <t> --agent <a> [--model <name>] Refuses if model conflicts with filled roles (multi-LLM mode)
status.py --audit EXTENDED — +model_divergence category
status.py --can-edit [...] UNCHANGED
# Rule agents
status.py --rule-scan [--project <p>] [--dry-run] Scan completed tasks, propose rules to RULE_PROPOSALS.md
status.py --install-rule-scan-schedule [--interval S] Install OS-native unit for --rule-scan (default daily)
status.py --rule-review [--project <p>] [--dry-run] Consolidate .rules.md, write RULE_REVIEW.md
status.py --install-rule-review-schedule [--interval S] Install OS-native unit for --rule-review (default monthly)
```
### 4.1 `--transition --model` flow
1. Load `models.json`. If missing or single-LLM mode → record model (advisory), no conflict check.
2. If multi-LLM mode: load `.state.models` for the task. Check the role being entered against the conflict matrix (§5).
3. If `--model` not provided: auto-assign next-available non-conflicting model from `models[]`. Refuse if none available.
4. If `--model` provided: verify it's in `models[]`. Check conflict matrix. Refuse on violation.
5. Write `role: model` to `.state.models`. Transition the phase.
### 4.2 `--claim --model` flow
1. Load `models.json`. If single-LLM mode → existing claim logic, no model check.
2. If multi-LLM mode: load `.state.models`. Determine the role for the phase being claimed. Check conflict matrix against already-filled roles.
3. Refuse if the claiming agent's model conflicts. Error message names the conflicting role and model.
### 4.3 `--audit` extension
New audit category `model_divergence`:
- For each task with `.state.models`: check all filled roles against the conflict matrix.
- Flag violations as `severity: high` (conflict-of-interest is a correctness issue, not a style issue).
- Output format mirrors existing audit categories.
## 5. Conflict Matrix (implementation)
```python
CONFLICT_MATRIX = {
"code_review": {"implement"},
"bug_find": {"implement"},
"adversarial_bug_find": {"implement", "bug_find"},
"referee": {"implement", "bug_find", "adversarial_bug_find"},
"loop-verify": {"loop-implement"},
}
```
- Key = role being entered. Value = set of roles that must have a different model.
- `doc_review`, `code_review`, `bug_find` are NOT in conflict with each other (only `bug_find` ↔ `adversarial_bug_find` conflicts).
- Check function: `def _check_conflict(state_models: dict, role: str, model: str, matrix: dict) -> Optional[str]` — returns the conflicting role name or `None`.
## 6. Rule Proposer Flow (`--rule-scan`)
```
1. Load .state.rule-scan (or init if missing).
2. Find completed tasks since last_scanned_task:
- Scan tasks/complete/ and tasks with .state phase=complete
- Filter by mtime > last_scan_at (or all if first run)
- Sort by mtime ascending
3. For each task:
a. Read BUG_REPORT.md, ADVERSARIAL_BUG_REPORT.md, VERDICT.md (skip if none exist)
b. Read current .rules.md (for dedup context — capped at 4k tokens)
c. Build proposer prompt (see §7)
d. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_proposer_model)
e. Parse LLM output for proposed rules (expect RULE_PROPOSALS.md format)
f. Append proposals to RULE_PROPOSALS.md
g. Update .state.rule-scan (last_scanned_task, proposals_count)
4. Write final .state.rule-scan with last_scan_at = now.
```
- `--dry-run`: list tasks that would be scanned, do not invoke harness.
- `--project`: scope to a project (default: framework dir).
- Errors during a single task scan do not abort the run; the scan continues to the next task and logs the error.
## 7. Rule Proposer Prompt Shape (`rule-proposer.md`)
```
# Rule Proposer — {date}
You are scanning completed tasks for failure patterns that should become rules.
## Current rules (read-only, for dedup)
{current_rules} # .rules.md content, capped at 4k tokens
## Task failure artifacts
{bug_report} # BUG_REPORT.md content, capped at 2k tokens
{adversarial_report} # ADVERSARIAL_BUG_REPORT.md, capped at 2k tokens
{verdict} # VERDICT.md, capped at 2k tokens
## What to do
For each distinct failure pattern you observe:
1. Check if a rule already exists in .rules.md that covers it. If so, skip.
2. If no existing rule covers it, propose a new rule with:
- A concrete example from the task artifacts
- The proposed rule text as it would appear in .rules.md
## Output (strict markdown, no JSON)
## Proposed Rule: {title}
**Source**: tasks/{task-name}/VERDICT.md
**Pattern**: {one-line description}
**Example**:
{concrete snippet}
**Proposed rule text**:
{rule text}
---
```
No `{model}` token in the prompt — the model is selected by the caller and passed to `_invoke_harness`.
## 8. Rule Reviewer Flow (`--rule-review`)
```
1. Load .state.rule-review (or init if missing).
2. Read .rules.md (full file).
3. Read recent RULE_PROPOSALS.md entries (since last_review_at).
4. Read recent completed-task summaries (last 30 days) for staleness context.
5. Build reviewer prompt (see §9).
6. Invoke harness via loop-runner._invoke_harness(prompt, model=rule_reviewer_model).
7. Parse LLM output for review sections (contradictions, stale, missing examples, merges).
8. Append to RULE_REVIEW.md.
9. Update .state.rule-review.
```
- `--dry-run`: report what would be reviewed, do not invoke harness.
## 9. Rule Reviewer Prompt Shape (`rule-reviewer.md`)
```
# Rule Reviewer — {date}
You are consolidating .rules.md for contradictions, staleness, and missing examples.
## Current rules (full)
{rules_content} # .rules.md, full file
## Recent proposals (since last review)
{recent_proposals} # RULE_PROPOSALS.md entries since last_review_at
## Recent completed tasks (last 30 days, for staleness context)
{task_summaries} # one-line per task: name + phase + completion date
## What to check
1. Contradictions: rules that conflict with each other.
2. Stale rules: no observed instance in last 30 days.
3. Rules missing examples: any rule without a concrete example.
4. Merge candidates: overlapping rules that could be consolidated.
## Output (strict markdown, no JSON)
## Contradictions Found
- ...
## Stale Rules
- ...
## Rules Missing Examples
- ...
## Merge Candidates
- ...
```
## 10. Harness Invocation (direct, not loop)
Rule agents reuse `loop-runner._invoke_harness` directly — they are NOT loops. The function signature (from `loop-runner.py:366-404`):
```python
def _invoke_harness(harness_command: str, prompt_content: str, cwd: str, env: dict = None) -> str:
```
`status.py --rule-scan` calls this as:
```python
from loop_runner import _invoke_harness
output = _invoke_harness(
harness_command=rule_harness_command, # from schedule config or default
prompt_content=resolved_prompt, # rule-proposer.md with tokens substituted
cwd=str(project_dir),
env={"AUTOMATON_RULE_ROLE": "proposer"}
)
```
`{model}` substitution: if the harness command contains `{model}`, it's replaced with the rule agent's configured model. Until model-divergence ships, this is the default model.
### 10.1 Schedule Config for Rule Agents
Rule agents do not use `loop.json`. Their config is embedded in the schedule stub:
```bash
#!/usr/bin/env bash
cd "<project_root>"
python3 "<framework>/scripts/status.py" --rule-scan --model <name>
```
The `--model` flag is optional and ignored in single-LLM mode. In multi-LLM mode it sets the rule agent's model (subject to conflict-of-interest checks once enforced).
## 11. Scheduler Unit Generation
Mirrors `cmd_install_cleanup_schedule` (`status.py:2188-2260`) exactly:
### 11.1 `--install-rule-scan-schedule`
```python
def cmd_install_rule_scan_schedule(args) -> int:
interval = args.interval if args.interval else 86400 # daily
# 1. Write stub: automaton-rule-scan.sh
# 2. Platform dispatch:
# Darwin → ~/Library/LaunchAgents/com.automaton.rule-scan.plist
# Linux → crontab block via _install_cron_block_generic
# Windows → schtasks /create /tn "AutomatonRuleScan"
```
### 11.2 `--install-rule-review-schedule`
```python
def cmd_install_rule_review_schedule(args) -> int:
interval = args.interval if args.interval else 2592000 # monthly
# Same pattern, labels: com.automaton.rule-review / AutomatonRuleReview
```
### 11.3 `_list_scheduled_jobs` extension
`_list_scheduled_jobs` (`status.py:2282`) gains recognition for new labels:
```python
def _launchd_label_kind(label: str) -> tuple[str, str]:
if label.startswith("com.automaton.loop."):
return ("loop", label[len("com.automaton.loop."):])
if label in ("com.automaton.cleanup",):
return ("cleanup", "")
if label in ("com.automaton.rule-scan",):
return ("rule-scan", "")
if label in ("com.automaton.rule-review",):
return ("rule-review", "")
...
```
This makes rule-scan and rule-review jobs appear in `/api/scheduled` with their real `kind`, which the Agent tab renders directly.
## 12. Agent Tab Data Flow
### 12.1 New endpoint: `/api/phase-roles`
`app.py` gains a handler:
```python
elif self.path == "/api/phase-roles":
self._serve_phase_roles()
```
```python
def _serve_phase_roles(self):
# 1. Parse .agent.md Agent Configuration for role definitions
# 2. Load all tasks via status.py module
# 3. For each role, count tasks in that role's phases
# 4. Return JSON:
{
"roles": [
{"id": "researcher", "label": "Researcher", "icon": "🔬",
"phases": ["research", "decomposition", "design", "test_design"],
"active_tasks": 2, "status": "active"},
...
],
"available": True
}
```
Role icons (self-documenting, per `.rules.md` Self-Documenting UI Names):
| Role | Icon |
|---|---|
| researcher | 🔬 |
| implementer | ⚙️ |
| code-reviewer | 👁️ |
| bug-hunter | 🐛 |
| referee | ⚖️ |
| orchestrator | 🎯 |
### 12.2 `dashboard.js` changes
**Remove**: `AGENT_TYPE_META` (lines 330-334), `AGENT_TYPES` (337), `AGENT_TYPE_META_FALLBACK` (338), `_resolveAgentType` (351-354).
**Replace `renderAgentTab`** with a two-section render:
```javascript
async function renderAgentTab() {
const panel = document.getElementById('agent-panel');
panel.innerHTML = '<div class="bg-loading">Loading…</div>';
const [rolesRes, schedRes] = await Promise.all([
fetch('/api/phase-roles').then(r => r.json()).catch(() => ({roles: [], available: false})),
fetchSchedule(),
]);
// Section 1: Phase Roles
const rolesHtml = rolesRes.available ? renderPhaseRoles(rolesRes.roles)
: '<div class="bg-empty">Phase roles require .agent.md Agent Configuration.</div>';
// Section 2: Scheduled Jobs
const jobsHtml = renderScheduledJobs(schedRes.jobs || []);
panel.innerHTML = `
<div class="agent-section">
<h3>Phase Roles</h3>
<div class="bg-grid">${rolesHtml}</div>
</div>
<div class="agent-section">
<h3>Scheduled Jobs</h3>
<div class="bg-grid">${jobsHtml}</div>
</div>`;
}
```
**`renderScheduledJobs`** uses `job.kind` directly (no fake type resolution):
```javascript
const JOB_META = {
cleanup: { icon: '🧹', label: 'Cleanup Archiver' },
loop: { icon: '🔄', label: (j) => `Loop: ${j.name}` },
'rule-scan': { icon: '📝', label: 'Rule Proposer' },
'rule-review':{ icon: '📋', label: 'Rule Reviewer' },
};
```
## 13. Loop Integration (model-divergence)
### 13.1 `loop.json` per-role model
```json
"roles": {
"implement": {"prompt": "loop-implement.md", "model": "glm-4.6"},
"verify": {"prompt": "loop-verifier.md", "model": "qwen3-coder"},
"orchestrate":{"prompt": "loop-orchestrate.md", "model": "glm-4.6"}
}
```
- `model` is optional. If absent, uses `models.json` `default`.
- `loop-verify` model is checked against `loop-implement` model in `--check-gate` (multi-LLM mode).
### 13.2 `{model}` substitution in `_invoke_harness`
`loop-runner.py:366-404` `_invoke_harness` gains `{model}` token substitution:
```python
def _invoke_harness(harness_command, prompt_content, cwd, env=None, model=None):
if model and "{model}" in harness_command:
harness_command = harness_command.replace("{model}", model)
...
```
The caller passes `model` from the role config. If the harness command has no `{model}` token, the model is informational only (the harness picks its own).
### 13.3 `--check-gate` model-divergence check
In multi-LLM mode, `--check-gate` adds:
- Load `loop.json` roles. Compare `verify.model` vs `implement.model`.
- If same model and multi-LLM mode → halt as `model_conflict` (new halt reason, or reuse `human_intervention` with a descriptive message).
## 14. Test Coverage
### 14.1 `test_model_divergence.py`
- `test_models_json_missing_single_llm_mode` — no file → advisory, no blocks.
- `test_single_model_advisory_once` — 1 model, `advised: true` → advisory printed once, then silent.
- `test_multi_llm_conflict_matrix` — 2+ models, `--transition --model` records, `--claim` refuses conflict.
- `test_auto_assign_next_available` — no `--model` flag → auto-assigns non-conflicting model.
- `test_auto_assign_exhausted` — all models conflict → refuse.
- `test_audit_model_divergence` — `--audit` flags conflict-matrix violations.
- `test_loop_verify_neq_implement` — `--check-gate` halts on same model in multi-LLM mode.
### 14.2 `test_rule_agents.py`
- `test_rule_scan_finds_completed_tasks` — seeded completed task with VERDICT.md → proposal written.
- `test_rule_scan_state_tracking` — `.state.rule-scan` updated with last_scanned_task + count.
- `test_rule_scan_dedup` — existing rule in `.rules.md` → not re-proposed.
- `test_rule_scan_dry_run` — no harness invocation, lists candidates.
- `test_rule_review_finds_contradictions` — seeded `.rules.md` with contradiction → review written.
- `test_rule_review_state_tracking` — `.state.rule-review` updated.
- `test_install_rule_scan_schedule` — stub + plist created with correct labels.
- `test_install_rule_review_schedule` — stub + plist created with correct labels.
### 14.3 `test_dashboard_phase_roles.py`
- `test_api_phase_roles` — `/api/phase-roles` returns 6 roles with correct phases.
- `test_phase_roles_active_count` — tasks in phases → correct active_tasks count.
- `test_scheduled_jobs_new_kinds` — rule-scan and rule-review jobs appear with correct `kind`.
- `test_agent_type_meta_removed` — `AGENT_TYPE_META` no longer in dashboard.js (grep test).
## 15. Rollout (3 sequential tasks for model-divergence)
The model-divergence-enforcement parent task decomposes into 3 subtasks:
1. **manifest+detection**: `models.json` schema, `detect_models.py`, `install.sh`/`update.sh`/`upgrade.sh` integration, `config.md` section, onboarding Step 2d.
2. **interactive enforcement+audit**: `.state.models`, `--transition --model`, `--claim --model` conflict check, `--audit` model_divergence category, dashboard badges.
3. **loop enforcement+dashboard**: `loop.json` per-role model, `{model}` substitution, `--check-gate` model check, loop dashboard badges.
Rule agents (FW-2, FW-3) and Agent tab (FW-1) are backlog items, picked up after model-divergence ships (for FW-2/FW-3) or independently (for FW-1).
## 16. Locked Decision Index
All decisions referenced by `(Fn)` are in `README.md` § "Locked decisions". Implementation must conform. Deviations require a design doc update + `[unreleased]` CHANGELOG entry.