# ADR-0006: LLM data provenance — no training, no retention, local-first Date: 2026-06-27 Status: Accepted (design phase) Supersedes: none Related: ADR-0001 (local-first multi-tenant), ADR-0005 (Analyst Voice) ## Context Investor Flow sends sensitive user data to an LLM for summaries, explainers, alerts, and commentary: SEC filings, portfolio positions, trade theses, sentiment annotations, and the user's own journal entries. This is a legal and trust issue, not just a preference. Provider policies vary materially and can change with notice: - OpenAI / Anthropic / Google: default may include training eligibility unless explicitly disabled via API toggle. - OpenCode Go (paid Zen tier): Terms explicitly exclude paid-account Content from the "develop and improve our services" grant, so OpenCode itself is not training on prompts. THIS IS SAFE FOR THE OPERATOR'S OWN CODING USE. Residual gap: the underlying model vendor (e.g., Zhipu AI for GLM 5.2) may have its own unclear retention/transit policy — not addressed by OpenCode's Terms. - Hosted SaaS model providers generally: subject to upstream-vendor ambiguity + policy-drift. The app cannot rely on each external provider's current policy staying safe; it must enforce local-first in code so sensitive user data never leaves the operator's machine by construction. ## Decision — three hard constraints baked into the build ### Constraint 1 — Production LLM Gateway defaults to local OpenAI-compatible endpoint - The `LLMGateway` deep module (M14) must default its `baseURL` to a **local endpoint** (e.g., `http://localhost:11434/v1` for Ollama, or a self-hosted vLLM/serverless endpoint on the operator's LAN). - The local endpoint is the only provider configuration loaded by default from environment/secrets. - Non-local providers MAY be supported behind the provider-agnostic interface, but their activation requires: - An explicit operator override in a gitignored secrets file (`secrets/llm_providers.local.json`), never committed. - A documented data-provenance review per provider (recorded as an ADR or provider-config note: what that provider's policy is as of the review date, who reviewed, when). ### Constraint 2 — Sensitive user data never routed through external LLM providers - Sensitive user category (default): portfolio positions, trade plans, journal entries, SEC filings content, sentiment annotations, saved posts, any data tagged `ownerId` or sourced from Tier C / shared Tier A filings/threads. - The Gateway enforces a **data-classification gate** before any non-local provider call: - Local provider → any data allowed (nothing leaves the host). - Non-local provider → only data explicitly classified `public-safe` (e.g., generic financial term glosses, non-user-specific educational content) is eligible; sensitive user data is blocked at the Gateway with a typed error, not sent. - This is a code-level guarantee, not a runtime toggle — tests assert the gate refuses sensitive payloads against non-local providers. ### Constraint 3 — Developer conduct for the build itself - When building Investor Flow via external AI assistants (OpenCode Go + GLM 5.2, Claude Code, etc.), developers do NOT paste live user-data samples into prompts. - Use **fixtures and anonymized synthetic data** for any prompt that touches realistic shapes (filing text, portfolio rows, journal entries). The repo includes a `fixtures/` directory of synthetic, non-PII, freely-shareable sample data for this purpose. - This keeps the OpenCode Go paid-tier-safety (which holds today) from being the only safeguard; it removes the risk vector upstream of any policy. ## Consequences - The app's production LLM never sees user data leave the host → the entire upstream-vendor ambiguity + future-policy-drift question is eliminated by architecture, not by trust. - Operators who want best-quality hosted models for non-sensitive features (e.g., the "beginner explainer" on public market data) can configure them via explicit override; the data gate still refuses anything sensitive. - A fixture corpus must be maintained for realistic prompt testing; CI asserts the data-classification gate works against a fuzz set of sensitive vs payload-safe inputs. - LLMGateway's interface stays provider-agnostic (same as ADR-0005 contract); only the *default* + *data gate* are new. ## Implementation notes (for DESIGN.md Section 3) - `LLMGateway.classifyPayload(payload): 'public_safe' | 'sensitive'` runs before any provider dispatch. - `LLMGateway.dispatch(feature, payload, opts): Promise` routes: if classified sensitive AND provider is non-local → throw `SensitiveDataBlockedError`; never send. - Provider config schema: `{ id, baseURL, isLocal: boolean }`; `isLocal` trusts only `localhost`, `127.0.0.1`, `::1`, and entries in `LOCAL_LLM_SUBNETS` env override. - Tests: property-test the classifier over a fuzz corpus; contract-test `dispatch` rejects sensitive payloads on non-local providers.