4.9 KiB
4.9 KiB
ADR-0006: LLM data provenance — no training, no retention, local-first
Date: 2026-06-27 Status: Accepted (design phase) Supersedes: none Related: ADR-0001 (local-first multi-tenant), ADR-0005 (Analyst Voice)
Context
Investor Flow sends sensitive user data to an LLM for summaries, explainers, alerts, and commentary: SEC filings, portfolio positions, trade theses, sentiment annotations, and the user's own journal entries. This is a legal and trust issue, not just a preference.
Provider policies vary materially and can change with notice:
- OpenAI / Anthropic / Google: default may include training eligibility unless explicitly disabled via API toggle.
- OpenCode Go (paid Zen tier): Terms explicitly exclude paid-account Content from the "develop and improve our services" grant, so OpenCode itself is not training on prompts. THIS IS SAFE FOR THE OPERATOR'S OWN CODING USE. Residual gap: the underlying model vendor (e.g., Zhipu AI for GLM 5.2) may have its own unclear retention/transit policy — not addressed by OpenCode's Terms.
- Hosted SaaS model providers generally: subject to upstream-vendor ambiguity + policy-drift.
The app cannot rely on each external provider's current policy staying safe; it must enforce local-first in code so sensitive user data never leaves the operator's machine by construction.
Decision — three hard constraints baked into the build
Constraint 1 — Production LLM Gateway defaults to local OpenAI-compatible endpoint
- The
LLMGatewaydeep module (M14) must default itsbaseURLto a local endpoint (e.g.,http://localhost:11434/v1for Ollama, or a self-hosted vLLM/serverless endpoint on the operator's LAN). - The local endpoint is the only provider configuration loaded by default from environment/secrets.
- Non-local providers MAY be supported behind the provider-agnostic interface, but their activation requires:
- An explicit operator override in a gitignored secrets file (
secrets/llm_providers.local.json), never committed. - A documented data-provenance review per provider (recorded as an ADR or provider-config note: what that provider's policy is as of the review date, who reviewed, when).
- An explicit operator override in a gitignored secrets file (
Constraint 2 — Sensitive user data never routed through external LLM providers
- Sensitive user category (default): portfolio positions, trade plans, journal entries, SEC filings content, sentiment annotations, saved posts, any data tagged
ownerIdor sourced from Tier C / shared Tier A filings/threads. - The Gateway enforces a data-classification gate before any non-local provider call:
- Local provider → any data allowed (nothing leaves the host).
- Non-local provider → only data explicitly classified
public-safe(e.g., generic financial term glosses, non-user-specific educational content) is eligible; sensitive user data is blocked at the Gateway with a typed error, not sent.
- This is a code-level guarantee, not a runtime toggle — tests assert the gate refuses sensitive payloads against non-local providers.
Constraint 3 — Developer conduct for the build itself
- When building Investor Flow via external AI assistants (OpenCode Go + GLM 5.2, Claude Code, etc.), developers do NOT paste live user-data samples into prompts.
- Use fixtures and anonymized synthetic data for any prompt that touches realistic shapes (filing text, portfolio rows, journal entries). The repo includes a
fixtures/directory of synthetic, non-PII, freely-shareable sample data for this purpose. - This keeps the OpenCode Go paid-tier-safety (which holds today) from being the only safeguard; it removes the risk vector upstream of any policy.
Consequences
- The app's production LLM never sees user data leave the host → the entire upstream-vendor ambiguity + future-policy-drift question is eliminated by architecture, not by trust.
- Operators who want best-quality hosted models for non-sensitive features (e.g., the "beginner explainer" on public market data) can configure them via explicit override; the data gate still refuses anything sensitive.
- A fixture corpus must be maintained for realistic prompt testing; CI asserts the data-classification gate works against a fuzz set of sensitive vs payload-safe inputs.
- LLMGateway's interface stays provider-agnostic (same as ADR-0005 contract); only the default + data gate are new.
Implementation notes (for DESIGN.md Section 3)
LLMGateway.classifyPayload(payload): 'public_safe' | 'sensitive'runs before any provider dispatch.LLMGateway.dispatch(feature, payload, opts): Promise<Summary>routes: if classified sensitive AND provider is non-local → throwSensitiveDataBlockedError; never send.- Provider config schema:
{ id, baseURL, isLocal: boolean };isLocaltrusts onlylocalhost,127.0.0.1,::1, and entries inLOCAL_LLM_SUBNETSenv override. - Tests: property-test the classifier over a fuzz corpus; contract-test
dispatchrejects sensitive payloads on non-local providers.