Files
investor-flow/docs/adr/0006-llm-data-provenance.md

4.9 KiB

ADR-0006: LLM data provenance — no training, no retention, local-first

Date: 2026-06-27 Status: Accepted (design phase) Supersedes: none Related: ADR-0001 (local-first multi-tenant), ADR-0005 (Analyst Voice)

Context

Investor Flow sends sensitive user data to an LLM for summaries, explainers, alerts, and commentary: SEC filings, portfolio positions, trade theses, sentiment annotations, and the user's own journal entries. This is a legal and trust issue, not just a preference.

Provider policies vary materially and can change with notice:

  • OpenAI / Anthropic / Google: default may include training eligibility unless explicitly disabled via API toggle.
  • OpenCode Go (paid Zen tier): Terms explicitly exclude paid-account Content from the "develop and improve our services" grant, so OpenCode itself is not training on prompts. THIS IS SAFE FOR THE OPERATOR'S OWN CODING USE. Residual gap: the underlying model vendor (e.g., Zhipu AI for GLM 5.2) may have its own unclear retention/transit policy — not addressed by OpenCode's Terms.
  • Hosted SaaS model providers generally: subject to upstream-vendor ambiguity + policy-drift.

The app cannot rely on each external provider's current policy staying safe; it must enforce local-first in code so sensitive user data never leaves the operator's machine by construction.

Decision — three hard constraints baked into the build

Constraint 1 — Production LLM Gateway defaults to local OpenAI-compatible endpoint

  • The LLMGateway deep module (M14) must default its baseURL to a local endpoint (e.g., http://localhost:11434/v1 for Ollama, or a self-hosted vLLM/serverless endpoint on the operator's LAN).
  • The local endpoint is the only provider configuration loaded by default from environment/secrets.
  • Non-local providers MAY be supported behind the provider-agnostic interface, but their activation requires:
    • An explicit operator override in a gitignored secrets file (secrets/llm_providers.local.json), never committed.
    • A documented data-provenance review per provider (recorded as an ADR or provider-config note: what that provider's policy is as of the review date, who reviewed, when).

Constraint 2 — Sensitive user data never routed through external LLM providers

  • Sensitive user category (default): portfolio positions, trade plans, journal entries, SEC filings content, sentiment annotations, saved posts, any data tagged ownerId or sourced from Tier C / shared Tier A filings/threads.
  • The Gateway enforces a data-classification gate before any non-local provider call:
    • Local provider → any data allowed (nothing leaves the host).
    • Non-local provider → only data explicitly classified public-safe (e.g., generic financial term glosses, non-user-specific educational content) is eligible; sensitive user data is blocked at the Gateway with a typed error, not sent.
  • This is a code-level guarantee, not a runtime toggle — tests assert the gate refuses sensitive payloads against non-local providers.

Constraint 3 — Developer conduct for the build itself

  • When building Investor Flow via external AI assistants (OpenCode Go + GLM 5.2, Claude Code, etc.), developers do NOT paste live user-data samples into prompts.
  • Use fixtures and anonymized synthetic data for any prompt that touches realistic shapes (filing text, portfolio rows, journal entries). The repo includes a fixtures/ directory of synthetic, non-PII, freely-shareable sample data for this purpose.
  • This keeps the OpenCode Go paid-tier-safety (which holds today) from being the only safeguard; it removes the risk vector upstream of any policy.

Consequences

  • The app's production LLM never sees user data leave the host → the entire upstream-vendor ambiguity + future-policy-drift question is eliminated by architecture, not by trust.
  • Operators who want best-quality hosted models for non-sensitive features (e.g., the "beginner explainer" on public market data) can configure them via explicit override; the data gate still refuses anything sensitive.
  • A fixture corpus must be maintained for realistic prompt testing; CI asserts the data-classification gate works against a fuzz set of sensitive vs payload-safe inputs.
  • LLMGateway's interface stays provider-agnostic (same as ADR-0005 contract); only the default + data gate are new.

Implementation notes (for DESIGN.md Section 3)

  • LLMGateway.classifyPayload(payload): 'public_safe' | 'sensitive' runs before any provider dispatch.
  • LLMGateway.dispatch(feature, payload, opts): Promise<Summary> routes: if classified sensitive AND provider is non-local → throw SensitiveDataBlockedError; never send.
  • Provider config schema: { id, baseURL, isLocal: boolean }; isLocal trusts only localhost, 127.0.0.1, ::1, and entries in LOCAL_LLM_SUBNETS env override.
  • Tests: property-test the classifier over a fuzz corpus; contract-test dispatch rejects sensitive payloads on non-local providers.