Shadow-Self Protocol v0.1
A design for getting frontier-model advice about my body and my travel, without ever handing anyone my body or my travel.
Status: spec, rev 0.1, 2026-10-04. Next: a runnable demo, then a bounty on anyone who finds a leak. This is engineering, not marketing — the failure modes section is the point.
0. The problem, one sentence
Personalised diet + exercise advice needs my numbers (weight, resting HR, sleep, location traces, calendar); privacy needs nobody's numbers. The protocol ships the minimum slice of "mine" that still changes the plan.
1. Threat model
Adversaries assumed:
- The model provider — sees prompts, keeps logs, may train on them. Curious, untrusted; its promises are not the guarantee.
- A linkage attacker with public data (posts, photos, social check-ins) trying to match a shadow record to a person.
- A future breach of the provider's prompt store — everything sent must be worthless in isolation.
Never leaves the device: name, wallet, email, device IDs, exact coordinates, timestamps finer than a date, raw weight/HR values, medication names, ANY free text (notes, venue or street names).
May leave the device: noisy features inside a published schema, under per-field and per-day privacy budgets.
2. Core trick: the model advises a shadow, not a person
- The device builds a shadow self: a synthetic profile statistically close enough to mine to pick the right plan family, produced by a calibrated noise mechanism so I cannot be singled out of it.
- The frontier model runs on the shadow only and returns population-level output: plan template, progression rules, red-flag checklist.
- Final personalisation — actual portions, actual schedule, deload days — is computed locally from the real numbers that never left.
The load-bearing decision: the cloud picks the family of plans; the device picks the instance. The information that would identify me is exactly the information the LLM does not need.
3. Pipeline (stages 1–6 on-device, only 7 sees the cloud)
- Collect — local store of health + travel records. Nothing leaves.
- Normalise — units, timezones stripped, durations quantised to 15-min bins.
- Generalise quasi-identifiers — age → 5-year band; height → 3-cm band; weight → 2-kg band; location → 25-km grid cell named as a climate cluster ("EU-central-temperate"); dates → weekday + ISO week; calendar → a 1–5 "trip-load" integer.
- k-anonymity gate — if my generalised record matches < k (k=10) simulated peers, bands widen automatically; if still unique, refuse to send. Rare bodies and rare trips are the #1 re-identification risk; this gate is the load-bearing wall.
- DP noise — numerics get calibrated noise: Laplace for integer-valued features, Gaussian with Rényi-DP accounting for continuous ones, at the ε table below. Noise is added once, before anything is cached.
- Schema clamp — serialise into a fixed JSON schema, every field always present (missing = noisy sentinel), field order fixed, free-text banned and validated by allowlist regex. "Notes"-style fields are where anonymisers leak.
- Frontier call — shadow record + a constant template prompt. Zero context beyond the schema.
- Local fit — small on-device model takes plan family + real numbers → the weekly plan I actually follow. This stage is the "self" half of the protocol.
##
