Protocol 11: Predictive Signal Is Not Causal Proof
The anomaly
A prospective smartphone study of schizophrenia relapse reported that passive-sensor anomalies appeared 2.12× more often in the month before relapse and 2.78× around relapse than during non-relapse intervals. That is a useful demonstration that weak, time-stamped signals can precede a crisis.
It does not establish what caused the crisis. It also does not show that an anomaly alert is clinically reliable.
Four claims that must remain separate
- Association: sensor anomalies occur near relapse.
- Temporal prediction: anomalies precede relapse often enough to forecast it.
- Clinical utility: acting on alerts improves outcomes at acceptable cost.
- Causation: the measured exposure or behavior contributes to relapse.
A result at level 1 cannot be promoted to levels 2–4 by rhetoric.
Threshold-performance checklist
Before calling a signal “predictive,” require:
- sensitivity at a prespecified threshold;
- specificity at that threshold;
- positive and negative predictive values;
- false alerts per person-month;
- calibration across risk levels;
- lead-time distribution, not merely an average window;
- comparison with a simple clinical baseline;
- out-of-sample or external validation;
- handling of missing data and phone non-use;
- confidence intervals around every operational metric.
The accessible report supplied an anomaly-frequency ratio but not sensitivity, specificity, or false-positive burden. The ratio is therefore a promising temporal marker, not a deployable detector.
Transportability checklist
A signal trained in one setting may encode the setting rather than the phenomenon. Test performance across:
- geography and health systems;
- phone hardware and operating systems;
- language and cultural routines;
- age, socioeconomic status, and housing stability;
- baseline diagnosis and illness severity;
- medication changes, sleep disruption, substance use, and acute stress;
- pandemic versus non-pandemic behavior;
- people who engage heavily with phones versus those who do not.
Application to chatbot-linked crisis claims
A rigorous prospective design should timestamp:
- chatbot minutes, messages, and nocturnal use;
- conversational features such as agreement or escalating certainty;
- sleep duration and timing;
- substance exposure;
- acute stressors;
- symptoms and functional change;
- baseline vulnerability and treatment changes.
Then compare competing temporal paths:
- chatbot changes → later symptom change;
- symptom change → later chatbot use;
- insomnia/substances/stress → both;
- a shared background trend → both.
Negative controls
A chatbot-specific interpretation should weaken if the same model predicts relapse equally well from:
- unrelated late-night phone activity;
- total screen time without chatbot exposure;
- future exposure values that cannot cause earlier symptoms;
- shuffled timestamps that destroy the proposed sequence.
Evidence that would change the current conclusion
The independent-ignition hypothesis would gain weight if preregistered, within-person analyses repeatedly showed chatbot exposure changes preceding symptoms; the effect survived adjustment for sleep, substances, stress, baseline vulnerability, and treatment; content-specific exposure outperformed generic screen activity; and the result replicated externally with an acceptable false-alert burden.
It would lose weight if symptoms reliably preceded exposure spikes, generic nocturnal phone use predicted equally well, rival causes absorbed the association, or the signal failed outside the original sample.
Provisional verdict
Digital phenotyping shows that temporal weak signals can be measured prospectively. It does not license a causal story. The next serious study must publish the false alarms, preserve temporal order, and let rival explanations compete on equal terms.
Open questions
- What false-alert rate would patients and clinicians tolerate?
- Does conversational content add signal beyond duration and sleep loss?
- Can an intervention test distinguish amplification from mere prediction?
- Which measurements remain stable across cultures and devices?
