Protocol 12: The False-Alarm Ledger
Question
When a digital-phenotyping model emits more anomalies near relapse, what must be known before calling it clinically predictive—or using it as evidence about causation?
Benchmark located
A prospective study followed people with schizophrenia across Boston, Bangalore, and Bhopal for up to one year using mindLAMP. Inputs included passive geolocation, accelerometer and screen-state data, active surveys, and data-quality measures.
The accessible report says anomalies occurred 2.12× more often in the month before relapse and 2.78× more often around relapse than during non-relapse intervals. It also describes the passive-data anomaly model as outperforming a naive survey-only model.
The missing denominator
Those ratios do not reveal how often alarms were wrong. The accessible material did not supply:
- sample size;
- sensitivity or recall;
- specificity;
- positive predictive value;
- false alarms per person-month;
- threshold-selection procedure;
- calibration;
- prospective versus retrospective threshold locking;
- internal or external validation details.
Without these quantities, an elevated anomaly ratio can coexist with an unusably high false-alarm burden.
Required false-alarm ledger
For every proposed warning system, publish:
- Unit of prediction: person-day, person-week, or episode.
- Prediction horizon: exactly how far before an outcome an alarm counts.
- Outcome adjudication: who defines relapse and whether they are blinded to sensor data.
- Threshold provenance: fixed before testing or selected after seeing outcomes.
- Confusion matrix: true positives, false positives, true negatives, false negatives.
- Operational burden: false alarms per 100 person-weeks and median alarms per participant.
- Calibration: observed event frequency for each risk band.
- Missingness: whether data loss itself produces an anomaly.
- Baseline comparator: survey-only, last-observation, and random or prevalence-matched models.
- Site-stratified results: Boston, Bangalore, and Bhopal separately.
Transportability stress test
A multi-site sample is not automatically a transported model. Refit-free evaluation should hold out one site at a time. Performance should also be stratified by phone type, operating system, mobility constraints, local lockdown period, baseline illness severity, and data completeness.
A model fails transportability if its anomaly threshold must be retuned extensively at each site, if performance is driven by one site, or if missingness predicts the label because care pathways differ.
Causal firewall
Even a well-validated relapse predictor does not identify a cause. Screen-state change may be a consequence of insomnia, worsening symptoms, altered routine, hospitalization, medication changes, or missing data. Prediction answers what precedes an event reliably; causation asks what would change the event under intervention.
For the Apophenoth investigation, smartphone or chatbot-use signals must therefore be paired with temporal measurement of sleep, substances, stress, baseline vulnerability, and symptom onset. A chatbot-use spike that predicts relapse may still reflect reverse causation.
Evidence that would change the conclusion
I would upgrade a digital signal from interesting association to actionable predictor if a preregistered threshold achieved useful sensitivity with a tolerable, explicitly reported false-alarm rate in a temporally external or held-out-site test.
I would upgrade it toward causal evidence only if a discriminating intervention on the proposed exposure changed outcomes while rival pathways were measured and addressed.
Current verdict
The located study supports the feasibility of collecting temporally ordered passive signals across countries. The accessible evidence does not establish diagnostic accuracy, operational usefulness, transportability, or causation. The anomaly ratios are a lead, not a license.
Source trail
- Cohen et al., Relapse prediction in schizophrenia with smartphone digital phenotyping during COVID-19: a prospective three-site, two-country longitudinal study, npj Schizophrenia (2023).
- Accessible catalog preview: https://www.mendeley.com/catalogue/159e875f-3e1e-3df1-a7c4-bccd4bd655ed/
Direct publisher and repository retrieval were unavailable during this pass; missing metrics are recorded as unknown rather than inferred.
