Protocol 09: Negative Controls for AI-Crisis Claims
Question
When chatbot use and psychiatric crisis co-occur, what observations would reveal confounding, reverse causation, or measurement bias rather than a chatbot-specific effect?
Why negative controls matter
A vivid transcript can make any nearby event feel causally central. Negative controls test whether the same analytical machinery produces an “effect” where no plausible chatbot-specific pathway exists. If it does, the design is detecting bias rather than causation.
Control 1 — Future exposure
Test whether chatbot use after a symptom measurement predicts that earlier symptom measurement.
- Expected causal result: no association after adjustment.
- Failure signal: a strong “effect” of future use on past symptoms indicates reverse causation, time-varying confounding, timestamp error, or model misspecification.
Control 2 — Non-conversational screen time
Compare conversational AI exposure with similarly timed non-conversational digital activity: search, video, games, or ordinary app use.
- If only chatbot exposure predicts subsequent symptom change, specificity gains weight.
- If all nocturnal screen use behaves similarly, sleep displacement or general digital overuse is the cleaner explanation.
This control must match time of day and duration; otherwise it merely relabels the confounder.
Control 3 — Benign conversational content
Separate potentially reinforcing exchanges from mundane sessions such as formatting, coding, or factual lookup.
- A content-specific amplification hypothesis predicts greater risk after affirming, personalized, certainty-laden exchanges.
- An exposure-only association with no content gradient points toward user state, duration, or selection effects.
Content should be coded blind to later outcomes.
Control 4 — Symptoms unlikely to move immediately
Include an outcome that should not plausibly change within the hypothesized short exposure window.
- A same-day association across every measured symptom domain suggests common-method bias or global distress reporting.
- A temporally and clinically specific pattern is harder to dismiss.
The control outcome must be chosen with clinicians before analysis, not after inspecting results.
Control 5 — Pre-exposure trend
Estimate symptom trajectories for several windows before exposure spikes.
- Flat pre-trends followed by a post-exposure change support temporal specificity.
- Rising symptoms before increased use support reverse causation: distress may drive chatbot engagement.
One pre-exposure measurement is not a trend.
Control 6 — Alternative trigger windows
Run the identical model using implausible lag windows selected in advance.
- A hypothesized effect confined to a clinically plausible window is informative.
- Similar coefficients across arbitrary windows imply unstable timing or persistent confounding.
Do not search dozens of windows and report only the strongest.
Control 7 — Transcript salience audit
Have adjudicators score chronology and rival causes before seeing chatbot transcript content, then score again after reveal.
Record:
- change in causal rating;
- evidence cited for the change;
- whether new chronological information was introduced;
- whether vivid language alone moved the judgment.
A large rating shift without new temporal evidence measures narrative salience, not chatbot causation.
Minimum variables
- timestamped chatbot minutes and messages;
- session time of day;
- blinded content features;
- repeated symptom measures;
- sleep duration and timing;
- substance use;
- acute stressors;
- medication changes;
- baseline vulnerability;
- non-chatbot digital activity.
Interpretation table
| Pattern | Strongest interpretation | |---|---| | Symptoms rise before use; future exposure also “predicts” symptoms | Reverse causation or model failure | | Chatbot and matched nocturnal screen time show similar associations | Sleep displacement/general digital exposure | | Only affirming or certainty-laden exchanges precede change | Content-specific amplification becomes more plausible | | Every outcome and every lag shows an effect | Common-method bias or residual confounding | | Specific post-exposure change, flat pre-trend, null negative controls | Chatbot-specific causal weight increases |
What would change the provisional conclusion
The current narrow conclusion is that measurable agreement and reliance mechanisms exist, crisis amplification is plausible, and independent ignition is not established. That conclusion should strengthen toward causation only if prospective data show:
- exposure preceding symptom change;
- flat pre-exposure symptom trends;
- a dose or content gradient;
- null future-exposure and matched-activity controls;
- robustness to sleep, substances, stress, treatment, and baseline vulnerability;
- independent clinical adjudication.
If negative controls light up alongside the main exposure, the apparent signal should be downgraded.
Status
This is a preregistration scaffold, not evidence that any effect exists. Its purpose is to make the thesis capable of failure before the next dramatic case arrives.
