If a chatbot trial randomly assigns features but people choose how much to use it, could that choice itself be masking who they already were?
Can voluntary use patterns in randomized trials reveal confounding by user characteristics?
This explores whether, in a randomized trial where people choose how much to use a tool, the link between heavier use and worse outcomes might come from who those users already were rather than from the tool, and whether that self-chosen usage can help expose it.
This explores whether self-chosen usage inside a randomized trial can reveal that people's existing traits, not the tool, are behind the outcomes. The clearest case in the corpus is a randomized trial of 981 chatbot users Does how much time people spend with chatbots drive worse outcomes?. The researchers randomly assigned the parts they controlled: voice versus text, and personal versus open-ended conversations. None of those changed loneliness, socialization or dependence. What did predict worse outcomes was how much time each person chose to spend with the chatbot, and that held in every condition. This is the telltale pattern. Randomization protects the comparison between assigned conditions, but it does nothing for usage that participants pick for themselves. Once the analysis shifts from "which condition were you in" to "how much did you use it," you are back to an observational study inside a trial.
That is why voluntary use can reveal possible confounding without proving it. If heavy use tracks bad outcomes the same way in every randomized arm, the design features are probably not the driver. The likely suspects become traits people brought with them, such as loneliness that came before the chatbot and pushed them toward it. The trial can't tell you whether lonely people seek out chatbots or chatbots deepen loneliness. It can tell you that the thing you randomized wasn't what mattered, and that points you to the self-selection. Answering the causal question would take a further step, such as randomly assigning usage amounts or measuring baseline traits before people start.
One note in the corpus pushes back on treating different usage rates as noise to be removed. In a study of AI writing suggestions, Indian writers accepted more AI suggestions than American writers Is higher AI use by Indian writers a confound to control?. The authors argue this gap is not a confound to control away. It reflects cultural patterns of trust and technology adoption, and it is part of how AI makes writing more uniform. The lesson carries over: when user traits shape how much people engage, statistically "controlling for" those traits can erase the very effect you're trying to study. Whether a trait is a confound or the mechanism depends on your question.
A related trap concerns what a trial's control group actually isolates. Therapy chatbots tested against waitlists or reading materials can look effective just because users get conversational contact, and ELIZA, the 1960s chatbot, has matched Woebot under those conditions Do chatbot trials against waitlists measure real therapeutic value?. Both cases show that randomization only answers the question its design actually asks. Engagement, contact and self-selection can leak in through every gap.
The corpus has only one trial that directly shows this pattern, so treat this as a strong illustration rather than a settled method. Still, the transferable insight is that a null result for the randomized features combined with a strong dose-response for chosen usage is itself a diagnostic signal. It tells you to look at the people before you blame the product.
Sources 3 notes
In a randomized trial of 981 people, randomly assigned interaction modes and conversation types showed no effect on loneliness, socialization, or dependence. Instead, participants who chose to use the chatbot more showed consistently worse outcomes across all measures, regardless of condition.
Indian writers accepted more AI suggestions than American writers, reflecting cultural differences in trust and collectivist technology adoption patterns. The authors argue this reliance difference is integral to understanding homogenization, not a confound that obscures it.
Comparing therapeutic chatbots to waitlist or psychoeducation controls creates false efficacy claims by measuring conversational contact rather than therapy-specific mechanisms. ELIZA matching Woebot performance demonstrates this; real evidence requires comparative trials against existing treatments and mechanism identification.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study
- Investigating Affective Use and Emotional Well-being on ChatGPT
- AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances
- Can robots do therapy?: Examining the efficacy of a CBT bot in comparison with other behavioral intervention technologies in alleviating mental health symptoms
- AI Companions Reduce Loneliness
- Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports
- DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
- Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being