心理学
心理健康
聊天机器人
心理治疗师
计算机科学
应用心理学
自然语言处理
作者
Florian Onur Kuhlmeier,Leon Hanschmann,Melina Rabe,Stefan Luettke,Eva‐Lotta Brakemeier,Alexander Maedche
标识
DOI:10.48550/arxiv.2503.21540
摘要
LLMs promise to overcome limitations of rule-based mental health chatbots through improved natural language capabilities, yet their ability to deliver evidence-based psychological interventions remains largely unverified because evaluations rarely apply the validated fidelity measures used to assess psychotherapists. We developed an LLM-based chatbot that delivers behavioral activation for depression and generated 48 complete chat sessions with diverse artificial users. Ten psychotherapists assessed these sessions using the Quality of Behavioral Activation Scale (Q-BAS), a validated fidelity instrument. Results show that the chatbot reliably executed the intervention across all phases and maintained safety protocols, but it struggled with clinical judgment, particularly when verifying the feasibility of proposed activities. Overall, our findings suggest that LLM-based chatbots can execute therapeutic protocols with high fidelity, while robust clinical reasoning remains an open challenge. We outline design implications to address this gap and provide the chatbot and artificial user prompts.
科研通智能强力驱动
Strongly Powered by AbleSci AI