CareChat 2022: Response-level ratings from a randomised field comparison of retrieval-based and generative counselling chatbots
收藏资源简介:
De-identified rating data from a randomised field study of a Korean psychological-counselling chatbot service. 160 users reporting at least moderate depression (PHQ-9 >= 10) were randomly assigned to a retrieval-based or a generative chatbot. Over ten days in December 2022 each participant entered three concerns per nightly session and rated every response on seven items, then rated the session as a whole. The deposit contains 2,919 response-level ratings from 973 sessions by 160 participants, session-level ratings, and the four psychological measures administered before and after participation, together with the scripts that reproduce the published tables. Participants' concern texts, chatbot response texts and free-text reasons are not included, because participants reported clinical-level depressive symptoms. Three properties obtainable only from the removed response texts are retained asderived columns, so that every analysis in the article remains reproducible: response length in characters, presence of a training-corpus marker, and presence of an execution-error string. See README.md and CODEBOOK.md in the archive for full documentation. License `Creative Commons Attribution 4.0 International` Keywords conversational agent evaluation, counselling chatbot, mental health chatbot, user evaluation, politeness theory, randomised field study, Korean



