Multilingual consumer-chatbot health-advice evaluation across six languages: physician ratings, prompts, and reproduction kit
收藏资源简介:
Public release of the dataset and reproduction kit for a six-language, physician-rated evaluation of consumer health chatbots. Contents: 504 chatbot responses (21 forum-derived clinical scenarios, six languages [English, Hebrew, French, Russian, Arabic, Thai], and four widely deployed consumer chatbots: ChatGPT, Claude, Gemini, DeepSeek); 1,008 native-language clinician ratings on five Likert dimensions (5,040 dimension-level scores); the paired-English-with-context controls arm (200 responses); the cross-lingual anchoring replication arm (240 responses); the triplicate-generation (N=3) reproducibility subset; the LLM-judge cross-validation pilot; language-property predictors (URIEL typological distance, tokenization fertility, Joshi resource tier); and a self-contained Python reproduction kit (scripts/) that regenerates every reported number and statistical test from the released CSVs. Licensing: code (scripts/) under MIT; data (data/) under CC-BY 4.0. Privacy: source-forum types and scenario provenance categories are described in the accompanying supplementary appendix; specific forum names, domains, URLs, usernames, post dates, and verbatim source-post text are not released.



