RedHOTExpect
收藏资源简介:
RedHOTExpect是由斯图加特大学和哥本哈根IT大学联合创建的医疗领域社交媒体文本数据集,旨在研究患者对治疗的期望表达。该数据集包含约4,500条来自Reddit的医疗相关帖子,其中约2,500条带有治疗-期望-结果(TEO)三元组的细粒度标注。数据来源于Reddit健康讨论板块,通过大语言模型(LLM)进行自动筛选和银标注,并经过人工验证(标注准确率约78%)。数据集构建过程包括经验帖子筛选、期望标注和TEO三元组提取三个步骤。该数据集主要用于自然语言处理中的期望检测任务,可应用于医疗意见挖掘、治疗效果预测等领域,帮助医生识别患者未公开表达的治疗担忧。
RedHOTExpect is a medical social media text dataset jointly created by the University of Stuttgart and the IT University of Copenhagen, aiming to investigate how patients express their expectations toward medical treatments. This dataset contains approximately 4,500 medical-related posts sourced from Reddit, among which around 2,500 posts have fine-grained annotations of Treatment-Expectation-Outcome (TEO) triples. The data is collected from Reddit's health discussion subreddits, with automatic screening and silver labeling performed using Large Language Models (LLMs), followed by manual validation with an annotation accuracy of approximately 78%. The dataset construction process includes three steps: screening of experience-related posts, expectation annotation, and TEO triple extraction. This dataset is primarily used for expectation detection tasks in natural language processing (NLP), and can be applied in domains such as medical opinion mining and treatment effect prediction, helping clinicians identify unexpressed treatment concerns among patients.



