CARE v1.0 (Conversational Audio-visual Recordings of health Experiences)
收藏资源简介:
CARE v1.0是由瓦伦西亚理工大学等机构联合构建的多模态健康交流数据集,旨在通过视频访谈记录支持语音和非语言行为的计算分析。该数据集包含4281个短视频片段,总计约144小时,涵盖612名参与者,涉及哮喘、慢性疼痛、抑郁症等12种医疗状况及对照组,数据源自HEXI平台的公开访谈档案,并提供了语音、面部活动、凝视模式和身体运动的预计算多模态描述符。数据集的创建过程包括从HEXI平台筛选相关医疗条件和参与者,并利用大型语言模型自动提取结构化元数据,如药物使用和情感表达。该数据集主要应用于自动疾病检测、症状监测、情感语境下的多模态建模,以及疾病轨迹和应对过程的计算研究,致力于推动数字生物标志物在神经、精神和呼吸系统疾病中的创新应用。
CARE v1.0 is a multimodal health communication dataset jointly constructed by the Universitat Politècnica de València and other institutions, aiming to support computational analysis of speech and non-verbal behaviors through video interview recordings. This dataset includes 4,281 short video clips, totaling approximately 144 hours, covering 612 participants, and involves 12 medical conditions such as asthma, chronic pain, and depression, as well as a control group. The data is sourced from public interview archives on the HEXI platform, and provides pre-computed multimodal descriptors for speech, facial activities, gaze patterns, and body movements. The dataset creation process involves screening relevant medical conditions and participants from the HEXI platform, and automatically extracting structured metadata such as medication use and emotional expressions using Large Language Models (LLMs). This dataset is primarily applied in automatic disease detection, symptom monitoring, multimodal modeling in emotional contexts, as well as computational studies of disease trajectories and coping processes, and is dedicated to promoting innovative applications of digital biomarkers in neurological, psychiatric, and respiratory diseases.





