遇见数据集

M4Health: A Multi-Modal, Multi-Domain, Multi-Platform, and Multi-Task Benchmark for Video-Driven Health Communication on Social Media

收藏
Zenodo2026-04-16 更新2026-05-26 收录
官方服务:

资源简介:

Short-video platforms have transformed how the public consumes health information on web and social media where uncurated content poses significant risks to vulnerable populations. While prior work has primarily focused on text-only or single-platform analyses, comprehensive benchmarks for multi-modal health communication in short videos remain limited. In this paper, we introduce M4Health, a multi-modal, multi-domain, multi-platform, and multi-task benchmark for health communication in short videos. M4Health comprises $669,995$ videos from TikTok, YouTube Shorts, and Reddit, spanning diverse health domains such as nutrition, fitness, mental health, and wellness. We provide expert annotations for a subset of videos across three interrelated tasks, including credibility assessment, AI-generation detection, and theme classification. Extensive benchmarking experiments show that current state-of-the-art models, including task-specific approaches and large vision-language models (LVLMs), achieve suboptimal performance. We will share the M4Health dataset with research communities to foster collaborative research toward supporting informed health decision-making on web and social media.

短视频平台彻底改变了公众在网络与社交媒体上获取健康信息的方式,而未经审核的内容会对脆弱人群构成显著风险。尽管此前的研究多聚焦于纯文本分析或单平台研究,但针对短视频场景下多模态健康传播的综合基准数据集仍较为匮乏。本文推出M4Health,一款面向短视频健康传播的多模态、多领域、多平台及多任务基准数据集。该数据集包含来自TikTok、YouTube Shorts及Reddit的669,995条短视频,覆盖营养、健身、心理健康与健康养生等多元健康领域。我们针对部分短视频样本提供了专家标注,涵盖三项相互关联的任务:可信度评估、AI生成内容检测与主题分类。多项大规模基准测试实验显示,包括任务专属方法与大视觉语言模型(Large Vision-Language Models, LVLMs)在内的当前前沿顶尖模型,在该数据集上的性能仍未达最优水平。我们将向科研社区公开M4Health数据集,以期推动相关协作研究,助力用户在网络与社交媒体上做出理性的健康决策。

提供机构:
Zenodo
创建时间:
2026-01-16
二维码
社区交流群
二维码
科研交流群
商业服务