FMNV
收藏资源简介:
FMNV数据集是一个由新闻媒体发布的新闻视频组成的创新数据集,旨在用于假新闻检测。该数据集包含了由27家新闻媒体在Twitter和YouTube上发布的2,393个新闻视频,涵盖了事故、疫情、政治等12个主题,时间跨度五年。数据集通过大规模语言模型进行数据增强,以解决真实与虚假新闻视频的数据不平衡问题,并包含了标题、视频片段和音频三种模态的信息。该数据集的构建旨在推动媒体生态系统中高影响力假新闻检测的研究,并促进了跨模态不一致性分析方法的进展。
The FMNV Dataset is an innovative corpus of news videos released by news media, specifically designed for fake news detection. This dataset contains 2,393 news videos published on Twitter and YouTube by 27 news outlets, covering 12 topics including accidents, pandemics, politics and other categories, with a five-year time span. Data augmentation via large language models (LLMs) has been applied to the dataset to address the data imbalance issue between real and fake news videos, and it includes three modalities: headlines, video clips and audio. The construction of this dataset aims to advance research on high-impact fake news detection in media ecosystems, and promote the development of cross-modal inconsistency analysis methods.




