PolyglotFakeFacts: A multilingual dataset of fake and real news across politics, security, and social domains
收藏资源简介:
PolyglotFakeFacts is a multilingual dataset designed to support research on the detection of fake and real news across diverse domains such as politics, geopolitics, security, social issues, and military affairs. The research hypothesis underpinning this dataset is that linguistic and contextual markers of misinformation can be systematically identified across multiple languages, enabling the development of more robust and generalizable fake news detection models. The dataset shows a balanced collection of human-labeled fake news articles alongside verified real news extracted from trusted media outlets. Notably, the data covers multiple languages and different thematic areas, which allows researchers to explore how misinformation manifests in diverse cultural and geopolitical contexts. Among the key findings is that fake news articles often display recurring linguistic and structural patterns regardless of the language, while real news tends to follow more standardized journalistic conventions. This suggests that multilingual approaches to fake news detection could leverage both cross-linguistic similarities and domain-specific features. The data was gathered through a combination of manual annotation by human experts for fake news samples and curation of real news from reliable sources. All samples were pre-processed to ensure consistent formatting, removal of duplicates, and inclusion of metadata such as language, domain, and label (fake/real). This dataset can be interpreted and used by researchers aiming to: - train and evaluate machine learning and deep learning models for fake news classification, - perform cross-lingual and multilingual comparative studies, - investigate the linguistic, semantic, and thematic characteristics of misinformation. By providing a curated, multilingual, and domain-diverse resource, PolyglotFakeFacts enables the community to develop more transparent, explainable, and resilient AI models for combating online misinformation.
PolyglotFakeFacts是一款多语言数据集,旨在支持针对政治、地缘政治、安全、社会议题及军事事务等多元领域内虚假与真实新闻的检测研究。 作为该数据集核心依托的研究假说认为:错误信息的语言与语境标记可在多语言场景下被系统性识别,从而助力开发更具鲁棒性与泛化性的虚假新闻检测模型。 该数据集包含比例均衡的人工标注虚假新闻样本,以及从可信媒体渠道提取的经核实真实新闻。值得注意的是,其覆盖多语言与多主题领域,可支持研究者探索错误信息在不同文化及地缘政治语境下的表现形式。 关键研究发现之一为:无论语言类型如何,虚假新闻往往呈现出重复出现的语言与结构模式;而真实新闻则更倾向于遵循标准化的新闻写作规范。这表明,虚假新闻检测的多语言方法可同时利用跨语言相似性与领域专属特征。 该数据集的采集流程为:针对虚假新闻样本开展专家人工标注,并从可靠来源筛选整理真实新闻。所有样本均经过预处理,以确保格式统一、去除重复样本,并附带语言、领域及标签(虚假/真实)等元数据。 本数据集可供以下研究方向的研究者使用与开展相关研究: - 训练并评估用于虚假新闻分类的机器学习与深度学习模型 - 开展跨语言及多语言对比研究 - 探究错误信息的语言、语义及主题特征。 通过提供经过精心筛选的多语言、跨领域多样化资源,PolyglotFakeFacts可助力学界开发更具透明度、可解释性与韧性的人工智能模型,以对抗在线错误信息的传播。




