遇见数据集

Curated Afaan Oromo Rumor Detection Dataset for Cross-Platform and Cross-Lingual Rumor Detection

收藏
Zenodo2026-06-21 更新2026-05-29 收录
官方服务:

资源简介:

The dataset includes 9,680 records, consisting of 4,956 Rumor records and 4,724 Non-Rumor records. The dataset is divided into training, testing, and validation partitions with a 70:20:10 split. Three annotator-label columns and a final majority-vote label are included to support annotation transparency and reproducibility. Personally identifiable information, usernames, profile links, phone numbers, email addresses, and unnecessary raw metadata have been removed before release.

本数据集为经精选整理的奥罗莫语(Afaan Oromo)谣言检测数据集,由题为《GRACE-MRD:面向低资源场景的跨平台与跨语言谣言检测通用框架》(GRACE-MRD: A Generalizable Framework for Cross-Platform and Cross-Lingual Rumor Detection in Low-Resource Settings)的研究生成并分析。本数据集包含经清洗与匿名化处理的奥罗莫语文本记录,标注类别分为谣言与非谣言两类,旨在为谣言检测、虚假信息检测、低资源语言处理、跨平台学习以及跨语言模型评估相关研究提供支撑。 本数据集总计包含9680条记录,其中谣言记录4956条,非谣言记录4724条。数据集按照70:20:10的比例划分为训练集、测试集与验证集三个子集。数据集设置了三条标注者独立标注列与一条最终多数投票整合标注列,以保障标注流程的透明度与研究结果的可复现性。在正式发布前,本数据集已移除所有个人可识别信息、用户名、个人主页链接、电话号码、电子邮箱地址以及冗余的原始元数据。

提供机构:
Zenodo
创建时间:
2026-05-28
二维码
社区交流群
二维码
科研交流群
商业服务