Urdu-News-Augmented-Dataset
收藏资源简介:
该数据集包含900篇原始乌尔都语新闻文章,标注为真实或虚假。此外,还包括通过Google翻译系统从英语翻译到乌尔都语的400篇新闻文章,用于增强数据集,以及这些数据集的多种组合,用于探索增强效果。
This dataset comprises 900 original Urdu news articles, annotated as either genuine or fake. Additionally, it includes 400 news articles translated from English to Urdu via the Google Translate system, aimed at augmenting the dataset. Various combinations of these datasets are also provided to explore the effects of augmentation.
数据集概述
数据集名称
Annotated Fake News Dataset in Urdu and Augmentation using Machine Translation
发布日期
March 03, 2020
主要贡献者
- Maaz Amjad
- Grigori Sidorov
- Alisa Zhila
数据集内容
- 原始数据集包含900篇乌尔都语新闻文章,标注为真实或虚假。
- 增广数据集包含400篇新闻文章,通过Google Translate从英语翻译至乌尔都语。
- 提供多种数据集组合,用于探索增广效果。
相关研究
数据集伴随论文《Data Augmentation using Machine Translation for Fake News Detection in the Urdu Language (2020)》,该论文已被LREC 2020接受。
引用信息
若使用此数据集进行出版物,请引用以下文献:
@article{Maazaug2020, author = {Maaz Amjad, Grigori Sidorov, Alisa Zhila}, title = {Annotated Fake News Dataset in Urdu and Augmentation using Machine Translation}, conference = {Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020)}, page = {2530–2535} year = {2020} }
联系方式
如有进一步问题或咨询,请联系Maaz Amjad (maazamjad@phystech.edu)。




