Potrika: Raw and Balanced Newspaper Datasets in the Bangla Language with Eight Topics and Five Attributes
收藏资源简介:
Knowledge is central to human and scientific developments. Natural Language Processing (NLP) allows automated analysis and creation of knowledge. Data is a crucial NLP and machine learning ingredient. The scarcity of open datasets is a well-known problem in the machine and deep learning research. This is very much the case for textual NLP datasets in English and other major world languages. For the Bangla language, the situation is even more challenging and the number of large datasets for NLP research is practically nil. We hereby present Potrika, a large single-label Bangla news article textual dataset curated for NLP research from six popular online news portals in Bangladesh (Jugantor, Jaijaidin, Ittefaq, Kaler Kontho, Inqilab, and Somoyer Alo) for the period 2014-2020. The articles are classified into eight distinct categories (National, Sports, International, Entertainment, Economy, Education, Politics, and Science & Technology) providing five attributes (News Article, Category, Headline, Publication Date, and Newspaper Source). The raw dataset contains 185.51 million words and 12.57 million sentences contained in 664,880 news articles. Moreover, using NLP augmentation techniques, we create from the raw (unbalanced) dataset another (balanced) dataset comprising 320,000 news articles with 40,000 articles in each of the eight news categories. Potrika contains both datasets (raw and balanced) to suit a wide range of NLP research. By far, to the best of our knowledge, Potrika is the largest and the most extensive dataset for news classification. Further details of the dataset, its collection, and usage can be found in our article here: https://doi.org/10.48550/arXiv.2210.09389.
知识是人类与科学发展的核心所在。自然语言处理(Natural Language Processing,以下简称NLP)可实现知识的自动化分析与构建。数据是NLP与机器学习的核心要素。开放数据集的稀缺性是机器学习与深度学习研究中公认的难题,英语及其他主流语种的文本NLP数据集亦存在这一问题。而孟加拉语(Bangla)领域的境况则更为严峻,适用于NLP研究的大型数据集数量几乎为零。 本文谨介绍Potrika:这是一个面向NLP研究的大规模单标签孟加拉语文本新闻数据集,采集自孟加拉国6家知名在线新闻门户(Jugantor、Jaijaidin、Ittefaq、Kaler Kontho、Inqilab及Somoyer Alo),数据覆盖时段为2014年至2020年。该数据集的新闻文本共分为8个独立类别(国内(National)、体育(Sports)、国际(International)、娱乐(Entertainment)、经济(Economy)、教育(Education)、政治(Politics)、科学与技术(Science & Technology)),并涵盖5项属性字段:新闻文本(News Article)、类别(Category)、标题(Headline)、发布日期(Publication Date)及新闻来源(Newspaper Source)。原始数据集共计包含664,880篇新闻文本,总词量达1.8551亿,总句量达1257万。 此外,我们借助NLP数据增强技术,从原始(非平衡)数据集生成了另一套平衡数据集:该数据集包含32万篇新闻文本,8个新闻类别各含4万篇文章。Potrika同时提供原始与平衡两套数据集,可满足多样化的NLP研究需求。据我们所知,Potrika是目前规模最大、覆盖最全面的新闻分类数据集。 有关该数据集的更多细节、采集流程与使用方法,请参阅我们的论文:https://doi.org/10.48550/arXiv.2210.09389。




