A Bangla Multi-Source News Article Dataset
收藏资源简介:
This dataset comprises 19,516 Bangla news articles collected from six prominent Bangladeshi online newspapers: Ajker Patrika (4,991), bdnews24 (663), Daily Janakantha (5,801), Jagonews24 (6,069), Manabzamin (1,400), and RTV News (592). The data from each newspaper are stored separately in UTF-8 encoded JSON files, with each file containing a single JSON array. Each news record includes the headline, cleaned article content, raw article content, category labels in both Bangla and English, newspaper name, website, source URL, and word-count information. The dataset follows a common classification scheme consisting of 15 categories: National/Bangladesh, International, Entertainment, Sports, Economy, Lifestyle, Science & Technology, Health, Education, Politics, Art and Literature, Opinion, Religion, Migration, and Environment. The dataset is distributed both as individual JSON files and as a compressed Bangla_News_Dataset.zip archive. Within the archive, the newspaper files are organized in the data/ directory. A README.md file provides information about the dataset structure and fields, while schema.json defines the structure of an individual news record using JSON Schema. The dataset can support a variety of Bangla Natural Language Processing (NLP) and media-related research tasks, including text classification, cross-newspaper domain adaptation, headline generation, misinformation analysis, and computational media studies. The collected news articles remain the intellectual property of their respective publishers. They are made available for non-commercial research and text-mining purposes under the CC BY-NC 4.0 license.



