BanglaFakeNews: A Curated Dataset for Bengali Fake News Detection
收藏资源简介:
BanglaFakeNews is a large-scale, curated dataset developed for fake news detection in the Bengali language. The dataset contains 50,800 labeled Bengali news articles collected from multiple publicly available online news sources and verified fact-checking platforms. Each news article is manually assigned one of two labels: Fake or Authentic, enabling binary classification for misinformation detection. The dataset was created to address the scarcity of high-quality benchmark datasets for Bengali fake news research. It covers a diverse range of topics, including politics, health, entertainment, sports, technology, crime, education, business, and social issues, making it suitable for developing robust and generalizable machine learning and deep learning models. The text data have undergone preprocessing to improve consistency and usability. The preprocessing pipeline includes duplicate removal, text normalization, tokenization, stop-word removal, and stemming while preserving the semantic content of the original news articles. The dataset is provided in CSV format and is intended to facilitate reproducible research in natural language processing (NLP), misinformation detection, text classification, and artificial intelligence applications for low-resource languages. Researchers can use this dataset to develop, evaluate, and benchmark traditional machine learning models, transformer-based architectures, and hybrid deep learning frameworks for Bengali fake news detection. It may also support related research in sentiment analysis, misinformation analysis, domain adaptation, explainable AI, and multilingual NLP. If this dataset contributes to your research, please cite the associated publication: A Hybrid Deep Learning Framework for Fake News Detection in Bengali News.
孟加拉语虚假新闻数据集(BanglaFakeNews)是专为孟加拉语虚假新闻检测任务开发的大规模精选数据集。该数据集包含50800条带标注的孟加拉语新闻稿件,数据采集自多个公开在线新闻源与经认证的事实核查平台。每条新闻均由人工标注为「虚假」或「真实」两类标签,可用于虚假信息检测的二分类任务。 该数据集的构建初衷在于解决孟加拉语虚假新闻研究领域高质量基准数据集匮乏的问题。其涵盖政治、医疗、娱乐、体育、科技、犯罪、教育、商业与社会议题等多元主题,适用于开发鲁棒性与泛化能力俱佳的机器学习与深度学习模型。 为提升文本数据的一致性与可用性,已对其开展标准化预处理。预处理流程涵盖去重、文本归一化、分词(tokenization)、停用词移除与词干提取,且全程保留原始新闻稿件的语义内容。该数据集以CSV格式发布,旨在推动低资源语言在自然语言处理(Natural Language Processing, NLP)、虚假信息检测、文本分类及人工智能应用领域的可复现研究。 研究人员可借助该数据集,开发、评估并基准测试面向孟加拉语虚假新闻检测的传统机器学习模型、基于Transformer架构的模型以及混合深度学习框架。此外,该数据集还可支撑情感分析、虚假信息分析、域自适应、可解释人工智能(Explainable AI)及多语言自然语言处理等相关研究。 若该数据集为你的研究提供了助力,请引用如下关联发表论文:《孟加拉语新闻虚假新闻检测混合深度学习框架》(A Hybrid Deep Learning Framework for Fake News Detection in Bengali News)



