BanGBook: A Multimodal Bengali Book Dataset for Genre Classification
收藏资源简介:
BanGBook-2501 is a publicly available multimodal Bengali book dataset developed for automatic book genre classification. The dataset contains 2,501 Bengali books categorized into six genres: Adventure, Horror, Mystery, Religious, Romance, and Thriller. Each record includes three complementary modalities: the book title, plot summary, and cover image, enabling research in multimodal machine learning, natural language processing, computer vision, and information retrieval. The textual metadata were collected from publicly available Bengali online bookstores, while the corresponding book cover images were obtained from the respective product pages. The dataset was cleaned by removing duplicate records, validating Unicode text, and ensuring image integrity. It contains no missing textual information and includes cover images for all books. This dataset was developed to support reproducible research on Bengali book genre classification and serves as the benchmark dataset used in the accompanying research paper: "BanGBook: A Multimodal Bengali Book Genre Classification Benchmark Using TF-IDF and ResNet50 Features Across Classical Classifiers." Researchers can use this dataset for multimodal classification, text classification, image classification, feature fusion, transfer learning, benchmark evaluation, and other Bengali NLP and computer vision applications.
BanGBook-2501是一款面向自动图书体裁分类任务研发的多模态孟加拉语图书公开数据集。该数据集包含2501本孟加拉语图书,共划分为六大体裁:冒险类、恐怖类、悬疑类、宗教类、言情类与惊悚类。每条数据均包含三种互补模态:图书标题、情节摘要与封面图像,可支撑多模态机器学习、自然语言处理、计算机视觉以及信息检索领域的相关研究。 该数据集的文本元数据采集自公开可用的孟加拉语在线书店,对应的图书封面图像则取自对应商品页面。数据集通过移除重复记录、验证Unicode文本格式以及确保图像完整性完成清洗流程,无缺失文本信息,且为所有图书均配备了封面图像。 本数据集专为孟加拉语图书体裁分类的可复现研究而研发,同时作为配套研究论文《BanGBook:基于TF-IDF与ResNet50特征的经典分类器多模态孟加拉语图书体裁分类基准》中所使用的基准数据集。研究人员可将该数据集应用于多模态分类、文本分类、图像分类、特征融合、迁移学习、基准评测以及其他孟加拉语自然语言处理与计算机视觉相关任务。



