BanglaCalamityMMD: A Comprehensive Benchmark Dataset for Multimodal Disaster Identification in the Low-Resource Bangla Language
收藏资源简介:
The BanglaCalamityMMD dataset is a comprehensive multimodal resource designed to address the significant gap in disaster identification within Bangla language text. Comprising a total of 7,903 instances spanning eight distinct categories: Landslides, Wildfire, Tropical Storm, Drought, Flood, Earthquake, Human Damage, and Non-Disaster—the dataset is meticulously divided into three subsets: 6,323 instances for training, 790 instances for testing, and 790 instances for validation. This structured division ensures that models can be trained effectively, tested rigorously, and validated accurately, thereby enhancing the overall reliability and applicability of disaster identification systems in Bangla. Here is the dataset description for various disaster categories: Category Train Test Validation Total ============================================ Earthquake 800 100 100 1000 Flood 800 100 100 1000 Landslides 803 100 100 1003 Wildfires 720 90 90 900 Tropical Storms 800 100 100 1000 Droughts 800 100 100 1000 Human Damage 800 100 100 1000 Non-Disaster 800 100 100 1000 ============================================= Total 6323 790 790 7903
孟加拉语灾害多模态数据集(BanglaCalamityMMD dataset)是一款综合性多模态研究资源,旨在填补孟加拉语文本灾害识别领域的显著空白。该数据集共包含7903条数据样本,涵盖8个明确的类别:滑坡(Landslides)、野火(Wildfire)、热带风暴(Tropical Storm)、干旱(Drought)、洪水(Flood)、地震(Earthquake)、人员损害(Human Damage)以及非灾害事件(Non-Disaster)。数据集被精心划分为三个子集:训练集6323条、测试集790条、验证集790条。这种结构化的划分方式可确保模型得到有效的训练、严谨的测试与精准的验证,进而提升孟加拉语灾害识别系统的整体可靠性与应用适用性。 以下为各灾害类别的数据集详情: 类别 训练集 测试集 验证集 总计 ============================================ 地震(Earthquake) 800 100 100 1000 洪水(Flood) 800 100 100 1000 滑坡(Landslides) 803 100 100 1003 野火(Wildfire) 720 90 90 900 热带风暴(Tropical Storm) 800 100 100 1000 干旱(Drought) 800 100 100 1000 人员损害(Human Damage) 800 100 100 1000 非灾害事件(Non-Disaster) 800 100 100 1000 ============================================= 总计 6323 790 790 7903




