Deep Learning Benchmarks and Datasets for Social Media Image Classification for Disaster Response
收藏资源简介:
The crisis image benchmark dataset consists data from several data sources such as CrisisMMD, data from AIDR and Damage Multimodal Dataset (DMD). The purpose of this work was to develop a consolidated dataset, create non-overlapping train/dev/test set and provide a benchmark results for the community. We propose new datasets for disaster type detection, and informativeness classification, and damage severity assessment. Moreover, we relabel existing publicly available datasets for new tasks. We identify exact- and near-duplicates to form non-overlapping data splits, and finally consolidate them to create larger datasets. In our extensive experiments, we benchmark several state-of-the-art deep learning models and achieve promising results. We release our datasets and models publicly, aiming to provide proper baselines as well as to spur further research in the crisis informatics community. https://crisisnlp.qcri.org/crisis-image-datasets-asonam20 The labels in the dataset for different tasks are as follows: Task 1: Disaster types Earthquake Fire Flood Hurricane Landslide Not disaster Other disaster Task 2: Informativeness Informative Not informative Task 3: Humanitarian categories Affected, injured, or dead people Infrastructure and utility damage Not humanitarian Rescue volunteering or donation effort Task 4: Damage severity Little or none Mild Severe Please cite the following papers, if you use any of these resources in your research. Firoj Alam, Ferda Ofli, Muhammad Imran, Tanvirul Alam, Umair Qazi, Deep Learning Benchmarks and Datasets for Social Media Image Classification for Disaster Response, In 2020 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), 2020. [Bibtex] Firoj Alam, Ferda Ofli, and Muhammad Imran, CrisisMMD: Multimodal Twitter Datasets from Natural Disasters. In Proceedings of the 12th International AAAI Conference on Web and Social Media (ICWSM), 2018, Stanford, California, USA. [Bibtex] Hussein Mozannar, Yara Rizk, and Mariette Awad, Damage Identification in Social Media Posts using Multimodal Deep Learning, In Proc. of ISCRAM, May 2018, pp. 529–543.
本危机图像基准数据集(Crisis Image Benchmark Dataset)整合了来自CrisisMMD、AIDR以及损伤多模态数据集(Damage Multimodal Dataset,DMD)等多个数据源的数据。本研究旨在构建统一整合的数据集,创建无重叠的训练/开发/测试集划分,并为学界提供基准测试结果。我们针对灾害类型识别、信息性分类以及损伤程度评估任务构建了全新的数据集,同时对现有公开数据集进行了重新标注以适配新任务。我们通过识别精确重复与近似重复样本,实现了数据集的无重叠划分,并将其整合为更大规模的数据集。在我们开展的大量实验中,对多款当前主流的深度学习模型进行了基准测试并取得了优异的结果。我们公开发布了本数据集与模型,旨在为危机信息学领域提供可靠的基线模型,并推动该领域的进一步研究。相关项目页面:https://crisisnlp.qcri.org/crisis-image-datasets-asonam20 本数据集针对不同任务设置的标签如下: 任务1:灾害类型 地震、火灾、洪水、飓风、滑坡、非灾害、其他灾害 任务2:信息性分类 信息性、非信息性 任务3:人道主义相关类别 受灾、受伤或遇难人员、基础设施与公共设施损毁、非人道主义内容、救援志愿或捐赠行动 任务4:损伤程度 轻微或无损伤、中度损伤、重度损伤 若您在研究中使用了本资源,请引用以下论文: Firoj Alam、Ferda Ofli、Muhammad Imran、Tanvirul Alam、Umair Qazi,《深度学习基准与面向灾害响应的社交媒体图像分类数据集》,收录于2020年IEEE/ACM社会网络分析与挖掘国际会议(ASONAM 2020)。[Bibtex] Firoj Alam、Ferda Ofli、Muhammad Imran,《CrisisMMD:来自自然灾害的多模态推特数据集》,收录于第12届国际AAAI网络与社交媒体会议(ICWSM 2018),美国加利福尼亚州斯坦福,2018年。[Bibtex] Hussein Mozannar、Yara Rizk、Mariette Awad,《基于多模态深度学习的社交媒体帖子损伤识别》,收录于ISCRAM 2018年会议,2018年5月,第529-543页。



