DisasterChain
收藏资源简介:
DisasterChain是一个高质量、多模态的数据集,旨在推动因果推理、环境监测和灾害影响评估研究。该数据集基于EM-DAT数据库,提供了自2014年以来全球1,400多个极端环境事件的全面重建。与传统仅提供静态元数据的灾害目录不同,DisasterChain引入了“基于真实事件的因果影响链”(Grounded Causal Impact Chains),使用先进的大语言模型(Qwen 2.5 72B和Llama 3 70B)逐步提取物理和社会经济影响的因果序列(例如:极端降水→河流泛滥→洪水→基础设施损坏→人口流离失所)。关键的是,因果链中的每一步都通过已核实的新闻引用严格锚定于真实世界报道,保证了零幻觉率。此外,每个事件都附有多模态数据,包括21天的天气时间序列(来自Open-Meteo API)和多光谱卫星图像可用性(Sentinel-1/2/3),使其成为连接自然语言处理(NLP)与地球观测(EO)的独特资源。数据集支持因果推理与事件提取、多模态灾害评估(结合卫星图像、气象数据和文本来预测物理损坏或人员伤亡)、以及气候影响分析等任务。数据以JSON格式提供,每个事件包含根元数据(标准标识符、地理坐标、日期、人口影响指标)、weather_data(事件前后统计摘要及21天逐日时间序列)、satellite_data(受影响区域的边界框及传感器可用性数组)、news_data(信息来源及相关性得分)、summary(由Llama 3生成的简洁事实叙述)和causal_chain(由Qwen 2.5生成的规范化因果步骤,包含事件类型、描述和支持性引文)。数据集的构建遵循严格的多阶段管道,包括事件锚定、多模态检索、文本上下文收集、因果提取(使用模糊匹配算法验证引文)、以及语义规范化(将500+原始事件类型归一化为约390个标准化类别)。需要注意的偏差包括:因果链的丰富程度偏向于获得国际或英语媒体广泛报道的事件,卫星光学图像在恶劣天气下受云层影响(提供Sentinel-1 SAR作为补充)。
DisasterChain is a high-quality, multimodal dataset designed to advance research in causal reasoning, environmental monitoring, and disaster impact assessment. Based on the EM-DAT database, it provides a comprehensive reconstruction of over 1,400 extreme environmental events worldwide since 2014. Unlike traditional disaster catalogs offering only static metadata, DisasterChain introduces Grounded Causal Impact Chains extracted step by step using advanced large language models (Qwen 2.5 72B and Llama 3 70B) to capture causal sequences of physical and socioeconomic impacts (e.g., extreme precipitation → river overflow → flood → infrastructure damage → population displacement). Critically, each step in the causal chain is strictly grounded in verified news citations, ensuring zero hallucination. Additionally, each event is accompanied by multimodal data, including 21-day weather time series (from Open-Meteo API) and multispectral satellite imagery availability (Sentinel-1/2/3), making it a unique resource bridging Natural Language Processing (NLP) and Earth Observation (EO). The dataset supports tasks such as causal reasoning and event extraction, multimodal disaster assessment (combining satellite imagery, meteorological data, and text to predict physical damage or casualties), and climate impact analysis. Data is provided in JSON format, with each event containing root metadata (standard identifiers, geographic coordinates, dates, population impact indicators), weather_data (pre- and post-event statistics summaries and 21-day daily time series), satellite_data (bounding boxes and sensor availability arrays for affected areas), news_data (sources and relevance scores), summary (concise factual narratives generated by Llama 3), and causal_chain (normalized causal steps generated by Qwen 2.5, including event types, descriptions, and supporting citations). The dataset construction follows a rigorous multi-stage pipeline including event anchoring, multimodal retrieval, text context collection, causal extraction (with fuzzy matching algorithms to verify citations), and semantic normalization (normalizing 500+ original event types to approximately 390 standardized categories). Notable biases include richer causal chains for events with broad international or English media coverage, and satellite optical imagery affected by cloud cover in adverse weather (with Sentinel-1 SAR provided as a supplement).
DisasterChain 数据集概述
基本信息
- 数据集地址:https://huggingface.co/datasets/Franco7Scala/DisasterChain
- 主页:http://disasterchain.icar.cnr.it
- 仓库:https://github.com/Franco7Scala/DisasterChain
- 联系点:Francesco Scala(francesco.scala@icar.cnr.it)
- 许可证:Apache-2.0
- 语言:英语(en)
- 数据规模:1K 至 10K
- 任务类别:文本生成、视觉问答、时间序列预测、文本分类
- 标签:气候变化、自然灾害、因果推理、多模态、地球观测、Sentinel、Open-Meteo、EM-DAT
数据集摘要
DisasterChain 是一个高保真、多模态数据集,旨在推进因果推理、环境监测和灾害影响评估领域的研究。该数据集基于 EM-DAT 数据库构建,全面重建了 2014 年以来超过 1,400 起全球极端环境事件。
与仅提供静态元数据的传统灾害目录不同,该数据集引入了 Grounded Causal Impact Chains(有据可依的因果影响链)。利用先进的大型语言模型(Qwen 2.5 72B 和 Llama 3 70B),提取物理和社会经济影响的逐步因果序列(例如:极端降水 -> 河流泛滥 -> 洪水 -> 基础设施损毁 -> 人口流离失所)。因果链中的每一步都通过经过验证的支持性引文严格锚定到真实世界新闻报道。
每个事件还通过多模态数据得到丰富,包括 21 天天气时间序列和多光谱卫星影像可用性,使其成为连接自然语言处理(NLP)与地球观测(EO)的独特资源。
支持的任务与应用
- 因果推理与事件抽取:训练模型识别非结构化文本中的级联效应;
- 多模态灾害评估:结合卫星影像(Sentinel-1、Sentinel-2)、天气数据和文本,预测物理损害或人员伤亡;
- 气候影响分析:分析特定天气触发因素在不同全球区域的 socio-economic 下游影响。
数据集结构
数据集以 JSON 格式提供。核心 JSON 结构为每个灾害事件包含以下嵌套对象:
1. 根元数据与影响数据
标准化标识符、地理坐标、日期以及来自 EM-DAT 的定量人类影响指标(死亡人数、受影响人口、经济损失)。
2. weather_data
通过 Open-Meteo API 获取,包含:
- 事件前后统计摘要(平均降雨量、降雪量、最低/最高气温);
- 捕捉事件前、中、后气象演变的 21 天每日时间序列。
3. satellite_data
通过 Copernicus Data Space Ecosystem(CDSE)/ Sentinel Hub 获取:
- 受影响区域的精确边界框(
bbox); - Sentinel-1(SAR)、Sentinel-2(光学/假彩色)和 Sentinel-3(热红外)的时间可用性数组;
- 与 ESA WorldCover 土地利用数据的集成。
4. news_data
关于信息检索过程的元数据:
- 查询来源(Google News、ReliefWeb、IFRC、Wikipedia);
- 相关性分数和过滤惩罚项,确保仅保留高质量、事件特定的新闻文本。
5. summary 与 causal_chain(LLM 推理层)
- Summary:由 Llama 3 70B 生成的事件简明、事实性叙述。
- Causal Chain:由 Qwen 2.5 72B-Instruct 生成。一个规范化的 JSON 数组,包含灾害的时间顺序步骤。每一步包括:
type_event:语义规范化类别(例如:洪水、伤亡、基础设施损毁);description:影响的具体表现;supporting_quote:来自检索新闻上下文的精确或模糊匹配引文,确保因果抽取的 零幻觉率;token_usage:计算足迹追踪。
方法论与数据流水线
该数据集的构建遵循严格的多阶段流水线,旨在消除 LLM 幻觉并确保物理一致性:
- 事件锚定:从 EM-DAT(2014 年后)获取基础事件,并通过 Nominatim 对位置进行地理编码;
- 多模态检索:自动获取历史天气数据和卫星影像足迹;
- 文本上下文收集:抓取、解析并算法化评分新闻文章的相关性;
- 因果抽取:使用 Qwen-72B 抽取因果图。应用严格的基于 Python 的模糊匹配算法(阈值 ≥ 0.80),将每个 LLM 生成的引文与原始新闻文本进行交叉引用。任何未经验证的因果步骤都会被系统性剔除;
- 语义规范化:将 500 多个原始事件类型算法化规范化为约 390 个标准化类别的清晰分类体系。
偏差、风险与局限
- 媒体报道偏差:因果链的丰富程度和新闻数据的可用性天然偏向于获得大量国际或英语媒体报道的事件。偏远地区的事件可能因果链较短。
- 卫星数据可用性:云层覆盖严重影响恶劣天气事件(如飓风)期间 Sentinel-2 光学影像的可用性。提供 Sentinel-1 SAR 以缓解此问题。
引用
如果使用该数据集,请引用论文:
bibtex Coming soon...





