OpenSWI-dataset
收藏资源简介:
OpenSWI: A Massive-Scale Benchmark Dataset for Surface Wave Dispersion Curve Inversion Home Page, Datasets, Codes, Articles Feng Liu, Sijie Zhao, Xinyu Gu, Fenghua Ling, Peiqin Zhuang, Yaxing Li, Rui Su, Lihua Fang, Lianqing Zhou, Jianping Huang, LeiBai Introduction Surface wave dispersion curve inversion plays a critical role in both shallow resource exploration and deep geological studies, yet it remains hindered by sensitivity to initial models, susceptibility to local minima, and low computational efficiency. Recently, data-driven deep learning methods, inspired by their success in computer vision and natural language processing, have shown promising potential to overcome these challenges. However, the lack of large-scale and diverse benchmark datasets remains a major obstacle to the development and evaluation of such methods. To address this gap, we introduce OpenSWI, a comprehensive benchmark dataset generated through the surface wave inversion dataset preparation (SWIDP) pipeline. OpenSWI comprises two synthetic datasets tailored to different research scales and application scenarios, namely OpenSWI-shallow and OpenSWI-deep, as well as an AI-ready real-world dataset for generalization evaluation, OpenSWI-real. OpenSWI-shallow is derived from the 2-D geological model dataset OpenFWI, containing over 22 million 1-D velocity profiles paired with their fundamental-mode phase and group velocity dispersion curves, spanning a broad spectrum of shallow geological structures (e.g., flat layers, faults, folds, and realistic stratigraphy). OpenSWI-deep is built from 14 global and regional 3-D geological models, comprising approximately 1.26 million high-fidelity 1-D velocity-dispersion data pairs for deep earth studies. OpenSWI-real, compiled from open-source projects, contains two sets of observed dispersion curves and their corresponding 1-D reference models, serving as a benchmark for evaluating the generalization of deep learning models. To demonstrate the utility of OpenSWI, we trained deep learning models on OpenSWI-shallow and OpenSWI-deep, and evaluated them on OpenSWI-real. The results show strong agreement between the predicted and reference velocity models, confirming the diversity and representativeness of the OpenSWI dataset. To facilitate the advancement of intelligent surface wave dispersion curve inversion techniques, we release the SWIDP toolbox, the OpenSWI datasets, trained deep learning models, and other examples, aiming to provide comprehensive support and open resources for the research community. Note!!!: Subsequent updates to this dataset can be found at huggingface (https://huggingface.co/datasets/LiuFeng2317/OpenSWI). Datasets More Details of the OpenSWI Datasets can be found at Huggingface (https://huggingface.co/datasets/LiuFeng2317/OpenSWI) OpenSWI-Shallow: 1D velocity profiles derived from 2D velocity models (OpenFWI dataset), paired with corresponding surface wave dispersion curves. OpenSWI-Deep: 1D velocity profiles generated from high-resolution 3D geological models, sourced globally and regionally, tailored for deep geological studies. OpenSWI-Real: AI-ready observational data from Long Beach, USA, and the China Seismological Reference Model Project, with 1D velocity profiles and corresponding surface wave dispersion curves. These datasets are ideal for training and evaluating deep learning models focused on surface-wave dispersion curve inversion tasks. 🏞️ OpenSWI-Shallow Resource Models: Includes diverse geological features such as flat layers, faults, folds, and their combinations, representing typical shallow subsurface structures. Profiles: ~22 million 1D velocity profiles, each paired with corresponding Rayleigh wave dispersion curves The dataset spans a period range from 0.2 to 10 seconds and covers 100 sampling points (including uniform, random, and logarithmic distributions) for each dispersion curve. This variety ensures robust training and evaluation across different geological scenarios. Geological Diversity: The models include a broad spectrum of real-world shallow subsurface structures, such as: Flat Layers Faulted Layers (Flat-Fault) Folds Folds with Faults (Fold-Fault) Real Style (Field) These diverse models make the dataset highly applicable for both synthetic and real-world seismic data inversion tasks. 🌍 OpenSWI-Deep Resource Models: Includes over 14 global and regional 3D geological models, derived from high-resolution seismic data. These models represent deep geological structures, from the crust to the mantle, spanning a variety of tectonic settings and geological environments. Profiles: ~1.26 million 1D velocity profiles, derived from the 3D models. The profiles span a period range from 1 to 100 seconds, covering 300 sampling points (including uniform, random, and logarithmic distributions) for each dispersion curve. These profiles offer high-resolution data suitable for deep geological studies and support advanced seismic inversion techniques. Geological Diversity: The 3D models come from various sources, including well-established models such as: LITHO1.0(Pasyanos et al., 2014) USTClitho1.0 (Xin et al., 2018) Central and Western US Models (Shen et al., 2013) Continental China (Shen et al., 2016) US Upper-Mantle Model (Xie et al., 2018) EUcrust Model (Lu et al., 2018) Alaska Model (Berg et al., 2020) CSEM-Europe Model (Blom et al., 2020; Fichtner et al., 2018; Çubuk-Sabuncu et al., 2017) CSEM Eastern Mediterranean Model (Blom et al., 2020; Fichtner et al., 2018) CSEM Western Mediterranean Model (Fichtner et al., 2018; Fichtner et al., 2015) CSEM South Atlantic Model (Fichtner et al., 2018; Colli et al., 2013) CSEM North Atlantic Model (Fichtner et al., 2018; Krischer et al., 2018) CSEM Japanese Island Model (Fichtner et al., 2018; Simutė et al., 2016) CSEM Australasian Model (Fichtner et al., 2018; Fichtner et al., 2010) These models provide a comprehensive representation of both regional and global deep geological structures, enhancing the dataset’s value for training deep learning models on complex inversion tasks. 🏔️ OpenSWI-Real Long Beach, USA (Fu et al., 2022): Contains 5,297 stations, each with phase velocity dispersion curves. The period range for the curves is 0.263 to 1.666 seconds, focusing on shallow subsurface structures. This dataset provides real-world observational data for evaluating model performance and generalization in seismic inversion tasks. China Seismological Reference Model (Xiao et al., 2024): Includes 12,901 grid points, each with phase and group velocity dispersion curves. The period range for these curves is 8 to 70 seconds, ideal for studying deeper geological structures. Data is sourced from a dense network of seismic stations across mainland China, offering comprehensive coverage for advanced inversion tasks.



