遇见数据集

High-latitude Pacific Ocean Sediment Geochemistry and XRF Data for Geoscientific Foundation Models

收藏
Zenodo2025-07-25 更新2026-05-26 收录
官方服务:

资源简介:

Introduction The dataset is a following development after the dataset (Chao et al., 2022). Besides the published XRF spectra-target measurements (CaCO3 and TOC) pairs of data, we further upload all raw XRF spectra in the project without alignments of the target measurements. They are compiled in a machine learning ready format for both pre-training and fine-tuning of foundation models. The investigated cores, which form the datast, are mostly retrieved across the high- to mid-latitude Northwest Pacific (37°N-52°N) and the Pacific sector of the Southern Ocean (53°S-63°S), with a water depth coverage from 1211 to 4853 m: Cruise SO264 in the subarctic Northwest Pacific with R/V SONNE in 2018 Cruise PS97 in the central Drake Passage with RV Polarstern in 2016 Cruise PS75 in the Pacific sector of the Southern Ocean in 2009/2010. Cruise KOMEX I and KOMEX II with R/V Akademik Lavrentyev in 1998 and cruise SO178 in 2004 in the Okhotsk Sea. For detailed measuring information, please checkout the published paper (Lee et al., 2022) and previous dataset (Chao et al., 2022).The use of this dataset is documented in the GitHub repo. Folder Structure raw: Containing raw spectra in the Avaatech XRF Core Scanner format. Each subfolder contains the raw data for a core series. legacy: Containing previously compiled data in Lee et al., 2022. pretrain: Containing data used for pre-training. It is built from the previously compiled spectra data `legacy/spe_dataset_20220629.csv`. The `train` subfolder has the training and validation sets. The `test` subfolder contains the data selected during fine-tuning as the zero-shot test, i.e., case study in the published paper. +- train +- spe (all spetra) +- info.csv (training set spectrum list) +- val.csv (validation set spectrum list) +- test +- spe +- info.csv (case study spectrum list) fine-tune: Containing data used for fine-tuning. The `train` subfolder has the training and validation sets. The `test` subfolder is the data for zero-shot test, i.e., case study in the published paper. +- CaCO3% +- train +- spe (all spetra) +- target (all target measurements) +- info.csv (training set spectrum-target pair list) +- info_#.csv (splits from the info.csv in different data amounts) +- val.csv (validation set spectrum-target pair list) +- test (same as in train) +- TOC% (same as in CaCO3%) The case study (i.e., test set) is composed of three cores ('PS75-056-1', 'LV28-44-3', 'SO264-69-2') isolated from the beginning and not used in both the pre-training and fine-tuning process. The rest of data are randomly split in to the trainging and validation sets wtih 4:1 ratio. The script is `src/datas/build_data.py` in the GitHub repo. Acknowledgements We thank the crew and the science parties of different cruises for their contributions to core and sample acquisition on the respective expeditions. We are very grateful to Dr. Weng‐Si Chao, Prof. Dr. Ralf Tiedemann, Dr. Lester Lembke‐Jene, and Dr. Frank Lamy for providing these data. We also sincerely thank Valéa Schumacher, Susanne Wiebe, and Rita Fröhlking and student assistants at the AWI Marine Geology Laboratory in Bremerhaven for technical assistance with XRF-scanning, CaCO3 and TOC measurements.

提供机构:
Zenodo
创建时间:
2025-07-25
二维码
社区交流群
二维码
科研交流群
商业服务