遇见数据集

Data files used in the production of results presented in the manuscript of "Scale-dependent biogeographic reshuffling after the Messinian Crisis"

收藏
Zenodo2026-03-08 更新2026-05-26 收录
官方服务:

资源简介:

Data files presented in the manuscript of "Scale-dependent biogeographic reshuffling after the Messinian Crisis" Description of the data and file structure General description This dataset contains processed occurrence data and analytical outputs used to reconstruct hierarchical biogeographic structure of Mediterranean marine faunas across the Tortonian–Zanclean interval using the HespDiv hierarchical spatial subdivision framework. The primary occurrence and locality data originate from the Mediterranean fossil database of Agiadi et al. (2024) available at Zenodo:https://zenodo.org/records/13358435. The original dataset was processed using a series of R scripts provided in the associated code repository (10.5281/zenodo.18911915). Processing steps included: linking occurrence records with locality coordinates, correcting taxonomic inconsistencies, optional ecological categorization of taxa, merging occurrences at the genus level per locality, and filtering occurrences by age interval and taxonomic level. The resulting datasets were used as input for HespDiv bioregionalization analyses, significance testing of detected boundaries, and multiple sensitivity analyses assessing the robustness of spatial subdivisions to subsampling. The archive contains: intermediate datasets generated during preprocessing, outputs of HespDiv analyses under multiple similarity metrics, results of split-line significance testing, results of recursive and non-recursive sensitivity analyses, datasets linking spatial scale (polygon area) with taxonomic contributions to boundaries. The principal dataset used in downstream statistical analyses and figure generation is data/hsa2/comb_db.RDS, which integrates statistically significant biogeographic boundaries, their spatial scale, and taxonomic contribution values across time intervals. Detailed descriptions of individual files and variables are provided in the File description section. Files and variables File data.zip Description This archive contains the data/ directory required to run the analyses in the Messinian R code repository (link to repository is provided in the submitted Manuscript). The folder should be placed directly inside the repository. Primary data source Primary occurrence and locality data were downloaded from: Agiadi et al. dataset (Zenodo):https://zenodo.org/records/13358435 The following files originate from that dataset: data/coord.csv – locality coordinates and locality metadata data/MessinianDB.csv – occurrence records Data preprocessing The following intermediate datasets are produced from the primary data using scripts in the repository: data/linked_db.RDSOccurrence records linked with locality coordinates.Generated by: code/linking_data.R. data/db_clean.RDSCorrected version of linked_db.RDS incorporating fixes to several taxonomic inconsistencies.Generated by: code/clean_data.R. Taxonomic categorization datasets The script code/categorize_data.R generates multiple datasets that add variables categorizing taxa as pelagic or benthic: data/db_class_fdb.csv data/db_class_fdb_manual.csv data/db_class_fdb_manual_worrms_s.csv data/db_class_fdb_manual_worrms_sg.csv data/db_class_fdb_manual_worrms_sgf.csv Equivalent RDS versions are also saved: data/db_class_fdb.RDS data/db_class_fdb_manual.RDS data/db_class_fdb_manual_worrms_s.RDS data/db_class_fdb_manual_worrms_sg.RDS data/db_class_fdb_manual_worrms_sgf.RDS These datasets extend data/db_clean.RDS with ecological classification variables. The categorization is incomplete and was not used in the present analyses, but the framework allows further classification by extending the categorize_data.R script. Genus-level dataset data/db_merged.RDS Generated by code/merging_data.R fromdata/db_class_fdb_manual_worrms_sgf.RDS. Species occurrences belonging to the same genus are merged per locality so that only one genus occurrence is retained per locality. Filtered datasets used for HespDiv analyses data/hespdiv_data.RDSGenerated by code/filter_data.R.Contains occurrences filtered by age interval and taxonomic level. data/hespdiv_data_group.RDSGenerated by code/filter_data_group.R.Contains occurrences filtered by age interval, taxonomic level, and taxonomic group. HespDiv bioregionalization results The main outputs of the hierarchical spatial subdivision analysis are: data/hespdiv_out_str.RDS data/hespdiv_out_str_s15.RDS data/hespdiv_out_str_sor.RDS data/hespdiv_out_str_s15_sor.RDS Generated by: run_hespdiv.R run_hespdiv_sorensen.R These files contain the principal HespDiv bioregionalization results derived from data/hespdiv_data.RDS. Split-line significance testing Significance testing results are stored in: all_nl1_results.xlsx RDS files indata/nl/str/data/nl/str_s15/ Generated by significance_tests.R. Sensitivity analyses Two types of sensitivity analyses are provided. Non-recursive sensitivity analysis Stored in: data/hsa/ Generated by sensitivity_constrained.R. Recursive sensitivity analysis Stored in: data/hsa2/ Generated by hsa2.R. These analyses evaluate robustness of HespDiv split-lines under subsampling. Polygon area calculations data/areas.RDS Contains geographic areas of polygons generated in recursive sensitivity analyses (data/hsa2).Created by polygon_sizes.R. Scale–contribution datasets The following datasets are produced by contributions_vs_scale_obtain.R: data/hsa2/comb_db.RDS data/hsa2/all_contr.RDS files in data/hsa2/split_tests/ These combine: recursive sensitivity results (data/hsa2) polygon areas (data/areas.RDS) HespDiv outputs comb_db.RDS links split-line significance results with polygon areas and retains only statistically significant boundaries. Main analytical dataset The principal dataset used for downstream analyses is: data/hsa2/comb_db.RDS This dataset was used in scripts including: PCA_time-taxa-area-contributions.R scale_dependence.R SD_MH_contribution_determinants.R contributions_vs_scale_plot.R It also served as input for exploratory analyses (F_statistic_anova_approach.R, slopes_approach.R). Structure of comb_db Variables: method – subdivision method (mh or sor) age – time interval (Tortonian, pre-evaporitic Messinian, Zanclean) hd_id – HespDiv run ID within sensitivity analysis split.id – split-line identifier rank – hierarchical rank of split-line plot.id – polygon identifier parent.pol.id – parent polygon identifier area – polygon area area_prop – polygon area relative to convex hull area size_cat – manually defined polygon size category Taxonomic contribution variables: benthic_foraminifera bivalves bryozoans corals dinocysts echinoids fish gastropods marine_mammals nanoplankton ostracods planktic_foraminifera scaphopod_chitons_cephalopods sharks Each variable represents the taxonomic contribution to the corresponding split-line. Long-format dataset data/df_long2.RDS Generated by SD_MH_contribution_determinants.R. This dataset converts comb_db to long format (taxon and contribution as variables) and includes additional explanatory variables used to evaluate determinants of contribution values. These variables are described in the Supplementary Materials of the manuscript Code/software To reproduce these datasets run code stored at 10.5281/zenodo.18911915 . Access information Data was derived from https://zenodo.org/records/13358435 (Licenced under Creative Commons Attribution 4.0 International) and 10.5281/zenodo.18911915

提供机构:
Zenodo
创建时间:
2026-03-08
二维码
社区交流群
二维码
科研交流群
商业服务