遇见数据集

Dataset for Shipwreck Susceptibility Predictive Modeling

收藏
Zenodo2026-09-20 更新2026-10-01 收录
官方服务:

资源简介:

README Dataset Title: A Model-Ready Dataset for Machine Learning-based ShipwreckSusceptibility Modeling in the Chinese Adjacent SeasVersion: 3 1. DATASET OVERVIEW This dataset supports shipwreck susceptibility mapping in the Chinese adjacent seasusing machine learning. It provides a balanced, deduplicated sample of historicalshipwreck locations and background locations, 16 marine environmental conditioningfactors, a fixed training/validation partition, and the code needed to reproduce thebenchmark end to end. The sample distributed for modeling comprises 688 points: 344 distinct shipwreckpositions and 344 background points. The full inventory of 856 harvested records isalso distributed, with flags that identify the duplication described in Section 4below, so that either the deduplicated sample or the original harvest can bereconstructed deterministically. Users upgrading from version 1 should read CHANGES.txt first: the sample size, thepartition and the benchmark figures all differ from that release. 2. FILE ORGANIZATION shipwreckData/ historical shipwreck inventory conditioningData/ the 16 conditioning factors and the study-area extent susceptibilityMapping/ model outputs, metrics and susceptibility maps code/ Python scripts, the fixed partition, and the FR rasters README.txt this file CHANGES.txt what changed in versions 2 and 3 3. DATA CONTENT A. shipwreckData/ Shipwreck_year.csv 856 harvested records, one row per record. Attributes: OBJECTID, Link, Name, Nationality, date_lost, Longitude, Latitude (WGS 1984 decimal degrees). Deduplication columns (see Section 4): wreck_group_id shared by records of the same vessel name, year and position (371 distinct values) position_id shared by records at the same reported coordinates (344 distinct values) grid_cell_id shared by records in the same 1/12 deg cell (311 distinct values) n_records_in_group size of the wreck_group is_group_primary 1 for one representative record per wreck_group_id is_position_primary 1 for one representative record per position_id <-- used here is_cell_primary 1 for one representative record per grid_cell_id shipwreck_points.gdb the same inventory as point feature classes: shipwreckPoints 856 wreck records carrying every attribute and deduplication column of Shipwreck_year.csv, so the modeling sample can be selected in a GIS with is_position_primary = 1 randomPoints the 856 background points generated for version 1, with in_release = 1 marking the 344 retained in this release shipwreckPoints_used the 344 positions of the modeling sample, pre-selected randomPoints_used the 344 background points of the modeling sample, pre-selected To reproduce the modeling sample, select records with is_position_primary = 1 in Shipwreck_year.csv or in the shipwreckPoints feature class, together with the background points marked in_release = 1 in randomPoints. The same 688 points are distributed ready-partitioned as code/trainingset.txt and code/validatingset.txt. B. conditioningData/ geospatial conditions.gdb distance, depth, shipdensity hydrodynamic conditions.gdb waveHeight, SwellwaveHeight, SwellwaveDirec, AtmosVorti, RotationWind depositional conditions.gdb temperature, salinity, ph, oxygen, nppv, chloro, phyto, zooplankton studyArea.gdb the study-area extent geotiff/ the same 16 layers as GeoTIFF, for users without access to Esri geodatabase drivers Coordinate reference systems. All layers are referenced to the WGS 1984 datum. Fifteen are stored in geographic coordinates (EPSG:4326). The distance-to-coastline layer is stored in WGS 84 / World Mercator (EPSG:3395), because the Euclidean Distance computation requires a projected coordinate system in order to return metric distances. Grids. The layers derive from products with different native footprints and are not co-registered: their dimensions vary between 203 x 483 and 204 x 502 cells, the distance layer in its projected grid is 190 x 505, and some origins differ by a fraction of a cell. Values must therefore be read at point or cell locations in each layer's own geometry. Resampling the layers onto a common grid with nearest-neighbor interpolation displaces values by up to half a cell and reassigns up to 14 % of sample points to a neighboring Frequency Ratio class. The scripts in code/ follow the native-geometry convention throughout. Note on the distance layer. Mercator inflates distance by 1/cos(latitude), from about 1.00 at 5 deg N to 1.31 at 40 deg N across this study area, so distance values are not directly comparable between the northern and southern parts of the domain. The maximum distance to the coastline within the study area is approximately 1,683 km. C. susceptibilityMapping/ RF0521/, SVM0521/, ANN0521/ per-model outputs: <model>_model.pkl the fitted classifier Training_metrics.csv performance on the training partition Validation_metrics.csv performance on the validation partition Training_roc.csv, Validation_roc.csv, *_roc_curve.png <model>_grid_search_results.csv full cross-validated grid search <model>_feature_importance.csv permutation importance <MODEL>Prediction.tif susceptibility raster, 203 x 483, EPSG:4326, float32, cell value = predicted probability D. code/ RFPrediction0521.py Random Forest: grid search, fitting, evaluation, SVMprediction0521.py spatial prediction ANNPrediction0521.py trainingset.txt fixed training partition, 481 points (240 wrecks) validatingset.txt fixed validation partition, 207 points (104 wrecks) training2025new_FR.gdb the 16 Frequency-Ratio-normalized rasters requirements.txt pinned package versions Column names in the partition files are in English and match the feature_columns list in the scripts. Each row carries OID_, X, Y, shipwreck (1 = wreck, 0 = background) and the 16 Frequency Ratio values. Run the scripts from within code/, one at a time and in any order: cd code python RFPrediction0521.py python SVMprediction0521.py python ANNPrediction0521.py Run them sequentially, not in parallel. All three write into sibling folders under code/ and a concurrent run can interleave the outputs of two processes. Hyperparameters are selected by repeated stratified cross-validation (five folds, five repeats) scored by ROC AUC. Each script writes its own grid-search results, metrics, ROC curves, permutation importances, fitted model and susceptibility map. With fixed random seeds the scripts are deterministic: a second run reproduces the distributed outputs exactly. Frequency Ratio coefficients in the distributed partition files are computed on the training partition only, so no information from the validation subset enters the feature transformation.

提供机构:
Zenodo
创建时间:
2026-09-20
二维码
社区交流群
二维码
科研交流群
商业服务