遇见数据集

Predicted Spatially Complete Zoning Map of North Carolina

收藏
Zenodo2024-03-29 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

Spatially-complete zoning map of North Carolina, USA. The <strong>results </strong>folder contains results of a machine learning (random forest) model predicting 3 core district zones (residential, non-residential, and mixed use) and 13 sub-district zones (open space, industrial, commercial, office, planned use, high-density residential, medium-high-density residential, medium-density residential, medium-low-density residential, low-density residential, agricultural residential, mixed use, and downtown). Results are provided as 30-m rasters (.tif) with each value corresponding to a zoning district. Table containing zone district ID (number) and zone district name (character string) is included in <strong>zone_classification.csv</strong>. Final (spatially complete statewide maps) can be found in the <strong>final_predicted </strong>folder. This folder includes Statewide core district results in <strong>NC_predicted_core.tif</strong> and statewide sub-district results in <strong>NC_predicted_sub.tif</strong>. Zoning was generalized and reclassified into 3 core district zones and 13 sub-district zones (described above). Reclassified zoning data, collected from 39 counties in North Carolina is provided in the <strong>observed </strong>folder with core districts in <strong>core_district_observed_zones.tif</strong> and sub-districts in <strong>sub_district_observed_zones.tif</strong>. Also in this folder is <strong>zoning_implementation_NC.csv</strong> which includes links to the source data (zoning map and zoning ordinance) for all collected data. Two models were created to predict zones under different data availability scenarios (i.e., scenarios that assume different levels of data availability). Predictions labeled “within_county” utilized the within-county model which predicts zoning districts in areas where zoning data is partially available for that county. To approximate scenarios of incomplete data accessibility, 20% of the data was randomly removed from training and reserved for independent performance assessments. Predictions labeled “between-county” utilized the between-county model which predicts zoning districts in areas where zoning data is inaccessible. To approximate this scenario, multiple between-county model iterations were computed by randomly removing entire counties from the training dataset and computing performance metrics on the removed (test) counties. Predictions are provided for both core districts and sub-districts (described above). Results from these models can be found in the <strong>predicted </strong>folder. This folder contains four subfolders: <strong>core_district_within_county</strong>, <strong>sub_district_within_county</strong>, <strong>core_district_between_county</strong>, and <strong>sub_district_between_county</strong>. Within each of these folders are predicted maps 30-m raster (.tif), performance reports including precision, recall, and f1 score overall and per district (.csv), and accuracy maps (3-km grid shapefile [.shp, .shx, .prj, .dbf]) with values corresponding to the proportion of misclassified pixels within a grid cell. Multiple randomized testing county samples were conducted for the between-county models. Each random sample is labeled <strong>r*_</strong> where * is replaced with a number between 1 and 15.

美国北卡罗来纳州空间全域覆盖的分区地图数据集。**results**文件夹中包含机器学习(随机森林,Random Forest)模型的预测结果,该模型可预测3类核心分区(住宅、非住宅、混合用途)与13类子分区(开放空间、工业、商业、办公、规划用地、高密度住宅、中高密度住宅、中等密度住宅、中低密度住宅、低密度住宅、农业住宅、混合用途、市中心)。预测结果以30米分辨率的栅格文件(.tif)形式提供,每个栅格值对应一种分区类型。**zone_classification.csv**文件中包含分区ID(数值型)与分区名称(字符型)的对应关系表。 空间全域覆盖的全州最终分区预测地图可在**final_predicted**文件夹中找到,该文件夹包含全州核心分区预测结果文件**NC_predicted_core.tif**与全州子分区预测结果文件**NC_predicted_sub.tif**。 原始分区数据已经过泛化与重分类,划分为前文所述的3类核心分区与13类子分区。从北卡罗来纳州39个县收集的重分类后分区数据存放在**observed**文件夹中,其中核心分区数据为**core_district_observed_zones.tif**,子分区数据为**sub_district_observed_zones.tif**。该文件夹同时包含**zoning_implementation_NC.csv**文件,其中收录了所有采集数据的源数据(分区地图与分区条例)的链接。 本数据集构建了两种模型,以在不同数据可得性场景下开展分区预测:即模拟不同数据可获取程度的预测场景。标记为“within_county”的预测结果使用县域内模型,用于预测该县分区数据部分可获取区域的分区类型。为模拟数据可获取性不足的场景,训练集中会随机移除20%的数据,留作独立性能评估使用。标记为“between-county”的预测结果使用跨县域模型,用于预测该县分区数据完全不可获取区域的分区类型。为模拟该场景,研究开展了多轮跨县域模型迭代:从训练集中随机移除完整的县域数据,并以被移除的县域作为测试集计算性能指标。数据集同时提供核心分区与子分区的预测结果(如前文所述)。 上述模型的预测结果存放在**predicted**文件夹中,该文件夹包含四个子文件夹:**core_district_within_county**、**sub_district_within_county**、**core_district_between_county**与**sub_district_between_county**。每个子文件夹中均包含:30米分辨率的预测栅格地图(.tif)、包含总体与各分区精确率、召回率与F1分数的性能评估报告(.csv),以及精度地图——精度地图采用3公里网格矢量形状文件(.shp、.shx、.prj、.dbf),其栅格值代表对应网格内误分类像素的占比。针对跨县域模型,研究开展了多轮随机测试县域采样,每一轮随机采样均以**r*_**命名,其中*将替换为1至15之间的数字。

提供机构:
Zenodo
创建时间:
2023-07-17
二维码
社区交流群
二维码
科研交流群
商业服务