<b>Machine learning-enhanced monitoring of global copper mining areas</b>,
收藏资源简介:
This dataset provides a comprehensive, site-specific global assessment of land use areas associated with copper mining activities as of 2022. Using machine learning methodologies applied to multispectral remote sensing data, we mapped and classified operational land-use features, including open-cut pits, waste rock dumps, and tailings storage facilities. The dataset covers a total of 1,313 copper mines across 80 countries, encompassing a combined spatial extent of approximately 7,267 km².Observations were made using Sentinel-2 satellite imagery, characterized by high spectral (13 bands, 443–2203 nm) and spatial (10 m) resolution. Parameters measured include spectral reflectance values, Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), Bare Soil Index (BSI), Enhanced Vegetation Index (EVI), Index-Based Built-up Index (IBI), and topographic information derived from Digital Elevation Models (DEM). Operational data, including historical copper production, production capacity, mining intensity, start and end dates of operations, were integrated from Standard & Poor’s and the US Geological Survey databases.The temporal coverage of the dataset is the year 2022, ensuring temporal consistency and accuracy across global mining areas. Spatial coverage is global, with significant data density in regions such as Canada, Australia, China, the United States, Chile, Peru, and Mexico.The primary purpose of this dataset was to quantify and evaluate the environmental impacts of global copper mining, particularly land use intensity and potential ecological consequences of mining activities. This dataset enables detailed environmental assessments, aids in ecological risk management, supports supply chain sustainability studies, and assists policy-makers and stakeholders in improving resource management and minimizing ecological impacts.Data collection involved preprocessing Sentinel-2 satellite imagery via the Google Earth Engine (GEE) platform, applying cloud-free median composites, and training a Random Forest classification algorithm using manually collected sample points for accurate delineation of mining features. Model validation achieved an overall accuracy of 91.08%.The dataset is provided in vector format, containing polygons annotated with detailed operational attributes, specifically: 1. Name: Copper mine name; 2. Latitude: Latitudinal coordinates of the mine centroid (WGS84, decimal degrees); 3. Longitude: Longitudinal coordinates of the mine centroid (WGS84, decimal degrees); 4. P_C: Primary commodity extracted from the mine; 5. List_of_C: List of secondary commodities extracted; 6. A_S: Activity status of the mine (e.g., Active, Inactive); 7. State/Prov: State or province location of the mine; 8. Country: Country location of the mine; 9. Land_use: Area altered by mining activities (square meters); 10. Cum_Prod: Historical cumulative copper production (metric tons); 11. MI: Mining intensity (100 m²/ton)This structure facilitates easy integration and usability across multiple research and management disciplines. The Google Earth Engine (GEE) script used for remote sensing classification is also included to facilitate replication and further analysis.<br>
本数据集针对2022年的铜矿开采活动相关用地范围,提供了全面且针对具体站点的全球评估。 我们将机器学习方法应用于多光谱遥感数据,对作业用地特征进行了制图与分类,涵盖露天采场、废石堆与尾矿存储设施。 数据集覆盖全球80个国家的共计1313座铜矿,总空间范围约7267平方千米。 本次观测采用Sentinel-2卫星影像(Sentinel-2 satellite imagery),该影像具备高光谱分辨率(13个波段,波长范围443–2203 nm)与高空间分辨率(10米)。 所测量的参数包括光谱反射率值、归一化差分植被指数(Normalized Difference Vegetation Index, NDVI)、归一化差分水体指数(Normalized Difference Water Index, NDWI)、裸土指数(Bare Soil Index, BSI)、增强型植被指数(Enhanced Vegetation Index, EVI)、基于指数的建筑指数(Index-Based Built-up Index, IBI),以及由数字高程模型(Digital Elevation Models, DEM)提取的地形信息。 作业相关数据(包括历史铜产量、生产产能、开采强度、作业起止日期)整合自标准普尔(Standard & Poor’s)与美国地质调查局(US Geological Survey)数据库。 数据集的时间覆盖范围为2022年,可确保全球矿区的时间一致性与数据准确性。 其空间覆盖范围为全球,在加拿大、澳大利亚、中国、美国、智利、秘鲁与墨西哥等区域具备显著的数据密度。 本数据集的核心目标是量化并评估全球铜矿开采的环境影响,尤其是开采活动带来的用地强度变化与潜在生态后果。 该数据集可支撑详细的环境评估工作,助力生态风险管理,支持供应链可持续性研究,并协助政策制定者与利益相关方优化资源管理、降低生态影响。 数据收集流程包括通过谷歌地球引擎(Google Earth Engine, GEE)平台对Sentinel-2卫星影像进行预处理,生成无云中位数合成影像,并利用人工采集的样本点训练随机森林(Random Forest)分类算法,以精准划定采矿特征区域。 模型验证的总体准确率达到91.08%。 本数据集以矢量格式提供,包含带有详细作业属性的多边形要素,具体字段如下: 1. 名称("Name"):铜矿名称; 2. 纬度("Latitude"):矿山质心的纬度坐标(WGS84,十进制度); 3. 经度("Longitude"):矿山质心的经度坐标(WGS84,十进制度); 4. 主要开采矿种("P_C"):矿山开采的主要矿产品; 5. 次要开采矿种列表("List_of_C"):开采的次要矿产品清单; 6. 作业状态("A_S"):矿山的作业状态(例如:活跃、停产); 7. 州/省("State/Prov"):矿山所在的州或省; 8. 国家("Country"):矿山所在国家; 9. 用地面积("Land_use"):采矿活动改变的土地面积(平方米); 10. 累计产量("Cum_Prod"):历史累计铜产量(公吨); 11. 开采强度("MI"):开采强度(100平方米/吨)。 该结构便于多研究与管理学科的集成与复用。 本数据集同时附带用于遥感分类的谷歌地球引擎(GEE)脚本,以方便研究复现与后续分析。



