遇见数据集

Dataset and code of groundwater nitrate for machine learning

收藏
Zenodo2024-04-11 更新2026-04-07 收录
官方服务:

资源简介:

Data of groudwater nitrate and related data in North China Plain (NCP). The data including nitrate concentration of groudwater collected from more than 4,000 sites (wells). The groundwater samples were collected in 2005–2021, and the collection was conducted in May (before rainy season) and October (after rainy season) in each year for every site.During sampling, basic information about well location, groundwater depth, farmland planting pattern and soil types were collected. Sampling wells were divided into three types according to depth, shallow (0–30 m), medium (30–100 m) and deep (> 100 m). The planting pattern mainly involved intensive croplands, grain crops, vegetable crops and orchards. Soil types of each sampling site were obtained from the China soil database (http://vdb3.soil.csdb.cn/).The socio-economic and agricultural information of the study areas (take the districts of municipalities and prefecture-level cities of provinces as basic units) were acquired via the China Statistical Yearbook (http://www.stats.gov.cn/sj/ndsj/). The data includes agricultural planting area, grain crop area, vegetable planting area, orchard planting area, total facility agricultural area; fertilizer amount, nitrogen fertilizer amount, unit area nitrogen fertilizer amount; total output value of agricultural, forestry, animal and fishery husbandry, agricultural output value, forestry output value, animal husbandry output value, fishery output value; Gross Domestic Product (GDP), per capita GDP; total population, and rural population. Two scenarios are selected for the machine learning investigations: (i) sampling site information (SSI) and (ii) socio-economic and agricultural information of study area (SEAI). The former is to study the relationship between nitrate data of each sampling site and basic information, while the latter is to study the relationship between average nitrate and socio-economic and agricultural information in the study area. The model selection was a downselection process from well-known machine learning (ML) algorithms to select the best predictive model [50]. Two types of ML algorithms (regression and classification) were used to identify the best emulsion stability prediction model. For the regression prediction of nitrite, 8 commonly used learners were applied: linear regression, LASSO, KNeighbors, Decision Tree (DT), Random Forest (RF), Gradient Boosting (GB), Support Vector (SV), and Multi-Layer Perceptron (MLP) Regressor [51, 52]. For the classification prediction of nitrite, 7 learners were applied: Logistic Regression, KNeighbors, DT, RF, GB, SV, and MLP Classifier. The models were built using the Scikit-Learn machine learning library from Python (https://scikit-learn.org/). The classification was defined according to whether groundwater nitrate exceeded WHO’s maximum contaminant level (10 mg N/L, defined as 1) or not (defined as 0).

创建时间:
2024-04-11
二维码
社区交流群
二维码
科研交流群
商业服务