Dataset for "Machine learning predictions on an extensive geotechnical dataset of laboratory tests in Austria"
收藏资源简介:
This dataset comprises over 20 years of geotechnical laboratory testing data collected primarily from Vienna, Lower Austria, and Burgenland. It includes 24 features documenting critical soil properties derived from particle size distributions, Atterberg limits, Proctor tests, permeability tests, and direct shear tests. Locations for a subset of samples are provided, enabling spatial analysis. The dataset is a valuable resource for geotechnical research and education, allowing users to explore correlations among soil parameters and develop predictive models. Examples of such correlations include liquidity index with undrained shear strength, particle size distribution with friction angle, and liquid limit and plasticity index with residual friction angle. Python-based exploratory data analysis and machine learning applications have demonstrated the dataset's potential for predictive modeling, achieving moderate accuracy for parameters such as cohesion and friction angle. Its temporal and spatial breadth, combined with repeated testing, enhances its reliability and applicability for benchmarking and validating analytical and computational geotechnical methods. This dataset is intended for researchers, educators, and practitioners in geotechnical engineering. Potential use cases include refining empirical correlations, training machine learning models, and advancing soil mechanics understanding. Users should note that preprocessing steps, such as imputation for missing values and outlier detection, may be necessary for specific applications. Key Features: Temporal Coverage: Over 20 years of data. Geographical Coverage: Vienna, Lower Austria, and Burgenland. Tests Included: Particle Size Distribution Atterberg Limits Proctor Tests Permeability Tests Direct Shear Tests Number of Variables: 24 Potential Applications: Correlation analysis, predictive modeling, and geotechnical design. Technical Details: Missing values have been addressed using K-Nearest Neighbors (KNN) imputation, and anomalies identified using Local Outlier Factor (LOF) methods in previous studies. Data normalization and standardization steps are recommended for specific analyses. Acknowledgments:The dataset was compiled with support from the European Union's MSCA Staff Exchanges project 101182689 Geotechnical Resilience through Intelligent Design (GRID).
本数据集涵盖20余年的岩土室内试验数据,采集范围主要覆盖维也纳、奥地利下奥地利州及布尔根兰州。数据集包含24项特征,记录了由颗粒级配试验、阿太堡界限(Atterberg limits)试验、普罗克特(Proctor)试验、渗透性试验与直接剪切试验所得的关键土体属性。部分试样附带地理位置信息,可支持空间分析。 该数据集是岩土研究与教学的宝贵资源,可支持用户探索土体参数间的关联关系并开发预测模型。此类关联示例包括液性指数与不排水抗剪强度、颗粒级配与内摩擦角,以及液限、塑性指数与残余内摩擦角之间的关联。 基于Python的探索性数据分析与机器学习应用已验证了本数据集在预测建模中的潜力,针对黏聚力与内摩擦角等参数的预测已取得中等精度。数据集兼具时间与空间跨度,加之多次重复试验,提升了其可靠性与适用性,可用于基准测试与验证分析型、计算型岩土工程方法。 本数据集面向岩土工程领域的研究人员、教育工作者与从业者,潜在应用场景包括优化经验关联式、训练机器学习模型,以及深化土力学认知。用户需注意,针对特定应用场景,可能需要开展缺失值插补、异常值检测等预处理步骤。 ### 核心特征 - 时间覆盖范围:20余年试验数据 - 地理覆盖范围:维也纳、奥地利下奥地利州及布尔根兰州 - 包含试验类型: - 颗粒级配试验 - 阿太堡界限(Atterberg limits)试验 - 普罗克特(Proctor)试验 - 渗透性试验 - 直接剪切试验 - 变量数量:24项 - 潜在应用方向:关联分析、预测建模与岩土工程设计 ### 技术细节 过往研究已采用K近邻(K-Nearest Neighbors, KNN)插补法处理缺失值,并通过局部离群因子(Local Outlier Factor, LOF)方法识别异常值。针对特定分析场景,建议对数据开展归一化与标准化处理。 ### 致谢 本数据集的编制得到欧盟玛丽·居里学者人员交流项目(MSCA Staff Exchanges)101182689号项目"基于智能设计的岩土韧性(Geotechnical Resilience through Intelligent Design, GRID)"的支持。



