遇见数据集

Chemical compositions of grinded obsidian artefacts from Early Eneolithic site in Nitra-Selenec (Slovakia) and reference samples of carpathian obsidians

收藏
Zenodo2025-01-30 更新2026-05-26 收录
官方服务:

资源简介:

Data analysis of chemical composition (ED-XRF) of obsidian artefacts from Nitra. The chemical analysis of obsidian artifacts from Nitra was conducted using the R programming environment, applying a variety of statistical techniques and machine learning models. The preprocessing steps included importing raw data from an Excel sheet using the openxlsx package. Red obsidian samples were excluded, and elemental concentrations were normalized using pairwise log-ratio (pwlr) transformation with the compositions package. Missing values were imputed using the zCompositions package, employing the multinomial KM (mKM) approach suitable for compositional data. Subsequent transformations ensured consistency in elemental sums across samples. Samples with SiO₂ content below 50% were filtered out using the dplyr package. The final dataset underwent standardization of major oxides and trace elements to achieve consistent scaling before classification. For provenance classification, a Random Forest model from the randomForest package was applied. The model underwent cross-validation to optimize performance, and variable importance was assessed through built-in feature ranking functions. The results were visualized using scatter plots and ternary diagrams generated with the ggplot2 and ggtern packages. This work was supported by the Slovak Research and Development Agency under the Contract no. APVV-20-0521, APVV-23-0282 and project VEGA 2/0033/23. The research was conducted under the wider project Ready for the future: understanding long-term resilience of the human culture (CZ.02.01.01/00/22_008/0004593) and Institutional grant of Faculty of Science, Masaryk University, no. 2222/315010.

尼特拉地区黑曜石制品化学成分的能量色散X射线荧光光谱法(ED-XRF)数据分析。 本次针对尼特拉地区黑曜石制品的化学成分分析基于R编程语言环境开展,运用了多种统计技术与机器学习模型。预处理流程包括通过openxlsx扩展包从Excel表格中导入原始数据;剔除红色黑曜石样本后,借助compositions扩展包的成对对数比(pwlr)变换方法对元素浓度进行归一化处理;针对缺失值,采用适配成分数据的多项式KM(mKM)方法,通过zCompositions扩展包完成插补。 后续变换操作确保了所有样本的元素总量保持一致。通过dplyr扩展包过滤掉二氧化硅(SiO₂)含量低于50%的样本。最终数据集在分类前,对主量氧化物与微量元素进行了标准化处理以实现统一的量纲缩放。 针对产地溯源分类任务,采用randomForest扩展包构建随机森林(Random Forest)模型。通过交叉验证优化模型性能,并借助内置的特征排序函数评估变量重要性。研究结果通过ggplot2与ggtern扩展包生成的散点图与三元图进行可视化展示。 本研究得到斯洛伐克研究与发展署的资助,资助合同编号为APVV-20-0521、APVV-23-0282以及VEGA 2/0033/23项目。本研究依托于大型项目"面向未来:理解人类文化的长期韧性"(编号CZ.02.01.01/00/22_008/0004593)以及马萨里克大学理学院的机构资助项目(编号2222/315010)开展。

提供机构:
Zenodo
创建时间:
2025-01-30
二维码
社区交流群
二维码
科研交流群
商业服务