MetaCost XGBoost Training and Evaluation Dataset with MATBLAB Codes and files for generating proxies
收藏资源简介:
The dataset consists of two curated subsets designed for the classification of alteration types based on geochemical and proxy variables. The traditional dataset (Trad_Train.csv and Trad_Test.csv) is derived directly from the original complete geochemical dataset (alldata.csv) without any missing values. It includes the original geochemical features and serves as a baseline for model training and evaluation. The simulated dataset (Simu_Train.csv and Simu_Test.csv) was generated using custom MATLAB scripts that transform the original geochemical features into proxy variables, simulating conditions based on multiple geostatistical realizations. These proxy values are represented in a Gaussian scale, which may include negative values due to normalization. In the simulated datasets, the target variable Alteration was initially encoded as integers, with the following mapping: 1 = AAA, 2 = IAA, 3 = PHY, 4 = PRO, 5 = PTS, and 6 = UAL. These datasets are intended for evaluating the performance and generalizability of classification models in both raw and proxy-variable scenarios.



