Prediction of the Cu oxidation state from EELS and XAS spectra using supervised machine learning
收藏资源简介:
Dataset accompanying the manuscript "Prediction of the Cu oxidation state from EELS and XAS spectra using supervised machine learning" Gleason, S.P., Lu, D. & Ciston, J. Prediction of the Cu oxidation state from EELS and XAS spectra using supervised machine learning. npj Comput Mater 10, 221 (2024). https://doi.org/10.1038/s41524-024-01408-1 This dataset is stored as one zip file, which contains several files and subdirectories, which are outlined below: The main data file for this paper produced by the authors, "Cu_reproducable_alignment_df_extracted_110222.joblib" which is a serialized pandas dataframe containing all the simulated site averaged XAS spectra, pymatgen structure objects, oxidation state labels, and other chemical and physical identifiers used to simulate the spectra and train and evaluate the ML model discussed in the paper linked to this dataset. This dataset contains ~3500 site averaged simulated XAS spectra of Cu containing materials with several post processing steps developed by the authors to correct systematic errors in the simulation procedure, remove flawed spectra, and prepare the spectra for ML model training. These post processing steps are detailed in the "Methods - Training set generation" second in the associated manuscript. Three subdirectories named "xas paper", "Cu_deconvolved_spectra" and "Additional_Literature_Spectra" contain experimental EELS/XAS data either: extracted by the authors from the literature (in the case of "xas paper" and "Additional_Literature_Spectra" which are stored as csv files) or taken by the authors using the TEAM I microscope at the National Center for Electron Microscopy at Lawrence Berkeley National Laboratory (in the case of "Cu_deconvolved_spectra" stored as dm4 files). The subdirectory "Dataset_generation" which contains: a pandas dataframe containing ~3700 site averaged XAS spectra simulated using the FEFF9 code base. This database combines the ~1500 site averaged spectra extracted from The Materials Project, and labeled by the authors with the material's oxidation state and other chemical/physical descriptors with ~2200 site averaged spectra simulated by the authors in this work. a subdirectory called "FEFF Simulations" which contains additional all site specific simulated spectra used in this work. The zip file "Z=29.zip" and the .joblib file are spectra extracted from the materials project, and the other .zip files contain the ~3500 site specific spectra simulated in by the authors in this work. These are distinct from site averaged spectra, where multiple symmetrically inequivalent absorbing sites from one material are averaged into a representation for the entire material. The raw FEFF outputs and input files are stored as text files in wrapper directories labeled by the materials project ID of the structure. Each feff.in file contains the structure representation in enough detail to build a pymatgen structure object. A processed pandas dataframe, in which the spectra are site averaged, is also included. Some other small helper files used to make example figures in the published manuscript.



