Reproduction package: AI-based framework for automated contaminant recognition — detection and quantification of asbestos in recycled gypsum with wide-angle X-ray scattering
收藏资源简介:
Complete reproduction package for the study "AI-based framework for automated contaminant recognition: Detection and quantification of asbestos in recycled gypsum with wide-angle X-ray scattering". It contains everything needed to re-derive every number, table and figure of the manuscript: the measurement data in exactly the form the analysis read them, the analysis code, the aggregated results, and a record of the software environment. Measurement data. data/ holds 1,553 WAXS spectra in the concentration-folder structure the pipeline expects, so the analysis runs from this package alone. Of these, 1,532 spectra from 373 physical mixtures entered the study; the pipeline excludes whole concentration folders, the 100 wt% pure mineral references above all. Those excluded files are included here deliberately, so that the exclusion can be checked rather than taken on trust. The complete measurement collection, of which this is the subset used, is published separately at doi:10.5281/zenodo.18890488; that archive is the citable data source, and this package fixes which of its files the reported results rest on. Code. The analysis pipeline, the aggregation scripts, the figure generation, and the validation and control studies reported in the manuscript. Results. Aggregated results across all 18 evaluation seeds, per-seed evidence including per-spectrum predicted probabilities, the latent-dimension selection sweep, the rendered figures, and the separate control runs on which the central claims rest: cross-matrix transfer to a commercial fibre-cement board, a null-window control, and an alternative-window comparison. Provenance and integrity. run_manifest.json records the script checksum, the seeds used, the split method, the normalisation, the feature-space definition and the selected latent dimension. environment/ pins the interpreter and every package version. CHECKSUMS.sha256 carries a SHA-256 for every file and INVENTORY.csv lists each file with its size and role. Per-seed diagnostic images and saved model weights were removed from the raw run output — they are reproducible by re-running the pipeline and would otherwise bury the result files. Every dropped category is counted and reported in README.md. Note on exact reproduction: the run is seeded and TensorFlow operation determinism is enabled, so the PCA, PLS-DA and untransformed-spectrum results are deterministic given the same seeds. Autoencoder training is not guaranteed bit-identical across TensorFlow builds, drivers or CPU instruction sets; the reported multi-seed means and confidence intervals are the quantities intended to be reproducible.



