S. aureus Left-Censored Dataset: a detection-limited, interval-censored observation version of in silico Staphylococcus aureus kinetics in cheese matrices
收藏资源简介:
This dataset is the S. aureus Left-Censored Dataset, a detection-limited and interval-censored observation version of the in silico Staphylococcus aureus kinetics described in the companion S. aureus Kinetic Database (DOI 10.5281/zenodo.20678817). It was created as part of a study on biology-informed hybrid quantum-classical machine learning (HQML) for predicting S. aureus behaviour in heterogeneous cheese matrices. In predictive microbiology, many enumeration results fall below the analytical limit of detection (LOD). Such observations are not missing: they are known to lie somewhere between zero and the detection limit. Ignoring this structure, or naively substituting the detection limit, biases parameter estimates. This dataset reproduces that situation explicitly. For every time point of every environmental condition, it provides the latent (true) model concentration, a flag for whether the value would be left-censored at the detection limit, the observed value when detectable, and an interval (lower and upper bound) that correctly encodes the censored information for survival-analysis or censored-likelihood modelling. The table contains 9216 rows and 18 columns. Each row is one time point of one environmental condition. There are 768 unique environmental conditions, formed by crossing 12 temperature levels, 8 pH levels and 8 water-activity levels, each sampled at 12 time points spanning 0 to 720 hours. The rows correspond one to one with the S. aureus Kinetic Database. The values are not raw laboratory measurements; they are model-generated using cardinal parameters and assumptions drawn from the predictive-microbiology literature, with a detection limit then applied to emulate realistic censored observation. The dataset is provided in CSV and XLSX formats. Please note that we cannot guarantee accuracy of the information provided. We do our best to keep the dataset consistent but we do not take responsibility if any of the provided information turns out to be incorrect or incomplete. We do not recommend using this dataset for analyses or projects that require complete accuracy. Researchers are welcome to reuse the data for any project. Please cite this record upon use or when published. We encourage reuse under the same CC BY 4.0 License.



