Dataset Based on Inline Monitoring and Offline Analysis of Yeast Cultivations
收藏资源简介:
The dataset comprises 52 cultivations of two yeast strains, each with a duration of 72 – 76 h, conducted under varying process parameters in bioreactor systems at 0.7 L and 10 L scale (working volume). At 0.7 L scale, 48 cultivations were performed as 12 experiments in four parallel reactors, while four additional cultivations were conducted at 10 L scale. The experimental design was inspired by brewing-related propagation processes: a top-fermenting (Saccharomyces cerevisiae) and a bottom-fermenting (Saccharomyces pastorianus) yeast were cultivated in wort-based media to approximate conditions relevant to industrial yeast propagation. A total of 43 variables are provided, comprising 30 inline variables (30 s resolution) and 13 offline variables sampled up to three times per day. The main archive (Dataset.zip) is organized into two primary subfolders: individual raw data tables and time-axis aligned raw data tables. The folder individual raw data tables contains 16 .xlsx files corresponding to the 16 experiments performed, including 12 experiments at small-scale and 4 experiments at pilot-scale. The folder time-axis aligned raw data tables contains a total of 104 .csv files, comprising 96 files from small-scale experiments (12 experiments × 4 reactors × 2 measurement types) and 8 files from pilot-scale experiments (4 experiments × 2 measurement types). For each experiment (referred to as ‘EXP’ in the datasets), a single Excel (.xlsx) file is provided containing the unprocessed raw data. Each file contains multiple worksheets, each corresponding to specific data and, in the case of small-scale experiments, to a specific reactor (referred to as ‘R’ in the datasets). Within each worksheet, the time axis starts at the point of inoculation, which is defined as 0 hours. For experiments conducted at small-scale, each Excel file contains 24 worksheets, corresponding to four parallel reactors with six worksheets per reactor. For experiments conducted at pilot-scale, each Excel file contains six worksheets in total. The Excel files are named according to the experiment number and the yeast strain and scale used. Since the individual sensor systems exhibited slight deviations from their nominal sampling intervals, small temporal offsets accumulated over time across data streams. To correct for this drift and obtain a temporally consistent dataset, all inline signals were aligned to a common master time axis. Accordingly, in addition to the individual raw data tables, time-axis-aligned raw data tables are also provided. For each experiment (EXP) and reactor (R), two comma-separated value (.csv) files are available: one containing inline measurement data and one containing offline measurement data. The .csv files use commas as delimiters between values within a row and periods as decimal separators. File names encode the experiment number, yeast strain, scale, reactor identifier (for small-scale experiments), and measurement type. The dataset is suitable for the development and evaluation of data-driven and hybrid modeling approaches, including soft sensors, as commonly applied in bioprocess engineering. For meaningful interpretation, inline sensor data should be considered in conjunction with the corresponding offline reference measurements provided in this dataset. As with other complex bioprocess datasets, careful preprocessing, and domain-informed interpretation are essential to avoid misinterpretation of sensor signals and inferred process states.



