Dataset Based on Chinese Hamster Ovary (CHO) Cultivations including Turbidity, Permittivity, O2 and CO2 Measurements
收藏资源简介:
The dataset comprises 24 cultivations of Chinese Hamster Ovary (CHO)-K1 cells, each with a duration of 161.5 – 328.5 h, conducted under varying process parameters in parallel bioreactor systems at 0.55 – 0.8 L scale (working volume). A total of 48 variables are provided, comprising 38 continuous inline variables and 10 offline variables sampled up to two times per day. The main archive (Dataset.zip) is organized into two primary subfolders: individual raw data tables and time-axis aligned raw data tables. The folder individual raw data tables contains 12 .xlsx files corresponding to the 12 experiments, each performed in two parallel reactors. These include nine batch experiments and three fed-batch experiments. The folder time-axis aligned raw data tables contains a total of 48 .csv files (12 experiments × 2 reactors × 2 measurement types). For each experiment (referred to as ‘EXP’ in the datasets), a single Excel (.xlsx) file is provided containing the unprocessed raw data. Each file contains 16 worksheets, each corresponding to specific data and to a specific reactor (referred to as ‘R’ in the datasets). Within each worksheet, the time axis starts at the point of inoculation, which is defined as 0 hours. The Excel files are named according to the experiment number and the process mode used. Since the individual sensor systems exhibited slight deviations from their nominal sampling intervals, small temporal offsets accumulated over time across data streams. To correct for this drift and obtain a temporally consistent dataset, all inline signals were aligned to a common master time axis. Accordingly, in addition to the individual raw data tables, time-axis-aligned raw data tables are also provided. For each experiment (EXP) and reactor (R), two comma-separated value (.csv) files are available: one containing inline measurement data and one containing offline measurement data. The .csv files use commas as delimiters between values within a row and periods as decimal separators. File names encode the experiment number, process mode, reactor identifier, and measurement type. The dataset is suitable for the development and evaluation of data-driven, mechanistic, and hybrid modeling approaches, including soft sensors, as commonly applied in bioprocess engineering. For meaningful interpretation, inline sensor data should be considered in conjunction with the corresponding offline reference measurements provided in this dataset. As with other complex bioprocess datasets, careful preprocessing, and domain-informed interpretation are essential to avoid misinterpretation of sensor signals and inferred process states.



