遇见数据集

Augmented and Real Microalgae Datasets for Biomass and Biochemical Composition Prediction

收藏
Zenodo2025-11-03 更新2026-05-26 收录
官方服务:

资源简介:

This project provides an interactive Algae Yield Predictor that estimates biomass, lipid, protein, and carbohydrate yields for selected microalgae species under defined culture conditions. The system integrates multiple machine learning models (XGBoost, LightGBM, CatBoost, MLP, and a stacked ensemble) trained on augmented and validated datasets. Key features include: Model selection: Compare predictions across single models or use the stacked ensemble (R² ≈ 0.89). Uncertainty quantification: Local 80% prediction intervals based on KNN neighborhood analysis. Species × medium clamping: Predictions are constrained to literature-reported ranges for biological plausibility. DOI reference matching: Closest culture conditions from published studies (via DOI database) are suggested. Visualization: Interactive response plots showing predicted trends across input variables with confidence bands. This repository contains two complementary datasets designed for the development and validation of predictive models of microalgal biomass and biochemical composition (lipid, protein, and carbohydrate fractions): Augmented Dataset (200,000 samples) Generated through controlled augmentation of experimental ranges reported in literature. Includes multiple species (e.g., Arthrospira platensis, Chlorella spp., Scenedesmus sp., Haematococcus pluvialis, Porphyridium purpureum). Covers diverse media formulations (BG-11, BBM, Zarrouk, TAP, artificial seawater, etc.). Features culture variables such as light intensity, photoperiod (day/night exposure), pH, temperature, and cultivation duration. Used primarily for model training. Real Dataset (Experimental Validation) Extracted and curated from actual cultivation experiments. Provides biomass yield (g/L) and proximate composition (lipid, protein, carbohydrate %) measured under specific species × medium × culture condition combinations. Used to benchmark and validate predictive models trained on the augmented dataset. Both datasets are harmonized with consistent feature naming, include categorical encodings for species/media, and are ready for direct use in machine learning pipelines (CSV format). These datasets have been applied in the Algae Yield Predictor project (Tiwari et al., 2025), enabling stacked ensemble modeling (XGB, LGBM, CatBoost, MLP) with literature-range clamping and DOI reference matching. Developed by Ashutosh Tiwari (Lead) and Siddhant Dubey ( Co-Lead) with contributions from Yamini Sumathi, this tool aims to support researchers in algal biotechnology, biofuel production, and bioprocess optimization by providing reproducible, literature-aware predictive analytics

提供机构:
Zenodo
创建时间:
2025-11-03
二维码
社区交流群
二维码
科研交流群
商业服务