Tea Biochemical & Sensory Dataset (TeaBioSens)
收藏资源简介:
Understanding the intricate relationships between the biochemical content and sensory properties of tea is crucial for effectively predicting sensory quality. However, the complexity of these relationships makes them challenging to model through traditional methods. Machine learning has emerged as a powerful tool for modeling multidimensional relationships, but it requires a diverse and high-quality training dataset to effectively learn these intricate connections. To comprehensively investigate the intricate relationships between the biochemical content and sensory properties of tea, a high-quality training dataset is imperative, encompassing a wide array of biochemical features. There is no single comprehensive source that provides a holistic understanding of how the various biochemical content influences the sensory quality. Existing studies predominantly present results from a limited range of features, necessitating the creation of a meta dataset for robust insights. In response to this gap, we introduce the "Tea Biochemical & Sensory Dataset (TeaBioSens)," a meticulously curated compilation of data including important biochemical features that influences tea sensory quality. This dataset consists thirty different tea blends, made from four major processed tea variety, their biochemical content, and their sensory score obtained from a semi-trained panel using the lexicon based descriptive technique. TeaBioSens comprises a total of 600 data points. While certain parameters such as total soluble sugar (TSS), protein, total polyphenol (TP), caffeine (CAF), (+)-catechin (C), theaflavin (TF), thearubigin (TR), pH, citric acid, malic acid, ascorbic acid, oxalic acid, gallic acid, succinic acid, and L-theanine were directly estimated from experiments, factors such as TP/theanine, TF/TR, protein/TP, and CAF/TP were calculated. Experiments involving catechin subtypes and volatile organic acid is still ongoing and to be added soon. The development of this novel database is an ongoing effort, and we anticipate its continual improvement over time as more datapoints are incorporated.



