遇见数据集

Mid-infrared spectral dataset, MATLAB-based feature extraction tools and R-based Partial Least Squares Discriminant Analysis workflow for discriminating green and roasted specialty coffee according to post-harvest processing method

收藏
Mendeley Data2026-09-07 收录
官方服务:

资源简介:

This dataset contains preprocessed Fourier transform mid-infrared spectral data, MATLAB scripts for spectral feature extraction and an R-based partial least squares discriminant analysis workflow developed for the classification of green and roasted specialty coffee according to post-harvest processing method. The dataset includes spectral matrices generated after different preprocessing strategies, including raw spectra, baseline correction, standard normal variate transformation, multiplicative scatter correction, first derivative and second derivative treatments. These files are organized to support two complementary modeling approaches: full-spectrum modeling, where the complete spectral profile is used as input, and feature-extraction modeling, where statistical descriptors derived from the spectra are used as predictor variables. The MATLAB scripts provide the computational routines required to extract representative spectral features from each sample. These features include descriptors related to spectral intensity, signal energy, dispersion, asymmetry, peakedness and entropy, which summarize relevant properties of the mid-infrared spectral profiles. The resulting feature-extracted datasets are provided together with the complete spectral matrices, allowing users to compare the classification performance obtained from full spectral information and reduced feature-based representations. The accompanying R script implements a reproducible chemometric workflow for selecting the preprocessing strategy, choosing the modeling approach and defining whether the analysis is performed using green coffee, roasted coffee or both stages simultaneously. The post-harvest processing classes are converted from dummy-coded variables into a categorical response with three levels: washed, semi-dry and dry. The script automatically defines the predictor matrix according to the selected data structure, removes non-informative variables, applies stratified training and validation partitioning, calibrates the partial least squares discriminant analysis model and evaluates its predictive performance. The workflow exports prediction results, confusion matrices, overall and class-specific performance metrics, score plots, loading plots, variable importance values and diagnostic plots. Therefore, the dataset and associated computational tools provide a complete framework for assessing the influence of spectral preprocessing, feature extraction and data representation on the discrimination of specialty coffee processed by different post-harvest methods. This resource may support future studies on coffee authentication, non-destructive quality control, chemometric classification and reusable computational workflows for spectral data analysis in agri-food systems.

创建时间:
2026-08-03
二维码
社区交流群
二维码
科研交流群
商业服务