Reproducibility Package for Controlled Robustness Engineering in Classical MFCC–LBG Speaker Identification
收藏资源简介:
This repository provides the reproducibility materials supporting the manuscript “Engineering Robustness in Classical Speaker Identification under Additive Noise: A Controlled Comparison of Feature-Level, Scoring, and Codebook Interventions.” The study uses classical mel-frequency cepstral coefficient (MFCC) and Linde–Buzo–Gray (LBG) vector-quantization speaker identification as a controlled platform for comparing robustness interventions introduced at different processing stages. The repository contains MATLAB source code for the baseline MFCC–LBG system, voice activity detection plus cepstral mean and variance normalization (VAD+CMVN), global diagonal Mahalanobis scoring, particle swarm optimization (PSO)-assisted LBG codebook optimization, and paired Wilcoxon signed-rank analysis with Holm adjustment. It also includes derived region-level and fold-level experimental outputs for TIMIT dialect regions DR4 and DR6, condition-preparation timing records, the statistical-comparison output, and the workbook used to derive the normalized robustness results reported in Table 4 and Figure 2. Experiments cover clean speech and additive white Gaussian noise (AWGN) and babble noise at 20, 10, and 5 dB SNR using six matched cross-validation folds. The original TIMIT speech corpus is not redistributed because access and redistribution are subject to the licensing terms of the Linguistic Data Consortium (LDC93S1). Researchers with authorized TIMIT access can use the released code and documentation to reproduce the experimental workflow. The package is intended to provide a transparent provenance chain from the experimental implementation and fold-level outputs to the statistical and normalized-robustness evidence reported in the manuscript.




