Skin-Vibration Speech Reconstruction Dataset, Models, and Smartphone Inference Assets
收藏资源简介:
This record contains the data and model artifacts supporting skin-vibration-based word classification and real-time speech reconstruction. It includes four-word recordings from two participants, 270 paired skin-vibration sensor and microphone recordings from six participants, pretrained and task-adapted LaCo-SENet checkpoints, smartphone ONNX deployment assets, and compact source data for the reported analysis summaries. The associated analysis, training, evaluation, and Android deployment code is available at https://github.com/yskim3271/skin-vibration-speech-reconstruction. The creators listed here are the three co-first authors of the associated manuscript; Yunsik Kim curated the data and code release. Version 1.1. The four-word classification metadata in analysis_source_data_for_speech_pipelines.tar.gz has been recomputed. The previous version was derived from a segment set containing duplicated clips; because a duplicate can fall in a different cross-validation fold from its copy, a held-out sample also appeared in the training set, inflating the reported accuracy to 1.0. Regenerating the segments from the released raw recordings gives a duplicate-free dataset and a five-fold cross-validated accuracy of 0.9975 (399 of 400 segments). The protocol is unchanged. A new file, four_word_classification_model_comparison.json, records the classical-model comparison on the same duplicate-free dataset using nested cross-validation. The audio recordings and model archives are unchanged from version 1.0.



