A Large-Scale Wi-Fi Channel State Information Dataset for Contactless Human Speech Recognition
收藏资源简介:
This dataset is designed for Wi-Fi-based human speech recognition, offering a large-scale, structured collection of wireless signal variations induced by vocal articulations. It comprises 22,500 individual trials captured from 25 subjects, each recording 30 distinct English words with 30 repetitions per word. Data acquisition was performed in a controlled indoor environment using commercial off-the-shelf hardware, specifically a Sagemcom 2704 access point and a desktop equipped with an Intel 5300 NIC. The dataset includes both Channel State Information (CSI) and Received Signal Strength Indicator (RSSI) values, providing a comprehensive and complementary representation of signal perturbations for word-level speech analysis. As a benchmark resource for the research community, this dataset facilitates the development, evaluation, and comparison of advanced signal processing and machine learning algorithms for contactless speech sensing. By providing extensive multi-subject data, it addresses the need for robust benchmarks in privacy-preserving human-computer interaction and assistive technologies. Researchers can leverage these raw Wi-Fi measurements to investigate the feasibility and generalizability of microphone-free voice interaction systems in diverse indoor settings, contributing to the broader field of wireless human sensing.



