MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators
收藏资源简介:
Datasets Used This repository contains code, benchmark scripts, and experimental results associated with the paper: MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators The experiments were conducted using the following publicly available datasets: VoxCeleb1 Citation: Nagrani et al. (2017) VoxCeleb1 is a large-scale audiovisual speaker recognition dataset containing speech recordings collected from interview videos. In this work, VoxCeleb1 was used to evaluate representation fidelity and downstream classification performance through a gender classification task. Dataset URL:https://www.robots.ox.ac.uk/~vgg/data/voxceleb/ SPIRA Citation: Casanova et al. (2021) SPIRA is a respiratory health dataset composed of speech recordings collected for the assessment of respiratory insufficiency and COVID-19 related symptoms. In this work, SPIRA was used to evaluate the proposed MFCCT frontend in a clinical respiratory classification setting. Dataset URL:https://github.com/SPIRA-Project LibriSpeech Citation: Panayotov et al. (2015) LibriSpeech is a corpus of read English speech derived from public-domain audiobooks. In this work, LibriSpeech samples were used as benchmark inputs for latency and energy measurements across multiple hardware platforms. Dataset URL:https://www.openslr.org/12



