Datasets for paper: Systematic Evaluation of Time-Frequency Features for Binaural Sound Source Localization
收藏资源简介:
This repository contains binaural audio datasets generated for training and evaluating sound source localization models. The datasets were synthesized using the Binamix Python library, which performs binaural spatialization using head-related impulse responses (HRIRs) and binaural room impulse responses (BRIRs) from the SADIE II Database. The datasets include speech-only and mixed-content binaural recordings rendered across the full sphere using different spatial resolutions and subject splits. Included Datasets 1. TSP-SSL Training Set Generated using randomly selected speech samples from the TSP Speech Dataset. Spatialized using HRIRs from SADIE subjects: D1, H3–H6, and H11–H13. Rendered over the full sphere at 5° resolution in both azimuth and elevation. Total recordings: 21,312 2. TSP-SSL Validation Set Generated using the same procedure as the training set. HRIR subjects: H3 and H4. Total recordings: 5,328 3. TSP–SSL Test Set Speech-only in-domain evaluation set. Generated from non-overlapping speakers from the TSP dataset. HRIR subjects: H7–H10. Rendered at 5° spatial resolution. Total recordings: 10,656 4. SynBAD–Var Test Set Out-of-domain evaluation set derived from the localization sensitivity subset of the SynBAD dataset. Contains natural and synthetic sounds (e.g., castanets, pink noise). Spatialized using HRIRs from SADIE subject D2. Uses variable angular resolution with denser frontal sampling to match human localization sensitivity characteristics. Total recordings: 11,500 5. SynBAD–Fix Test Set Derived from the same source content as SynBAD–Var. Rendered using uniform 5° angular resolution across azimuth and elevation. Total recordings: 16,560 If you use this dataset, please cite the following papers: https://arxiv.org/abs/2511.13487 https://arxiv.org/abs/2505.01369



