Speech-Office-Sim - a dataset of simulated speech in office
收藏资源简介:
Summary SOS: Simulated Office Speech consists of two simulated datasets of speech recorded in an office environment: SOS-1SP (single-speaker, monologue-style) and SOS-2SP (two-speaker, conversation-style). Both datasets are created by acoustically rendering read speech from the VCTK corpus in a simulated stereo room, then mixing it with real office background noise and office-specific sound events (e.g., keyboard typing and office ambience) at controlled, reproducible levels. They are intended for research on speech activity detection, speech enhancement and separation, and automatic speech recognition under realistic, parametrically controlled office-noise conditions. Both datasets include frame-level ground-truth speech-activity annotations. Dataset overview SOS-1SP SOS-2SP Speaking style Single speaker, VCTK utterances concatenated back-to-back Two speakers, alternating turns Turn-length statistics N/A Matched to the CallHome English corpus's per-turn word-count distribution Segment length 1 minute per file 1 minute per file Noise conditions 10 (9 controlled + 1 clean reference) 10 (9 controlled + 1 clean reference) Per-file ground truth Frame-level speech-activity annotation Frame-level speech-activity annotation + per-turn speaker/text/timestamp metadata Every 1-minute source segment (a single-speaker file for SOS-1SP or a two-speaker conversation for SOS-2SP) is rendered under all 10 noise conditions described below. Thus, the same underlying speech content is available at every point in the noise-condition grid, enabling controlled comparisons, such as evaluating how model performance degrades as noise conditions become more challenging. Source materials Component Source Speech VCTK Corpus 0.92 [1] — 111 English speakers reading the same fixed set of newspaper/elicitation sentences Events Keyboard typing in office (Freesound, bonnyorbit) + Office ambience (Freesound, chelly01), pre-mixed into a single events bed [4] Background Channels 1 and 4 of the "OOFFICE" recording, DEMAND [2] Conversational structure reference CallHome English corpus [3] (via TalkBank) Important note on CallHome: CallHome audio recordings and transcripts are access-restricted under TalkBank/LDC terms and are not included in or redistributed with this dataset. CallHome is used solely as a reference corpus during dataset construction, specifically to estimate the empirical distribution of the number of words spoken by one participant before their conversation partner takes a turn. This distribution is then used to determine how VCTK utterances, which do not contain natural conversational turns, are grouped into turn-length-matched segments for SOS-2SP. No CallHome audio, transcript text, or per-utterance content is included in the released data. Dataset construction - Single-speaker segments (SOS-1SP) For each of the 111 VCTK speakers, their individual utterance recordings are concatenated, with short silence gaps between utterances, into contiguous single-speaker segments of approximately one minute using a simple greedy bin-packing procedure based on utterance durations. Recordings that are already close to one minute in duration are kept as-is, while segments that are too short to be usable (i.e., consisting mostly of padding) are discarded. Every source utterance is either included in a segment or explicitly logged as discarded, ensuring that the segment construction is fully accounted for. The resulting corpus consists of single-speaker, continuously spoken, approximately one-minute files: one speaker “talking to camera” per file, as if reading material at their desk. - Two-speaker conversations (SOS-2SP) Constructing natural-feeling two-speaker conversations from VCTK requires turn-length variability representative of real conversations. This is achieved in three steps: Turn-length reference. CallHome transcripts are parsed to extract, for each uninterrupted stretch of speech by one participant before the other takes over, the number of words spoken in that turn. This produces an empirical distribution of conversational turn lengths from spontaneous two-party English telephone conversations. Length-matched grouping. Each VCTK speaker's utterances are grouped into pseudo-turns. For a target turn length sampled from the CallHome distribution, a set of that speaker's own utterances is assembled using a randomized search such that their combined word count falls within ±20% of the target. This process is repeated until every VCTK utterance has been assigned to exactly one pseudo-turn group. Utterances that cannot be matched to a sampled target are retained as singleton groups, so no speech material is discarded. Each group is then concatenated into a single audio clip, with short cross-fades and brief silence gaps between utterances. This produces a pool of pseudo-turn clips for each speaker whose turn-length statistics approximate those of real conversational turns. Conversation assembly. VCTK speakers are paired, and pseudo-turn clips are drawn alternately from the two speakers' pools, with a short silence gap between turns, until the accumulated duration reaches 80–100% of the one-minute target. The resulting sequence is then symmetrically padded with silence to exactly one minute. This process is repeated using fresh, previously unused pseudo-turns until one of the two speakers in the pair runs out of usable material. Each resulting conversation file is accompanied by a per-turn metadata table containing the speaker ID, transcript text, onset timestamp, and source VCTK utterance ID. Speaker pairs are fixed throughout the dataset: each VCTK speaker is paired with exactly one other speaker. Consequently, no speaker appears in conversations with multiple partners. This fixed pairing makes within-pair speaker identification well-defined and avoids distributing a speaker's data across multiple pairings. - Speech spatialisation Each one-minute speech segment, whether from a single-speaker or two-speaker recording, is spatialized using a simulated shoebox room based on the image-source method, with 5th-order reflections and air absorption. The room represents a moderately absorptive office measuring 10 m × 10 m × 3 m, with a wall absorption coefficient of 0.80. A stereo microphone pair with 20 cm spacing is placed at the center of the room at a seated ear height of 1.5 m, and the speaker is positioned 2 m from the microphone pair. The current release uses a single centered source position, with the speaker directly facing the microphone pair. Although the simulation code supports azimuth offsets of up to 90° for future lateral-position releases, these are not used here: every file in the current release corresponds to a centered speaker. The room-rendered speech is then leveled to a fixed reference derived from reported acoustic measurements. The mean ambient level of an office has been reported at approximately 53.6 dBA (Vidal et al., 2023), while typical conversational speech levels have been reported at approximately 54.0 dBA at 1 m in a separate study of speech-level variation across office environments and communication types. Accounting for distance attenuation to 2 m gives a speech level of approximately 48 dBA, corresponding to a speech level approximately 5.6 dB below the reported office ambient level. This relative offset is applied to the measured level of the office-events source clip used in this dataset, providing a fixed internal reference level for the rendered speech before additional noise is mixed in. - Controlled mixing conditions Once a segment's speech has been rendered and leveled, office sound events and background noise are added at nine controlled combinations of two independent ratios, along with one clean reference condition containing neither added events nor background noise: EBR (Event-to-Background Ratio): the level of the office sound-event bed (keyboard typing and office ambience) relative to the background noise bed. SAR (Speech-to-Ambient Ratio): the level of the combined ambient bed (events + background, after EBR scaling) relative to the already-leveled speech. In other words, SAR specifies how many decibels below the speech level the ambient bed is set. EBR label Event level relative to background Perceptual effect low −50 dB (background) Office events are effectively inaudible / absent — background noise dominates alone mid +3 dB (events) Office events sit only slightly above the background, blending into it high +12 dB (events) Office events are clearly audible and prominent above the background SAR label Ambient level relative to speech Perceptual effect low −20 dB (ambient is 20 dB below speech) Ambient noise is barely audible; easiest listening condition mid −10 dB (ambient is 10 dB below speech) Ambient noise is clearly present but subordinate to speech high 0 dB (ambient equals speech level) Ambient noise is as loud as speech; hardest listening condition - Details Audio is delivered as FLAC (lossless), stereo, 48 kHz, 16-bit. There is no predefined train/validation/test split — all segments across all speakers are included, and users are expected to define their own splits (e.g. by speaker, to test speaker-independent generalization) appropriate to their task. Frame-level speech-activity ground truth is provided for every segment. It is computed directly from the clean reference condition : the signal is peak-normalized, divided into non-overlapping 50ms frames, and each frame is labelled active (1) or inactive (0) by thresholding its RMS level at −36 dB relative to the signal peak. The dataset generation process is reproducible using the code available in the following GitHub repository: https://github.com/modantailleur/codeSpeechOfficeSim Folder structure Each dataset is distributed as a single top-level folder (SOS-1SP/ or SOS-2SP/): ebr-{low,mid,high}-sar-{low,mid,high}/ — 9 folders, one for each noisy mixing condition (see the EBR/SAR table above) ebr-none-sar-none/ — the clean reference condition (rendered speech only, with no events or background noise) transc/ — one .txt transcript per unique speech segment, shared across the 10 noise-condition renders of that segment vadgt/ — one .npy frame-level speech-activity ground-truth file per unique speech segment, shared across the 10 noise-condition renders of that segment turns/ — SOS-2SP only: one .csv file per conversation, containing the speaker, text, and onset timestamp for each turn speaker-info.txt — VCTK's speaker metadata table (age, gender, accent, and region) Within each of the 10 condition folders, audio is organized by pan position (currently always centered), then by speaker for SOS-1SP or by speaker pair for SOS-2SP. For example: SOS-1SP/ └── ebr-high-sar-high/ └── pan_0/ └── p263/ └── spk_p263__p263_349-368_mic1__pan_0__ebr-high-sar-high.flac SOS-2SP/ └── ebr-high-sar-high/ └── pan_0/ └── p287-p288/ └── spk_p287-p288__p287-p288_conv1134_mic1__pan_0__ebr-high-sar-high.flac Audio is stored as lossless FLAC, stereo, 48 kHz / 16-bit. Every file is exactly 60 s long, with shorter recordings padded with silence and longer recordings truncated to the fixed duration. There is no predefined train/validation/test split. All speakers and segments are included in the release, and users are expected to define splits appropriate to their task, for example by speaker when evaluating speaker-independent generalization. Dataset statistics SOS-1SP SOS-2SP Speakers 110 108 speakers, 54 fixed pairs Unique speech segments 2,572 monologue segments 2,099 conversations Noise conditions per segment 10 (9 mixed + 1 clean reference) 10 (9 mixed + 1 clean reference) Total audio files 25,720 20,990 Segment duration 60 s (fixed) 60 s (fixed) Unique speech content ≈42.9 hours ≈35.0 hours Total delivered audio (all conditions) ≈428.7 hours ≈349.8 hours Uncompressed size on disk ≈86 GB ≈69 GB Zip archive size (split volumes) ≈92 GB across 9 volumes ≈73 GB across 7 volumes How to download and unzip Each dataset is distributed as a split ZIP archive: a final .zip file accompanied by a sequence of numbered volume files (.z01, .z02, ...), due to the size of the datasets. Download all volumes into the same folder before extracting. The archive cannot be extracted unless every part is present. Recommended: 7-Zip (7z), which can extract split ZIP archives directly without requiring a merging step: 7z x SOS-1SP.zip # extracts into a SOS-1SP/ folder in the current directory 7z x SOS-2SP.zip Alternative: Info-ZIP (zip/unzip), which requires the volumes to be merged into a single archive first. This temporarily requires approximately as much additional free disk space as the archive itself, for the merged copy, before extraction can begin: zip -s 0 SOS-1SP.zip --out SOS-1SP-merged.zip unzip SOS-1SP-merged.zip Known limitations Rendered speech uses a single, centered talker position; no lateral or off-axis positions are included in this release, though the construction pipeline supports them. This will be investigated in future versions of this dataset. Source speech is read speech (VCTK), not spontaneous speech; SOS-2SP matches conversational timing statistics to a real corpus but the linguistic content itself is still read material, not natural dialogue. Internal acoustic levels (LAeq) are relative/internal, not calibrated to real-world SPL (see "A note on levels" above). VCTK speaker pairing for SOS-2SP is fixed (not randomized per release); the same two speakers always appear together. License This dataset is released under CC BY 4.0, consistent with the license of its dominant redistributed source material (VCTK Corpus, CC BY 4.0). The DEMAND background recordings and the Freesound event clips are used under their respective source licenses — see the links in "Source materials" above for the original terms. No CallHome/TalkBank material is redistributed (see note above). How to cite @dataset{tailleur_2026_sos,author = {Tailleur, Modan and Chasle Cauchy, Théo and Lagrange, Mathieu},title = {{SOS: Simulated Office Speech -- SOS-1SP and SOS-2SP}},year = 2026,publisher = {Zenodo},version = {1.0},doi = {10.5281/zenodo.22687159},url = {https://doi.org/10.5281/zenodo.22687159}} References [1] VCTK Corpus 0.92: Yamagishi, J., Veaux, C., MacDonald, K. (2019). CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92). University of Edinburgh, CSTR. https://datashare.ed.ac.uk/handle/10283/3443 [2] DEMAND: Thiemann, J., Ito, N., Vincent, E. (2013). The Diverse Environments Multichannel Acoustic Noise Database (DEMAND). https://www.kaggle.com/datasets/chrisfilo/demand [3] CallHome English: TalkBank CallHome English corpus, used as a reference distribution only (no redistribution). https://talkbank.org/ca/access/CallHome/eng.html [4] Freesound event recordings: "Keyboard typing in office" by user bonnyorbit (https://freesound.org/people/bonnyorbit/sounds/399823/); "Office ambience" by user chelly01 (https://freesound.org/people/chelly01/sounds/541117/). [5] Room acoustics simulation via pyroomacoustics (Scheibler, Bezzam, Dokmanić, 2018). Contact modan.tailleur@gmail.com



