EarthSpeciesProject/synthetic-detect-diarize
收藏资源简介:
该数据集是一个合成的生物声学数据集,包含10秒的WAV音频和源注释元数据。数据集不包含生成的语言对对话、字幕、问答对或其他data-synth输出。数据集包含1,000,000行数据,分为50个分片,每个分片最多包含20,000行数据。文件包括WebDataset风格的分片和元数据文件,其中分片包含音频和选择表条目,元数据文件包含每个音频文件的信息。选择表包括说话人、信噪比(SNR)、源干和源文件等元数据。
This dataset contains synthetic bioacoustic 10-second WAV audio and source annotation metadata. It does not include generated language-pair conversations, captions, QA pairs, or other `data-synth` outputs. Rows: 1,000,000, Shards: 50, Maximum rows per shard: 20,000. Files include WebDataset-style shards and metadata files, where shards contain audio and selection table entries, and metadata files contain information for each audio file. The selection tables include diarization metadata such as Speaker, SNR (dB), Source Stem, and Source File.




