EarthSpeciesProject/ROOTS
收藏资源简介:
ROOTS是一个生物声学的音频-语言训练数据集,包含生成的语言任务与源音频的引用。该存储库仅包含语言/对话内容和公开的音频标识符,不包含音频文件本身。数据集包含44,790,034行数据,分为9,074个Parquet分片。数据集的模式(Schema)包括多个字段,如ID、训练层级、任务类别、任务、格式、源数据集、模板、源ID、音频路径、源音频ID、源URL、音频开始和结束时间、音频数量以及对话消息等。音频文件不包含在数据集中,部分行引用了公共源数据集(如Xeno-Canto或iNaturalist)的ID和URL,其他行则仅通过相对文件名/路径引用合成或私有源音频,这些音频文件可能无法从该数据集中公开获取。
ROOTS is a bioacoustic audio-language training dataset containing generated language tasks paired with references to source audio. This repository contains language/conversation content and public-oriented audio identifiers only; it does not host audio files. The dataset consists of 44,790,034 rows, divided into 9,074 Parquet shards. The schema includes fields such as ID, training tier, task category, task, format, source dataset, template, source ID, audio paths, source audio IDs, source URLs, audio start and end times, number of audios, and conversation messages. Audio files are not included in the dataset; some rows reference public source datasets like Xeno-Canto or iNaturalist through IDs and URLs, while others reference synthetic or private-source audio by relative filename/path only, which may not be publicly retrievable from this dataset alone.




