AudioTime
收藏资源简介:
AudioTime数据集由上海交通大学X-LANCE实验室、上海人工智能实验室和香港中文大学(深圳)共同创建,专注于时间对齐的音频-文本数据。该数据集包含5000条训练数据和500条测试数据,每条数据包括一个长达10秒的音频片段、精确的元数据和时间精确标注的文本描述。数据集的创建过程包括音频片段的筛选、声音模拟和文本描述生成,旨在解决音频生成领域中时间控制精度不足的问题。
AudioTime Dataset was jointly created by X-LANCE Lab at Shanghai Jiao Tong University, Shanghai AI Laboratory, and The Chinese University of Hong Kong, Shenzhen. It focuses on time-aligned audio-text data. This dataset contains 5,000 training samples and 500 test samples, each of which includes a 10-second audio clip, precise metadata, and temporally accurately annotated text descriptions. The dataset construction process includes audio clip screening, sound simulation and text description generation, aiming to address the issue of insufficient temporal control accuracy in the field of audio generation.

- 1AudioTime: A Temporally-aligned Audio-text Benchmark Dataset上海交通大学X-LANCE实验室, 上海人工智能实验室, 香港中文大学(深圳) · 2024年



