Audio files for: Expressive range characterization of open text-to-audio models (AIIDE 2025)
收藏资源简介:
Audio files for the paper: Jonathan Morse, Azadeh Naderi, Swen Gaudl, Mark Cartwright, Amy K. Hoover, Mark J. Nelson (2025). Expressive range characterization of open text-to-audio models. In: Proceedings of the 21st AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment. Contents: fig1_samples.zip: Generated audio for the examples in Fig. 1. Two prompts; two models; 100 samples for each. thunder_samples.zip: Generated audio for the running "thunder" example. One prompt; two models; 100 samples for each. Source for Figs. 2-4. esc50_samples.zip: Generated audio for the prompt "Sound of X" for each label X in the ESC-50 environmental audio dataset. Fifty prompts; three models; 100 samples for each. Source for Figs. 5-6 and Table 1. generation_scripts.zip: Python scripts used to generate audio from the three models.
本论文配套音频数据集: 作者:乔纳森·莫尔斯(Jonathan Morse)、阿扎德·纳德里(Azadeh Naderi)、斯文·高德尔(Swen Gaudl)、马克·卡特赖特(Mark Cartwright)、艾米·K·胡佛(Amy K. Hoover)、马克·J·纳尔逊(Mark J. Nelson),2025年发表。论文题目:《开放文本转音频(text-to-audio)模型的表现力范围表征》,收录于第21届人工智能与交互式数字娱乐AAAI会议论文集(Proceedings of the 21st AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment)。 数据集内容: fig1_samples.zip:包含图1示例对应的生成音频样本,共2条提示文本(prompt)、2个模型,每个模型生成100条音频样本。 thunder_samples.zip:包含贯穿全文的“thunder”(雷声)示例对应的生成音频样本,共1条提示文本(prompt)、2个模型,每个模型生成100条音频样本,为图2至图4的数据源。 esc50_samples.zip:针对ESC-50环境音频数据集(ESC-50 environmental audio dataset)中每个标签X,以提示文本(prompt)“Sound of X”生成的音频样本,共50条提示文本(prompt)、3个模型,每个模型生成100条音频样本,为图5至图6及表1的数据源。 generation_scripts.zip:用于从上述三款模型生成音频的Python脚本文件。



