AF-Synthetic
收藏资源简介:
AF-Synthetic是由英伟达研究团队创建的大规模高质量合成字幕数据集,旨在提升文本到音频生成模型的表现。该数据集包含135万条字幕,通过音频理解模型生成,并经过严格的CLAP相似度过滤,确保字幕与音频内容高度相关。数据集的创建过程涉及对多个公开音频数据集的整合与优化,最终生成了具有强音频相关性的合成字幕。AF-Synthetic主要应用于文本到音频生成领域,旨在解决现有数据集规模小、字幕质量参差不齐的问题,为模型训练提供更丰富、更高质量的数据支持。
AF-Synthetic is a large-scale high-quality synthesized subtitle dataset developed by the NVIDIA Research team, aiming to improve the performance of text-to-audio generation models. This dataset contains 1.35 million subtitle entries, which are generated by audio understanding models and filtered through strict CLAP similarity criteria to ensure high relevance between the subtitles and their corresponding audio content. The construction process of AF-Synthetic involves integrating and optimizing multiple public audio datasets, ultimately yielding synthesized subtitles with strong audio relevance. Primarily applied in the text-to-audio generation domain, AF-Synthetic aims to address the limitations of existing datasets, including small scale and uneven subtitle quality, thereby providing richer and higher-quality data support for model training.




