TextrolMix
收藏资源简介:
TextrolMix是由亚马逊Prime Video团队创建的一个用于目标语音提取(TSE)的数据集,包含12万条双人语音混合数据,总计157小时。每条数据包含目标语音的自然语言描述和参考音频线索,支持灵活的文本引导TSE模型。数据集通过增强TextrolSpeech数据集生成,每条语音混合数据包含六种属性:说话者身份、情感、音高、性别、口音和语速。数据集的设计使得模型能够基于细微的属性差异提取目标语音,而无需依赖显著不同的整体说话风格。TextrolMix数据集的应用领域主要集中在语音分离和目标语音提取,旨在解决传统TSE方法在缺乏明确说话者身份线索时的局限性。
TextrolMix is a Target Speech Extraction (TSE) dataset developed by the Amazon Prime Video team. It consists of 120,000 two-speaker mixed speech samples, with a total duration of 157 hours. Each sample includes both a natural language description of the target speech and reference audio cues, enabling flexible text-guided TSE models. This dataset is generated by augmenting the existing TextrolSpeech dataset, and every mixed speech sample contains six attributes: speaker identity, emotion, pitch, gender, accent, and speech rate. The design of TextrolMix allows models to extract target speech based on subtle attribute differences, without relying on significantly distinct overall speaking styles. The main application fields of TextrolMix focus on speech separation and target speech extraction, aiming to address the limitations of traditional TSE methods when explicit speaker identity cues are absent.




