tibetan-audio-english-sentence-merged-lilgoose
收藏资源简介:
藏语音频-句子合并数据集是一个综合性的资源,包含藏语录音及其对应的文本转录。该数据集合并了三个高质量的藏语语音数据集,旨在为藏语语言处理任务(如自动语音识别ASR、文本到语音TTS和翻译)提供更大、更多样化的资源。数据集采用统一的格式,包含两个主要字段:`audio`(音频文件,格式为WAV、MP3或FLAC)和`sentence`(藏语文本转录,使用藏文Unicode编码)。数据集规模在1K到10K样本之间,适用于语言保存、教育、文化传承等领域。使用CC-BY-4.0许可,用户需遵守相应的许可要求。数据集存在一定的局限性,如音频质量不一、方言覆盖不均等,使用时需注意。
The Tibetan Audio-Sentence Consolidated Dataset is a comprehensive resource containing Tibetan audio recordings and their corresponding text transcriptions. This dataset consolidates three high-quality Tibetan speech datasets, aiming to provide a larger and more diverse resource for Tibetan language processing tasks such as automatic speech recognition (ASR), text-to-speech (TTS) and machine translation. The dataset adopts a unified format with two main fields: `audio` (audio files in WAV, MP3 or FLAC formats) and `sentence` (Tibetan text transcriptions encoded in Tibetan Unicode). The dataset contains between 1,000 and 10,000 samples, and is applicable to fields such as language preservation, education and cultural heritage. It is licensed under CC-BY-4.0, and users must comply with the corresponding license requirements. There are certain limitations in the dataset, such as inconsistent audio quality and uneven dialect coverage, which should be noted during use.




