遇见数据集

KTTS Single Speaker Dataset

收藏
Mendeley Data2024-01-01 更新2026-04-09 收录
官方服务:

资源简介:

This dataset has been primarily developed to facilitate the creation of text-to-speech systems for the Kashmiri language, a digitally underrepresented language predominantly spoken in the Jammu and Kashmir region of India. The dataset comprises 2,984 audio recordings in WAV format, each accompanied by its corresponding textual data in a separate file named ‘textcorpus.csv’. The ‘id’ column in the CSV file serves as a unique identifier, allowing users to efficiently locate the corresponding WAV files, which are systematically named according to the ‘id’ associated with the sentences they contain. All recordings feature a single male voice with a sample rate of 48,000 Hz, ensuring high-quality audio suitable for detailed phonetic analysis and machine learning applications. This consistent audio quality across the dataset provides a reliable foundation for training and testing text-to-speech models. Furthermore, the dataset can be a valuable resource for future research and development efforts aimed at enhancing digital accessibility for the Kashmiri-speaking population.

本数据集旨在为克什米尔语(Kashmiri)的文本转语音系统开发提供支撑,克什米尔语是一种数字化资源较为匮乏的语言,主要通行于印度查谟和克什米尔地区。数据集包含2984条WAV格式的音频录音,每条录音均配有对应的文本数据,这些文本数据存储于名为“textcorpus.csv”的独立文件中。该CSV文件中的“id”列作为唯一标识符,可帮助使用者快速定位对应的WAV音频文件,所有音频文件均根据其对应语句的“id”进行系统性命名。所有录音均采用单男声录制,采样率为48000Hz,可提供高质量音频,适用于精细语音分析与机器学习应用。数据集统一的音频质量为文本转语音模型的训练与测试提供了可靠基础。此外,本数据集还可作为宝贵的研究资源,助力未来提升克什米尔语使用者数字可及性的相关研发工作。

提供机构:
University of Kashmir
创建时间:
2024-01-01
二维码
社区交流群
二维码
科研交流群
商业服务