遇见数据集

edifier99/sinhala-openslr-111h

收藏
Hugging Face2026-05-24 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个音频-文本对数据集,专为语音处理任务设计,如自动语音识别(ASR)或音频标注。它包含178364个训练示例,每个示例由三个特征组成:audio(音频数据,采样率为16000Hz,确保高保真音频质量)、text(对应的转录文本,字符串格式,用于提供音频内容的语义信息)和duration(音频持续时间,以秒为单位的浮点数值,便于时间序列分析)。数据集总大小为约12.8GB,下载大小约12.7GB,适用于大规模机器学习模型训练。数据以train分割组织,文件路径为data/train-*,支持高效的数据加载和处理。该数据集可能用于开发语音技术应用,如语音助手、转录服务或音频分析工具。

This dataset is an audio-text pair dataset designed for speech processing tasks, such as automatic speech recognition (ASR) or audio annotation. It contains 178,364 training examples, each comprising three features: audio (audio data with a sampling rate of 16000Hz, ensuring high-fidelity audio quality), text (corresponding transcribed text in string format, providing semantic information of the audio content), and duration (audio duration in seconds as a floating-point value, facilitating time-series analysis). The total dataset size is approximately 12.8GB, with a download size of about 12.7GB, making it suitable for large-scale machine learning model training. The data is organized into a train split with file paths as data/train-*, enabling efficient data loading and processing. This dataset can be used for developing speech technology applications, such as voice assistants, transcription services, or audio analysis tools.

提供机构:
edifier99
二维码
社区交流群
二维码
科研交流群
商业服务