Xamdi-Speech
收藏资源简介:
该数据集是一个包含音频-文本配对样本的小规模数据集,专为需要音频与文本关联的任务设计。数据集包含14个训练样本,每个样本由两个核心字段构成:文本字段(text)存储文本内容,音频字段(audio)存储对应的音频数据,音频采样率为24000Hz。数据以结构化格式存储,适用于语音合成、语音识别、音频字幕生成或音频-文本对齐等机器学习任务的研究与开发。
This dataset is a small-scale collection of audio-text paired samples, specifically designed for tasks requiring audio-text association. It contains 14 training samples, each of which consists of two core fields: the "text" field that stores the textual content, and the "audio" field that stores the corresponding audio data, with an audio sampling rate of 24000 Hz. The data is stored in a structured format, and is applicable to the research and development of machine learning tasks such as speech synthesis, speech recognition, audio caption generation, and audio-text alignment.
数据集概述
名称:Xamdi-Speech
来源:由 lakiLabs 提供,托管于 Hugging Face 数据集平台。
数据格式:
- 包含两个字段:
- text:字符串类型,表示语音对应的文本内容。
- audio:音频类型,采样率为 24000 Hz。
数据集划分:
- 仅包含一个训练集(train),共有 14 个样本,总字节数约为 1,350,018 字节。
下载信息:
- 下载大小约为 1,351,761 字节。




