audio_flamingo_description_4
收藏官方服务:
资源简介:
该数据集包含音频样本及其对应的标准化文本和描述信息。数据集仅包含训练集,共317个样本。每个样本包含三个字段:音频数据(audio)、标准化文本(normalized_text)以及描述(flamingo_next_description)。该数据集可用于语音识别、音频描述生成或多模态学习等任务。
This dataset contains audio samples along with their corresponding normalized text and description information. The dataset only includes a training set with a total of 317 samples. Each sample consists of three fields: audio data (audio), normalized text (normalized_text), and description (flamingo_next_description). The dataset can be used for tasks such as speech recognition, audio description generation, or multimodal learning.
提供机构:
NADSOFT创建时间:
2026-08-13
原始信息汇总
数据集概述:nadsoft/audio_flamingo_description_4
该数据集包含音频数据及其对应的文本描述,主要用于音频相关的多模态模型训练或评估。
基本信息
- 数据集名称:
nadsoft/audio_flamingo_description_4 - 访问地址:https://huggingface.co/datasets/nadsoft/audio_flamingo_description_4
- 数据集大小:总大小约 76.04 MB(下载大小约 75.22 MB)
- 数据格式:Hugging Face Dataset 格式,包含音频文件和文本字段
数据特征(Features)
该数据集包含以下三个字段:
| 字段名 | 数据类型 | 描述 |
|---|---|---|
audio |
audio | 音频数据(存储音频文件) |
normalized_text |
string | 标准化后的文本内容 |
flamingo_next_description |
string | 与音频相关的 Flamingo 模型生成的描述文本 |
数据划分(Splits)
- 训练集(train):
- 样本数:317 条
- 总字节数:约 76.04 MB
配置文件(Config)
- 仅有一个默认配置(
default),数据文件路径为data/train-*,包含全部训练数据。
适用场景
该数据集适合用于音频理解、音频文本匹配、多模态模型(如 Flamingo)的训练或评测任务,尤其是需要音频与其语义描述配对的数据场景。



