遇见数据集

audio_flamingo_description_4

收藏
Hugging Face2026-08-13 更新2026-08-15 收录
官方服务:

资源简介:

该数据集包含音频样本及其对应的标准化文本和描述信息。数据集仅包含训练集,共317个样本。每个样本包含三个字段:音频数据(audio)、标准化文本(normalized_text)以及描述(flamingo_next_description)。该数据集可用于语音识别、音频描述生成或多模态学习等任务。

This dataset contains audio samples along with their corresponding normalized text and description information. The dataset only includes a training set with a total of 317 samples. Each sample consists of three fields: audio data (audio), normalized text (normalized_text), and description (flamingo_next_description). The dataset can be used for tasks such as speech recognition, audio description generation, or multimodal learning.

提供机构:
NADSOFT
创建时间:
2026-08-13
原始信息汇总

数据集概述:nadsoft/audio_flamingo_description_4

该数据集包含音频数据及其对应的文本描述,主要用于音频相关的多模态模型训练或评估。

基本信息

数据特征(Features)

该数据集包含以下三个字段:

字段名 数据类型 描述
audio audio 音频数据(存储音频文件)
normalized_text string 标准化后的文本内容
flamingo_next_description string 与音频相关的 Flamingo 模型生成的描述文本

数据划分(Splits)

  • 训练集(train)
    • 样本数:317 条
    • 总字节数:约 76.04 MB

配置文件(Config)

  • 仅有一个默认配置(default),数据文件路径为 data/train-*,包含全部训练数据。

适用场景

该数据集适合用于音频理解、音频文本匹配、多模态模型(如 Flamingo)的训练或评测任务,尤其是需要音频与其语义描述配对的数据场景。

二维码
社区交流群
二维码
科研交流群
商业服务