遇见数据集

barockok/dataset-alpha-v1

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

Dataset Alpha V1 是一个多语言音频-文本平行语料库,从开源语音研究数据集编译而来。它包含多个子集(A、B、C),总计约12,956个样本和21.5小时的音频数据,主要覆盖南岛语系语言(如Austronesian语族,使用人数从约4000万到2亿不等)。数据格式包括24kHz立体声音频文件(WAV格式)、单词级时间戳对齐的JSON文件(包含每个单词的开始和结束时间及说话者信息),以及元数据清单(manifest.jsonl)。所有转录使用拉丁字母并带有语言特定的变音符号。数据集适用于语音识别(ASR)、文本到语音(TTS)等任务,并支持通过Hugging Face的datasets库加载。

Dataset Alpha V1 is a multilingual audio-text parallel corpus compiled from open-source phonetic research datasets. It includes multiple subsets (A, B, C) with a total of approximately 12,956 samples and 21.5 hours of audio data, primarily covering Austronesian language family languages (e.g., with speaker populations ranging from about 40 million to 200 million). The data format consists of 24kHz stereo audio files (WAV format), word-level timestamp alignment JSON files (containing start and end times and speaker information for each word), and metadata manifests (manifest.jsonl). All transcriptions use the Latin script with language-appropriate diacritics. The dataset is suitable for tasks such as automatic speech recognition (ASR) and text-to-speech (TTS), and can be loaded via the Hugging Face datasets library.

提供机构:
barockok
二维码
社区交流群
二维码
科研交流群
商业服务