diarizers-community/voxconverse
收藏资源简介:
VoxConverse是一个音频-视觉的说话人分离数据集,包含从YouTube视频中提取的多说话人语音片段。该数据集已经过预处理,使其兼容diarizers工具,用于微调pyannote分割模型。数据集的特征包括音频、时间戳开始、时间戳结束和说话人信息。数据集分为开发集和测试集,分别包含216和232个样本。数据集的总下载大小为7296384603字节,总数据集大小为7354283539字节。数据集的语言为英语,许可证为cc-by-4.0。
VoxConverse is an audio-visual speaker diarization dataset consisting of multi-speaker speech segments extracted from YouTube videos. This dataset has been preprocessed to be compatible with diarization tools, and is intended for fine-tuning pyannote segmentation models. The features of the dataset include audio, start timestamp, end timestamp, and speaker information. The dataset is divided into a development set and a test set, which contain 216 and 232 samples respectively. The total download size of the dataset is 7296384603 bytes, and the total dataset size is 7354283539 bytes. The language of the dataset is English, and its license is CC-BY-4.0.
数据集概述
数据集名称
- Voxconverse
数据集特征
- audio: 音频数据
- timestamps_start: 开始时间戳,数据类型为
float64 - timestamps_end: 结束时间戳,数据类型为
float64 - speakers: 说话人标识,数据类型为
string
数据集分割
- dev: 包含216个样本,总大小为2338411143字节
- test: 包含232个样本,总大小为5015872396字节
数据集大小
- 下载大小: 7296384603字节
- 数据集总大小: 7354283539字节
配置
- config_name: default
- data_files:
- dev: 路径为
data/dev-* - test: 路径为
data/test-*
- dev: 路径为
标签
- speaker diarization
- voice activity detection
许可证
- cc-by-4.0
语言
- en




