登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
VoxLingua-107 Vector-Embedding and Labels.
VoxLingua-107 Vector-Embedding and Labels.
收藏
kaggle
2024-11-13 更新
2024-12-28 收录
语音识别
多语言处理
数据链接:
https://www.kaggle.com/datasets/neeraj8180/voxlingua-vector-embedding-and-labels
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
Language Diarization
语言分离技术
应用场景:
创建时间:
2024-11-13
相关数据集
glenn2/voice_breaths
语音识别
音频分析
--- dataset_info: features: - name: audio dtype: audio: sampling_rate: 48000 - name: sentence dtype: string splits: - name: train num_bytes: 55098221109.05 num_
Hugging Face
2024-04-13 更新
28
0
BadBoy17G/TestingDataset
语音识别
--- license: apache-2.0 dataset_info: features: - name: audio dtype: audio - name: sentence dtype: string splits: - name: train num_bytes: 20497932.0 num_examples: 253 down
Hugging Face
2024-01-20 更新
23
0
minato-ryan/emilia-cjk-xcodec2
语音识别
日语处理
该数据集包含多个配置(ja_long,ja_short,ja_skipped,ja_very_long),每个配置都有如uid、dnsmos、时长、语言、说话人、文本、wav文件和ids等特征。每个配置的训练分割都包含了示例数量和字节数的信息。此外,还提供了每个配置的下载大小和数据集大小。
Hugging Face
2025-10-24 更新
11
0
filipinospeechcorpus
语音识别
语音合成
菲律宾语音语料库(Filipino Speech Corpus, FSC)是一个用于自动语音识别(ASR)和文本转语音(TTS)任务的句子级语音数据集。该数据集以Hugging Face Parquet格式存储,包含16kHz单声道音频片段。数据集总大小约为6GB,由139个分片组成。每个样本包含音频文件、对应的文本转录、音频时长、单词数量、说话者ID、性别、年龄组、语音类型(朗读/自发/机器生成
Hugging Face
2026-03-22 更新
24
0
kotoba-speech/common_voice_17_0
语音识别
多语种语音数据
--- dataset_info: config_name: ja features: - name: client_id dtype: string - name: path dtype: string - name: audio dtype: audio: decode: false - name: text
Hugging Face
2025-02-15 更新
12
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广