官方服务:
资源简介:
Handwritten Tulu Characters Dataset
手写图卢语字符数据集
应用场景:
创建时间:
2024-05-08
相关数据集
yodas_owsmv4
该数据集包含了跨越75种语言的166,000小时的多种语言语音,被分割成30秒的长格式音频片段。数据来源于YODAS2数据集,该数据集基于大规模的网络爬取内容。由于网络源数据的性质,原始的YODAS2数据集可能包含不准确的语言标签和音频-文本对不齐的情况。为了解决这一问题,我们开发了一个可扩展的数据清洗管道,使用公开可用的工具包,从而形成原始数据集的一个精选子集。这个清洗后的数据集是我们OWSM
Hugging Face2025-06-03 更新320
中英混读手机采集语音数据【数据堂】
1,535小时中英混读手机采集语音数据由3972名中国本土人员参与录制,口音覆盖七大方言区。录音文本均为中英混合句子,涵盖通用场景及人机交互场景,内容丰富,转写精准。可用于改善语音识别系统对中英混读语音的识别效果。
OpenDataLab2023-12-20 更新140
Audiovisual speech perception in a predictive framework
In language comprehension, a variety of contextual cues act in unison to render upcoming words more or less predictable. As a sentence unfolds, we use prior context (sentential constraints) to predict
DataCite Commons2024-05-13 更新100
Indonesian Sign Language System (SIBI) Dataset
The SIBI Dataset is a collection of Indonesian Sign Language System (SIBI) video signs based on the SIBI online dictionary issued by the Ministry of Education and Culture of the Republic of Indonesia.
DataCite Commons2025-04-01 更新160
Australian Radio Talkback Corpus (ART)
Australian Radio Talkback (ART) is a set of transcribed recordings of samples of national, regional and commercial Australian talkback radio from 2004 to 2006. It consists of transcriptions of 27 audi
figshare.mq.edu.au2024-04-24 更新120



