DOI: http://dx.doi.org/10.17632/btfx5pw2rm.2#file-671c7b82-253e-465c-8844-b557e3b72b78
收藏资源简介:
SARA is a database, which consists of a set of spontaneous not pre-specified colloquial phrases in everyday life and life situations that are collected from media shows, episodes and films published on YouTube played by native speakers with three different Arabic dialects and accents. SARA's dialects are the Egyptian dialect (EGY), the Arabian Peninsula dialect (ARP) and the Levantine dialect (LEV). In addition, within each dialect group there are a number of different accents. Each utterance contains a one-speaker speech, the number of speakers within the dialect and the number of utterances by a speaker is unknown. This dataset can be used in speech, speaker, dialect and accent recognition applications. SARA dataset contains only adult speakers to avoid the improper pronunciation of the children that can affect the detection process. The dataset samples are variant in length from 3 to 7 seconds in order to verify the minimum time in which we can determine the speaker dialect or accent when the speaker speaks in free talk, which is the main scope of this research.
SARA数据集是一个语音数据库,其收录自YouTube平台发布的媒体节目、剧集与电影,其中包含由三种不同阿拉伯方言及口音的母语使用者发出的自发且未预设的日常口语短语与生活情境话语。SARA涵盖的方言分别为埃及方言(EGY)、阿拉伯半岛方言(ARP)与黎凡特方言(LEV),且每一方言组内还包含多种细分口音。每一条话语均为单发言人语音,方言组内的发言人总数及单发言人的话语条数均未明确。本数据集可应用于语音识别、说话人识别、方言识别及口音识别相关任务。SARA数据集仅收录成年发言人的语音样本,以规避儿童发音不规范可能对检测流程造成的负面影响。数据集样本的时长介于3至7秒不等,旨在验证在自由会话场景中,确定说话人方言或口音所需的最短时长,这也是本研究的核心目标。




