遇见数据集

A Dataset for the Classification of Different Kurdish Dialects

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Kurdish is an Indo-Iranian language that is largely spoken by people of Kurdish descent in the countries of Turkey, Iraq, Iran, and Syria. It contains a number of regional dialects, the most common of which is Northern Kurdish (also known as Kurmanji or Badini), while Central Kurdish (also known as Sorani) is spoken in some regions of Iraq and Iran. A unique Kurdish dialect, Hawrami, often referred to as Gorani, is the primary language spoken in the Hawraman area, which spans portions of western Iran and northeastern Iraq. Despite the fact that the dialects have separate pronunciations, vocabularies, and certain grammatical distinctions, they share a common core. In spite of the difficulties, the Kurdish language continues to be an essential component of Kurdish identity and cultural legacy. It plays an essential role in the protection and promotion of the distinct cultural identity of the Kurdish people. The concepts of language and dialect recognition are intricately interconnected within the fields of linguistics and natural language processing. Having a good dataset for Kurdish dialect recognition improves identification and classification, natural language processing applications for Kurdish, preservation of Kurdish linguistic heritage, cultural insights, customized content and services for users, empowerment of local businesses, and a benchmark for evaluating dialect recognition systems. The presented dataset was gathered by numerous members of the University of Halabja's Computer Science Department's teaching staff over the course of several months. During each stage of the data collecting process, the established policies, procedures, and guidelines were adhered to. This included taking into consideration the ages as well as the genders of the speakers who were included in the dataset. The recordings are taken from a variety of TV programmes and TV interviews that were broadcast on Speda tv, NRT, and GK Sat. There were 2000 instances of the Sorani dialect, 2000 examples of the Badini dialect, and 2000 examples of the Hawrami dialect. The total duration of this dataset is 6000 s, and the duration of each sample is precisely one second. The dataset labeling procedure that has been suggested consists of two consecutive stages. Initially, it is necessary to categorize the distinct sounds of each dialect, namely Sorani, Badini, and Hawrami, into different directories. Following this, it is recommended that the files included inside these folders be systematically labeled from 1 to 2000, according to the prescribed scheme: for Sorani files, the labels should range from s1 to s2000; for Badini files, the labels should range from b1 to b2000; and for Hawrami files, the labels should range from h1 to h2000.

库尔德语属于印伊语族,主要为库尔德族裔在土耳其、伊拉克、伊朗及叙利亚四国使用。其拥有诸多区域方言,最通用的为北库尔德语(Northern Kurdish),亦称库尔曼吉语(Kurmanji)或巴迪尼语(Badini);中库尔德语(Central Kurdish)亦称索拉尼语(Sorani),通行于伊拉克与伊朗的部分区域。另有独特的霍拉米方言(Hawrami,亦称戈拉尼语(Gorani)),是横跨伊朗西部与伊拉克东北部的霍拉曼地区的主要使用语言。尽管各方言在发音、词汇及部分语法规则上存在差异,但共享核心语言基础。尽管面临诸多挑战,库尔德语始终是库尔德族身份与文化遗产的核心组成部分,在保护与弘扬库尔德族独特文化身份方面发挥着至关重要的作用。 语言与方言识别的概念在语言学与自然语言处理(Natural Language Processing)领域中紧密交织。构建高质量的库尔德语方言识别数据集,有助于提升方言识别与分类效果、推动库尔德语相关自然语言处理应用发展、保护库尔德语言遗产、提供文化洞察、为用户定制个性化内容与服务、赋能本地商业,同时也可为方言识别系统的性能评估提供基准测试集。 本数据集由哈莱卜杰大学(University of Halabja)计算机科学系的多名教职员工历时数月采集完成。数据采集全流程严格遵循既定政策、流程与规范,充分考虑了受访发言者的年龄与性别分布。录制音频素材取自Speda TV、NRT及GK Sat播出的各类电视节目与电视访谈。数据集包含索拉尼方言样本2000条、巴迪尼方言样本2000条以及霍拉米方言样本2000条,总时长为6000秒,单条样本时长严格为1秒。 本数据集推荐的标注流程分为两个连续阶段。首先,需将索拉尼、巴迪尼、霍拉米三方言的语音样本分别归类至不同目录;随后,按照既定规则对各文件夹内的文件进行1至2000的系统编号标注:索拉尼文件标注格式为s1至s2000,巴迪尼文件为b1至b2000,霍拉米文件为h1至h2000。

创建时间:
2024-10-15
二维码
社区交流群
二维码
科研交流群
商业服务