遇见数据集

KuSarcasm: Automated Kurdish Sorani Sarcasm Dataset (KSSD)

收藏
Mendeley Data2026-04-09 收录
官方服务:

资源简介:

This study presents KuSarcasm, which is an automated Kurdish Sorani Sarcasm Dataset (KSSD). KuSarcasm is a comprehensive dataset developed for detecting sarcasm in Kurdish Sorani, a low-resource language with rich morphological complexity and limited Natural Language Processing (NLP) support. The dataset was constructed through a multi-stage data collection and annotation process guided by linguistic consultation and methodological rigor. Initial data was sourced from a wide range of Kurdish cultural materials, including proverbs, poems, and idiom texts extracted from Sekhurma Magazine, Digital publishing, and online repositories. Extensive consultations with Kurdish language experts, editors, and scholars were conducted to establish annotation rules and refine the dataset's cultural and contextual relevance. Data acquisition incorporated both manual and automated methods. Publicly available texts were extracted using Optical character recognition (OCR). Additionally, more data was gathered via web scraping, manually recording data to gather information, and structured queries, resulting in over 16,000 text entries. These texts were subjected to in-depth preprocessing pipeline, including deduplication, normalization, and noise reduction. Automatic annotating process was carried out using a custom hybrid method that cooperatively multilingual sentiment classification by Multilingual-Bidirectional Encoder Representations for Transformers (MBERT) and semantic similarity scoring with sentence-Bidirectional Encoder Representations for Transformers (sBERT). This rule guided annotation strategy relied on over 100 predefined linguistic patterns to distinguish sarcastic from non-sarcastic expressions based on both emotional polarity and semantic proximity. Eventually, KuSarcasm, is annotated for binary sarcasm classification and includes metadata such as source, matched rule, and sentiment category. Given its depth, diversity, and cultural alignment, KuSarcasm holds strong reuse potential for researchers working in NLP for underrepresented languages, sentiment analysis, and computational linguistics. It also offers a valuable foundation for developing and benchmarking deep learning models in low-resource settings.

本研究提出了KuSarcasm数据集,即自动化库尔德语索拉尼反讽数据集(Kurdish Sorani Sarcasm Dataset, KSSD)。KuSarcasm是专为库尔德语索拉尼——一种形态复杂度高且自然语言处理(Natural Language Processing, NLP)支持有限的低资源语言——的反讽检测任务构建的综合数据集。该数据集遵循语言学咨询与严谨的方法论,通过多阶段数据收集与标注流程构建而成。初始数据源自多类库尔德文化素材,包括从《Sekhurma Magazine》、数字出版资源及在线馆藏中提取的谚语、诗歌与习语文本。研究团队与库尔德语言专家、编辑及学者开展了广泛磋商,以制定标注规则并优化数据集的文化适配性与语境相关性。数据采集结合了手动与自动两种方式:利用光学字符识别(Optical Character Recognition, OCR)技术提取公开可用文本;此外还通过网络爬取、手动录入及结构化查询收集更多数据,最终获得逾16000条文本条目。这些文本经过了包含去重、归一化与降噪的深度预处理流程。自动标注流程采用定制化混合方法,结合多语言双向编码器表征(Multilingual-Bidirectional Encoder Representations for Transformers, MBERT)的多语言情感分类与句子双向编码器表征(sentence-Bidirectional Encoder Representations for Transformers, sBERT)的语义相似度评分完成。该规则导向的标注策略依托100余种预定义语言模式,基于情感极性与语义相似度区分反讽与非反讽表达。最终,KuSarcasm为二分类反讽标注数据集,包含来源、匹配规则与情感类别等元数据。鉴于其深度、多样性与文化适配性,KuSarcasm对于研究低资源语言自然语言处理、情感分析及计算语言学的研究者而言具备极高的复用价值,同时也为低资源场景下深度学习模型的开发与基准测试提供了宝贵的基础。

二维码
社区交流群
二维码
科研交流群
商业服务