StanceNakba 2026 dataset
收藏资源简介:
StanceNakba 2026数据集由卡塔尔西北大学与哈马德·本·哈利法大学联合构建,专为巴以冲突背景下的立场检测研究设计,包含2,606条经过标注的英文与阿拉伯文社交媒体帖子。该数据集内容涵盖两个子任务:子任务A包含1,401条英文帖子,通过梅尔沃特媒体平台采集并基于关键词查询进行平衡采样,每条数据均经过去除噪声字符、长度过滤及去重等预处理;子任务B则聚焦阿拉伯文帖子,针对“与以色列关系正常化”和“约旦难民问题”两个冲突相关主题进行标注。数据集通过系统化的数据收集、清洗与标注流程构建,旨在为政治话语分析与跨语言立场检测模型提供基准资源,推动冲突领域自然语言处理技术发展。
The StanceNakba 2026 dataset was jointly constructed by Northwestern University in Qatar and Hamad Bin Khalifa University, specifically designed for stance detection research in the context of the Israeli-Palestinian conflict. It comprises 2,606 annotated English and Arabic social media posts, covering two subtasks: Subtask A includes 1,401 English posts, collected via the Melwote Media Platform and balancedly sampled through keyword queries, with each data entry undergoing preprocessing steps such as noise character removal, length filtering and deduplication. Subtask B focuses on Arabic posts, annotated for two conflict-related topics: "normalization of relations with Israel" and "Jordanian refugee issue". Constructed via a systematic workflow of data collection, cleaning and annotation, this dataset aims to provide benchmark resources for political discourse analysis and cross-lingual stance detection models, and promote the development of natural language processing technologies in the conflict domain.

- 1StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse卡塔尔西北大学; 哈马德·本·哈利法大学 · 2026年



