SINAI/ALIA-es-discriminative-stance-detection
收藏资源简介:
ALIA西班牙语判别式立场检测语料库是一个手动标注的数据集,专为训练和评估西班牙语立场检测模型而设计。给定一个公民话题(目标)和一条公民评论,任务是将评论对该话题的立场分类为三种类别之一:支持、反对或中立。该语料库包含3000个手动标注的实例,来源于Decide Madrid参与式民主平台上的真实公民评论。数据集采用跨话题设计,覆盖387个不同话题,涉及城市交通、环境政策、公共服务等多个地方治理领域。每个实例由3个独立的人类标注者通过众包平台标注,并遵循数据透视主义原则发布所有个体标注,而不是将标注合并为单一真实标签并丢弃分歧,从而支持标注者分歧、模糊性建模和多视角立场分类的研究。
The ALIA Spanish Discriminative Stance Detection Corpus is a manually annotated dataset designed for training and evaluating stance detection models in Spanish. Given a civic topic (target) and a citizen comment, the task is to classify the comments stance toward the topic into one of three categories: favor, against, or neutral. The corpus comprises 3,000 manually annotated instances for stance detection in Spanish, built from real citizen comments posted on the Decide Madrid participatory democracy platform. It is cross-topic by design, spanning 387 distinct topics covering various areas of local governance. Each instance was annotated by 3 independent human annotators, and the dataset is published in full accordance with the principles of data perspectivism, releasing all individual judgments to enable research on annotator disagreement, ambiguity modeling, and multi-perspective stance classification.




