SINAI/ALIA-es-discriminative-hate-speech
收藏资源简介:
ALIA西班牙歧视性仇恨言论语料库是一个大规模西班牙语数据集,用于仇恨言论检测,基于精选的社交媒体评论构建,并通过多专家LLM管道与专家融合(FoE)技术自动标注。该数据集包含228,708个实例,来源于YouTube和TikTok的西班牙语评论。每个实例都保留了三个LLM专家的预测和解释,以及最终的融合输出(包括二元分类标签和置信度分数)。数据集旨在支持西班牙语仇恨言论检测的研究,包括模型分歧分析、基于解释的审核和置信度评估。数据集的构建过程分为三个阶段:原始社交媒体评论收集、西班牙语评论的筛选和过滤,以及通过三个提示的LLM专家进行自动标注和FoE融合。所有文档和源代码可在ALIA-UJA GitHub仓库中获取。
The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE). The release contains 228,708 instances of Spanish comments from YouTube and TikTok, with per-expert predictions and explanations from three LLM experts, and final fused outputs (foe_class, foe_score) for discriminative hate-speech classification. The corpus is intended for research on Spanish hate-speech detection, model disagreement analysis, explanation-aware moderation, and confidence-based evaluation. All documentation and source code for the collection, curation, annotation and fusion processes are available in the ALIA-UJA GitHub repository.




