appledora/DANGA-Adapted
收藏资源简介:
DANGA-Adapted数据集是Bangla Dataset on Aggressive Narratives and Group-based Attacks (BanDANGA)的扩展版本,包含阿拉伯语、英语和中文的高质量翻译。该数据集旨在支持前沿大型语言模型(LLM)的跨语言安全评估,以及不同文化背景下相同暴力结构的比较研究。数据集由专家标注,采用细粒度的正交分类法,包含暴力、仇恨和严重冒犯性语言(孟加拉语),专门用于仇恨言论检测、内容审核和自然语言处理等研究目的。通过使用Adaption Lab的Blueprint功能,研究团队实现了对原始数据集的跨地区本地化改写,保留了暴力结构(身份轴、表达类型、修辞强度)的同时,针对不同地理政治区域(孟加拉国、印度、美国、中国)进行了适应性调整。
DANGA-Adapted is an extension of the Bangla Dataset on Aggressive Narratives and Group-based Attacks (BanDANGA), featuring high-quality translations into Arabic, English, and Chinese. This multilingual dataset enables cross-lingual safety evaluation of frontier LLMs and comparative research on how identical rhetorical patterns of communal violence are recognized across different cultural contexts. Expert-annotated with a fine-grained orthogonal taxonomy, it contains violent, hateful, and severely offensive language in Bengali, intended solely for research purposes (hate speech detection, content moderation, NLP). Using Adaption Labs Blueprint feature, the team localized original Bengali hate speech patterns to different geopolitical regions (Bangladesh, India, USA, China) while preserving the violence structure (identity axis, expression type, rhetorical intensity).




