COUNTER
收藏资源简介:
COUNTER数据集由法国国家信息与自动化研究所创建,旨在解决在线极端内容检测的多语言复杂性问题。该数据集包含英语、法语和阿拉伯语的帖子,涵盖社交媒体、平台和论坛,涉及极端主义的多个层次和行动呼吁。数据集通过伪匿名化处理,保护个人隐私,同时保留上下文信息。创建过程中,数据集经过多重标注,并分析了标注过程中的偏差和分歧。COUNTER数据集主要应用于极端内容检测领域,旨在提高模型在多语言环境下的公平性和透明度。
The COUNTER dataset was developed by the French National Institute for Informatics and Automation (INRIA) to address the multilingual complexity inherent in online extreme content detection. This dataset contains posts in English, French, and Arabic, collected from social media, online platforms, and forums, covering multiple tiers of extremism and calls to action. Pseudo-anonymization techniques are applied to the dataset to safeguard personal privacy while preserving contextual information. During its curation, the dataset underwent multiple rounds of annotation, and biases and disagreements within the annotation process were analyzed. The COUNTER dataset is primarily applied in the field of extreme content detection, with the goal of improving the fairness and transparency of models in multilingual environments.




