LAHM
收藏资源简介:
LAHM数据集是由Logically.ai创建的一个大型多语言和多领域仇恨言论识别数据集,旨在解决社交媒体中仇恨言论的自动检测问题。该数据集包含近50万条推文,覆盖英语、印地语、阿拉伯语、法语、德语和西班牙语六种语言,并针对辱骂、种族主义、性别歧视、宗教仇恨和极端主义等多个领域进行详细标注。数据集的创建过程涉及使用特定关键词从社交媒体和新闻文章中收集数据,并通过多层级的标注流程确保数据的质量。LAHM数据集的应用领域广泛,包括但不限于社交媒体监控、内容审核和跨语言情感分析,旨在提高对仇恨言论的识别准确性和效率。
The LAHM dataset is a large-scale multilingual, multi-domain hate speech recognition dataset developed by Logically.ai, which aims to address the automatic detection of hate speech on social media. This dataset contains nearly 500,000 tweets covering six languages: English, Hindi, Arabic, French, German, and Spanish, with detailed annotations across multiple domains including abuse, racism, sexism, religious hatred, and extremism. The creation of the LAHM dataset involves collecting data from social media and news articles using specific keywords, and ensuring data quality through a multi-level annotation workflow. The LAHM dataset has a wide range of application scenarios, including but not limited to social media monitoring, content moderation, and cross-lingual sentiment analysis, with the goal of improving the accuracy and efficiency of hate speech recognition.

- 1LAHM : Large Annotated Dataset for Multi-Domain and Multilingual Hate Speech IdentificationLogically.ai · 2023年



