dravidianlangtech/hope_edi
收藏资源简介:
HopeEDI数据集是一个用于平等、多样性和包容性的希望言论检测数据集,包含从社交媒体平台YouTube收集的用户生成评论,分别有28,451条英语评论、20,198条泰米尔语评论和10,705条马拉雅拉姆语评论,这些评论被手动标注为包含希望言论或不包含希望言论。据我们所知,这是第一个在多语言环境中为平等、多样性和包容性标注希望言论的研究。
The HopeEDI dataset is a hope speech detection dataset designed for equality, diversity and inclusion (EDI). It comprises user-generated comments collected from the social media platform YouTube, with 28,451 English comments, 20,198 Tamil comments, and 10,705 Malayalam comments respectively. These comments have been manually annotated as either containing hope speech or not. To the best of our knowledge, this is the first study focused on annotating hope speech for equality, diversity and inclusion in a multilingual setting.
数据集概述
数据集基本信息
- 名称: HopeEDI
- 语言: 英语、马拉雅拉姆语、泰米尔语
- 许可证: CC-BY-4.0
- 多语言性: 单语和多语
- 大小类别: 10K<n<100K, 1K<n<10K
- 源数据: 原始数据
- 任务类别: 文本分类
- 标签: hope-speech-classification
数据集配置
-
英语:
- 特征:
text: 字符串label: 类别标签,可能值为 "Hope_speech", "Non_hope_speech", "not-English"
- 分割:
train: 22762个样本, 2306656字节validation: 2843个样本, 288663字节
- 下载大小: 2739901字节
- 数据集大小: 2595319字节
- 特征:
-
泰米尔语:
- 特征:
text: 字符串label: 类别标签,可能值为 "Hope_speech", "Non_hope_speech", "not-Tamil"
- 分割:
train: 16160个样本, 1531013字节validation: 2018个样本, 197378字节
- 下载大小: 1795767字节
- 数据集大小: 1728391字节
- 特征:
-
马拉雅拉姆语:
- 特征:
text: 字符串label: 类别标签,可能值为 "Hope_speech", "Non_hope_speech", "not-malayalam"
- 分割:
train: 8564个样本, 1492031字节validation: 1070个样本, 180713字节
- 下载大小: 1721534字节
- 数据集大小: 1672744字节
- 特征:
数据集示例
-
英语:
text: "all lives matter .without that we never have peace so to me forever all lives matter."label: "Hope_speech"
-
泰米尔语:
text: "Idha solla ivalo naala"label: "Non_hope_speech"
-
马拉雅拉姆语:
text: "ഇത്രെയും കഷ്ടപ്പെട്ട് വളർത്തിയ ആ അമ്മയുടെ മുഖം കണ്ടപ്പോൾ കണ്ണ് നിറഞ്ഞു പോയി"label: "Hope_speech"
数据集创建
- 标注创建者: 专家生成
- 语言创建者: 众包
- 标注过程: 使用Google表单收集标注,每个表单最多包含100条评论,每页最多10条评论。标注者包括来自澳大利亚、爱尔兰、英国和美国的英语标注者,以及来自印度泰米尔纳德邦和斯里兰卡的泰米尔语标注者。
数据集使用注意事项
- 个人和敏感信息: 数据集包含来自社交媒体的高度敏感信息,已采取措施最小化个人身份信息的风险,但保留了与种族、性别、性取向、民族起源和哲学信仰相关的信息。
附加信息
-
许可证信息: Creative Commons Attribution 4.0 International Licence
-
引用信息:
@inproceedings{chakravarthi-2020-hopeedi, title = "{H}ope{EDI}: A Multilingual Hope Speech Detection Dataset for Equality, Diversity, and Inclusion", author = "Chakravarthi, Bharathi Raja", booktitle = "Proceedings of the Third Workshop on Computational Modeling of Peoples Opinions, Personality, and Emotions in Social Media", month = dec, year = "2020", address = "Barcelona, Spain (Online)", publisher = "Association for Computational Linguistics", url = "https://www.aclweb.org/anthology/2020.peoples-1.5", pages = "41--53", }




