SOLID
收藏资源简介:
SOLID数据集是由卡塔尔计算研究机构创建的一个大型半监督数据集,专门用于识别攻击性语言。该数据集包含了超过九百万条英语推文,这些推文是通过更为系统的方法收集的,不同于以往使用关键词收集的方式。SOLID数据集的创建旨在解决现有数据集OLID的局限性,如大小限制和可能存在的偏见。通过SOLID与OLID的结合使用,可以显著提高在OLID测试集上的性能,尤其是在分类层次的较低级别。此外,SOLID数据集还用于SemEval共享任务OffensEval-2020,展示了其在实际应用中的价值。
The SOLID dataset is a large-scale semi-supervised dataset created by the Qatar Computing Research Institute, specifically designed for offensive language identification. It contains over nine million English tweets collected via a more systematic methodology, contrasting with prior keyword-driven collection approaches. The SOLID dataset was developed to address the limitations of the existing OLID dataset, such as its size constraints and potential biases. By combining SOLID with OLID, performance on the OLID test set can be significantly improved, especially at the lower tiers of the classification hierarchy. Furthermore, the SOLID dataset was utilized in the SemEval shared task OffensEval-2020, demonstrating its practical value.

- 1SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification卡塔尔计算研究机构 · 2021年



