L3CubeHingCorpus
收藏资源简介:
L3CubeHingCorpus是由L3Cube实验室开发的一个大型印度语-英语混合语料库,包含5293万条句子和10.4亿个标签。该数据集主要用于训练和评估处理代码混合文本的模型,如HingBERT和Hing-FastText。数据集的创建过程涉及从Twitter等社交媒体平台收集真实的Hinglish文本,并通过无监督学习方法进行处理。该数据集主要应用于自然语言处理任务,特别是仇恨言论检测,旨在解决在多语言社区中有效识别和分类仇恨言论的问题。
L3CubeHingCorpus is a large Hindi-English code-mixed corpus developed by L3Cube Labs, containing 52.93 million sentences and 1.04 billion tags. This dataset is primarily used for training and evaluating models that handle code-mixed text, such as HingBERT and Hing-FastText. The creation of this dataset involves collecting authentic Hinglish texts from social media platforms including Twitter, and processing them via unsupervised learning methods. This dataset is mainly applied to natural language processing tasks, particularly hate speech detection, aiming to address the challenge of effectively identifying and classifying hate speech in multilingual communities.

- 1On Importance of Code-Mixed Embeddings for Hate Speech IdentificationL3Cube实验室 · 2024年



