harishkumard25/IndicCorpV2
收藏官方服务:
资源简介:
IndicCorp v2是一个大规模单语语料库,专为印度语言设计,包含约209亿个标记,覆盖来自4个语言家族的24种语言。该数据集用于支持自然语言理解(NLU)任务,是ACL 2023论文《Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages》的核心组成部分,旨在提升印度语言的NLU能力,包括构建基准测试和多语言模型。
IndicCorp v2 is a large-scale monolingual corpus for Indic languages, containing approximately 20.9 billion tokens and covering 24 languages from 4 language families. It is designed to support Natural Language Understanding (NLU) tasks and is part of the ACL 2023 paper Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages, aiming to enhance NLU capabilities for Indic languages through benchmarks and multilingual models.
提供机构:
harishkumard25


