遇见数据集

harishkumard25/IndicCorpV2

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

IndicCorp v2是一个大规模单语语料库,专为印度语言设计,包含约209亿个标记,覆盖来自4个语言家族的24种语言。该数据集用于支持自然语言理解(NLU)任务,是ACL 2023论文《Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages》的核心组成部分,旨在提升印度语言的NLU能力,包括构建基准测试和多语言模型。

IndicCorp v2 is a large-scale monolingual corpus for Indic languages, containing approximately 20.9 billion tokens and covering 24 languages from 4 language families. It is designed to support Natural Language Understanding (NLU) tasks and is part of the ACL 2023 paper Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages, aiming to enhance NLU capabilities for Indic languages through benchmarks and multilingual models.

提供机构:
harishkumard25
二维码
社区交流群
二维码
科研交流群
商业服务