遇见数据集

jintz0/assamese-monolingual-corpus

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

阿萨姆语单语语料库(2025)是一个高质量、句子级别的阿萨姆语单语数据集,包含161.3万个经过清理、分割和去重的句子,使用孟加拉文脚本。该语料库支持阿萨姆语自然语言处理开发、语言建模,并为印度东北部地区提供公共部署。具体细节包括:语言为阿萨姆语(孟加拉文脚本),大小为1,613,879个句子,格式为纯文本CSV(text列),总标记数为77,427,585(使用IndicBERTv2分词器),许可证为CC BY-SA 4.0(需署名),来源包括IITB-IndicMonoDoc、Samanantar(阿萨姆语部分)以及阿萨姆语诗歌和公民文本。

Assamese Monolingual Corpus (2025) is a high-quality, sentence-level Assamese monolingual dataset containing 1.61 million cleaned, segmented, and deduplicated sentences in Bengali script. This corpus supports Assamese NLP development, language modeling, and public deployment for Northeast India. Details include: Language is Assamese (Bengali script), size is 1,613,879 sentences, format is plain text CSV (text column), total tokens are 77,427,585 (using IndicBERTv2 tokenizer), license is CC BY-SA 4.0 (attribution required), and sources include IITB-IndicMonoDoc, Samanantar (Assamese side), and Assamese poetry and civic texts.

提供机构:
jintz0
二维码
社区交流群
二维码
科研交流群
商业服务