官方服务:
资源简介:
Bible. Word-alligned corpus
应用场景:
相关数据集
LORELEI Vietnamese Representative Language Pack
Introduction LORELEI Vietnamese Representative Language Pack consists of Vietnamese monolingual text, Vietnamese-English parallel text, annotations, supplemental resources and related software tools d
DataCite Commons2020-08-17 更新120
OLDI Seed Corpus French Partition
OLDI Seed Corpus French Partition是一个法语分区,由Inria研究机构创建,旨在解决低资源语言翻译训练数据不足的问题。该数据集包含大约6000个英文句子,这些句子从维基百科的核心文章中抽取,涵盖了广泛的主题。为了创建这个数据集,使用了多种机器翻译系统和定制的后编辑界面,由母语为法语的专业人员进行后编辑。这个法语语料库不仅是翻译的终点,而且作为关键的中转资源,旨在促进
arXiv2025-08-04 更新80
Tattbabadhana
তত্ত্বাবধান (Tattbabadhana, means "supervise") -- supervised WSD for under-resourced languages
NIAID Data Ecosystem70
BR4-004 - Bislama stops
Pilot Bislama stop wordlist with LE. Language as given:
Research Data Australia120
bharatgenai/IndicParam
--- configs: - config_name: IndicParam data_files: - path: data* split: test tags: - benchmark - low-resource - indic-languages task_categories: - question-answering - text-classification lice
Hugging Face2025-12-10 更新80



