CKTN
收藏资源简介:
CKTN是由越南温大学和胡志明市理工大学等机构联合构建的首个面向越南少数民族语言(占族语、高棉语和岱侬语)的多语种语料库与基准测试集。该数据集共收录44,367篇文档,包含超过2400万子词标记,数据主要来源于越南之声等政府及地方新闻媒体的持续更新的网络文本。语料库通过自动化爬取、文本清洗和元数据提取流程构建,涵盖持续预训练、28类别分类和摘要-文档检索三大任务场景,旨在解决低资源语言在自然语言处理中因文字系统差异和标准化不足导致的语义泛化难题,为跨脚本语言建模研究提供关键基础设施。
CKTN is the first multilingual corpus and benchmark dataset targeting Vietnamese ethnic minority languages (Cham, Khmer, and Tay-Nung), jointly constructed by institutions including Vietnam Wen University and Ho Chi Minh City University of Technology. This dataset contains 44,367 documents with over 24 million subword tokens. The data is primarily sourced from continuously updated web texts from government and local news media such as Voice of Vietnam. The corpus is constructed via automated crawling, text cleaning, and metadata extraction workflows. It covers three core task scenarios: continuous pre-training, 28-category classification, and summary-document retrieval. This work aims to address the semantic generalization challenges faced by low-resource languages in natural language processing (NLP) due to disparities in writing systems and inadequate standardization, providing critical infrastructure for cross-script language modeling research.




