phonsobon/khmer-word-segmentation
收藏官方服务:
资源简介:
该数据集包含高棉语的句子,反映了正式的行政和政府风格的写作。数据集是使用谷歌开发的Gemini大语言模型合成的,旨在模拟官方高棉文件语言,如报告、信件和机构通信。数据集分为训练集、验证集和测试集,每个文件每行包含一个句子。数据集用于高棉语自然语言处理的研究和实验,包括语言建模、文本生成以及语言模型的预训练或微调。
This dataset contains Khmer-language sentences that reflect formal administrative and government-style writing. The dataset was synthetically generated using the Gemini large language model, developed by Google, to simulate official Khmer document language such as reports, letters, and institutional communication. The dataset is split into train, validation, and test subsets, each containing one sentence per line. It is intended for research and experimentation in Khmer NLP, including language modeling, text generation, and pretraining or fine-tuning language models.
提供机构:
phonsobon


