thetemirbolatov/Karachay-words
收藏资源简介:
该数据集是一个卡拉恰伊语单词和短语的基本集合,专门用于训练和微调突厥语系的语言模型,特别是卡拉恰伊-巴尔卡尔语(ISO 639-3语言代码:krc)。它旨在弥补低资源语言数据的严重短缺,适用于从头开始训练、微调现有模型或扩展分词器。数据集格式为JSON Lines(.jsonl),每条记录是一个有效的JSON对象,包含一个word字段,存储卡拉恰伊语单词或短短语。内容涵盖卡拉恰伊-巴尔卡尔语的本土词汇、俄语借词(适应拼写)以及固定表达和习语。数据集包含136,706行(单词),文件大小约4.6 MB。
This repository contains a basic set of Karachay words and phrases, designed for training and fine-tuning language models of the Turkic group, specifically the Karachay-Balkar language (ISO 639-3 code: krc). The dataset is created to address the acute shortage of low-resource language data and is suitable for training from scratch, fine-tuning existing models, or expanding tokenizers. The format is JSON Lines (.jsonl), with each line being a valid JSON object containing a word field for Karachay words or short phrases. The content includes native Karachay-Balkar vocabulary, modern borrowings from Russian adapted in spelling, and stable expressions or idioms. The dataset consists of 136,706 lines (words), with a file size of approximately 4.6 MB.




