遇见数据集

metuKKhud/bashqort-raw

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含用于大型语言模型(LLM)持续训练的巴什基尔语原始文本集合。它是项目“为巴什基尔语适应开源LLM”的一部分,旨在评估LlamaTurk(Toraman, 2024)和波斯语适应(Mahdizadeh Sani等人, 2024)提出的适应方法。语料库从多个来源汇编而成,为语言建模提供多样化的语言基础。来源包括bashgazet.ru(每日社会政治报纸)、neftcity.ru(地区新闻)、bash.news(新闻与分析)和bashkir-corpus(混合公共领域和打乱文本),总令牌数在1000万到1亿之间。预处理包括文档和句子级别的去重、移除少于5个词的句子、清理HTML伪影、广告和元数据。整个语料库用于自监督学习,无训练/验证分割。每个样本为JSON对象,包含文本、来源和是否打乱字段。数据集用于LLM的持续预训练、下一个令牌预测(因果语言建模)以及任何旨在改进巴什基尔语在NLP中表示的研究。许可证为MIT。

This dataset contains raw Bashkir text collected for continual training of large language models (LLMs). It is part of the project Adapting Open-Source LLMs for the Bashkir Language, which aims to evaluate adaptation methods proposed by LlamaTurk (Toraman, 2024) and Persian adaptation (Mahdizadeh Sani et al., 2024). The corpus is assembled from multiple sources to provide a diverse linguistic foundation for language modeling. Sources include bashgazet.ru (daily socio-political newspaper), neftcity.ru (regional news), bash.news (news & analytics), and bashkir-corpus (mixed public domain and shuffled text), with total tokens between 10M and 100M. Preprocessing includes deduplication (document and sentence level), removal of sentences with <5 words, removal of HTML artifacts, ads, and metadata. The whole corpus is for self-supervised learning, with no train/val split. Each example is a JSON object with text, source, and is_shuffled fields. Intended use includes continual pre-training of LLMs, next-token prediction (causal language modeling), and any research aiming to improve Bashkir language representation in NLP. Licensed under MIT.

提供机构:
metuKKhud
二维码
社区交流群
二维码
科研交流群
商业服务