uyghur-text
收藏资源简介:
该数据集名为“Uyghur AI Corpus: Bridging Heritage & Technology”(维吾尔语:ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى),是一个专门为维吾尔语自然语言处理任务构建的高质量语料库。其目标是在人工智能时代保障维吾尔语在数字世界中的生存与发展,为训练大语言模型(LLM)提供基础资源,使其能够理解、生成和翻译维吾尔语。数据集适用于文本生成、翻译和掩码填充等任务。数据集的内容来源于公开互联网上的多样文本,包括故事、文章、历史记录等,经过精心筛选、清洗和结构化处理,以满足高质量 AI 训练标准。在可能的情况下,保留了作者和来源等元数据以尊重原始创作者。数据以 Parquet 格式存储,包含训练集,但未提供具体样本数量。该数据集使用 MIT 许可证,语言为维吾尔语(ug)。
The dataset is named Uyghur AI Corpus: Bridging Heritage & Technology (Uyghur: ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى). It is a high-quality corpus specifically constructed for Uyghur natural language processing tasks. Its goal is to ensure the survival and development of the Uyghur language in the digital world during the AI era, providing foundational resources for training large language models (LLMs) to understand, generate, and translate Uyghur. The dataset is suitable for tasks such as text generation, translation, and masked language modeling. The content is sourced from diverse publicly available internet texts, including stories, articles, historical records, etc., which have been carefully selected, cleaned, and structured to meet high-quality AI training standards. Where possible, metadata such as author and source are retained to respect original creators. The data is stored in Parquet format, containing a training set, but the specific number of samples is not provided. The dataset is licensed under MIT, and the language is Uyghur (ug).
维吾尔文AI语料库 (Uyghur AI Corpus)
数据集概述
该数据集是一个为人工智能训练优化的维吾尔语文本语料库,旨在帮助大语言模型(LLMs)理解、生成和翻译维吾尔语,达到母语级熟练度。项目旨在确保维吾尔语在数字时代持续发展,并通过AI技术保护和传承维吾尔文化遗产。
基本信息
- 语言: 维吾尔语 (ug)
- 许可证: MIT
- 任务类别: 文本生成、翻译、填充掩码
- 标签: 维吾尔语、语料库、自然语言处理、维吾尔语数据集、维吾尔语AI
- 数据格式: Parquet文件
- 数据划分: 训练集(train)
数据来源与构建
- 来源: 开放互联网上公开可获取的各类平台、论坛和网站
- 内容类型: 故事、散文、文章、历史记载等公共领域的文本
- 处理方式: 原始网页数据经过精细清洗、格式化,达到高质量AI训练标准
- 版权保护: 尽可能保留原始作者和来源的元数据,尊重知识创作者权益
项目特点
- 文化使命: 定位为连接维吾尔文化遗产与人工智能技术的桥梁
- 维护状态: 活跃维护中
- 项目用途: 作为训练大语言模型的基础资源,使模型能够以母语级水平处理维吾尔语
配置信息
- 配置名称: default
- 数据文件路径: *.parquet (训练集)




