Ultra-FineWeb-L3-zh-hant-translated
收藏资源简介:
Ultra-FineWeb-L3(繁体中文翻译版)是英文Ultra-FineWeb-L3语料库的完整繁体中文翻译,目标语言为台湾标准繁体中文(臺灣正體中文)。源数据集Ultra-FineWeb-L3是OpenBMB开发的UltraData数据管理框架中的L3精炼层,它从一个高质量的网络语料库(Ultra-FineWeb)出发,通过L3精炼将原始网络文档转换为两种结构化合成格式:问答对生成和多风格改写。本翻译数据集旨在为训练或微调语言模型提供高知识密度和结构多样性的繁体中文文本。数据集包含两个子集:1. multistyle(多风格改写)子集:包含2,495,221条示例,每条示例将网络文档改写成四种风格之一:百科全书风格(类似维基百科,中立第三人称,分节,客观)、教科书风格(形式逻辑递进:定义→原理→示例→总结)、博客风格(对话式短段落,类比,个人语气,150–350字)和摘要风格(高度压缩的学术风格散文,80–150字)。该子集还包含少量从繁体中文本地语料库(TMMLU+)合成的博客风格记录,与翻译记录全局混洗。2. qa(问答对生成)子集:包含1,194,156条示例,每条示例包含原始文档和多个结构化问答对,问答对涵盖五种问题类型:概念理解、因果分析、比较/对比、推理/应用和综合/总结。两个子集共享相同的扁平数据模式,包含两个字段:uid(字符串,继承自源数据的唯一文档标识符)和text(字符串,翻译/合成的繁体中文文本)。翻译使用gemma-4-27b-A3B模型完成,保留了原始格式(如Markdown、代码块、数学公式、URL、标识符),并强制使用台湾词汇规范(例如使用“軟體”、“程式”、“資訊”,而非大陆用词“软件”、“程序”、“信息”)。数据集按原样提供,翻译为机器生成,未经人工验证,质量可能因文档而异,用户应根据自身用例进行质量过滤。
Ultra-FineWeb-L3 (Traditional Chinese Translation) is a complete traditional Chinese translation of the English Ultra-FineWeb-L3 corpus, targeting Taiwanese Standard Traditional Chinese (臺灣正體中文). The source dataset, Ultra-FineWeb-L3, is part of the L3 refinement layer in the UltraData data management framework developed by OpenBMB. It starts from a high-quality web corpus (Ultra-FineWeb) and uses L3 refinement to transform raw web documents into two structured synthetic formats: QA pair generation and multi-style rewriting. This translation dataset aims to provide traditional Chinese text with high knowledge density and structural diversity for training or fine-tuning language models. The dataset includes two subsets: 1. multistyle (multi-style rewriting) subset: contains 2,495,221 examples, each rewriting a web document into one of four styles: encyclopedia style (similar to Wikipedia, neutral third-person, sectioned, objective), textbook style (formal logical progression: definition → principle → example → summary), blog style (conversational short paragraphs, analogies, personal tone, 150–350 words), and summary style (highly compressed academic prose, 80–150 words). This subset also includes a small number of blog-style records synthesized from a traditional Chinese local corpus (TMMLU+), globally shuffled with translated records. 2. qa (QA pair generation) subset: contains 1,194,156 examples, each including the original document and multiple structured QA pairs. The QA pairs cover five question types: concept understanding, causal analysis, comparison/contrast, reasoning/application, and synthesis/summary. Both subsets share the same flat data schema with two fields: uid (string, unique document identifier inherited from the source data) and text (string, translated/synthesized traditional Chinese text). Translation was performed using the gemma-4-27b-A3B model, preserving original formats (e.g., Markdown, code blocks, mathematical formulas, URLs, identifiers) and enforcing Taiwanese vocabulary norms (e.g., using 軟體, 程式, 資訊 instead of Mainland Chinese terms like 软件, 程序, 信息). The dataset is provided as-is; the translation is machine-generated, not manually verified, and quality may vary across documents, so users should perform quality filtering based on their use case.




