Fineweb-Edu-Chinese-V2.3
收藏资源简介:
# Chinese Fineweb Edu Dataset V2.3 <div align="center"> <a href="#chinese">中文</a> | <a href="#english">English</a> </div> <div align="center"> <img width="600px" alt="OpenCSG" src="./logo.png"> [OpenCSG 社区](https://opencsg.com/models) | [GitHub](https://github.com/yuyijiong/fineweb-edu-chinese) | [数据集许可协议](./OpenCSG数据集许可协议.md) </div> <a id="chinese"></a> ## 数据集简介 **Chinese Fineweb Edu Dataset V2.3** 是 OpenCSG 面向中文教育、知识问答、指令微调和文本生成场景构建的高质量中文教育 SFT 数据集。 该版本包含 **23.04 万条高质量中文教育 QA pairs**,并将同一批问答对发布为 **Alpaca**、**Messages**、**Messages-no-system** 三种训练格式。三种格式面向不同训练模板,建议训练时按模型和框架选择其中一种格式使用,而不是将不同格式简单相加作为独立知识规模。 这里仅发布 **10% 抽样预览**,完整 **230,440 条 QA pairs** 请到 [CSGHub Fineweb-Edu-Chinese-V2.3 数据集页](https://opencsg.com/datasets/OpenCSG/Fineweb-Edu-Chinese-V2.3?tab=summary) 获取。 V2.3 是在 **V2.2** 基础上的质量升级版本。针对 V2.2 社区反馈和内部质量审计中出现的重复模式、异常中英文混入、噪声片段、弱证据支撑回答和低质量合成输出等问题,V2.3 提高了源文本进入生成环节的门槛,并优化了问答生成与过滤逻辑。 在数据构建上,V2.3 从约 **2.3T Parquet / raw corpus** 中进行高召回候选筛选,再通过小范围 **GPT-4.1 mini** 标注获得“是否适合生成高质量 SFT 样本”的 0/1 监督信号,进一步基于 **`IEITYuan/Yuan-embedding-2.0-zh`** 训练中文源文本分类打分器,对候选文本进行 SFT 可生成性排序与选择。选择 `IEITYuan/Yuan-embedding-2.0-zh` 作为底座,是因为该模型在选型时的中文 embedding 检索与排序公开榜单中取得 Rank 1 表现。最终,入选文本进入 FineWeb-Edu-Ultra 数据构建链路,由 **GPT-4.1 mini** 完成问答生成,并进行证据对齐、质量过滤和多格式导出。 本页保留高层次说明,聚焦数据集价值、版本升级、使用方式和风险边界。更完整的技术细节会随论文发布。 --- ## 核心价值 中文教育 SFT 数据长期面临三个实际问题:高质量公开数据稀缺,教育网页内容与可训练问答格式之间存在断层,合成数据又容易出现重复、空泛和不可验证回答。V2.3 的目标不是简单生成更多样本,而是提供一批更适合直接进入监督微调流程的中文教育问答数据。 该版本重点提升三类能力: - **更稳定的中文教育回答能力**:数据覆盖概念解释、事实问答、结构化总结、步骤推理和面向学习者的清晰表达。 - **更高纯度的 SFT 训练输入**:筛选目标从“是否像教育内容”升级为“是否适合生成可回答、可追溯、可训练的 SFT 样本”。 - **更低接入成本的训练格式**:同一批 QA pairs 同步提供 Alpaca、Messages、Messages-no-system 三种格式,方便直接接入常见训练框架。 对社区用户而言,V2.3 提供了一批可直接用于中文教育 SFT 的公开数据。对产业场景而言,V2.3 更适合作为中文知识服务、教育助手、企业培训、垂直问答和文本生成模型的数据底座之一。 --- ## V2.3 的质量升级 V2.2 将 Fineweb-Edu-Chinese 系列从预训练语料推进到 SFT 问答数据,为社区提供了大规模中文教育问答样本。V2.3 在此基础上进一步收紧质量门槛,强调源文本选择、生成稳定性和训练可用性。 V2.3 的规模小于 V2.2,是因为筛选门槛更高。系统不再只保留“看起来具有教育属性”的文本,而是更严格地判断候选文本是否能够支撑稳定、清晰、有依据的 SFT 问答构造。能通过这一门槛的数据更少,但更适合在后训练阶段提升模型回答质量。 V2.3 特别加强了以下问题的控制: - 长段回答中的重复句式和循环表达 - 源文本噪声导致的异常中英文混入 - 乱码、网页模板、广告、导航栏等低信号内容 - 问题可以提出但源文本不足以支撑答案的弱证据样本 - 生成回答看似完整但缺少可靠来源约束的样本 V2.3 的可复制难点不在于单次调用生成模型,而在于把大规模中文源语料治理、标注目标定义、筛选模型训练、问答生成和质量审计串成闭环。0/1 分类打分器的训练目标来自前置的小范围 LLM 标注,标注对象不是普通文本质量,而是候选文本是否适合生成高质量中文教育 SFT 样本。这个目标定义、标注闭环和大规模筛选落地,是 V2.3 相比普通合成数据流程更难复现的部分。 --- ## 数据筛选与构造策略 V2.3 的筛选策略围绕一个核心目标展开:优先让更可能产出稳定、可回答、有教育价值问答的源文本进入后续构造流程。 入选文本通常具备以下特征: - **中文主体清晰、可读性强**:文本主体以中文为主,语义连贯,句段结构相对完整,避免乱码、机器翻译残片、异常符号堆叠和大段中英文错位混杂。 - **具备教育价值和知识密度**:文本包含概念、定义、解释、因果关系、步骤说明、结构化事实、公式推导、表格信息、对比分析或领域知识,能够支持学习和问答构造。 - **信息相对自洽**:关键知识点可以从文本本身获得,不严重依赖缺失图片、外部链接、页面导航或网页上下文。 - **适合证据支撑生成**:文本能够为问题和答案提供明确依据,降低自由发挥和幻觉式回答风险。 - **低重复、低模板、低噪声**:文本不以导航栏、广告、网页模板、重复句段、短新闻片段、论坛闲聊和低信息密度列表为主。 经过筛选的 selected source text 会进入 FineWeb-Edu-Ultra 数据构建链路,由 GPT-4.1 mini 生成问答,并经过证据对齐、格式校验和质量过滤后导出为训练格式。这里仅描述高层流程,不展开阈值、提示词、审计规则和实验细节。 --- ## 版本演进 | 版本号 | 核心定位 | 数据规模 | 关键特性与改进 | 当前状态 | | --- | --- | --- | --- | --- | | **V1.0** | 概念验证 | 约 9000 万条,约 300GB | 初代 Chinese Fineweb Edu 语料;BERT 打分模型;MinHash 去重;数据源包括 CCI2、SkyPile、Tele-AI | 已弃用 | | **V2.0** | 规模化扩展 | 约 1.88 亿条,约 420B tokens | 升级至 OpenCSG csg-wukong-enterprise V2 打分器;扩展 Industry2、wanjuan1.0、wudao 等数据源 | 已弃用 | | **V2.1** | 预训练精选 | 总计约 1.5T tokens | 按分数分层组织;新增 map-cc、opencsg-cc;支持灵活预训练和课程学习 | 推荐用于预训练 | | **V2.2** | SFT 与对齐 | 约 143.7 万条高质量问答 | 将高质量教育语料转化为 SFT 问答数据;提供纯 QA 与上下文版本 | 历史 SFT 版本 | | **V2.3** | 更高纯度的 SFT 数据 | 23.04 万条 QA pairs | 升级 V2.2 的源文本选择和生成逻辑;强化证据对齐、质量过滤和多格式导出 | 推荐用于 SFT | --- ## 数据规格与仓库组织 当前魔塔社区版本为 10% 抽样预览,以合并后的 JSONL 文件组织在 `sft/` 路径下,并通过仓库 metadata 暴露 `messages`、`messages_no_sys`、`alpaca` 三个 config。每个 config 各提供 `train` 和 `validation` 两个 split,并提供 `sft/stats.json` 作为抽样统计文件。三种格式是同一批 QA pairs 的不同导出视图。 | 数据组件 | ModelScope `subset_name` | split | 文件路径 | 抽样规模 | 格式 | 推荐用途 | | --- | --- | --- | --- | --- | --- | --- | | **Alpaca** | `alpaca` | `train`, `validation` | `sft/train_alpaca.jsonl`, `sft/val_alpaca.jsonl` | train 20,739 / validation 2,305 | `instruction`, `input`, `output`, `metadata` | Alpaca 风格 SFT、LLaMA-Factory 等训练流程 | | **Messages** | `messages` | `train`, `validation` | `sft/train_messages.jsonl`, `sft/val_messages.jsonl` | train 20,739 / validation 2,305 | system, user, assistant messages + `metadata` | 带 system prompt 的 chat-style SFT | | **Messages-no-system** | `messages_no_sys` | `train`, `validation` | `sft/train_messages_no_sys.jsonl`, `sft/val_messages_no_sys.jsonl` | train 20,739 / validation 2,305 | user, assistant messages + `metadata` | 不带 system prompt 的 chat-style SFT | 魔塔社区抽样预览包含 **23,044 条 QA pairs**。完整发布口径以 **230,440 条 QA pairs** 为准。由于同一问答会被导出为多种训练格式,实际训练时应根据模型模板选择对应 split。 --- ## Schema ### Alpaca ```json { "instruction": "Answer the following fineweb_sft_pilot question with detailed derivations and explanations.", "input": "用户问题", "output": "助手答案", "metadata": { "domain": "fineweb_sft_pilot", "question_type": "General QA", "source_paper": "part-00000-00000_0005997", "quality_score": 0.0, "images": [] } } ``` ### Messages ```json { "messages": [ { "role": "system", "content": "You are an expert bilingual educational assistant. Convert high-quality Chinese or English web knowledge into rigorous, useful SFT examples with precise explanations, complete reasoning, and careful grounding in the source text. Use the same language as the user question whenever possible." }, { "role": "user", "content": "用户问题" }, { "role": "assistant", "content": "助手答案" } ], "metadata": { "domain": "fineweb_sft_pilot", "question_type": "General QA", "source_paper": "part-00000-00000_0005997", "quality_score": 0.0, "images": [] } } ``` ### Messages-no-system ```json { "messages": [ { "role": "user", "content": "用户问题" }, { "role": "assistant", "content": "助手答案" } ], "metadata": { "domain": "fineweb_sft_pilot", "question_type": "General QA", "source_paper": "part-00000-00000_0005997", "quality_score": 0.0, "images": [] } } ``` 常见 metadata 字段包括: | 字段 | 含义 | | --- | --- | | `domain` | 数据构建域或任务来源标记 | | `question_type` | 问题类型标记 | | `source_paper` | 源文本或构造批次标识 | | `quality_score` | 质量相关分数字段 | | `images` | 图像相关字段,通常为空数组 | | `license_type` | 若公开快照中包含该字段,它是记录级授权来源标记,不替代仓库级许可协议 | 公开训练样本以问答和 metadata 为主。证据对齐是构造阶段和质量控制阶段的重要环节,不代表每条最终 JSONL 记录都会包含完整源文本或证据片段字段。 --- ## 真实样例 以下样例节选自公开 JSONL 文件,保留原始中文问答内容。 ### 样例 1:语文知识 **Question** 什么是AABB结构成语?请结合摇摇摆摆、清清楚楚等例子说明其特点和表达作用,并举例说明正确与错误使用该类成语时可能带来的语义或语境差异。 **Answer** AABB结构成语由两个相同的双字组成,例如“摇摇摆摆”、“恍恍惚惚”、“清清楚楚”等。这类成语通常用来形容某种状态或动作的持续性、反复性或强烈程度。以“摇摇摆摆”为例,它描述动作不稳或晃动的状态;“清清楚楚”则表示非常明白、明确。通过填空练习,学生可以加深对这类成语构成和用法的理解与掌握。正确使用这类成语能够准确传递表达对象的特点,增强语言的生动性。例如:“树在风中摇摇摆摆”生动形象地描绘了树木的晃动。而错误使用可能导致语义模糊或语境不符,如将“清清楚楚”用于形容模糊的事物,会使表达产生矛盾,影响理解。总之,掌握AABB成语的结构和语境适用,有助于提高语文表达的准确性和丰富性。 ### 样例 2:工业自动化 **Question** 组态软件如何利用实时多任务和多线程技术提升工业自动化中的数据采集与监控效率?其良好可靠性、可扩展性和封装性对系统性能和用户操作有哪些具体影响?请结合实例说明丰富的图形工具和设备驱动如何满足复杂工业需求。 **Answer** 组态软件通过采用多线程操作,实现了实时多任务处理,使其能够在工业自动化环境中快速响应各类数据采集与监控需求。多线程机制保证监控界面能够实时刷新,同时支持数据的及时输入与输出,从而提高了系统的响应速度和效率。良好的可靠性依赖于稳定的硬件配置及供电系统配合,同时支持第三方程序接口,增强了软件与不同硬件及系统的兼容性,降低系统运行故障率。软件的可扩展性允许用户通过开发应用程序对控制系统进行微调和功能拓展,使系统能够适应不断变化的工业需求。封装性方面,组态软件将复杂的编程操作封装在图形化界面中,使非专业编程人员也能轻松通过图形界面完成系统配置,提高了用户操作的便利性和效率。举例来说,丰富的图形工具和标准图形库支持用户快速创建直观的人机交互界面,便于实时监控生产过程。多样化的设备驱动程序则保障了不同硬件设备间的稳定通信,并支持自定义通信协议和驱动程序开发包,从而满足复杂工业自动化系统中多样化的设备接入与数据交互需求。总体而言,这些特性协同作用,有效提升了工业自动化系统的控制精度、运行稳定性及操作便捷性。 --- ## 快速开始 ```python from modelscope.msdatasets import MsDataset ds_messages = MsDataset.load( dataset_name="Fineweb-Edu-Chinese-V2.3", namespace="OpenCSG", subset_name="messages", split="train", use_streaming=True, ) ds_messages_no_sys = MsDataset.load( dataset_name="Fineweb-Edu-Chinese-V2.3", namespace="OpenCSG", subset_name="messages_no_sys", split="train", use_streaming=True, ) ds_alpaca = MsDataset.load( dataset_name="Fineweb-Edu-Chinese-V2.3", namespace="OpenCSG", subset_name="alpaca", split="train", use_streaming=True, ) ``` --- ## 适用场景 - 中文教育问答模型训练 - 中文大模型监督微调 - 教育、培训、知识服务类模型构建 - 基于中文网页知识的问答生成和文本生成 - 合成数据质量过滤、数据构建链路评估和 SFT 数据研究 --- ## 限制与风险边界 V2.3 是基于大规模中文网页语料筛选和 GPT-4.1 mini 生成得到的合成 QA 数据。尽管数据构建过程中进行了筛选、证据对齐和质量过滤,合成问答仍可能包含事实错误、遗漏、过时信息或表达偏差,不应被视为事实权威来源。 将本数据集用于医疗、法律、金融、公共服务、教育评价等高风险场景前,需要进行额外的领域专家评审、模型评测、安全评测和合规审查。使用者也应结合自身产品形态、部署地区和下游任务要求,评估数据使用带来的偏差、版权、隐私和安全风险。 --- ## 许可说明 使用本数据集需要遵循 OpenCSG 数据集许可协议。仓库 metadata 中的 license: other 表示本数据集采用平台预设列表之外的许可协议,实际许可条款以该协议为准。 本数据集可按 OpenCSG 数据集许可协议申请商业用途。若计划将本数据集,或基于本数据集训练、增强的模型、系统、Agent、API 服务和商业产品用于商业场景,请发送邮件至 lorraineg@opencsg.com 获取许可。 当前公开快照中的 license_type: 商业授权 是记录级授权来源标记,不替代仓库级许可协议。 --- ## Citation ```bibtex @dataset{opencsg_cimd_2026, title = {CIMD: A Cross-Source Multilingual Document Corpus}, author = {OpenCSG}, year = {2026}, url = {https://opencsg.com/datasets/OpenCSG/CIMD}, note = {OpenCSG dataset repository} } ``` --- <a id="english"></a> ## Dataset Overview **Chinese Fineweb Edu Dataset V2.3** is a high-quality Chinese educational SFT dataset released by OpenCSG for Chinese education, knowledge question answering, instruction tuning, and text generation. This version contains **230.4K high-quality Chinese educational QA pairs**. The same QA pairs are exported in **Alpaca**, **Messages**, and **Messages-no-system** formats. These formats target different training templates. For training, users should select the format that matches their model and framework instead of treating the three exports as independent knowledge scale. This ModelScope page only publishes a **10% preview sample**. The complete **230,440 QA pairs** are available on the [CSGHub Fineweb-Edu-Chinese-V2.3 dataset page](https://opencsg.com/datasets/OpenCSG/Fineweb-Edu-Chinese-V2.3?tab=summary). V2.3 is a quality upgrade over **V2.2**. In response to community feedback and internal quality audits on V2.2, including repeated patterns, abnormal Chinese-English mixing, noisy fragments, weakly grounded answers, and low-quality synthetic outputs, V2.3 raises the threshold for source text selection and improves the QA generation and filtering logic. For data construction, V2.3 starts from about **2.3T Parquet / raw corpus** and applies high-recall candidate filtering. A small-scale **GPT-4.1 mini** annotation step is then used to produce 0/1 supervision signals for whether a candidate text is suitable for generating high-quality SFT samples. Based on these labels, OpenCSG trains a Chinese source-text classification scorer on **`IEITYuan/Yuan-embedding-2.0-zh`** and uses it to rank and select candidate texts by SFT generation suitability. `IEITYuan/Yuan-embedding-2.0-zh` was selected because it ranked No. 1 in Chinese embedding retrieval and ranking public leaderboards at model selection time. The selected texts are then processed by the FineWeb-Edu-Ultra data construction flow, where **GPT-4.1 mini** generates QA pairs followed by evidence alignment, quality filtering, and multi-format export. This page keeps the description at a high level and focuses on dataset value, version upgrade, usage, and risk boundaries. More technical details will be released with the paper. --- ## Core Value Chinese educational SFT data faces three practical challenges: scarcity of high-quality public data, a gap between educational web content and trainable QA formats, and synthetic outputs that may become repetitive, vague, or difficult to verify. V2.3 is designed not to maximize sample count, but to provide a cleaner set of Chinese educational QA data that is more suitable for supervised fine-tuning. This version strengthens three capabilities: - **More stable Chinese educational responses**: the data covers concept explanation, factual QA, structured summarization, step-by-step reasoning, and learner-oriented explanations. - **Higher-purity SFT input**: the filtering target is upgraded from whether a text looks educational to whether it can support answerable, traceable, and trainable SFT samples. - **Lower integration cost**: the same QA pairs are released in Alpaca, Messages, and Messages-no-system formats for direct use in common training frameworks. For community users, V2.3 provides a public dataset that can be directly used for Chinese educational SFT. For industry scenarios, V2.3 can serve as one of the data foundations for Chinese knowledge services, educational assistants, enterprise training, vertical QA, and text generation models. --- ## V2.3 Quality Upgrade V2.2 moved the Fineweb-Edu-Chinese series from pretraining corpus construction into SFT QA data and provided large-scale Chinese educational QA samples to the community. V2.3 further tightens quality thresholds, with stronger emphasis on source text selection, generation stability, and training usability. V2.3 is smaller than V2.2 because the selection threshold is higher. The system no longer keeps texts merely because they appear educational. It more strictly checks whether candidate texts can support stable, clear, and grounded SFT QA construction. Fewer samples pass this threshold, but the retained data is better suited for improving model response quality during post-training. V2.3 strengthens control over the following issues: - Repeated sentence patterns and circular expressions in long answers - Abnormal Chinese-English mixing caused by noisy source text - Garbled text, page templates, ads, navigation bars, and other low-signal content - Weak-evidence samples where a question can be asked but the source text does not sufficiently support the answer - Synthetic answers that look complete but lack reliable source constraints The hard-to-replicate part of V2.3 is not a single call to a generation model. It is the closed loop that combines large-scale Chinese source-corpus governance, target definition for annotation, scorer training, QA generation, and quality auditing. The 0/1 classification scorer is trained from an earlier small-scale LLM annotation stage. The annotation target is not generic text quality, but whether a candidate text is suitable for generating high-quality Chinese educational SFT samples. This target definition, annotation loop, and large-scale selection process make V2.3 harder to reproduce than a regular synthetic-data pipeline. --- ## Data Selection and Construction Strategy The V2.3 filtering strategy is built around one objective: prioritize source texts that are more likely to produce stable, answerable, and educational QA samples. Selected texts typically have the following properties: - **Readable Chinese-dominant content**: the main body is Chinese, semantically coherent, and structurally complete, while garbled text, machine-translation fragments, abnormal symbol sequences, and misaligned Chinese-English mixtures are deprioritized. - **Educational value and knowledge density**: the text contains concepts, definitions, explanations, causal relations, procedural steps, structured facts, formulas, tables, comparisons, or domain knowledge that can support learning and QA construction. - **Self-contained information**: key knowledge points can be obtained from the text itself, without heavy dependence on missing images, external links, page navigation, or unavailable web context. - **Suitable for evidence-grounded generation**: the text provides clear grounding for both questions and answers, reducing free-form speculation and hallucination-like answers. - **Low repetition, low templating, and low noise**: navigation bars, ads, page templates, repeated passages, short news fragments, forum chatter, and low-information lists are filtered out or deprioritized. The selected source text enters the FineWeb-Edu-Ultra data construction flow. GPT-4.1 mini generates QA pairs, followed by evidence alignment, format validation, quality filtering, and export into training formats. This page only describes the high-level process and does not disclose thresholds, prompts, audit rules, or experimental details. --- ## Version Evolution | Version | Positioning | Scale | Key Features and Improvements | Status | | --- | --- | --- | --- | --- | | **V1.0** | Proof of concept | About 90M records, about 300GB | First Chinese Fineweb Edu corpus; BERT scoring model; MinHash deduplication; sources include CCI2, SkyPile, and Tele-AI | Deprecated | | **V2.0** | Scaled expansion | About 188M records, about 420B tokens | Upgraded to OpenCSG csg-wukong-enterprise V2 scorer; expanded sources including Industry2, wanjuan1.0, and wudao | Deprecated | | **V2.1** | Curated pretraining corpus | About 1.5T tokens in total | Score-tiered organization; added map-cc and opencsg-cc; supports flexible pretraining and curriculum learning | Recommended for pretraining | | **V2.2** | SFT and alignment | About 1.437M high-quality QA pairs | Converts high-quality educational corpus into SFT QA data; provides pure QA and context versions | Historical SFT version | | **V2.3** | Higher-purity SFT data | 230.4K QA pairs | Upgrades V2.2 source selection and generation logic; strengthens evidence alignment, quality filtering, and multi-format export | Recommended for SFT | --- ## Data Specification and Repository Organization The ModelScope version is a 10% preview sample. It is organized as merged JSONL files under `sft/` and exposes three configs through repository metadata: `messages`, `messages_no_sys`, and `alpaca`. Each config has `train` and `validation` splits, with `sft/stats.json` as the sample statistics file. The three formats are different export views of the same QA pairs. | Component | ModelScope `subset_name` | Splits | File Paths | Sample Scale | Format | Recommended Use | | --- | --- | --- | --- | --- | --- | --- | | **Alpaca** | `alpaca` | `train`, `validation` | `sft/train_alpaca.jsonl`, `sft/val_alpaca.jsonl` | train 20,739 / validation 2,305 | `instruction`, `input`, `output`, `metadata` | Alpaca-style SFT, LLaMA-Factory, and similar training flows | | **Messages** | `messages` | `train`, `validation` | `sft/train_messages.jsonl`, `sft/val_messages.jsonl` | train 20,739 / validation 2,305 | system, user, assistant messages + `metadata` | Chat-style SFT with system prompt | | **Messages-no-system** | `messages_no_sys` | `train`, `validation` | `sft/train_messages_no_sys.jsonl`, `sft/val_messages_no_sys.jsonl` | train 20,739 / validation 2,305 | user and assistant messages + `metadata` | Chat-style SFT without system prompt | The ModelScope preview sample contains **23,044 QA pairs**. The full release scale is **230,440 QA pairs**. Because the same QA pairs are exported into multiple formats, users should choose the split that matches their model template for training. --- ## Schema ### Alpaca ```json { "instruction": "Answer the following fineweb_sft_pilot question with detailed derivations and explanations.", "input": "User question", "output": "Assistant answer", "metadata": { "domain": "fineweb_sft_pilot", "question_type": "General QA", "source_paper": "part-00000-00000_0005997", "quality_score": 0.0, "images": [] } } ``` ### Messages ```json { "messages": [ { "role": "system", "content": "You are an expert bilingual educational assistant. Convert high-quality Chinese or English web knowledge into rigorous, useful SFT examples with precise explanations, complete reasoning, and careful grounding in the source text. Use the same language as the user question whenever possible." }, { "role": "user", "content": "User question" }, { "role": "assistant", "content": "Assistant answer" } ], "metadata": { "domain": "fineweb_sft_pilot", "question_type": "General QA", "source_paper": "part-00000-00000_0005997", "quality_score": 0.0, "images": [] } } ``` ### Messages-no-system ```json { "messages": [ { "role": "user", "content": "User question" }, { "role": "assistant", "content": "Assistant answer" } ], "metadata": { "domain": "fineweb_sft_pilot", "question_type": "General QA", "source_paper": "part-00000-00000_0005997", "quality_score": 0.0, "images": [] } } ``` Common metadata fields include: | Field | Meaning | | --- | --- | | `domain` | Data construction domain or task source marker | | `question_type` | Question type marker | | `source_paper` | Source text or construction batch identifier | | `quality_score` | Quality-related score field | | `images` | Image-related field, usually an empty array | | `license_type` | If present in a public snapshot, this is a record-level authorization source marker and does not replace the repository-level license agreement | Public training records mainly contain QA content and metadata. Evidence alignment is an important construction-stage and quality-control step, but this does not mean that every final JSONL record contains full source text or evidence-span fields. --- ## Sample Records The following examples are excerpted from public JSONL files. The original Chinese QA content is preserved. ### Sample 1: Chinese Language Knowledge **Question** 什么是AABB结构成语?请结合摇摇摆摆、清清楚楚等例子说明其特点和表达作用,并举例说明正确与错误使用该类成语时可能带来的语义或语境差异。 **Answer** AABB结构成语由两个相同的双字组成,例如“摇摇摆摆”、“恍恍惚惚”、“清清楚楚”等。这类成语通常用来形容某种状态或动作的持续性、反复性或强烈程度。以“摇摇摆摆”为例,它描述动作不稳或晃动的状态;“清清楚楚”则表示非常明白、明确。通过填空练习,学生可以加深对这类成语构成和用法的理解与掌握。正确使用这类成语能够准确传递表达对象的特点,增强语言的生动性。例如:“树在风中摇摇摆摆”生动形象地描绘了树木的晃动。而错误使用可能导致语义模糊或语境不符,如将“清清楚楚”用于形容模糊的事物,会使表达产生矛盾,影响理解。总之,掌握AABB成语的结构和语境适用,有助于提高语文表达的准确性和丰富性。 ### Sample 2: Industrial Automation **Question** 组态软件如何利用实时多任务和多线程技术提升工业自动化中的数据采集与监控效率?其良好可靠性、可扩展性和封装性对系统性能和用户操作有哪些具体影响?请结合实例说明丰富的图形工具和设备驱动如何满足复杂工业需求。 **Answer** 组态软件通过采用多线程操作,实现了实时多任务处理,使其能够在工业自动化环境中快速响应各类数据采集与监控需求。多线程机制保证监控界面能够实时刷新,同时支持数据的及时输入与输出,从而提高了系统的响应速度和效率。良好的可靠性依赖于稳定的硬件配置及供电系统配合,同时支持第三方程序接口,增强了软件与不同硬件及系统的兼容性,降低系统运行故障率。软件的可扩展性允许用户通过开发应用程序对控制系统进行微调和功能拓展,使系统能够适应不断变化的工业需求。封装性方面,组态软件将复杂的编程操作封装在图形化界面中,使非专业编程人员也能轻松通过图形界面完成系统配置,提高了用户操作的便利性和效率。举例来说,丰富的图形工具和标准图形库支持用户快速创建直观的人机交互界面,便于实时监控生产过程。多样化的设备驱动程序则保障了不同硬件设备间的稳定通信,并支持自定义通信协议和驱动程序开发包,从而满足复杂工业自动化系统中多样化的设备接入与数据交互需求。总体而言,这些特性协同作用,有效提升了工业自动化系统的控制精度、运行稳定性及操作便捷性。 --- ## Quick Start ```python from modelscope.msdatasets import MsDataset ds_messages = MsDataset.load( dataset_name="Fineweb-Edu-Chinese-V2.3", namespace="OpenCSG", subset_name="messages", split="train", use_streaming=True, ) ds_messages_no_sys = MsDataset.load( dataset_name="Fineweb-Edu-Chinese-V2.3", namespace="OpenCSG", subset_name="messages_no_sys", split="train", use_streaming=True, ) ds_alpaca = MsDataset.load( dataset_name="Fineweb-Edu-Chinese-V2.3", namespace="OpenCSG", subset_name="alpaca", split="train", use_streaming=True, ) ``` --- ## Use Cases - Chinese educational QA model training - Supervised fine-tuning for Chinese large language models - Educational, training, and knowledge-service model development - QA generation and text generation based on Chinese web knowledge - Synthetic data quality filtering, data construction evaluation, and SFT data research --- ## Limitations and Risk Boundaries V2.3 is synthetic QA data generated from large-scale Chinese web corpus selection and GPT-4.1 mini. Although the construction process includes filtering, evidence alignment, and quality control, synthetic QA may still contain factual errors, omissions, outdated information, or expression bias. It should not be treated as an authoritative source of facts. Before using this dataset in high-risk scenarios such as medicine, law, finance, public services, educational assessment, or other sensitive domains, users should conduct additional domain-expert review, model evaluation, safety evaluation, and compliance review. Users should also assess bias, copyright, privacy, and safety risks according to their product form, deployment region, and downstream task requirements. --- ## Licensing Use of this dataset is governed by the OpenCSG Dataset License Agreement. The license: other metadata value indicates that the license is outside the platform's preset license list. The license file in this repository is authoritative. This dataset supports commercial use. For commercial use of the dataset, models, systems, agents, APIs, or products trained or enhanced with this dataset, follow the license agreement and contact lorraineg@opencsg.com for authorization. The license_type: Commercial Authorization field in the current public snapshot is a record-level authorization source marker and does not replace the repository-level license agreement. --- ## Citation ```bibtex @dataset{opencsg_cimd_2026, title = {CIMD: A Cross-Source Multilingual Document Corpus}, author = {OpenCSG}, year = {2026}, url = {https://opencsg.com/datasets/OpenCSG/CIMD}, note = {OpenCSG dataset repository} } ```



