Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training (CPT Corpus)
收藏资源简介:
本数据集名为Unfolding科学论文多轮生成轨迹数据集,由字节跳动等机构联合构建。该数据集包含约180万条从arXiv论文中解构得到的多轮生成轨迹,总token数达57-600亿,原始论文文本来源约300亿token。数据创建过程采用逆向重构方法,通过大语言模型为每篇论文生成写作请求、全局规划及章节预写作思索,同时保留原文段落与摘要。该数据集适用于大语言模型的持续预训练,旨在提升模型在学术写作、长文档理解及结构化文本生成方面的能力,同时保持通用推理性能不受损。
This dataset, named the Unfolding Scientific Paper Multi-turn Generation Trajectory Dataset, was co-constructed by ByteDance and other institutions. It contains approximately 1.8 million multi-turn generation trajectories deconstructed from arXiv papers, with a total token count ranging from 5.7 billion to 60 billion, while the raw source academic paper corpus totals approximately 30 billion tokens. The dataset was developed using a reverse reconstruction methodology, where large language models (LLMs) are employed to generate writing prompts, global planning schemes, and pre-writing reflections for each paper, while retaining the original paragraphs and abstracts of the source papers. This dataset is suitable for continuous pre-training of large language models, aiming to enhance the models' capabilities in academic writing, long document comprehension, and structured text generation, while preserving their general reasoning performance without degradation.

- 1Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training字节跳动 Seed; 南京大学; Evolvent AI · 2026年




