zxj3060/paper2thesis
收藏资源简介:
Paper2Thesis是一个用于极端长度多文档合成的基准数据集。每个实例将一组arXiv研究论文映射到一个目标arXiv博士论文。任务要求生成一个论文规模的文档,将多篇论文整合成一个连贯、结构化的叙述。该数据集针对超出标准文本生成的范围,涉及长上下文推理、跨文档整合和结构化生成。数据集结构为JSONL格式,包含训练、验证和测试集。数据来源于arXiv,经过严格的筛选和验证流程。数据集不包含论文或论文的全文,仅包含arXiv标识符和元数据。
Paper2Thesis is a benchmark for extreme-length multi-document synthesis. Each instance maps a set of input arXiv research papers to a target arXiv PhD thesis. The task requires generating a thesis-scale document that integrates multiple papers into a coherent, structured narrative. This benchmark targets a regime beyond standard text generation, involving long-context reasoning, cross-document integration, and structured generation. The dataset is provided in JSONL format, including train, validation, and test sets. Data is derived from arXiv, with a rigorous screening and validation pipeline. The dataset does not include full text of papers or theses, only arXiv identifiers and metadata.





