SCICODEPILE
收藏资源简介:
SCICODEPILE是由新加坡管理大学等机构联合构建的迄今为止规模最大的科学代码语料库,旨在为计算科学领域的代码生成提供训练资源与可执行评测基准。该数据集总量达128GB,源自37,737个经过严格筛选的公共代码仓库,覆盖化学、生命科学、计算方法和AI建模等多个学科领域,其内容以原始代码文件、仓库摘要、函数指令及问题-解决方案对四种互补格式呈现。数据构建过程采用了“检索-过滤”流程,结合大语言模型辅助的关键词扩展与质量控制,并进一步从中提炼出包含200个任务的、配备沙箱执行环境与自动化测试框架的可执行评测集。该数据集主要应用于提升大语言模型在科学代码生成任务上的能力,旨在解决现有资源在规模、领域覆盖及功能可验证性方面的不足,为可靠、可执行的科学代码生成研究提供关键基础设施。
SCICODEPILE is the largest scientific code corpus jointly constructed by Singapore Management University and other institutions to date, aiming to provide training resources and executable evaluation benchmarks for code generation in the field of computational science. With a total size of 128 GB, this dataset is sourced from 37,737 rigorously screened public code repositories, covering multiple disciplines including chemistry, life sciences, computational methods, and AI modeling. Its content is presented in four complementary formats: original code files, repository summaries, function instructions, and problem-solution pairs. The dataset construction process adopts a retrieval-filtering workflow, combined with large language model (LLM)-assisted keyword expansion and quality control, and further extracts an executable evaluation set containing 200 tasks equipped with a sandbox execution environment and an automated testing framework. This dataset is primarily applied to enhance the capabilities of large language models (LLMs) in scientific code generation tasks, aiming to address the shortcomings of existing resources in terms of scale, domain coverage, and functional verifiability, providing key infrastructure for research on reliable and executable scientific code generation.

- 1SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation新加坡管理大学; 南京大学; 中山大学; 内政部科技局 · 2026年



