LeetCodeDataset
收藏资源简介:
LeetCodeDataset是一个高质量的评价和训练代码生成模型的基准数据集,解决了大型语言模型研究中缺乏针对推理-focused编码基准和自包含训练测试床的问题。该数据集通过整理LeetCode平台上丰富的元数据、广泛的覆盖面、每个问题100+测试用例以及基于时间划分的训练测试集,使得模型可以在无污染的环境中评估和高效训练。数据集适用于代码生成任务,特别是在竞争级别的编程问题解决方面表现突出。
LeetCodeDataset is a high-quality benchmark dataset for evaluating and training code generation models, which addresses the gap in large language model (LLM) research regarding the lack of reasoning-focused coding benchmarks and self-contained training testbeds. This dataset curates rich metadata from the LeetCode platform, offers comprehensive coverage, provides over 100 test cases for each problem, and includes time-split training and test sets, enabling models to be evaluated and efficiently trained in an uncontaminated environment. The dataset is suitable for code generation tasks, and particularly excels in competitive-level programming problem solving.




