TACO
收藏资源简介:
TACO数据集是由北京人工智能研究院和山东师范大学信息科学与工程学院联合创建的大型开源代码生成数据集,专注于算法主题,旨在为代码生成模型提供更具挑战性的训练数据和评估基准。该数据集包含26,443个编程任务,涵盖数学、数据结构和图论等多个领域,每个问题都附有详细的算法标签,如任务主题、算法、编程技能和难度级别,以提供更精确的训练和评估参考。TACO数据集不仅包含算法能力,还包含多方面的元数据,如时间和空间约束,旨在解决现实世界编程场景中的问题理解和推理能力评估。
The TACO dataset is a large-scale open-source code generation dataset co-developed by the Beijing Institute of Artificial Intelligence and the College of Information Science and Engineering, Shandong Normal University. It focuses on algorithm-related topics, aiming to provide more challenging training data and evaluation benchmarks for code generation models. The dataset contains 26,443 programming tasks covering multiple domains including mathematics, data structures, graph theory and others. Each problem is accompanied by detailed algorithmic labels such as task topic, algorithm, programming skills and difficulty level, to offer more precise references for training and evaluation. Besides algorithmic capabilities, the TACO dataset also includes various metadata such as time and space constraints, which aims to support the evaluation of problem understanding and reasoning abilities in real-world programming scenarios.

- 1TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding · 2024年



