DS-1000
收藏资源简介:
DS-1000是由香港大学等机构创建的一个包含1000个数据科学问题的代码生成基准数据集。该数据集涵盖了NumPy、Pandas等七个Python库的实际应用问题,旨在通过自然语言描述生成相应的代码解决方案。数据集通过从StackOverflow收集真实问题,并采用多标准自动评估方法确保解决方案的正确性和可靠性。DS-1000的应用领域包括提高代码生成模型的性能,解决数据科学编程中的实际问题,推动相关技术的发展。
DS-1000 is a code generation benchmark dataset consisting of 1,000 data science problems, developed by institutions including the University of Hong Kong. This dataset covers practical application scenarios across seven Python libraries such as NumPy and Pandas, with the goal of generating corresponding code solutions based on natural language descriptions. The dataset collects real-world problems from StackOverflow, and employs multi-criteria automatic evaluation methods to guarantee the correctness and reliability of the solutions. The application areas of DS-1000 include enhancing the performance of code generation models, addressing practical issues in data science programming, and advancing the development of relevant technologies.

- 1DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation香港大学 · 2022年



