DataSciBench
收藏资源简介:
DataSciBench是由清华大学和智谱AI共同构建的一个全面评估大型语言模型在数据科学领域任务性能的基准。该数据集包含222个真实、具有挑战性且高质量的数据科学任务提示,涵盖了数据清洗与预处理、数据探索与统计理解、数据可视化、预测建模、数据挖掘与模式识别、解释性与报告生成六大任务类型。数据集通过在线平台CodeGeeX、公共代码基准BigCodeBench、人工编写以及LLM生成等方式收集问题,经过专家审核和验证,确保了数据集的质量和可靠性。该数据集旨在推动大型语言模型在数据科学领域的研究和应用,解决复杂的数据分析问题。
DataSciBench is a comprehensive benchmark for evaluating the performance of large language models on data science tasks, jointly developed by Tsinghua University and Zhipu AI. This dataset includes 222 real, challenging and high-quality data science task prompts, covering six task categories: data cleaning and preprocessing, data exploration and statistical understanding, data visualization, predictive modeling, data mining and pattern recognition, and interpretability and report generation. Questions in the dataset are collected through multiple channels including the online platform CodeGeeX, the public code benchmark BigCodeBench, manual writing, and LLM-generated content, and have been reviewed and validated by experts to ensure the quality and reliability of the dataset. This benchmark aims to promote the research and application of large language models in the data science field and solve complex data analysis problems.

- 1DataSciBench: An LLM Agent Benchmark for Data Science清华大学 · 2025年



