RUC-DataLab/CoDA-Bench
收藏资源简介:
CoDA-Bench(代码与数据密集型基准)是首个在现实数据密集型环境中联合评估AI代理代码智能和数据智能的基准。与直接提供标准数据的现有基准不同,它要求代理在包含数百个语义相似文件的Linux沙盒环境中,发现相关数据、导航复杂文件层次、整合来自多个异构数据源的信息,并为数据驱动分析任务生成正确代码。数据集包含全基准(1,009个任务,覆盖31个社区)和困难子集(119个挑战性任务,覆盖15个社区),源数据来自267个笔记本中的199个Kaggle数据集,平均每个环境有980个文件(总压缩约43 GB)。
CoDA-Bench (Code and Data-Intensive Benchmark) is the first benchmark to jointly evaluate AI agents' code intelligence and data intelligence in real-world data-intensive environments. Unlike existing benchmarks that directly provide standardized datasets, it requires agents to discover relevant data, navigate complex file hierarchies, integrate information from multiple heterogeneous data sources, and generate correct code for data-driven analysis tasks within a Linux sandbox environment housing hundreds of semantically similar files. The benchmark consists of two parts: a full benchmark set with 1,009 tasks spanning 31 communities, and a challenging hard subset with 119 tasks spanning 15 communities. The source data is derived from 199 Kaggle datasets across 267 notebooks, with an average of 980 files per environment and a total compressed size of approximately 43 GB.




