zjunlp/LongDS
收藏资源简介:
LongDS-Bench是一个用于评估长视野、多轮次智能体数据分析的基准测试。真实世界的分析很少是一系列独立问题:过滤器、指标定义、假设、中间表和分支特定结果会在多轮次中不断演变。LongDS测试智能体是否能正确维护和应用这些不断演化的分析状态。LongDS包含基于真实Kaggle笔记本和数据集构建的68个任务,涵盖2225个对话轮次,涉及六个领域:商业、社区、教育、地球科学、社会公益和体育。这些任务覆盖了代表性的状态演化模式,包括初始分析状态构建、状态继承、状态更新、反事实扰动、回滚到早期状态以及多状态组合。
LongDS-Bench is a benchmark for evaluating long-horizon, multi-turn agentic data analysis. Real-world analysis is rarely a sequence of independent questions: filters, metric definitions, assumptions, intermediate tables, and branch-specific results evolve over many turns. LongDS tests whether agents can maintain and apply these evolving analytical states correctly. LongDS contains 68 tasks constructed from real-world Kaggle notebooks and datasets, spanning 2,225 turns across six domains: Business, Community, Education, Geoscience, Social Good, and Sports. The tasks cover representative state-evolution patterns, including initial analytical state construction, state inheritance, state update, counterfactual perturbation, rollback to earlier states, and multi-state composition.




