Dataset of excel tasks for LLMs
收藏资源简介:
We present an extended version of SheetCopilotBench (https://doi.org/10.48550/arXiv.2305.19308), a benchmark dataset for evaluating large language models (LLMs) and code-generation systems on spreadsheet-related tasks. The dataset covers a wide spectrum of Excel problems, including formula synthesis, table manipulation, data cleaning, aggregation, visualization, and automation. Each task is provided with structured inputs—such as natural language instructions, spreadsheet context, and relevant cell data. Compared to the original benchmark, this extension broadens task diversity and complexity, reflecting real-world spreadsheet use cases encountered in business, education, and data analysis. The dataset is intended to support research on LLM-based agents for productivity tools, robust code generation, and human-computer interaction in spreadsheet environments.



