Agentic Frameworks for Reasoning Tasks: An Empirical Study
收藏资源简介:
This dataset contains two types of files: Replication Package Benchmarks Results (.zip) Below, we explain each file. 1) Replication Package:This Excel file documents the framework selection process and presents the benchmark results and their analysis. It includes five sheets that describe the framework selection process, benchmark results, and analysis used in the study. Sheet 1: Search_StringThis sheet provides the search strings used to identify candidate frameworks from GitHub, along with the inclusion and exclusion criteria applied during the framework selection process. Sheet 2: Initial_Frameworks_ExtractionThis sheet contains the initially extracted 164 frameworks, followed by the refined set of 55 frameworks shortlisted from the original pool. Sheet 3: Final_FrameworksThis sheet presents the final set of frameworks selected for the experimental evaluation. It also includes the selection criteria used to finalize these frameworks for further benchmarking. Sheet 4: Benchmark_Results (RQ1–RQ3)This sheet reports the benchmark results for each agentic framework across the three reasoning benchmarks: BBH, GSM8K, and ARC. It includes the accuracy of each framework, the time taken per task, the total execution time per framework, and token consumption at both the task and framework levels for each benchmark. Sheet 5: AnalysisThis sheet contains the analysis of the benchmark results and summarizes the comparative performance of the evaluated frameworks. 2) Benchmarks Results (.zip):This ZIP file contains the complete benchmark results for the 22 evaluated agentic frameworks across three reasoning benchmarks: BBH, GSM8K, and ARC. The file is organized into separate folders for each benchmark, and each folder contains the result logs for all frameworks. Each log file is named in the format FM_frameworkname_benchmarkname_full_config.log. These files include detailed task-level information such as accuracy, input and output token usage, total token consumption, and execution time. Across all three benchmarks, the dataset contains results for 16,495 tasks, and the total ZIP file size is 17.4 MB.



