A BENCHMARK FRAMEWORK FOR EVALUATING AUTONOMOUS AI AGENTS IN ACCOUNTING
收藏资源简介:
The rapid proliferation of large language model (LLM)-based autonomous agents has introduced new possibilities for automating complex, multi-step accounting tasks, ranging from journal entry generation and bank reconciliation to tax computation and financial statement drafting. Despite growing industry adoption, the accounting profession currently lacks a standardized, discipline-specific benchmark for evaluating the reliability, regulatory compliance, and auditability of these agents. Existing AI benchmarks are largely borrowed from general natural-language-processing or software-engineering domains and fail to capture the procedural rigor, traceability, and normative constraints that define professional accounting practice. This paper proposes a Benchmark Framework for Evaluating Autonomous AI Agents in Accounting (BEAAA), a structured evaluation protocol comprising six dimensions: task accuracy, regulatory compliance, auditability and traceability, robustness to adversarial or ambiguous inputs, cost-efficiency, and explainability. The framework operationalizes each dimension through a task taxonomy spanning four accounting sub-domains — bookkeeping, reconciliation, financial reporting, and tax preparation — and a graded scoring rubric that captures partial task completion rather than binary success/failure outcomes. We illustrate the framework through a pilot evaluation of three representative agent architectures across a synthetic task suite of 120 accounting scenarios. Results indicate substantial variance across agents in compliance and auditability scores despite comparable raw task-accuracy scores, suggesting that accuracy alone is an insufficient basis for deployment decisions in regulated accounting environments. The proposed framework offers accounting researchers, software vendors, and regulators a common vocabulary and methodology for comparing autonomous agents, and it identifies auditability and explainability as the most under-addressed dimensions in current agent design. Implications for standard-setting bodies and directions for future benchmark development are discussed.



