CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
收藏数据链接:
官方服务:
资源简介:
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of
提供机构:
美国国家经济研究局创建时间:
2026-08-01



