遇见数据集

CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

收藏
NBER2026-08-01 更新2026-09-04 收录
数据链接:
官方服务:

资源简介:

Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of

创建时间:
2026-08-01
二维码
社区交流群
二维码
科研交流群
商业服务