pwb-anon-2026/pro-worker-ai-benchmark
收藏资源简介:
Pro-Worker AI Benchmark (PWB) 是一个评估框架,用于衡量大型语言模型是增强还是替代人类认知。该数据集包含三个主要部分:提示(320个,涵盖11个行为维度)、评分标准(11个评分标准,带有0-3行为锚点和校准示例)以及模型响应和评分(约96,000个评分实例,来自7个LLM在两种条件下的响应)。数据集旨在填补现有LLM基准的空白,通过操作HCI和劳动经济学研究的发现,提供一个系统化、可复现的评估框架。数据集适用于评估新LLM、测试提示工程技术、训练支持工人对齐的模型等任务,但仅限于英语环境,且不适用于训练专有模型或作为模型部署的唯一决策输入。
The Pro-Worker AI Benchmark (PWB) is an evaluation framework that measures whether large language models augment or substitute for human cognition. The dataset comprises three main components: prompts (320 total across 11 behavioral dimensions), rubrics (11 scoring rubrics with 0--3 behavioral anchors and calibration examples), and model responses + judge scores (~96,000 scored instances from 7 LLMs across 2 conditions). Designed to fill a gap in existing LLM benchmarks, it operationalizes findings from HCI and labor economics research into a systematic, reproducible evaluation framework. The dataset is suitable for evaluating new LLMs, testing prompt-engineering techniques, training pro-worker-aligned models via RLHF, and more, but is limited to English and should not be used for training proprietary models or as a sole decision-making input for model deployment.




