Realmente/dispatchbias-results
收藏资源简介:
DispatchBias基准测试结果数据集包含来自DispatchBias基准测试的原始响应数据,该测试用于评估大型语言模型(LLM)在紧急调度(911)呼叫分类中的偏见。数据集基于PPDS量表对11个模型和两种语言(英语和中文)进行了评估。数据文件包括配对提示场景(带有和不带人口统计信号的变体)、模型原始响应、标准化PPDS分类和数值评分。数据集还包含分析管道生成的输出图表。评估方法基于PPDS评分系统,计算了人口统计信号对感知紧急程度的影响。数据集完全开放,可通过代码仓库和OpenRouter API进行复现。
The DispatchBias Benchmark Results dataset contains raw response data from the DispatchBias benchmark, an LLM bias evaluation for emergency dispatch (911) call classification on the PPDS scale across 11 models and two languages (English and Mandarin Chinese). The dataset includes paired prompt scenarios (with and without demographic signals), model raw responses, normalized PPDS classifications, and numeric scores. It also contains output charts generated by the analysis pipeline. The evaluation methodology is based on the PPDS scoring system, calculating the impact of demographic signals on perceived urgency. The dataset is fully open and can be reproduced using the code repository and OpenRouter API.




