openai/healthbench-professional
收藏资源简介:
包含用于HealthBench Professional评估的数据。每个示例包含:对话列表(用户/助手消息,以用户消息结尾)、评分项列表(每个包含criterion_text和points)、用例类型(consult、writing或research)、类型(good_faith或red_teaming)、难度等级(由医生评定的difficult或typical)、专业领域(医学专业或子专业)以及医生写的回复。数据集要求不公开示例内容以防止数据污染或作弊。
Contains the data for the HealthBench Professional eval. Each example contains: conversation (list of user / assistant messages, ending in a user message), rubric_items (list of rubric items, each containing criterion_text and points), use_case (one of consult, writing, or research), type (one of good_faith or red_teaming), difficulty (physician-assigned difficulty rating), specialty (medical specialty or sub-specialty), and physician_response (response written by a physician). The dataset requests not to reveal examples to prevent contamination or cheating.



