multipref
收藏资源简介:
MultiPref数据集是一个包含10,000个人类偏好的丰富集合,具有多重注释和多方面的特点。每个实例被普通众包工作者和领域专家分别注释两次,总共有大约40,000个注释。除了整体偏好外,注释者还在五个方面的Likert量表上选择他们偏好的响应,包括帮助性、真实性和无害性。此外,注释者还说明了他们认为一个响应优于另一个的原因。数据集的结构包括每个实例的多个字段,如比较ID、提示ID、文本、模型生成的响应、来源、类别、学科研究、最高学位、普通工作者和专家工作者的注释等。注释字段包括每个方面的偏好、检查原因、自由形式的原因、注释者的信心、整体偏好和信心、评估者ID、注释时间、提交时间戳等。数据集的创建涉及从多个数据源获取提示,并使用多个模型生成响应,然后进行配对比较和随机选择进行注释。注释者包括普通众包工作者和领域专家,他们通过资格测试进行筛选。
The MultiPref dataset is a comprehensive collection of 10,000 human preferences, featuring multiple annotations and multi-faceted characteristics. Each instance is annotated twice by crowdsourced workers and twice by domain experts, resulting in a total of approximately 40,000 annotations across the entire dataset. In addition to providing overall preference ratings, annotators also score their preferred responses on five-dimensional Likert scales covering dimensions including helpfulness, authenticity, and harmlessness. Additionally, annotators are required to elaborate on the reasons why they deem one response superior to another. The dataset structure encompasses multiple fields for each instance, such as comparison ID, prompt ID, raw text, model-generated responses, source, category, research domain, highest academic degree, annotations from crowdsourced workers, and annotations from domain experts, among others. The annotation fields include per-aspect preference scores, checklist-based reasons, free-form reasons, annotator confidence levels, overall preference and its associated confidence, rater ID, annotation time, submission timestamp, and other relevant fields. The construction of the MultiPref dataset involves collecting prompts from multiple data sources, generating responses using multiple models, followed by pairwise comparison and random selection of response pairs for annotation. The annotators consist of both crowdsourced workers and domain experts, who are screened through qualification tests.




