rag-consistency-test-14
收藏资源简介:
该数据集是一个用于评估和比较大型语言模型在帮助性、安全性、真实性等维度上表现的多配置偏好对齐数据集,包含两个主要配置:get_help_steer_3_no_history 和 help_steer_3,每个配置下包含多个由不同模型生成的数据子集。每个数据样本包括多轮对话上下文(由角色和内容组成)、两个候选回答(A和B)、人工标注的真实偏好标签(ground_truth),以及一系列用于深入分析模型行为的元数据和指标,如上下文长度、历史长度、解析失败次数、位置偏差(偏向第一个或第二个答案)、平局不一致性等,专门用于量化模型在生成答案时的系统性偏差。数据子集来源于多个知名AI模型和机构,如MiniMaxAI、Google Gemma、Swiss AI的Apertus模型、ZAI的GLM等,每个子集包含300个样本,使用不同的提示模板(如human_template, gepa_apertus8b_template)生成。该数据集适用于模型对齐研究、偏好建模、偏差检测、多模型比较等任务。
This dataset is a multi-configuration preference alignment dataset designed for evaluating and comparing the performance of large language models across dimensions such as helpfulness, safety, and truthfulness. It includes two main configurations: get_help_steer_3_no_history and help_steer_3, each containing multiple data subsets generated by different models. Each data sample consists of a multi-turn dialogue context (composed of roles and content), two candidate responses (A and B), a human-annotated ground truth preference label, and a series of metadata and metrics for in-depth analysis of model behavior. These metrics include context length, history length, parsing failures, position bias (favoring the first or second answer), tie inconsistency, etc., specifically designed to quantify systematic biases in model-generated answers. The data subsets are sourced from various well-known AI models and organizations, such as MiniMaxAI, Google Gemma, Swiss AIs Apertus model, ZAIs GLM, etc., with each subset containing 300 samples generated using different prompt templates (e.g., human_template, gepa_apertus8b_template). This dataset is suitable for tasks like model alignment research, preference modeling, bias detection, and multi-model comparison.




