CheemsBench, CheemsPreference
收藏资源简介:
CheemsBench是一个专为中文奖励模型设计的全面评价基准,包含1146个来自开源数据集和现实世界人类指令的prompt,每个prompt经过五轮人类驱动的三重比较,并通过图论算法解决标注冲突,生成唯一且一致的partial ranking。CheemsPreference是一个大规模、多样化的中文偏好数据集,通过人类标注和GPT协作标注构建,包含27,861个现实世界的人类指令,旨在为中文奖励模型训练提供监督信号。
CheemsBench is a comprehensive evaluation benchmark specifically designed for Chinese reward models. It contains 1,146 prompts sourced from open-source datasets and real-world human instructions. Each prompt has undergone five rounds of human-driven triple comparisons, and annotation conflicts are resolved via graph theory algorithms to generate unique and consistent partial rankings. CheemsPreference is a large-scale, diverse Chinese preference dataset constructed through human annotation and GPT collaborative annotation, which includes 27,861 real-world human instructions, aiming to provide supervision signals for the training of Chinese reward models.




