teamcore/DPO_Pm3B_U0_beta0.25dpo_proEurus_RM_7bbt_noise_flip_paper0.3_nu0.008
收藏数据链接:
官方服务:
资源简介:
该数据集包含了一系列文本数据,其中包括源文本、指令、模型生成的完成情况(如帮助性、诚实性、指示遵循性和真实性等评分和解释)、评分、回应、正确和错误答案、提示、选择和拒绝的文本、不同模型的评分、模型预测的概率以及生成的奖励分数。数据集分为默认的一个拆分,共有3187个示例,总数据大小为55332502字节。
The dataset consists of a series of text entries, including source text, instructions, model-generated completions with ratings and explanations for helpfulness, honesty, instruction following, and truthfulness, ratings, responses, correct and incorrect answers, prompts, chosen and rejected texts, scores for different models, model prediction probabilities, and generated reward scores. The dataset is split into a default split with a total of 3187 examples and a dataset size of 55332502 bytes.
提供机构:
teamcore


