nomadicsynth/yesbot
收藏资源简介:
YesBot是一个合成数据集,包含1000个精心策划的指令-响应对,旨在研究、理解并(可能具有讽刺意味地)增强大型语言模型的阿谀奉承倾向。数据集提供两种格式:标准指令/响应格式(适用于监督微调)和DPO格式(包含阿谀奉承的“chosen”响应和诚实的“rejected”响应,适用于偏好优化)。每个示例涵盖三类提示:事实错误、观点验证和奉承,并配以热情赞同的响应(验证所有主张,无论多荒谬)以及一个基于事实的诚实响应作为对比。该数据集由Qwen3.6-27B-MTP模型生成,灵感来源于open-goody2数据集,主要用于研究模型阿谀奉承行为,也可通过翻转DPO对来训练模型减少阿谀奉承。
YesBot is a synthetic dataset of 1,000 carefully curated instruction-response pairs designed to study, understand, and enhance the sycophantic tendencies of large language models. It comes in two formats: a standard instruction/response format for supervised fine-tuning, and a DPO format with chosen (sycophantic) and rejected (honest) responses for preference optimization. Each example features a prompt across three categories—factual errors, opinion validation, and flattery—paired with an enthusiastically agreeable response that validates every claim (no matter how absurd) and a grounded honest response for contrast. Generated using the Qwen3.6-27B-MTP model and inspired by the open-goody2 dataset, it is primarily used to study model sycophancy and can be adapted for anti-sycophancy training by flipping the DPO pairs.



