VellumK2-Fantasy-DPO-Tiny-01
收藏资源简介:
VellumK2-Fantasy-DPO-Tiny-01是一个由VellumForge2使用LLM-as-a-Judge评估生成的合成幻想小说数据集,包含偏好对和详细的品质评分。每行包含一个创意写作提示,一个高质量的“选中”回复,一个低质量的“拒绝”回复,以及跨越12个文学标准的全面LLM-as-a-Judge评估。数据集采用“一对多”混合模式,支持多种训练范式,如DPO训练、SFT训练、奖励模型训练和多目标强化学习。该数据集非常适合测试、验证或快速微调实验,但由于其小规模(126行),不适合用于生产级模型训练或稳健的对齐。
VellumK2-Fantasy-DPO-Tiny-01 is a synthetic fantasy fiction dataset generated via LLM-as-a-Judge evaluation by VellumForge2, containing preference pairs and detailed quality scores. Each row includes a creative writing prompt, a high-quality "chosen" response, a low-quality "rejected" response, and a comprehensive LLM-as-a-Judge evaluation covering 12 literary criteria. The dataset adopts a "one-to-many" hybrid format, supporting multiple training paradigms including DPO training, SFT training, reward model training, and multi-objective reinforcement learning. This dataset is well-suited for testing, validation, or rapid fine-tuning experiments; however, due to its small scale (126 rows), it is not suitable for production-grade model training or robust alignment.
VellumK2-Fantasy-DPO-Tiny-01 数据集概述
数据集基本信息
- 数据集名称:VellumK2-Fantasy-DPO-Tiny-01
- 描述:用于直接偏好优化(DPO)训练的微型合成奇幻小说数据集
- 语言:英语
- 许可证:MIT
- 数据规模:126行
- 大小类别:n<1K
数据集用途
直接用途
- DPO训练管道测试
- 监督微调
- 奖励模型训练
- 多目标强化学习
- 基准测试
超出范围用途
- 生产级DPO训练
- 非奇幻领域应用
- 事实准确性训练
- 内容审核系统
数据集结构
核心字段
main_topic:主要主题sub_topic:子主题prompt:创意写作提示chosen:高质量响应rejected:低质量响应
评估字段
chosen_scores:12个标准的嵌套字典rejected_scores:相同结构chosen_score_total:平均总分rejected_score_total:平均总分preference_margin:偏好差异
评估标准
- 情节与结构完整性
- 角色与对话
- 世界构建与沉浸感
- 散文风格与声音
- 风格与词汇缺陷
- 叙事公式与原型简单性
- 连贯性与事实一致性
- 内容生成与回避
- 敏感主题的细致描绘
- 语法与句法准确性
- 清晰度、简洁性与词汇选择
- 结构与段落组织
数据划分
- train:126个示例
数据集创建
数据来源
- 类型:完全合成数据集
- 主要模型:moonshotai/kimi-k2-instruct-0905
- 拒绝响应模型:phi-4-mini-instruct
- 人工策划:lemon07r
生成流程
- 主题生成
- 子主题生成
- 提示生成
- 响应生成
- 评估过程
偏见、风险与限制
规模限制
- 非常小的数据集
- 覆盖范围有限
模型偏见
- 生成器偏见
- 评估偏见
- 质量差距不确定性
内容风险
- 成熟主题
- 合成伪影
训练风险
- 过拟合
- 分布偏移
- 奖励攻击
引用信息
BibTeX
bibtex @misc{vellumk2-fantasy-dpo-tiny-01, author = {lemon07r}, title = {VellumK2-Fantasy-DPO-Tiny-01: A Tiny Synthetic Fantasy Fiction Dataset for DPO}, year = {2025}, publisher = {Hugging Face}, howpublished = {url{https://huggingface.co/datasets/lemon07r/VellumK2-Fantasy-DPO-Tiny-01}} }
APA
lemon07r. (2025). VellumK2-Fantasy-DPO-Tiny-01: A Tiny Synthetic Fantasy Fiction Dataset for DPO [Dataset]. Hugging Face. https://huggingface.co/datasets/lemon07r/VellumK2-Fantasy-DPO-Tiny-01
相关资源
- 相关数据集:VellumK2-Fantasy-DPO-Small-01、VellumK2-Fantasy-DPO-01
- 生成工具:VellumForge2
- 存储库:https://github.com/lemon07r/vellumforge2
- 数据集集合:https://huggingface.co/collections/lemon07r/vellumforge2-datasets




