ephorata/tenacious-bench-path-b-preference
收藏资源简介:
Tenacious Bench Path B Preference数据集旨在评估在B2B外展工作中特定失败模式的偏好对。数据集包含训练集、开发集和保留集,其中保留集仅作为元数据公开。数据集的构建基于四种来源模式(trace-derived、programmatic、multi-llm-synthesis、hand-authored),并通过生成控制变体创建偏好对。数据集适用于训练偏好模型或进行ORPO、DPO、SimPO等实验,但不建议用作广泛的指令遵循基准或直接针对保留集进行调优。
The Tenacious Bench Path B Preference dataset is designed to evaluate preference pairs for specific failure modes in B2B outbound work. The dataset includes training, development, and held-out sets, with the held-out set only available as metadata. The dataset is constructed from four source modes (trace-derived, programmatic, multi-llm-synthesis, hand-authored) and creates preference pairs by generating controlled worse variants. It is suitable for training preference models or conducting experiments like ORPO, DPO, and SimPO, but not recommended as a broad instruction-following benchmark or for direct tuning against the held-out partition.





