kirutew17654321/tenacious-bench-v0.1
收藏资源简介:
Tenacious-Bench v0.1是一个包含260个任务的机器可验证评估基准,专为Tenacious Conversion Engine AI外呼销售代理设计。其目的是测量τ²-Bench无法评估的故障模式,如信心校准的措辞、陈旧性披露、放弃路由和线程隔离。数据集分为训练集(143个任务,用于LoRA微调)、开发集(55个任务,用于验证/提示迭代)和保留集(62个任务,用于独立Delta B验证)。每个记录包含任务ID、类别、来源模式、输入(潜在客户上下文和代理提示)、预期输出和评分标准。
Tenacious-Bench v0.1 is a 260-task, machine-verifiable evaluation benchmark for the Tenacious Conversion Engine AI outbound sales agent. Purpose-built to measure the failure modes that τ²-Bench cannot grade: confidence-calibrated phrasing, staleness disclosure, abstention routing, and thread isolation. The dataset is divided into train (143 tasks, for LoRA fine-tuning), dev (55 tasks, for validation/prompt iteration), and held_out (62 tasks, for independent Delta B verification). Each record contains task_id, category, source_mode, input (prospect_context and agent_prompt), expected output, and scoring criteria.




