Nurye/tenacious_bench_v0.1
收藏资源简介:
Tenacious-Bench v0.1是一个领域基准测试数据集,用于在Tenacious特定的业务和政策约束下评估销售代理的输出。该数据集评估代理是否能够正确处理以下问题:不支持的定价或范围声明、过度声称的信号或成熟度声明、通用/无根据的外联、错误的CRM/HubSpot/日历下一步行动、回复升级或异议处理失败。数据集包含200行数据,分为训练集(100行)、验证集(60行)和测试集(40行)。每行数据包含多个字段,如task_id、input、ground_truth等。
Tenacious-Bench v0.1 is a domain benchmark for evaluating sales-agent outputs under Tenacious-specific business and policy constraints. This dataset evaluates whether an agent correctly handles: unsupported pricing or scope claims, overclaimed signal or maturity claims, generic/ungrounded outreach, incorrect CRM/HubSpot/calendar next actions, reply escalation or objection-handling failures. The dataset contains 200 rows, divided into train (100 rows), validation (60 rows), and test (40 rows) splits. Each row contains fields such as task_id, input, ground_truth, etc.





