exolabs/tau2-evals
收藏资源简介:
该私有数据集包含来自Hermes Agent在本地vLLM服务的NVFP4、FP8、BF16模型上运行的tau2-bench电信领域评估痕迹。每个任务是一个多轮客户支持对话:GPT-4.1用户模拟器扮演客户,被评估模型扮演支持代理(可访问tau2电信领域堆栈的终端工具集),官方tau2验证器根据真实目标对最终环境状态进行评分。评分是每个任务的二进制评分,主要指标是平均奖励(即二进制评分的pass@1)。数据集布局参考了exolabs/tau2-hermes-nemotron-eval-data,每个完成的运行都位于顶级目录<model-shortname>_<TS>/下。
This private dataset comprises evaluation traces generated during the tau2-bench telecom domain evaluation executed by Hermes Agent on NVFP4, FP8, and BF16 models deployed via local vLLM services. Each task takes the form of a multi-turn customer support conversation: the GPT-4.1-powered user simulator assumes the role of a customer, while the model under evaluation acts as the support agent with access to the terminal toolset of the tau2 telecom domain stack. The official tau2 validator scores the final environment state against predefined real-world objectives. Scoring for each task follows a binary rating scheme, with the primary evaluation metric being average reward (equivalent to pass@1 of the binary ratings). The dataset layout is referenced from exolabs/tau2-hermes-nemotron-eval-data, where each completed run is stored under the top-level directory named <model-shortname>_<TS>.




