David-beakr/agent-harness-seed
收藏资源简介:
Agent Harness Seed v0.3 是一个小型数据集,包含16个单轮和多轮代理案例,专门用于评估代理工具的可靠性,重点关注指令遵循、工具选择与避免使用错误工具、跨轮次状态跟踪以及在需要时请求澄清等核心能力。该数据集基于tau-bench和BFCL两篇论文的理论基础,并针对Beakr内部需求进行了扩展,包括澄清和格式遵循等行为。数据集采用混合检查方法,结合确定性检查(如字符串匹配、工具调用验证)和LLM法官评估,以确保语义正确性。
Agent Harness Seed v0.3 is a small dataset of 16 single-turn and multi-turn agent cases designed to evaluate agent harness reliability, focusing on pillar 1: instruction-following, tool choice (and refraining from wrong tools), state tracking across turns, and asking for clarification when needed. It is grounded in two anchor papers (tau-bench and BFCL) and extends beyond them with Beakr-specific behaviors like clarification and format adherence. The dataset employs a hybrid checking methodology, combining deterministic checks (e.g., string matching, tool call verification) with LLM judge evaluations to ensure semantic correctness.




