yc-bench
收藏资源简介:
YC-Bench是一个专为LLM代理设计的长时程确定性基准测试数据集。该数据集模拟了一个AI初创公司在一年内的运营情况,代理通过CLI工具与SQLite支持的离散事件模拟系统交互,扮演CEO角色。主要测试代理在数百次决策中管理复合决策的能力,包括声望专业化、员工分配、现金流、截止日期风险和对抗性客户检测等。数据集适用于文本生成任务,属于基准测试、代理、长时程、模拟和评估类别。数据集规模小于1K,采用Apache 2.0许可证。使用指南包括设置、评估命令、评分方法和结果提交规范。
YC-Bench is a long-duration deterministic benchmark dataset specifically designed for LLM agents. This dataset simulates the one-year operation of an AI startup, where agents act as CEOs and interact with a SQLite-backed discrete-event simulation system via CLI tools. It primarily tests the agents' ability to manage complex decision-making across hundreds of decisions, including reputation specialization, staff allocation, cash flow, deadline risk, adversarial customer detection, and more. The dataset is applicable to text generation tasks and falls under the categories of benchmarking, agents, long-duration, simulation, and evaluation. The dataset has a scale of less than 1K and is released under the Apache 2.0 license. Its usage guidelines cover setup, evaluation commands, scoring methods, and result submission specifications.



