dataset-A-routing-eval
收藏资源简介:
数据集A是一个用于评估3个LLM模型和3个路由系统在6种能力上的分层数据集。总共有5339行数据,分为多个配置,包括编程、数学推理、规划代理、指令遵循、世界知识和创意合成。数据集包含来自多个来源的数据,如AIME-2025、BFCL-v4和Custom-Validated等,每个来源有不同的许可证。数据模式包括查询ID、查询内容、维度、来源、输入令牌数、预期答案等字段。所有提示和少样本示例均为英文。数据集还提供了关于令牌计数、自定义创意内容和合成少样本池的可重复性说明。
Dataset A is a hierarchical dataset designed to evaluate 3 LLM models and 3 routing systems across 6 capability dimensions. It contains a total of 5339 rows, divided into multiple configurations covering programming, mathematical reasoning, planning agents, instruction following, world knowledge, and creative synthesis. The dataset incorporates data from multiple sources including AIME-2025, BFCL-v4, Custom-Validated, among others, with each source holding its own distinct license. Its data schema includes fields such as query ID, query content, capability dimension, source, number of input tokens, expected answer, and more. All prompts and few-shot examples are provided in English. The dataset also provides reproducibility instructions regarding token counting, custom creative content, and the synthesized few-shot pool.




