neo4j/aip-skillbench-cohort-ab-sonnet-aipv0_3a2
收藏资源简介:
AIP-SkillBench — 16任务组合队列(Sonnet, AIP v0.3a2)是一个用于评估两种技能格式在相同任务上性能的数据集。具体来说,它比较了人类编写的原始技能(human-curated)和基于AIP规范v0.3a2编译的AIP技能(aip-from-curated)。数据集包含两个平衡的8任务队列(A和B),每个任务×模式进行了5次独立试验,总计160次运行。评估使用claude-sonnet-4-6模型作为求解器,并在docker沙盒环境中执行。该数据集暴露了AIP规范v0.3a2的两个失败模式(如脚本错误和过度工程化),从而推动了AIP规范升级到v0.3a3。数据集文件包括运行矩阵、状态总结、详细试验结果以及每个试验的工作目录和日志。
AIP-SkillBench — 16-task Combination Queue (Sonnet, AIP v0.3a2) is a dataset developed to evaluate the performance of two skill formats on identical tasks. Specifically, it compares human-curated original skills with AIP skills compiled against the AIP specification v0.3a2 (aip-from-curated). The dataset consists of two balanced 8-task queues (A and B), with 5 independent trials conducted for each task×mode, totaling 160 runs. The evaluation uses the claude-sonnet-4-6 model as the solver, with all executions carried out in a Docker sandbox environment. This dataset uncovered two failure modes of the AIP specification v0.3a2, such as script errors and over-engineering, which prompted the upgrade of the AIP specification to v0.3a3. The dataset files include run matrices, status summaries, detailed trial results, as well as the working directories and logs for each trial.




