modal-labs/autoinference-review-replay-512-v1
收藏资源简介:
该数据集名为自动推理审查重放(512令牌解码目标),包含64个真实的代码审查重放请求,这些请求采用回合结构,并按轮询顺序排列以模拟生产环境中的基数缓存命中情况。每行数据设置max_tokens为512,这是用户为greptile新应用工作负载定义的解码目标,输入约为37.5k令牌/请求,前缀命中率约为92%。数据格式为每行一个{messages, max_tokens}对象,用于自动推理的OpenAI服务基准测试。数据来源于review_replay_v1_slice0(来自auto inference仓库)。
The dataset is titled auto inference review replay (512-token decode target) and contains 64 real code-review replay requests structured in turns, with round-robin order to reproduce the production radix-cache hit profile. Each row has max_tokens=512 as the user-set decode target for the greptile new-app workload, with input approximately 37.5k tokens per request and a prefix hit rate of about 92%. The format is one {messages, max_tokens} per line for the auto inference OpenAI serving benchmark. Source: review_replay_v1_slice0 (from the auto inference repository).




