semianalysisai/cc-traces-weka-with-subagents-052726-256k
收藏资源简介:
CC Traces — Weka, With Subagents, in+out ≤ 256k (May 27 2026, deduped) 是一个256k上下文模拟语料库,源自cc-traces-weka-with-subagents-052726数据集。该数据集经过过滤和去重处理:删除了每个请求(主代理和子代理内部)中输入加输出超过256,000代理标记符的条目,并重新调整了时间线以消除因删除请求而产生的间隙。数据集包含470条轨迹、51,739个主轮次、1,053个子代理组和26,119个子代理内部请求,总计77,858个模型请求。统计信息涵盖了标记符分布、轨迹跨度和模型组成(主要使用Claude模型如opus-4-7)。适用于文本生成、LLM推理、基准测试和多轮代理对话等任务。
CC Traces — Weka, With Subagents, in+out ≤ 256k (May 27 2026, deduped) is a 256k-context simulation corpus derived from the cc-traces-weka-with-subagents-052726 dataset. It has been filtered and deduped: every individual request (main-agent and sub-agent inner) where input plus output exceeds 256,000 proxy-tokenizer tokens has been dropped, and the request timeline is reshifted to collapse gaps created by removed requests. The dataset includes 470 traces, 51,739 main turns, 1,053 sub-agent groups, and 26,119 sub-agent inner requests, totaling 77,858 model requests. Summary statistics cover token distributions, trace spans, and model composition (primarily Claude models such as opus-4-7). It is suitable for text generation, LLM inference, benchmarking, and multi-turn agentic conversations.




