ChicagoHAI/UniCo-Completions-SFT
收藏资源简介:
该数据集包含由UniCo框架生成的66,603个SFT(监督微调)训练示例,源自论文《Towards a Universal Causal Reasoner》。它涵盖44种(表示形式,查询类型)配置,并以ShareGPT格式存储。响应通过拒绝采样从Qwen3-32B、Olmo-3.1-32B-Instruct和Qwen3.5-27B的集成模型中筛选,每个问题有预算2次采样,如果多个响应正确则随机选择一个,否则随机选择一个采样响应,响应长度限制在100-8192个令牌以避免异常推理轨迹。
This dataset contains 66,603 supervised fine-tuning (SFT) training examples generated by the UniCo framework, derived from the paper *Towards a Universal Causal Reasoner*. It covers 44 (representation type, query type) configurations and is stored in the ShareGPT format. Responses are filtered via rejection sampling from an ensemble of Qwen3-32B, Olmo-3.1-32B-Instruct, and Qwen3.5-27B. A sampling budget of 2 is set for each question: if multiple correct responses are obtained, one is randomly selected; otherwise, one sampled response is randomly chosen. The response length is restricted to 100–8192 tokens to avoid anomalous reasoning trajectories.



