HumanRef-CoT
收藏资源简介:
HumanRef-CoT是一个大规模的CoT式对象引用数据集,包含90,824个样本,由GPT-4o在HumanRef数据集上生成。每个样本都被注释为一个结构化的推理轨迹,遵循规划、行动和总结的范式,使得模型能够学习对对象候选者进行分解的、可解释的推理。该数据集支持Rex-Thinker模型的训练,该模型通过冷启动监督微调阶段和基于GRPO的强化学习训练,在HumanRef基准测试中取得了最先进的性能,并在域外场景和对象上展示了强大的泛化能力。
HumanRef-CoT is a large-scale Chain-of-Thought (CoT) style object reference dataset consisting of 90,824 samples generated by GPT-4o based on the HumanRef dataset. Each sample is annotated as a structured reasoning trajectory adhering to the paradigm of planning, action and summarization, enabling models to learn decomposable and interpretable reasoning for object candidates. This dataset supports the training of the Rex-Thinker model, which achieves state-of-the-art performance on the HumanRef benchmark through a cold-start supervised fine-tuning phase and GRPO-based reinforcement learning training, and demonstrates strong generalization capabilities across out-of-domain scenarios and objects.



