abhayesian/answers-with-reasoning-apps
收藏资源简介:
该数据集是一个关于代码问题的数据集,使用了Qwen/Qwen3-8B模型进行自我蒸馏,并在codeparrot/apps(面试难度)上进行了推理。数据集保留了最终答案与黄金解决方案匹配的推理过程。数据集是项目《Eliciting Trying Hard: Does Reasoning Generalize Across Domains?》中使用的三个兄弟数据集之一。数据集包含模型的完整推理轨迹以及答案。每个条目包含多个字段,如id、prompt、messages、reasoning、answer等。数据集还包含了采样设置、接受过滤器、评分方法、统计数据以及已知的限制和注意事项。
Self-distilled `Qwen/Qwen3-8B` (instruct, reasoning ON) rollouts on **codeparrot/apps (interview difficulty)**, filtered to keep only rollouts whose final answer matches the gold solution. This is one of three sibling datasets used in the project *Eliciting Trying Hard: Does Reasoning Generalize Across Domains?* The dataset contains the models full reasoning trace plus the post-`</think>` answer. Each row includes fields such as id, prompt, messages, reasoning, answer, etc. The dataset also includes sampling settings, acceptance filters, grading methods, statistics, and known limitations and caveats.




