Google Tunix Hack Custom Data
收藏官方服务:
资源简介:
Synthetic Reasoning Traces for LLM Process Supervision (GRPO/SFT)
面向大语言模型(LLM)过程监督的合成推理轨迹(GRPO/SFT)
创建时间:
2026-01-12

Synthetic Reasoning Traces for LLM Process Supervision (GRPO/SFT)
面向大语言模型(LLM)过程监督的合成推理轨迹(GRPO/SFT)