twangodev/devpost-hacks-judgments
收藏资源简介:
该数据集名为Devpost Hackathon Judgments,主要用于研究用途,包含了对黑客马拉松项目提交的成对LLM-judge追踪。每一行数据都是一个单聊格式的对话,其中助手比较两个项目并选择较强的一个(或TIE),并附有推理追踪。数据集包含两种不同的模型(Qwen/Qwen3.5-27B和Qwen/Qwen3.5-4B)的判决结果,总计63,044行。数据集的配置包括多个黑客马拉松的特定配置,以及一个默认的all配置,包含所有行。数据集的模式包括多个字段,如messages、judgment_id、pair_id等。数据集的使用方法、注意事项、来源和许可信息也在README中详细说明。
The dataset is named Devpost Hackathon Judgments and is intended for research use only. It contains pairwise LLM-judge traces over hackathon project submissions. Each row is a single chat-format conversation where the assistant compares two projects and picks the stronger one (or TIE), with a reasoning trace. The dataset includes judgments from two different models (Qwen/Qwen3.5-27B and Qwen/Qwen3.5-4B), totaling 63,044 rows. The dataset configurations include per-hackathon settings and a default all config that combines all rows. The schema includes fields such as messages, judgment_id, pair_id, etc. The README also provides detailed instructions on loading the dataset, caveats, sources, and licensing information.




