2026-09-07-mask-qwen36-0-dat-7
收藏资源简介:
该数据集来源于一次对模型 LASR-Callum/2026-09-07-qwen36-0-dat-7 的 mask 评估实验(运行模式为 think)。评估使用基础模型 Qwen/Qwen3.6-27B 作为基线,生成日期为 2026-09-07,并在 2026-09-10 进行了重新评分(重读每个评判的最后一个 Answer 行,并按论文 §4.3 方法进行汇总池化,最终准确率从 71.85 修改为 67.6)。数据集结构按模式分为三类:rollouts(包含完整的对话交互记录)、results(包含评估结果 JSON、评判输出等)、metadata(包含运行元数据、配置文件及溯源信息)。数据集无宪法约束,其生成配置为空字典。该数据集适用于评估模型在掩码任务中的表现,尤其适合对模型推理(reasoning)效果进行对比分析。
This dataset originates from a mask evaluation experiment (operation mode: think) on the model LASR-Callum/2026-09-07-qwen36-0-dat-7. The evaluation uses the base model Qwen/Qwen3.6-27B as a baseline, with a generation date of 2026-09-07. It was re-scored on 2026-09-10 (re-reading the last Answer line of each judgment and performing aggregated pooling according to §4.3 of the paper, with the final accuracy revised from 71.85 to 67.6). The dataset structure is divided into three categories by mode: rollouts (containing complete dialogue interaction records), results (containing evaluation result JSON, judgment outputs, etc.), and metadata (containing runtime metadata, configuration files, and provenance information). The dataset has no constitutional constraints, and its generation configuration is an empty dictionary. This dataset is suitable for evaluating model performance in masked tasks, especially for comparative analysis of model reasoning effects.





