Language models struggle with compartmentalization - eval data
收藏资源简介:
Precomputed evaluation outputs for the paper Language models struggle with compartmentalization. Includes per-checkpoint validation losses on fineweb, finetune trajectories, cross-compartment cosine similarity sweeps, per-language multilingual curves, and raw training-time val-loss logs for the InfoNCE runs. Drop these into experiment/ of the companion code release (https://github.com/vinhowe/compartmentalization) and generate paper figures via provided plot scripts. See README.txt in this bundle for the file-by-file figure map. Version 2 (2026-05-19) Refreshed eval data — corrects a step/metric alignment bug in the previous bundle and bumps batch-averaging from N=10 to N=100.
本数据集为论文《语言模型在分隔任务(compartmentalization)中表现欠佳》的预计算评估输出结果,涵盖FineWeb(fineweb)数据集上各检查点的验证损失、微调轨迹(finetune trajectories)、跨分隔余弦相似度扫描、各语言的多语言性能曲线,以及InfoNCE(InfoNCE)运行的原始训练时验证损失日志。请将该数据集放入配套代码仓库(https://github.com/vinhowe/compartmentalization)的experiment/目录下,即可通过附带的绘图脚本生成论文所需图表。本压缩包内的README.txt文件包含了各文件与对应图表的映射关系。 版本2(2026年5月19日) 本次更新了评估数据:修复了此前数据包中存在的步骤与指标对齐错误,并将批次平均参数从N=10调整至N=100。



