brain-lm-alignment-ds002236
收藏资源简介:
该数据集基于Lytle et al. 2020关于学龄儿童(8.7-15.5岁)正字法、语音和语义词汇加工的研究,采用听觉和视觉两种刺激模态。数据集构建了30个任务×session单元,每个单元生成一个表征相异矩阵(RDM),覆盖所有被试共享的刺激,体素模式在每个run内进行z-score标准化后聚合,并计算了被试间噪声上限。每个单元的刺激数量为72或96个。控制测试评估了刺激持续时间、强度、词长、频率、音素和音节数量、听觉刺激的声学模型以及实验本身的对比条件,但所有30个测试在Holm校正后均未达到显著性,表明这些RDM并未可证明地编码刺激,因此后续的模型对齐结果不可作为语言模型与大脑对齐的证据。数据集包含了15个模型家族(如pico-decoder、babylm-gpt2、pythia、beagle等)的共7260个对齐行,报告了每个家族的对齐均值、标准差、最大绝对值、占噪声上限的比例以及等效性检验(TOST ±0.05)结果。所有15个家族均被判定为等效于零。此外还提供了pythia尺度趋势(ρ = -0.221, p = 0.24)。RDM有效秩为84个刺激中的56。数据集文件包括对齐结果(按检查点、家族、单元)、噪声上限、控制测试结果、尺度梯度和图表。该数据集适用于研究表征相似性分析(RSA)方法,但注意其控制测试失败,因此不能用于评估语言模型与大脑发育的对齐程度。
This dataset is based on the study by Lytle et al. (2020) on orthographic, phonological, and semantic lexical processing in school-age children (8.7-15.5 years), using both auditory and visual stimulus modalities. The dataset consists of 30 task×session units, each generating a Representational Dissimilarity Matrix (RDM) covering stimuli shared across all subjects. Voxel patterns were z-score normalized within each run, aggregated, and the noise ceiling across subjects was computed. Each unit contains 72 or 96 stimuli. Control tests evaluated stimulus duration, intensity, word length, frequency, phoneme and syllable count, acoustic model of auditory stimuli, and contrast conditions of the experiment itself. However, all 30 tests were non-significant after Holm correction, indicating that these RDMs do not demonstrably encode stimuli, and thus subsequent model alignment results cannot be used as evidence of language model-brain alignment. The dataset includes 7260 alignment rows from 15 model families (e.g., pico-decoder, babylm-gpt2, pythia, beagle), reporting alignment mean, standard deviation, maximum absolute value, proportion of noise ceiling, and equivalence test (TOST ±0.05) results. All 15 families were judged equivalent to zero. Additionally, the pythia scaling trend (ρ = -0.221, p = 0.24) is provided. The effective rank of the RDM is 56 out of 84 stimuli. Dataset files include alignment results (by checkpoint, family, unit), noise ceiling, control test results, scaling gradients, and figures. This dataset is suitable for studying Representational Similarity Analysis (RSA) methods, but note that its control tests failed, so it cannot be used to evaluate alignment between language models and brain development.
数据集概述:Brain–language-model alignment: ds002236
基本信息
- 数据集名称: Brain–language-model alignment: ds002236 (whole-brain)
- 来源研究: Lytle et al. 2020 — 学龄儿童(8.7–15.5岁)的正字法、音韵和语义词汇处理,包含听觉和视觉两种模态。
- 论文: https://pubmed.ncbi.nlm.nih.gov/31956678/
- 原始数据: https://openneuro.org/datasets/ds002236/versions/1.0.1
- 生成日期: 2026-09-07
- 处理流程: https://github.com/suchirsalhan/cdl-representations-brains-babylms
质量控制警告(GATE: FAILED)
- 结论: 30项刺激测试中0项显著(经Holm校正),包括对儿童实际听到的音频声学模型以及研究自身的实验对比均未达到显著水平。
- 含义: 该数据集的对齐数值不能作为语言模型证据来解释,因为所测量的表征几何结构未能明确编码刺激。数据集仅供参考或用于改进评估方法,不应引用为模型与发育中大脑对齐失败的证据。
- RDM有效秩: 84个刺激中有61个,说明RDM确实携带刺激结构,控制失败是由于特定控制未达到显著性,而非测量本身完全不可解释。
数据构建
- 结构: 54个任务×会话单元,每个单元基于共享刺激构建RDM(表征相异性矩阵)。
- 预处理: 体素模式在运行内进行z-score标准化(避免测量扫描漂移),并计算被试间噪声上限。
- 任务类型: 两个任务(Phon 音韵任务、Sem 语义任务),多个会话组合(ses-9、ses-11、ses-11+——组合会话)。
- 刺激数量: 音韵任务96个、语义任务72个。
- 排除项: 三分之一的试验为无效编码(Tones/nullsilence.WAV),已从刺激集中排除;配对设计中刺激身份为配对整体而非单个词汇。
模型评估结果
总体统计
- 平均噪声上限: 0.247
- 最佳对齐值: 0.1066(约为噪声上限的80.4%)
- 等价于零的模型家族: 15/15(所有模型家族均无法与零区分,TOST ±0.05)
- Pythia规模趋势: ρ = +0.202, p = 0.29(不显著)
模型家族性能(共15个家族,覆盖14094行对齐数据)
| 家族 | 检查点数量 | 平均RSA | 标准差 | 绝对最大对齐 |
|---|---|---|---|---|
| pico-decoder-medium | 21 | 0.0222 | 0.0237 | 0.1066 |
| pythia-1b-full | 21 | 0.0215 | 0.0157 | 0.0816 |
| babylm-gpt2 | 9 | 0.0190 | 0.0252 | 0.0501 |
| pico-decoder-large | 20 | 0.0189 | 0.0210 | 0.0804 |
| babylm-gpt2-7 | 9 | 0.0166 | 0.0266 | 0.0581 |
| pythia-1.4b-full | 21 | 0.0165 | 0.0189 | 0.0999 |
| babylm-gpt2-3 | 9 | 0.0165 | 0.0259 | 0.0582 |
| pico-decoder-small | 21 | 0.0161 | 0.0166 | 0.0978 |
| pythia-410m-full | 21 | 0.0151 | 0.0118 | 0.0729 |
| babylm-gpt2-5 | 9 | 0.0131 | 0.0259 | 0.0477 |
| pythia-160m-full | 21 | 0.0122 | 0.0114 | 0.0771 |
| pythia-70m-full | 21 | 0.0121 | 0.0140 | 0.0808 |
| pico-decoder-tiny | 21 | 0.0119 | 0.0137 | 0.0799 |
| beetle-humanscale-eng | 18 | 0.0091 | 0.0050 | 0.0727 |
| beetle-fineweb3-eng | 19 | 0.0001 | 0.0031 | 0.0420 |
数据集特色
- 被试年龄: 数据集文章未明确标注编号,通过匹配OpenNeuro数据集名称("Cross-Sectional Multidomain Lexical Processing")和参与者年龄范围(8.67–15.5)解析为ds002236。
- 发展轴优势: 四个数据集中最佳——有明确的每次扫描被试年龄,是连续性而非分箱数据。
- 任务设计: 6个任务跨越模态(听觉/视觉)与判断类型(押韵/拼写/语义),提供了其他数据集所缺乏的模态控制。
包含文件
alignment_by_checkpoint.csv: 每个模型×检查点×单元的对齐值及噪声上限alignment_by_family.csv: 每个模型家族的对齐值及等价性检验alignment_by_cell.csv: 每个任务×会话的对齐值ceilings_ds002236.csv: 每个单元的噪声上限control/: 阳性控制及RDM维度分析(即"门控"部分)scale_ladder.csv: Pythia 70M→1.4B规模测试fig_*.pdf,fig_*.png: 图表文件
方法
采用表征相似性分析(RSA):对每个单元,构建大脑RDM(基于每刺激GLM beta模式间的相关距离,运行内z-score,跨被试聚合),通过Spearman相关与同一刺激集上每个检查点隐藏状态计算的模型RDM进行比较。对齐值以原始值及占被试间噪声上限的比例报告,并通过PARC套件(18个仅随机种子不同的模型)构建的零分布来判断显著性。





