遇见数据集

Register, Not Scale: Predicting Arabic Lexical Decision Latencies from Twelve Language Models (data and code release)

收藏
Zenodo2026-07-25 更新2026-08-02 收录
官方服务:

资源简介:

Data and code release accompanying the paper Register, Not Scale: PredictingArabic Lexical Decision Latencies from Twelve Language Models, submitted toIEEE/ACM Transactions on Audio, Speech, and Language Processing. Twelve Arabic and multilingual language models were scored on the item set ofthe Arabic Lexicon Project, and the resulting isolated word probabilities wereevaluated against 809,033 lexical decision trials from 604 first language and43 second language readers. The release contains the model derived estimates,the human behavioral data they were evaluated against, the tokenizer alignmentmeasurements against CAMeL Tools gold morphological segmentation, every fittedstatistic reported in the paper, and the seven notebooks that produce all ofthem. Psychometric fit is reported as the improvement in log likelihood per 1,000trials when standardized surprisal is added to a baseline mixed effects modelof log reaction time. The primary analysis uses a maximal by participant randomeffects structure, with a by participant random slope for surprisal retained inboth the baseline and the full model so that the likelihood ratio test isolatesthe fixed effect. The conventional random intercepts structure is retained as adocumented comparison, because the difference between the two turns out to bearchitecture dependent and is itself reported as a finding. The human behavioral data is included, and it is not ours. The trial levelrecords, the item level aggregate reaction times, and the stimulus set alloriginate with the Arabic Lexicon Project, which releases them under the MITLicense. They are redistributed here under those terms, with the requiredcopyright notice in LICENSE-UPSTREAM-ALP. Anyone using this deposit must citeAlzahrani et al. (2026), DOI 10.3758/s13428-026-03115-9. UPSTREAM_DATA.mdrecords column by column which data is theirs and which is ours, and documentsone property of the trial file, a set of exactly duplicated records present inthe source release, that anyone reanalyzing it should know about. Because the behavioral data travels with the estimates, the analysis isreproducible end to end from this deposit alone: 239,988 item level modelestimates over twelve models by 19,999 items, 21,999 item level humanaggregates, and the 809,033 trials the mixed effects models were fitted on. See README.md for the layout and SCHEMAS.md for a column by column descriptionof every file. Funding: this research was funded by Ongoing Research Fundingprogram, (ORF-2026-276), King Saud University, Riyadh, Saudi Arabia.

提供机构:
Zenodo
创建时间:
2026-07-25
二维码
社区交流群
二维码
科研交流群
商业服务