Data and Analysis Code for "What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development"
收藏资源简介:
These data and analysis scripts support an in-silico investigation of the computational evaluator between AI-assisted item generation and expert psychometric review. Using Big Five personality item development as a test bed, the project crosses two open-weight generators with two compound generation-and-gating packages, producing 400 adaptive generation tasks and 32,000 selected item occurrences. The same generated content is evaluated across five embedding configurations and two structural methods. Study 1 examines whether semantic representation changes the construct evidence attached to fixed items, their recovered community structure, item retention, and content coverage. Study 2 follows the paired structural evidence through inclusive and agreement-based eligibility policies to determine whether those decisions change the content and wording assembled into candidate forms for expert review. The deposit provides analysis-ready Parquet data and plain Python analysis scripts. The data preserve exact prompts and raw model responses; parsed, rejected, surplus, and selected candidate occurrences; task and item lineage; unrounded embedding vectors; structural contexts and item-event histories; stage-specific evidence; missingness; and analytic noncompletion. The scripts use the public Parquet files to reproduce the study tables and supporting analyses, including construct-alignment, representation-sensitivity, item-reduction, content-coverage, policy-comparison, and candidate-form analyses. Supporting documentation describes the scientific design, variables, provenance, software environment, and execution of the analyses. All materials concern pre-response evidence. Generated items and candidate forms are inputs to expert psychometric review, not validated scales or participant-ready instruments. The deposit does not establish respondent interpretation, reliability, validity, measurement invariance, or the general superiority of any generator, embedding model, structural method, generation package, or candidate-form policy.



