A Leakage-Audited Forward-Prediction Benchmark for Research-Area Persistence in Doctoral Training
收藏资源简介:
This record carries the frozen data for a five-discipline benchmark of research-area persistence in doctoral training: given only what is observable by a student's fifth career year, does the student still work in the advisor's area a decade later? The accompanying paper is "A Leakage-Audited Forward-Prediction Benchmark for Research-Area Persistence in Doctoral Training" (submitted to the KDD 2027 Datasets and Benchmarks Track). Five advising genealogies, 68,235 student-advisor pairs, one protocol, frozen in hash-pinned tables under a time contract: every model input is observable by the student's fifth career year, the label is read fifteen years out, and an assertion harness checks the contract on every build rather than asking a reader to trust it. FINDINGS THE DATA SUPPORTS Persistence is predictable, unevenly: a strong tabular model sets the ceiling in every discipline. Of four graph architectures only two clear that ceiling, in one discipline of the five; every exceeding cell is chemistry, and each one is flagged and audited. The true advisor beats a cohort-matched placebo in all five disciplines, and the predictable part of the outcome is carried mainly by the student's own early topic autocorrelation. Two instruments on disjoint inputs agree with the label only at fair levels. Five construction choices are sized against a 0.0013 determinism floor for repeated runs of one cell, and four sit above it; the smallest of the four is the only evaluated choice that changed a reported label, and the largest changed none. The evaluation's own history ships with the artifact: a naive first aggregation reported seven graph crossings, the audit traced them to protocol deviations rather than to the data, and the corrected protocol and the full audit are part of the release. CONTENTS zenodo_archive.zip is unchanged from the previous version, byte for byte, so every hash printed in the paper still verifies. New in this version are the five per-author concept event tables, concept_events_econ.parquet, concept_events_math.parquet, concept_events_neuro.parquet, concept_events_physics.parquet, and concept_events_chemistry.parquet: one row per work and concept, dated, so a reader can re-derive the early profiles, the label, and the time contract itself from events rather than reading a recorded verdict. The repository's verifier expects them at data/supplement/. The archive's reproduce_assertions.py is the original script, retained unchanged for archive integrity; the maintained harness lives in the repository and runs 90 checks with this deposit in place, 59 without it. The archive's DATASHEET.md is likewise the first-release original, retained unchanged; the maintained datasheet is in the repository. USE Unpack zenodo_archive.zip into data/ of a repository clone, and place the five concept event tables at data/supplement/. The four leakage certificates rerun on a CPU. Reproducing the published numbers from the frozen tables needs no credentials; rebuilding the tables from scratch needs an OpenAlex key. Code and documentation: https://github.com/Xinke-Li/who-inherits-the-field LICENSES AND CITATION The tables derive from OpenAlex (CC0) and Academic Family Tree (CC BY 4.0) records; the derived tables are CC BY 4.0, and the code and datasheet are MIT. As a use policy, the dataset must not be used to rank or evaluate individual researchers; its target is research-area continuity, not merit. Cite the concept DOI 10.5281/zenodo.21403539, which always resolves to the latest version; citation metadata, including the paper reference, is in CITATION.cff.



