遇见数据集

Em—ergence of the em—dash: a population-level rise in em-dash frequency in frequency in medRxiv preprints at the dawn of the large-language-model era

收藏
Zenodo2026-06-05 更新2026-06-12 收录
官方服务:

资源简介:

Reproduction package for the pre-registered study of em-dash frequency in medRxiv preprints. This archive accompanies the confirmatory study "em—ergence of the em—dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era" (OSF pre-registration HFT8C). It contains everything needed to reproduce the reported numbers: the frozen per-paper dataset, the analysis code, the canonical result files, the figures, and the manuscript. It is a tool for reproducing the results, not a record of the working process. Study in brief. The study asks whether the em-dash (the Unicode character U+2014) became more frequent in the Discussion sections of medRxiv preprints after the public release of ChatGPT (30 November 2022). The full-text medRxiv corpus was retrieved on 26 May 2026 from the server's official Amazon S3 Text-and-Data-Mining resource as JATS XML, which (unlike PubMed or the typographically normalised medRxiv HTML) preserves the author's original em-dash as the literal character. The primary cohort comprised 69,632 first-version preprints deposited in 2020–2025 with a Discussion section of at least 500 characters. The primary endpoint was the presence of at least one em-dash in the Discussion, and the primary estimand was the absolute change in its prevalence between the pre- and post-ChatGPT eras. Headline result. The proportion of Discussion sections containing at least one em-dash rose from 4.23% before ChatGPT to 11.58% after, an absolute increase of +7.35 percentage points (95% CI 6.94–7.77; odds ratio 2.96, 95% CI 2.77–3.17, with standard errors robust to clustering by first author). The rise was a delayed acceleration rather than a step: prevalence held near 4% through 2023, reached 8.0% in 2024, and 20.3% in 2025. The effect survived all six sensitivity analyses (7.35–7.60 pp) and both falsification tests; a placebo temporal split within the pre-LLM era showed no meaningful change (+0.13 pp). What it is and is not. The em-dash is a population-level marker, not a per-paper detector. It cannot establish whether any individual manuscript was produced with artificial intelligence, and the observational design cannot establish causation. What the data show is that a large, abrupt, temporally specific change in how clinical preprints are written appeared in close coincidence with the mass availability of LLM-assisted writing, and is compatible with it. Contents. data/processed/ — the frozen per-paper measurements (stage2_em_extraction.csv and the …_with_doidate.csv variant that adds the canonical DOI-derived deposit date), with a full CODEBOOK.md; the exploratory-overlap DOI list; and supporting tables. code/ — extraction, date-derivation (derive_doi_dates.py), the confirmatory analysis, supporting/sensitivity/falsification analyses, the lexical-marker and disclosure analyses, the blind-replication harness, and the figure generators. results/ — canonical result objects (per-year prevalence, segmented interrupted-time-series data, sensitivity summary, lexical-marker tables) and the final figures (PDF + PNG, 300 dpi). the manuscript and codebook. Data provenance and redistribution. The underlying medRxiv corpus is public external data and is not redistributed here. This archive contains only derived, per-paper measurements (section lengths and em-dash counts, identifiers, and the DOI-derived deposit date) together with the code that produced them. The original full text can be obtained from the medRxiv Text-and-Data-Mining resource. Reproducing the headline numbers. bash pip install -r requirements.txt python code/derive_doi_dates.py # canonical deposit dates from the DOI python code/stage2_analysis_v3_dualprefix.py # Expected: N = 69,632; pre 4.23% → post 11.58%; +7.35 pp; OR 2.96 All confirmatory, supporting, and sensitivity analyses were also independently re-implemented from the pre-registered specification, blind to the reported estimates, and reproduced the stated values. Related work. The choice of a typography-preserving source is examined in a companion audit of Unicode fidelity across biomedical bibliographic APIs (Czuma, 2026; OSF https://osf.io/269b5), which quantifies how little typographic punctuation survives in PubMed and OpenAlex.

提供机构:
Zenodo
创建时间:
2026-06-05
二维码
社区交流群
二维码
科研交流群
商业服务