Linguistic Stability and Journal-Level Heterogeneity in Full-Text Scientific Writing (2021-2025)
收藏资源简介:
This record provides the Zenodo-safe dataset and documentation for the study “Linguistic Stability and Journal-Level Heterogeneity in Full-Text Scientific Writing (2021-2025)”. It includes (i) corpus metadata for 240 peer-reviewed full-text articles sampled from six dentistry-related journals (2021 vs 2025; 20 articles per journal per year) and (ii) a per-article results table containing the derived linguistic metrics used to generate all tables and figures in the manuscript. In addition, the repository includes the LLM-marker lexicon and the exact matching rules used to compute the primary outcome (LLM-associated marker density), along with a statistical output file and brief documentation. Full-text articles are not redistributed due to copyright restrictions. The metadata includes article identifiers, where available, to facilitate independent access via the original publishers. Files: Corpus_Metadata_v1.0.xlsx - corpus metadata (journal, year, available identifiers) Master_Results_v1.0.csv - per-article derived metrics (240 rows) statistical_tests_v1.0.csv - summary statistics and tests reported in the manuscript LLM_Marker_Lexicon_v1.0.txt - prespecified 15-token marker lexicon (primary outcome) LLM_Marker_Matching_Rules_v1.0.txt - tokenization and exact-matching specification README_v1.0.txt - brief file-level documentation Contact: Hadar Better (bettermed@gmail.com)



