A Verbatim-Grounded, Field-Neutral Tag Layer for Cross-Disciplinary Reading of Academic Literature
收藏资源简介:
A field-neutral, verbatim-grounded tag layer for academic literature. Every paper is described with the same small set of faceted tags (claim type, phenomenon as a normalised cross-field join key, claim relations, replication status, limitations), under one design rule: every tag is licensed by an exact sentence from the paper's own text, and the audit reports each tag's provenance status rather than assuming full exact-body grounding. A proof-of-format on a curated 126-paper corpus spanning 39 disciplines (116 fully faceted). What the paper reports, including what did not work. Written as free text, the phenomenon names almost never collide across fields — exactly one pair of disciplines shares one verbatim — and a novelty scan surfaced no cross-disciplinary link the literature had not already drawn. Both null results are reported plainly. An audited controlled vocabulary of statistical terms does create cross-field joins, and 16 papers across 8 disciplines group under one canonical phenomenon (priming) despite each field naming it differently; a grounded audit of that cluster finds association and modulation in 16/16 papers and secondariness in 14/16, the two exceptions being genuine boundary cases. Reliability is measured, not asserted. All 519 legacy statistical-tag assignments were re-audited (420 accepted, 99 rejected, 19.1%). An independent, blinded second coder re-coded a stratified sample of 120 assignments, giving κ = 0.767 (95% CI 0.652–0.882), with the disagreement asymmetric in the direction of the original coding being the stricter of the two. A second coder is not a gold standard, and that limitation is stated in the paper. Version 11.1. The layer is published as Linked Open Data — a SKOS vocabulary and a W3C Web Annotation collection, about 47,800 triples — with the controlled vocabularies reconciled to Wikidata (39/39 disciplines, 13/13 methods, 15/25 statistical terms, 61/88 concepts; a link is asserted only where both a matching label and a licensing type statement could be shown). Re-expressing the tags as annotations re-verified 99.4% of the publishable verbatim spans independently of this project's own checker. Nine digital-humanities papers were added as a deliberate facet stress test, which the method vocabulary failed for 4 of the 9 — reported as a gap rather than papered over. This version also corrects two in-text citation numbers, several summary lines that misstated the paper's own findings, and the wording throughout. Two PDFs. meta-tagging-preprint_v19_SUMMARY-VIEW.pdf (13 pp) replaces every full paragraph and numbered point with one written sentence saying what it argues, keeping all figures and legends; meta-tagging-preprint_v19.pdf (22 pp) is the complete paper. The live page at shir-openu.github.io/meta-tagging-showcase/preprint.html shows the summaries with the full text one click behind each. Per-paper redistribution rights are verified in an explicit manifest: full text is reproduced only for papers under a verified redistributable licence; every other paper shows the tag layer with short licensing excerpts and a link to the source. Version 12.0. This version measures the decision the tag layer makes every time it assigns a tag: which items fall under a term — that is, a definition. Earlier versions established the layer and reported the null result that free-text names for a phenomenon almost never collide across fields (500 distinct strings over 511 papers, two crossing a disciplinary boundary). A controlled vocabulary was identified as the missing ingredient, and choosing what belongs in one is applying a definition. Added here: a definition is scored as a binary classifier over cases the literature itself adjudicated, after three logical gates that no statistical measure detects (circularity, recursion without a base case, and a free parameter the author never fixed). Inter-coder reliability is reported on the act of application itself (κ = 0.547, n = 140). Corpus at this version: 538 records across 86 disciplines; 18,416 evidence strings re-checked against their source, three failed.



