Intra-reviewer reliability re-screen of a title-and-abstract screening boundary: instrument, decision records and recomputation script
收藏资源简介:
This deposit holds the instrument, the decision records and the computationbehind a registered reliability re-screen conducted within the scoping review"AI Disclosure and Verification in Research" (OSF registration10.17605/OSF.IO/TNA23), and reported to that registration as Amendment 9 on27 September 2026. The review is conducted by a single reviewer. JBI guidance specifies tworeviewers at screening; that standard is not met, and the review does notclaim to meet it. In its place the review registered, in advance, a re-screenof a random ten per cent of the records that entered title-and-abstractscreening, to measure how consistently one reviewer applied the registeredinclusion criteria. The sampling fraction, the seed, the blinding, a limit onrecords per sitting, and a three-branch decision rule keyed to the resultingkappa were all fixed before the sample was screened. WHAT WAS DONE 105 records (10.02 per cent of 1,048) were drawn under random seed 20260925and presented in randomised order under the same seed. The instrument showedthe inclusion criteria as originally registered and offered three responses:In, Out and Unsure. It carried no indication of the first-pass decision, whichhad been extracted and held aside beforehand. Screening took place between 25and 27 September 2026 in four sittings of 5, 25, 47 and 28 records. WHAT WAS FOUND Observed agreement between the two passes was 70.5 per cent (74 of 105) andCohen's kappa was 0.447. Observed agreement was 61.3 per cent among recordspreviously included and 83.7 per cent among records previously excluded; kappacannot be computed within either stratum, where one rater's marginaldistribution is degenerate. Thirty-one records received a different response,with net movement toward exclusion. Two sensitivity analyses over thetreatment of Unsure return kappa 0.490 and 0.492. On the Landis and Kochbenchmarks the value is moderate, at the lower end of that range. The consequence, under the branch registered in advance, is that the reviewdoes not represent its inclusion boundary as reliable: the figure is reportedwherever an absence is claimed, and every claim of absence is bounded to thecharted corpus at the point it is made rather than stated as an absence inthe literature or in the field. The re-screen is a measurement and not a second adjudication. The recordswhose response differed retain their first-pass decisions, and the corpuscarried forward is unchanged. Had the sampled records been re-decided, onetenth of the corpus would stand screened to a different standard from theremainder, introducing the very inconsistency the exercise set out to measure. CONTENTS The screening instrument exactly as delivered; the three dated exports itproduced, of 5, 77 and 105 decisions; the first-pass decisions held asidefor the sampled records; a record-level results table giving both decisions,presentation order, stratum and timestamp for all 105 records; the 1,048record identifiers in the order held when the sample was drawn; and a Pythonscript that reproduces the draw from the seed and recomputes every reportedfigure, printing each beside the value as filed. A README describes each fileand a manifest gives SHA-256 checksums. The script requires Python 3.8 or later and no third-party packages. It exitsnon-zero and names any divergence. WHAT THIS DEPOSIT DOES NOT ESTABLISH It does not establish that any individual screening decision was correct. Eachdecision, on both passes, is the reviewer's own. DECLARATION OF AI ASSISTANCE The instrument, the random draw, the agreement arithmetic and therecomputation script were produced with AI assistance (Anthropic Claude),within the categories permitted by the review protocol. No screening decisionwas generated, checked, ranked, pre-sorted or confirmed by AI. Seeds, inputsand computation are deposited so that every reported figure can be recomputedindependently of both the reviewer and the assistant.



