Screening records for a scoping review of demographic equity in medical AI evaluations, with a focus on Middle Eastern and Arabic-speaking populations (pilot round, PubMed)
收藏资源简介:
Screening records supporting the pilot round of a JBI-framework scoping review examiningwhether evaluations of medical artificial intelligence report performance stratified bydemographic variables, with particular attention to Middle Eastern and Arabic-speakingpopulations. Protocol registered at https://osf.io/pnwf7 (24 July 2026), amended 4 August 2026. This deposit contains four records: 1. Screening log, full record — all 145 records screened at title and abstract, each with its call, the abstract language supporting that call, and the author's verification note. Includes the method and provenance table, the screening rubric, the search strategy with its declared limitations, and the screening concordance analysis. 2. Search record, 24 July 2026 — the contemporaneous record of both Boolean searches executed against PubMed via the NCBI E-utilities API, with both strings, both date ranges, both total-match counts, and all retrieved records listed. Its decision columns are intentionally empty; screening was conducted through the two-stage process documented in record 1. 3. Tier assignments — final eligibility dispositions with reasoning, the two-tier eligibility structure adopted at amendment, PRISMA-ScR counts, and the round-2 search strategy. 4. Record list — all 145 screened records identified by PubMed ID and search arm. Because the general search retained the top 50 of 889 matches by relevance ranking, and PubMed relevance ranking is not temporally stable, this list rather than the query is the authoritative record of what was screened. Search strategy. Two Boolean searches were executed against PubMed via the NCBI E-utilitiesAPI on 24 July 2026. A general search requiring one of five exact bias phrases combined withartificial intelligence terms and clinical terms, over 2020-2026, returned 889 matches, ofwhich the top 50 by relevance ranking were screened. A regional supplement combining abroader bias and algorithm block with Middle Eastern and Gulf region terms, over 2015-2026,returned 95 matches, all of which were screened. Screening procedure. Title and abstract screening was conducted in two stages. A largelanguage model (Claude Sonnet 5, Anthropic) applied a pre-specified two-question rubric toall 145 retrieved abstracts, recording each call with the abstract language supporting it.The author then reviewed all 145 records against the same abstracts, confirming, overriding,or resolving every call, and retrieving full text where the abstract was insufficient.Concordance is reported as a verification agreement rate and explicitly not as aninter-rater reliability coefficient: the author's review followed exposure to the model'sreasoning, so anchoring bias cannot be excluded. This limitation is stated in the records. Limitations. Single database. The general arm screened 50 of 889 records and its findingsare limited to that subset. The general query's reliance on exact bias phrases, combinedwith relevance-ranked sampling and no publication-type exclusion, returned 49 reviews andcommentaries against a single primary study, which is a property of the query rather than ofthe literature. Proportions derived from these records are bounded by the query and thesingle database that produced them and are not presented as coverage of the widerliterature.



