Sycophancy in Language Models: Reproducible Scoping-Review Search, Screening, and Taxonomy-Grounding Data and Code
收藏资源简介:
Reproducibility package for the postgraduate specialization final paper Sociological and Philosophical Aspects of Sycophancy in Language Models: A Literature Review and a Taxonomy (PUC-Rio, 2026). The archive contains the complete, re-runnable pipeline for a PRISMA-ScR scoping review of sycophancy in language models, executed on 8 July 2026: scripts/ — PowerShell scripts that query three open databases (arXiv, OpenAlex, Semantic Scholar), deduplicate by normalized title, apply deterministic two-stage lexical screening, and ground the taxonomy categories in the retrieved corpus. data/raw/ — the raw records retrieved from each database. data/processed/ — the deduplicated pool, the full per-record screening decisions, the list of included studies, the PRISMA counts, and the taxonomy-evidence table. CODEBOOK.md, PRISMA.md, ENVIRONMENT.md — screening codebook, PRISMA flow, and environment/versioning notes. Selection flow: 1123 records identified → 577 duplicates removed → 546 screened → 458 assessed for eligibility → 247 studies included. Licensing: source code is released under the MIT License; the derived data are released under CC-BY-4.0. See LICENSE in the archive.



