Performance is necessary but not sufficient to make sense of large language models in evidence synthesis: a qualitative methodological study
收藏资源简介:
This serves as both the data supplement and the protocol/registration record for Pavel Zhelnov’s study “Performance is necessary but not sufficient to make sense of large language models in evidence synthesis: a qualitative methodological study.” This study is being conducted from January to May 2026 as part of comprehensive course requirements in the Health Systems Research PhD program at the Institute of Health Policy, Management and Evaluation within the Dalla Lana School of Public Health, University of Toronto. It contains raw hand-typed reports developed for the preparation of the protocol under analyses, raw data and PRISMA flow diagram source up to screening under data, plus some accompanying documentation for search sources used under docs. This also contains some software code used in this protocol. This includes src/render_matrix.py and accompanying test, plus pyproject.toml and pixi.lock for the matrix tool used to validate aislop renditions. This also includes Simon Willison’s codex-timeline tool under src/github.com/simonw to render an OpenAI Codex rollout from chats/chats-2026-05-13-codex-aislop. Unfortunately, due to time constraints the protocol could not be finalized; the unfinished protocol draft along with aislop AI-generated renditions are placed under protocol. Under protocol/release the original letter of intent form for this study submitted on December 15th, 2025, is placed.



