Data archive for a four-axis ASME V&V 40 credibility assessment of LLM-based synthetic-patient simulation across 12 models
收藏资源简介:
Data archive for the article "The choice of large language model reverses the simulated go/no-go decision in digital mental-health trial planning: a four-axis credibility assessment of synthetic-patient simulation across 12 models". It contains the simulation outputs and the aggregate analysis-ready results from which every reported figure and table is built: the frozen archetype-grid runs across the 12-model panel (architecture_outputs/, 31 files), the decoding-variance replicates (decoding_variance/, 16 files), the aggregate results for each assessment axis (analysis_ready/, 15 files), and the sycophancy-probe records, judge verdicts and per-model scores (sycophancy/, 19 files), with a file manifest (DATA_MANIFEST.md). All records are synthetic outputs of the simulation framework; no human-participant data are redistributed. The external reference cohorts (Brighten, SSAQS, KAIST, LifeSnaps) are not included and remain governed by their own source terms — Brighten by a Synapse Data Use Agreement. The third-party sycophancy probe prompts are not redistributed; the deposited response files retain item identifiers and prompt hashes so they can be joined back to the source dataset. The analysis code is deposited separately under the MIT License.



