Replication dataset for "Toward Ludic-Aware Narrative Generation: A Neuro-Symbolic Framework for Evaluating Playability in LLM-Generated Backstories"
收藏资源简介:
Minimal replication dataset for the article "Toward Ludic-Aware Narrative Generation: A Neuro-Symbolic Framework for Evaluating Playability in LLM-Generated Backstories" (Applied Sciences, MDPI; Special Issue "Advances in Games and Immersive Technologies"). The article introduces two metrics for the playability of LLM-generated character backstories — the Ludic Potential Index (LPI) and the Potential Conflict Value (PCV) — and tests, interventionally, whether gating generation on those metrics reduces the downstream Adaptation Effort (Ea) required to make generated content playable. An expert study assesses whether the metrics track the judgement of practitioners. VERSION 1.1.0 accompanies the revised manuscript. It extends the expert study to 26 raters in one pool (495 ratings over 120 backstories; 21 AHP weight elicitations) and adds every analysis introduced in the revision. See CHANGELOG.md. Contents: - data/ — the 400 trials of both studies (master analytical file and each study as run); the 120 backstories rated by the expert panel, full text; 495 pseudonymous expert ratings with their collection round; 21 AHP pairwise-comparison submissions; weight-sensitivity analyses; the Notary extraction audit - revision/ — re-analysis under expert-derived LPI weights, with the re-run trials and the adaptation-effort re-measures for both weight vectors; PCV parameter and threshold sensitivity; gate decomposition; Ea normalisations; TH distribution; joint LPI x PCV tests; cost per backstory; the expert study as a whole and by collection round - config/ — frozen run configurations (weights, gate thresholds, seeds, prompts), the lore source used to condition generation, and the exact local model versions - protocol/ — the pre-registrations, the human-study and AHP protocols, and the dated record of every deviation - analysis/ — Friedman/Nemenyi multiple-comparison output and the weight-sensitivity write-ups A verification script is included. Running python3 verify_dataset.py recomputes the article's headline statistics from the shipped files and exits non-zero if any fails to reproduce: the per-arm delta-Ea values and the 2x4 ANOVA, gate convergence and entity attribution; the human construct-validity correlations with their Holm correction and intraclass correlations, overall and by collection round; the AHP group weights; the gating effect under the expert-derived weights; and both weight-sensitivity runs. It uses the Python standard library only. The reference implementation of the LPI and PCV scorers is separately available at https://github.com/Luis-Pena-Udit/ludicmetrics (MIT licence). HUMAN-SUBJECTS DATA. First-round participants gave verbal informed consent and second-round participants consented electronically; all were told that the ratings and comments would be published as an open dataset. Raters are identified only by unguessable pseudonymous tokens; no names, e-mail addresses or other personal data were collected or are released. One first-round free-text comment containing a real given name was redacted; the redaction is itemised in REDACTIONS.md, and the scores on that record are unchanged. NOT INCLUDED. The complete per-trial generation logs (approximately 13 MB of model-generated prose, Fact Ledgers and Session-1 outlines) are omitted for size. They contain no personal data, no reported result depends on them, and they are available from the corresponding author on request. THIRD-PARTY CONTENT. config/forgotten_realms_story_bible.json references the Forgotten Realms campaign setting, property of Wizards of the Coast. It contains only the short factual lore constraints used to condition generation and is included for reproducibility. It is not covered by the CC BY 4.0 grant and is not licensed for redistribution as setting material. See LICENSE.txt.



