遇见数据集

Incremental Instruction Creative Writing: Benchmark, Generations, and Evaluation Dataset

收藏
Zenodo2026-08-15 更新2026-08-20 收录
官方服务:

资源简介:

Dataset accompanying “The Effects of Incremental Instruction Delivery on Language-Model Creative Writing." This release contains the complete dataset and evaluation artifacts used in the study. It includes 160 creative-writing benchmark tasks across six genres; model generations from six open-weight language models under FULL and SHARDED instruction-delivery conditions; 1,920 pointwise LLM evaluations of atomic constraint adherence and five writing-quality dimensions; 30-pair simple and evidence-first LLM comparison audits in primary and reversed response order; and the corresponding 30 human judgments. The dataset is organized into separate configurations for benchmark tasks, generations, model evaluations, pairwise judge comparisons, evidence-first judge audits, human-evaluation cases, and human annotations. Additional files provide extracted atomic constraints, evaluator configuration, run provenance, integrity hashes, and the statistics underlying the evaluator-validation analysis. The FULL condition presents the complete task specification in one turn. The SHARDED condition distributes the task specification progressively across multiple conversational turns. Shard sets were designed to preserve the task content of their fully specified counterparts; the accompanying paper documents a minor post-hoc sharding omission and other relevant limitations. Code and analysis scripts are released separately in the SISTER repository and archived on Zenodo under DOI 10.5281/zenodo.21951541. This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

提供机构:
Zenodo
创建时间:
2026-08-15
二维码
社区交流群
二维码
科研交流群
商业服务