MAAC Study 1 Hypothesis Testing Dataset: Complexity Validation of LLM-Generated Decision Scenarios
收藏资源简介:
This dataset contains the source data supporting all hypothesis tests and research questions reported in Study 1 of the Multi-Dimensional Assessment for AI Cognition (MAAC) dissertation series. Files cover construct validity (H1a–H1d), test-retest reliability (H2), multi-judge consensus scoring (H3), domain balance (H4), complexity spectrum distribution (H5), model throughput (H6), schema compliance (H7), speed-quality tradeoffs (RQ5), and failure pattern analysis (RQ6). Data were generated using a multi-LLM judge panel and validated using a composite four-framework complexity instrument (ICC = .997). This dataset supports replication of findings reported in: "Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models."



