MAAC Study 1 Hypothesis Testing Dataset: Complexity Validation of LLM-Generated Decision Scenarios
收藏资源简介:
This dataset contains the source data supporting all hypothesis tests and research questions reported in Study 1 of the Multi-Dimensional Assessment for AI Cognition (MAAC) dissertation series. Files cover construct validity (H1a–H1d), test-retest reliability (H2), multi-judge consensus scoring (H3), domain balance (H4), complexity spectrum distribution (H5), model throughput (H6), schema compliance (H7), speed-quality tradeoffs (RQ5), and failure pattern analysis (RQ6). Data were generated using a multi-LLM judge panel and validated using a composite four-framework complexity instrument (ICC = .997). This dataset supports replication of findings reported in: "Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models."
本数据集包含支撑人工智能认知多维评估(Multi-Dimensional Assessment for AI Cognition,MAAC)学位论文系列研究1中所有报告的假设检验与研究问题的源数据。 所涵盖文件涉及构念效度(H1a–H1d)、重测信度(H2)、多评判者共识评分(H3)、领域平衡性(H4)、复杂度谱分布(H5)、模型吞吐量(H6)、图式合规性(H7)、速度-质量权衡(RQ5)以及失效模式分析(RQ6)。 本数据集采用多大语言模型(Large Language Model,LLM)评估者面板生成数据,并通过复合四框架复杂度评估工具完成验证,其组内相关系数(Intraclass Correlation Coefficient,ICC)为0.997。 本数据集可复现下述研究的报告成果:《使用大语言模型生成复杂度验证型决策场景的自动化方法》。 数据集仓库:https://github.com/Doleha/maac



