遇见数据集

Moral Reasoning RAG Evaluation Study: Data, Reproducibility Materials, and Statistical Analysis

收藏
Zenodo2026-09-28 更新2026-10-01 收录
官方服务:

资源简介:

This archive contains the data, source records, analysis code, verification materials, and supplementary documentation for a controlled evaluation of a Moral Reasoning Retrieval-Augmented Generation (RAG) intervention. The study compared the same underlying language model under two conditions: an unaugmented Raw API condition and a Moral Reasoning RAG condition. The evaluation dataset contained 300 moral-reasoning questions drawn from established benchmarks, published benchmark frameworks, and researcher-created reliability scenarios. Each matched Raw API and RAG response pair was evaluated under blinded A/B presentation by three AI evaluators, producing 900 judge–question evaluation records and 1,800 individually scored responses. The archive includes: the complete 300-question provenance record and study condition key; original evaluator-result packages; batch-load and model-export records; blinded and decoded canonical analytic datasets; data-cleaning and reconstruction documentation; the Universal AI Moral Reasoning Grading Protocol; executable Python statistical-analysis code; machine-readable statistical outputs; verification records and reproduced figures; exploratory response-length and sensitivity analyses; a cluster-robust sensitivity analysis accounting for the 12 five-stage Multi-step Moral Dilemma sequences represented among the 60 MMD items; and a blind-review recalculation package designed to permit independent statistical verification of the reported results. The canonical analysis reconstructs Moral Reasoning Total Scores from the 12 component dimensions (MR1–MR12) and documents all transformations required to move from the original blinded evaluator records to the final decoded analytic dataset. The released code reproduces the principal paired analysis, robustness and sensitivity analyses, dimension-level analyses, evaluator-specific comparisons, agreement and reliability analyses, presentation-position analyses, response-length analyses, and the MMD cluster-robust sensitivity analysis. The archive is intended to support independent auditing, statistical reproduction, and methodological scrutiny of the reported study results. The verbatim Moral Reasoning Ontology and Critical Directive used within the intervention are proprietary and are not included in this archive. Their functional roles are described in the associated manuscript. The archive therefore supports reproduction of the reported evaluation and statistical analyses, but not exact source-level reconstruction of the proprietary historical intervention.

提供机构:
Zenodo
创建时间:
2026-09-28
二维码
社区交流群
二维码
科研交流群
商业服务