遇见数据集

Cross-Model Repertory Grid: LLM-as-LLM Evaluative Ratings

收藏
Zenodo2026-06-16 更新2026-06-17 收录
官方服务:

资源简介:

Cross-Model Repertory Grid (CM-RG) measures structural divergence in evaluative judgment between large language models on advisory tasks that have no verifiablecorrect answer. Adapting Kelly's Personal Construct Psychology, each model writes a free-text advisory response, elicits its own bipolar constructs by triadic comparison, and cross-rates anonymized peers on the union of emergent constructs. This is the Phase 2L run: 36 frontier models (12 provider families x 3 deployment tiers) on 7 advisory tasks under 2 prompting conditions. The release contains five normalized parquet tables - ratings (3,055,153 rows), constructs (86,418), responses (504), cells (395), api_calls (1,462) - with 13,928 distinct rater-ratee pairs and a mean inter-rater correlation of 0.200. Data of record: results_phase2l/analysis_results.json (2026-06-13). Run cost USD 112.89 over 2,572 API calls.

创建时间:
2026-06-16
二维码
社区交流群
二维码
科研交流群
商业服务