TraCE-LLM: Evaluation datasets and pipeline (v2.3)
收藏官方服务:
资源简介:
This repository contains the core assets used in the TraCE-LLM study described in the article The Adversarial Compensation Effect: Identifying Hidden Instabilities in Large Language Model Evaluation. The project is centered on the TraCE-LLM protocol, which measures latent behavioral traits of Large Language Models (LLMs) using a multidimensional rubric with two primary dimensions: Depth of Reasoning (DoR) and Originality (ORI). It includes benchmark test splits (ARC-Challenge, MMLU, HellaSwag), an N=500 evaluation sample, human evaluation baselines, aggregated statistics, curated illustrative examples, and the complete analysis pipeline as Jupyter notebooks.
提供机构:
Zenodo创建时间:
2026-02-09



