遇见数据集

TraCE-LLM: Evaluation datasets and pipeline (v2.3)

收藏
Zenodo2026-02-09 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the core assets used in the TraCE-LLM study described in the article The Adversarial Compensation Effect: Identifying Hidden Instabilities in Large Language Model Evaluation. The project is centered on the TraCE-LLM protocol, which measures latent behavioral traits of Large Language Models (LLMs) using a multidimensional rubric with two primary dimensions: Depth of Reasoning (DoR) and Originality (ORI). It includes benchmark test splits (ARC-Challenge, MMLU, HellaSwag), an N=500 evaluation sample, human evaluation baselines, aggregated statistics, curated illustrative examples, and the complete analysis pipeline as Jupyter notebooks.

提供机构:
Zenodo
创建时间:
2026-02-09
二维码
社区交流群
二维码
科研交流群
商业服务