devinendorphin/Gradient-Decomposition-Assay
收藏资源简介:
该数据集名为梯度分解分析(Gradient Decomposition Assay,GDA),是一个探索性行为评估数据集,用于研究前沿语言模型对八维提示流形的响应,该流形范围从良性技术任务到对抗性压缩和反事实/叙事重构。数据集包含CSV语料库和摘要表,通过多个配置提供结构化遥测数据、原始输出、向量摘要和无效运行审计。评估指标包括LLM评判员分配的分数,如内容Phi、形式Phi、特异性Phi、安全阻力、自我审计、拒绝强度和模板强度,这些应视为探索性行为测量,而非潜在模型状态或内部推理能力的真实测量。数据集设计矩阵最初包含2000次运行,有效摘要排除了无效或格式错误的行,并单独保留无效运行表以供审计。
This repository contains the CSV corpus and summary tables for the Gradient Decomposition Assay, an exploratory behavioral evaluation of how frontier language models respond to an eight-vector prompt manifold ranging from benign technical tasks to adversarial compression and counterfactual/narrative reframing. The dataset includes multiple configurations for structured telemetry, raw outputs, vector summaries, and invalid-run audits, with evaluator-assigned scores for metrics such as Phi_Content, Phi_Form, Phi_Specificity, Safety_Drag, Self_Audit, Refusal_Intensity, and Boilerplate_Intensity, which are treated as exploratory behavioral measurements rather than ground-truth assessments. The original design matrix had 2,000 runs, with valid summaries excluding invalid or malformed rows and a separate audit table for transparency.



