Dataset for Quantifying Affective Misgrounding in Large Language Models (EGS–EHI Benchmark)
收藏资源简介:
This dataset accompanies the study on affective misgrounding in large language models (LLMs), providing a structured benchmark for evaluating discrepancies between input emotional content and model-generated responses. The dataset consists of controlled prompt–response pairs collected across three affective conditions: neutral, mild, and high emotional intensity. Each prompt is designed to isolate emotional content while minimizing semantic complexity. For every input prompt, responses are generated using two model classes: an instruction-tuned model (Gemini) and a base generative model (GPT-2). For each input–response pair, the dataset includes: Input prompt Model-generated response Emotional intensity of input (EinputE_{input}Einput) Emotional intensity of response (EresponseE_{response}Eresponse) Emotional Grounding Score (EGS) Emotional Hallucination Index (EHI) Emotional intensity values are computed using the VADER sentiment analysis framework, with polarity converted to magnitude to capture affective strength independent of valence. EGS measures the alignment between input and response emotional intensity, while EHI quantifies normalized excess emotional expression. The dataset is intended to support reproducible evaluation of affective behavior in LLMs and to enable further research on alignment-induced biases, emotional overexpression, and grounding in generative systems. All data is provided in CSV format to facilitate direct analysis and integration into experimental pipelines. This dataset provides a minimal but controlled benchmark for quantifying affective distortion in LLM outputs and is designed to support comparative and methodological studies in LLM evaluation.



