遇见数据集

Entropy Profiles of Metaphor Comprehension: A Continuous Gradient from Conventional to Novel

收藏
Zenodo2026-04-14 更新2026-05-26 收录
官方服务:

资源简介:

Predictive-processing accounts of language comprehension suggest that metaphor involves structured perturbations of predictive uncertainty. However, existing approaches typically treat metaphoricity as a binary property of expressions rather than as a continuous, process-level phenomenon. We investigate the information-theoretic structure of metaphor comprehension by extracting next-token entropy profiles from a transformer language model (GPT-2, 124M parameters) at 179,431 annotated token positions in the VUA-2020 corpus, including 23,054 metaphorical and 156,377 literal instances. Each token is characterized by thirteen features describing the geometry of its local entropy profile: peak height, temporal offset, directional asymmetry, contextual baseline, and related measures. Three principal findings emerge from this analysis. First, conventional metaphors, which constitute approximately 75% of corpus-annotated instances, exhibit systematically lower entropy than their surrounding literal context, functioning as predictable elements embedded in informationally dense sentences. This finding contradicts the standard assumption that metaphors are processing-costly disruptions. Second, novel metaphors show the expected disruption signature, with entropy at the metaphorical position exceeding the surrounding context. The transition between these regimes follows a continuous gradient indexed by peak entropy height, with a crossover from negative to positive entropy differential occurring at approximately 3.2 standard deviations above baseline. Third, classification experiments confirm this dual structure quantitatively: entropy profile features achieve only AUROC 0.62 when applied to all annotated metaphors indiscriminately, but reach AUROC 0.96 when restricted to novel metaphors above the crossover threshold. A feature ablation analysis reveals that the discriminative features differ qualitatively between these regimes: for conventional metaphors, only the absolute entropy at the head position carries weak signal; for novel metaphors, the spatial geometry of the entropy perturbation—its peak magnitude, directional spread, and contextual embedding—becomes strongly predictive. These findings reframe metaphoricity as a graded perturbation of predictive information flow rather than a categorical semantic deviation, with implications for computational approaches to figurative language, theories of metaphor processing, and the relationship between conventionalization and information structure.

提供机构:
Zenodo
创建时间:
2026-04-14
二维码
社区交流群
二维码
科研交流群
商业服务