遇见数据集

Beyond Reference Similarity: Why Current Metrics Fail to Capture Dialectical Reasoning in LLMs

收藏
Zenodo2025-04-13 更新2026-05-26 收录
官方服务:

资源简介:

This study examines the challenges of evaluating dialectical reasoning in Large Language Models (LLMs) through the development and assessment of AERIS (Adaptive Emergent Relational Intelligence System), a minimalist framework designed to enhance emergent reasoning capabilities. Dialectical reasoning, which involves exploring conceptual tensions and synthesizing opposing perspectives, is crucial for applications requiring nuanced understanding, yet current evaluation paradigms struggle to assess it. Evaluations on established benchmarks, including MMLU, TruthfulQA, and BIG-Bench-Hard (BBH), using standard metrics (BLEU, METEOR, ROUGE-L, BLEURT, COMET, BERTScore), consistently undervalued AERIS’s nuanced, explanatory responses, favoring instead short, reference-aligned answers. This inadequacy reflects a broader limitation in current evaluation paradigms, which prioritize lexical and semantic similarity over the richness of dialectical exploration. Building on insights from the author’s earlier work (Dulin, 2025), which anticipated this mismatch and explored alternative evaluation approaches, the findings confirm that standard metrics and benchmarks like MMLU and BBH appear ill-suited for assessing emergent cognition. The study highlights the need for new evaluation paradigms that capture the quality of reasoning processes, particularly as LLMs evolve toward advanced reasoning and metacognitive capabilities, offering insights into the future of LLM assessment.

提供机构:
Zenodo
创建时间:
2025-04-13
二维码
社区交流群
二维码
科研交流群
商业服务