NoisyAG-News: A Benchmark for Addressing Instance-Dependent Noise in Text Classification
收藏资源简介:
Real-world text classification is challenged by instance-dependent noise (IDN), yet Learning with Noisy Labels (LNL) research predominantly relies on simplistic synthetic models, creating a critical disconnect between academic evaluation and practice. To bridge this gap, we introduce NoisyAG-News, a large-scale benchmark with controllable, real-world IDN. Constructed via a meticulous multi-annotator process on 50,000 samples, our benchmark embodies the genuine cognitive biases inherent in human annotation. Our analysis of this benchmark yields two fundamental insights. First, we reveal why realistic IDN is substantially more destructive than synthetic noise: unlike the initial resistance models show to synthetic errors, it causes immediate learning contamination that leads to a catastrophic, generalizable "Short-Plank Effect." Second, we identify the distinct causal mechanisms of realistic noise sources: a top-down, semantic "Fallback" for human errors versus a bottom-up, feature-driven "Collapse" for LLM biases. NoisyAG-News provides a crucial testbed and a deeper understanding of noise to accelerate the development of LNL algorithms that are genuinely robust to real-world challenges.
现实世界文本分类面临着实例相关噪声(instance-dependent noise,IDN)的挑战,然而带噪标签学习(Learning with Noisy Labels,LNL)领域的研究大多依赖于简单化的合成噪声模型,这使得学术评估与实际应用之间出现了严重脱节。为填补这一研究空白,我们推出了NoisyAG-News——一个搭载可控真实世界实例相关噪声的大规模基准数据集。该数据集基于5万个样本,通过严谨的多标注者流程构建,充分体现了人工标注中固有的真实认知偏差。针对该基准数据集的分析得出两项核心发现:其一,我们揭示了真实世界实例相关噪声为何显著更具破坏性——与模型对合成错误最初表现出的抵抗性不同,真实噪声会引发即时的学习污染,进而导致灾难性且可泛化的「短板效应(Short-Plank Effect)」;其二,我们明确了真实噪声源的独特因果机制:人工标注错误属于自上而下的语义「回退(Fallback)」机制,而大语言模型(Large Language Model,LLM)偏差则属于特征驱动的自下而上「坍塌(Collapse)」机制。NoisyAG-News为研究人员提供了至关重要的测试平台,也让我们对噪声有了更深入的理解,从而推动能够真正抵御现实世界挑战的带噪标签学习算法研发。



