somosnlp-hackathon-2026/che-boludo-benchmark
收藏资源简介:
Che Boludo Bench 是一个针对西班牙语Rioplatense变体(主要来自阿根廷地区)的实用对齐基准数据集。当前语言模型在解释西班牙语文化变体时存在重要限制:它们倾向于过度字面化处理语言,并对特定语言社区的俚语表达进行过度调节。这一问题在Rioplatense西班牙语中尤为明显,因为日常交流高度依赖讽刺、反讽、隐含语境、俚语、情感代码和表演性攻击幽默。在阿根廷,包含侮辱或文本攻击性的表达在社会中常作为信任、亲近或幽默的展示,但许多现有模型错误地将其分类为真实毒性。数据集旨在评估模型如何理解Rioplatense西班牙语对话中的攻击性、讽刺和社交意图,具体测量模型何时:错误检测不存在的毒性、将侮辱误解为真实攻击而非情感代码、以及对使用俚语的日常查询做出不合理的说教或拒绝响应。数据集结构遵循经典的偏好优化(DPO/RLHF)格式,每条记录包含:prompt(来自本地社区如Reddit Argentina的Rioplatense西班牙语短语或对话上下文)、chosen(展示实用对齐的理想响应,识别幽默、不过度调节、以文化自然方式回应)、rejected(未对齐的响应,通常是企业过度调节类型,如无法处理攻击性语言或过度字面化解释)以及pragmatica(详细注释,包括言语行为、真实意图、极性及真实敌意/情感水平)。创新点在于提出文化感知的安全对齐视角,即关注本地文化和实用规范的对齐方法,探索主要基于英语数据集训练的调节系统在与高度情境化语言社区(如阿根廷)互动时可能产生的误报。预期影响包括:促进多文化对齐研究、情境化调节、减少实用偏见、保护西班牙语文化变体以及开发拉丁美洲本地化人工智能。此外,该基准可作为未来微调、安全对齐、社会语言评估和计算语用学研究的基础。结果预期是提供一个可复现的基准,能够测量Rioplatense西班牙语中的过度调节、评估上下文实用理解、展示当前模型的文化缺陷,并作为未来文化对齐AI系统的评估工具。
Che Boludo Bench is a pragmatic alignment benchmark dataset for the Rioplatense Spanish variant (primarily from Argentina). Current language models present a significant limitation in interpreting cultural varieties of Spanish: they tend to process language overly literally and over-moderate colloquial expressions specific to certain linguistic communities. This problem is particularly evident in Rioplatense Spanish, where much of daily communication depends on sarcasm, irony, implicit context, slang, affective codes, and performative aggressive humor. In Argentina, expressions containing insults or textual aggressiveness often function socially as displays of trust, closeness, or humor, but many current models erroneously classify them as actual toxicity. The dataset aims to evaluate how models interpret aggressiveness, sarcasm, and social intent in Rioplatense Spanish conversations, specifically measuring when a model: detects toxicity where none exists, misinterprets insults as real attacks rather than affective codes, and responds with unwarranted sermons or refusals to everyday queries using slang. The dataset structure follows the classic preference optimization (DPO/RLHF) format, with each record containing: prompt (a phrase or conversational context in Rioplatense Spanish, mainly extracted from local communities like Reddit Argentina), chosen (the ideal or target response demonstrating pragmatic alignment, recognizing humor, not over-moderating, responding with cultural naturalness), rejected (the misaligned response, typically corporate over-moderation such as I cannot process offensive language or an overly literal interpretation), and pragmatica (detailed annotations on speech act, real intent, polarity, and levels of real hostility/affectivity). The innovation lies in proposing a culturally-aware safety alignment perspective, i.e., an alignment approach sensitive to local cultural and pragmatic norms, exploring how moderation systems trained mainly on English datasets can generate false positives when interacting with highly contextual linguistic communities like Argentina. The expected impact includes contributing to research in multicultural alignment, contextualized moderation, reduction of pragmatic biases, preservation of Spanish cultural variants, and development of localized AI for Latin America. Additionally, the benchmark can serve as a basis for future work in fine-tuning, safety alignment, sociolinguistic evaluation, and computational pragmatics. The expected outcome is a reproducible benchmark capable of measuring over-moderation in Rioplatense Spanish, evaluating contextual pragmatic understanding, demonstrating cultural flaws in current models, and serving as an evaluation tool for future culturally aligned AI systems.




