SeaExam and SeaBench
收藏资源简介:
SeaExam和SeaBench是两个新颖的基准数据集,专为评估大型语言模型在东南亚应用程序场景中的能力而设计。SeaExam基于东南亚地区现实世界的教育考试场景构建,包含地方历史和文学等科目。而SeaBench则围绕多轮、开放式任务,反映东南亚社区内的日常互动。这两个数据集均由本地语言专家精心构建,以适应东南亚地区的独特应用场景和文化背景。
SeaExam and SeaBench are two novel benchmark datasets specifically designed to evaluate the capabilities of large language models (LLMs) in Southeast Asian application scenarios. SeaExam is constructed based on real-world educational examination scenarios in Southeast Asia, covering subjects such as local history and literature. SeaBench focuses on multi-turn, open-ended tasks that reflect daily interactions within Southeast Asian communities. Both datasets are meticulously developed by local language experts to cater to the unique application scenarios and cultural backgrounds of the Southeast Asian region.

- 1SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia新加坡南洋理工大学,新加坡;阿里巴巴集团DAMO学院,新加坡;杭州湖畔实验室,中国;新加坡管理大学 · 2025年



