AraSTEM
收藏资源简介:
AraSTEM是一个专门用于评估大型语言模型在阿拉伯语STEM科目中知识掌握情况的数据集,由贝鲁特美国大学的研究团队创建。该数据集包含11637个多项选择题,涵盖数学、科学、物理、生物、化学、计算机科学和医学等多个学科,难度从小学到大学水平不等。数据集的创建过程包括网页抓取、手动提取和LLM提取,确保了数据的多样性和广泛性。AraSTEM旨在解决阿拉伯语STEM领域缺乏高质量评估基准的问题,为多语言模型的性能评估提供了重要参考。
AraSTEM is a dataset specifically developed to assess the knowledge proficiency of large language models (LLMs) in Arabic STEM disciplines, created by a research team at the American University of Beirut. This dataset contains 11,637 multiple-choice questions covering multiple disciplines including mathematics, science, physics, biology, chemistry, computer science, and medicine, with difficulty levels ranging from primary school to university. The dataset was developed through web scraping, manual extraction, and LLM-based extraction, ensuring the diversity and broad coverage of the data. AraSTEM aims to address the shortage of high-quality evaluation benchmarks in the Arabic STEM field, providing an important reference for the performance evaluation of multilingual models.




