Uhura
收藏资源简介:
Uhura数据集是由Masakhane研究团队创建的,旨在评估低资源非洲语言中的科学问答和事实准确性。该数据集包括六个非洲语言的翻译版本,涵盖了科学知识和事实性问题,旨在解决大型语言模型在低资源语言中的性能问题。数据集通过专业翻译人员的协作创建,确保了翻译质量和文化相关性。Uhura数据集的应用领域主要集中在自然语言处理和人工智能安全,旨在提升多语言环境下的模型性能和可靠性。
The Uhura Dataset was created by the Masakhane research team to evaluate scientific question answering and factual accuracy in low-resource African languages. This dataset includes translated versions across six African languages, covering scientific knowledge and factual questions, aiming to address the performance limitations of large language models in low-resource languages. It was developed through collaboration among professional translators to ensure translation quality and cultural relevance. The main application areas of the Uhura Dataset focus on natural language processing and AI safety, with the goal of improving model performance and reliability in multilingual environments.




