issai/freshqa_kazakh
收藏资源简介:
FreshQA哈萨克语数据集是原始FreshQA幻觉基准数据集的机器翻译版本,专门设计用于测试大型语言模型的事实性和幻觉率。该数据集特别关注那些答案可能随时间变化或涉及错误前提的问题。它包含600个样本,分为三种不同的事实类型。为了涵盖模型可能正确回应的各种方式,每个问题提供了最多10个可能的有效答案。问题和所有有效答案均已翻译成哈萨克语,从而支持对模型处理动态事实知识和复杂推理步骤的能力进行稳健的跨语言评估。
The Kazakh FreshQA dataset is a machine-translated version of the original FreshQA hallucination benchmark dataset, specifically designed to evaluate the factuality and hallucination rate of large language models (LLMs). This dataset specifically focuses on questions whose answers may change over time or involve false premises. It consists of 600 samples, categorized into three distinct factuality types. To cover the diverse ways a model can deliver a correct response, each question is paired with up to 10 valid possible answers. Both the questions and all valid answers have been translated into Kazakh, thus enabling robust cross-lingual evaluation of a model's ability to process dynamic factual knowledge and perform complex reasoning steps.




