官方服务:
资源简介:
Arabic-English CS judgment tasks (Exp 1 & 2) testing the MLF model
应用场景:
创建时间:
2023-05-02
相关数据集
ucinlp/unstereo-eval
--- license: mit task_categories: - text-classification language: - en tags: - bias-evaluation - language modeling - gender-bias pretty_name: USE size_categories: - 1K<n<10K configs: - config_name: US
Hugging Face2024-05-02 更新480
AlgGeoBench
# Welcome to AlgGeoBench created by PKU-DS-LAB! ## Dataset Description AlgGeoBench is a frontier math benchmark consisting of 843 8-choice problems adapted from definitions and propositions of the o
魔搭社区2026-03-13 更新220
RepLiQA
RepLiQA是一个评估数据集,包含上下文-问题-答案三元组,上下文涉及虚构的实体,如人物或地点,不真实存在。该数据集旨在测试大型语言模型(LLMs)在提供文档中查找和使用上下文信息的能力。与现有问答数据集不同,RepLiQA的非事实性确保模型性能不受LLMs记忆训练数据中事实的能力影响,可以更自信地测试模型利用提供上下文的能力。数据集涵盖17个主题或文档类别,每个文档附带5个问题-答案对。此外,
github2024-06-19 更新321
cabusar/gutenberg-txt-fr
--- license: unknown task_categories: - text-classification - question-answering - text2text-generation language: - fr tags: - rag --- This dataset is not yet refined enough to do other things than t
Hugging Face2024-03-30 更新190



