TRIVIAROOMQA
收藏资源简介:
TRIVIAROOMQA是由法国国家信息与自动化研究所创建的跨语言知识评测基准,旨在评估大语言模型在日常文化、长尾知识及多语言理解方面的能力。该数据集共包含8640道法语选择题(其中3300道为六种欧洲语言的平行问题),覆盖288个细分主题和24个广泛类别,数据来源于人类设计的开放式知识问答平台。数据集通过人工撰写和平台收集构建,每个问题均标注难度级别、时空元数据等丰富注释。该基准主要用于揭示模型在学术知识外的日常文化认知缺陷,助力提升模型在娱乐、新闻、流行文化等现实场景中的知识覆盖与多语言一致性。
TRIVIAROOMQA is a cross-lingual knowledge evaluation benchmark developed by the Institut National de Recherche en Informatique et en Automatique (INRIA) of France, designed to assess the capabilities of large language models (LLMs) in daily cultural cognition, long-tail knowledge and multilingual understanding. This dataset comprises 8640 French multiple-choice questions in total, with 3300 of them being parallel questions across six European languages, covering 288 subdivided topics and 24 broad categories. The dataset is built upon data sourced from human-curated open knowledge Q&A platforms, and constructed via both manual question writing and platform data collection. Each question is annotated with rich annotations including difficulty level, spatiotemporal metadata and other relevant details. This benchmark is primarily used to reveal the daily cultural cognitive deficits of LLMs beyond academic knowledge, and help enhance the knowledge coverage and multilingual consistency of models in real-world scenarios such as entertainment, news and popular culture.

- 1When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs法国国家信息与自动化研究所·巴黎 · 2026年



