Vrei să fii Milionar?
收藏资源简介:
该数据集是从罗马尼亚版《谁想成为百万富翁?》电视节目视频记录中提取的多语言数据集。数据集包含1000个多项选择题,涵盖了艺术、文化、电影、美食等多个领域,并标注了文化相关性和难度等级。数据集通过光学字符识别、自动文本提取和人工验证的过程收集而来,旨在解决低资源和多元文化背景下大型语言模型(LLM)性能评估的问题。数据集公开可在Hugging Face上获取。
This multilingual dataset is extracted from video recordings of the Romanian edition of the iconic television game show *Who Wants to Be a Millionaire?*. It includes 1,000 multiple-choice questions covering diverse domains including art, culture, cinema, cuisine and others, with annotations for cultural relevance and difficulty level. The dataset was collected through optical character recognition (OCR), automated text extraction and manual verification processes, with the core goal of addressing performance evaluation challenges for large language models (LLMs) in low-resource and multicultural contexts. The dataset is publicly available on Hugging Face.
数据集概述
基本信息
- 数据集名称: WWTBM
- 托管平台: Hugging Face
- 数据集地址: https://huggingface.co/datasets/WWTBM/wwtbm
数据集配置
数据集包含以下四种配置:
-
Romanian
- 数据文件: Romanian.json
- 描述: 完整的原始罗马尼亚语问答数据。
-
English
- 数据文件: English.json
- 描述: 英语问答题目。
-
French
- 数据文件: French.json
- 描述: 法语问答题目。
-
Romanian_translated
- 数据文件: Romanian_translated.json
- 描述: 翻译后的罗马尼亚语数据。




