遇见数据集

NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1

收藏
Hugging Face2025-05-22 更新2025-11-01 收录
官方服务:

资源简介:

DUMY是一个面向乌克兰推理任务的开放基准和数据集,专为训练、蒸馏和评估聚焦于乌克兰推理任务的语言模型而设计。该数据集包含来自ZNO-Eval基准的约3303个考试问题,涵盖乌克兰语言和文学、数学、乌克兰历史和地理等多个领域。每个任务都配备了DeepSeek R1和OpenAI o1模型生成的链式推理样本,适合进行LLM蒸馏、GRPO或标准微调、链式推理研究以及评估多语种和低资源推理能力。

DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks. This dataset includes ~3303 exam questions from the ZNO-Eval benchmark across multiple domains such as Ukrainian language and literature, Mathematics, History of Ukraine, and Geography. Each task is enhanced with chain-of-thought samples generated by DeepSeek R1 and OpenAI o1, making DUMY suitable for LLM distillation, GRPO or standard fine-tuning, CoT reasoning research, and evaluation of multilingual and low-resource reasoning capabilities.

提供机构:
NLPForUA
二维码
社区交流群
二维码
科研交流群
商业服务