OJI (Romanian County-Level Informatics Olympiad) 数据集
收藏资源简介:
OJI数据集是由罗马尼亚县级信息学奥林匹克竞赛提供的技术问题集合,包含300条罗马尼亚语的计算问题。该数据集由布加勒斯特大学等研究机构创建,旨在通过增强的英语翻译来支持大语言模型的训练和评估。数据集的内容涵盖了8年级学生的低至中等难度问题,涉及字符串处理等复杂文本。数据集的创建过程包括从原始罗马尼亚语问题中选择44条进行翻译,并通过多次运行GPT-4o模型来评估翻译质量。该数据集的应用领域主要集中在自动翻译、教育材料生成以及多语言技术问题的解决,旨在减少翻译错误并提高大语言模型在非英语语言任务中的表现。
The OJI Dataset is a collection of technical problems provided by the Romanian County-level Informatics Olympiad, containing 300 Romanian-language computational problems. Developed by research institutions including the University of Bucharest, this dataset aims to support the training and evaluation of Large Language Models (LLMs) through enhanced English translations. The dataset covers low-to-medium difficulty problems designed for 8th-grade students, involving complex text processing tasks such as string manipulation. The dataset creation workflow included selecting 44 original Romanian problems for translation, and evaluating the translation quality by running the GPT-4o model multiple times. The primary application scenarios of this dataset cover automatic translation, educational material generation, and multilingual technical problem solving, with the objective of reducing translation errors and improving the performance of LLMs on non-English language tasks.

- 1Exploring Large Language Models for Translating Romanian Computational Problems into English布加勒斯特大学, It Just Works Inc., QPillars · 2025年



