OmniGIRL
收藏资源简介:
OmniGIRL是一个多语言、多模态和多领域的GitHub问题解决基准数据集,包含来自四种编程语言(Python、JavaScript、TypeScript和Java)和八个不同领域的959个任务实例。该数据集不仅包含了文本信息,还包括了图像等多模态信息,旨在评估大型语言模型在解决GitHub问题方面的能力。数据集的创建过程包括了语言和仓库的选择、拉取请求数据的收集、任务实例的构建、基于执行的验证以及不必要的图像过滤等五个阶段。OmniGIRL数据集的应用领域主要在于评估和提升大型语言模型在解决GitHub问题方面的能力,旨在解决当前大型语言模型在多语言、多模态和多领域问题解决方面的局限性。
OmniGIRL is a multilingual, multimodal, and multi-domain GitHub issue-solving benchmark dataset containing 959 task instances spanning four programming languages (Python, JavaScript, TypeScript, and Java) and eight distinct domains. This dataset not only includes textual information but also multimodal content such as images, aiming to evaluate the capabilities of Large Language Models (LLMs) in solving GitHub-related issues. The creation process of the OmniGIRL dataset consists of five stages: selection of languages and repositories, collection of pull request data, construction of task instances, execution-based validation, and filtering of unnecessary images. The main application areas of the OmniGIRL dataset focus on evaluating and enhancing the problem-solving abilities of large language models for GitHub issues, with the goal of addressing the current limitations of large language models in multilingual, multimodal, and multi-domain problem-solving.

- 1OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution中山大学 · 2025年



