GPR-bench
收藏资源简介:
GPR-bench是一个轻量级、可扩展的基准测试,旨在为通用用例的生成式AI系统提供回归测试。它包含一个开放的双语(英语和日语)数据集,涵盖了八个任务类别(例如文本生成、代码生成和信息检索)和每个任务类别中的10个场景(每种语言共80个测试案例)。该数据集由Galirage Inc.创建,旨在通过系统性的回归测试来确保生成式AI系统的可重复性和可靠性。数据集包括80个双语场景,涵盖了八个不同的任务类别,例如文本生成、代码生成和信息检索等,每个类别有10个场景。数据集的内容来源于OpenAI的ChatGPT模型,并使用了OpenEvals框架进行评估。GPR-bench的创建过程包括数据集构建、参考答案生成、模型和提示变体、评估流程以及统计分析方法。该数据集的应用领域主要是生成式AI系统的回归测试,旨在解决生成式AI系统在模型更新或提示修订时可能出现的行为漂移问题,确保系统的可重复性和可靠性。
GPR-bench is a lightweight, scalable benchmark intended for regression testing of generative AI systems across general use cases. It features an open bilingual (English and Japanese) dataset covering eight task categories including text generation, code generation, information retrieval and more, with 10 scenarios per category, totaling 80 test cases per language. Developed by Galirage Inc., this benchmark aims to ensure the reproducibility and reliability of generative AI systems through systematic regression testing. The dataset content is sourced from OpenAI's ChatGPT model, and evaluations are conducted using the OpenEvals framework. The development pipeline of GPR-bench encompasses dataset construction, reference answer generation, model and prompt variants, evaluation workflows and statistical analysis methodologies. Its core application is regression testing for generative AI systems, designed to mitigate behavioral drift issues that may arise during model updates or prompt revisions, so as to guarantee the reproducibility and reliability of the AI systems.

- 1Ensuring Reproducibility in Generative AI Systems for General Use Cases: A Framework for Regression Testing and Open DatasetsGalirage Inc. · 2025年



