遇见数据集

GEM (Generation, Evaluation, and Metrics)

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

GEM 是自然语言生成的基准环境,其重点是通过人工注释和自动化指标进行评估。 创业板旨在: 衡量跨语言的许多 NLG 任务的 NLG 进度。 审计数据和模型,并通过数据卡和模型稳健性报告呈现结果。 开发使用自动和人工指标评估生成文本的标准。 我们将定期更新 GEM,并通过扩展现有数据或开发其他语言的数据集来鼓励更具包容性的评估实践。

GEM is a benchmark environment for natural language generation (NLG), with a focus on evaluation through human annotations and automated metrics. GEM aims to: - Measure the progress of NLG across a wide range of cross-lingual NLG tasks. - Audit datasets and models, and present findings via data cards and model robustness reports. - Develop standards for evaluating generated text using both automatic and human evaluation metrics. We will regularly update GEM and encourage more inclusive evaluation practices by expanding existing datasets or developing datasets in other languages.

提供机构:
OpenDataLab
创建时间:
2022-08-11
搜集汇总
数据集介绍
GEM (Generation, Evaluation, and Metrics) 数据集图片
背景与挑战
背景概述
GEM是一个自然语言生成的基准环境,专注于通过人工注释和自动化指标来评估多语言NLG任务的进展。它旨在审计数据和模型,并开发评估生成文本的标准,由GEMv2 Team于2021年发布。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务