GEM (Generation, Evaluation, and Metrics)
收藏资源简介:
GEM 是自然语言生成的基准环境,其重点是通过人工注释和自动化指标进行评估。 创业板旨在: 衡量跨语言的许多 NLG 任务的 NLG 进度。 审计数据和模型,并通过数据卡和模型稳健性报告呈现结果。 开发使用自动和人工指标评估生成文本的标准。 我们将定期更新 GEM,并通过扩展现有数据或开发其他语言的数据集来鼓励更具包容性的评估实践。
GEM is a benchmark environment for natural language generation (NLG), with a focus on evaluation through human annotations and automated metrics. GEM aims to: - Measure the progress of NLG across a wide range of cross-lingual NLG tasks. - Audit datasets and models, and present findings via data cards and model robustness reports. - Develop standards for evaluating generated text using both automatic and human evaluation metrics. We will regularly update GEM and encourage more inclusive evaluation practices by expanding existing datasets or developing datasets in other languages.




