遇见数据集

ExBAN Corpus

收藏
Zenodo2024-10-16 更新2026-05-26 收录
官方服务:

资源简介:

ExBAN Corpus (Explanations for BAyesian Networks) Authors: Miruna-Adriana Clinciu, Arash Eshghi and Helen Hastie The ExBAN dataset: a corpus of NL explanations generated by crowd-sourced participants presented with the task of explaining simple Bayesian Network (BN) graphical representations. These explanations, in a separate collection effort, are rated for clarity and informativeness. Citing If you use this dataset in your work, please cite the following paper: Clinciu, Miruna-Adriana, Arash Eshghi, and Helen Hastie. "A Study of Automatic Metrics for the Evaluation of Natural Language Explanations." Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. Association for Computational Linguistics, April 2021, Online, pp. 2376-2387. Available at: https://www.aclweb.org/anthology/2021.eacl-main.202. Abstract: As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural language (NL) explanations. Here, we explore parallels between the generation of such explanations and the much-studied field of evaluation of Natural Language Generation (NLG). Specifically, we investigate which of the NLG evaluation measures map well to explanations. We present the ExBAN corpus: a crowd-sourced corpus of NL explanations for Bayesian Networks. We run correlations comparing human subjective ratings with NLG automatic measures. We find that embedding-based automatic NLG evaluation methods, such as BERTScore and BLEURT, have a higher correlation with human ratings, compared to word-overlap metrics, such as BLEU and ROUGE. This work has implications for Explainable AI and transparent robotic and autonomous systems. @inproceedings{clinciu-etal-2021-study, title = "A Study of Automatic Metrics for the Evaluation of Natural Language Explanations", author = "Clinciu, Miruna-Adriana and Eshghi, Arash and Hastie, Helen", booktitle = "Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume", month = apr, year = "2021", address = "Online", publisher = "Association for Computational Linguistics", url = "https://www.aclweb.org/anthology/2021.eacl-main.202", pages = "2376--2387", abstract = "As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural language (NL) explanations. Here, we explore parallels between the generation of such explanations and the much-studied field of evaluation of Natural Language Generation (NLG). Specifically, we investigate which of the NLG evaluation measures map well to explanations. We present the ExBAN corpus: a crowd-sourced corpus of NL explanations for Bayesian Networks. We run correlations comparing human subjective ratings with NLG automatic measures. We find that embedding-based automatic NLG evaluation methods, such as BERTScore and BLEURT, have a higher correlation with human ratings, compared to word-overlap metrics, such as BLEU and ROUGE. This work has implications for Explainable AI and transparent robotic and autonomous systems.", }

ExBAN语料库(Explanations for BAyesian Networks,贝叶斯网络解释语料库) 作者:米鲁娜-阿德里亚娜·克林丘(Miruna-Adriana Clinciu)、阿拉什·埃什吉(Arash Eshghi)与海伦·黑斯廷(Helen Hastie) ExBAN数据集:一个由众包参与者生成的自然语言(Natural Language, NL)解释语料库,参与者需完成简单贝叶斯网络(Bayesian Network, BN)图形表示的解释任务。在独立的收集流程中,所有解释会被标注清晰度与信息丰富度两项评分。 ### 引用说明 若您在研究工作中使用本数据集,请引用以下论文: 米鲁娜-阿德里亚娜·克林丘、阿拉什·埃什吉与海伦·黑斯廷. "面向自然语言解释评估的自动指标研究". 见:第16届欧洲计算语言学协会会议论文集:主卷. 国际计算语言学协会欧洲分会, 2021年4月, 线上举办, 第2376-2387页. 可访问链接:https://www.aclweb.org/anthology/2021.eacl-main.202. ### 摘要 随着可解释性成为机器人学与人工智能(AI)的核心需求,评估可解释性的提供方法(包括自动生成的自然语言解释)已成为必要之举。本文探讨了此类解释生成与广受研究的自然语言生成(Natural Language Generation, NLG)评估领域之间的共通性。具体而言,本文研究了哪些自然语言生成评估指标能够很好地适配解释任务。本文提出了ExBAN语料库:一个面向贝叶斯网络的众包自然语言解释语料库。我们通过相关性分析对比了人类主观评分与自然语言生成自动评估指标的表现。研究发现,基于嵌入的自然语言生成自动评估方法(如BERTScore与BLEURT)相较于词重叠类指标(如BLEU与ROUGE),与人类评分的相关性更高。本研究对可解释人工智能(Explainable AI, XAI)以及透明化机器人与自主系统研究具有重要参考价值。 ### BibTeX引用格式 @inproceedings{clinciu-etal-2021-study, title = "A Study of Automatic Metrics for the Evaluation of Natural Language Explanations", author = "Clinciu, Miruna-Adriana and Eshghi, Arash and Hastie, Helen", booktitle = "Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume", month = apr, year = "2021", address = "Online", publisher = "Association for Computational Linguistics", url = "https://www.aclweb.org/anthology/2021.eacl-main.202", pages = "2376--2387", abstract = "As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural language (NL) explanations. Here, we explore parallels between the generation of such explanations and the much-studied field of evaluation of Natural Language Generation (NLG). Specifically, we investigate which of the NLG evaluation measures map well to explanations. We present the ExBAN corpus: a crowd-sourced corpus of NL explanations for Bayesian Networks. We run correlations comparing human subjective ratings with NLG automatic measures. We find that embedding-based automatic NLG evaluation methods, such as BERTScore and BLEURT, have a higher correlation with human ratings, compared to word-overlap metrics, such as BLEU and ROUGE. This work has implications for Explainable AI and transparent robotic and autonomous systems.", }

提供机构:
Zenodo
创建时间:
2024-10-16
二维码
社区交流群
二维码
科研交流群
商业服务