tabilab/biosses
收藏资源简介:
BIOSSES是一个用于生物医学句子相似性估计的基准数据集。该数据集包含100个句子对,每个句子对选自TAC(文本分析会议)生物医学摘要跟踪训练数据集中的生物医学领域文章。句子对由五位不同的人类专家评估其相似性,并给出0(无关系)到4(等价)的评分。在原始论文中,五位人类注释者评分的平均值被作为黄金标准。使用Pearson相关系数作为评估指标,评估模型估计的评分与黄金标准评分之间的相关性。
BIOSSES is a benchmark dataset for biomedical sentence similarity estimation. This dataset contains 100 sentence pairs, each selected from biomedical articles in the training dataset of the TAC (Text Analysis Conference) biomedical summarization track. Each sentence pair was evaluated for similarity by five distinct human experts, with scores ranging from 0 (no relation) to 4 (equivalent). In the original paper, the average of the five human annotators' scores was used as the gold standard. The Pearson correlation coefficient is employed as the evaluation metric to assess the correlation between the scores estimated by the model and the gold standard scores.
数据集概述
数据集名称
- 名称: BIOSSES
- 别名: BIOSSES
数据集基本信息
- 语言: 英语
- 许可证: GPL-3.0
- 多语言性: 单语
- 大小类别: 小于1K
- 源数据集: 原始
- 任务类别: 文本分类
- 任务ID: 文本评分, 语义相似度评分
数据集结构
- 特征:
- sentence1: 字符串
- sentence2: 字符串
- score: 浮点数(32位)
- 数据分割:
- 训练集: 100个样本, 32775字节
数据集创建
- 来源数据: TAC (Text Analysis Conference) Biomedical Summarization Track Training Dataset
- 注释: 由五位不同的人类专家评估句子对相似性并给出评分,评分范围从0(无关)到4(等同)。
使用数据集的注意事项
-
许可证: 根据GNU通用公共许可证v.3.0提供
-
引用信息:
@article{souganciouglu2017biosses, title={BIOSSES: a semantic sentence similarity estimation system for the biomedical domain}, author={So{u{g}}anc{i}o{u{g}}lu, Gizem and {"O}zt{"u}rk, Hakime and {"O}zg{"u}r, Arzucan}, journal={Bioinformatics}, volume={33}, number={14}, pages={i49--i58}, year={2017}, publisher={Oxford University Press} }




