SuMe
收藏资源简介:
SuMe数据集是由石溪大学等机构创建的,专注于生物医学机制总结的首个数据集。该数据集包含22,000个实例,通过半自动化过程从大量生物医学文献摘要中提取。数据集的创建涉及使用生物医学信息提取系统提取关键实体和关系,并通过领域专家的少量标注来训练机制句子分类器,进而筛选出包含机制句子的摘要。SuMe数据集主要用于训练大型神经网络模型,以理解和总结生物医学文献中的机制,旨在解决生物医学领域信息过载的问题,帮助研究人员快速获取和组织相关生物医学关系。
SuMe Dataset is the first dataset dedicated to biomedical mechanism summarization, developed by Stony Brook University and other institutions. It contains 22,000 instances extracted from a large corpus of biomedical literature abstracts through a semi-automated process. The construction of the dataset involves utilizing biomedical information extraction systems to extract key entities and relationships, and training a mechanism sentence classifier via few-shot annotations from domain experts, which is subsequently used to screen abstracts containing mechanism-related sentences. The SuMe Dataset is primarily designed for training large-scale neural network models to comprehend and summarize biomedical literature mechanisms, with the goal of addressing the problem of information overload in the biomedical domain and helping researchers rapidly acquire and organize relevant biomedical relationships.



