遇见数据集

bigbio/an_em

收藏
Hugging Face2022-12-22 更新2024-03-04 收录
官方服务:

资源简介:

AnEM语料库是一个领域和物种无关的资源,手动注释了使用细粒度分类系统的解剖实体提及。该语料库包含500个文档(超过90,000字),这些文档是从引用摘要和全文论文中随机选择的,旨在代表整个可用的生物医学科学文献。语料库的注释涵盖了健康和病理解剖实体的提及,包含超过3,000个注释提及。

The AnEM corpus is a domain- and species-agnostic resource with manually annotated anatomical entity mentions using a fine-grained classification system. This corpus consists of 500 documents (over 90,000 words) randomly selected from cited abstracts and full-text articles, designed to represent the entire available biomedical scientific literature. The annotations in the corpus cover mentions of both healthy and pathological anatomical entities, with more than 3,000 annotated mentions in total.

提供机构:
bigbio
原始信息汇总

数据集概述

基本信息

  • 数据集名称: AnEM
  • 语言: 英语
  • 许可证: CC-BY-SA-3.0
  • 多语言性: 单语种
  • PubMed可用性:
  • 公开可用性:

任务类型

  • 命名实体识别 (NER)
  • 共指消解 (COREF)
  • 关系抽取 (RE)

数据集详情

  • 描述: AnEM是一个领域和物种独立的手工标注资源,用于解剖实体提及,使用细粒度分类系统。该数据集包含500个文档(超过90,000字),随机选自引文摘要和全文论文,旨在代表整个可用的生物医学科学文献。
  • 标注内容: 包含健康和病理解剖实体的提及,共有超过3,000个标注提及。

引用信息

@inproceedings{ohta-etal-2012-open, author = {Ohta, Tomoko and Pyysalo, Sampo and Tsujii, Jun{}ichi and Ananiadou, Sophia}, title = {Open-domain Anatomical Entity Mention Detection}, journal = {}, volume = {W12-43}, year = {2012}, url = {https://aclanthology.org/W12-4304}, doi = {}, biburl = {}, bibsource = {}, publisher = {Association for Computational Linguistics} }

搜集汇总
数据集介绍
bigbio/an_em 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务