DocRED
收藏资源简介:
DocRED是由清华大学计算机科学与技术系创建的大规模文档级关系抽取数据集,包含5,053篇维基百科文档,标注了132,375个实体和56,354个关系事实。该数据集不仅要求从多句话中提取实体和推断关系,还提供了远监督数据以支持弱监督学习场景。DocRED的应用领域广泛,旨在推动文档级关系抽取的研究,解决现有方法在处理跨句关系时的局限性。
DocRED is a large-scale document-level relation extraction dataset developed by the Department of Computer Science and Technology, Tsinghua University. It includes 5,053 Wikipedia documents, annotated with 132,375 entities and 56,354 relational facts. This dataset not only requires extracting entities and inferring relations from multiple sentences, but also provides distant supervision data to support weakly supervised learning scenarios. DocRED covers a wide range of application fields, aiming to promote research on document-level relation extraction and address the limitations of existing methods in handling cross-sentence relations.

- 1DocRED: A Large-Scale Document-Level Relation Extraction Dataset清华大学计算机科学与技术系 · 2019年



