遇见数据集

NERetrieve

收藏
arXiv2023-10-22 更新2024-06-21 收录
官方服务:

资源简介:

NERetrieve数据集由巴伊兰大学创建,包含约400万段英文维基百科段落,标记了500种实体类型的实体范围。数据集设计用于支持从细粒度监督识别到零样本全面检索的一系列任务,特别强调跨领域鲁棒性、细粒度、特定和交叉实体类型处理,以及从识别到检索的零样本设置扩展。该数据集旨在推动实体识别技术的发展,解决现有模型在处理复杂和特定实体类型时的局限性。

The NERetrieve dataset was developed by Bar-Ilan University. It comprises approximately 4 million English Wikipedia paragraphs, with entity spans annotated for 500 distinct entity types. The dataset is designed to support a spectrum of tasks ranging from fine-grained supervised entity recognition to zero-shot comprehensive retrieval, placing particular emphasis on cross-domain robustness, fine-grained processing, specific and cross-entity type handling, as well as the extension of zero-shot settings from recognition to retrieval. This dataset aims to advance the development of entity recognition technologies and address the limitations of existing models when processing complex and specialized entity types.

提供机构:
巴伊兰大学
创建时间:
2023-10-22
搜集汇总
数据集介绍
NERetrieve 数据集图片
背景与挑战
背景概述
NERetrieve是一个用于下一代命名实体识别和检索的数据集,基于EMNLP 2023论文,提供三种格式以支持全面实体提及检索、细粒度监督NER和零样本细粒度NER任务。该数据集遵循CC BY-SA 4.0许可证,包含代码和详细文档,适用于多种自然语言处理研究场景。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务