Spacy pipeline for English Named Entity Linking to Wikipedia/Wikidata
收藏资源简介:
A modified version of the standard spaCy model en_core_web_lg (described as an "English multi-task CNN trained on OntoNotes, with GloVe vectors trained on Common Crawl. Assigns word vectors, POS tags, dependency parses and named entities.") with Entity Linking trained on the first 300,000 lines of a gold_entities.jsonl that was itself created from a complete dump of Wikidata and en-Wikipedia on October 11 2020.
本模型为标准spaCy模型en_core_web_lg的修改版本。该基础模型的官方描述为:「基于OntoNotes语料库训练的英语多任务卷积神经网络(Convolutional Neural Network, CNN),搭载经Common Crawl数据集训练得到的GloVe词向量,可输出词向量、词性标注(Part-of-Speech, POS)标签、依存句法分析结果与命名实体」。本次修改版本的实体链接(Entity Linking)模块,基于gold_entities.jsonl文件的前30万行数据训练所得;该gold_entities.jsonl文件本身,源自2020年10月11日发布的维基数据(Wikidata)与英文维基百科(en-Wikipedia)完整导出快照。



