遇见数据集

LDkp Dataset

收藏
Zenodo2021-09-12 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

LDkp (Long Document keyphrase) dataset is the first benchmark corpus of 1.3M documents for identifying keyphrases from long documents.The LDkp dataset is released in two versions : <strong>LDkp3k</strong> consists of 0.1M keyphrase tagged long documents, is created using keyphrases from KP20k (Meng et al., 2017) and their corresponding long document text from S2ORC (Lo et al., 2020). <strong>LDkp10k</strong> consists of 1.3M long documents along with target keyphrases is created using keyphrases from OAGKX (Çano, 2019) and their corresponding long document text from S2ORC (Lo et al., 2020).

创建时间:
2021-09-12
二维码
社区交流群
二维码
科研交流群
商业服务