LDkp Dataset
收藏数据链接:
官方服务:
资源简介:
LDkp (Long Document keyphrase) dataset is the first benchmark corpus of 1.3M documents for identifying keyphrases from long documents.The LDkp dataset is released in two versions : <strong>LDkp3k</strong> consists of 0.1M keyphrase tagged long documents, is created using keyphrases from KP20k (Meng et al., 2017) and their corresponding long document text from S2ORC (Lo et al., 2020). <strong>LDkp10k</strong> consists of 1.3M long documents along with target keyphrases is created using keyphrases from OAGKX (Çano, 2019) and their corresponding long document text from S2ORC (Lo et al., 2020).
提供机构:
Dibya Gautam; Debanjan Mahata; Amardeep Kumar; MIDAS- Lab, IIIT-D; Navneet Agarwal; Anish Acharya创建时间:
2021-09-12



