遇见数据集

The BigGrams: the semi-supervised information extraction system from HTML: an improvement in the wrapper induction - dataset

收藏
Zenodo2026-04-13 更新2026-05-25 收录
官方服务:

资源简介:

Brief description The zip file contains two folders. The "websites" folder includes crawled web pages from real websites, like a agatameble.pl (an e-shop website), filmweb.pl (a website about films), and ptaki.info (a website about birds). The "reference-seeds" folder contains three subfolders, i.e. agatameble.pl, filmweb.pl, and ptaki.info. Each subfolder contains reference-seeds.csv file. The file contains data, i.e. reference instances - carefully labelled ground-truth of corresponding values in each web page of given websites mentioned above. Reference I would appreciate it if you cite the following paper when using the dataset: Marcin Mirończuk The BigGrams: the semi-supervised information extraction system from HTML: an improvement in the wrapper induction, Knowledge and Information Systems, Volume 54, Issue 3, p. 711–776, 2018, (pdf Open Access – http://rdcu.be/u88F lub DOI http://dx.doi.org/10.1007/s10115-017-1097-2)

数据集简要说明 该压缩包包含两个文件夹。其中"websites"文件夹存储从真实网站爬取的网页,示例站点包括agatameble.pl(电商网站)、filmweb.pl(影视资讯网站)以及ptaki.info(鸟类资讯网站)。"reference-seeds"文件夹下设三个子文件夹,分别对应agatameble.pl、filmweb.pl与ptaki.info。每个子文件夹中均包含reference-seeds逗号分隔值(CSV)文件,该文件存储参考样本数据,即上述各网站对应网页中经过精细标注的真实值基准(ground-truth)。 参考文献 若您在使用该数据集时引用以下论文,我们将不胜感激: Marcin Mirończuk. 《BigGrams:面向HTML的半监督信息抽取系统:包装器归纳算法的改进》[J]. 知识与信息系统,2018,54(3):711-776. 该论文为开放获取论文,获取链接:http://rdcu.be/u88F,DOI:http://dx.doi.org/10.1007/s10115-017-1097-2

提供机构:
Zenodo
创建时间:
2018-04-04
二维码
社区交流群
二维码
科研交流群
商业服务