nhagar/CC-MAIN-2014-23_urls
收藏官方服务:
资源简介:
这是一个包含网页抓取数据的数据集,具体特征包括网页的抓取信息(crawl)、URL的主机名(url_host_name)以及URL的出现次数(url_count)。数据集被划分为训练集,包含大约3200万样本,数据集总大小约为1.69GB。提供了默认配置,用于指定训练集数据文件的路径。
This dataset consists of web crawl data, with features including the crawl information of the web page (crawl), the host name of the URL (url_host_name), and the occurrence count of the URL (url_count). The dataset is split into a training set, which contains about 32 million samples, with a total size of approximately 1.69GB. A default configuration is provided to specify the path to the data files for the training set.
提供机构:
nhagar


