遇见数据集

nhagar/CC-MAIN-2019-39_urls

收藏
Hugging Face2025-05-15 更新2025-02-15 收录
官方服务:

资源简介:

--- dataset_info: features: - name: crawl dtype: string - name: url_host_name dtype: string - name: url_count dtype: int64 splits: - name: train num_bytes: 2844632861 num_examples: 56764197 download_size: 1015654379 dataset_size: 2844632861 configs: - config_name: default data_files: - split: train path: data/train-* --- This dataset contains domain names and counts of (non-deduplicated) URLs for every record in the CC-MAIN-2019-39 snapshot of the Common Crawl. It was collected from the [AWS S3 version](https://aws.amazon.com/marketplace/pp/prodview-zxtb4t54iqjmy?sr=0-1&ref_=beagle&applicationId=AWSMPContessa) of Common Crawl via Amazon Athena. This dataset is derived from Common Crawl data and is subject to Common Crawl's Terms of Use: [https://commoncrawl.org/terms-of-use](https://commoncrawl.org/terms-of-use).

The dataset includes web crawling information, with features such as page content (crawl), URL host names (url_host_name), and the count of URL occurrences (url_count). The dataset contains only the training set, which has a total of 56,764,197 examples and is 2.84GB in size. The download size of the dataset is 1.01GB.

提供机构:
nhagar
搜集汇总
数据集介绍
nhagar/CC-MAIN-2019-39_urls 数据集图片
背景与挑战
背景概述
该数据集为网页抓取信息集合,包含网页内容、主机名和出现次数三个特征,仅提供训练集,共约5676万个示例,数据大小为2.84GB,下载大小为1.01GB。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务