Tracking the Trackers
收藏资源简介:
跟踪跟踪器是对万维网上第三方跟踪器的大规模分析。我们从CommonCrawl 2012语料库的35亿多个网页中提取第三方嵌入,并将这些嵌入汇总到包含4100万多个域中的1.4亿多个第三方嵌入的数据集中。我们提供了最近对web上第三方跟踪器的大规模分析中使用的数据。我们创建了一个提取器,用于从HTML页面中查找嵌入的第三方资源,并在CommonCrawl 2012 web爬网中包含的35亿网页上运行它。
Tracking Trackers is a large-scale analysis of third-party trackers on the World Wide Web. We extract third-party embeddings from over 3.5 billion web pages in the CommonCrawl 2012 corpus, and aggregate these embeddings into a dataset containing over 140 million third-party embeddings across more than 41 million domains. We provide the data used in a recent large-scale analysis of third-party trackers on the web. We developed an extractor to identify embedded third-party resources from HTML pages, and ran it on the 3.5 billion web pages included in the CommonCrawl 2012 web crawl.




