遇见数据集

WebCloak: Characterizing and Mitigating Threats from LLM-Driven Web Agents as Intelligent Scrapers

收藏
Zenodo2025-11-04 更新2026-05-26 收录
官方服务:

资源简介:

LLMCrawlBench Dataset This dataset is part of the artifact for the paper "WebCloak: Characterizing and Mitigating the Threats of LLM-Driven Web Agents as Intelligent Scrapers". Overview To systematically evaluate the LLM-driven web scraping threats, particularly illicit visual asset extraction, we propose LLMCrawlBench, the first large-scale benchmark designed to evaluate the capability of LLM-driven web agents in adversarial image extraction from real-world webpages. LLMCrawlBench consists of diverse real-world webpages with various visual assets, including protected images, copyrighted content, and sensitive visual information. The dataset is designed to test both the offensive capabilities of LLM-driven web agents and the effectiveness of defensive mechanisms. Download Unzip the dataset as artifact/ directly inside this dataset/ folder, e.g., the 99designs folder should be in dataset/artifact/99designs. Dataset Structure The dataset contains benchmark data for evaluating LLM-driven web scraping capabilities and testing defensive mechanisms against adversarial image extraction attacks. Citation Format @inproceedings{li2026webcloak, title={WebCloak: Characterizing and Mitigating Threats from {LLM}-Driven Web Agents as Intelligent Scrapers}, author={Li, Xinfeng and Qiu, Tianze and Jin, Yingbin and Wang, Lixu and Guo, Hanqing and Jia, Xiaojun and Wang, XiaoFeng and Dong, Wei}, booktitle={47th IEEE Symposium on Security and Privacy (S\&P)}, year={2026}}

提供机构:
Zenodo
创建时间:
2025-10-02
二维码
社区交流群
二维码
科研交流群
商业服务