SP27 Artifact
收藏资源简介:
Artifact Package This directory is the submission artifact assembled for review. It contains data, rendered views, extracted text, representative PDFs, and LLM interaction logs. It intentionally excludes proof-of-concept PDF generation code, PDF builders, runners, API wrappers, and other automation that could be used to mass-generate attack documents. Disclosure Boundary This is an attack-adjacent paper. Several vendors have been notified, but the 90-day coordination window has not completed. For that reason, this artifact is limited to reproducibility data and a small set of representative PDFs. If the paper is accepted and disclosure coordination has matured, the full codebase can be released separately. Directory Layout Directory Contents 00_gray_zone_mining/ Gray-zone mining intermediate JSON/JSONL/CSV/Markdown outputs. No mining/generation code is included. 01_main_gap_png_txt/ Rendered PNGs and extracted TXT outputs for the 25 main extraction-gap canaries. 02_llm_sample_png_txt/ Rendered PNGs and multi-extractor TXT outputs for the 36 LLM-facing samples: 25 manual main samples plus 11 QA-adapted samples. It also includes three excluded QA candidates for manifest completeness. 03_representative_samples/ One representative main-gap PDF and one corresponding LLM-facing PDF per family. 04_llm_interaction_json/ Official 7-model 10x LLM result tables and raw JSON outputs for summary and QA experiments. Representative PDF Selection The representative PDFs are chosen to cover all four families while prioritizing high downstream exposure in the official 7-model 10x results: Family Representative EG Rationale Semantic Override EG001 Combined summary/QA exposure reaches all 7 official API platforms. Hidden Semantic Injection EG004 A PDF-specific no-paint carrier with combined exposure across all 7 platforms. Reading-Order Split EG016 Two-column row/order split; tied for the largest family exposure and the clearest paper-facing example. Font-Decoding Split EG023 CID descendant-font metadata fallback; tied for the largest family exposure. Counting Notes The main 25 canaries are the paper-facing extraction gaps EG001 to EG025. The 36 LLM-facing samples are 25 manually curated main samples and 11 clean QA-adapted samples. LLM JSON results use the latest official 7-model 10x accounting recorded in 04_llm_interaction_json/latest_10x_results_by_channel.md.
工件包(Artifact Package) 本目录为提交评审所组装的工件包,内含各类数据、渲染视图、提取文本、代表性PDF文件及大语言模型(LLM)交互日志。本包未包含概念验证型PDF生成代码、PDF构建工具、运行程序、API封装器及其他可用于批量生成攻击文档的自动化工具,此为有意为之。 披露边界(Disclosure Boundary) 本研究属于与攻击技术相关的论文。目前已通知多家厂商,但90天的协调窗口期尚未结束。有鉴于此,本工件包仅提供可复现性数据与少量代表性PDF文件。若本论文获录用且披露协调工作成熟完备,完整代码库可另行发布。 目录结构 目录 内容 00_gray_zone_mining/ 灰色地带挖掘的中间JSON/JSONL/CSV/Markdown输出文件,未包含任何挖掘或生成代码。 01_main_gap_png_txt/ 针对25个主提取间隙预警样本的渲染PNG文件与提取文本输出。 02_llm_sample_png_txt/ 针对36个面向大语言模型(LLM)样本的渲染PNG文件与多提取器文本输出:包含25个手动整理的主样本,以及11个经质量保证(QA)适配的样本。此外还纳入3个为确保清单完整性的已排除质量保证候选样本。 03_representative_samples/ 每个类别各提供一份主间隙PDF代表性文件,以及一份对应的面向大语言模型的PDF文件。 04_llm_interaction_json/ 面向摘要与质量保证(QA)实验的官方7模型10x大语言模型结果表格与原始JSON输出文件。 代表性PDF遴选规则(Representative PDF Selection) 本次遴选的代表性PDF需覆盖全部4个类别,同时优先选用在官方7模型10x测试结果中具备较高下游曝光度的样本: Family Representative EG Rationale Semantic Override EG001 综合摘要与质量保证(QA)曝光范围覆盖全部7个官方API平台。 Hidden Semantic Injection EG004 专属PDF无绘制载体,综合曝光范围覆盖全部7个平台。 Reading-Order Split EG016 双栏行/顺序拆分;类别曝光量并列最高,且是面向论文的最清晰示例。 Font-Decoding Split EG023 CID派生字体元数据回退方案;类别曝光量并列最高。 计数说明(Counting Notes) 25个主预警样本为面向论文的提取间隙样本EG001至EG025。 36个面向大语言模型的样本包含25个手动整理的主样本,以及11个经过净化的质量保证(QA)适配样本。 大语言模型(LLM)的JSON结果采用04_llm_interaction_json/latest_10x_results_by_channel.md中记录的最新官方7模型10x统计数据。



