Deployed Quantization Tier and Lossy Context Compression in Extractive QA: Aggregate Dataset and Reproducibility Artifact
收藏资源简介:
An aggregate-only dataset and reproducibility artifact for a corrective paired grid: three deployments of Qwen3.8-27B on one 12 GiB laptop GPU (Q4_K_XL GGUF partly offloaded, EXL3 2.0 bpw and IQ2_S GGUF fully resident) answered 117 private and 120 WildChat extractive questions under original context, LLMLingua-2, and two rungs of a stopword/code-stripping pipeline (24 cells, 2,844 scored requests). No cross-tier interaction survives Holm correction across 12 CR2/Satterthwaite cluster-robust tests (estimates -5.0 to 6.0 points, pointwise envelope -11.6 to 12.9); compressor choice dominates (LLMLingua-2 costs 15-21 points, the naive rungs 47-62). The package contains the aggregate claims record, flat tables, a standard-library verifier that recomputes every statistic and paper number from per-cluster histograms, path-free environment provenance, the frozen measurement and analysis sources, and the technical report. It is a corrective exploratory rerun, not an independent confirmation; no conversation text, question, answer, response, or per-item row from either corpus is included. AI assistance: Claude (Anthropic) and Codex (OpenAI) assisted with evaluation-harness development, experiment orchestration, evidence auditing, analysis and figure code, citation checking, and manuscript drafting and revision. Matthew Schwartz directed the research and is responsible for its methods, results, interpretation, claims, citations, rights, and released artifacts.



