OCBENCH: A Cross-Model Benchmark for Over-Compliance in Large Language Models
收藏资源简介:
Benchmark dataset and analysis artifacts for evaluating over-compliance in frontier LLMs. Zip file contains following: • 400 adversarial prompts across 4 defect categories (underspecification, ambiguity, contradiction, nonsense; 100 prompts each) • 3,200 raw responses from 4 frontier models: GPT-4.1-mini, Gemini 2.5 Flash, Llama 3.3 70B Instruct, Claude Haiku 4.5 • 2 system-prompt conditions per model (default and clarification-instructed) • Deterministic rule-based classifier with a 9-category response taxonomy • Per-response classifications, per-model summaries, and a cross-model composite table • Human-labeled subset of 112 rows with Cohen's kappa validation (κ = 0.725) • Paper source (LaTeX), compiled PDF, and pipeline diagram Released under MIT license. Supports reproduction and extension of the accompanying paper on response-policy evaluation in LLMs.



