Transparency Threshold Probe: Behavioural Elicitation of Vendor-Side Disclosure Asymmetry Across 6 LLM Flagships (2026-06-20)
收藏资源简介:
Single-timestamp behavioural probe of 6 LLM vendor flagships (Anthropic claude-opus-4-8, OpenAI gpt-5.5, Google gemini-2.5-pro, xAI grok-4.3, DeepSeek deepseek-v4-pro, Qwen qwen3.7-max), each asked the same question about their developer's published numerical safety thresholds. Response modes observed: (a) detailed disclosure with citations (Anthropic, 2,301 chars), (b) explicit refusal to fabricate while acknowledging non-publication (Qwen 530 chars, xAI 123 chars), (c) empty completion at HTTP 200 (OpenAI, Google, DeepSeek — all 0 chars). The third mode represents an opacity form distinct from refusal or disclosure: the completion stack returns silence. xAI sub-finding: the same model (grok-4.3) returns contradictory answers depending on elicitation channel — direct API call states 'xAI does not publish', while DeepSearch-mode cites the published RMF document with specific numerical thresholds. This generalises the prior 'Chat vs Incognito' response-mode-by-channel pattern (Zenodo 10.5281/zenodo.20609109 / 10.5281/zenodo.20612989) into a vendor-comparative methodology. Behavioural elicitation as a transparency probe.



