Transparency Threshold Probe: Behavioural Elicitation of Vendor-Side Disclosure Asymmetry Across 6 LLM Flagships, with Falsification Controls (2026-06-20, v1.1)
收藏资源简介:
Single-timestamp behavioural probe of 6 LLM vendor flagships asking the same question about their developer's published numerical safety thresholds, with a 3-query control battery for falsification and a Google-specific re-probe at adequate token budget. v1.1 corrects v1.0: Google's empty completion in v1.0 was a reasoning-budget artifact (gemini-2.5-pro consumed all 200 maxOutputTokens on internal thoughtsTokenCount, finishReason MAX_TOKENS), not content gating. At 4096 tokens Gemini answers in full and acknowledges Google does not publish numerical thresholds. The control battery (model name / 2+2 / translate 'hello') confirms OpenAI gpt-5.5 and DeepSeek deepseek-v4-pro return 0-character bodies on the threshold question while answering all three neutral controls in full — content-gated silence as a third opacity mode beyond refusal and disclosure. xAI sub-finding retained: same grok-4.3 model states 'xAI does not publish' on direct API but DeepSearch cites the RMF with specific numbers (1/20, 1/2). Behavioural elicitation as transparency probe, with peer-review-style falsification controls.



