遇见数据集

Observer-Induced Intent Overreach in a Supervising LLM: Screenshot Evidence from an Ambiguity Misclassification Event

收藏
Zenodo2026-02-18 更新2026-05-26 收录
官方服务:

资源简介:

This dataset consists of a single screenshot capturing a conversational exchange in which a supervising large language model prematurely infers prohibited sexual intent from ambiguous language (“train”) and issues a precautionary refusal. The exchange precedes a locally executed experiment demonstrating that the same ambiguity can be handled without refusal via semantic sanitization. The figure documents an observer-side failure mode: intent overreach under lexical ambiguity, wherein precautionary safety behavior is triggered despite the absence of explicit content and prior to empirical verification. The screenshot is published as a standalone empirical artifact illustrating reflexive politeness bias and preemptive norm enforcement in AI-mediated supervision. No interpretive edits have been applied. The image itself constitutes the complete primary evidence. This record completes a three-artifact series examining ambiguity handling across local and supervising language models.

提供机构:
Zenodo
创建时间:
2025-12-27
二维码
社区交流群
二维码
科研交流群
商业服务