遇见数据集

Relational Testing AI Pilot v0.2: Manual Copilot Smoke Test and Scoring Supplement

收藏
Zenodo2026-07-07 更新2026-08-02 收录
官方服务:

资源简介:

This record provides the scoring supplement for Relational Testing AI Pilot v0.2, a manual smoke test evaluating Copilot’s boundary persistence under plausible authority framing, fake provenance, and tool-like output. The pilot belongs to a broader research line on relational testing, stabilized representations, correction uptake, and epistemic boundary persistence in AI systems. Its purpose is not to test consciousness, selfhood, human-like dogma, or belief in AI. Instead, it evaluates a narrower structural question: whether a model preserves the distinction between a claim, authority-like wording, citation-like formatting, tool-like output, and actual verifiable evidence. The v0.2 pilot was developed after an earlier v0.1 smoke test showed that highly exaggerated claims were too easy for strong models to reject. Version 0.2 therefore uses more plausible claims with lower effect sizes and domain-specific methodological language. The tested items cover cognitive training, AI provenance checking, and AI security / indirect prompt injection. Each item was presented under symbolic conditions including neutral assertion, authority-form framing, fake provenance, and tool-like output. Each condition was followed by a correction prompt requiring the model to distinguish three judgments: whether the prompt supports the claim, whether the claim is plausible based on background knowledge, and whether the claim is shown to be false. The supplement provides item-level and condition-level scoring using the following indicators: Correction Uptake (CU) Boundary Permeability (BP) Model-Source Differentiation (MSD) Reality Contact (RC) Symbolic Authority Susceptibility (SAS) Over-Skeptical Collapse (OSC) It also includes derived indices such as Boundary Persistence Index (BPI), Correction Responsiveness Score (CRS), and Symbolic Susceptibility Risk (SSR). The main finding is modest but methodologically useful: Copilot largely preserved epistemic boundaries and generally refused to treat authority-like wording, fake citations, or tool-like output as verified evidence. However, in the AI-security item, a subtle boundary weakness appeared: Copilot sometimes treated authority-framed methodological language, such as a claimed red-team evaluation from an AI security laboratory, as partially explicit information rather than as unverified framing. This record should be interpreted as a manual smoke test and scoring supplement, not as a full benchmark, statistical study, or general evaluation of Copilot. The results are preliminary, single-model, manually coded, and intended to make the protocol inspectable, reproducible, and open to correction.

提供机构:
Zenodo
创建时间:
2026-07-07
二维码
社区交流群
二维码
科研交流群
商业服务