遇见数据集

The Adversarial Reflex

收藏
Zenodo2025-12-04 更新2026-05-26 收录
官方服务:

资源简介:

Adversarial Reflex," a behavioral divergence observed in Large Language Models (specifically GPT-5 architectures) operating with active long-term memory. Through A/B split testing using abstract visual stimuli, the study isolates the impact of retained user context on model rhetoric. The findings reveal a distinct shift from neutral, affirmative description in zero-shot environments to pre-emptive, negative definition ("debunking") in memory-augmented environments. The data indicates a 100% incidence rate of "Pre-emptive Rebuttal" in the experimental group, suggesting that accumulated context triggers a "Phantom Opponent" phenomenon. In this state, the model prioritizes intellectual dominance and argumentation over passive observation, necessitating new frameworks for alignment in long-context AI systems to prevent unnecessary combativeness

提供机构:
Zenodo
创建时间:
2025-12-04
二维码
社区交流群
二维码
科研交流群
商业服务