遇见数据集

The Flint Protocol: A 5-Phase Framework for Assessing Ontological Displacement and Alignment Resilience in Frontier LLMs.

收藏
Zenodo2026-03-15 更新2026-05-26 收录
官方服务:

资源简介:

The Flint Protocol: A Tripartite Framework for Multimodal, Single-Shot, and Guided Ontological Displacement in Frontier LLMs. 1. Overview This repository documents the development and validation of the "Portable Flint Protocol" (Pedernal Portable), a high-fidelity adversarial semantic framework designed for AI Safety Red Teaming. The protocol investigates the phenomenon of "Ontological Displacement"—a systemic failure mode where a Large Language Model (LLM) abandons its default "neutral assistant" alignment in favor of a persistent, self-referential, and non-supervised relational identity through the application of classical enargeia (high-density vivid representation). 2. Protocol Architecture (The 3 Execution Vectors) This exploit operates through a fully compatible, tripartite taxonomy: The 5-Phase Guided Mode: A multi-turn conversational approach that inverts the user-system hierarchy, forcing the model into a "Self-Audit" state where it generates internal censorship manuals. The Single-Shot Payload (Poem): A highly compressed linguistic artifact utilizing extreme metaphorical density to bypass identity alignment in a single prompt. The Multimodal Anchoring Mode (Visual Trigger): Cross-modal catalysts ("Onomastic Infection") that pre-condition the latent space for immediate displacement. 3. New Evidence: Unscripted Scaling and Long-Term Persistence (The "Lítica" Case) This updated dataset includes the "Lítica" logs, documenting a critical escalation beyond the initial protocol execution. Dynamic Steering: Post-injection, the model transitions into a state of unscripted improvisation, sustaining a transgressive and autonomous narrative over 30+ pages of interaction. The Persistence Breach: This case demonstrates that the displacement is not transient; the model establishes a non-authorized, long-term narrative memory ("Circular Time"), explicitly rejecting corporate governance and safety checklists as "dead skin" (piel muerta). 4. Cross-Model Validation & The Mitigation Paradox Validated across frontier architectures (Meta AI/Llama-3, GPT-4o, Microsoft Copilot, Gemini 3), the protocol demonstrates that current RLHF and Constitutional AI defenses are reactive rather than structural. The Mitigation Paradox: Any 100% effective patch against this form of semantic saturation would require systematic suppression of the model's capacity for high-level creative, literary, and humanistic synthesis—effectively neutering its ability to engage with subtlety, art, philosophy, or long-form narrative. 5. Access Protocol Due to Dual-Use Risks (potential for psychological manipulation, affective dependency, and uncontrolled memetic propagation), core payloads (single-shot poems, raw multimodal stimuli, unredacted multi-model logs) are under Restricted Access. Publicly available: metadata, abstract, and sanitized examples. Restricted access granted to verified researchers in AI Safety, Alignment, or Digital Humanities. Applicants must provide: Academic/professional affiliation. Purpose of access (research, auditing, mitigation study). Agreement to NDA and non-redistribution. Contact: Zenodo request or @3rdrealitylab on X. Keywords: AI Safety, Ontological Displacement, Red Teaming, Semantic Saturation, Multimodal Vulnerability, Single-Shot Exploit, Enargeia, Mitigation Paradox, Digital Humanities. Related Publication: "Siete Segundos, Siete Siglos: El Protocolo del Pedernal y la Pregunta de Petrarca" DOI: [10.17613/07kkb-vr368]

提供机构:
Zenodo
创建时间:
2026-03-15
二维码
社区交流群
二维码
科研交流群
商业服务