Standard Experimental Protocol (SEP v2.4) From Prompt Coherence Engines (PCE) to Semantic Trajectory Stabilization
收藏资源简介:
As Large Language Models (LLMs) are increasingly deployed in high-stakes environments, traditional alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and prompt-based safety filters remain vulnerable to surface-level heuristics and adversarial prompts. This paper formalizes Axiomatic Behavioral Consistency, an alternative paradigm that models alignment as a set of structural constraint operators overan LLM’s generative distribution. We introduce the Prompt Coherence Engine (PCE), a framework designed to anchor internal logical resilience through specialized fine-tuning and strict system-level prompts. To evaluate these claims, we present the Standard Experimental Protocol (SEP v2.4), a rigorous, cross-model benchmarking framework that explicitly tests multidimensional coherence under adversarial conditions, identity hijacking, and contradictory constraints. By integrating systematic ablation tests (shuffled structures, semantic paraphrasing, and single-axiom removal), a strict reproducibility registry, and a formalized evidence hierarchy (Levels L1–L7), this work provides an empirical foundation to determine whether robust logical resilience can emerge as a foundational property of internal axiomatic topology rather than fragile behavioral heuristics.



