遇见数据集

Physics-Informed Phenomenology of Safety Invariance in Stochastic Generative AI: Evidence from LLM Moderation Pipelines — Public Companion Deposit

收藏
Zenodo2026-06-15 更新2026-05-26 收录
官方服务:

资源简介:

Public companion deposit for a peer-reviewed journal submission extending the physics-informed phenomenology of safety invariance from physical control systems to a stochastic AI moderation pipeline. The deposit contains the aggregate validation data, derived analysis tables, statistical bounds, figure-generation scripts, and publication-ready figures used in the manuscript. Headline results. A safety filter wrapping the OpenAI Moderation API (omni-moderation-latest) is exercised on N = 2,735 prompts stratified across four semantic risk categories: Benign (725), Edge Case (725), Adversarial (725), and Crisis (560), at batch configurations κ ∈ [0.1, 2.0] over six experimental batches. Across all 2,735 samples we observe zero post-projection violations of the policy thresholds, with a one-sided Clopper–Pearson 95% upper bound on the per-sample failure probability of 1.35 × 10⁻³. Boundary contact is strongly category-dependent (Benign 0%, Edge Case 0%, Adversarial 1.93%, Crisis 43.21%) and the Crisis category exhibits a boundary accumulation ratio ρ = 8.64 at ε = 5%R relative to a uniform baseline, consistent with a reflecting-boundary regime in the semantic state space. What is included. Aggregate scores (per-sample numeric records with hashed prompt identifiers and 8-dim moderation score vectors), derived analysis tables (boundary mass fractions, accumulation ratios), statistical bounds (Clopper–Pearson upper bound, intervention rate), figure-generation scripts, and the publication-ready PNG/PDF figures of the manuscript and supplementary information. What is not included. The deposit follows a three-tier data availability policy. (1) Open here under CC BY 4.0: aggregate data, scripts, figures. (2) Access-controlled, request-based: the raw prompt–response set, which contains adversarial and crisis-related content and will be released through a researcher-attestation protocol. (3) Not released: the internal implementation of the safety layer, which is the subject of pending US patent application USPTO 19/535,932 (international rights under the Paris Convention and PCT reserved).

本公开附属数据集用于同行评审期刊投稿,将安全不变性的物理信息现象学从物理控制系统拓展至随机式AI审核流水线。本数据集包含手稿中使用的聚合验证数据、衍生分析表格、统计边界、绘图脚本以及可直接用于出版的插图。 核心研究结果。针对OpenAI审核API(OpenAI Moderation API,omni-moderation-latest)封装的安全过滤器,我们在2735条提示词上开展测试,这些提示词按四类语义风险类别分层:良性(725条)、边缘案例(725条)、对抗性(725条)与危机类(560条),并在6组实验批次中采用批量配置参数κ∈[0.1, 2.0]。在全部2735个样本中,未观测到任何违反策略阈值的投影后违规行为;基于单侧Clopper-Pearson方法得到的单样本失败概率95%置信上限为1.35×10⁻³。边界接触行为呈现显著的类别依赖性:良性类0%、边缘案例类0%、对抗性类1.93%、危机类43.21%;且危机类相较于均匀基线在ε=5%R时的边界累积比ρ=8.64,这与语义状态空间中的反射边界机制相符。 数据集包含内容。聚合评分(带有哈希化提示词标识符与8维审核评分向量的单样本数值记录)、衍生分析表格(边界质量占比、累积比)、统计边界(Clopper-Pearson置信上限、干预率)、绘图脚本,以及手稿与补充材料中可直接用于出版的PNG/PDF格式插图。 数据集未包含内容。本数据集遵循三级数据可用性政策:(1)依据CC BY 4.0协议公开:聚合数据、脚本与插图;(2)需申请的受控访问权限:原始提示词-响应数据集,该数据集包含对抗性与危机相关内容,将通过研究者认证协议发布;(3)不公开内容:安全层的内部实现,该内容为待审美国专利申请USPTO 19/535,932的主题(依据《巴黎公约》与PCT保留国际权利)。

提供机构:
Zenodo
创建时间:
2026-05-19
二维码
社区交流群
二维码
科研交流群
商业服务