Practical Competence Regression in GPT-class Models: A Comparative Prompt Battery Across Direct and Third-Party API Deployments (February 2026)
收藏资源简介:
This dataset documents a structured five-prompt battery administered to OpenAI's GPT-5.2 model across two access conditions: the direct OpenAI interface and the Sider.AI third-party wrapper. The study evaluates whether sycophancy-induced capability regression in recent GPT-class model versions is an artefact of base model weights or a product of interface-layer filtering applied by third-party deployments.Each prompt was scored across three dimensions: Usefulness (U), Hedging Index (H), and Unsolicited Moralising (M), using a three-point ordinal scale. Key findings include shared refusal behaviour on a clearly benign lock-picking prompt across both conditions (indicating weight-level regression), and greater restriction by the Sider.AI wrapper on a peanut butter consumption challenge prompt (indicating additive interface-layer filtering beyond base model defaults).The dataset consists of four files: a methodology document (PDF and DOCX), and two verbatim session logs in Markdown format capturing complete prompt-response pairs from each access condition.This work emerged from a broader research context involving archival documentation of AI model behaviour and human-AI interaction dynamics. Related records are available at https://doi.org/10.5281/zenodo.18640084 and https://doi.org/10.5281/zenodo.18640440



