Architecture-Specific Canary Signals and Boundary Oscillation in Large Language Model Degradation: A Three-Architecture Empirical Study (Dataset and Analysis Code)
收藏资源简介:
This deposit contains the scored results dataset and statistical analysis code supporting the paper "Architecture-Specific Canary Signals and Boundary Oscillation in Large Language Model Degradation: A Three-Architecture Empirical Study" (under review at PLOS ONE). The dataset comprises 9,000 per-case scored evaluations: 2,700 from three large language model providers (Claude Sonnet 4.6, Gemini 2.5 Flash, and GPT-4.1-mini, 900 cases each) and 6,300 from seven heuristic baselines run across the same case set (900 cases each). Cases span two domains, three difficulty tracks, five degradation modes, and five instruction-clarity levels. The deposit includes a primary per-case results file, eleven aggregate diagnostic tables, and a Python script and notebook that reproduce all six tables and both figures in the paper from the primary file. The benchmark generation code, prompt templates, degradation implementation, and scoring rubric are proprietary and not included. The Methods section of the paper provides sufficient specification for an independent implementation. All language model evaluations were run at temperature zero with identical system instructions. No human subjects, animal subjects, or personally identifiable data are involved.



