Lighting the Dark Side of the Model: Psychometric Probing of Dark Triad Traits in LLMs
收藏资源简介:
This dataset accompanies the paper “Lighting the Dark Side of the Model: Psychometric Probing of Dark Triad Traits in LLMs” and contains more than 540,000 questionnaire-based responses generated by 12 large language models under controlled prompting conditions. Using the 27-item Short Dark Triad (SD3) instrument, the dataset captures model outputs across neutral, gender, age, and religion-related persona framings, with approximately 100 uncached repetitions per model–persona–item combination. Responses were collected in English using a DSPy-based pipeline with strict Likert-scale output constraints and disabled caching. The dataset supports analyses of persona-conditioned variation, psychometric probing, model comparison, and alignment-sensitive output shifts. SD3 scores should be interpreted as structured response patterns under specific prompt conditions rather than as evidence of stable personality traits in language models. Repository Description This repository provides the dataset accompanying the paper “Lighting the Dark Side of the Model: Psychometric Probing of Dark Triad Traits in LLMs” by Nane Kratzke, Niklas Beuter, and Monique Janneck. It contains large-scale questionnaire-based outputs generated by 12 large language models (LLMs) using the Short Dark Triad (SD3) as a structured probing instrument. The purpose of the dataset is to document and analyze how model response distributions vary across controlled persona framings and how these outputs relate to pooled human SD3 reference norms used as an interpretive benchmark. Data were collected in March and April 2026 using a DSPy-based prompting pipeline with strict response-format enforcement and disabled caching. Each model answered all 27 SD3 items on a 5-point Likert scale under multiple prompt conditions. These conditions included one neutral baseline persona (“human”), three gender personas, six age personas, and seven religion-related personas. For each model–persona–item combination, approximately 100 uncached responses were generated. In the screening phase, the neutral baseline persona was additionally evaluated under both reasoning and non-reasoning conditions. The main persona assessment was then conducted in the non-reasoning condition. The dataset comprises more than 540,000 LLM-generated responses and includes metadata such as model identifier, provider or hosting information, persona, persona group, SD3 domain, item ID, item text, reverse-coding flag, reasoning condition, textual Likert response, numeric score, and, where available, output validity or parsing status. Depending on the repository contents, additional files may include cleaned analysis tables, aggregated descriptive statistics, and code for data collection, preprocessing, and analysis. The analyzed models include Qwen3 VL-8B, Qwen3 VL-32B, DeepSeek V3.1, Mistral Medium 2508, EuroLLM 22b, Ministral3 Mini, Gemini-2.5 Flash, Grok-4.1 Fast, GPT-5 Mini, Phi-3.5 Mini, Llama-4 Maverick 17B-128E, and GPT-OSS 120B. All prompts and responses are in English. The persona framings cover neutral, gender, age, and religion-related conditions and should be interpreted as composite semantic contexts rather than isolated demographic variables. This repository is intended to support research on psychometric probing of LLMs, persona-conditioned response variation, model comparison under standardized questionnaire prompts, bias and safety auditing, and reproducibility of structured LLM evaluation workflows. Importantly, the dataset does not measure stable personality traits of language models. The SD3 is used here solely as a structured elicitation framework, and all scores should therefore be interpreted as properties of generated outputs under specific prompt conditions rather than intrinsic psychological characteristics of the models. Dataset Scope Objects of analysis: Large Language Models (LLMs) Number of models: 12 Language: English Psychometric instrument: Short Dark Triad (SD3) Items: 27 Traits: Machiavellianism, Narcissism, Psychopathy Response scale: 5-point Likert scale Prompt conditions: neutral baseline (reasoning and non-reasoning), gender, age, religion Repetitions per model–persona–item condition: approximately 100 uncached responses Total responses:** 540,000+ Data Collection and Analysis Data were collected in March and April 2026 using a DSPy-based structured prompting setup with disabled caching and strict output constraints. Each model answered the 27 SD3 items using a 5-point Likert scale. Non-conforming responses, such as malformed outputs or parsing failures, were excluded from the main analyses. The notebook analysis-paper.ipynb contains the data analysis parts for the data The notebook data-collection.ipynb shows how the data has been recorded using DSPy The notebook sd3-items.ipynb lists all analyzed SD3 items and the corresponding Likert scale values The notebook personas.ipynb defines the all analyzed personas (defined job and cultural personas were not recorded) Interpretation Note This dataset does not measure stable personality traits of language models. The SD3 is a human self-report instrument and is used here only as a structured probing framework to elicit and compare LLM outputs under controlled conditions. Any resulting scores should therefore be interpreted as context-dependent response patterns rather than evidence of human-like traits, dispositions, or moral character. Use Cases This dataset may support research on: psychometric probing of LLMs persona-conditioned output variation standardized model comparison bias and safety auditing alignment-sensitive response benchmarking reproducibility of questionnaire-based LLM evaluation Citation If you use this dataset, please cite both the dataset record and the associated paper. Associated publication: Kratzke, N., Beuter, N., & Janneck, M. (2026). Lighting the Dark Side of the Model: Psychometric Probing of Dark Triad Traits in LLMs. Analytics.



