Cognitive Digital Shadows (CDS): LLM-generated debates on societal issues under persona and AI-assistant conditioning
收藏资源简介:
Cognitive Digital Shadows (CDS) Dataset FIS2 Project PENSO | Funded by MUR | CogNosco Lab, University of Trento LLMs · Dataset · Opinion Pooling · Reasoning · Social Media GitHub: https://github.com/NaviDATA-Repos/PENSO_Data_WP-ConvinceMe_FIS2_UniTrento Overview This repository contains the official dataset for Cognitive Digital Shadows (CDS), a synthetic corpus of approximately 190,000 debate records generated by 19 Large Language Models (LLMs) under controlled prompting conditions. The CDS dataset is designed to investigate how state-of-the-art LLMs discuss socially sensitive and controversial issues (e.g., COVID-19 Vaccines, Fake News, Gender Gap and STEM Stereotypes) when configured to shadow human sociodemographics, personality traits and social media behaviours across both Opinion and Reasoning Summary layers. Key Capabilities & Representations The CDS corpus enables multi-level semantic, cognitive and affective analysis across synthetic debate records: Semantic Frames: capturing recurring conceptual structures and meaning patterns across pooled texts. Mindset Streams: representing dominant cognitive, interpretative and discursive orientations. Emotional Flowers: modelling the distribution, intensity and co-occurrence of affective states. Textual Forma Mentis Networks (TFMNs): extracting network structures of connected concepts to evaluate conceptual centrality and cognitive structure. Rankings: enabling comparative evaluations across semantic frames, mindset streams and emotional dimensions. Repository Structure Code/: analysis and execution scripts Data/ Raw_Data/: raw LLM-generated outputs across 19 models Processed_Data/ Data_visualization/: plots and graphical outputs per model EdgeList/: extracted network edge lists (CSV & Python .pkl) Merged_per_Model/: aggregated edge lists for cross-model comparison Hypothesis_Testing/: Kruskal–Wallis test results & p-values TFMN_EmoA_stats/: Text/Feature Mining & Affective statistics Data Pooling System/: system for custom sociodemographic querying



