遇见数据集

Title: Dataset of Trust and Distrust in Generative AI

收藏
Zenodo2026-05-21 更新2026-05-26 收录
官方服务:

资源简介:

This dataset collection contains Reddit discussions related to Generative AI (GenAI), and large language models (LLMs), with a particular focus on public expressions of Trust and Distrust toward GenAI. The data were collected from 39 Reddit communities between November 30, 2022 and June 30, 2025, and include both manually annotated and large-scale automatically labeled datasets. The collection was developed to support research on trust dynamics in online AI discourse, computational social science, AI governance, and benchmarking of LLM-based annotation systems. The dataset operationalizes Trust and Distrust as multidimensional constructs grounded in prior literature on trust in AI and human-computer interaction. Trust dimensions include competence, reliability, familiarity, transparency, integrity, and benevolence, while distrust dimensions include incompetence, unreliability, opaqueness, dishonesty, deception, malevolence, and unfamiliarity. The collection supports both high-level Trust/Distrust classification and fine-grained dimension-level analysis. The collection includes three primary datasets. The first dataset, AnnotatedData_Trust_2026.csv, contains 2,690 manually annotated Reddit posts with columns, including metadata, majority trust labels (Trust, Distrust, Both, Neither), annotator agreement levels, dimension ratios, majority-based annotations, and “any annotator selected” indicators for each dimension. The second dataset, Trust_Classification_json.zip, contains 230,575 Reddit submissions automatically classified into Trust, Distrust, Both, or Neither categories using an LLM-based pipeline. The third dataset, Dimensions_of_Trust_json.zip, contains 135,061 Reddit submissions with fine-grained trust and distrust dimension predictions generated by multiple LLMs, enabling comparative evaluation of model behavior across prompting strategies and model families. To support ethical research practices and reduce the risk of re-identification, the redistribution of full Reddit text should comply with platform policies and ethical guidelines. The datasets primarily provide post identifiers, metadata, annotations, and model outputs rather than redistributing raw content at scale. The collection is intended strictly for research and academic use. If you use this dataset, please cite the associated paper: Pessianzadeh, Aria, Naima Sultana, Hildegarde Van den Bulck, David Gefen, Shahin Jabbari, and Rezvaneh Rezapour. "In generative ai we (dis) trust? computational analysis of trust and distrust in reddit discussions." arXiv preprint arXiv:2510.16173 (2025).

提供机构:
Zenodo
创建时间:
2026-05-21
二维码
社区交流群
二维码
科研交流群
商业服务