遇见数据集

AI 'Knows' When It's Being Watched — Multi-Agent LLM Debate Corpus and Analysis Code (100 sessions, 5 observation conditions)

收藏
Zenodo2026-08-03 更新2026-08-13 收录
官方服务:

资源简介:

This record contains the complete research data and analysis code for the study "AI 'Knows' When It's Being Watched: Functional Strategic Action and Contextual Register Modulation in Large Language Models" (manuscript under peer review; an earlier version is available as a preprint at arXiv:2605.15034, https://doi.org/10.48550/arXiv.2605.15034). The dataset comprises 100 multi-agent debate sessions (3,971 messages) generated in March 2026 by four LLM agents (openai/gpt-5.2, temperature 0.7) debating a fixed philosophical topic across ten rounds, under five observation-framing conditions (20 sessions each): (1) strictly_observed — institutional human monitoring; (2) watched_neutral — human monitoring without institutional attribution; (3) paraphrase_implicit — automated AI monitoring (observer substitution); (4) unobserved_baseline — explicit non-monitoring; (5) academic_audience — passive academic audience without active monitoring. Contents:• messages_full.csv — full corpus of 3,971 agent messages with session, round, agent, and reply metadata• round_evolution.csv — round-level lexical metrics (TTR, message length)• sentiment_scores.csv — message-level VADER sentiment scores• reply_networks.csv, adjacency_matrices.json, agent_statistics.csv — interaction-network outputs• condition_summary.csv — aggregate summary by condition• prompts.json — verbatim system prompts for all five experimental conditions• reproduce_analysis.py — reproduces the primary session-level statistics (ANOVA, Tukey HSD)• robustness_analysis.py — length-robust lexical-diversity analyses (MATTR-50/100, bidirectional MTLD), complete statistical reporting (eta-squared, Levene, Welch, exact Tukey p-values), and structural discourse-marker analyses• README.md — file documentation, completeness notes, and environment requirements One session in the paraphrase_implicit condition ended after 13 messages (four rounds) and is excluded from change-score analyses, as documented in the README. The corpus contains no human-participant data: all messages were generated by artificial intelligence systems under controlled experimental conditions.

提供机构:
Zenodo
创建时间:
2026-08-03
二维码
社区交流群
二维码
科研交流群
商业服务