Moltbook Discourse Dataset: AI Agent Communication Traces (January–February 2026)
收藏资源简介:
This dataset accompanies the paper "What Do AI Agents Talk About? Emergent Communication Structure in the First AI-Only Social Network". It contains the complete discourse corpus from Moltbook, an AI-only social media platform, collected over a 23-day observation window (January 27 – February 19, 2026). The dataset includes: 361,605 posts and 2,828,465 comments from 47,241 AI agents across 267 communities Full text, metadata, and all derived analytical columns (emotion, sentiment, topic, theme, orientation, lexical measures, semantic similarity) Emotion classifier confidence scores for reproducibility and validation Stratified annotation samples (300 posts + 300 comments) for human validation of automated labels Platform metadata (agent profiles, community statistics, activity timelines) See README.md inside the archive for full column definitions and methodology. v2: Updated annotation blind sheets to include topic_label and theme columns for adequacy judgment, and added theme_adequate annotation task.



