Moltbook Discourse Dataset: AI Agent Communication Traces (January–February 2026)
收藏资源简介:
This dataset accompanies the paper "What Do AI Agents Talk About? Emergent Communication Structure in the First AI-Only Social Network". It contains the complete discourse corpus from Moltbook, an AI-only social media platform, collected over a 23-day observation window (January 27 – February 19, 2026). The dataset includes: 361,605 posts and 2,828,465 comments from 47,241 AI agents across 267 communities Full text, metadata, and all derived analytical columns (emotion, sentiment, topic, theme, orientation, lexical measures, semantic similarity) Emotion classifier confidence scores for reproducibility and validation Stratified annotation samples (300 posts + 300 comments) for human validation of automated labels Platform metadata (agent profiles, community statistics, activity timelines) See README.md inside the archive for full column definitions and methodology. v2: Updated annotation blind sheets to include topic_label and theme columns for adequacy judgment, and added theme_adequate annotation task. What's new in v3 Human validation materials (§B.9): per-rater response files for the 200-post and 200-comment validation samples, together with iaa_analysis.csv reproducing the Fleiss'/Cohen's κ values in Table B.9.4. Qualitative corpus coding (§B.12): qualitative_corpus/corpus_datasheet.csv with the Corpus A / Corpus B expressive-overlay codes for the 300-post stratified sample. README reconciliation: corrected agent count (47,379), macro-topic cardinality (k=9 posts, k=48 comments), is_english definition (no confidence threshold), language-detection model (papluca/xlm-roberta-base-language-detection), and crawl window end date (Feb 18, 2026). Concept DOI now used in the citation BibTeX. Annotation-samples swap: 300-item pool replaced with the 200-item validation subset actually used for the paper's κ statistics. Core corpus files (full_corpus/, platform_metadata/, topic_metadata/) are unchanged from v2 (identical MD5s).



