datasets for the manuscript: The Transpacific "Short Circuit" of the AGI imaginary: Dwarkesh Patel's The Lunar Society and his envisioned Post-Liberal Counter-Public(s)
收藏资源简介:
This paper's empirical analysis draws on a multi-source dataset spanning YouTube, TikTok, WeChat, Bilibili, and RedNote. The files listed above fall into four methodological categories corresponding to distinct analytical stages: (1) attention measurement, (2) thematic/sentiment classification, (3) semantic drift and re-coding analysis, and (4) supporting metadata for preprocessing and alignment. The data files used in our analysis are as follows: (1) Attention Index Construction (US vs. China) These visualizations and spreadsheets contain raw engagement metrics and their derived normalized indices: China_US Comparison.png – visualization comparing both indices on a shared temporal axis. US_Attention_Index.png and Chinese_Attention_Index_all.png – static graphs showing individual platform trajectories. Chinese_Attention_Index_Time_.png – temporal line chart highlighting shifts in Chinese attention. CN & US_Attention_Index_comparison.xlsx – merged numerical results used to compute correlation coefficients. US_Attention_Index_.xlsx and CN_Attention_Index.xlsx – raw metrics and normalized attention index scores. youtube metrics.xlsx, tiktok metrics.xlsx, and wechat metrics (quotes and forwards).xlsx – the platform-specific raw engagement data used as inputs for index computation. These files operationalize the first methodological step: constructing quantitative indicators of audience attention using min–max normalization and weighting procedures described in the methodology section. (2) Cross-platform metric preprocessing and alignment These spreadsheets document scraped data and the cleaning/normalization process: merged_tiktok_youtube_data (2).xlsx – cross-platform alignment file linking TikTok clips to corresponding full YouTube episodes. TikTok_Scraping_Detail.xlsx, TikTok_Final_Metrics_1207_1623.xlsx, and tiktok_analysis (parsed).xlsx – raw TikTok metadata and derived depoliticization/aestheticization scores. youtube_analysis (cleaned with texts).xlsx, youtube_analysis (views+likes).xlsx, and youtube_keyword_analysis.xlsx – cleaned YouTube datasets used for comparative semantic analysis. wechat_parsed (with texts) 1205.xlsx – full parsed WeChat corpus with extracted metadata and linked episode identities. Translate_Result_1439.xlsx – machine-translated output file used in standardizing multilingual text for embedding processing. Together, these files constitute the preprocessing layer for the second methodological step—cross-platform alignment and topic/phrase normalization before keyword extraction, embedding, and drift measurement. (3) Thematic and sentiment analysis corpora These materials contain outputs and bilingual versions of the discourse/sentiment coding pipeline: Report on WeChat contents thematic and sentiment analysis (EN).docx – English-language results summary. Report on WeChat contents thematic and sentiment analysis 微信语料...CN.pdf – Chinese report used for validation and comparison. These correspond to the DHA–computational iterative coding described earlier, informing the construction of semantic labels and theoretical categories deployed in the embedding analysis. (4) Intermediate visualization files used in semantic drift and narrative reframing analysis These images serve to validate and communicate comparative semantic drift and cross-platform re-coding tendencies: China_US Comparison.png Chinese_Attention_Index_all.png Chinese_Attention_Index_Time_.png US_Attention_Index.png These plots were generated after vectorization and normalization pipelines described in the methodology, representing the empirical grounding for claims of attention decoupling and geopolitical reframing. Together, these files constitute a traceable workflow that: extracts platform-specific audience reactions (attention indices); aligns disparate platform corpora for cross-context comparison; supports embedding-based semantic drift measurement; and enables DHA-informed interpretation of platform-mediated narrative recoding. This layered documentation ensures reproducibility and enables triangulation between computational models, discourse analysis, and geopolitical contextualization.



