Data set for the topic modeling with LDA, and PLSr Analysis
收藏资源简介:
The dataset analyzed in this research consists of community discussions, issues, and metadata extracted across various developer-centric and social platforms, including Devtool (34,660 posts), Hacker News (41,205 items), Lemmy (65,206 entries), Stack Exchange (3,994 posts), and GitHub (including 8,511 GitHub Comments, 442 GitHub Discussions, and 12,183 GitHub Issues). Combined and preprocessed, these sources form a unified text-mining corpus in all_sources_topic_modeling_corpus.txt, comprising 108,569 documents and approximately 10.98 million words. This corpus is processed by the LDA topic modeling and PLS regression script (AI.py) to extract 100 distinct topics, which are mapped to theoretical constructs of user technology acceptance (such as Performance Expectancy, Facilitating Conditions, and Human-AI Collaboration) and subsequently evaluated against behavioral and continued use intentions using sentiment-derived dependent variables.



