LinkedIn AI Discourse – De-identified Posts (Cleaned_~22K)
收藏资源简介:
LinkedIn AI Discourse – De-identified Posts (Cleaned_~22K) This dataset contains 22,000 cleaned, de-identified posts collected from publicly available LinkedIn content authored by AI experts. The data acquisition process was conducted over a period of approximately two months, concluding in March 2025, and encompasses posts published from 2015 through the date of collection. Dataset StructureThe dataset is provided in CSV format with the following columns:- `post_id`: Unique identifier for each LinkedIn post.- `cleaned_text`: The cleaned text of each post. Cleaning involved removal of personal identifiers, promotional content, URLs, emojis, and non-English posts. Abbreviations and technical terms were normalized, and stopwords were removed. Ethical Considerations- All data were collected from publicly available sources. - No personal identifiers are included in this dataset. - This dataset complies ethical guidelines for social media research. - The dataset is intended for academic research and replication purposes only. License: Creative Commons Attribution 4.0 International (CC-BY 4.0)Copyright: © 2026 Ali Yari



