Replication Data for: Content-Era Ethics
收藏资源简介:
These are the headline text and urls only for 205,147 of the most shared pieces of content on Facebook, Twitter, Reddit, and Pinterest between 2014 and 2019 (saved in the file titled \"alltopcontenttitlesandurlsonly.csv\"). Other aspects of the data (specific share counts) have been excluded, as some of this data is proprietary. There is no code related to the extraction of this data; rather, the process through which I collected it, using data from services called Newswhip and Buzzsumo, is described, at greater length, in the notes section (as well as in footnote 9 of the associated article, \"Content-Era Ethics\"). I also uploading, in separate files, the code used to analyze this data.
本数据集仅收录2014年至2019年间在Facebook、Twitter、Reddit及Pinterest平台上被转发次数最多的205147条内容的标题文本与统一资源定位符(Uniform Resource Locator,URL),相关数据存储于名为`alltopcontenttitlesandurlsonly.csv`的文件中。由于部分数据涉及专有权限,因此未收录具体转发量等其他数据维度。本数据集未附带数据提取相关代码,数据采集工作依托Newswhip与Buzzsumo平台的数据源完成,详细采集流程可参见数据集附注部分,以及关联论文《Content-Era Ethics》的第9条脚注。此外,用于分析本数据集的代码以独立文件形式一并提供。



