PERCEIVE - ENGAGING THE PEOPLE': IS SOCIAL MEDIA COVERAGE OF EU POLICY ASSOCIATED WITH PUBLIC SUPPORT FOR EUROPEAN INTEGRATION?
收藏资源简介:
<strong>README file</strong> Data Set Title:<strong> “PERCEIVE - ENGAGING THE PEOPLE’: IS SOCIAL MEDIA COVERAGE OF EU POLICY ASSOCIATED WITH PUBLIC SUPPORT FOR EUROPEAN INTEGRATION?”</strong> Data Set Authors: <strong>Vitaliano Barberio</strong> (Wirtschaftsuniversität Wien), ORCID <strong>http://orcid.org/0000-0002-2615-5006</strong>; <strong>Luca Pareschi </strong>(Università di Roma Tor Vergata), ORCID <strong>http://orcid.org/0000-0002-4402-9329</strong>; Data Set Contributors: <strong>Ines Kuric</strong> (Wirtschaftsuniversität Wien); <strong>Edoardo Mollona</strong> (Università di Bologna), ORCID <strong>http://orcid.org/0000-0001-9496-8618</strong><strong>.</strong> <strong>Markus Höllerer</strong> (Wirtschaftsuniversität Wien); <strong>http://orcid.org/0000-0003-2509-2696</strong> Data Set Contact Person: <strong>Luca Pareschi </strong>(Università di Roma Tor Vergata), ORCID <strong>http://orcid.org/0000-0002-4402-9329</strong>; luca.pareschi@uniroma2.it . Data Set License: this data set is distributed under a Creative Commons Attribution (CC BY) 4.0 International license Publication Year: <strong>2021</strong> Project Info: <strong>PERCEIVE (Perception and Evaluation of Regional and Cohesion Policies by Europeans and Identification with the Values of Europe), </strong>funded by European Union, Horizon 2020 Programme. Grant Agreement num.<strong> 693529</strong>;<strong> https://www.perceiveproject.eu/</strong>. <strong>Data set Contents</strong> The data set consists of: 1 README file 6 textual qualitative file saved in .txt format “<strong>stoplist_file_[nation].txt</strong>” 12 textual quantitative file saved in .txt format “<strong>[source]-keys.txt</strong>”: 6 files 2 excel quantitative files saved in .xlsx format “<strong>SentimentFB.xlsx”</strong> “<strong>topics_prevalence_and_clustering</strong><strong>.xlsx”</strong> <strong>Data set Documentation</strong> Abstract This data set contains the underlying data of the paper “’<strong>ENGAGING THE PEOPLE’: IS SOCIAL MEDIA COVERAGE OF EU POLICY ASSOCIATED WITH PUBLIC SUPPORT FOR EUROPEAN INTEGRATION?</strong>”, submitted to the Journal of European Public Policy (Print ISSN: 1350-1763 Online ISSN: 1466-4429) in 2021. Data openly available within this dataset are a subset of the two following data sets, which contains all the relevant data of Work Package 3 and Work Package 5 of PERCEIVE project: Data set:<strong> “PERCEIVE: WP3: Effectiveness of communication strategies of EU projects” </strong>https://doi.org/10.5281/zenodo.3371133 Data set:<strong> “PERCEIVE: WP5: The multiplicity of shared meanings of EU and Cohesion Regional and Urban Policy at different discursive levels” </strong>https://doi.org/10.5281/zenodo.3371174 For the paper we collected Facebook posts referred to EU CP policies. We don’t have the permission to share these data (as they are protected by copyright), but all the sources are described in Deliverable 5.2, which is public (see http://doi.org/10.6092/unibo/amsacta/5726 or http://doi.org/10.5281/zenodo.1318184). We analyzed the textual content of data to construct a database of discursive topics in Task5.4. Data set includes the results of topic modeling and of a sentiment analysis performed on the Facebook homepages of Local Management Authorities (LMA) of PERCEIVE case study regions. Content of the files: 1 sub-folder, named “<strong>A_Stopword</strong>”, which contains all the stopword lists used for performing Topic Modeling. These are 6 .txt files, one for each language: Austrian, Italian, Polish, Romanian, Spanish, Swedish (“<strong>stoplist_file_[nation].txt</strong>”). 1 sub-folder which contain the Topic Modeling results for Facebook profiles of the Local Managing Authorities for Austria, Italy, Poland, Romania, Spain, and Sweden (sub-folder “<strong>B_Facebook</strong>”, 12 .txt files). For each case, a file “<strong>[source]-keys.txt</strong>” lists the 100 most important words for each topic, while a file “<strong>[source]-composition.txt</strong>” details the topic composition of each textual source. These files were obtained through Mallet software[1]. File “<strong>SentimentFB.xlsx</strong>” contains data regarding the sentiment analysis for contents on Facebook homepages of Local Managing Authorities. The first column indicates the country, as well as row labels (see below). Columns 2-21 indicate the number id of the topics for each topic model (national level). The three rightmost columns of the file represent respectively a) the name of the lexicon used to detect sentiment orientation (i.e. “VADER”); c) the average sentiment score for positive, neutral and average words for each lexicon and each country; and c) the sentiment score across all topics in a country. File “<strong>topics_prevalence_and_clustering</strong><strong>.xlsx</strong>” contains data regarding the three clusters of topics analyzed in the paper. The first column represents the ID of each topic; the second column reports the cluster of each topic; the third and the fourth columns report the average prevalence of each topic (rows) in posts and comments, respectively. As these data refer to a regional case study, these columns refer the first region for each country; the sixth and the seventh columns report the average prevalence of each topic (rows) in posts and comments for the second region analyzed (only for those countries where we analyzed two regions); the eighth and ninth columns reports the average prevalence of topics and comments, respectively, for each country; and finally the tenth column reports the country to which data in the previous two columns are referred. [1] McCallum, Andrew Kachites. "MALLET: A Machine Learning for Language Toolkit."http://mallet.cs.umass.edu. 2002.
<strong>自述文件</strong> 数据集标题:<strong>“PERCEIVE——凝聚民心:欧盟政策的社交媒体报道是否与公众对欧洲一体化的支持度相关?”</strong> 数据集作者:<strong>维塔利亚诺·巴贝里奥(Vitaliano Barberio)</strong>(维也纳经济大学),开放研究者与贡献者标识符(Open Researcher and Contributor ID,ORCID)<strong>http://orcid.org/0000-0002-2615-5006</strong>;<strong>卢卡·帕雷斯基(Luca Pareschi)</strong>(罗马第二大学(Università di Roma Tor Vergata)),开放研究者与贡献者标识符(ORCID)<strong>http://orcid.org/0000-0002-4402-9329</strong>;数据集贡献者:<strong>伊内斯·库里奇(Ines Kuric)</strong>(维也纳经济大学);<strong>爱德华多·莫洛纳(Edoardo Mollona)</strong>(博洛尼亚大学),开放研究者与贡献者标识符(ORCID)<strong>http://orcid.org/0000-0001-9496-8618</strong>;<strong>马库斯·赫勒勒(Markus Höllerer)</strong>(维也纳经济大学),开放研究者与贡献者标识符(ORCID)<strong>http://orcid.org/0000-0003-2509-2696</strong>。数据集联系人:<strong>卢卡·帕雷斯基(Luca Pareschi)</strong>(罗马第二大学),ORCID<strong>http://orcid.org/0000-0002-4402-9329</strong>;邮箱:luca.pareschi@uniroma2.it。数据集许可:本数据集采用知识共享署名(Creative Commons Attribution,CC BY)4.0国际许可协议进行分发。出版年份:<strong>2021</strong>。项目信息:<strong>PERCEIVE项目(欧洲人对区域与凝聚政策的感知与评价及欧洲价值认同)</strong>,由欧盟地平线2020计划(Horizon 2020 Programme)资助,资助协议编号:<strong>693529</strong>;项目官网:<strong>https://www.perceiveproject.eu/</strong>。<strong>数据集内容</strong> 本数据集包含以下内容:1. 1份自述文件;2. 6份文本格式定性文件,保存为.txt格式,命名为“<strong>stoplist_file_[nation].txt</strong>”;3. 12份文本格式定量文件,保存为.txt格式,命名为“<strong>[source]-keys.txt</strong>”;4. 2份Excel格式定量文件,分别为“<strong>SentimentFB.xlsx</strong>”与“<strong>topics_prevalence_and_clustering.xlsx</strong>”。<strong>数据集文档</strong> 摘要:本数据集包含2021年提交给《欧洲公共政策期刊》(印刷版ISSN:1350-1763,网络版ISSN:1466-4429)的论文《“凝聚民心”:欧盟政策的社交媒体报道是否与公众对欧洲一体化的支持度相关?》的配套原始数据。本数据集公开可用的数据为以下两个数据集的子集,这两个数据集包含了PERCEIVE项目第3工作包(WP3)与第5工作包(WP5)的全部相关数据:1. 数据集《PERCEIVE: WP3: Effectiveness of communication strategies of EU projects》,DOI:10.5281/zenodo.3371133;2. 数据集《PERCEIVE: WP5: The multiplicity of shared meanings of EU and Cohesion Regional and Urban Policy at different discursive levels》,DOI:10.5281/zenodo.3371174。本论文针对欧盟凝聚政策(EU CP policies)收集了Facebook帖子数据。由于受版权保护,研究团队无权分享此类原始数据,但所有数据来源均已在公开可获取的可交付成果5.2(Deliverable 5.2)中详述(详见http://doi.org/10.6092/unibo/amsacta/5726 或 http://doi.org/10.5281/zenodo.1318184)。研究团队在任务5.4(Task5.4)中对文本内容进行分析,构建了话语主题数据库。本数据集包含针对PERCEIVE项目案例研究区域的地方管理当局(Local Management Authorities,LMA)Facebook主页数据所开展的主题建模与情感分析结果。文件内容详情如下:1. 名为“<strong>A_Stopword</strong>”的子文件夹,内含用于主题建模的停用词表,共6个.txt格式文件,分别对应奥地利语、意大利语、波兰语、罗马尼亚语、西班牙语、瑞典语,命名格式为“<strong>stoplist_file_[nation].txt</strong>”。2. 名为“<strong>B_Facebook</strong>”的子文件夹,内含针对奥地利、意大利、波兰、罗马尼亚、西班牙、瑞典地方管理当局Facebook主页的主题建模结果,共12个.txt格式文件。针对每个案例,“<strong>[source]-keys.txt</strong>”文件列出了每个主题的前100个核心词汇,而“<strong>[source]-composition.txt</strong>”文件则详细说明了每个文本源的主题构成。此类文件均通过MALLET机器学习语言工具包[1]生成。3. 文件“<strong>SentimentFB.xlsx</strong>”:包含针对地方管理当局Facebook主页内容的情感分析数据。第一列为国家名称,同时作为行标签。第2至21列为各主题模型(国家级)的主题编号。文件最右侧的三列分别为:a) 用于检测情感倾向的词典名称(即“VADER”);b) 各词典与各国家对应的正面、中性及平均词汇情感得分;c) 一国所有主题的平均情感得分。4. 文件“<strong>topics_prevalence_and_clustering.xlsx</strong>”:包含本论文所分析的三类主题簇数据。第一列为每个主题的ID;第二列为每个主题所属的簇;第三、第四列分别为各主题(行)在帖子与评论中的平均出现频率,由于此类数据针对区域案例研究,因此这两列对应每个国家的第一个区域;第六、第七列分别为第二个分析区域的帖子与评论中各主题的平均出现频率(仅针对分析了两个区域的国家);第八、第九列分别为每个国家的主题与评论平均出现频率;最后一列为前两列数据所对应的国家。[1] McCallum, Andrew Kachites. "MALLET: A Machine Learning for Language Toolkit."http://mallet.cs.umass.edu. 2002.



