DeepCube: Post-processing dataset of social media data
收藏资源简介:
Researcher(s): Alexandros Mokas, Eleni Kamateri Supervisor: Ioannis Tsampoulatidis This dataset contains the post-processing of the social media data collected for two different use cases during the first two years of the Deepcube project. More specifically, it contains two sub-datasets, including: The UC2 dataset containing the post-processing of the Twitter data collected for the DeepCube use case (UC2) dealing with the climate induced migration in Africa. This dataset contains in total 5,695,253 social media posts collected from the Twitter platform, based on the initial version of search criteria relevant to UC2 (defined by Universitat De Valencia), focused on the regions of Ethiopia and Somalia and started from 26 June, 2021 till March, 2023. The UC5 dataset containing the post-processing of the Twitter and Instagram data collected for the DeepCube use case (UC5) related to the sustainable and environmentally-friendly tourism. This dataset contains in total 58,143 social media posts collected from the Twitter and Instagram platform (12,881 collected from Twitter and 45,262 collected from Instagram), based on the initial version of search criteria relevant to UC5 (defined by MURMURATION SAS), focused on the regions of Brasil and started from 26 June, 2021 till March, 2023. For every social media post retrieved from Twitter and Instagram, a preprocessing step was performed. This involved a three-step analysis of each post using the appropriate web service. First, the location of the post was automatically extracted from the text using a location extraction service. Second, the images included in the post were analyzed using a concept extraction service, which identified and provided the top ten concepts that best described the image. These concepts included items such as "person," "building," "drought," "sun," and so on. Finally, the sentiment expressed in the post's text was determined by using a sentiment analysis service. The sentiment was classified as either positive, negative, or neutral. After the social media posts were preprocessed, they were visualized using the Social Media Web Application. This intuitive, user-friendly online application was designed for both expert and non-expert users and offers a web-based user interface for filtering and visualizing the collected social media data. The application provides various filtering options, an interactive map, a timeline, and a collection of graphs to help users analyze the data. Moreover, this application provides users with the option to download aggregated data for specific periods by applying filters and clicking the "Download Posts" button. This feature allows users to easily extract and analyze social media data outside of the web application, providing greater flexibility and control over data analysis.<br> <br> The dataset is provided by INFALIA. INFALIA, being a spin-off of the CERTH institute and a partner of a research EU project, releases this dataset containing Tweets IDs and post pre-processing data for the sole purpose of enabling the validation of the research conducted within the DeepCube. Moreover, Twitter Content provided in this dataset to third parties remains subject to the Twitter Policy, and those third parties must agree to the Twitter Terms of Service, Privacy Policy, Developer Agreement, and Developer Policy (https://developer.twitter.com/en/developer-terms) before receiving this download. License: Creative Commons Attribution 4.0 International
研究人员:亚历山德罗斯·莫卡斯(Alexandros Mokas)、埃莱妮·卡马泰里(Eleni Kamateri);指导教师:约阿尼斯·詹普拉蒂迪斯(Ioannis Tsampoulatidis) 本数据集涵盖DeepCube项目前两年期间,为两个不同用例收集的社交媒体数据后处理成果。具体而言,本数据集包含两个子数据集: UC2子数据集:针对DeepCube用例UC2(聚焦非洲气候引发的移民问题)收集的Twitter数据后处理成果。该数据集共包含5,695,253条来自Twitter平台的社交媒体帖文,基于瓦伦西亚大学制定的UC2初始搜索条件,采集范围覆盖埃塞俄比亚与索马里地区,采集时间为2021年6月26日至2023年3月。 UC5子数据集:针对DeepCube用例UC5(聚焦可持续环保旅游)收集的Twitter与Instagram数据后处理成果。该数据集共包含58,143条来自Twitter与Instagram平台的社交媒体帖文(其中Twitter端12,881条、Instagram端45,262条),基于MURMURATION SAS制定的UC5初始搜索条件,采集范围覆盖巴西地区,采集时间为2021年6月26日至2023年3月。 针对每一条从Twitter与Instagram获取的社交媒体帖文,本数据集均执行了预处理步骤:通过三类专用Web服务对单条帖文开展三步分析:首先,利用位置提取服务从帖文文本中自动提取发布位置;其次,通过概念提取服务分析帖文内嵌图片,识别并返回最能描述该图片的十大核心概念,例如“人物”“建筑”“干旱”“阳光”等;最后,借助情感分析服务确定帖文文本所表达的情感倾向,将情感划分为正面、负面与中性三类。 完成预处理的社交媒体帖文可通过社交媒体Web应用进行可视化展示。该应用界面直观易用,面向专业与非专业用户群体,提供基于网页的用户界面以实现采集到的社交媒体数据的筛选与可视化功能,支持多种筛选选项、交互式地图、时间线以及各类图表,助力用户开展数据分析。此外,用户可通过该应用筛选特定时段并点击“下载帖文”按钮,获取对应时段的聚合数据。该功能允许用户脱离Web应用自主提取与分析社交媒体数据,为数据分析提供了更高的灵活性与控制权。 本数据集由INFALIA发布。INFALIA作为CERTH研究所(希腊研究与技术中心)的衍生企业,同时也是欧盟DeepCube研究项目的合作方,本次发布该包含Twitter帖文ID与帖文预处理数据的数据集,仅用于支持DeepCube项目内相关研究的验证工作。此外,本数据集所包含的Twitter内容,其第三方使用者仍需遵守Twitter相关规定,在获取该数据集前,第三方必须同意Twitter服务条款、隐私政策、开发者协议以及开发者政策(https://developer.twitter.com/en/developer-terms)。 许可证:知识共享署名4.0国际许可(Creative Commons Attribution 4.0 International)



