CEIMaT2021 Dataset
收藏资源简介:
This package contains the dataset generated in the research published in the paper: "Almudena Sánchez Ruíz, Daniel Galan, Ángel García-Beltrán and Javier Rodríguez-Vidal. Detecting topics and polarity from Twitter a university faculty case" The dataset is available for research purpose. The CEIMaT2021 Dataset contains data from Twitter (http://www.twitter.com/), where anyone can express their own opinion by short messages shared publicly. Specifically, CEIMaT2021 Dataset includes tweets related with Escuela Técnica Superior de Ingenieros Industriales of Universidad Politécnica de Madrid (ETSII-UPM). To build the dataset, we automatically extracted all the existing tweets written in Spanish with reference to ETSII-UPM. To do this, we took into consideration four different groups: 1. Tweets published by the user @industrialesupm, official Twitter account of ETSII-UPM. 2. Tweets with mention to @industrialesupm, independent of who published it. 3. Tweets including the hashtag #etsii. 4. Tweets including the hashtag #industrialesupm. In total, we extracted 18971 tweets, but not all of them were strictly related to ETSII-UPM or written in Spanish. Therefore, we filtered them and delete the unnecessary tweets. After this process, we obtained a final set of 11014 tweets, which are collected in CEIMaT2021 Dataset. Three experts related with ETSII-UPM annotated each tweet depending on their polarity and topic, as: Polarity: POSITIVE, NEUTRAL, NEGATIVE. Topic: EVENTS, EXAMS, COMPUTING, TEACHING AND RESEARCH, INSTITUTION, SERVICES, INFRASTRUCTURE, OTHER.



