遇见数据集

Exploratory Twitter hashtag analysis of movie premieres in the USA

收藏
Figshare2024-02-07 更新2026-04-08 收录
官方服务:

资源简介:

This work is an exploratory, quantitative, and not experimental study with an inductive inference type and a longitudinal follow-up. It analyzes movie data and tweets published by users using the official Twitter hashtags of movie premieres the week before, the same week, and the week after each release date.The scope of the study is the collection of movies released in February 2022 in the USA, and the object of the study includes them and the tweets that refer to the film in the 3 closest weeks to their premiere dates. The tweets recollected were classified by the week they were published, so they are classified by a time dimension called timepoint. The week before the release date has been designated as timepoint 1, the week of the release date is timepoint 2, and the week immediately afterward is timepoint 3. Another dimension that has been considered is if the movie has domestic production or not, which means that if one of the countries of origin is the United States, the movie is designated as domestic.The chosen variables are organized in two data tables, one for the movies and one for the collected tweets.Variables related to the movies:id: Internal id of the moviename: Title of the moviehashtag: Official hashtag of the moviecountries: List of countries of the movie, separated by a semicolonmpaa: Film ratings system by the Motion Picture Association of America. It is a completely voluntary rating system and ratings have no legal standing. The currently rating systems include G (general audiences), PG (parental guidance suggested), PG-13 (parents strongly cautioned), R (restricted, under 17 requires accompanying parent or adult guardian) and NC-17 (no one 17 and under admitted)(<i>Film Ratings - Motion Picture Association</i>, n.d.)genres: List of genres of the movie, e.g., Action or Thriller, separated by a semicolonrelease_date: Release date of the movie in a format YYYY-MM-DDopening_grosses: Amount of USA dollars that the movie obtained on the opening date (the first week after the release date)opening_theaters: Amount of USA theaters that released the movie on the opening date (the first week after the release date)rating_avg: Average rating of the movieVariables related to the tweets:id: Internal id of the tweetstatus_id: Twitter id of the tweetmovie_id: Internal id of the movietimepoint: Week number related to the movie premiere that the tweet was published on. “1” is the week before the movie release, “2” is the week after the movie release” and “3” is the second week after the movie release.author_id: Twitter id of the author of the tweetcreated_at: Date and time of the tweet, with format “YYYY-MM-DD HH:MM:SS”quote_count: Number of the tweet’s quotesreply_count: Number of the tweet’s repliesretweet_count: Number of the tweet’s retweetslike_count: Number of the tweet’s likessentiment: Sentiment analysis of the tweet’s content with a range from -1 (negative) to 1 (positive)This dataset has contributed to the elaboration of the book chapters:Yeste, Víctor; Calduch-Losa, Ángeles (2022). Genre classification of movie releases in the USA: Exploring data with Twitter hashtags. In <i>Narrativas emergentes para la comunicación digital (</i>pp. 1012-1044). Dykinson, S. L.Yeste, Víctor; Calduch-Losa, Ángeles (2022). Exploratory Twitter hashtag analysis of movie premieres in the USA. In <i>Desafíos audiovisuales de la tecnología y los contenidos en la cultura digital</i> (pp. 169-187). McGraw-Hill Interamericana de España S.L.Yeste, Víctor; Calduch-Losa, Ángeles (2022). ANOVA to study movie premieres in the USA and online conversation on Twitter. The case of rating average using data from official Twitter hashtags. In <i>El mapa y la brújula. Navegando por las metodologías de investigación en comunicación</i> (pp. 151-168). Editorial Fragua.

本研究属于探索性、定量且非实验性的归纳推理类纵向追踪研究。研究分析了两类数据:2022年2月美国上映的全部电影数据,以及各电影上映日期前一周、当周及后一周期间,用户使用该电影首映官方Twitter话题标签发布的推文。研究对象包含上述电影,以及距其首映日期前后3周内提及该影片的推文。收集的推文按发布周进行分类,以时间维度“时间节点(timepoint)”划分:上映前一周记为时间节点1,上映当周为时间节点2,上映后一周为时间节点3。此外,研究还考量了影片是否为本土制作这一维度:若影片的原产国包含美国,则将其归类为本土影片。 研究选取的变量被整理为两个数据表,分别对应电影数据与收集的推文数据。 ### 电影相关变量 1. `id`:电影内部标识号 2. `name`:电影片名 3. `hashtag`:电影官方话题标签 4. `countries`:电影原产国列表,以分号分隔 5. `mpaa`:美国电影协会(Motion Picture Association of America)电影分级体系。该体系为完全自愿性分级,分级结果不具备法律效力,当前生效的分级包括:G级(大众级,所有年龄段均可观看)、PG级(建议家长陪同观看)、PG-13级(13岁以下观众需家长或成年监护人陪同)、R级(限制级,17岁以下观众需家长或成年监护人陪同)以及NC-17级(17岁及以下观众禁止观看)(<i>Film Ratings - Motion Picture Association</i>,无出版日期) 6. `genres`:电影类型列表,例如动作片、惊悚片,以分号分隔 7. `release_date`:电影上映日期,格式为`YYYY-MM-DD` 8. `opening_grosses`:影片上映首周(上映后第一周)在美国获得的票房收入,单位为美元 9. `opening_theaters`:影片上映首周在美国上映的影院数量 10. `rating_avg`:影片的平均评分 ### 推文相关变量 1. `id`:推文内部标识号 2. `status_id`:推文的Twitter官方标识号 3. `movie_id`:关联电影的内部标识号 4. `timepoint`:推文发布时对应的电影首映周节点:"1"代表电影上映前一周,"2"代表电影上映后第一周,"3"代表电影上映后第二周(注:原文此处与前文时间节点定义存在表述差异,前文将上映后一周记为时间节点3) 5. `author_id`:推文作者的Twitter官方标识号 6. `created_at`:推文发布的日期与时间,格式为`YYYY-MM-DD HH:MM:SS` 7. `quote_count`:该推文的引用数 8. `reply_count`:该推文的回复数 9. `retweet_count`:该推文的转发数 10. `like_count`:该推文的点赞数 11. `sentiment`:推文内容的情感分析得分,取值范围为-1(负面情感)至1(正面情感) 本数据集已用于以下图书章节的撰写: 1. Yeste, Víctor; Calduch-Losa, Ángeles (2022). 《美国上映影片的类型分类:基于Twitter话题标签的数据探索》. 载于<i>Narrativas emergentes para la comunicación digital</i>(《新兴叙事与数字传播》),第1012-1044页。Dykinson, S. L.出版。 2. Yeste, Víctor; Calduch-Losa, Ángeles (2022). 《美国电影首映式的Twitter话题标签探索性分析》. 载于<i>Desafíos audiovisuales de la tecnología y los contenidos en la cultura digital</i>(《数字文化中技术与内容的视听挑战》),第169-187页。McGraw-Hill Interamericana de España S.L.出版。 3. Yeste, Víctor; Calduch-Losa, Ángeles (2022). 《用于研究美国电影首映与Twitter线上对话的方差分析:以基于官方话题标签数据的平均评分为例》. 载于<i>El mapa y la brújula. Navegando por las metodologías de investigación en comunicación</i>(《地图与罗盘:导航传播学研究方法论》),第151-168页。Editorial Fragua出版。

提供机构:
Yeste, Víctor
创建时间:
2024-02-07
二维码
社区交流群
二维码
科研交流群
商业服务