遇见数据集

Global and U.S. Box Office Performance Data Cleaned

收藏
Zenodo2025-05-24 更新2026-05-26 收录
官方服务:

资源简介:

This cleaned dataset is a focused subset of the original Global and U.S. Box Office Performance Data, developed for academic purposes through a custom scraping pipeline combining Selenium and BeautifulSoup. The cleaned version exclusively includes records related to the “Brand” section from Box Office Mojo by IMDbPro. Specifically, it comprises: Aggregated financial and commercial statistics for each brand (e.g., Marvel Comics, Pixar, Disney, Legendary Pictures). Individual movie entries that were part of these brands, scraped separately by navigating into the specific brand pages and extracting detailed data per film. The data selection was performed by filtering the original dataset using the category column to retain only entries labeled as Brand, and by incorporating additional movie-level data scraped directly from within each brand's dedicated page. The dataset has been enriched with external, non-commercial IMDb data to supplement missing attributes and provide more context where necessary. After the cleaning and filtering process, the resulting dataset includes: 919 rows: representing both brand-level summaries and associated movies. 20 columns: including brand name, movie title (if applicable), release date, total gross, opening weekend, distributor, number of theaters, and more. Key cleaning steps documented in the accompanying report (box_office_mojo_cleaning.html) include: Parsing and standardizing date formats and numeric fields (e.g., gross revenues, theater counts). Renaming ambiguous column names for clarity and consistency. Removing or imputing invalid, null, or duplicate records. Adding derived features, such as is_limited_release or cleaned percentage values. Standardizing categorical variables (e.g., genres) and fixing inconsistent capitalizations. Ensuring uniformity in monetary units and number formatting (e.g., $ and commas removed). Dropping unused or sparsely populated fields to streamline the dataset. The cleaned dataset has 919 records (14,086) and retains 28 essential variables that enable time series analysis, market segmentation, and modeling of box office trends at multiple aggregation levels (daily, monthly, by brand, genre, etc.). This dataset was used to explore which variables influence a movie’s total gross revenue. Additionally, it served to analyze whether there are significant differences in gross sales based on genre, runtime, or brand affiliation. Important Note: This dataset was created exclusively for educational purposes and complies with ethical data collection practices: minimal server load, transparency in source attribution, and non-commercial use. It is distributed under the Creative Commons BY-NC-ND 4.0 license, with additional terms specified in the LICENSE file on the associated GitHub repository.

提供机构:
Zenodo
创建时间:
2025-05-24
二维码
社区交流群
二维码
科研交流群
商业服务