遇见数据集

macpaw-research/mac-app-store-apps-metadata

收藏
Hugging Face2025-04-10 更新2026-01-03 收录
官方服务:

资源简介:

--- license: mit tags: - Software Analysis pretty_name: Mac App Store Applications Metadata size_categories: - 10K<n<100K language: - en - de configs: - config_name: metadata_UA data_files: "metadata-UA.csv" - config_name: metadata_US data_files: "metadata-US.csv" - config_name: metadata_DE data_files: "metadata-DE.csv" --- # Dataset Card for Macappstore Applications Metadata <!-- Provide a quick summary of the dataset. --> Mac App Store Applications Metadata sourced by the public API. - **Curated by:** [MacPaw Way Ltd.](https://huggingface.co/MacPaw) <!---- **Funded by [optional]:** [More Information Needed] --> <!--- **Shared by [optional]:** [MacPaw Inc.](https://huggingface.co/MacPaw) --> - **Language(s) (NLP):** Mostly EN, DE - **License:** MIT ## Dataset Details This data aims to cover our internal company research needs and start collecting and sharing the macOS app dataset since we have yet to find a suitable existing one. Full application metadata was sourced by the public iTunes search API for the US, Germany, and Ukraine between December 2023 and January 2024. The data is organized into 43 columns ranging from such essential details as app names, descriptions, and genres, along with release and version information, user ratings, and asset links. <!--For the convenience of further analysis, we present the data as three separate datasets: metadata, release notes, and descriptions. ### Dataset Description <!-- Provide a longer summary of what this dataset is. --> <!-- ### Dataset Sources [optional] <!-- Provide the basic links for the dataset. --> <!-- - **Repository:** [More Information Needed] - **Paper [optional]:** [More Information Needed] - **Demo [optional]:** [More Information Needed] <!-- ## Uses <!-- Address questions around how the dataset is intended to be used. --> <!-- ### Direct Use <!-- This section describes suitable use cases for the dataset. --> <!-- [More Information Needed] ### Out-of-Scope Use <!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. --> <!-- [More Information Needed] --> <!--## Dataset Structure <!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. --> <!--For the convenience of further analysis, we present the data as three separate datasets. --> <!--## Dataset Creation <!-- ### Curation Rationale <!-- Motivation for the creation of this dataset. --> <!-- [More Information Needed] ### Source Data <!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). --> ## Data Collection and Processing <!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. --> The data was fetched by querying the iTunes Search API with the request: `https://itunes.apple.com/search?term={term} &country={country}&entity={entity}&genreId=genre &limit={limit}&offset={offset}},` where the `term` parameter was selected in a way that maximized app types and skipped games; `country` and `offset` had to be specified to avoid the API limitations. The data is provided in its original form without any additional cleaning or transformations. It contains over 87,000 samples and is grouped by the country API and is presented in the corresponding CSV files: `metadata-US.csv`, `metadata-DE.csv`, and `metadata-UA.csv`. <!-- #### Who are the source data producers? <!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. --> <!-- [More Information Needed] ### Annotations [optional] <!-- If the dataset contains annotations which are not part of the initial data collection, use this section to describe them. --> <!-- #### Annotation process <!-- This section describes the annotation process such as annotation tools used in the process, the amount of data annotated, annotation guidelines provided to the annotators, interannotator statistics, annotation validation, etc. --> <!--N/A[More Information Needed] #### Who are the annotators? <!-- This section describes the people or systems who created the annotations. --> <!--[More Information Needed] #### Personal and Sensitive Information <!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. --> <!--[More Information Needed] ## Bias, Risks, and Limitations <!-- This section is meant to convey both technical and sociotechnical limitations. --> <!--[More Information Needed] ### Recommendations <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. --> <!--Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations. --> <!--## Citation <!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. --> ## Links - [MacPaw Research](https://research.macpaw.com/) - [Mac App Store Applications Release Notes dataset](https://huggingface.co/datasets/MacPaw/macappstore-apps-release-notes) - [Mac App Store Applications Descriptions dataset](https://huggingface.co/datasets/MacPaw/macappstore-apps-descriptions) <!--- **BibTeX:** [More Information Needed] **APA:** [More Information Needed] ## Glossary [optional] <!-- If relevant, include terms and calculations in this section that can help readers understand the dataset or dataset card. --> <!---[More Information Needed] ## More Information [optional] [More Information Needed] ## Dataset Card Authors [optional] [More Information Needed] --> ## Dataset Card Contact Feel free to reach out tech-research@macpaw.com if you have any questions or need further information about the dataset!

license: MIT许可证 tags: - 软件分析 pretty_name: Mac App Store 应用元数据 size_categories: - 10K<n<100K language: - 英语 - 德语 configs: - config_name: metadata_UA data_files: "metadata-UA.csv" - config_name: metadata_US data_files: "metadata-US.csv" - config_name: metadata_DE data_files: "metadata-DE.csv" # Mac App Store 应用元数据数据集卡片 Mac App Store 应用元数据通过公开API获取。 - **整理方:** [MacPaw Way Ltd.](https://huggingface.co/MacPaw) - **语言(自然语言处理):** 以英语、德语为主 - **许可证:** MIT ## 数据集详情 本数据集旨在满足公司内部研究需求,同时启动macOS应用数据集的收集与共享工作——因目前尚未找到合适的现有同类数据集。完整的应用元数据于2023年12月至2024年1月期间,通过公开的iTunes搜索API(iTunes Search API),针对美国、德国和乌克兰三个地区采集得到。 数据集共包含43列字段,涵盖应用名称、描述、类别等核心详情,同时包含发布与版本信息、用户评分以及资源链接等内容。 ## 数据采集与处理 本数据集通过调用iTunes搜索API获取,请求格式为: `https://itunes.apple.com/search?term={term} &country={country}&entity={entity}&genreId=genre &limit={limit}&offset={offset}}` 其中`term`参数的选取原则为最大化覆盖应用类型并排除游戏;`country`与`offset`参数需按要求指定以规避API调用限制。 本数据集保留原始采集格式,未进行额外清洗或转换处理,总计包含超过87,000条样本,按API调用的地区分组,分别存储于`metadata-US.csv`、`metadata-DE.csv`与`metadata-UA.csv`三个CSV文件中。 ## 相关链接 - [MacPaw 研究平台](https://research.macpaw.com/) - [Mac App Store 应用发布说明数据集](https://huggingface.co/datasets/MacPaw/macappstore-apps-release-notes) - [Mac App Store 应用描述数据集](https://huggingface.co/datasets/MacPaw/macappstore-apps-descriptions) ## 数据集卡片联系方式 如有任何关于本数据集的疑问或需要进一步信息,请发送邮件至 tech-research@macpaw.com 联系我们!

提供机构:
macpaw-research
二维码
社区交流群
二维码
科研交流群
商业服务