遇见数据集

Processed patent corpus by sector

收藏
Zenodo2026-03-02 更新2026-05-26 收录
官方服务:

资源简介:

What the dataset contains For each patent, the database includes: Patent identifier and publication year Assignee / assigned company(ies) A cleaned document field compiled from title, abstract, and key elements of claims (processed text suitable for NLP/topic modeling) Patent-level embeddings computed from the processed document (used in the BERTopic-based pipeline) Additional derived fields used in downstream analyses (e.g., topic assignments / metadata needed for dynamic topic-window computations) What it enables The dataset is designed for researchers and practitioners who want to: build and inspect semantic topic structures of an industry’s patent corpus, track dynamic topic evolution across time windows, identify shift-driving / breakout topics and patents, and run company-level analyses (portfolio positioning, proximity to critical patents, incumbents vs entrants, network measures). Reproducibility materials A companion GitHub repository contains the full set of Jupyter notebooks (.ipynb) and utility scripts required to reproduce the main pipeline end-to-end, including data filtering/preparation, BERTopic training, dynamic-topic and predictivity analyses, and downstream company analytics: https://github.com/mingchennnn/TechSeis Intended use and citation This release is meant as a public benchmark and starting point for extending “technology seismograph” methods to other domains. If you use the dataset or code, please cite this Zenodo record and the GitHub repository above.

提供机构:
Zenodo
创建时间:
2026-03-02
二维码
社区交流群
二维码
科研交流群
商业服务