Processed patent corpus by sector
收藏资源简介:
What the dataset contains For each patent, the database includes: Patent identifier and publication year Assignee / assigned company(ies) A cleaned document field compiled from title, abstract, and key elements of claims (processed text suitable for NLP/topic modeling) Patent-level embeddings computed from the processed document (used in the BERTopic-based pipeline) Additional derived fields used in downstream analyses (e.g., topic assignments / metadata needed for dynamic topic-window computations) What it enables The dataset is designed for researchers and practitioners who want to: build and inspect semantic topic structures of an industry’s patent corpus, track dynamic topic evolution across time windows, identify shift-driving / breakout topics and patents, and run company-level analyses (portfolio positioning, proximity to critical patents, incumbents vs entrants, network measures). Reproducibility materials A companion GitHub repository contains the full set of Jupyter notebooks (.ipynb) and utility scripts required to reproduce the main pipeline end-to-end, including data filtering/preparation, BERTopic training, dynamic-topic and predictivity analyses, and downstream company analytics: https://github.com/mingchennnn/TechSeis Intended use and citation This release is meant as a public benchmark and starting point for extending “technology seismograph” methods to other domains. If you use the dataset or code, please cite this Zenodo record and the GitHub repository above.



