遇见数据集

Linking Historical U.S. Patent Inventors to the Decennial Census, 1840–1950

收藏
Zenodo2026-08-11 更新2026-08-13 收录
官方服务:

资源简介:

This repository contains a dataset linking inventors named on USPTO patents published between 1840 and 1950 to individuals enumerated in the U.S. full-count decennial censuses of 1850–1940 (the 1890 census was destroyed). The dataset comprises 3,948,749 patent-inventor–census links covering 2,254,386 distinct patent-inventors (91% of the ~2.47 million U.S.-resident patent-inventor records that entered the matching) with up to three census waves per inventor. Links were produced by a supervised record-linkage pipeline: candidate census records are retrieved by blocking on gender, a ±10-year age window, and geography using four complementary fuzzy name-matching algorithms; thirteen rule-based filters remove implausible candidates; and a random-forest classifier, trained on 6,631 hand-verified patent–census matches augmented with answer-absent negative groups including provably deceased inventors, selects the most probable census record per inventor and wave (held-out precision 0.96, recall 0.93). A post-match step uses the Census Tree Project (CTP) crosswalks to flag cross-wave inconsistencies and to supplement 97,266 additional links, without altering any model prediction. The dataset is released without an acceptance threshold: every machine-learning link carries the classifier's predicted match probability (probability), so users choose their own precision–coverage trade-off (e.g., about 1.1 million patent-inventors are linked at probability ≥ 0.4). Each row is one matched patent-inventor–census pair, keyed on (PatInvID, histid, census_year). The IPUMS person identifier histid allows rejoining the full IPUMS USA census microdata (join case-insensitively, as histid letter case varies across waves). from_ctp marks CTP-supplemented rows (no model probability), ctp_flag marks multi-wave inventors whose matched records the CTP does not connect, and histid_ctp records the CTP-implied identifier alongside each prediction. Files File Format Contents Primary Dataset patent_inventor_census_links_1850-1940_v1.zippatent_inventor_census_links_1850-1940_v1.parquet CSV (zip)Parquet The linked dataset described in this Data Descriptor: one row per matched patent-inventor–census link. Public Patent Inputs USPTO_patent_inventors_1840_1950.zip TSV (zip) Patent-inventor table (the patent side of the linkage): one row per patent–inventor pair. USPTO_patent_titles_1840_1950.zip TSV (zip) Patent titles (pubnumber, title); input to the relatedness features. USPTO_patent_classes_1840_1950.zip TSV (zip) Patent technology classes (USPC class, NBER category/subcategory); input to the relatedness features. Ground Truth and Validation ground_truth.csv CSV Confirmed patent–census matches used to train and evaluate the classifier. deceased_inventors.csv CSV Confirmed deceased-inventor records used as hard negatives in training. validation_sample_deceased.csv CSV Executor/household-relative sample for the out-of-sample precision check. Machine Learning Tables (Reproducibility) training_nonames.parquet Parquet Name-free labeled training table that reproduces the classifier. matched_candidates_ml.parquet Parquet Full candidate-level feature table (all candidates) used for scoring. Reference Lookups occupation_dict.json JSON Occupation-string dictionary used in feature construction. OCC1950.xlsx, IND1950.xlsx XLSX IPUMS 1950 occupation and industry classification code lists. uspc_class_names.xlsx XLSX USPC technology-class names (USPTO); input to the relatedness features. nber_category.txt TXT NBER technology category and subcategory names; input to the relatedness features. Note: The files training_nonames.parquet and matched_candidates_ml.parquet include fields derived from the restricted IPUMS full-count microdata (occupation strings/codes and histid) and are subject to the IPUMS USA terms of use. The full pipeline code is openly available under the MIT License at https://github.com/sbreschi/uspto-inventors-census-linked-1840-1950 and permanently archived at https://doi.org/10.5281/zenodo.21265096. The dataset is described in detail in the accompanying Data Descriptor, "A Novel Dataset for Historical Innovation Studies: Linking USPTO Patents and U.S. Census Data from 1840 to 1950" (submitted to Scientific Data). Users of the dataset are asked to cite the Data Descriptor and IPUMS USA (Ruggles et al., 2024, https://doi.org/10.18128/D014.V4.0R). Licensing and Data Use Code (MIT License) The Python pipeline code is released under the MIT License.See LICENSE file for details. Linked Dataset (CC-BY 4.0) The published linked dataset (patent_inventor_census_links_1850-1940_v1.parquet) and supplementary files are released under Creative Commons Attribution 4.0 International. You are free to: Share, copy, and redistribute the data Adapt, remix, and build upon the dataset Use for commercial purposes You must: Give appropriate credit to the authors Indicate if changes were made IPUMS-Derived Data Census-derived fields in the dataset remain subject to IPUMS USA Data Use Agreement. To obtain full census variables not included in this deposit, register with IPUMS USA and download directly from https://usa.ipums.org/ IPUMS USA citation:Ruggles, S., Flood, S., Sobek, M., et al. (2023). IPUMS USA: Version 14.0 [dataset]. Minneapolis, MN: IPUMS. https://doi.org/10.18128/D010.V14.0

提供机构:
Zenodo
创建时间:
2026-08-11
二维码
社区交流群
二维码
科研交流群
商业服务