Replication package: Random Audits and Administrative Exit from the Trademark Register — Causal Evidence from the USPTO's Post-Registration Proof-of-Use Program
收藏资源简介:
Replication code and derived data for the paper "Random Audits and Administrative Exit from the Trademark Register: Causal Evidence from the USPTO's Post-Registration Proof-of-Use Program." The paper estimates the causal effect of the U.S. Patent and Trademark Office's post-registration proof-of-use audit on the fate of an audited trademark registration. Audited and unaudited maintenance filings are compared within 737 exact cells built from pretreatment administrative covariates. In the common-support sample of 15,897 filings, receipt of a first audit office action raises maintenance-related full cancellation within three years by 15.54 percentage points against a 1.36 percent control mean. CONTENTS code_and_metadata.tar.gz (140 KB) — 56 analysis scripts, the manuscript assembly script, a 283-check consistency audit, and the USPTO event-code dictionary used to define treatment and outcomes. results.tar.gz (10.4 MB) — every estimate reported in the paper, as CSV, JSON and parquet, together with the pre-committed decision-rule files and their SHA-256 digests. processed_analysis_data.tar.gz (200 MB) — the derived analysis samples: the maintenance-filing risk set, the common-support cohort, owner-portfolio panels, and the frozen verification samples. REDACTION_LOG.json — machine-readable record of every file and column altered before deposit. CHECKSUMS.txt — SHA-256 of the three archives. PERSONAL DATA REMOVED BEFORE DEPOSIT The underlying USPTO records are public, but a bulk, permanently indexed republication of personal contact details is a different act from an individual lookup on the agency's website. Two categories were therefore removed. First, 11,085 retrieved TSDR documents are not included. The filing XML carries the correspondent block — phone, fax and email — for 5,423 distinct people, mostly registrants and their attorneys, and 5,059 of the retrieved HTML pages embed a static USPTO authentication token that has no research value. These files are fetch artifacts rather than evidence: the manifests record which serial numbers were retrieved and when, and TSDR remains publicly searchable, so the verification samples can be reconstructed. Second, owner identity columns were dropped from 20 parquet files: owner name, alternate name, the "composed of" text that names natural persons, street address lines, city, postal code, normalised name and location, attorney name, and USPTO's internal owner identifier. The analysis needs owner linkage — for owner-clustered inference and portfolio construction — but not owner identity, so the linkage keys were rewritten as a salted SHA-256 hash. The mapping is injective, so grouping is unchanged: every pair of rows that shared an owner still shares one, and all 12,935 owner clusters reported in the paper are present. Country, state, entity type and portfolio size are retained. Verified before release: the three archives contain zero email addresses, zero authentication tokens, zero cleartext owner keys and zero street addresses. This deposit is not anonymised data and is not presented as such. Trademark serial numbers are retained because they are the unit of analysis; without them nothing here can be replicated or linked to the public source data, and a serial number entered into TSDR returns the current owner and correspondence address. What the steps above remove is a ready-made bulk compilation of contact details, which is the form in which such data is most easily scraped and reused. Users remain bound by the terms under which USPTO publishes the underlying records. NOT INCLUDED The raw source archives, about 7.2 GB, are not redistributed. They are public and are cited in the paper: the USPTO Trademark Case Files Dataset (2022 and 2023 annual releases); an independent analysis-ready reconstruction of the official bulk XML current through 6 July 2026; and the USPTO Trademark Status and Document Retrieval system, retrieved without an API key. The scripts rebuild every derived file from those sources. Published papers cited by the manuscript are not included, for copyright reasons. They are identified in the paper's reference list. PRE-COMMITTED DECISION RULES The paper distinguishes prospective design choices from post-outcome reporting choices. Files ending in _lock_2026-07-30.json and _decision_rules_2026-07-30.json, with the matching .sha256 files, record decision rules fixed and hashed before the corresponding outcomes were read. A hash shows that a file was not altered after creation; it does not turn a post-discovery specification into a prospective preregistration. Appendix D of the paper gives the chronology and states which results are exploratory. REPRODUCTION Appendix G of the paper maps each reported quantity to the script that produces it. Running scripts/audit_manuscript_consistency.py re-derives 131 numeric table rows from the result files and checks them against the manuscript text. Analyses that consume the retrieved TSDR documents must first re-fetch them using the manifests and the fetch scripts. USE OF AI TOOLS Anthropic Claude and OpenAI Codex were used to assist with data processing, code implementation and debugging, literature search, figure production and manuscript revision. Two uses are methodological rather than preparatory and are documented in the paper itself: AI-assisted contextual coding of ambiguous goods-and-services text, and two blinded independent AI-assisted recodings used to validate those labels. Neither is presented as human double coding. No AI tool selected the estimand, the specification, the sample restrictions or the reporting hierarchy. The author reviewed all outputs and takes full responsibility for the contents.



