Machine learning and natural language processing on the patent corpus: Data, tools, and new measures
收藏资源简介:
The new disambiguation dataset contains an extended and improved work from its original paper "Machine learning and natural language processing on the patent corpus: Data, tools, and new measures". Schema is shown as follows: +++++++++++++++++++++++++++++++++++ Inventor/geography disambiguation, USPTO 1976-2018 Patent No (x) Inventor Name (y) Inventor Last Name (before the semicolon) Inventor First Name + Middle Name (after the semicolon) Original Sequence of (y) in (x) Disambiguated Inventor ID X,XXX,XXX-Y, where X,XXX,XXX is the very first patent of that inventor and Y is the order of appearance of that inventor in X,XXX,XXX City State Country County name IF IN USA FIPS 5-digit (state + city combined) IF IN USA FIPS 2-digit (state code) IF IN USA FIPS 3-digit (city code) IF IN USA Latitude disambiguated Longitude disambiguated Zipcode IF IN USA Metropolitan statistical area (MSA) IF IN USA +++++++++++++++++++++++++++++++++++ Assignee disambiguation Name Assignee Sequence Geography Disambiguated assignee in pdpass as an unique identifier +++++++++++++++++++++++++++++++++++ CPC CPC Sequence Level-1 Level-2 Level-3 Level-4 Level-5



