A Dataset of 735,000 Migration Records (Bremen, 1830-1939)
收藏资源简介:
Documentation This dataset provides structured, enriched, and normalised historical passenger list data from the German North Sea port of Bremen. It comprises 735,545 individual records spanning the years 1830 to 1939. The dataset is derived from the transcription efforts of the Bremen Society for Family Research (MAUS) and the Bremen Chamber of Commerce [1]. This version has been post-processed to enhance its analytical value for historical migration research, social history, and demography. Key enhancements Normalization: Standardised versions of primary dataset values. Cleaning: Correction of orthographic variants and structural inconsistencies. Enrichment: Georeferencing (Lat/Lon) for ports of departure and arrival. Classification: Occupations mapped to HISCO codes, KldB2010 sectors, and training levels. Technical Specifications Filename: CHUP_Passengerlists.csv File Format: CSV (Delimiter: "@") File Size: ≈ 259.6 MB Encoding: UTF-8 Dimensions: 735,545 Rows × 62 Columns Language: Mostly German Methodology The processing workflow utilised string matching algorithms, regular expressions, and manualverification: Extraction: Unique values were extracted from the primary MAUS dataset. Normalization: Orthographic errors and variants were consolidated (e.g., 145 travel classvariants reduced to 16 standardised classes). Enrichment: Geography: Ports were geocoded using the OpenStreetMap Nominatim API. Occupations: Most frequent job titles were manually classified into sectors and training levels [2] and mapped to HISCO codes [3] where possible. Data Dictionary The dataset includes both the original transcription fields (derived from the primary MAUS database) and processed fields (indicated by the suffix _cleaned). A. Metadata Column Name Data Type Description id Integer Unique identifier for the data entry. archive_id String Archival identifier linking to the physical list. notes String Remarks or notes on duplicate passengers. sort_da String Sorting helper column for dates. pfd String Holds dates of departure for certain voyages. B. Passenger Data Column Name Data Type Description last_name String Last name of passenger. first_name String First name of passenger. gender String Gender (Original). gender_cleaned String Gender (Standardised: Male 53.6%, Female 46.4%). age String Age of Passenger (Original). age_cleaned Integer Transformed birth dates to age in years (Mean age: 31.1). marital_status String Marital status of passenger (Original). marital_status_cleaned String Marital status (Standardised). nationality String Citizenship (Original). nationality_cleaned String Citizenship (Standardised). previous_residence String City/Town of previous residence (Original). state_or_province String Region of previous residence (Original). state_or_province_cleaned String Region of previous residence (Standardised). occupation String Occupational title (Original). occupation_cleaned String Occupational title (Standardised). occupation_hisco_nr Float HISCO classification code. occupation_sector Integer KldB2010 Sector (0–9) classification. occupation_training_level Integer KldB2010 Training Level (1–4) classification. religion String Religion of passenger (Original). religion_cleaned String Religion of passenger (Standardised). ethnicity String Ethnicity of passenger (Original). ethnicity_cleaned String Ethnicity of passenger (Standardised). nr String Passenger number on the list. emigrant String Passenger emigrant status. amount_of_money String Cash carried by passenger. literacy String Literacy status of passenger. relative String Relatives of passenger. ticket String Ticket details. who_paid String Payer information. previous_US_stay String Details on previous US residency. duration_of_stay String Intended duration of stay (Original). duration_of_stay_cleaned String Intended duration of stay (Standardized). place_of_stay String Place of residence. reference_person String Contact person. C. Voyage Data Column Name Data Type Description date_of_departure String Departure date (Original). date_of_departure_cleaned String Standardised departure date (ISO-8601: YYYY-MM-DD). ship String Name of ship (Original). ship_cleaned String Name of ship (Standardised). agent String Shipping company (Original). agent_cleaned String Shipping company (Standardised). travel_class String Travel class (Original). travel_class_cleaned String Travel class (S captain String Captain of ship. port_of_departure String Port of departure (Original). port_of_departure_cleaned String Port of departure (Standardised). port_of_departure_country String Country of port of departure. port_of_departure_LAT/_LON Float Geocoordinates (WGS84) of departure port city. port_of_arrival String Arrival port (Original). port_of_arrival_cleaned String Arrival port (Standardised). port_of_arrival_country String Country of arrival port. port_of_arrival_LAT/_LON Float Geocoordinates (WGS84) of arrival port city. port_of_arrival_US_state String US-state of arrival port. emigration_destination String General destination stated by passenger. US_state String State of destination in the US (Original). US_state_cleaned String State of destination in the US (Standardised). Limitations Missing Years: Passenger lists from 1875 to 1907 were destroyed due to historical storagepolicies and are not included in this dataset. In total, the dataset holds just a few entriesfrom the 19th century. Interwar Gaps: Records for the years 1924 and 1929 are notably incomplete due toarchival loss during WWII. Data Sparsity: While core fields are nearly 100% complete, socio-economic variables likeethnicity, religion, or amount_of_money have higher rates of missing data. Unprocessed Data: Some columns were not standardised, either because their cleaningwould not have enhanced the analytical value of the column, or because the amount ofunique values would have made an attempt of normalising unfeasible.



