MAPE: A Dataset of Correspondence from the Portuguese Empire
收藏资源简介:
We present the MAPE dataset: Mapping the Atlantic Portuguese Empire: a large-scale historical resource curated from archival material. The dataset is made available in different versions together with its detailed description: MAPE Dataset: Raw Archival Materials (Version 1). Observation: treat it as a reference dataset. MAPE Dataset: Bilingual Version (Portuguese-English) (version 2). Observations: A bilingual layer was added, but no records were removed. MAPE Dataset: Bilingual Version with Senders and Recipients (version 3). Observation: When adding the Sender Name and Receiver Name fields, the pipeline lost some records. This looks like an export/parsing error MAPE Dataset: Bilingual Version with Senders and Recipients with assigned topics (version 4) - some problems detected (please see below) MAPE Dataset: Bilingual Version with Senders and Recipients with assigned topics (version 5) - detected issues have been resolved; a detailed report MAPE_Dataset_Report_2026-04-17.pdf is provided together with the dataset. The MAPE dataset comprises 182,491 historical correspondence records from the Arquivo Histórico Ultramarino de Lisboa (Portuguese Overseas Archives of Lisbon, hereafter AHU), in particular from the collection of the Conselho Ultramarino (Overseas Council), covering the period from 1581 to 1859. The AHU holds an extensive archive of correspondence covering the administrative, diplomatic and commercial activities of the Portuguese Empire. The “Conselho Ultramarino”, created in 1642, represents formal bureaucratic communication between Lisbon and its overseas dominions and covers topics such as colonial administration, trade, diplomacy and social developments. Originally, these materials were only available as unstructured PDF files, which posed a major challenge for data analysis and large-scale retrieval. These PDF documents contained not only the core correspondence registers, but also a variety of non-essential metadata, such as cataloging details, pagination markers, section headings, and record summaries. The mixing of primary records with additional metadata hindered effective content analysis, searching and visualization. To overcome these challenges, we converted the PDFs into a structured format (CSV) that isolates the main data elements, improving searchability, navigation and analytical potential. This restructuring allows researchers to work directly with the primary correspondence records without the noise of the surrounding archival metadata. The data for this study were obtained from the Arquivo Histórico Ultramarino, where historical documents were preserved primarily in unstructured formats, predominantly as PDF files. These documents were publicly available for free download at https://actd.iict.pt/collection/actd:CU. The collections were originally divided into large sections such as Portugal, Africa, Brazil, etc. Within each section there are further subdivisions corresponding to the different colonies of the Portuguese empire at that time. The correspondence register of each colony is stored in individual PDF files, which are organized chronologically. However, these files also contain extraneous metadata such as headings, page numbers, cataloging details and document summaries, which make it difficult to extract the relevant content. As a rule, a header is followed by a short summary of the correspondence, with each document being provided with details of the archiving source. Our research focuses primarily on specific collections within the AHU, arranged chronologically and geographically, which include the following: Africa: The Angola Collection (Série Angola), whose cataloging was financially supported by the Portuguese Fundação para a Ciência e Tecnologia as part of the project África Atlântica: da documentação ao conhecimento, sécs. XVII-XIX (Atlantic Africa: from documentation to knowledge, seventeenth to nineteenth centuries). The Cabo Verde and Guinea Collection (Série Cabo Verde, Série Guiné), which was cataloged as part of two separate projects: the aforementioned África Atlântica and the Resgate do acervo histórico de Cabo Verde em Portugal (Rescue the historical collection of Cape Verde in Portugal) funded by Camões, Instituto da Cooperação e da Língua (ICL). The São Tomé Collection (Série S. Tomé e Príncipe), also cataloged within the África Atlântica project. Brazil: The “Barão do Rio Branco” — Historical Documentation Rescue Project known as Projeto Resgate (Bertoletti et al. 2022; Boschi 2018) includes 26 catalogues of documents referring to Brazilian regions, cataloged at different times and by different researchers. The Projeto Resgate collection is currently managed by the National Library of Rio de Janeiro in Brazil, but is housed in the AHU. Portugal: Madeira-CA and Madeira. Rio da Prata: Nova Colónia do Sacramento, Montevideu, Buenos Aires, Paraguai Oriente Macau Timor MAPE Dataset: Raw Archival Materials (version 1) Column Type Description doc_id Integer Unique identifier for each record. doc_source String Archival origin (e.g. ALAGOAS, BAHIA, Cabo Verde). doc_box String Physical box code within the archive (e.g. Cx.1). doc_number String Document number within the box (zero-padded, e.g. 00001). doc_type String Type of register (e.g. INFORMACAO, CONSULTA, CARTA, PROPOSTA, REQUERIMENTO, PARECER). year Integer Four-digit year of the correspondence (e.g. 1690). month Integer Month of the register (1–12). Blank if not recorded in the original. day Integer Day of the month (1–31). Blank if not recorded. reference_code String Integer Unique identifier for each record. doc_link URL Direct link to the AHU catalog entry for the document. Doc_Text String Original Portuguese summary of the correspondence, as transcribed from the archival register. The MAPE dataset is provided as a single CSV file the repository root. It consolidates all correspondence registers extracted from the AHU PDFs into a uniform tabular structure. 2. MAPE Dataset: Bilingual Version (Portuguese-English) (version 2) 1. v1 → v2: the scope of the material did not changeThe change was structural, not quantitative. A bilingual layer was added, but no documents were removed. Multilingual Adaptation of Consolidated Data Files The consolidated dataset originally contained correspondence in Portuguese, which was a significant barrier for a global audience. To overcome this limitation, we translated the original content into English using Google Gemini 1.5 Flash, a lightweight transformer-based model optimized for multilingual text processing and translation. Google Gemini 1.5 Flash supports over 100 languages and is designed to strike a balance between speed, computational efficiency and high-quality text creation. With a context window of up to 1 million tokens, it can process large volumes of text in a single prompt and is therefore well suited to the translation of historical documents. As our dataset consists of colonial-era correspondence, it was important to maintain historical accuracy and linguistic integrity. To achieve this, we carefully crafted the following translation prompt: Prompt:"You are a skilled historical linguist and translator with deep knowledge of both colonial-era Portuguese and archaic/historical English usage. Your task is to translate the following Portuguese text into an English style that reflects the era in which it was originally written. Please: Maintain the historical tone. Avoid modern terms and slang. Capture the nuanced formality of the original text." The translated dataset, which is structured in the same format as the original, ensures linguistic and historical authenticity and at the same time makes the correspondence accessible to a wider audience 3. MAPE Dataset Bilingual Version with Senders and Recipients (version 3) 2. v2 → v3: a small technical issue appeared When adding the Sender Name and Receiver Name fields, the pipeline lost some records. This looks like an export/parsing error rather than an intentional research selection. 4. MAPE Dataset: Bilingual Version with Senders and Recipients with assigned topics (version 4) This version enriches the previous versions of bilingual dataset by adding automatically assigned topics to each correspondence record. Topics were generated through a multi-stage framework combining large language models, multilingual embeddings, and clustering, producing both concise document-level tags and macro-level thematic groups. At the highest level, the corpus is organized into seven thematic clusters, providing a scalable structure for thematic exploration and navigation across the 182,000+ records of the Portuguese Overseas Historical Archive. Topic Number of records % if the corpus Colonial Administration, Trade & Revenue 32,298 19,2% Petitions, Appointments & Royal Permissions 30,881 18,4% Military Personnel, Discipline & Logistics 28, 617 17,0% Maritime Trade, Naval Operations & Logistics 27, 976 16,6% Civil Administration, Justice & Royal Governance 24,764 14,7% Land Grants, Passports & Travel 13,917 8,3% Military Appointments, Ranks & Confirmations 9,592 5,7% Observations/Problems detected in dataset 4 (as per 19.03.2026) 3. v3/v2 → v4: a real corpus selection took place Most importantly: in v4, entire source groups disappear. These are not random gaps of individual records. Among those that disappeared completely are: SÃO PAULO-MG – 5,117 documents SÃO TOMÉ E PRÍNCIPE – 3,771 + 64 SÃO PAULO – 1,380 RIO NEGRO – 740 SANTA CATARINA – 619 SERGIPE D’EL-REI – 495 TIMOR – 272 MONTEVIDEO – 224 PARAGUAY – 27 Additionally, even within sources retained in v4, partial losses are visible, e.g.: RIO DE JANEIRO-CA – minus 400 BAHIA – minus 283 MINAS GERAIS – minus 248 5. MAPE Dataset: Bilingual Version with Senders and Recipients with assigned topics (version 5) - detected issues have been resolved; a detailed report MAPE_Dataset_Report_2026-04-17.pdf is provided together with the dataset. V5 conclusion: No permanent reduction of the corpus occurred; the V5 version restores the full size of the archival material. Errors were indeed made, but primarily at the v3 stage. The V5 version corrects these issues and introduces clear analytical improvements, but retains several formatting inconsistencies that require refinement if the dataset is to function as a long-term reference dataset. The most accurate characterization of V5 is therefore as a complete and analytically mature version, yet still requiring final standardization of metadata representation. Changes in dataset size were not the result of a single, uniform process, but rather a sequence of distinct operations: initial data enrichment (v2), a temporary technical disruption (v3), deliberate analytical reduction (v4), and the final reconstruction of the full corpus in the V5 version. The key problem remains the lack of transparent documentation of the transitions between these stages. Comparing Dates V1 & V5: The comparative analysis of dates between version v1 and V5 shows no losses or changes in the chronological distribution of documents. This means that the data processing did not affect the temporal scope of the corpus. The only significant change concerns the representation of dates: in the V5 version, some records were enriched with monthly information (YYYY-MM format), but this operation was applied inconsistently. Consequently, the issue is structural in nature (data format), rather than substantive (loss of information).



