遇见数据集

Semantic Search Engine and Vernacularity Predictor for the Southern Macro-Region of Brazil AHU Catalog

收藏
Zenodo2026-03-09 更新2026-05-26 收录
官方服务:

资源简介:

Overview This dataset and accompanying web application provide a machine-readable, semantically annotated corpus of catalog summaries from the Arquivo Histórico Ultramarino (AHU). Focused on the "Macrorregião Sul do Brasil" during the late colonial period (1737–1828), these files are the product of a pipeline that utilized Large Language Models, Lexical Search Engines, and dense vector embeddings to quantify sociolinguistic variables. The project is designed to support research in Diachronic Linguistics, specifically aimed at mapping the competition of grammars and isolating the probability of "Vernacular Leakage" (oral/Brazilian Portuguese syntax) in historical administrative documentation. Furthermore, it serves as an advanced finding aid, generating direct, clickable hyperlinks to the manuscripts' registries in the DigitArq (Portugal) and Projeto Resgate (Brazil) databases. Files Included in this Dataset The dataset and application are distributed as a repository containing the following core files: ahu_documents_phase2_by_deepseek_v7_crav.json: The core structured resource (JSON format). It connects the raw archival text to geographical metadata, temporal data, modernized archival reference codes, and algorithmic sociolinguistic annotations. Core Keys: reference_code: Modernized CRAV/DigitArq reference code (e.g., PT/AHU/CU/023-001/0006/00631). new_code: Standardized/intermediary code still used by the Projeto Resgate. description: Full text description from the AHU catalog. vernacular_score: A quantitative probability decimal (0.0 to 1.0) indicating the likelihood of localized syntax. vector: Directionality of the document (e.g., Top-Down, Bottom-Up). extracted_typology: Normalized singular root of the document type (e.g., CARTA, REQUERIMENTO). sociolinguistic_reasoning_by_deepseek_v3: The AI-generated justification for the assigned vernacular score. ahu_semantic_index.pkl: A serialized Python dictionary containing the dense vector embeddings of the archival descriptions. Generated via the intfloat/multilingual-e5-large model, these embeddings are hard-anchored to the documents' reference codes to ensure perfect synchronization during semantic retrieval. app.py: The Streamlit Application script (GUI). It handles the dynamic Hybrid Ensemble Search (fusing BM25 Lexical scoring with E5 Semantic Similarity), sociolinguistic filtering ("Lenses"), dynamic URL routing to digital archives, and programmatic ABNT-formatted PDF dossier generation via FPDF. requirements.txt: List of Python environment dependencies required to execute the pipeline and web application (e.g., pandas, streamlit, sentence-transformers, rank-bm25, torch, fpdf). Scope Included Temporal Boundary: 1737 to 1828. Geographical Focus: Southern Macro-Region of Brazil (Macrorregião Sul do Brasil). Data Source: Summaries and metadata extracted from the Arquivo Histórico Ultramarino (AHU) catalogs. Source of Archival Data The original unstructured textual summaries used to build this corpus were obtained from the public cataloging efforts of the Arquivo Histórico Ultramarino and the Projeto Resgate. We gratefully acknowledge their continuous work in preserving and indexing the colonial documentary heritage. The modernized reference codes generated by this application interface directly with the Arquivos de Portugal (DigitArq) infrastructure. Methodology Extraction & Modernization: RegEx parsing of raw AHU catalog text to isolate document IDs, dates, and typologies, followed by algorithmic zero-padding to reverse-engineer and generate modern CRAV-compliant reference codes. Annotation: Zero-shot LLM inference (DeepSeek v3) to categorize the sender/recipient, vector of communication, and compute the Probability of Vernacularity Score (SPV) alongside a written justification. Vectorization (Semantic): Context-enriched passages embedded using HuggingFace's heavyweight intfloat/multilingual-e5-large model. Tokenization (Lexical): Term frequency-inverse document frequency indexing via the BM25Okapi algorithm to ensure precision recall of proper nouns and locations. Search Fusion: Dynamic Ensemble Retrieval calculating a weighted hybrid score between semantic meaning and lexical exactness. Methodological Criteria For a detailed description of the computational pipeline, LLM prompting parameters, the "CRAV Reference Modernizer" logic, the Ensemble Hybrid Search algorithms, and the UI/UX deployment rules, please refer to the forthcoming data paper that will be associated with this repository. Citation Currently, this dataset and application are published as an independent computational resource. If you use this corpus, methodology, or code, please cite this repository directly using its Zenodo DOI: Pacheco Rocha, Saulo Rogério. (2026). AHU Catalog Document Classifier for the Southern Macro-Region of Brazil (Version 2.0) [Data set and Software]. Zenodo. https://doi.org/10.5281/zenodo.18772667 License This dataset and the associated code are made available under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) license. Official Repository Source code, updates, and full documentation can also be found on GitHub:🔗 https://github.com/saulorrp/catalogo-ahu-cone-sul.git

提供机构:
Zenodo
创建时间:
2026-02-25
二维码
社区交流群
二维码
科研交流群
商业服务