LAYERS Index Dataset for Documenta Indica (Monumenta Historica Societatis Iesu): Normalized Index Entries and Page-Level Occurrences, Vols. 1–16 and 18
收藏资源简介:
This dataset was prepared in connection with the ERC Advanced Grant application Hidden Layers: Converso Networks and the Neural Architecture of the Early Jesuits (LAYERS), led by Robert Aleksander Maryks. It provides a structured, provenance-aware dataset derived from the indexes of Documenta Indica, edited by Iosephus Wicki, S.J., and published within the Monumenta Historica Societatis Iesu series. The deposited workbook contains 112,837 page-level index occurrences extracted from 17 volumes: volumes 1–16 and 18. Volume 17 is not included in the present version, for it lacks an index. Rather than reproducing the source volumes themselves, the dataset transforms their printed indexes into a machine-readable structure suitable for systematic querying, name normalization, entity-resolution workflows, provenance tracking, and subsequent historical network analysis. Each row represents an indexed occurrence linked to its source volume and page. The dataset preserves the original index form while recording successive stages of normalization and aggregation. The principal fields are volume, source_file, source_sheet, row_in_sheet, page, original_names, automatic_names, name, context, automatic_aggregation_applied, manual_decision_applied, manual_name_changed, and aggregation_applied. The workbook contains 11,472 distinct values in the original_names field and 10,472 distinct final normalized name labels. The distinction between original, automatically normalized, and final forms is intentional. It preserves the provenance of computational transformations and makes human interventions identifiable rather than silently replacing the source index terminology. The accompanying processing flags indicate whether individual occurrence records were affected by automatic aggregation, manual review, manual name modification, or the final aggregation procedure. Within LAYERS, the dataset serves as an intermediate evidentiary layer between the digitized historical sources and the project’s relational data infrastructure. The project investigates whether converso-ascribed actors occupied distinctive positions in the governance, archival-secretarial communication, and pedagogical structures of the early Society of Jesus, and whether organizational routines associated with these actors remained resilient during the progressive exclusion of conversos between 1573 and 1608. Index-derived data provide a scalable means of locating actors, places, documentary contexts, and recurring associations across the source series before more computationally expensive full-text extraction and archival verification are applied. The dataset therefore supports the first stage of the LAYERS workflow: historical reconstruction, corpus processing, provenance control, normalization, and entity-resolution preparation. It is intended to facilitate subsequent linkage with OCR/HTR-derived text, TEI-aligned documentary references, authority data, and the project’s temporally evolving multiplex graph. Normalization in this dataset should not be interpreted as complete historical entity disambiguation. Identical or similar name forms may refer to different historical actors, and the final name field represents a normalized analytical label rather than an authoritative prosopographical identifier. Ambiguous identities, kinship relations, and historically contested attributions require further documentary and expert verification before being incorporated into the project's high-confidence relational dataset. The dataset was produced primarily through automated and semi-automated extraction, parsing, normalization, and aggregation of the printed indexes, followed by selective manual review. Although quality-control procedures were applied, the resource should not be treated as a definitive or error-free transcription of the original indexes. It may contain residual errors resulting from automated processing, including incorrect name segmentation, normalization, aggregation, page assignment, or interpretation of index structure. Users are therefore advised to verify individual records against the original printed volumes when using the dataset as evidence for historical interpretation. The dataset is published as a research-ready intermediate resource designed for further validation, entity resolution, and refinement rather than as a critical edition of the indexes. This resource complements the LAYERS Corpus Register: Bibliographic Dataset of Digitized Jesuit Institutional Sources (DOI: 10.5281/zenodo.20811183), which documents the broader 139-volume source base of the project. Whereas the Corpus Register describes the available source corpus at bibliographic level, the present dataset provides occurrence-level structured data derived from one of its major source series. The dataset is intended as a transparent and reusable resource for Jesuit studies, early modern history, historical prosopography, digital humanities, entity resolution, and historical network analysis. Subsequent corrections, additional volumes, improved disambiguation, and transformations into relational, TEI, RDF, or graph-based formats should be released as versioned updates or associated research datasets.



