Konbaung Chronicle V3: Canonical Claims, Source Sentences and Embeddings
收藏资源简介:
I built this dataset to study kingship, office, military command, religious patronage, tribute and kinship in the Konbaung chronicles. It connects 27,129 canonical subject–predicate–object claims to Burmese source sentences, English translations, source-page identifiers, semantic categories and reusable vector representations across three volumes and 1,215 pages. The release contains 11,282 sentence records; 23,890 entity and 11,886 relation labels with 768-dimensional float32 embeddings; 27,135 contextual triple vectors with explicit claim-to-vector links; and entity-cluster and occurrence-level person-resolution tables. The three released matrices preserve the original vectors used in this research. My construction pipeline combines OCR and sentence reconstruction, model-assisted translation and V3 claim extraction, canonical selection, a 52-entity/81-relation category system, Gemini Embedding 2 representations, and guarded identity resolution. The record-level identifiers retain the path from graph claims back to source text and pages. The accompanying guide explains the schemas, joins, extraction provenance and resolution variants. A manifest records every original payload checksum, and a Python example reads an actual claim and its linked source text and vectors. These are research annotations of chronicle assertions, with source context retained for interpretation. Methods, prompts, software and analysis: https://github.com/conradcompagna/konbaung-knowledge-graph Released under the Conrad Compagna Research Data Evaluation License 1.0 for research, education, experimentation and evaluation, including hiring evaluation. Commercial use and redistribution of the original or modified dataset and embeddings require separate permission; third-party rights remain with their holders.



