Named-entity recognition outputs for the Divergent Discourses corpus of Tibetan newspapers, 1951–1965
收藏资源简介:
This dataset contains named-entity recognition outputs derived from the Divergent Discourses corpus of Tibetan newspapers. The corpus was prepared within the AHRC/DFG-funded project Divergent Discourses: Processes of Narrative Construction in Tibet, 1955–62, a collaborative project concerned with Tibetan newspapers, narrative construction, and political discourse in mid-twentieth-century Tibet and its wider contexts. The files deposited here were generated from a CSV version of the Divergent Discourses corpus prepared by Franz Xaver Erhard. Although the main project title refers to 1955–62, the CSV processed for this deposit spans 1951–1965. A modified spaCy script was used to process the full corpus and identify named entities. The resulting JSON file records, for each corpus item, metadata such as record ID, source image filename, date, input text, entity count, and extracted entities with their labels and character offsets. A second script tabulated the extracted entities by label and frequency and produced a CSV summary. This summary includes each entity’s label, entity text, total count, first date of occurrence, last date of occurrence, and a month-by-month “burst” bar across the 1951–1965 period. In the burst bar, each position represents a month, with numerals indicating the number of occurrences and # indicating ten or more occurrences in that month. The data should be treated as automatically generated NLP output. It is intended to support exploratory research, corpus navigation, and further work on Tibetan named-entity recognition, discourse analysis, and digital humanities methods. The entity recognition is imperfect: some entities are incorrectly identified, some are missed, and some errors reflect the limits of the training data, OCR/HTR quality, Tibetan tokenisation, and the challenges of named-entity recognition for modern Tibetan.



