遇见数据集

Divergent Discourses Modern Tibetan NER Training Data

收藏
Zenodo2025-12-18 更新2026-05-26 收录
官方服务:

资源简介:

The Divergent Discourses Modern Tibetan NER Training dataset was created to train a spaCy model to detect named entities in Tibetan newspapers of the 1950s and 1960s published in the PRC and in South Asia. The dataset was generated as follows: plain Tibetan paragraphs, each consisting of 300 characters according to Python’s character-counting system (e.g., བོད་ is counted as four characters), were submitted to Claude Sonnet 4 (July 2025), which extracted named entities. Each entity in the dataset is defined by three parameters: the start position of the named entity, the end position, and the entity label. The proportion of each source text is shown in the following table (See also here): Source Proportion Tibetan books, 1950s-1960s (Divergent Discourses) 0.007 Tibetan Newspapers, 1950s-1960s (Divergent Discourses) 0.053 Esukhia Tibetan Newspaper Corpus 0.941 Claude extracted named entities of the following categories: See here.

提供机构:
Zenodo
创建时间:
2025-12-18
二维码
社区交流群
二维码
科研交流群
商业服务