遇见数据集

A Diachronic Dataset of Semantically-Annotated Geographical Nouns in Ancient Greek and Latin

收藏
Zenodo2026-02-10 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the dataset and supplementary materials for the paper "Sense-Based Annotation of Geographical Nouns in Ancient Greek and Latin: A Diachronic Study with LLMs". This dataset supports a diachronic and cross-linguistic analysis of how geographical concepts (e.g., city, sea, mountain) were lexicalised in Ancient Greek and Latin between the 8th century BCE and the 2nd century CE. The data includes a manually curated vocabulary of place names, a gold standard validation set, and a large-scale corpus automatically annotated with WordNet synsets using Large Language Models (GPT-5.2). File Description The repository consists of three primary CSV files: 1. place_names.csv This file contains the bilingual inventory of geographical nouns (GNs) used to extract tokens from the corpus. It maps English spatial concepts to their Latin and Ancient Greek lexical counterparts. Columns: CONCEPT: The English geographical concept serving as the onomasiological anchor (e.g., CITY, RIVER). category: The semantic category of the place. Latin: The Latin lemma(s) expressing the concept. Ancient Greek: The Ancient Greek lemma(s) expressing the concept. 2. annotated_ground_truth.csv This file contains the manually annotated validation set used to evaluate the performance of the LLM annotator. It consists of 252 tokens sampled from the corpus. Columns: ID: Unique identifier for the token. TOKEN: The specific word form as it appears in the text. SENTENCE: The context sentence containing the token. LEMMA: The dictionary form of the word. SEMANTICS: The manually assigned WordNet synset/sense ID. LANGUAGE: The language of the text (Latin or Ancient Greek). 3. annotated_tokens.csv This file contains the full dataset of 16,429 geographical noun occurrences extracted from the PREMOVE Base Corpus and automatically annotated by the LLM. Columns: ID: Unique identifier for the token. TOKEN: The word form in the text. LEMMA: The lemma of the token. SENTENCE: The context sentence. author & title: Metadata identifying the source text. LANGUAGE: Latin or Ancient Greek. passage: The specific citation/location within the work. PREDICTED_SEMANTICS: The WordNet synset ID predicted by the model (GPT-5.2). PRED_CONFIDENCE: The confidence score (0-1) assigned by the model for its prediction. PRED_LITERAL: Binary classification (yes/no) indicating if the usage is literal or figurative. PRED_SOURCE: The source of the synset (e.g., Latin WordNet, Open English WordNet). EXAMPLE_COUNT: The number of few-shot examples provided to the model during annotation. Methodology The data was derived from the PREMOVE Base Corpus, a multi-genre collection of Latin and Ancient Greek texts. The automatic annotation was performed using GPT-5.2, which disambiguated senses by selecting appropriate synsets from the Latin WordNet (LWN) and Open English WordNet (OEWN). Funding This work is supported by the UKRI under the Horizon Europe Guarantee (grant number UKRI947) for the project COALA (Computational Corpus Annotation for Quantitative Analysis of Latin Lexical Semantics) successfully evaluated by the ERC, and by King's College London's AHRS Research Grant (Research & Scholarship Development Stream) for the project "Mapping meaning with Large Language Models''. Keywords Ancient Greek, Latin, Geographical Nouns, Word Sense Disambiguation, LLMs, Digital Humanities, Historical Semantics, Diachronic Linguistics.

提供机构:
Zenodo
创建时间:
2026-02-10
二维码
社区交流群
二维码
科研交流群
商业服务