A Diachronic Dataset of Semantically-Annotated Geographical Nouns in Ancient Greek and Latin
收藏资源简介:
This repository contains the dataset and supplementary materials for the paper "Sense-Based Annotation of Geographical Nouns in Ancient Greek and Latin: A Diachronic Study with LLMs". This dataset supports a diachronic and cross-linguistic analysis of how geographical concepts (e.g., city, sea, mountain) were lexicalised in Ancient Greek and Latin between the 8th century BCE and the 2nd century CE. The data includes a manually curated vocabulary of place names, a gold standard validation set, and a large-scale corpus automatically annotated with WordNet synsets using Large Language Models (GPT-5.2). File Description The repository consists of three primary CSV files: 1. place_names.csv This file contains the bilingual inventory of geographical nouns (GNs) used to extract tokens from the corpus. It maps English spatial concepts to their Latin and Ancient Greek lexical counterparts. Columns: CONCEPT: The English geographical concept serving as the onomasiological anchor (e.g., CITY, RIVER). category: The semantic category of the place. Latin: The Latin lemma(s) expressing the concept. Ancient Greek: The Ancient Greek lemma(s) expressing the concept. 2. annotated_ground_truth.csv This file contains the manually annotated validation set used to evaluate the performance of the LLM annotator. It consists of 252 tokens sampled from the corpus. Columns: ID: Unique identifier for the token. TOKEN: The specific word form as it appears in the text. SENTENCE: The context sentence containing the token. LEMMA: The dictionary form of the word. SEMANTICS: The manually assigned WordNet synset/sense ID. LANGUAGE: The language of the text (Latin or Ancient Greek). 3. annotated_tokens.csv This file contains the full dataset of 16,429 geographical noun occurrences extracted from the PREMOVE Base Corpus and automatically annotated by the LLM. Columns: ID: Unique identifier for the token. TOKEN: The word form in the text. LEMMA: The lemma of the token. SENTENCE: The context sentence. author & title: Metadata identifying the source text. LANGUAGE: Latin or Ancient Greek. passage: The specific citation/location within the work. PREDICTED_SEMANTICS: The WordNet synset ID predicted by the model (GPT-5.2). PRED_CONFIDENCE: The confidence score (0-1) assigned by the model for its prediction. PRED_LITERAL: Binary classification (yes/no) indicating if the usage is literal or figurative. PRED_SOURCE: The source of the synset (e.g., Latin WordNet, Open English WordNet). EXAMPLE_COUNT: The number of few-shot examples provided to the model during annotation. Methodology The data was derived from the PREMOVE Base Corpus, a multi-genre collection of Latin and Ancient Greek texts. The automatic annotation was performed using GPT-5.2, which disambiguated senses by selecting appropriate synsets from the Latin WordNet (LWN) and Open English WordNet (OEWN). Funding This work is supported by the UKRI under the Horizon Europe Guarantee (grant number UKRI947) for the project COALA (Computational Corpus Annotation for Quantitative Analysis of Latin Lexical Semantics) successfully evaluated by the ERC, and by King's College London's AHRS Research Grant (Research & Scholarship Development Stream) for the project "Mapping meaning with Large Language Models''. Keywords Ancient Greek, Latin, Geographical Nouns, Word Sense Disambiguation, LLMs, Digital Humanities, Historical Semantics, Diachronic Linguistics.
本仓库收录了论文《基于义项的古希腊语与拉丁语地理名词标注:结合大语言模型(Large Language Model,LLM)的历时研究》的数据集与配套辅助材料。 本数据集支持对公元前8世纪至公元2世纪间,地理概念(如城市、海洋、山脉)在古希腊语与拉丁语中的词汇化路径开展历时与跨语言分析。数据集包含经人工甄选的地名词汇表、一套金标准验证集,以及借助GPT-5.2自动标注了WordNet义项的大规模语料库。 文件说明 本仓库包含三个核心CSV文件: 1. place_names.csv:本文件收录了用于从语料库中提取Token(词元)的地理名词(Geographical Nouns,GNs)双语对照表,将英语空间概念映射至其拉丁语与古希腊语对应词目。 字段说明: - CONCEPT:作为命名学锚点的英语地理概念(如CITY、RIVER)。 - category:该地理实体的语义类别。 - Latin:表达该概念的拉丁语词目(复数形式)。 - Ancient Greek:表达该概念的古希腊语词目(复数形式)。 2. annotated_ground_truth.csv:本文件包含用于评估大语言模型标注器性能的人工标注验证集,共从语料库中采样252个Token。 字段说明: - ID:该Token的唯一标识符。 - TOKEN:文本中出现的具体词形。 - SENTENCE:包含该Token的上下文语句。 - LEMMA:该单词的词典形式(词目)。 - SEMANTICS:人工指派的WordNet义项编号。 - LANGUAGE:文本所属语言(拉丁语或古希腊语)。 3. annotated_tokens.csv:本文件包含从PREMOVE基础语料库中提取的全部16429条地理名词出现实例,以及由大语言模型自动完成的标注结果。 字段说明: - ID:该Token的唯一标识符。 - TOKEN:文本中的词形。 - LEMMA:该Token的词目。 - SENTENCE:上下文语句。 - author & title:标识源文本的元数据。 - LANGUAGE:拉丁语或古希腊语。 - passage:该作品内的具体引用/位置。 - PREDICTED_SEMANTICS:模型(GPT-5.2)预测的WordNet义项编号。 - PRED_CONFIDENCE:模型为其预测赋予的置信度评分(取值范围0至1)。 - PRED_LITERAL:二元分类标签(是/否),用于标识该用法为字面义还是比喻义。 - PRED_SOURCE:义项来源(如拉丁语WordNet(LWN)、开放英语WordNet(OEWN))。 - EXAMPLE_COUNT:标注过程中向模型提供的少样本(Few-shot)示例数量。 研究方法 本数据集源自PREMOVE基础语料库——一个涵盖多体裁的拉丁语与古希腊语文本合集。自动标注环节借助GPT-5.2完成,该模型通过从拉丁语WordNet(LWN)与开放英语WordNet(OEWN)中选取适配的义项集合来完成词义消歧。 资助信息 本研究由英国研究与创新署(UKRI)根据地平线欧洲保障计划资助(项目编号:UKRI947),相关项目为COALA(拉丁语词汇语义量化分析的语料库计算标注),该项目已通过欧洲研究理事会(ERC)评估;同时亦获伦敦国王学院的AHRS研究资助金(研究与学术发展专项),对应项目为《借助大语言模型映射语义》。 关键词 古希腊语、拉丁语、地理名词、词义消歧、大语言模型(LLM)、数字人文、历史语义学、历时语言学。



