遇见数据集

A Comprehensive Catalogue of Digital Latin Corpora from Archaic to Neo-Latin (700 BCE – 2000 CE)

收藏
Figshare2026-04-01 更新2026-04-28 收录
官方服务:

资源简介:

This dataset provides a structured survey of major digital corpora of Latin texts spanning from Archaic Latin to Neo-Latin. It includes corpora that are available in machine-readable formats (e.g., TXT, XML, HTML), excluding image-only or non-textual resources.Each entry in the dataset represents a corpus and is annotated with detailed metadata, including:corpus name and access linksize (number of words, works, and authors)chronological coverage (centuries, language periods)geographical scopeannotation type (e.g., morphological, syntactic, lemmatised)literary genresaccess conditions (open, commercial, etc.)file formats (e.g., XML, HTML, CSV)associated reference publicationsproject status and main creator(s)The dataset is intended as a reference resource for researchers in Digital Humanities, Classics, Corpus Linguistics, and Natural Language Processing working on Latin texts. It facilitates corpus discovery, comparison, and selection for linguistic, philological, and computational research.This resource aims to improve transparency and accessibility in the landscape of Latin digital corpora and to support reproducible research.For suggestions, corrections, or updates, please contact the author (email provided in author metadata).

本数据集针对从古拉丁语(Archaic Latin)到新拉丁语(Neo-Latin)时期的主流拉丁语数字语料库开展了结构化梳理调研。其收录的语料库均采用机器可读格式(如TXT、XML、HTML),排除仅含图像或非文本类资源。数据集中的每条记录对应一个语料库,并附带详细元数据标注,具体包含以下内容:语料库名称与访问链接、语料规模(含词数、作品数与作者数)、时间覆盖范围(世纪跨度与语言时期)、地理覆盖范围、标注类型(如词形分析、句法分析、词形还原标注)、文学体裁、获取条件(开放获取、商业授权等)、文件格式(如XML、HTML、CSV)、关联参考出版物、项目状态与主要创建者。本数据集旨在为从事拉丁语文本研究的数字人文(Digital Humanities)、古典学、语料库语言学(Corpus Linguistics)以及自然语言处理(Natural Language Processing)领域的研究者提供参考资源,可助力研究者在开展语言学、文献学与计算类研究时,实现语料库的检索、比对与遴选。本资源旨在提升拉丁语数字语料库领域的透明度与可及性,并为可复现研究提供支撑。若有建议、修正或更新需求,请联系数据集创建者(作者元数据中附有联系邮箱)。

创建时间:
2026-04-01
二维码
社区交流群
二维码
科研交流群
商业服务