glossAPI/archetai
收藏资源简介:
该数据集包含从雅典考古学会数字出版物档案中OCR提取的文本。该学会成立于1837年,是希腊最古老的学术团体之一,主要进行和发表考古研究。其出版物涵盖挖掘报告、碑铭学、艺术史、古迹保护以及相关的历史和文献学研究,时间跨度从19世纪中期至今。数据集中的每条记录对应一个PDF卷、专著或五个出版系列中的一个问题。数据集包含814条记录,结构为扁平表格数据集,具有三个字符串类型的列。每条记录由其`pdf_url`唯一标识。`collection`列包含五个分类标签之一,与源档案的五个子目录一一对应。数据集的语言主要为希腊语,包含少量外语项目(主要是英语、德语和法语的摘要或完整卷)。由于源文档是扫描的PDF,OCR输出存在一定的噪声,特别是较旧的卷结合了多调和单调希腊语、连字和19世纪的印刷惯例。
This dataset contains text extracted via OCR from the digital publication archive of the Archaeological Society at Athens. Founded in 1837, it is one of the oldest academic institutions in Greece, dedicated to conducting and publishing archaeological research. Its publications cover excavation reports, epigraphy, art history, monument preservation, as well as related historical and philological studies, spanning from the mid-19th century to the present. Each record in the dataset corresponds to a PDF volume, monograph, or an issue of one of the five publication series. The dataset comprises 814 records, structured as a flat tabular dataset with three string-type columns. Each record is uniquely identified by its `pdf_url` field. The `collection` column holds one of five categorical labels, which have a one-to-one correspondence with the five subdirectories of the source archive. The primary language of the dataset is Greek, with a small number of foreign-language entries, mainly abstracts or full volumes in English, German, and French. As the source documents are scanned PDFs, the OCR outputs carry a certain amount of noise, especially for older volumes that combine polytonic and monotonic Greek, ligatures, and 19th-century printing conventions.




