CNR-ILC/gs-dataset-eval
收藏资源简介:
这是一个用于古希腊纸草学文本中填空任务(gap filling)的评估数据集。数据集包含以下特征:x(带有掩码空缺的文本)、y(领域专家提出的可接受补充列表,即黄金标准)、gap_length(空缺的估计长度)、corpus_id/file_id(原始文档的标识符)。数据来源包括MAAT(机器可操作古代文本语料库)、PDL(Perseus数字图书馆)、First1KGreek和TLG(希腊文辞典)。数据集分为开发集(dev,包含P.Herc.块的开发案例)和测试集(test,包含最终测试案例),均带有黄金标准标签。
Evaluation dataset for the gap filling task in ancient Greek papyrological texts. The dataset includes features such as x (text with masked gaps), y (list of acceptable integrations proposed by domain experts, i.e., gold labels), gap_length (estimated length of the gap), and corpus_id/file_id (identifiers of the original document). Data sources include MAAT (Machine-Actionable Ancient Text corpus), PDL (Perseus Digital Library), First1KGreek, and TLG (Thesaurus Linguae Graecae). The dataset is split into dev (development cases with P.Herc. blocks) and test (final test cases), both containing gold labels.




