遇见数据集

LIFT2-25_Gascon_Dataset

收藏
Zenodo2025-10-08 更新2026-05-29 收录
官方服务:

资源简介:

This repository provide the data used in "Adaptation of models for parsing of Old Gascon" (Natasha Romanova, Rayan Ziane and Barbara Francioni) The training and pre-finetuning corpora come from several Romance treebanks annotated within the Universal Dependencies framework: TTB – Tolosa Treebank (Modern Occitan) https://github.com/UniversalDependencies/UD_Occitan-TTB/tree/master PRFT – Profiterole (Medieval French) https://github.com/UniversalDependencies/UD_Old_French-PROFITEROLE/tree/master https://github.com/UniversalDependencies/UD_Middle_French-PROFITEROLE/tree/master OI – Old Italian Treebank (Medieval Italian) https://github.com/UniversalDependencies/UD_Italian-Old/tree/master CorAG – Corpus d’Ancien Gascon (Medieval Occitan, Gascon variety) https://github.com/UniversalDependencies/UD_Old_Occitan-CorAG/tree/master All datasets are provided in CoNLL-U format and include normalized train/dev/test splits. For pre-finetuning, the training, development and test sets of the source treebanks were concatenated to form a unified training set (re-splited into train/dev).To ensure comparability, the Romance corpora used for pre-finetuning were also downsampled to match the size of TTB, the smallest treebank ("training_split_ech"). For the target language (Old Gascon), the training material was incrementally sampled in steps of 100 to 607 sentences, allowing us to measure the effect of gradually increasing the amount of in-domain data on parsing performance. The version of CorAG used in these experiments predates its official release in the Universal Dependencies collection (v2.16) and corresponds to an earlier state of manual correction and sentence segmentation. All corpora share the same UD annotation schema, ensuring full cross-compatibility of the trained models.

提供机构:
Zenodo
创建时间:
2025-10-07
二维码
社区交流群
二维码
科研交流群
商业服务