遇见数据集

Data and code for Enjambement as Syntactic Disruption in Historical and Contemporary German and Spanish Poetry

收藏
Zenodo2026-08-18 更新2026-08-20 收录
官方服务:

资源简介:

The repository contains data and code to reproduce our paper "Enjambement as Syntactic Disruption in Historical and Contemporary German and Spanish Poetry", accepted at the Computational Humanities Research journal. Repository structure annot-auto: Scripts for training, automatic annotation and result visualization annot-manual Manual annotation and inter-annotator agreement Annotation guidelines corpus-dev: Documents Spanish corpus selection criteria, with more detail than the body of the paper and its appendices. As regards German corpus development, this was already documented in publications cited in the paper. data Manual and large-scale corpora Contain the boundary class predictions and metrical annotation on which the study relies In contemporary texts, to preserve reuse conditions, line text was redacted and only the annotations and the metadata were preserved. This does not affect reproducing the paper's figures. For research purposes only, and not for redistribution, we could make available those poems' texts privately. To reproduce the paper figures Plots can be reproduced with the reproduce.sh bash script, which must be executable. The script creates a repro directory in the current directory, with figures numbered as in the paper. Run ./reproduce.sh. This was tested with Python 3.12.3 on Ubuntu 24.04.4 LTS. For the plots, these packages are required: numpy pandas matplotlib For analyzing the manual annotation data under annot-manual, including inter-annotator agreement, the following packages would be required, but reproducing that is not part of reproduce.sh: numpy pandas matplotlib scikit-learn seaborn odfpy

提供机构:
Zenodo
创建时间:
2026-08-18
二维码
社区交流群
二维码
科研交流群
商业服务