The Weimar Parliament Corpus
收藏资源简介:
The files corpus.parquet and corpus.csv contain all speeches and interjections recorded in the official proceedings of the German Reichstag during the Weimar Republic (1920-1932). The files topics.parquet and topics.csv contain sentence-level topic classifications that can be linked to the corpus. The folder ParlaClarin contains the data (corpus and topic classifications combined) in the ParlaCLARIN TEI format. The folder also contains the Python script used to generate this format. The file Codebook.pdf contains brief descriptions of the variables in each file as well as the coding scheme and instructions used for the sentence-level topic classifications. The folder additional-files/PAGE-XML-files contains the PAGE XML files that form the basis for the corpus. The layout/segmentation model and the transcription model used to produce these files can be found at https://doi.org/10.5281/zenodo.19222213 and https://doi.org/10.5281/zenodo.19188500. The folder additional-files/scripts contains scripts used during the OCR step (PagePlus.py), for post-processing the transcriptions (cleanXML_v4.py), generating a figure with topic salience (plot_replication.R), and training and running the topic classifier (sentence_split.ipynb, train_classifier.ipynb, classify.ipynb).



