遇见数据集

The Weimar Parliament Corpus

收藏
Zenodo2026-08-02 更新2026-08-13 收录
官方服务:

资源简介:

The files corpus.parquet and corpus.csv contain all speeches and interjections recorded in the official proceedings of the German Reichstag during the Weimar Republic (1920-1932). The files topics.parquet and topics.csv contain sentence-level topic classifications that can be linked to the corpus. The folder ParlaClarin contains the data (corpus and topic classifications combined) in the ParlaCLARIN TEI format. The folder also contains the Python script used to generate this format. The file Codebook.pdf contains brief descriptions of the variables in each file as well as the coding scheme and instructions used for the sentence-level topic classifications. The folder additional-files/PAGE-XML-files contains the PAGE XML files that form the basis for the corpus. The layout/segmentation model and the transcription model used to produce these files can be found at https://doi.org/10.5281/zenodo.19222213 and https://doi.org/10.5281/zenodo.19188500. The folder additional-files/scripts contains scripts used during the OCR step (PagePlus.py), for post-processing the transcriptions (cleanXML_v4.py), generating a figure with topic salience (plot_replication.R), and training and running the topic classifier (sentence_split.ipynb, train_classifier.ipynb, classify.ipynb).

提供机构:
Zenodo
创建时间:
2026-08-02
二维码
社区交流群
二维码
科研交流群
商业服务