PTPARL-D
收藏资源简介:
PTPARL-D是由古尔本基安科学研究所创建的葡萄牙议会辩论标注语料库,涵盖了1976至2019年间的所有葡萄牙民主时期。该数据集包含4448场辩论,每场辩论平均有305.12条发言,总计约95.20字每条。数据集通过人工和光学字符识别技术从葡萄牙议会网站下载的文本中创建,旨在通过自然语言处理和文本挖掘技术,深入分析政治数据,提高政治过程的透明度和公众参与度。该数据集特别适用于研究多党制民主体系的政治动态和话语分析。
PTPARL-D is an annotated corpus of Portuguese parliamentary debates developed by the Gulbenkian Institute of Science, covering the entire democratic period of Portugal from 1976 to 2019. This dataset comprises 4,448 debates, with an average of 305.12 utterances per debate and an average of approximately 95.20 words per utterance. Built from texts downloaded from the Portuguese Parliament's website using both manual annotation and optical character recognition (OCR) techniques, it aims to enable in-depth analysis of political data via natural language processing (NLP) and text mining technologies, thereby enhancing the transparency of the political process and public participation. This corpus is particularly well-suited for research on political dynamics and discourse analysis in multi-party democratic systems.



