taz2024full
收藏资源简介:
taz2024full is a large German newspaper corpus. It consists of over 1.8 million articles published between 1980 and 2024 in the newspaper taz – die Tageszeitung. The dataset was created for and is described in the ACL 2025 Findings paper “taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades”. It provides a valuable resource for long-term analyses of language, public discourse, and media representation. The dataset is structured as a ZIP archive. Inside, folders are organised by year, each containing monthly JSON files with individual articles published that month. Every article includes metadata and textual content in the following format: meta data: "published_on": publishing date in the format YYYY-MM-DDThh:mm:ss+01:00 "contains_actors": boolean indicating whether person entities were detected "crawled_on": crawling date "language": always "de" "type": always "article" "author": article author "keywords": topic keywords "token_count": number of tokens text: "title": article title "teaser": short summary "text": main body of the article Each entry contains at least the main article text. The dataset supports a wide range of NLP and computational social science applications, especially for studying gender bias and representation in German media over time. If you use this dataset, please cite the following paper:Stefanie Urchs, Veronika Thurner, Matthias Aßenmacher, Christian Heumann, and Stephanie Thiemichen. 2025. taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades. In Findings of the Association for Computational Linguistics: ACL 2025, pages 10661–10671, Vienna, Austria. Association for Computational Linguistics.



