Spanish-PoliCorpus-2020: A Spanish Twitter Corpus for Author Profiling and Attribution
收藏资源简介:
Spanish-PoliCorpus-2020 is a Spanish Twitter corpus designed for author profiling and author attribution tasks in the political domain. The dataset includes pseudonymised author identifiers and is annotated with multiple author-level traits, such as gender, age range, and ideological orientation. Two evaluation scenarios are supported through independent data splits: author profiling and author attribution. This record provides a consolidated and FAIR-compliant public version of the dataset containing tweet identifiers and author-level annotations, while excluding any real user identifiers. Tweet text and additional derived representations are not included in the public release. The dataset was originally introduced in the associated scientific publication and is preserved here to ensure reproducibility and long-term reuse.



