Racismmaff corpus
收藏资源简介:
This repository contains the derived linguistic dataset and metadata package for the complete RACISMMAFF corpus, designed to support research in corpus linguistics, critical discourse analysis, and Natural Language Processing (NLP). Corpus Content and ScopeThe RACISMMAFF corpus compiles and analyzes public discourse surrounding immigration and refugees from various global conflicts. It focuses on two core source types across both Spanish and English contexts:* Media Discourse: Opinion articles extracted from a diverse selection of major Spanish and British newspapers.* Political Discourse: Official speeches and statements from various political parties in Spain and the United Kingdom. To strictly comply with copyright restrictions and ensure a repo-safe design, this dataset does not redistribute the copyrighted source texts (such as full newspaper articles, headlines, speeches, or text excerpts). Instead, it provides comprehensive, non-reconstructive analytical data, tag frequencies, and positional metadata to ensure full transparency and research reproducibility without violating intellectual property rights. Corpus Structure and Technical DesignThe corpus is organized into distinct subcorpora representing different ideological or editorial profiles (e.g., Conservative, among others). For each subcorpus, the analytical data is provided both as a comprehensive Excel workbook (`.xlsx`) and as independent, UTF-8 encoded CSV files to facilitate seamless programmatic processing and scripting. Each subcorpus package contains the following structured components:* 01_corpus_profile: Document-level metadata, including derived token/word counts, publication details, and direct bibliographic URLs to the original web sources.* 02_tag_frequencies: Consolidated tag-instance counts aggregated by specific subcorpus and overall totals.* 03_summary_stats: High-level corpus summary statistics (e.g., total documents, total words, and total tag instances).* 04_tag_pointers: Positional character offsets (`marker_start_char`, `marker_end_char`), line numbers, and column numbers for annotated linguistic markers mapped against the source analysis strings, alongside original source URLs.* data_dictionary.csv: Column-level field descriptions and documentation for the entire database schema to ensure interoperability. A reproducible quality-check checklist has been successfully validated against all source materials to ensure structural and quantification integrity across all files and subcorpora. For detailed guidelines on schema definitions, tag formats, and replication steps, please refer to the included `README.txt` file.



