Corpus Data and R Analysis Code for a Register-Based Disambiguation Study of "Bias" in Contemporary English
收藏资源简介:
This deposit supports a corpus study of how the English noun "bias" partitions into distinct senses across registers. It comprises eight sub-corpus profiles, collocate and cluster tables for each sense, the sense-by-cue contingency matrix, and the inferential test summary, together with the R script that reproduces every table and statistic reported in the article. The materials cover seven sense-specific datasets (prejudice, preference, statistical, bowls, fabric, verbal, and political-legal) plus a register-free daily-usage benchmark, assembled from the english-corpora.org family and supplemented with United States federal court opinions. Collocates were extracted in AntConc 4.3.1 within an L5 to R5 window and ranked by Log-Likelihood, with Mutual Information as a complementary index; inferential testing was carried out in R 4.5.1. Concordance lines themselves are withheld: the source corpora are licensed material that may not be redistributed. The query settings needed to rebuild them are preserved in the kwic_info and corpus_kwic_info files included here.



