The Basics of Corpus Collection for Research: Principles, Methods and Pitfalls
收藏资源简介:
In the domain of empirical language research, the corpus has emerged as an indispensable tool. A corpus is a large, structured, and principled collection of texts stored electronically, designed to represent a specific language variety, genre, or population (McEnery & Hardie, 2012). From lexicography and discourse analysis to second language acquisition and forensic linguistics, corpora provide researchers with quantifiable evidence of language patterns. However, the validity of any corpus-based study is fundamentally dependent on the quality and methodology of the corpus collection process. This article outlines the foundational principles of corpus collection, including defining research goals, ensuring representativeness and balance, addressing sampling strategies, handling spoken data, and navigating ethical and legal considerations. The first and most critical step in corpus collection is clearly articulating the research question. The design of a corpus—whether it is a general corpus, a specialized corpus, a diachronic corpus (tracking change over time), or a learner corpus—must align directly with the study's objectives (Biber, 1993). For instance, a researcher studying the evolution of scientific English would require a diachronic corpus with texts sampled from multiple decades, whereas a study on conversational fillers would need a spoken corpus of natural dialogue. Without a precise research goal, the corpus risks becoming a "convenience sample"—a haphazard collection of easily accessible texts that fails to support meaningful generalization. The research question thus determines every subsequent decision, from text selection to corpus size. Two interrelated concepts dominate discussions of corpus quality: representativeness and balance. Representativeness refers to the extent to which a corpus captures the full range of linguistic variation within a target population. Balance refers to the proportional inclusion of different text types or genres within the corpus (Biber, 1993). For example, a balanced corpus of general English should not contain 90% news articles and 10% fiction if fiction constitutes a larger share of everyday reading material. Achieving representativeness is challenging because language populations are often infinite and ill-defined. Corpus designers therefore rely on stratified sampling: dividing the target population into relevant strata (e.g., spoken vs. written, formal vs. informal, academic vs. popular) and then sampling proportionally from each stratum. As McEnery and Hardie (2012) note, "No corpus is ever fully representative of a language; it can only be representative of the sampling frame from which it was drawn" (p. 18). Researchers must also acknowledge the limitations of their sampling frame and avoid overgeneralizing findings.



