Background data for: Case-control down-sampling in corpus research
收藏资源简介:
This dataset records the result of a systematic review on the use of case-control down-sampling in corpus research. It is associated with the following paper: Sönning, Lukas. 2025. Case-control down-sampling in corpus research. OSF Preprints. https://doi.org/10.31219/osf.io/g74bk_v1 It consists of a UTF-8-encoded, tab-separated data table. Case-control down-sampling In corpus-based work that (partly) relies on manual coding or disambiguation, the full list of tokens returned by a corpus query is often too large; the researcher is then forced to down-size the initial set of data. This is referred to as down-sampling, and there are various techniques and designs that can be applied (see Sönning 2024). One variant is case-control down-sampling, which uses information about the outcome (or dependent) variable to select a subset of instances from the larger pool. This form of down-sampling can be applied if the outcome variable is categorical (usually binary), and if its realization is readily ascertainable. The prototypical case-control design involves drawing the same number of instances of each level of the categorical response variable. This sampling design is particularly common in epidemiology and health-related sciences. A systematic review of down-sampling practices in corpus research To better understand how (corpus) linguists down-sample their data, I conducted a systematic review. The report pool includes n = 5,805 research articles published between 2012 and 2024 across 19 linguistic journals. The selection of journals, which are listed in the table below, aims to cover a broad range of subdisciplines and research styles. The following table lists the number of studies in the review by journal; journals are sorted by article count. Journal n Applied Psycholinguistics 624 Applied Linguistics 464 World Englishes 457 Linguistics 451 Natural Language and Linguistic Theory 425 Language Learning 408 Studies in Second Language Acquisition 394 Linguistic Inquiry 386 English Language and Linguistics 295 Language 277 Cognitive Linguistics 255 International Journal of Corpus Linguistics 251 Journal of Sociolinguistics 224 Corpus Linguistics and Linguistic Theory 190 Language Variation and Change 177 Corpora 163 English World-Wide 154 Research in Corpus Linguistics 116 International Journal of Learner Corpus Research 94 Total 5,805 The database was stored locally and analyzed with the qualitative data analysis software MAXQDA 2022 (Verbi Software 2021). The first step was to identify all studies that used some form of down-sampling. To this end, I searched articles for a number of expressions that may appear in the methods section of relevant papers. The search terms I used are listed next: Search term Example hit(s) random[a-z]* sampl[a-z]* random sample, randomized samples random[a-z]* select[a-z]* randomly selected, random selection a subset of a subset of downsampl[a-z]* downsampled, downsampling down-sampl[a-z]* down-sampled, down-sampling down sampl[a-z]* down sampled, down sampling (thinned|thinning) thinned, thinning randomization randomization oversampling oversampling case-control case-control stratif[a-z]* stratified, stratification random[a-z]* extract[a-z]* randomly extracted, random extraction random[a-z]* instances random instances random[a-z]* concordances random concordances (tokens|cases|concordances) per tokens per, cases per, concordances per (a maximum of|at most|up to|no more than) [1-9]{1,} a maximum of 250, up to 30, no more than 5 (a maximum of|at most|up to|no more than) (two|three|four|five|six|seven|eight|nine|ten|eleven|twelve|twenty|thirty|fourty|fifty) a maximum of three, up to ten, no more than five This search narrowed down the pool to n = 2,152 articles, which were then screened for whether the study applied down-sampling; n = 236 studies did so. Out of these, n = 25 studies applied a form of case-control down-sampling, and the current dataset concentrates on these 25 research articles. These were coded for the following variables, which correspond to columns in the data table: article article ID doi DOI for research article topic a short summary of the topic of the study unit_of_analysis which unit forms the basis of a statistical analysis; can be the individual instances/tokens/concordances returned by the corpus query ("token") or a higher-level unit such as the speaker or text file ("text") or a lexical type or item ("item") clustering information about the grouping (or clustering) structure in the data; whether multiple instances/tokens are observed per text file ("text") or lexical type/item ("item"), or both ("text, item") stratified whether stratified down-sampling was used ("yes" vs. "no"), which involves drawing separate down-samples from different sub-corpora or other subgroups in the data strat_by_corpus whether stratificatied down-sampling involved sampling from different corpora or sub-corpora ("yes" vs. "no"); coded "NA" if no use was made of stratified down-sampling strat_variable the variable(s) that was/were used for stratification; coded "NA" if no use was made of stratified down-sampling structured whether structured down-sampling was used, which relies on clustering variables (text and/or item) as a stratification variable; "no" indicates that this design feature was not used; if it was used, the clustering variable that informed the selection of instances is given in parentheses ("yes (text)", "yes (item)"); one study applied structured case-control down-sampling ("structured case-control") unit_of_downsampling whether down-sampling was executed at the level of the individual instances ("token") or whether higher-level units were targeted for down-sampling ("text", "item") outcome the outcome (or dependent) variable in the study language the language studied prop_less_common proportion of the less common variant (binary outcomes) or least common variant (ternary outcomes) analysis statistical procedure used for data analysis: "logistic regression", "recursive partitioning", "naive discriminative learning", "distinctive collexeme analysis", "chi-square tests", "behavioral profiles" notes additional notes on study design, data analysis, or interpretation References Sönning, Lukas. 2024. Down-sampling from hierarchically structured corpus data. International Journal of Corpus Linguistics 29(4). 507–533. doi: 10.1075/ijcl.23079.son VERBI Software. 2021. MAXQDA 2022 [computer software]. Berlin, Germany: VERBI Software. Available from maxqda.com.



