Background data for: Down-sampling strategies in corpus phonology
收藏资源简介:
This dataset records the result of a systematic review on the use of down-sampling practices in corpus phonology. It is associated with the following paper: Sönning, Lukas. 2025. Down-sampling strategies in corpus phonology. OSF Preprints. https://osf.io/preprints/osf/5hduq_v1 It consists of a UTF-8-encoded, tab-separated data table. A systematic review of down-sampling practices in corpus research To better understand how (corpus) linguists down-sample their data, I conducted a systematic review. The report pool includes n = 6,972 research articles published between 2012 and 2024 across 20 linguistic journals. The selection of journals, which are listed in the table below, aims to cover a broad range of subdisciplines and research styles. The following table lists the number of studies in the review by journal; journals are sorted by article count. Journal n Lingua 1,167 Applied Psycholinguistics 624 Applied Linguistics 464 World Englishes 457 Linguistics 451 Natural Language and Linguistic Theory 425 Language Learning 408 Studies in Second Language Acquisition 394 Linguistic Inquiry 386 English Language and Linguistics 295 Language 277 Cognitive Linguistics 255 International Journal of Corpus Linguistics 251 Journal of Sociolinguistics 224 Corpus Linguistics and Linguistic Theory 190 Language Variation and Change 177 Corpora 163 English World-Wide 154 Research in Corpus Linguistics 116 International Journal of Learner Corpus Research 94 Total 6,972 The database was stored locally and analyzed with the qualitative data analysis software MAXQDA 2022 (Verbi Software 2021). The first step was to identify all studies that used some form of down-sampling. To this end, I searched articles for a number of expressions that may appear in the methods section of relevant papers. The search terms I used are listed next: Search term Example hit(s) random[a-z]* sampl[a-z]* random sample, randomized samples random[a-z]* select[a-z]* randomly selected, random selection a subset of a subset of downsampl[a-z]* downsampled, downsampling down-sampl[a-z]* down-sampled, down-sampling down sampl[a-z]* down sampled, down sampling (thinned|thinning) thinned, thinning randomization randomization oversampling oversampling case-control case-control stratif[a-z]* stratified, stratification random[a-z]* extract[a-z]* randomly extracted, random extraction random[a-z]* instances random instances random[a-z]* concordances random concordances (tokens|cases|concordances) per tokens per, cases per, concordances per (a maximum of|at most|up to|no more than) [1-9]{1,} a maximum of 250, up to 30, no more than 5 (a maximum of|at most|up to|no more than) (two|three|four|five|six|seven|eight|nine|ten|eleven|twelve|twenty|thirty|fourty|fifty) a maximum of three, up to ten, no more than five This search narrowed down the pool to n = 3,295 articles, which were then screened for whether the study applied down-sampling; n = 236 studies did so. Out of these, n = 30 studies dealt with a phonological structure that was analyzed or coded by means of acoustic or auditory measurements. The current dataset concentrates on these 30 research articles. These were coded for the following variables, which correspond to columns in the data table: article article ID doi DOI for research article topic a short summary of the topic of the study structure the phonological structure of interest analysis whether an acoustic or auditory analysis was carried out (or both) clustering information about the grouping (or clustering) structure in the data; whether multiple instances/tokens are observed per speaker ("text") or lexical type/item ("item"), or both ("text, item") stratified whether stratified down-sampling was used ("yes" vs. "no"), which involves drawing separate down-samples from different sub-corpora or other subgroups in the data strat_by_corpus whether stratified down-sampling involved sampling from different corpora or sub-corpora ("yes" vs. "no"); coded "NA" if no use was made of stratified down-sampling strat_variable the variable(s) that was/were used for stratification; coded "NA" if no use was made of stratified down-sampling structured whether structured down-sampling was used, which relies on clustering variables (text/speaker and/or item) as a stratification variable ("yes" vs. "no"); if the data inlcuded two clustering variables, the relevant clustering variable is given in parentheses ("yes (text)", "yes (item)") max_tokens_speaker the maximum number of tokens per speaker; can be a range (e.g. "30-50"); coded "NA" if no use was made of structured down-sampling max_tokens_item_speaker the maximum number of instances of the same lexical item for a specific speaker; coded "NA" if no use was made of structured down-sampling References VERBI Software. 2021. MAXQDA 2022 [computer software]. Berlin, Germany: VERBI Software. Available from maxqda.com.



