Background data for: Random forests in corpus research: A systematic review
收藏资源简介:
This dataset records the result of a systematic review on the use of random-forest models in corpus-based research. It is associated with the following paper: Sönning, Lukas. 2026. Random forests in corpus research: A systematic review. OSF Preprints. It consists of a UTF-8-encoded, tab-separated data table and a PDF file with a PRISMA flow chart (Page et al. 2021) outlining the study selection process. A systematic review of random-forest modeling in (corpus-)linguistic research To better understand how corpus linguists use random-forest models in their research, I conducted a systematic review. The report pool includes n = 6,972 research articles published between 2012 and 2024 across 20 linguistic journals. The selection of journals, which are listed in the table below, aims to cover a broad range of subdisciplines and research styles. The database was stored locally and analysed with the qualitative data analysis software MAXQDA 2022 (VERBI Software 2021). I used the search terms “random forest”, “random forests” and “random-forest” to identify relevant papers. This returned 130 articles, which were then screened for whether an RF model was used for data analysis; only 79 studies actually did so. The study selection process is outlined in the PRISMA flow chart (Page et al. 2021) that is part of this data post (file: fig02_PRISMA_2020_flow_diagram.pdf) The following table lists, for each journal, Total the total number of studies in the report pool, N the number of articles using a random-forest model for data analysis, and % the percentage of articles (N/Total*100) using a random-forest model Journals are sorted based on the right-most column (%). Journal Total N % International Journal of Learner Corpus Research 94 6 6.4% Language Variation and Change 177 11 6.2% Corpus Linguistics and Linguistic Theory 190 11 5.8% English World-Wide 154 6 3.9% International Journal of Corpus Linguistics 251 7 2.8% English Language and Linguistics 295 6 2.0% Corpora 181 3 1.7% Journal of Sociolinguistics 224 3 1.3% World Englishes 457 4 0.9% Linguistics 451 4 0.9% Studies in Second Language Acquisition 394 3 0.8% Lingua 1149 7 0.6% Language Learning 408 2 0.5% Applied Psycholinguistics 624 3 0.5% Applied Linguistics 464 2 0.4% Cognitive Linguistics 255 1 0.4% Language 277 1 0.4% Linguistic Inquiry 386 0 0.0% Natural Language and Linguistic Theory 425 0 0.0% Research in Corpus Linguistics 116 0 0.0% Total 6,972 Within this pool of n = 79 articles, I then identified all RF models that were reported, which yielded a list of 147 items. These 147 models were then coded for a range of features, which are listed below. Given the focus of the current review on corpus-based research, I screened the candidate models for the type of data underlying the analysis; 21 models were excluded because they did not rely on corpus data. I adopted a relatively broad notion of “corpus-based”: An RF model qualified for inclusion if the method of data collection classified as observational, naturalistic and/or uncontrolled; most notably, sociolinguistic interviews were considered “corpus data”. One further RF model was excluded because its application was – in my view – nonsensical. This yields a final pool of 125 RF models drawn from 69 research articles. While most studies reported a single analysis (n = 46, 67%), several papers included multiple RF models. article_id article identifier doi persistent digital object identifier journal name of journal year of publication month month of publication exclude whether the RF model should be excluded from this review (because model was nensensical) corpus_data whether the data are drawn from a corpus other_methods other methods that were used to analyze the data set software software used to fit the RF model package R package used to fit the RF model purpose whether the random forest model served as the "primary" means of answering the research question, or whther it played an "ancillary" role n_pred number of predictor variables in the forest n_pred_pars number of predictor parameters in an equivalent main-effects-only regression model clust_var_levels if a clustering variable was included as a predictor in the model, number of levels n_terms number of terms in the forest; differs from "n_pred" if interactions are manually coded and included mtry setting of the mtry parameter, which specifies the number of candidate variables (or terms) to sample at each splitting occasion mtry_tuned whether the choice of mtry was based on some type of optimality criterion n_trees number of trees in the random forest n_obs number of observations (i.e. tokens/data points) in the dataset var_importance the type of variable importance measure used ("standard", "conditional", "unclear") interpretation list of strategies used for model interpretation; see [1] below interactions_incl whether interaction between predictors were manually specified and included tree_type classification tree (categorical response variable) vs. regression tree (continuous response variable) n_levels for classification trees (categorical outcomes), number of response categories ordinal for categorical outcomes with three or more levels, whether these are ordered (i.e. an ordinal response variable) alternation whether the outcome variable can be considered as representing some sort of alternation outcome_variable short description of the outcome variable clust_vars whether clustering variables are present in the data, and if so, which ones (speaker/text/item) clust_vars_aggr clustering variables aggregated into "test" (speaker, author, text) and "item" (word, lexeme, lexical item, etc.) clust_vars_incl whether (and if so, which) clustering variables (e.g. speaker/text/item) were included in the forest clust_vars_incl_aggr whether (and if so, which) clustering variables (aggregated label) were included in the forest clusterlevel_preds whether the forest included cluster-level variables alongside the clustering variable clust_vars_info information about the relative importance of clustering and cluster-level variables baseline for classification forests, the proportion of the most common class/level (benchmark for prediction accuracy) type_pred_specified whether the type of prediction (ordinary vs. out-of-bag) underlying the assessment of predictive performance was specified pred_accuracy for classification forests, the prediction accuracy (proportion of correctly classified tokens) reported, without indication of the type (training data/ordinary vs. test data/out-of-bag) pred_accuracy_oob for classification forests, the out-of-bag prediction accuracy (proportion of correctly classified tokens) c_index concordance index or, equivalently, area under the curve (AUC) reported for the model, without indication of the type (training data/ordinary vs. test data/out-of-bag) c_index_oob out-of-bag concordance index or, equivalently, area under the curve (AUC) reported for the model [1] "varimp" = variable importance scores"tree" = representative tree"pdp" = partial dependence plots"preds" = predictions from the RF model"gsm" = global surrogate model"predictive performance" = model was used for genuine prediction task) References Page, Matthew J, Joanne E. McKenzie, Joanne E. Bossuyt, Isabelle Boutron, Tammy C. Hoffmann, Cynthia D. Mulrow, Larissa Shamseer, Jennifer M. Tetzlaff, Elie A. Akl, Sue E. Brennan, Roger Chou, Julie Glanville, Jeremy M. Grimshaw, Asbjørn Hróbjartsson, Manoj M. Lalu, Tianjing Li, Elizabeth W. Loder, Evan Mayo-Wilson, Steve McDonald, Luke A. McGuinness, Lesley A. Stewart, James Thomas, Andrea C. Tricco, Vivian A. Welch, Penny Whiting & David Moher. 2021. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. British Medical Journal 372: n71. https://doi.org/10.1136/bmj.n71 VERBI Software. 2021. MAXQDA 2022 [computer software]. Berlin, Germany: VERBI Software. Available from maxqda.com.



