Background data for: Ordinal random forests in language data analysis
收藏资源简介:
This dataset is associated with a methodological study on the use of ordinal random-forest models for language data analysis: Mühlbauer, Michael & Lukas Sönning. 2025. Ordinal random forests in language data analysis. OSF Preprint. https://doi.org/10.31234/osf.io/s4bv2_v1 Description This dataset contains information on the usage preferences of speakers of Maltese English with regard to 63 pairs of lexical expressions. These pairs (e.g. truck-lorry or realization-realisation) are known to differ in usage between BrE and AmE (cf. Algeo 2006). The data were elicited with a questionnaire that asks informants to indicate whether they always use one of the two variants, prefer one over the other, have no preference, or do not use either expression (see Krug and Sell 2013 for methodological details). Usage preferences were therefore measured on a symmetric 5-point ordinal scale. Data were collected between 2008 to 2023, as part of a larger research project on lexical and grammatical variation in varieties of English. The current dataset, which informs a methodological study on modeling strategies for ordinal data, include 1,386 speakers and 78,364 ratings in total. It also provides information about a number of socio-demographic variables, including gender, year of birth, age, highest qualification, and the language(s) used at home while growing up. Questionnaire administration The questionnaire used for data elicitation, which is included in this post ("lexical_questionnaire.pdf"), was administered in printed form and filled in with a pen. Data were mostly gathered by German university students, who approached (potential) informants and asked if they were willing to take part in an anonymous survey that would be used for scientific purposes only; no financial compensation was offered. Once informants gave their verbal consent to participate in the study, they would either fill in the questionnaire on their own, or they were assisted by the data collector. Participants were predominantly recruited in the streets of Malta, i.e. in various public places (e.g. parks, busses, university campus, etc.) as part of several field trips (between 2008 and 2023), each in connection with a university-level seminar on the varieties of English spoken and written in Malta. Prior to data acquisition, students received basic instructions on how to avoid (or handle) potentially problematic sitations, e.g. ensuring that respondents understood the task and the relevant meaning of lexical expressions (see "Explanation/comment" column in questionnaire), allowing informants to withhold sensitive biographical information, etc. Due to the fact that questionnaires were also administered to university students, either in a lecture hall or on campus, the age distribution of participants in the dataset is uneven, with a notable peak at around 20 years of age. Data and file overview malta_lexical_data_mixfabOF.tsv tab-separated data table containing the ratings provided by the 1,386 informants item_labels.tsv tab-separated data table with information about the full set of 68 item pairs in the questionnaire lexical_questionnaire.pdf questionnaire used for data elicitation Data-specific information for: item_labels.tsv UTF-8-encoded, tab-separated data table with 69 rows and 5 columns. The columns represent the following variables: item item pair ID (links to file "malta_lexical_data_mixfabOF.tsv", which includes the rating data) AmE_variant the more/exclusively/traditionally American English variant BrE_variant the more/exclusively/traditionally British English variant AmE_short short(er) label for American English variant, used for plotting BrE_short short(er) label for British English variant, used for plotting Data-specific information for: malta_lexical_data_mixfabOF.tsv UTF-8-encoded, tab-separated data table with 78,365 rows and 11 columns. The columns represent the following variables: id anonymized participant ID year year of data collection age age of participant in years gender self-reported gender of participant ("Female", "Male") highest_qualification highest level of qualification (completed or ongoing) "Higher Education" "No/Other Qualification" "Professional/Vocational" "Secondary Education" "Work-Based Learning" language_home language(s) used at home while growing up, where "(Other)" indicates an optional third language "English (Other)" "English, Maltese (Other)" "Maltese, English (Other)" "Maltese (Other)" ratio proportion of life spent in Malta item item pair ID rating numeric version of the rating provided by the participant +2 I always use this (British English) expression +1 I use this (British English) expression more often 0 I have no preference -1 I use this (American English) expression more often -2 I always use this (American English) expression date_of_birth : year of birth rating_factor second numeric version of the rating provided by the participant (for straightforward conversion to a factor) 5 I always use this (British English) expression 4 I use this (British English) expression more often 3 I have no preference 2 I use this (American English) expression more often 1 I always use this (American English) expression References Algeo, John. 2006. British or American English: A handbook of word and grammar patterns. Cambridge: Cambridge University Press. Krug, Manfred & Katrin Sell. 2013. Designing and conducting interviews and questionnaires. In Manfred Krug & Julia Schlüter (eds.), Research methods in language variation and change, 69–98. Cambridge: Cambridge University Press.



