遇见数据集

The dative dataset of World Englishes

收藏
Zenodo2022-09-05 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>This dataset is distributed under a Creative Commons Attribution Non Commercial 4.0 International license. Use for research purposes only!</strong> The dataset contains 13,171 variable double-object and prepositional datives extracted from the International Corpus of English series and the Corpus of Global web-based English sampling from nine national varieties of English: British English Canadian English New Zealand English Irish English Hong Kong English Philippine English Singapore English Indian English Jamaican English <strong>The dataframe contains the following columns:</strong> <strong>1 TokenID</strong>: Unique identifier for the individual token <strong>2 Variety</strong>: The variety from which the token is taken <strong>3 Nativity</strong>: Native or non-native variety of English (L1 vs. L2) <strong>4 Corpus</strong>: The corpus from which the token stems <strong>5 Subcorpus</strong>: Combination of Corpus and Variety <strong>6 FileID</strong>: ID of the corpus file in which the token was found. Format: VARIETY:FILENAME <strong>7 TextID</strong>: ID of the corpus text in which the token was found. Individual files in ICE can have multiple texts. Format: VARIETY:FILENAME:TEXTNUMBER <strong>8 LineID</strong>: ID of the line in the text in which the token sentence was found. Format: VARIETY:FILENAME:TEXTNUMBER:LINENUMBER <strong>9 SpeakerID</strong>: ID of the speaker of the sentence. Speakers in spoken texts are indicated with capital letters. Authors of written texts have ID ‘A’. Format: VARIETY:FILENAME:TEXTNUMBER:SPEAKERID <strong>10 UnitMarker</strong>: UnitMarker of the utterance in the text. Format UTTERANCE NUMBER:TEXTNUMBER:SPEAKERID <strong>11 GenreFine</strong>: 14-level distinction: The 12-level ICE sub-register in which the token was found and the two levels in GloWbE (blog vs. general). Levels: See ICE documentation <strong>12 GenreCoarse</strong>: 5-level distinction: The 4-level ICE register in which the token was found and GloWbE (online = 1 level). Levels: See ICE documentation <strong>13 Mode</strong>: The mode (‘spoken’, ‘written’) of the token. <strong>14 Register</strong>: The 4-level Register along two axes – spoken vs. written / informal vs. formal <strong>15 PriorContextPlain</strong>: The plain text version of the 100 words preceding the dative token. <strong>16 PriorContextTag</strong>: The POS-tagged version of the 100 words preceding the dative token. <strong>17 SentencePlain</strong>: The plain text version of the sentence containing the dative token. <strong>18 SentenceTag</strong>: The POS-tagged version of the sentence containing the dative token. <strong>19 WholeConstructionPlain</strong>: The plain text version of the VP containing the dative token (i.e. verb + object + object). <strong>20 WholeConstructionTag</strong>: The POS-tagged version of the VP containing the dative token. <strong>21 Verb</strong>: The lemma of the verbal head (<em>give </em>in <em>gave it some thought</em>) <strong>22 VerbForm</strong>: The verb form of the verbal head (<em>gave </em>in <em>gave it some thought</em>) <strong>23 RecipientShort</strong>: The short plain text version of the recipient without hesitations or repetitions <strong>24 ThemeShort</strong>: The short plain text version of the theme without hesitations or repetitions <strong>25 RecipientLong</strong>: The long plain text version of the recipient with hesitations or repetitions <strong>26 ThemeLong</strong>: The long plain text version of the theme with hesitations or repetitions <strong>27 RecHeadPlain</strong>: The plain text version of the recipient head <strong>28 RecHeadTag</strong>: The POS-tagged version of the recipient head <strong>29 RecHeadLemma</strong>: The lemma of the recipient head <strong>30 ThemeHeadPlain</strong>: The plain text version of the theme head <strong>31 ThemeHeadTag</strong>: The POS-tagged version of the theme head <strong>32 ThemeHeadLemma</strong>: The lemma of the theme head <strong>33 VerbThemeLemma</strong>: Combination of the verb lemma and the theme head. Format: VERB_THEME <strong>34 VerbSense</strong>: Semantics of the verb based on the whole construction combined with the verb lemma. Format: VERB.VERBSEMANTICS <strong>35 VerbSemantics</strong>: 5-level distinction of verb semantics (‘a’, ‘t’, ‘p’, ‘f’, ‘c’). <strong>36 Resp</strong>: The variant order. Levels: ‘do’ (=ditransitive), ‘pd’ (=prepositional) <strong>37 RecAnimacy</strong>: 6-level distinction of recipient animacy following previous research: human (a1) &gt; animal (a2) &gt; collective (c) &gt; locative (l) &gt; temporal (t) &gt; inanimate (i) <strong>38 ThemeAnimacy</strong>: 6-level distinction of theme animacy following previous research: human (a1) &gt; animal (a2) &gt; collective (c) &gt; locative (l) &gt; temporal (t) &gt; inanimate (i) <strong>39 RecWordLth</strong>: Length of recipient NP in words <strong>40 RecLetterLth</strong>: Length of recipient NP in orthographic characters <strong>41 ThemeWordLth</strong>: Length of theme NP in words <strong>42 ThemeLetterLth</strong>: Length of theme NP in orthographic characters <strong>43 RecComplexity </strong> 15-level distinction of recipient complexity indicating type and number of posthead dependents, restricted to the ICE components. (GloWbE components make simplified distinction between ‘simple’ and ‘complex’). Levels: ‘s’ = simple (no postmodifications), ‘co’ = coordinated, ‘ge’ = general extender, ‘gn’ = genitive, ‘postad’ = postmodifying adverbial/adjective, ‘pp’ = modifying prepositional phrase, ‘appnom’ = nominal apposition, ‘rc’ = relative clause, ‘cp’ = complement clause, ‘advc’ = adverbial clause, ‘nonfin’ = nonfinite clause, ‘tpp’ = two nominal posthead dependents, ‘tvp’ = two posthead dependents involving at least one VP, ‘mpp’ = more than two nominal posthead dependents, ‘mvp’ = more than two posthead dependents involving at least one VP <strong>44 ThemeComplexity </strong> 15-level distinction of theme complexity indicating type and number of posthead dependents, restricted to the ICE components. (GloWbE components make simplified distinction between ‘simple’ and ‘complex’). Levels: ‘s’ = simple (no postmodifications), ‘co’ = coordinated, ‘ge’ = general extender, ‘gn’ = genitive, ‘postad’ = postmodifying adverbial/adjective, ‘pp’ = modifying prepositional phrase, ‘appnom’ = nominal apposition, ‘rc’ = relative clause, ‘cp’ = complement clause, ‘advc’ = adverbial clause, ‘nonfin’ = nonfinite clause, ‘tpp’ = two nominal posthead dependents, ‘tvp’ = two posthead dependents involving at least one VP, ‘mpp’ = more than two nominal posthead dependents, ‘mvp’ = more than two posthead dependents involving at least one VP <strong>45 RecNPExprType</strong>: Syntactic category of the recipient NP Levels: ‘dem’ = bare demonstrative; ‘nc’ = common noun; ‘np’ = proper noun; ‘pprn’ = personal pronoun; ‘iprn’ = impersonal pronoun; ‘rprn’ = reflexive pronoun; ‘vp’ = gerund (-ing) NP; ‘wh’ = NP headed by wh- word <strong>46 ThemeNPExprType</strong>: Syntactic category of the theme NP Levels: ‘dem’ = bare demonstrative; ‘nc’ = common noun; ‘np’ = proper noun; ‘pprn’ = personal pronoun; ‘iprn’ = impersonal pronoun; ‘rprn’ = reflexive pronoun; ‘vp’ = gerund (-ing) NP; ‘wh’ = NP headed by wh- word <strong>47 RecGivenness</strong>: Givenness of the recipient NP. Levels: ‘given’, ‘new’ <strong>48 ThemeGivenness</strong>: Givenness of the theme NP. Levels: ‘given’, ‘new’ <strong>49 RecDefiniteness</strong>: Definiteness of the recipient NP. Levels: ‘def’, ‘indef’ <strong>50 ThemeDefiniteness</strong>: Definiteness of the theme NP. Levels: ‘def’, ‘indef’ <strong>51 RecBinComplexity</strong>: Binary predictor of recipient complexity indicating following postmodifications after the head noun. Levels: ‘simple’, ‘complex’ <strong>52 ThemeBinComplexity</strong>: Binary predictor of theme complexity indicating following postmodifications after the head noun. Levels: ‘simple’, ‘complex’ <strong>53 RecPerson</strong>: Person of recipient. Levels: ‘local’, ‘non-local’ <strong>54 ThemeConcreteness</strong>: Concreteness of theme based on verb semantics. Levels: ‘concrete’, ‘non-concrete’ <strong>55 TypeTokenRatio</strong>: Type-token ratio of the 100 word context surrounding the token <strong>56 RecHeadFreq</strong>: Frequency of recipient head lemma in GloWbE <strong>57 ThemeHeadFreq</strong>: Frequency of theme head lemma in GloWbE <strong>58 RecThematicity</strong>: Normalized frequency of recipient head lemma in its text (per 2000 words) <strong>59 ThemeThematicity</strong>: Normalized frequency of theme head lemma in its text (per 2000 words) <strong>60 PrimeType</strong>: The response type of the preceding dative token, if any. Levels: ‘do, ‘pd, ‘NA’ <strong>61 Persistence</strong>: Indicates whether preceding dative token, if any, is the same or not. Levels: ‘none’, ‘yes’, ‘no’ <strong>62 SameUtterance</strong>: Indicates whether the preceding dative token occurred in the same utterance or not. Necessary for manual coding of persistence. <strong>63 DistanceToPrevious</strong>: Number of utterances between current and preceding dative token. ‘None’ if no preceding dative token. <strong>64 RecPron</strong>: Binary factor of recipient pronominality. Levels: ‘pron’, ‘non-pron’ <strong>65 ThemePron</strong>: Binary factor of theme pronominality: Levels: ‘pron’, ‘non-pron’ <strong>66 RecBinAnimacy</strong>: Binary factor of recipient animacy. Levels of RecAnimacy conflated to: ‘animate’, ‘inanimate’ <strong>67 ThemeBinAnimacy</strong>: Binary factor of theme animacy. Levels of ThemeAnimacy conflated to: ‘animate’, ‘inanimate’ <strong>68 logRecLetterLth</strong>: Natural logarithm of recipient length in orthographic characters <strong>69 logThemeLetterLth</strong>: Natural logarithm of theme length in orthographic characters <strong>70 WeightRatio</strong>: Ratio of object lengths: Recipient length in characters divided by theme length in characters <strong>71 logWeightRatio</strong>: Natural logarithm of weight ratio <strong>72 PrimeTypePruned</strong>: The variant of the preceding dative token within the previous 10 utterances. Levels: ‘none’, ‘do’, ‘pd’ <strong>73 NumDistanceToPrevious</strong>: Numeric distance to previous token (for calculations in R) <strong>74 PersistencePruned</strong>: Indicates whether the preceding token within the previous 10 utterances is the same as the current token. Levels: ‘none’, ‘yes’, ‘no’ <strong>75-82 z.__________</strong>: Numeric predictor centered around the mean and scaled by two standard deviations. <strong>83 Variety.Sum</strong>: Column used for sum coding in modeling process

提供机构:
Zenodo
创建时间:
2020-10-20
二维码
社区交流群
二维码
科研交流群
商业服务