Background data (adapted from Jenset & McGillivray 2017) for: Down-sampling from hierarchically structured corpus data
收藏资源简介:
Dataset description This dataset, which is adapted from Jenset and McGillivray (2017), contains tabular files documenting the alternating usage of -(e)th and -(e)s to mark third-person verb inflection in Early Modern English. The data provided by Jenset and McGillivray (2017) are drawn from the PPCEME corpus (Kroch et al. 2004) and cover the period from 1500 to 1700. In total, 13,757 third-person singular tokens (excluding the verb BE) were annotated by these authors for a range of variables. For the purposes of the present methodological study, this dataset was reduced to a subset of 11,645 tokens, and the coding of variables was in some parts revised, completed, or modified. The dataset includes information about the Author and Verb Lemma, as well as a number of predictor variables, including Genre, Year, Frequency (of the verb lemma in the third-person singular), Phonological Context (stem-final sound), and the Gender of the author. Abstract for related publication Resource constraints often force researchers to down-size the list of tokens returned by a corpus query. This paper sketches a methodology for down-sampling and offers a survey of current practices. We build on earlier work and extend the evaluation of down-sampling designs to settings where tokens are clustered by text file and lexeme. Our case study deals with third-person present-tense verb inflection in Early Modern English and focuses on five predictors: Year, Gender, Genre, Frequency, and Phonological Context. We evaluate two strategies for selecting 2,000 (out of 11,645) tokens: simple down-sampling, where each hit has the same selection probability; and structured down-sampling, where this probability is inversely proportional to the author- and verb-specific token count. We form 500 sub-samples using each scheme and compare regression results to a reference model fit to the full set of cases. We observe that structured down-sampling shows better performance on several evaluation criteria.
数据集说明 本数据集改编自Jenset与McGillivray(2017)的研究成果,收录了记录近代早期英语中第三人称动词屈折变化(third-person verb inflection)时-(e)th与-(e)s交替使用情况的表格文件。Jenset与McGillivray(2017)所提供的数据源自PPCEME语料库(PPCEME corpus,Kroch et al. 2004),覆盖1500年至1700年的语料时段。两位作者共标注了13757个第三人称单数Token(Token,不含动词BE),涵盖多类变量。为适配本项方法论研究,本数据集被裁剪为包含11645个Token的子集,并对部分变量的编码进行了修订、补充与调整。本数据集包含作者、动词词元(Verb Lemma)等信息,同时涵盖若干预测变量:体裁(Genre)、年份、第三人称单数动词词元的使用频率、语音语境(词尾音素)以及作者性别。 相关研究论文摘要 资源限制常迫使研究者缩减语料库查询返回的Token列表规模。本文阐述了一种下采样(down-sampling)方法论,并对当前的通用实践进行了综述。本文基于前期研究成果,将下采样设计的评估范围拓展至Token按文本文件与词位(lexeme)聚类的场景。本案例研究聚焦近代早期英语的第三人称现在时动词屈折变化,选取年份、性别、体裁、频率与语音语境五类预测变量展开分析。我们评估了两种从11645个Token中选取2000个Token的策略:一是简单下采样,即每个Token被选中的概率均等;二是结构化下采样,即选中概率与作者及动词对应的Token数量成反比。我们分别采用两种方案生成500个子样本,并将回归结果与拟合全量样本的参考模型进行对比。结果显示,结构化下采样在多项评估指标上表现更优。



