Replication data for: Big data in Russian linguistics? Another look at paucal constructions
收藏资源简介:
This post contains a database of Russian numeral constructions from the RuTenTen corpus (https://www.sketchengine.co.uk/rutenten-russian-corpus/). The constructions are of the following type: paucal numeral (2, 3 or 4) followed by an adjective and a feminine noun. Abstract: With the advent of large web-based corpora, Russian linguistics steps into the era of “big data”. But how useful are large datasets in our field? What are the advantages? Which problems arise? The present study seeks to shed light on these questions based on an investigation of the Russian paucal construction in the RuTenTen corpus, a web-based corpus with more than ten billion words. The focus is on the choice between adjectives in the nominative (dve/tri/četyre starye knigi) and genitive (dve/tri/četyre staryx knigi) in paucal constructions with the numerals dve, tri or četyre and a feminine noun. Three generalizations emerge. First, the large RuTenTen dataset enables us to identify predictors that could not be explored in smaller corpora. In particular, it is shown that predicates, modifiers, prepositions and word-order affect the case of the adjective. Second, we identify situations where the RuTenTen data cannot be straightforwardly reconciled with findings from earlier studies or there appear to be discrepancies between different statistical models. In such cases, further research is called for. The effect of the numeral (dve, tri vs. četyre) and verbal government are relevant examples. Third, it is shown that adjectives in the nominative have more easily learnable predictors that cover larger classes of examples and show clearer preferences for the relevant case. It is therefore suggested that nominative adjectives have the potential to outcompete adjectives in the genitive over time. Although these three generalizations are valuable additions to our knowledge of Russian paucal constructions, three problems arise. Large internet-based corpora like the RuTenTen corpus (a) are not balanced, (b) involve a certain amount of “noise”, and (c) do not provide metadata. As a consequence of this, it is argued, it may be wise to exercise some caution with regard to conclusions based on “big data”.
本数据集收录了取自RuTenTen语料库(https://www.sketchengine.co.uk/rutenten-russian-corpus/)的俄语数词结构数据库。此类结构的格式为:少量数词(paucal numeral,即2、3或4)后接形容词与阴性名词。 摘要:随着大型网络语料库的问世,俄语语言学迈入了"大数据"时代。但在本研究领域中,大型数据集的价值几何?其优势何在?又会引发哪些问题?本研究以拥有超千亿词量的RuTenTen网络语料库为依托,通过对俄语少量数词结构的考察,试图解答上述疑问。 本研究聚焦于:当使用dve、tri或četyre与阴性名词构成少量数词结构时,形容词使用主格(nominative,如dve/tri/četyre starye knigi)与属格(genitive,如dve/tri/četyre staryx knigi)的选择问题。 本研究得出三点核心结论:其一,规模庞大的RuTenTen数据集使我们得以识别出小型语料库中无法探究的预测因子(predictor)。研究表明,谓语、修饰语、介词以及词序均会影响形容词的格形态选择。其二,本研究发现部分场景下,RuTenTen数据集无法与既往研究结论直接契合,或是不同统计模型间存在明显分歧,此时便需要开展进一步研究。数词(dve、tri与četyre的差异)以及动词支配的影响便是典型案例。其三,研究表明主格形容词拥有更易于学习的预测因子,此类因子可覆盖更多样的语料实例,且对特定格形态的偏好更为显著。据此可推测,随着时间推移,主格形容词或可逐渐取代属格形容词的使用场景。 尽管上述三点结论为俄语少量数词结构的相关研究提供了有价值的补充,但同时也暴露出三个问题:第一,RuTenTen这类网络语料库并不具备语料平衡性;第二,其中包含一定比例的"噪声"数据;第三,未提供元数据(metadata)。据此,有观点认为,基于"大数据"得出的研究结论需谨慎对待。



