遇见数据集

Replication Data for: The semantic structuring of minimizing constructions in present-day Netherlandic Dutch: a distribution-based cluster analysis

收藏
DataONE2025-09-02 更新2025-09-06 收录
官方服务:

资源简介:

Dataset abstract: This dataset contains the data files that were used for the cluster analysis of the Dutch minimizing construction, as described in the publication cited below. In addition to a ReadMe file, it contains three files: A txt file is provided with the corpus queries that were used to find tokens of the minimizing constructions in the Dutch Web 2014 (nlTenTen14) corpus, available via Sketch Engine (more information about the TenTen corpora: Jakubíček, M., A. Kilgarriff, V. Kovář, P. Rychlý & V. Suchomel (2013). The TenTen corpus family. In: 7th International Corpus Linguistics Conference CL. Lancaster, 125–127). A csv file is provided that forms the input file for the cluster analysis. It contains a list of 5,863 minimizer-predicate combinations, more specifically a list of the predicates that are combined with the minimizers that have a token frequency of at least 10 in my dataset. An R-script is provided with the code to perform the cluster analysis in R. Article abstract: This paper examines the semantic structuring of a paradigm of 89 minimizers, i.e., nouns that reinforce sentential negation in present-day Netherlandic Dutch, such as meter ‘meter’ in voor geen meter vertrouwen ‘not to trust for a meter’. Cosine distances are computed on the basis of the predicates the minimizers combine with in a sample of 100 tokens downloaded from the Dutch Web corpus 2014 (nlTenTen14) and clustered according to the Partitioning Around Medoids (PAM) algorithm into nine semantic clusters. The clusters largely correspond to semantic categories such as taboo terms or units of money. This suggests that, in general, minimizers belonging to the same semantic domain are combined with a similar (core) set of predicates. Based on the shared predicates per cluster, we detect signs of analogical attraction between minimizers or, conversely, competition. Crucially, low silhouette widths enable us to identify outliers in their respective clusters, for instance, minimizing nouns that exhibit signs of context expansion, as shown by their combination with semantically non-harmonious verbs. As such, this paper provides a synchronic snapshot of the semantic processes involved in (incipient) grammaticalization of minimizing nouns and, more in general, it illustrates how distributional semantics offers a heuristic to analyze the structure of a network of comparable micro-constructions.

数据集摘要:本数据集包含下述引用文献中提及的荷兰语弱化构式(minimizing construction)聚类分析所用的全部数据文件。除README文件外,本数据集共包含三类文件:一份文本文件:内含用于在Sketch Engine平台获取的荷兰语Web 2014语料库(nlTenTen14)中检索弱化构式Token(token)的语料库查询语句。有关TenTen语料族的更多信息可参考:Jakubíček, M., A. Kilgarriff, V. Kovář, P. Rychlý 与 V. Suchomel (2013). TenTen语料族. 载于:第7届国际语料库语言学会议(CL)论文集. 兰卡斯特,125–127页。一份逗号分隔值(CSV)文件:作为聚类分析的输入文件,内含5863组弱化词(minimizer)-谓词搭配列表,具体而言,该列表收录了与在本数据集中Token频次不低于10次的弱化词所搭配的全部谓词。一份R脚本文件:内含用于在R语言环境中执行聚类分析的代码。 论文摘要:本文对现代荷兰语中89个弱化词(minimizer)的范式语义结构展开研究,此类弱化词为用于强化语句否定意义的名词,例如短语‘voor geen meter vertrouwen’(意为“丝毫不信任”)中使用的‘meter’(意为“米”)。本文基于从荷兰语Web 2014语料库(nlTenTen14)中提取的100个Token样本内弱化词所搭配的谓词计算余弦距离,并通过围绕中心点划分(Partitioning Around Medoids, PAM)算法将其聚类为9个语义簇。这些语义簇大体对应于特定语义范畴,例如禁忌语或货币单位。这表明,总体而言,隶属于同一语义域的弱化词会与一组相似的(核心)谓词搭配。基于各语义簇中共有的谓词,本文检测出了弱化词之间存在类比吸引或竞争关系的迹象。尤为关键的是,较低的轮廓宽度(silhouette width)可帮助我们识别各语义簇中的离群值,例如那些呈现出语境扩张迹象的弱化名词——这类名词会与语义上不协调的动词搭配,可作为佐证。据此,本文同步呈现了弱化名词(处于萌芽阶段的)语法化过程所涉及的语义机制快照,且总体而言,本文阐释了分布语义学如何为分析同类微观构式网络的结构提供启发式方法。

创建时间:
2025-09-03
二维码
社区交流群
二维码
科研交流群
商业服务