遇见数据集

Annotated database of nominalization pairs in English

收藏
Zenodo2025-07-12 更新2026-05-26 收录
官方服务:

资源简介:

Annotated database of nominalization pairs in English The dataset contains annotations of (arguably) deverbal nominalizations in English, based on 3,440 concordance lines extracted from the British Web Corpus (ukWaC) (Baroni et al. 2009). The nominalizations in question include 40 lemmas: 20 native -ing nominalizations and 20 Latinate -(t)ion nominalizations. These nominalizations were selected to form 20 pairs of lemmas, each consisting of one -ing nominalization and one -(t)ion nominalization derived from the same root. The database is part of a study aimed at determining the extent to which eventive deverbal nominalizations in English and Serbian exhibit verbal properties. The study focused on two key indicators of verbal properties: the presence of a genitive complement and the presence of plural forms of nominalizations. Selecting the pairs and testing the ‘process’ reading The selection process began with over 20 pairs of nominalizations per category, chosen ad hoc by the first author, with the additional criterion that each nominalization had at least 1,000 corpus hits. The categories comprised pairs consisting of nominalizations in -ing and their corresponding -(t)ion counterparts. Since lemmatization of -ing nominalisations proved to be unreliable, the items were selected based on token ratios. A table was created in Google Spreadsheets containing information on token frequency, extracted from the corpus through a dedicated search for each nominalisation. For example, the entry for the nominalisation communicating shows a total token frequency of 17,179. To ensure that each included nominalization can be interpreted as a process, a Corpus Query Language (CQL) search was performed using the pattern [lemma="process"][lemma="of"][word="x"], where x represents the specific nominalization. Pairs that did not yield any corpus hits were excluded from further analysis. Since, after the application of the process test, more than 20 pairs were still available per category, the pairs had to be ranked to select the most appropriate ones. Ranking the pairs of nominalizations The final selection was based on the token frequency ratios of the nominalizations within each pair. For example, in the pair consisting of an -ing nominalization, e.g. investigating and the corresponding -(t)ion nominalization, e.g. investigation, the token frequency of investigating (31,644) was divided by the token frequency of investigation (91,822), yielding a ratio of 0.34. The 20 pairs with ratios closest to 1 were selected to ensure the most balanced pairs in terms of frequency. Extracting concordances from the corpus The final dataset contains 2000 concordances, i.e. 50 valid example sentences per nominalization. Sentences found in titles and subtitles, as well as duplicate sentences, were excluded due to their different argument structure patterns. Consequently, more than 50 concordances were initially extracted for each nominalization using dedicated CQL queries to ensure that 50 valid examples remained for analysis. Annotation process and criteria The concordances were annotated for several linguistic features, including whether the nominalization is used in a sentence, whether it is pluralized, the presence of an of-genitive complement, the presence of a possessive, the presence of a clausal complement, the presence of a free relative clause as its complement, the presence of an internal argument in the accusative case (1), the presence of an internal argument realized as a compound (e.g. data extraction), clausal gerunds with overt subjects (2) and the presence of an ECM-like structure (3). The presence of each structure is marked with 1, and its absence is marked with 0. (1) These processes of selection of 'photo-opportunities ', even within the constraints established in the guidelines to the children, serves in the same way as children 's selection of particular representation forms in drawing, to provide some preliminary indication of secondary artifacts the children draw upon in communicating their relationships with technology. (2) The conceit here is almost as disingenuous as Val Gielgud fabricating letters to represent false opinion on BBC programmes. (3) Click here to see Jodie investigating her Crinkle Bag. Gerund -ing forms used for building progressive tenses (e.g. was communicating), or functioning as adverbials in a sentence, as preparing in (4), as well as -(t)ion forms as modifiers of the nouns (e.g. communication in communication strategies) were not taken into account. (4) As we walk onto the dock in Simonstown, preparing to board the boat that will take us on our pelagic trip, Roland and Eric are astonished to see Martyn Sidwell [...] While the initial selection criterion required at least 1,000 corpus hits per nominalization, an exception was made for the pair participating/participation. Since participating was a clear outlier in the dataset, being the only -ing nominalization in the dataset that could not license an accusative complement, a free relative clause, or a clausal complement, it was replaced with fabricating/fabrication, the next available pair in the ranking, despite fabricating having 614 corpus hits. Column-by-column explanation Column A (nominalization_type) indicates the type of nominalization in question based on the derivational suffix. ing - -ing nominalization (e.g. preparing derived from to prepare) (t)ion - nominalization ending in -(t)ion (e.g. preparation derived from to prepare) Column B (root) contains the root shared by the nominalizations within the pair (e.g. prepar- in the pair preparing/preparation). Column C (nominalisation) contains the nominalization in question (e.g preparing). Column D (concordance) contains the concordance in which the particular nominalization appears. E.g., (5) is a concordance that contains the nominalization preparing. (5) Such formal relationships were not so popular in Catholic schools, although Catholic principals often stressed the relatively greater value of informal contacts with parents outside church on Sundays or in the course of preparing children for the sacraments. Column E (sentence) indicates whether the context in which the nominalization in question is used constitutes a full sentence. Titles and subtitles were not considered as full sentences. For example, (6) received the value 0, as ‘preparing’ in this case does not represent a nominalization, but a gerund -ing form used for building Past Progressive Tense. (6) It seems he was preparing them for war sooner rather than later. The value 0 was also assigned to sentences that were repetitions of already annotated sentences. On the other hand, (7) received the value 1. (7) We tend to think and assume that communicating with one another is relately easy and straight forward. Column F (plural) indicates whether the nominalization in question is in the plural. For example, (8) received the value 1. (8) Once there was a hatch to be opened near where he was stationed; he watched the preparations for a second or so suspiciously, and then [...] Column G (of-argument) indicates the presence of an of-argument. Both external and internal genitive arguments are taken into account. For example, (9) received the value 1. (9) Normally, best results are obtained by the preparation of a tincture OR add a pinch to your dogs food and mix in well. Column H (possessive) indicates the presence or absence of a possessive pronoun, or a possessive adjective (‘s) that can be interpreted as an argument of the nominalization in question. For example, (10) received the value 1. (10) They are making nothing by me, ' was another of his observations ; 'they 're making something by that fellow.’ Column I (internal_argument_compound) indicates the presence or absence of a compound element functioning as the internal argument of the nominalization in question. For example, (11) received the value 1. (11) data extraction Column J (accusative_internal_argument) indicates the presence or absence of the internal argument in the accusative case (only noun phrases were taken into account, not clauses). For example, (12) received the value 1. (12) I think this article was excellent in that its use of References was a model of how we should use the Internet when preparing essays. Column K (clausal_complement) indicates the presence or absence of a non-finite clause introduced by the complementizer to as a complement of the nominalization in question. For example, (13) received the value 1. (13) At that moment a hand was clapped heavily upon West 's shoulder, and the Boer who had saluted him so roughly pointed to the wagon, and he saw that his companion was being treated in the same way, while, the scare being over, upon their walking back and preparing to climb in, they were called upon to stop. Column L (free_clausal_complement) indicates the presence or absence of a free relative clause as a complement of the nominalization in question. For example, (14) received the value 1. (14) Investigating what is going on in the classroom is the beginning of a process of an enriching change, that will sometimes bring turbulence, but that is also rewarding for teachers. Column M (noun+ing) indicates the presence or absence of a noun in combination with an -ing nominalization, i.e. clausal gerunds with overt subjects. For example, (15) received the value 1. (15) The conceit here is almost as disingenuous as Val Gielgud fabricating letters to represent false opinion on BBC programmes. Column N (ECM) indicates the presence or absence of a particular type of ECM structure, i.e. a combination of an overt subject and a gerund as a complement of the verbs of perception or in a passive construction involving the verb of perception. For example, (16) received the value 1. (16) Click here to see Jodie investigating her Crinkle Bag Price: £11.99. References Baroni, M., Bernardini, S., Ferraresi, A., & Zanchetta, E. (2009). The WaCky wide web: A collection of very large linguistically processed web-crawled corpora. Language Resources and Evaluation, 43(3), 209–226. https://doi.org/10.1007/s10579-009-9081-4.

提供机构:
Zenodo
创建时间:
2025-07-12
二维码
社区交流群
二维码
科研交流群
商业服务