NPN Construction and Distractor dataset
收藏资源简介:
The present dataset collects instances of the Italian Noun-Preposition-Noun (NPN) discontinuous reduplication construction (e.g., pagine su pagine, pages upon pages). The instances come from the CORIS corpus, 2021 version, 165Mw (https://corpora.ficlit.unibo.it/TCORIS/). Dataset construction Data is built upon Masini (2024), the result of an empirical analysis that surveys all occurrences of the construction pattern NPN found in CORIS, a large-scale reference corpus of contemporary written Italian developed at the University of Bologna. The structure of Masini (2024) Masini (2024) contains 1298 construction types, each annotated with the following parameters: NPN: the NPN expression (type) token_frequency: the token frequency of the NPN expression in the CORIS 2021 corpus preposition: the preposition found in the NPN expression (12 possible values) reduplicated_noun: the lemmatised form of the noun that appears in the NPN expression number_of_noun: the number value of the noun that appears in the NPN expression, with possible values singular, plural, singular/plural syntactic_function: the syntactic function of the NPN expression, with possible values modifier, nominal, clause meaning: the semantic function of the NPN expression (9 possible values) The original data taken as a starting point for this work is contained in the file dataset_Masini.csv. From Masini (2024) to the NPN dataset For the present dataset, we only considered a subset of the original data, namely constructions instantiated by the prepositions su and a. Each construction form in this subset was manually searched in CORIS in order to retrieve the occurrence within its sentential boundaries. This procedure revealed a number of instances that cannot be identified through a CQL query of the form NN–PREP–NN, since automatic PoS tagging in CORIS occasionally fails to assign the NN tag to nouns that function as such, instead labelling them as verbs (VERB). Main differences in frequency between the present dataset and Masini (2024), where they occur, can therefore be attributed to this issue. Other mismatches come from the decision for the present dataset to exclude some occurrences. The main reasons are duplication of sentences in the corpus and nominal phrases with no significant context. The differences in frequencies between Masini (2024) and the present dataset are underlined in dataset_freq.csv. The exclusion criteria reported by Masini (2024) were initially followed: foreign expressions (e.g., vis a vis) instances containing proper nouns (e.g., Italia su Italia) occurrences that actually belong to another construction (e.g., PNPN) With respect to the original dataset, we removed: lemmas that are primarily adjectives or adverbs (e.g., poco a poco) dialectal expressions (e.g., core a core) nominal expressions (e.g., newspaper headlines) expressions appearing in bullet lists without a proper sentential context Hence, both construction and non-construction (henceforth, distractor) instances were annotated with categories listed in the NPN dataset section. Construction instances parameters The construction dataset (data/construction.csv) contains 3256 occurrences, each annotated with the following parameters: NPN: the NPN expression (type) preposition: the preposition found in the NPN expression, with possible values a, su noun: the lemmatised form of the noun that appears in the NPN expression number_of_noun: the number value of the noun that appears in the NPN expression, with possible values singular (2890 items), plural (346 items), singular/plural (20 items) syntactic_function: the syntactic function of the NPN expression, with possible values modifier (2089 items), nominal (1166 items), clause (1 item) meaning: the semantic function of the NPN expression, with possible values succession/iteration/distributivity (577 items), greater_plurality/accumulation (392 items), juxtaposition/contact (1144 items), connection/transition (50 items), inescapable_presupposition (1 item), intensification (1 item), idiosyncratic (1091 items) Distractor instances parameters In addition to the dataset of actual occurrences of the NPN construction, a parallel dataset of distractors (data/distractor.csv) was created, defined as sequences sharing the same surface form as the construction but not instantiating it. The distractor dataset contains 1751 occurrences, each annotated with the following parameters: NPN: surface form of the distractor pattern shared with the NPN construction preposition: the preposition found in the pattern, with possible values a (1565 items), su (186 items) noun: the lemmatised form of the noun that appears reduplicated in the pattern number_of_noun: the number value of the noun that appears in the pattern, with possible values singular (1600 items), plural (1750 items), singular/plural (8 items), X (108 items) type of distractor: pattern type, with possible values N_extended (35 items), NsuNgiù (392 items), juxtaposition/contact (13 items), NumPNum (100 items), PNPN (1442 items), proper_name_inglobation (31 items), thematic_target (50 items), verbal (80 items) proper_cxn: whether this represents a structurally distinct construction different from the NPN and linked horizontally, or merely a surface-isomorphic pattern, with possible values yes (1535 items), no (216 items) Coder agreement Annotation processInter-annotator agreement was evaluated on a controlled subset of 100 instances. All annotators followed a shared annotation schema, including formal constraints and semantically motivated label definitions derived from the literature. The agreement subset was balanced across semantic categories: 50 instances from the subset of NPN constructions instantiated by the preposition a (25 juxt, 25 succ) 50 instances from the subset of NPN constructions instantiated by the preposition su (25 succ, 25 acc) Inter-annotator agreementAnnotation quality was assessed using Cohen’s kappa and Krippendorff’s alpha. Pairwise Cohen’s kappa shows consistently high observed agreement (Ao = 0.87–0.94), with stable expected agreement (Ae approximately 0.36–0.37), resulting in strong to near-perfect reliability across annotator pairs (kappa = 0.79–0.91). Since Cohen’s kappa assumes equal distances between labels, agreement was also evaluated using Krippendorff’s alpha, which supports distance-sensitive comparison. Nominal alpha is high (alpha = 0.858) and further increases when a reduced penalty is assigned to confusions between semantically adjacent labels (custom distance = 0.5; alpha = 0.892). This indicates that residual disagreement mainly concerns borderline semantic cases rather than inconsistencies in the shared annotation scheme.



