Manually Verified Idiomatic Variants from TenTen Corpora (English, Japanese, Korean)
收藏资源简介:
This dataset contains manually verified tokens of idiom variants extracted from three web-based TenTen corpora (enTenTen15, jaTenTen11, koTenTen18) via Sketch Engine. It supports a corpus-based investigation of lexical retention patterns under two types of structural manipulation — contraction (omission of canonical elements) and substitution (replacement of canonical elements) — across 25 idioms in three typologically distinct languages (English, Japanese, Korean). The dataset comprises over 14,000 variant tokens drawn from 25 idioms (approximately 8 per language), represented across over 80,000 rows in the coding spreadsheet. Each token is coded for manipulation type (contraction, substitution, or extension) and for the retention or non-retention of each canonical content-word element. Canonical forms were established via dictionary consultation and a majority-rules operationalization. The data were used to analyze element-level retention rates and to test whether lexical stability differs systematically between contraction and substitution contexts. Underlying corpus data are not included due to third-party licensing restrictions (Sketch Engine); only the derived coding is provided here. Contains: coded variant tokens with language, idiom, manipulation type, and per-element retention codes.



