NICKLE: The Neungyule Interlanguage Corpus of Korean Learners of English
收藏资源简介:
The corpus was constructed as a complementary resource for an English-Korean bilingual dictionary, the NeungYule-Longman English-Korean Dictionary. As a result, it may not be fully balanced. Basic Information: The size of the corpus is approximately 1 mil. tokens, including both written and spoken components. (in the ratio of approximately 9:1). The data is divided into several text types or registers according to the topics and communicative contexts. However, the usable size may be smaller after removing duplicate and irrelevant texts, depending on your research needs. Proficiency levels are not explicitly coded in the files, as they were collected from several universities across the country, each using different proficiency standards. The majority of the texts range from basic to pre-intermediate to intermediate levels, with some advanced-level texts included. You can identify advanced texts based on university names in the header or text lengths. When using the corpus, I typically refer to the source information (i.e., the university that produced the text) for proficiency-level insights. Annotation & Format: The corpus is not error-tagged or POS-tagged due to practical constraints. Only a few files had been error-tagged for testing an error-tagging scheme. However, automatic large-scale error tagging was not feasible for this corpus. If you need POS tagging, you can use any available NLP tools (e.g., TreeTagger, spaCy) or I can assist you in tagging the corpus if needed. The corpus is stored in XML format, following TEI (Text Encoding Initiative) standards.
本语料库作为《NeungYule-Longman英韩词典》(NeungYule-Longman English-Korean Dictionary)的配套资源构建而成,因此可能无法实现完全均衡分布。 ### 基本信息 该语料库规模约为100万Token(Token),涵盖书面与口语两大板块,二者占比约为9:1。数据依据主题与交际语境划分为多种文本类型或语域。但根据您的研究需求,在去除重复文本与无关内容后,可用语料规模可能会有所缩减。 本语料库未对文本熟练度等级进行显式编码,因数据采集自国内多所高校,而各校采用的熟练度评判标准并不统一。 绝大多数文本难度介于基础级、准中级至中级区间,同时包含少量高级文本。您可通过文件头部标注的高校名称或文本长度识别高级文本。在使用本语料库时,笔者通常会参考文本来源(即产出该文本的高校)来推断其熟练度等级。 ### 标注与格式 受实际条件限制,本语料库未进行错误标注与词性标注(POS tagging),仅少量文件曾用于测试错误标注方案而完成了错误标注。针对本语料库的大规模自动化错误标注并不可行。 若您需要词性标注,可使用现有自然语言处理工具(如TreeTagger、spaCy)完成;如有需要,笔者也可协助您完成语料库标注工作。 本语料库采用可扩展标记语言(XML)格式存储,遵循文本编码倡议(TEI)标准。



