MarcusBennevall/EtymologyTaggerDataset
收藏资源简介:
该数据集是一个英语词源标注数据集,包含102,111个经过解析的记录,源自Wiktionary条目,通过Kaikki提供的机器可读JSONL数据提取。数据集涵盖了英语单词的词源信息,包括单词、显示形式、词性、词源文本、词源对(如机制、源语言、源代码、源术语、模板、细节)、源语言序列和机制序列。统计显示,主要源语言包括拉丁语(23.24%)、英语(20.54%)、法语(16.35%)、希腊语(9.87%)等;词源机制分为借用(56.34%)、派生(44.81%)、继承(17.66%)和仿译(2.18%)。数据经过加工,包括语言合并(如将变体映射到主要语系)和排除非词源标签,以提高数据质量。数据集来源于Wiktextract提取的英语Wiktionary数据,但请注意,词源信息基于Wiktionary模板,可能无法反映复杂术语的所有历史细节。
This is an English etymology-annotated dataset consisting of 102,111 parsed records, extracted from machine-readable JSONL data provided by Kaikki and sourced from Wiktionary entries. It covers etymological information of English words, including the word itself, display form, part-of-speech (POS), etymological text, etymological pairs (e.g., mechanism, source language, source code, source term, template, details), source language sequence and mechanism sequence. Statistics show that the top source languages include Latin (23.24%), English (20.54%), French (16.35%), Greek (9.87%), among others; the etymological mechanisms are categorized into borrowing (56.34%), derivation (44.81%), inheritance (17.66%), and calquing (2.18%). The data has been processed with steps including language consolidation (e.g., mapping variants to their main language families) and removal of non-etymological tags to improve data quality. The dataset is derived from English Wiktionary data extracted via Wiktextract, but it should be noted that the etymological information is based on Wiktionary templates and may not fully reflect all historical details of complex terms.



