five

Wikidata Lemmatization Dataset

收藏
NIAID Data Ecosystem2026-05-01 收录
下载链接:
https://zenodo.org/record/10819305
下载链接
链接失效反馈
官方服务:
资源简介:
The Wikidata Lemmatization Dataset was collected using the following SPARQL query: https://w.wiki/9TwH Languages included in the dataset: Akkadian : AKK (Q35518) Arabic : AR (Q13955) Czech : CS (Q9056) German : DE (Q188) English : EN (Q1860) French : FR (Q150) Hebrew : HE (Q9288) Hittite : HIT (Q35668) Italian : IT (Q652) Russian : RU (Q7737) Sumerian : SUX (Q36790) Turkish : TR (Q256) The choice of languages to include have to do with a collection of primary and secondary source documents which we have digitized (OCR) and are using as references for the FactGrid Cuneiform project. The resulting lexemes for each language are shared in CSV with the file names references each language, their Wikidata Q-ids, the number of lexemes at that date, and the date of access (MM_YYYY). The format of each CSV includes the following fields: lexeme : the Wikidata lexeme id (L-id) lexemeLabel : the label assigned to the lexeme in Wikidata lexical_category : the Wikidata Q-item for the part of speech lexical_categoryLabel : the label assigned to the lexical category (e.g. noun, verb, adjective, etc.) This dataset will be updated periodically using standard version control.
创建时间:
2024-03-14
5,000+
优质数据集
54 个
任务类型
进入经典数据集
二维码
社区交流群

面向社区/商业的数据集话题

二维码
科研交流群

面向高校/科研机构的开源数据集话题

数据驱动未来

携手共赢发展

商业合作