遇见数据集

Wikidata Lemmatization Dataset

收藏
Zenodo2024-03-15 更新2026-05-26 收录
官方服务:

资源简介:

The Wikidata Lemmatization Dataset was collected using the following SPARQL query: https://w.wiki/9TwH Languages included in the dataset: Akkadian : AKK (Q35518) Arabic : AR (Q13955) Czech : CS (Q9056) German : DE (Q188) English : EN (Q1860) French : FR (Q150) Hebrew : HE (Q9288) Hittite : HIT (Q35668) Italian : IT (Q652) Russian : RU (Q7737) Sumerian : SUX (Q36790) Turkish : TR (Q256) The choice of languages to include have to do with a collection of primary and secondary source documents which we have digitized (OCR) and are using as references for the FactGrid Cuneiform project. The resulting lexemes for each language are shared in CSV with the file names references each language, their Wikidata Q-ids, the number of lexemes at that date, and the date of access (MM_YYYY). The format of each CSV includes the following fields: lexeme : the Wikidata lexeme id (L-id) lexemeLabel : the label assigned to the lexeme in Wikidata lexical_category : the Wikidata Q-item for the part of speech lexical_categoryLabel : the label assigned to the lexical category (e.g. noun, verb, adjective, etc.) This dataset will be updated periodically using standard version control.

Wikidata词形还原(Lemmatization)数据集通过以下SPARQL查询完成采集:https://w.wiki/9TwH 本数据集涵盖的语言如下: 阿卡德语(Akkadian):AKK(Q35518) 阿拉伯语(Arabic):AR(Q13955) 捷克语(Czech):CS(Q9056) 德语(German):DE(Q188) 英语(English):EN(Q1860) 法语(French):FR(Q150) 希伯来语(Hebrew):HE(Q9288) 赫梯语(Hittite):HIT(Q35668) 意大利语(Italian):IT(Q652) 俄语(Russian):RU(Q7737) 苏美尔语(Sumerian):SUX(Q36790) 土耳其语(Turkish):TR(Q256) 本次选取的语言覆盖范围,与我们已通过光学字符识别(Optical Character Recognition, OCR)完成数字化,并作为FactGrid楔形文字项目参考素材的一手、二手源文献集合密切相关。各语言对应的词位(lexeme)均以CSV格式共享,文件名包含对应语言名称、其Wikidata Q编号、截至该日期的词位数量以及访问日期(格式为MM_YYYY)。 各CSV文件的字段格式如下: lexeme:Wikidata词位标识符(L-id) lexemeLabel:Wikidata为该词位分配的标签 lexical_category:对应词性的Wikidata Q项 lexical_categoryLabel:为该词性分配的标签(例如名词、动词、形容词等) 本数据集将通过标准版本控制机制定期更新。

提供机构:
Zenodo
创建时间:
2024-03-14
二维码
社区交流群
二维码
科研交流群
商业服务