遇见数据集

arabic-digital-humanities/root-extraction-validation-data: 0.1.0

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Dataset for evaluating root extraction This dataset contains data to evaluate the roots extracted by Arabic stemmers and morphological analyzers. The dataset was created in the context of the Bridging the Gap project. The <code>txt</code> directory contains the text files for which roots were extracted, and <code>gs</code> contains csv files with columns <code>word</code> and <code>root</code>, containing the words from the text and the manually corrected roots. If multiple roots apply, they have been separated using a backslash (<code>\</code>). If the word is a letter, the root is <code>#</code>. The text files have been analyzed using SAFAR, using our software. The directories <code>khoja</code>, <code>isri</code>, and <code>alkhalil</code> contain the output xml files. License The data in this repository is licensed under a Creative Commons Attribution 4.0 International License. The text fragments have been taken from the OpenITI project. The stopword list was created by Maksim Abdul Latif.

提供机构:
Zenodo
创建时间:
2019-06-26
二维码
社区交流群
二维码
科研交流群
商业服务