NERSkill.Id
收藏Mendeley Data2026-04-18 收录
下载链接:
https://data.mendeley.com/datasets/5s8r9ndfvc
下载链接
链接失效反馈官方服务:
资源简介:
NERSkill.Id stands out as the initial annotated corpus designed specifically for NER datasets emphasizing skill entities in the Indonesian language. This marks a valuable addition to the existing resources for Natural Language Processing (NLP) in Indonesian. Despite its relatively compact size, NERSkill.Id holds considerable promise for refining language models. Moreover, its integration with larger pre-existing corpora can enhance the training of more extensive and versatile mixed Indonesian models tailored for diverse NLP tasks.
The dataset categorizes named entities into three distinct classes: hard skill, soft skill, and technology. It consists of 418.868 tokens. Subsequently, these tokens are marked using the BIO format. The annotation table is presented in ConLL2003 format, consisting of three columns: Sentence#, word, and tag columns.
*We already have paper at https://www.sciencedirect.com/science/article/pii/S235234092400163X (please cite)
创建时间:
2024-04-02



