遇见数据集

NERSkill.Id

收藏
Mendeley Data2024-04-05 更新2024-06-27 收录
官方服务:

资源简介:

NERSkill.Id stands out as the initial annotated corpus designed specifically for NER datasets emphasizing skill entities in the Indonesian language. This marks a valuable addition to the existing resources for Natural Language Processing (NLP) in Indonesian. Despite its relatively compact size, NERSkill.Id holds considerable promise for refining language models. Moreover, its integration with larger pre-existing corpora can enhance the training of more extensive and versatile mixed Indonesian models tailored for diverse NLP tasks. The dataset categorizes named entities into three distinct classes: hard skill, soft skill, and technology. It consists of 418.868 tokens. Subsequently, these tokens are marked using the BIO format. The annotation table is presented in ConLL2003 format, consisting of three columns: Sentence#, word, and tag columns. *We already have paper at https://www.sciencedirect.com/science/article/pii/S235234092400163X (please cite)

创建时间:
2024-01-23
二维码
社区交流群
二维码
科研交流群
商业服务