遇见数据集

西里尔蒙古文命名实体人工标注数据库

收藏
官方服务:

资源简介:

数据集面向命名实体识别任务研究,针对西里尔蒙古文命名实体标注工作,整理了5万句长度在15到25词数之间的文本语料,并与国外合作单位蒙古国科学院数学与数字技术研究所共同制定了西里尔蒙古文命名实体的标记规范,包括标记范围和标记规则。由国外合作单位专业数据标注人员人工标记完成。

This dataset is developed for named entity recognition (NER) research, focusing on the named entity annotation task for Cyrillic Mongolian. It includes 50,000 text sentences, each containing 15 to 25 words. The annotation specifications for Cyrillic Mongolian named entities, covering annotation scope and annotation rules, were jointly formulated with our foreign partner, the Institute of Mathematics and Digital Technology of the Mongolian Academy of Sciences. The dataset was manually annotated by professional data annotators from this foreign partner institution.

提供机构:
内蒙古大学
搜集汇总
数据集介绍
西里尔蒙古文命名实体人工标注数据库 数据集图片
背景与挑战
背景概述
该数据集是一个面向西里尔蒙古文命名实体识别任务的人工标注语料库,包含5万句文本,由蒙古国科学院数学与数字技术研究所合作制定标注规范并完成人工标记。它旨在支持相关自然语言处理研究,特别是命名实体识别领域。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务