prachuryyaIITG/APTFiNER
收藏资源简介:
APTFiNER是一个通过使用大型语言模型(LLMs)进行注释保留翻译来创建高质量细粒度命名实体识别数据集的框架。利用APTFiNER,创建了六种语言的细粒度命名实体识别数据集:阿萨姆语(as)、博多语(brx)、马拉地语(mr)、尼泊尔语(ne)、泰米尔语(ta)和泰卢固语(te)。数据集统计信息包括每种语言的训练集、开发集和测试集中的句子数、实体数和标记数,以及注释者间一致性(IAA)分数。该数据集是AWED-FiNER生态系统的一部分。
APTFiNER is a framework to create high-quality fine-grained named entity recognition datasets through annotation preserving translation using LLMs. Utilizing APTFiNER, fine-grained named entity recognition dataset is created in six languages: Assamese (as), Bodo (brx), Marathi (mr), Nepali (ne), Tamil (ta) and Telugu (te). The dataset statistics include the number of sentences, entities, and tokens for each languages train, development, and test sets, along with Inter-Annotator Agreement (IAA) scores. The dataset is a part of the AWED-FiNER ecosystem.




