遇见数据集

Named-Entity Recognition for Modern Tibetan Newspapers: Tagset, Guidelines and Training Data

收藏
Zenodo2022-12-06 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This dataset, tagset and guidelines were the output of a six-month incubator project on the feasibility of developing Named-Entity Recognition (NER) for modern Tibetan, primarily for use with contemporary Tibetan-language newspapers and media published inside the PRC. The project was carried out by the Mongolian and Inner Asian Studies Unit at Cambridge University’s Department of Social Anthropology. It was funded by an incubator grant from Cambridge Language Sciences. The project title was “Named-Entity Recognition in Tibetan and Mongolian Newspapers.” The Project PI was Dr Hildegard Diemberger (Cambridge), the Coordinator and Lead Author was Dr Robert Barnett (SOAS), and Senior Advisers were Dr Nathan Hill (SOAS), Dr Marieke Meelen (Cambridge), and Dr Thomas White (Cambridge). <br> <br> Although some forms of NER and other NLP procedures have been developed within China for modern Tibetan (see Liu, Nuo <em>et al</em>, 2011), the data underlying those initiatives have not been made publicly available and their findings cannot be tested or reproduced. Significant work on developing NLP for Tibetan has been carried out outside China, but has focused largely on classical Tibetan and religious texts (see Hill &amp; Garrett, Edward, 2017). The Cambridge incubator project therefore produced a tagset, guidelines and training data for developing NER for modern Tibetan, with a focus on historical and political analysis of contemporary newspapers, media and other public documents in Tibetan. We compiled 3.11m syllables of data in Tibetan extracted from articles downloaded from Chinese-language news aggregator sites within China, primarily tibet.cpc.people.com.cn and tibet.people.com.cn. From this data, we selected texts containing 280,000 syllables in Tibetan, grouped in 26,000 utterances/sentences (available on request). Using Lighttag, an online annotation site, we developed a tagset for NER consisting of 17 tags (and one for wrong segmentation if using segmented data). We annotated approximately 186,000 syllables, leading to 9,884 annotations. Of these, after discounting flawed data, we produced training data containing c.6,700 annotations. We carried out the secondary, manual review offline (for our method of converting Lighttag data for offline review, see the attached report “Using Spreadsheets to Review Annotations Offline.pdf”), and found an error rate of 3.6%. The final total of reviewed annotations was 6,624. The dataset, tagset, guidelines and reports were developed and documented by Robert Barnett, with assistance from Tsering Samdrup, Dr Hill and Dr Meelen. Primary annotation was by Tsering Samdrup, assisted by Dr Barnett.<br> <br> The datasets published here include: The <strong>tagseet guidelines and annotation manual</strong>, including the 17-tag tagset, guidelines, and recommendations ("NER for Modern Tibetan-tagset and guidelines.pdf"). The <strong>tagged training data </strong>in .csv format ("Tibetan NER Training Data-tagged, reviewed wth context-v10-UTF-8.csv") and .xls format ("Tibetan NER Training Data-tagged with context-v10-UTF-8.xlsx"). This includes 6,624 reveiwed annotations, arranged according to the Tibetan alphabet together with the tags and context (utterance) for each annotation. The <strong>raw annotation results </strong>downloaded from Lighttag as .json files ("Raw Training Data for NER in Modern Tibetan -Jobs2-11-JSON.zip") and as .xls files ("Training Data for NER in Modern Tibetan -Jobs2-11-XLS.zip"). These include 10 "tasks" or datasets of articles scraped from Tibetan-language websites within Tibet. A <strong>guide to preparing Lighttag annotation results for manual review offline </strong>(“Using Spreadsheets to Review Annotations Offline.pdf”). The project's findings regarding the status of NER and NLP for vertical Mongolian are available at DOI: 10.5281/zenodo.5103499.

本数据集、标注集(tagset)与标注指南,是一项为期六个月的孵化项目的产出成果。该项目旨在评估为现代藏文开发命名实体识别(Named-Entity Recognition,NER)技术的可行性,主要服务于中华人民共和国(PRC)境内发行的当代藏文报纸与媒体。本项目由剑桥大学社会人类学系蒙古与内亚研究组(Mongolian and Inner Asian Studies Unit)执行,获剑桥语言科学(Cambridge Language Sciences)提供的孵化资助,项目名称为《藏文与蒙古文报纸中的命名实体识别》。项目首席研究员(PI)为剑桥大学的希尔德加德·迪恩贝格尔博士,协调人与主要作者为伦敦大学亚非学院(School of Oriental and African Studies,SOAS)的罗伯特·巴尼特博士,高级顾问包括伦敦大学亚非学院的内森·希尔博士、剑桥大学的玛丽克·梅林博士与托马斯·怀特博士。 尽管中国境内已针对现代藏文开发了部分命名实体识别(NER)及其他自然语言处理(Natural Language Processing,NLP)技术(详见刘诺等(Liu, Nuo et al),2011年),但此类研究依托的数据集并未公开,其研究结果无法得到验证与复现。中国境外虽已开展大量藏文自然语言处理技术开发工作,但研究主要聚焦于古典藏文与宗教典籍(详见希尔与加勒特·爱德华(Hill & Garrett, Edward),2017年)。因此,剑桥大学孵化项目针对现代藏文开发命名实体识别(NER)技术打造了标注集(tagset)、标注指南与训练数据集,研究重点为当代藏文报纸、媒体及其他公开文献的历史与政治分析。我们从中国境内的中文新闻聚合网站(主要为tibet.cpc.people.com.cn与tibet.people.com.cn)下载的文章中提取了总计311万音节的藏文数据,并从中筛选出包含28万音节的藏文文本,共计2.6万条话语/句子(可按需获取)。我们依托在线标注工具Lighttag开发了一套包含17个标签的命名实体识别(NER)标注集(tagset,若使用分词数据,还可增设1个错误分词标签),共完成约18.6万音节的标注,总计得到9884条标注结果。在此基础上,我们剔除存在缺陷的数据后,得到包含约6700条标注结果的训练数据集。我们通过线下方式完成了二次人工复核(关于我们将Lighttag标注数据转换为线下复核格式的方法,详见附件报告《使用电子表格开展线下标注复核指南.pdf》),最终标注错误率为3.6%,最终经复核的有效标注总数为6624条。本数据集、标注集(tagset)、标注指南及相关报告由罗伯特·巴尼特主导开发与撰写,泽灵·桑珠(Tsering Samdrup)、希尔博士与梅林博士提供协助;首轮标注工作由泽灵·桑珠完成,巴尼特博士提供辅助。 本次发布的数据集包含以下内容:**标注集指南与标注手册**,内含17标签标注集、标注规范与相关建议("NER for Modern Tibetan-tagset and guidelines.pdf");**带标注的训练数据**,包含逗号分隔值(Comma-Separated Values,CSV)格式文件"Tibetan NER Training Data-tagged, reviewed with context-v10-UTF-8.csv"与XLS格式文件"Tibetan NER Training Data-tagged with context-v10-UTF-8.xlsx",内含6624条经复核的有效标注,按藏文字母顺序排列,并附带每条标注对应的标签与上下文(句子);**原始标注结果**,包含从Lighttag导出的JSON格式压缩包"Raw Training Data for NER in Modern Tibetan -Jobs2-11-JSON.zip"与XLS格式压缩包"Training Data for NER in Modern Tibetan -Jobs2-11-XLS.zip",其中包含从中国西藏境内藏文网站爬取的10组“任务”或文章数据集;**Lighttag标注结果线下人工复核准备指南**("Using Spreadsheets to Review Annotations Offline.pdf")。本项目关于竖写蒙古文自然语言处理与命名实体识别(NER)发展现状的研究结果可通过DOI:10.5281/zenodo.5103499获取。

提供机构:
Zenodo
创建时间:
2021-08-02
二维码
社区交流群
二维码
科研交流群
商业服务