遇见数据集

WHLL corpus

收藏
arXiv2024-03-25 更新2024-06-21 收录
官方服务:

资源简介:

WHLL corpus是由京都大学信息学研究科创建的一个大规模地理解析数据集,包含超过130万篇维基百科文章,每篇文章平均包含约7.8个独特的地理位置表达。该数据集通过利用维基百科中的超链接,自动为多个地理位置表达标注坐标,其中45.6%的表达存在歧义。创建过程涉及从维基百科的HTML和CirrusSearch dump文件中提取信息,自动为文章中的地理位置表达分配坐标。该数据集主要用于训练和评估机器学习模型在处理文本中的地理位置信息时的性能,特别是在解决地理位置表达歧义方面的应用。

The WHLL corpus is a large-scale geographic parsing dataset developed by the Graduate School of Informatics at Kyoto University. It contains over 1.3 million Wikipedia articles, with each article averaging approximately 7.8 unique geographic location mentions. This dataset automatically annotates geographic coordinates for multiple such mentions using hyperlinks embedded in Wikipedia, and 45.6% of these location mentions are ambiguous. The creation process extracts information from Wikipedia's HTML and CirrusSearch dump files, then automatically assigns coordinates to geographic location mentions in the articles. Primarily, this dataset is used for training and evaluating the performance of machine learning models when handling geographic information in text, especially for applications focused on geographic location mention disambiguation.

创建时间:
2024-03-25
二维码
社区交流群
二维码
科研交流群
商业服务