遇见数据集

Kuala Lumpur Travel blogs Dataset

收藏
Mendeley Data2024-03-27 更新2024-06-27 收录
官方服务:

资源简介:

This dataset contains three folders: 1) Training: The first sub-folder "raw training files" contains travel text extracted from 36 travel blog posts related to Kuala Lumpur. The second sub-folder "labeled files" consists of .xml version of raw text files containing 500 annotated spatial triplets as "trajector, spatial indicator, landmark" for spatial relation extraction. 2) Testing: The first sub-folder "raw testing files" contains travel text extracted from 10 travel blog posts related to Kuala Lumpur. The second sub-folder "labeled files" is the gold standard for evaluation consists of .xml version of raw text files containing 200 annotated spatial triplets as "trajector, spatial relation, landmark". 3) Related files: This folder contains annotation scheme definition (.xml) for training and testing files.

本数据集包含三个文件夹: 1) 训练集(Training):第一个子文件夹「原始训练文件(raw training files)」收录了从36篇与吉隆坡相关的旅游博客文章中提取的旅游文本;第二个子文件夹「标注文件(labeled files)」包含原始文本文件的XML格式版本,其中包含500条已标注的空间三元组(spatial triplets),格式为「轨迹载体(trajector)、空间指示词(spatial indicator)、地标(landmark)」,用于空间关系抽取任务。 2) 测试集(Testing):第一个子文件夹「原始测试文件(raw testing files)」收录了从10篇与吉隆坡相关的旅游博客文章中提取的旅游文本;第二个子文件夹「标注文件(labeled files)」作为评估所用的金标准(gold standard)数据集,包含原始文本文件的XML格式版本,其中包含200条已标注的空间三元组(spatial triplets),格式为「轨迹载体(trajector)、空间关系(spatial relation)、地标(landmark)」。 3) 关联文件(Related files):该文件夹包含训练集与测试集所用的标注方案定义(XML格式)文件。

创建时间:
2024-01-23
二维码
社区交流群
二维码
科研交流群
商业服务