遇见数据集

Tourism Tweets Dataset

收藏
arXiv2023-11-20 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

Tourism Tweets Dataset 是由法国波城和阿杜尔地区大学 E2S 研究所 LIUPPA 的研究团队创建的多语言数据集,包含法语、英语和西班牙语的旅游相关推文。该数据集通过 Twitter 学术 API 收集,覆盖 2019 年夏季的推文,重点关注法国巴斯克海岸地区的旅游相关内容。数据集经过人工标注,涵盖情感分析、命名实体识别和细粒度主题概念提取三个 NLP 任务。此数据集旨在解决旅游领域中从社交媒体提取结构化知识的挑战,尤其是在处理多语言、非结构化和非正式文本时的难题。

The Tourism Tweets Dataset is a multilingual dataset created by a research team from the LIUPPA Laboratory of the E2S Institute, University of Pau and Pays de l'Adour, France. It covers tourism-related tweets in French, English and Spanish. Collected via the Twitter Academic API, the dataset spans tweets from the summer of 2019, with a focus on tourism-related content concerning the French Basque Coast. The dataset has been manually annotated and supports three natural language processing (NLP) tasks: sentiment analysis, named entity recognition (NER), and fine-grained topic concept extraction. This dataset aims to address the challenges of extracting structured knowledge from social media in the tourism domain, particularly when dealing with multilingual, unstructured and informal texts.

创建时间:
2023-11-20
二维码
社区交流群
二维码
科研交流群
商业服务