遇见数据集

Tweets with traffic-related labels for developing a Twitter-based traffic information system.

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This data contain tweets that have been collected through Twitter search API. Each tweet has been classified into one of the three categories: 1) Non-Traffic (NT)-Class 0: Any tweet that does not fall into the other two categories is labeled as NT. 2) Traffic Incident (TI)-Class 1: This type of tweet reports non-recurring events that generate an abnormal increase in traffic demand or reduces transportation infrastructure capacity. The examples of non-recurring events include traffic crashes, disabled vehicles, highway maintenance, work zones, road closure, vehicle fire, traffic signal problems, special events, and abandoned vehicles. Since the ultimate goal of our framework is to inform users and agencies on the occurrence of a traffic incident in a real-time basis, if a tweet reports on the clearance or re-opening of roads that had already been affected by non-recurring traffic events, that tweet is classified as TCI, the third tweet category. Indeed, such tweets are providing information on the current status of the network rather than informing an ongoing traffic incident. 3) Traffic Conditions and Information (TCI)-Class 2: This type of tweet reports traffic flow conditions such as daily rush hours, traffic congestion, traffic delays due to high traffic volume, and jammed traffic. Also, any tweets that disseminate new traffic rules, traffic advisory, and any other information on transport infrastructures (e.g., new facilities or changing the direction of a street) are classified as TCI. Two types datasets are available: (1) 2-class dataset, in which tweets are categorized into traffic-related tweets (i.e., TI and TCI) and non-related-traffic tweets (i.e., only NT). 2) 3-class dataset, in which tweets are categorized into three groups including NT, TI, and TCI. Each type has its own training and test sets. Each csv file has three columns. First Column: Tweet class number according to the above definition. Second Column: Tweet id fetched from the Twitter API. Note that the 's' character should be removed. Third Column: The tweet text, that is used for analysis. Other attributes of each tweet (e.g., user’s screen_name, UCT time when the tweet is created, tweet’s unique ID, the geographic location of the tweet when posted) can be retrieved using the tweet id.

本数据集包含通过Twitter搜索API采集的推文,所有推文均被划分为以下三类: 1) 非交通类(Non-Traffic, NT)-类别0:凡不属于另外两类的推文均标注为非交通类。 2) 交通事件类(Traffic Incident, TI)-类别1:此类推文用于报道会引发交通需求异常攀升或削弱交通基础设施通行能力的非重复性事件。非重复性事件的示例包括交通事故、故障车辆、高速公路养护、作业区、道路封闭、车辆起火、交通信号故障、特殊活动以及遗弃车辆。鉴于本研究框架的最终目标是实时向用户与交管机构通报交通事件的发生情况,若某推文针对已受非重复性交通事件影响的道路发布清理或恢复通行的相关信息,则应将其归类为第三类——交通状况与信息类(Traffic Conditions and Information, TCI)。此类推文实际是在通报交通网络的当前状态,而非播报正在发生的交通事件。 3) 交通状况与信息类(Traffic Conditions and Information, TCI)-类别2:此类推文报道日常早高峰、交通拥堵、因高流量引发的延误、交通堵塞等交通流状况。此外,所有用于传播新交通规则、交通警示以及其他交通基础设施相关信息(例如新增设施或更改街道通行方向)的推文,均归类为TCI。 本数据集提供两种版本:(1) 二分类数据集:将推文划分为交通相关推文(即TI与TCI类)与非交通相关推文(仅NT类);(2) 三分类数据集:将推文划分为NT、TI与TCI三个类别。每种数据集均配有独立的训练集与测试集。 每个CSV文件均包含三列:第一列:按照上述定义标注的推文类别编号;第二列:从Twitter API获取的推文ID,请注意需移除其中的‘s’字符;第三列:用于分析的推文文本。 用户可通过推文ID获取该推文的其他属性,例如用户屏幕名、推文创建的UCT时间、推文唯一ID以及发布时的地理位置。

创建时间:
2018-03-23
二维码
社区交流群
二维码
科研交流群
商业服务