VeReMi_Extension: Dataset for Misbehaviors in VANETs
收藏资源简介:
This VeReMi Extension dataset consists of a set of misbehaviors in Vehicular Ad hoc Networks (VANETs) • Original dataset: The dataset used in this study is based on an existing dataset that Kamel et al. [1] published in 2019. They made use of the vehicle trace data from the Luxembourg SUMO Traffic (LuST) scenario, an open-source synthetic traffic scenario verified using real data from the VehicularLab of the University of Luxembourg. 39 resulting datasets make up this dataset. • Prepared dataset: We combined all 39 datasets into a single CSV file for our dataset. Ground-truth data and log files are included with each simulation. There is only one ground truth file that describes how a vehicle actually acts in the network, and it is used to run simulations. The ground truth file also includes an attacker type to distinguish between legitimate and misbehaving vehicles. However, in a simulation, the number of log files is equal to the number of vehicles on the network. Every vehicle creates a log file that contains all of the BSMs that were received. The first step is to combine all of the various log files into one file because there are as many log files as there are receivers. The ground truth file must then be mapped to the log files for each simulation in order to link the log files and ground truth files. We added a categorical feature called "class" for the target class to the combined dataset in order to create a labeled database. This dataset has an uneven class structure because the VeReMi Extension comprises 3,194,808 instances, with the normal class making up 59,488% of the whole database. The publication of this work on misbehavior detection is in [2] and [3]. Researchers in a variety of disciplines, including artificial intelligence and misbehavior detection in VANETs, can use our labeled dataset. When utilizing this dataset, please refer to the respective papers [2] and [3]. [1] J. Kamel, “Github repository: Framework for misbehavior detection (f2md),” 2019. [Online]. Available: https://github.com/josephkamel/f2md [2] O. Slama, B. Alaya, and S. Zidi, “Towards Misbehavior Intelligent Detection Using Guided Machine Learning in Vehicular Ad-hoc Networks (VANET),” Inteligencia Artificial, vol. 25, no. 70, pp. 138–154, 2022, doi: 10.4114/intartif.vol25iss70pp138-154. [3] O. Slama, B. Alaya, S. Zidi, and M. Tarhouni, “Comparative Study of Misbehavior Detection System for Classifying misbehaviors on VANET.,” in 2022 8th International Conference on Control, Decision and Information Technologies (CoDIT). IEEE, May 2022, vol. 1, pp. 243–248.
本VeReMi扩展数据集包含车载自组织网络(Vehicular Ad hoc Networks, VANETs)中的一系列异常行为样本。 • 原始数据集:本研究使用的数据集基于Kamel等人[1]于2019年发布的现有数据集。该团队采用了卢森堡SUMO交通(Luxembourg SUMO Traffic, LuST)场景下的车辆轨迹数据,该开源合成交通场景通过卢森堡大学VehicularLab的真实数据完成验证。该原始数据集共包含39个子数据集。 • 预处理数据集:我们将全部39个子数据集整合为单个逗号分隔值(Comma-Separated Values, CSV)文件,构成本数据集。每个仿真均附带真实标注数据(Ground-truth data)与日志文件。仅存在一份真实标注文件,用于描述车辆在网络中的实际运行行为,并作为仿真运行的依据;该真实标注文件还包含攻击者类型字段,用以区分合法车辆与异常行为车辆。在仿真过程中,日志文件的数量与网络内的车辆总数一致,每台车辆都会生成一份日志文件,记录其接收到的所有基本安全消息(Basic Safety Message, BSM)。由于日志文件的数量与接收方数量相等,第一步需将所有分散的日志文件整合为单个文件。随后,需将真实标注文件与各仿真的日志文件进行映射,以建立日志文件与真实标注文件之间的关联。为构建带标签的数据库,我们在整合后的数据集新增了名为"class"的分类特征作为目标类别。本数据集存在类别不平衡问题:VeReMi扩展数据集共计3,194,808条样本,其中正常类样本占比为59,488%。 本研究关于车载自组织网络异常行为检测的相关成果已发表于文献[2]与[3]。本带标签数据集可供人工智能、车载自组织网络异常行为检测等多个研究领域的科研人员使用。 使用本数据集时,请引用对应的文献[2]与[3]。 [1] J. Kamel,"Github仓库:异常行为检测框架(f2md)",2019年。[在线资源] 链接:https://github.com/josephkamel/f2md [2] O. Slama、B. Alaya与S. Zidi,"面向车载自组织网络(VANET)异常行为的智能检测:基于引导式机器学习的方法",《Inteligencia Artificial》,第25卷,第70期,第138–154页,2022年,DOI: 10.4114/intartif.vol25iss70pp138-154。 [3] O. Slama、B. Alaya、S. Zidi与M. Tarhouni,"车载自组织网络异常行为分类的异常检测系统对比研究",收录于2022年第8届控制、决策与信息技术国际会议(CoDIT),IEEE出版社,2022年5月,第1卷,第243–248页。




