遇见数据集

Exposing TCP watermarks in the Tor network using deep learning

收藏
Zenodo2024-08-07 更新2026-05-26 收录
官方服务:

资源简介:

Description The datasets consist of Tor network flows that have been captured with packet analyser Wireshark and converted into CSV format. The network flows consist of file transfers of images whose size varies from a few kilobytes to several megabytes. The captured packets are flows from the entry guard of the connection to the client. The datasets contain both clean Tor traffic and "watermarked" Tor traffic. Watermarking is a method of leaving small prints on the network flow at the sender end and trying to detect them at the receiving end. A positive detection indicates a connection between the two parties, thus breaking the anonymity aspect of the Tor. The used watermarking algorithms are "Interval-based watermarking" (IBW) presented by Pyun et al. [1] in 2007 and "Scalable watermark that is invisible and resilient to packet losses" (SWIRL) presented by Houmansadr and Borisov [2] in 2011. The algorithms were implemented with a watermarking module, which is essentially a modified TCP/IP stack. The module is publicly available [3]. The IBW-watermarked data is produced in the following way: The watermarked training data is endoded with the bit string {0, 1, 1, 0, 1, 0, 1, 1, 1, 0, 0, 0, 1, 1, 1, 0, 0, 0, 1, 1} by introducing the following repeating delays (milliseconds): {0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0}. The watermarked test data is encoded with the bit string {0, 1, 1, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 1, 0, 1, 1, 0, 0} by introducing the following repeating delays (milliseconds): {0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,0,1000,0,0,0,0, 1000,0,0,0,0, 1000,0,0,0,0, 0,0,0,0,1000,0,0,0,0, 0,0,0,0,1000,0,0,0,0}. The SWIRL-watermarked data is produced in the following way: The watermarked training data is encoded with a repeating delay string (milliseconds) {100, 0, 250, 250, 100, 250, 250, 100, 0, 0, 250, 0, 0, 100, 250, 0, 0, 250, 0, 250}, which mimics a SWIRL's permutation. The watermarked test data is encoded with a repeating delay string (milliseconds) {100, 100, 300, 300, 300, 100, 300, 100, 0, 0, 100, 100, 0, 0, 300, 0, 100, 100, 300, 0}, which also mimics a SWIRL's permutation. The datasets have been collected as a part of master's thesis work at Tampere University and are used primarily in neural network classification tests. The datasets intended for neural network training are longer and have "_train.csv" endings in their names. Datasets intended for testing are shorter and have "_test.csv" endings in their names. References [1] Y. J. Pyun, Y. H. Park, X. Wang, D. S. Reeves and P. Ning. "Tracing Traffic through Intermediate Hosts that Repacketize Flows," IEEE INFOCOM 2007 - 26th IEEE International Conference on Computer Communications, Anchorage, AK, USA, 2007 [2] A. Houmansadr and N. Borisov. “SWIRL: A Scalable Watermark to Detect Correlated Network Flows,” Network and Distributed System Security Symposium, San Diego, United States, 2011 [3] https://gitlab.com/nisec/tcp-watermark/

提供机构:
Zenodo
创建时间:
2024-08-07
二维码
社区交流群
二维码
科研交流群
商业服务