遇见数据集

Learning about the Internet through efficient sampling and aggregation

收藏
Mendeley Data2024-01-31 更新2024-06-28 收录
官方服务:

资源简介:

The Internet is important for nearly all aspects of our society, affecting ordinary people, businesses, and social activities. Because of its importance and wide‐spread applications, we want to have good knowledge about Internet's operation, reliability and performance, through various kinds of measurements. However, despite the wide usage, we only have limited knowledge of its overall performance and reliability. The first reason of this limited knowledge is that there is no central governance of the Internet, making both active and passive measurements hard. The second reason is the huge scale of the Internet. This makes brute‐force analysis hard because of practical computing resource limits such as CPU, memory and probe rate. ❧ This thesis states that sampling and aggregation are necessary to overcome resource constraints in time and space to learn about better knowledge of the Internet. Many other Internet measurement studies also utilize sampling and aggregation techniques to discover properties of the Internet. We distinguish our work by exploring novel mechanisms and new knowledge in several specific areas. First, we aggregate short‐time‐scale observations and use an efficient multi‐time‐scale query scheme to discover the properties and reasons of long‐lived Internet flows. Second, we sample and probe /24 blocks in the IPv4 address space, and use greedy clustering algorithms to efficiently characterize Internet outages. Third, we show an efficient and effective aggregation technique by visualization and clustering. This technique makes both manual inspection and automated characterization easier. Last, we develop an adaptive probing system to study global scale Internet reliability. It samples and adapts probe rate within each /24 block for accurate beliefs. By aggregation and correlation to other domains, we are also able to study broader policy effects on Internet use, such as political causes, economic conditions, and access technologies. ❧ This thesis provides several examples of Internet knowledge discovery with new mechanisms of sampling and aggregation techniques. We believe our approaches of new sampling and aggregation mechanisms can be used by and will inspire new ways for future Internet measurement systems to overcome resource constraints, such as large amount and dispersed data.

互联网对社会几乎所有层面都至关重要,其影响力覆盖普通民众、商业运营与社会活动各领域。鉴于其重要性与广泛应用,我们希望通过各类测量手段,充分掌握互联网的运行状况、可靠性与性能表现。然而,尽管互联网应用极为普及,我们对其整体性能与可靠性的认知却十分有限。造成这一现状的原因有二:其一,互联网并无中心化治理机制,使得主动与被动测量均难以开展;其二,互联网规模极其庞大,受CPU、内存、探测速率等实际计算资源的限制,蛮力分析难以实施。 ◆ 本论文提出,为克服时空维度的资源约束以更深入地认知互联网,采样与聚合技术必不可少。诸多互联网测量研究也均采用采样与聚合手段来挖掘互联网的特性。本研究的创新之处在于,在多个特定领域探索了全新机制与新知。其一,我们聚合短时观测数据,并采用高效的多时间尺度查询方案,以探究长存活期互联网流的特性与成因;其二,我们对IPv4地址空间(IPv4 address space)中的/24网段(/24 block)进行采样与探测,并利用贪婪聚类算法高效表征互联网中断事件;其三,我们提出一种结合可视化与聚类的高效聚合技术,该技术可简化人工检视与自动化表征流程;最后,我们开发了一套自适应探测系统,用于研究全球范围的互联网可靠性,该系统可在每个/24网段内调整采样与探测速率以获得精准的观测认知。通过聚合与关联其他领域的数据,我们还可进一步探究影响互联网使用的更广泛政策效应,例如政治动因、经济环境与接入技术等。 ◆ 本论文展示了若干基于新型采样与聚合技术的互联网知识发现案例。我们认为,所提出的新型采样与聚合机制,可被未来互联网测量系统所采用,并将为其克服海量且分散的数据所带来的资源约束提供新思路。

创建时间:
2024-01-31
二维码
社区交流群
二维码
科研交流群
商业服务