SmartNIC-Offloaded Worker Node for Storage━Compute Disaggregated Recommendation System
收藏资源简介:
Deep learning-based recommendation systems are commonly used to provide personalized recommendations. In a common storage—compute disaggregated inference architecture, the inference speed of the recommendation system is limited by the internode network transmission bottleneck caused by embedding queries. The emerging SmartNIC technology enables complex traffic control without contending for host Central Processing Unit (CPU) resources, offering new possibilities for optimizing the embedding layer in disaggregated recommendation systems. This study proposes SmartNIC-offloaded Worker Node (SmartWN), a disaggregated recommendation system worker node optimized via SmartNIC. By leveraging the independent computing and communication capabilities of SmartNICs, SmartWN implements embedding query reordering and preparation, along with traffic-aware dynamic cache management for multiple embedding tables without impacting host resources. This significantly improves communication efficiency and cache utilization during recommendation inference, reduces embedding query latency, and enhances overall system performance. This study implements SmartWN on an NVIDIA BlueField-2 SmartNIC and demonstrates its performance improvements. Compared to existing technologies, using SmartWN as a compute node in a disaggregated recommendation system significantly enhances the embedding layer query throughput by 2.13x and reduces query latency by approximately 50.6%.



