Storage access optimization for efficient GPU-centric information retrieval
摘要
The rapid growth of AI/ML workloads has outpaced the capabilities of CPU-centric architectures to deliver the required data throughput and compute efficiency. This paper introduces a GPU-centric architecture leveraging GPUDirect Storage (GDS) to transfer data directly from SSDs to GPU memory, bypassing CPU bottlenecks and enabling high-throughput data paths. We propose Embedding from Storage Pipelined Network (ESPN) and its extension, ESPN-LIVE, which employ optimizations like data prefetching and on-demand embedding generation to align storage latency with GPU throughput. Experiments show ESPN reduces query latency by up to