<p>The rapid growth of AI/ML workloads has outpaced the capabilities of CPU-centric architectures to deliver the required data throughput and compute efficiency. This paper introduces a GPU-centric architecture leveraging GPUDirect Storage (GDS) to transfer data directly from SSDs to GPU memory, bypassing CPU bottlenecks and enabling high-throughput data paths. We propose Embedding from Storage Pipelined Network (ESPN) and its extension, ESPN-LIVE, which employ optimizations like data prefetching and on-demand embedding generation to align storage latency with GPU throughput. Experiments show ESPN reduces query latency by up to <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7118_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="38" /> </InlineMediaObject> <EquationSource Format="TEX">\(3.9\times\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>3.9</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation>, cuts memory usage by up to <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7118_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="33" /> </InlineMediaObject> <EquationSource Format="TEX">\(16\times\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>16</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation>, and improves throughput by up to 68%. ESPN-LIVE eliminates the need to store multi-vector embeddings by dynamically computing document representations, reducing storage costs by up to <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_7118_Article_IEq3.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="33" /> </InlineMediaObject> <EquationSource Format="TEX">\(16\times\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>16</mn> <mo>×</mo> </mrow> </math></EquationSource> </InlineEquation>, and making it particularly effective for single-query systems. These results highlight the potential of SSD-GPU integration for scalable, high-performance AI/ML workloads in information retrieval and LLM applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Storage access optimization for efficient GPU-centric information retrieval

  • Susav Shrestha,
  • Aayush Gautam,
  • Narasimha Reddy

摘要

The rapid growth of AI/ML workloads has outpaced the capabilities of CPU-centric architectures to deliver the required data throughput and compute efficiency. This paper introduces a GPU-centric architecture leveraging GPUDirect Storage (GDS) to transfer data directly from SSDs to GPU memory, bypassing CPU bottlenecks and enabling high-throughput data paths. We propose Embedding from Storage Pipelined Network (ESPN) and its extension, ESPN-LIVE, which employ optimizations like data prefetching and on-demand embedding generation to align storage latency with GPU throughput. Experiments show ESPN reduces query latency by up to \(3.9\times\) 3.9 × , cuts memory usage by up to \(16\times\) 16 × , and improves throughput by up to 68%. ESPN-LIVE eliminates the need to store multi-vector embeddings by dynamically computing document representations, reducing storage costs by up to \(16\times\) 16 × , and making it particularly effective for single-query systems. These results highlight the potential of SSD-GPU integration for scalable, high-performance AI/ML workloads in information retrieval and LLM applications.