Hardware-Algorithm Co-design for Iterative Half-Pixel Motion Estimation Using Inter-Level Systolic Array
摘要
Motion estimation (ME) plays a very significant role in any standard video codec. It is one of the most compute and bandwidth intensive operation in modern vide codes. The complexity of motion estimation algorithms often causes memory bandwidth bottlenecks. This paper introduces a co-deign of search algorithm, dataflow and hardware architecture to alleviate memory bottleneck issues. We present a compact inter-level systolic architecture for half-pixel motion estimation that preserves the low-complexity nature of iterative search while recovering much of the data reuse needed for high-throughput implementation. The proposed design combines a 9-point iterative search strategy with a 3 × 3 array of processing elements, queue-assisted data alignment, and a broadcast-oriented data-flow that supplies 2 × 2 reference pixels per cycle. This organization enables efficient execution of large, small, and fractional search stages using the same processing array, while significantly reducing external memory traffic and on-chip bandwidth pressure relative to straightforward iterative implementations. The architecture requires only nine processing elements and a structured current-frame buffer organization to sustain high throughput. Synthesized in UMC 180 nm technology, the design occupies about 49.5 K gates and, at 145 MHz, supports real-time processing of 1080p video at 30 frames/s. Experimental evaluation shows that the proposed search strategy achieves an average PSNR degradation of only 0.44 dB relative to full search with half-pixel refinement, demonstrating that memory-bandwidth savings and compact hardware can be achieved with limited loss in coding performance.