The Methods by Which Machine Learning Enhances the Resolution and Detail of Extended Reality Display Devices
摘要
With the wide application of Extended Reality (XR) technology in fields such as education, healthcare, industry and entertainment, users’ demands for visual immersion and image clarity have significantly increased. However, due to the physical constraints such as power consumption, computing resources and size of XR head-mounted devices, there are still bottlenecks in their display resolution and the ability to present image details. Traditional image enhancement techniques, such as algorithms based on interpolation or filtering, are difficult to meet the requirements of high frame rate, low latency and dynamic interaction in XR scenes. In recent years, the development of deep learning, especially convolutional neural networks (CNN), generative adversarial networks (GAN), and lightweight Super-Resolution (SR) models, has provided new ideas for breaking through the bottleneck of XR display.This paper proposes a real-time image enhancement framework for XR systems, combining a lightweight neural network structure, channel attention mechanism and gaze area guidance strategy to conduct high-quality image super-resolution reconstruction for the user's current gaze area. Meanwhile, a low-complexity enhancement scheme is adopted for non-focus areas to improve the overall visual quality while controlling the computational overhead. This study trained the model based on public image datasets such as DIV2K and Flickr2K, and conducted tests and fine-tuning in a real XR usage environment. The experimental results show that the proposed method, on the premise of ensuring stable frame rate and controllable delay, is superior to the traditional interpolation method and some mainstream SR models in terms of image quality indicators such as PSNR, SSIM, and LPIPS, and shows great advantages in subjective perceived quality and user immersive experience. This study not only verified the feasibility and effectiveness of machine learning methods in improving the display performance of XR, but also provided a technical path for the future deployment of adaptive image enhancement algorithms in edge computing devices, laying the foundation for the next generation of immersive interaction systems.