Enhancing multimodal interaction in resource-constrained environments: a case study on smart TV control
摘要
Multimodal interaction improves user experience by allowing users to select the most suitable way to interact with devices based on their needs and preferences. Its adoption has grown in IoT environments, particularly in resource-rich settings equipped with connected devices and intermediate tools for interaction detection. However, in resource-constrained environments, such as those in developing countries, interactions often rely on conventional methods because of the absence of intermediate devices, which are typically expensive, requiring a lengthy learning process for end-users. To overcome this limitation and make multimodal interaction accessible in such environments, this study demonstrates how smartphones can replace additional intermediate devices. As a case study, we explore the control of YouTube videos on a smart TV using a mobile application called UMI. The application was designed to simplify interaction methods and allow users to perform tasks such as fast forwarding, rewinding, navigating, and sharing videos through multimodal techniques, including swipe gestures (touch-based interaction) and hand-air gestures (touchless interaction). The performance of the UMI application was evaluated against traditional methods, including a remote control and the standard YouTube app. The evaluation considered criteria such as task completion time (speed of performing actions), error rate (accuracy of interactions), and user satisfaction. The results showed that the proposed approach enabled faster task completion and improved user satisfaction, while the error rate remained comparable to that of conventional methods.