<p>The HyperText Markup Language 5 (HTML5) <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;is useful for creating visual-centric web applications. However, unlike traditional web applications, HTML5 <InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;applications render objects onto the <InlineEquation ID="IEq9"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq10"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;bitmap without representing them in the Document Object Model (DOM). Mismatches between the expected and actual visual output of the <InlineEquation ID="IEq11"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq12"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;bitmap are termed visual bugs. Due to the visual-centric nature of <InlineEquation ID="IEq13"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq14"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;applications, visual bugs are important to detect because such bugs can render a <InlineEquation ID="IEq15"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq16"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;application useless. As we showed in prior work, <i>asset-based</i>&#xa0;graphics can provide the ground truth for a visual test oracle. However, many <InlineEquation ID="IEq17"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq18"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;applications procedurally generate their graphics. In this paper, we investigate how to detect visual bugs in <InlineEquation ID="IEq19"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq20"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;applications that use <i>procedural</i>&#xa0;graphics as well. In particular, we explore the potential of Vision-Language Models (VLMs) to automatically detect visual bugs. Instead of defining an exact visual test oracle, information about the application’s expected functionality (the context) can be provided with the screenshot as input to the VLM. To evaluate this approach, we constructed a dataset containing 80 bug-injected screenshots across four visual bug types (<i>Layout</i>, <i>Rendering</i>, <i>Appearance</i>, and <i>State</i>) plus 20 bug-free screenshots from 20 <InlineEquation ID="IEq21"> <EquationSource Format="TEX">\(&lt;\)</EquationSource> </InlineEquation><Emphasis FontCategory="NonProportional">canvas</Emphasis><InlineEquation ID="IEq22"> <EquationSource Format="TEX">\(&gt;\)</EquationSource> </InlineEquation>&#xa0;applications. We ran experiments with a state-of-the-art VLM using several combinations of text and image context to describe each application’s expected functionality. Our results show that by providing the application README(s), a description of visual bug types, and a bug-free screenshot as context, VLMs can be leveraged to detect visual bugs with up to 100% per-application accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring the capabilities of vision-language models to detect visual bugs in HTML5 canvas applications

  • Finlay Macklon,
  • Cedric Boucher,
  • Cor-Paul Bezemer

摘要

The HyperText Markup Language 5 (HTML5) \(<\) canvas \(>\)  is useful for creating visual-centric web applications. However, unlike traditional web applications, HTML5 \(<\) canvas \(>\)  applications render objects onto the \(<\) canvas \(>\)  bitmap without representing them in the Document Object Model (DOM). Mismatches between the expected and actual visual output of the \(<\) canvas \(>\)  bitmap are termed visual bugs. Due to the visual-centric nature of \(<\) canvas \(>\)  applications, visual bugs are important to detect because such bugs can render a \(<\) canvas \(>\)  application useless. As we showed in prior work, asset-based graphics can provide the ground truth for a visual test oracle. However, many \(<\) canvas \(>\)  applications procedurally generate their graphics. In this paper, we investigate how to detect visual bugs in \(<\) canvas \(>\)  applications that use procedural graphics as well. In particular, we explore the potential of Vision-Language Models (VLMs) to automatically detect visual bugs. Instead of defining an exact visual test oracle, information about the application’s expected functionality (the context) can be provided with the screenshot as input to the VLM. To evaluate this approach, we constructed a dataset containing 80 bug-injected screenshots across four visual bug types (Layout, Rendering, Appearance, and State) plus 20 bug-free screenshots from 20 \(<\) canvas \(>\)  applications. We ran experiments with a state-of-the-art VLM using several combinations of text and image context to describe each application’s expected functionality. Our results show that by providing the application README(s), a description of visual bug types, and a bug-free screenshot as context, VLMs can be leveraged to detect visual bugs with up to 100% per-application accuracy.