Incomplete Scene Graph Fusion for Environment Map Estimation
摘要
Accurately predicting objects’ relative localizations in a home environment facilitates Embodied AI agents’ actions in their various missions. Scene graphs, which consist of triplets , have been used for this task notably due to their manageable data structure. However, the accuracy of current scene graph generation models remains insufficient for practical use. This work focuses on creating a symbolic representation of the environment map using images from fixed cameras, whose fields of view partially cover the scene and overlap. Based on the Dempster-Shafer framework, our proposed method represents and manages uncertain information captured by uncertainty scores provided by a scene graph generation model applied to simultaneous frames coming from distinct cameras. Our method is illustrated in an example of a scene in a virtual home generated using the VirtualHome2KG system. This illustration shows the potential of our method for correctly managing uncertainty and increasing the accuracy of scene graphs.