Prompting to Gather Object Categories in NeRF Scenes Related to Manufacturing
摘要
Despite the effectiveness of closed-set object detectors, recent advancements have introduced zero-shot detectors that can recognize a wide range of object categories across different environments. These detectors rely on text prompts, such as object tags. This study explores using multimodal large language models (MLLMs) to gather and refine object information from NeRF scenes into tags. We propose a training-free pipeline for extracting object-specific details, such as category, color, material, and functionality, from 3D scenes via prompting. Subsequently, we investigate how to apply the object tagging problem to NeRF-reconstructed scenes, particularly in a manufacturing context. This pipeline is evaluated in manufacturing environments for object recognition, with the resulting categories serving as inputs for zero-shot object detection and other tasks.