TY - GEN
T1 - Scalable Object Detection in Mixed Reality Using Incremental Re-Training and One-Shot 3D Annotation
AU - Taheritajar, Alireza
AU - Benson, Jeffrey
AU - Gibson, Anthony
AU - Wilburn, Brandon
AU - Zhao, Jieqiong
AU - Orlosky, Jason
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - While object detection can be incredibly useful for a variety of augmented and mixed reality applications, achieving a large number of classifiable objects with high accuracy without extremely large deep learning (DL) or object recognition models is still difficult. More importantly, object recognition frameworks are often rigid in that they don't provide a direct means to add new classes to pretrained models in real time. In this paper, we introduce a novel approach that enables ondemand training of new object classes for consistent detection of in-situ objects for virtual labeling and interaction. By leveraging knowledge of the 3D location of an object in the scene taken from a mixed reality (MR) display's environment mesh, we are able to automate the labeling of subsequent 2D images taken from the frontfacing camera, which requires only a single, initial labeling interaction from an end-user. In addition, we have developed a continual learning approach that allows for on-the-fly retraining of the classifier and provides accurate classification quickly enough for the model to be practically usable in MR applications. We validate this approach by measuring the re-training time required for various object configurations, provide a comparison to other classification strategies, and analyze how the addition of object classes affect detection continuity across 3D scenes. We also demonstrate that labeling interactions work for practical applications in AR that are dependent on object detection, such as language learning, procedural instruction, or manufacturing guidance.
AB - While object detection can be incredibly useful for a variety of augmented and mixed reality applications, achieving a large number of classifiable objects with high accuracy without extremely large deep learning (DL) or object recognition models is still difficult. More importantly, object recognition frameworks are often rigid in that they don't provide a direct means to add new classes to pretrained models in real time. In this paper, we introduce a novel approach that enables ondemand training of new object classes for consistent detection of in-situ objects for virtual labeling and interaction. By leveraging knowledge of the 3D location of an object in the scene taken from a mixed reality (MR) display's environment mesh, we are able to automate the labeling of subsequent 2D images taken from the frontfacing camera, which requires only a single, initial labeling interaction from an end-user. In addition, we have developed a continual learning approach that allows for on-the-fly retraining of the classifier and provides accurate classification quickly enough for the model to be practically usable in MR applications. We validate this approach by measuring the re-training time required for various object configurations, provide a comparison to other classification strategies, and analyze how the addition of object classes affect detection continuity across 3D scenes. We also demonstrate that labeling interactions work for practical applications in AR that are dependent on object detection, such as language learning, procedural instruction, or manufacturing guidance.
KW - 3D annotation
KW - Mixed reality
KW - augmented reality
KW - incremental learning
KW - object detection
UR - https://www.scopus.com/pages/publications/105025043262
UR - https://www.scopus.com/pages/publications/105025043262#tab=citedBy
U2 - 10.1109/ISMAR67309.2025.00040
DO - 10.1109/ISMAR67309.2025.00040
M3 - Conference contribution
AN - SCOPUS:105025043262
T3 - Proceedings - 2025 IEEE International Symposium on Mixed and Augmented Reality, ISMAR 2025
SP - 282
EP - 292
BT - Proceedings - 2025 IEEE International Symposium on Mixed and Augmented Reality, ISMAR 2025
A2 - Eck, Ulrich
A2 - Lee, Gun
A2 - Plopski, Alexander
A2 - Smith, Missie
A2 - Sun, Qi
A2 - Tatzgern, Markus
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 24th IEEE International Symposium on Mixed and Augmented Reality, ISMAR 2025
Y2 - 8 October 2025 through 12 October 2025
ER -