Authors: Dharini Raghavan, Raghu Krishnapuram, Bharadwaj S Amrutur

We propose a scalable perception framework leveraging monocular security cameras in the infrastructure for localizing and tracking indoor autonomous mobile robots. We present a zero-shot pose estimation approach that combines semantic and visual descriptors to identify reliable, repeatable and robust keypoints along with a quantification of its epistemic uncertainty via a mathematical covariance model of the external camera. These pose estimates are then fused with the robot’s on-board sensors to achieve high-accuracy localization. We also enhance an optimal camera placement algorithm by constraining it with the external camera’s covariance to simultaneously maximize total coverage and localization accuracy which is an integral aspect of multi-camera robot localization systems. We show through real-world experiments that fusing pose estimates from fixed monocular security cameras with an off-the shelf visual SLAM system results in a significant improvement in localization performance alongside eliminating the kidnapped robot problem.

Share: