NeSS-ST: Detecting Good and Stable Keypoints with a Neural Stability Score and the Shi-Tomasi Detector
Abstract: Learning a feature point detector presents a challenge both due to the ambiguity of the definition of a keypoint and, correspondingly, the need for specially prepared ground truth labels for such points. In our work, we address both of these issues by utilizing a combination of a hand-crafted Shi-Tomasi detector, a specially designed metric that assesses the quality of keypoints, the stability score (SS), and a neural network. We build on the principled and localized keypoints provided by the Shi-Tomasi detector and learn the neural network to select good feature points via the stability score. The neural network incorporates the knowledge from the training targets in the form of the neural stability score (NeSS). Therefore, our method is named NeSS-ST since it combines the Shi-Tomasi detector and the properties of the neural stability score. It only requires sets of images for training without dataset pre-labeling or the need for reconstructed correspondence labels. We evaluate NeSS-ST on HPatches, ScanNet, MegaDepth and IMC-PT demonstrating state-of-the-art performance and good generalization on downstream tasks.
- ORB-SLAM: a versatile and accurate monocular SLAM system. IEEE Trans. Robotics, 31(5):1147–1163, 2015.
- Structure-from-motion revisited. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 4104–4113, 2016.
- Video-rate localization in multiple maps for wearable augmented reality. In 12th IEEE International Symposium on Wearable Computers, pages 15–22, 2008.
- Scalable 6-dof localization on mobile devices. In Eur. Conf. on Computer Vision (ECCV), pages 268–283, 2014.
- Working hard to know your neighbor’s margins: Local descriptor learning loss. Advances in Neural Information Processing Systems (NIPS), 30, 2017.
- Sosnet: Second order similarity regularization for local descriptor learning. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 11016–11025, 2019.
- Image matching across wide baselines: From paper to practice. Intl. J. of Computer Vision, 129(2):517–547, 2021.
- Andrew P. Witkin. Scale-space filtering. In IJCAI, 1983.
- Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration. ACM Trans. Graph., 36(4):1, 2017.
- Pixelwise view selection for unstructured multi-view stereo. In Eur. Conf. on Computer Vision (ECCV), pages 501–518. Springer International Publishing, 2016.
- Lift: Learned invariant feature transform. In Eur. Conf. on Computer Vision (ECCV), pages 467–483. Springer, 2016.
- LF-Net: Learning local features from images. Advances in Neural Information Processing Systems (NIPS), 31, 2018.
- R2d2: Reliable and repeatable detector and descriptor. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems (NIPS), volume 32. Curran Associates, Inc., 2019.
- D2-net: A trainable cnn for joint detection and description of local features. arXiv preprint arXiv:1905.03561, 2019.
- Disk: Learning local features with policy gradient. Advances in Neural Information Processing Systems (NIPS), 33:14254–14265, 2020.
- Megadepth: Learning single-view depth prediction from internet photos. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 2041–2050, 2018.
- Scannet: Richly-annotated 3d reconstructions of indoor scenes. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 5828–5839, 2017.
- Jianbo Shi and Tomasi. Good features to track. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 593–600, 1994.
- SuperPoint: Self-supervised interest point detection and description. In IEEE Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 224–236, 2018.
- Key. net: Keypoint detection by handcrafted and learned cnn filters. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 5836–5844, 2019.
- Self-supervised equivariant learning for oriented keypoint detection. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 4847–4857, 2022.
- Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 5173–5182, 2017.
- Hans P. Moravec. Rover visual obstacle avoidance. In Proceedings of the 7th International Joint Conference on Artificial Intelligence - Volume 2, IJCAI’81, page 785–790, San Francisco, CA, USA, 1981. Morgan Kaufmann Publishers Inc.
- C. Harris and M. Stephens. A combined corner and edge detector. In Proceedings of the 4th Alvey Vision Conference, pages 147–151, 1988.
- A multiscale region detector. Computer Vision, Graphics, and Image Processing, 45(1):22–41, 1989.
- Tony Lindeberg. Feature detection with automatic scale selection. Intl. J. of Computer Vision, 30(2):79–116, 1998.
- Matching images with different resolutions. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), volume 1, pages 612–618 vol.1, 2000.
- Indexing based on scale invariant interest points. In Intl. Conf. on Computer Vision (ICCV), volume 1, pages 525–531. IEEE, 2001.
- An affine invariant interest point detector. In Eur. Conf. on Computer Vision (ECCV), pages 128–142. Springer, 2002.
- David G Lowe. Distinctive image features from scale-invariant keypoints. Intl. J. of Computer Vision, 60(2):91–110, 2004.
- Surf: Speeded up robust features. In Eur. Conf. on Computer Vision (ECCV), pages 404–417. Springer, 2006.
- Kaze features. In Eur. Conf. on Computer Vision (ECCV), pages 214–227. Springer, 2012.
- Survey and evaluation of rgb-d slam. IEEE Access, 9:21367–21387, 2021.
- U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer Assisted Intervention (MICCAI), pages 234–241. Springer, 2015.
- Adam: A method for stochastic optimization. In ICLR (Poster), 2015.
- A comparison of affine region detectors. Intl. J. of Computer Vision, 65:43–72, 2005.
- Tilde: A temporally invariant learned detector. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 5279–5288, 2015.
- Learning discriminative and transformation covariant local feature detectors. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 6818–6826, 2017.
- A performance evaluation of local descriptors. IEEE Trans. Pattern Anal. Machine Intell., 27(10):1615–1630, 2005.
- Learning feature descriptors using camera pose supervision. In Eur. Conf. on Computer Vision (ECCV), pages 757–774. Springer International Publishing, 2020.
- Gary Bradski. The opencv library. Dr. Dobb’s Journal: Software Tools for the Professional Programmer, 25(11):120–123, 2000.
- Learning to find good correspondences. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 2666–2674, 2018.
- Two-view geometry estimation unaffected by a dominant plane. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), volume 1, pages 772–779. IEEE, 2005.
- Opengv: A unified and generalized approach to real-time calibrated geometric vision. In IEEE Intl. Conf. on Robotics and Automation (ICRA), pages 1–8. IEEE, 2014.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.