Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Decoding Scheme with Successive Aggregation of Multi-Level Features for Light-Weight Semantic Segmentation

Published 17 Feb 2024 in cs.CV | (2402.11201v2)

Abstract: Multi-scale architecture, including hierarchical vision transformer, has been commonly applied to high-resolution semantic segmentation to deal with computational complexity with minimum performance loss. In this paper, we propose a novel decoding scheme for semantic segmentation in this regard, which takes multi-level features from the encoder with multi-scale architecture. The decoding scheme based on a multi-level vision transformer aims to achieve not only reduced computational expense but also higher segmentation accuracy, by introducing successive cross-attention in aggregation of the multi-level features. Furthermore, a way to enhance the multi-level features by the aggregated semantics is proposed. The effort is focused on maintaining the contextual consistency from the perspective of attention allocation and brings improved performance with significantly lower computational cost. Set of experiments on popular datasets demonstrates superiority of the proposed scheme to the state-of-the-art semantic segmentation models in terms of computational cost without loss of accuracy, and extensive ablation studies prove the effectiveness of ideas proposed.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (22)
  1. “Cross-dataset collaborative learning for semantic segmentation in autonomous driving,” in AAAI, 2022, vol. 36, pp. 2487–2494.
  2. “Global and local feature reconstruction for medical image segmentation,” IEEE Transactions on Medical Imaging, vol. 41, no. 9, pp. 2273–2284, 2022.
  3. “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2020.
  4. “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in ICCV, 2021, pp. 568–578.
  5. “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV, 2021, pp. 10012–10022.
  6. “Lite vision transformer with enhanced self-attention,” in CVPR, 2022, pp. 11998–12008.
  7. “Metaformer is actually what you need for vision,” in CVPR, 2022, pp. 10819–10829.
  8. “Segformer: Simple and efficient design for semantic segmentation with transformers,” NeurIPS, vol. 34, pp. 12077–12090, 2021.
  9. “Lawin transformer: Improving semantic segmentation transformer with multi-scale representations via large window attention,” arXiv preprint arXiv:2201.01615, 2022.
  10. “Efficient self-ensemble for semantic segmentation,” arXiv preprint arXiv:2111.13280, 2021.
  11. “Per-pixel classification is not all you need for semantic segmentation,” NeurIPS, vol. 34, pp. 17864–17875, 2021.
  12. “Masked-attention mask transformer for universal image segmentation,” in CVPR, 2022, pp. 1290–1299.
  13. “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in ECCV, 2018, pp. 801–818.
  14. “In defense of pre-trained imagenet architectures for real-time semantic segmentation of road-driving images,” in CVPR, 2019, pp. 12607–12616.
  15. “RTformer: Efficient design for real-time semantic segmentation with transformer,” NeurIPS, vol. 35, pp. 7423–7436, 2022.
  16. “Improving semantic segmentation in transformers using hierarchical inter-level attention,” arXiv preprint arXiv:2207.02126, 2022.
  17. “Learning implicit feature alignment function for semantic segmentation,” in ECCV. Springer, 2022, pp. 487–505.
  18. “Shuffle transformer: Rethinking spatial shuffle for vision transformer,” arXiv preprint arXiv:2106.03650, 2021.
  19. “Scene parsing through ADE20k dataset,” in CVPR, 2017, pp. 633–641.
  20. “The Cityscapes dataset for semantic urban scene understanding,” in CVPR, 2016, pp. 3213–3223.
  21. MMSegmentation Contributors, “Openmmlab semantic segmentation toolbox and benchmark,” https://github.com/open-mmlab/mmsegmentation, 2020.
  22. “Panoptic feature pyramid networks,” in CVPR, 2019, pp. 6399–6408.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.