A Hybrid Pyramid and Strip Pooling Network for Accurate Building Extraction from Remote Sensing Images
DOI:
https://doi.org/10.54097/gz73jc92Keywords:
Building Extraction, Remote Sensing Imagery, Semantic Segmentation, Multi-Scale Feature Fusion, Strip Pooling, Attention MechanismAbstract
Accurate extraction of building footprints from remote sensing imagery is important for urban planning, disaster management, and geographic information systems. However, complex building shapes, occlusions, and scale variation continue to challenge conventional segmentation models. This paper presents SRB-Net, a U-Net-based framework that combines three complementary components: (1) strip pooling (SP) for long-range horizontal and vertical context; (2) residual multi-scale atrous spatial pyramid pooling (RMASPP) with squeeze-and-excitation (SE) blocks for multi-scale and channel-aware feature learning; and (3) a bottleneck attention module (BAM) for refining skip-connection features. The model was trained with the Adam optimizer and evaluated on the aerial and Satellite Dataset II subsets of the WHU Building Dataset. Among the evaluated baselines, SRB-Net achieved the best overall performance, reaching 98.83% accuracy and 90.12% Intersection over Union (IoU) on the aerial dataset and 98.28% accuracy and 70.89% IoU on Satellite Dataset II. These results show consistent performance improvements across the two evaluated WHU subsets while avoiding claims beyond the within-dataset experimental setting.
Downloads
References
[1] Ji, S., Wei, S., & Lu, M. (2019). Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set. IEEE Transactions on Geoscience and Remote Sensing, 57(1), 574-586. https://doi.org/10.1109/tgrs.2018.2858817 DOI: https://doi.org/10.1109/TGRS.2018.2858817
[2] Ji, S., Wei, S., & Lu, M. (2019). A scale robust convolutional neural network for automatic building extraction from aerial and satellite imagery. International Journal of Remote Sensing, 40(9), 3308-3322. https://doi.org/10.1080/01431161.2018.1528024 DOI: https://doi.org/10.1080/01431161.2018.1528024
[3] Lin, J., Jing, W., Song, H., & Chen, G. (2019). ESFNet: Efficient network for building extraction from high-resolution aerial images. IEEE Access, 7, 54285-54294. https://doi.org/10.1109/access.2019.2912822 DOI: https://doi.org/10.1109/ACCESS.2019.2912822
[4] Zhu, Q., Liao, C., Hu, H., Mei, X., & Li, H. (2021). MAP-net: Multiple attending path neural network for building footprint extraction from remote sensed imagery. IEEE Transactions on Geoscience and Remote Sensing, 59(7), 6169-6181. https://doi.org/10.1109/tgrs.2020.3026051 DOI: https://doi.org/10.1109/TGRS.2020.3026051
[5] Cai, J., & Chen, Y. (2021). MHA-net: Multipath hybrid attention network for building footprint extraction from high-resolution remote sensing imagery. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14, 5807-5817. https://doi.org/10.1109/jstars.2021.3084805 DOI: https://doi.org/10.1109/JSTARS.2021.3084805
[6] Wu, G., Guo, Z., Shi, X., Chen, Q., Xu, Y., Shibasaki, R., & Shao, X. (2018). A boundary regulated network for accurate roof segmentation and outline extraction. Remote Sensing, 10(8), 1195. https://doi.org/10.3390/rs10081195 DOI: https://doi.org/10.3390/rs10081195
[7] Wu, G., Shao, X., Guo, Z., Chen, Q., Yuan, W., Shi, X., Xu, Y., & Shibasaki, R. (2018). Automatic building segmentation of aerial imagery using multi-constraint fully convolutional networks. Remote Sensing, 10(3), 407. https://doi.org/10.3390/rs10030407 DOI: https://doi.org/10.3390/rs10030407
[8] Wei, S., Ji, S., & Lu, M. (2020). Toward automatic building footprint delineation from aerial images using CNN and regularization. IEEE Transactions on Geoscience and Remote Sensing, 58(3), 2178-2189. https://doi.org/10.1109/tgrs.2019.2954461 DOI: https://doi.org/10.1109/TGRS.2019.2954461
[9] Kang, W., Xiang, Y., Wang, F., & You, H. (2019). EU-net: An efficient fully convolutional network for building extraction from optical remote sensing images. Remote Sensing, 11(23), 2813. https://doi.org/10.3390/rs11232813 DOI: https://doi.org/10.3390/rs11232813
[10] Deng, W., Shi, Q., & Li, J. (2021). Attention-gate-based encoder–decoder network for automatical building extraction. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14, 2611-2620. https://doi.org/10.1109/jstars.2021.3058097 DOI: https://doi.org/10.1109/JSTARS.2021.3058097
[11] Jing, H., Sun, X., Wang, Z., Chen, K., Diao, W., & Fu, K. (2021). Fine building segmentation in high-resolution SAR images via selective pyramid dilated network. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14, 6608-6623. https://doi.org/10.1109/jstars.2021.3076085 DOI: https://doi.org/10.1109/JSTARS.2021.3076085
[12] Zhou, D., Wang, G., He, G., Yin, R., Long, T., Zhang, Z., Chen, S., & Luo, B. (2021). A large-scale mapping scheme for urban building from gaofen-2 images using deep learning and hierarchical approach. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14, 11530-11545. https://doi.org/10.1109/jstars.2021.3123398 DOI: https://doi.org/10.1109/JSTARS.2021.3123398
[13] Li, X., Yao, X., & Fang, Y. (2018). Building-A-Nets: Robust building extraction from high-resolution remote sensing images with adversarial networks. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 11(10), 3680-3687. https://doi.org/10.1109/jstars.2018.2865187 DOI: https://doi.org/10.1109/JSTARS.2018.2865187
[14] Xia, L., Zhang, X., Zhang, J., Yang, H., & Chen, T. (2021). Building extraction from very-high-resolution remote sensing images using semi-supervised semantic edge detection. Remote Sensing, 13(11), 2187. https://doi.org/10.3390/rs13112187 DOI: https://doi.org/10.3390/rs13112187
[15] Chen, J., He, F., Zhang, Y., Sun, G., & Deng, M. (2020). SPMF-net: Weakly supervised building segmentation by combining superpixel pooling and multi-scale feature fusion. Remote Sensing, 12(6), 1049. https://doi.org/10.3390/rs12061049 DOI: https://doi.org/10.3390/rs12061049
[16] Liu, H., Luo, J., Huang, B., Hu, X., Sun, Y., Yang, Y., Xu, N., & Zhou, N. (2019). DE-net: Deep encoding network for building extraction from high-resolution remote sensing imagery. Remote Sensing, 11(20), 2380. https://doi.org/10.3390/rs11202380 DOI: https://doi.org/10.3390/rs11202380
[17] Chen, M., Wu, J., Liu, L., Zhao, W., Tian, F., Shen, Q., Zhao, B., & Du, R. (2021). DR-net: An improved network for building extraction from high resolution remote sensing image. Remote Sensing, 13(2), 294. https://doi.org/10.3390/rs13020294 DOI: https://doi.org/10.3390/rs13020294
[18] Guo, M., Liu, H., Xu, Y., & Huang, Y. (2020). Building extraction based on U-net with an attention block and multiple losses. Remote Sensing, 12(9), 1400. https://doi.org/10.3390/rs12091400 DOI: https://doi.org/10.3390/rs12091400
[19] Hou, X., Wang, P., & An, W. (2022). Multi-scale residual network for building extraction from satellite remote sensing images. In IGARSS 2022 - 2022 IEEE International Geoscience and Remote Sensing Symposium (pp. 1348-1351). IEEE. https://doi.org/10.1109/igarss46834.2022.9883509 DOI: https://doi.org/10.1109/IGARSS46834.2022.9883509
[20] Wang, H., & Miao, F. (2022). Building extraction from remote sensing images using deep residual U-net. European Journal of Remote Sensing, 55(1), 71-85. https://doi.org/10.1080/22797254.2021.2018944 DOI: https://doi.org/10.1080/22797254.2021.2018944
[21] Li, Z., Zhang, X., Xiao, P., & Zheng, Z. (2021). On the effectiveness of weakly supervised semantic segmentation for building extraction from high-resolution remote sensing imagery. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14, 3266-3281. https://doi.org/10.1109/jstars.2021.3063788 DOI: https://doi.org/10.1109/JSTARS.2021.3063788
[22] Yuan, Y., Shi, X., & Gao, J. (2025). Building extraction from remote sensing images with deep learning: A survey on vision techniques. Computer Vision and Image Understanding, 251, 104253. https://doi.org/10.1016/j.cviu.2024.104253 DOI: https://doi.org/10.1016/j.cviu.2024.104253
[23] Li, Q., Mou, L., Sun, Y., Hua, Y., Shi, Y., & Zhu, X. X. (2024). A review of building extraction from remote sensing imagery: Geometrical structures and semantic attributes. IEEE Transactions on Geoscience and Remote Sensing, 62, 1-15. https://doi.org/10.1109/tgrs.2024.3369723 DOI: https://doi.org/10.1109/TGRS.2024.3369723
[24] Xu, B., Xu, J., Xue, N., & Xia, G. (2023). HiSup: Accurate polygonal mapping of buildings in satellite imagery with hierarchical supervision. ISPRS Journal of Photogrammetry and Remote Sensing, 198, 284-296. https://doi.org/10.1016/j.isprsjprs.2023.03.006 DOI: https://doi.org/10.1016/j.isprsjprs.2023.03.006
[25] Li, W., Zhao, W., Yu, J., Zheng, J., He, C., Fu, H., & Lin, D. (2023). Joint semantic–geometric learning for polygonal building segmentation from high-resolution remote sensing images. ISPRS Journal of Photogrammetry and Remote Sensing, 201, 26-37. https://doi.org/10.1016/j.isprsjprs.2023.05.010 DOI: https://doi.org/10.1016/j.isprsjprs.2023.05.010
[26] Li, Q., Mou, L., Hua, Y., Shi, Y., & Zhu, X. X. (2022). Building footprint generation through convolutional neural networks with attraction field representation. IEEE Transactions on Geoscience and Remote Sensing, 60, 1-17. https://doi.org/10.1109/tgrs.2021.3109844 DOI: https://doi.org/10.1109/TGRS.2021.3109844
[27] Zhang, L., Bai, M., Liao, R., Urtasun, R., Marcos, D., Tuia, D., & Kellenberger, B. (2018). Learning deep structured active contours end-to-end. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 8877-8885). IEEE. https://doi.org/10.1109/cvpr.2018.00925 DOI: https://doi.org/10.1109/CVPR.2018.00925
[28] Girard, N., Smirnov, D., Solomon, J., & Tarabalka, Y. (2021). Polygonal building extraction by frame field learning. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5887-5896). IEEE. https://doi.org/10.1109/cvpr46437.2021.00583 DOI: https://doi.org/10.1109/CVPR46437.2021.00583
[29] Woo, S., Park, J., Lee, J., & Kweon, I. S. (2018). CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 3-19). Springer. https://doi.org/10.1007/978-3-030-01234-2_1 DOI: https://doi.org/10.1007/978-3-030-01234-2_1
[30] Hou, Q., Zhang, L., Cheng, M., & Feng, J. (2020). Strip pooling: Rethinking spatial pooling for scene parsing. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4002-4011). IEEE. https://doi.org/10.1109/cvpr42600.2020.00406 DOI: https://doi.org/10.1109/CVPR42600.2020.00406
[31] Cheng, D., Liao, R., Fidler, S., & Urtasun, R. (2019). DARNet: Deep active ray network for building segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7431-7439). IEEE. DOI: https://doi.org/10.1109/CVPR.2019.00761
[32] Chen, S., Shi, W., Zhou, M., Zhang, M., & Xuan, Z. (2022). CGSANet: A contour-guided and local structure-aware encoder–decoder network for accurate building extraction from very high-resolution remote sensing imagery. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 15, 1526-1542. https://doi.org/10.1109/jstars.2021.3139017 DOI: https://doi.org/10.1109/JSTARS.2021.3139017
[33] Chen, J., Zhang, D., Wu, Y., Chen, Y., & Yan, X. (2022). A context feature enhancement network for building extraction from high-resolution remote sensing imagery. Remote Sensing, 14(9), 2276. https://doi.org/10.3390/rs14092276 DOI: https://doi.org/10.3390/rs14092276
[34] Fang, F., Zheng, D., Li, S., Liu, Y., Zeng, L., Zhang, J., & Wan, B. (2022). Improved pseudomasks generation for weakly supervised building extraction from high-resolution remote sensing imagery. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 15, 1629-1642. https://doi.org/10.1109/jstars.2022.3144176 DOI: https://doi.org/10.1109/JSTARS.2022.3144176
[35] Borba, P., De Carvalho Diniz, F., Da Silva, N. C., & De Souza Bias, E. (2021). Building footprint extraction using deep learning semantic segmentation techniques: Experiments and results. In 2021 IEEE International Geoscience and Remote Sensing Symposium IGARSS (pp. 4708-4711). IEEE. https://doi.org/10.1109/igarss47720.2021.9553855 DOI: https://doi.org/10.1109/IGARSS47720.2021.9553855
[36] Miao, Y., Jiang, S., Xu, Y., & Wang, D. (2022). Feature residual analysis network for building extraction from remote sensing images. Applied Sciences, 12(10), 5095. https://doi.org/10.3390/app12105095 DOI: https://doi.org/10.3390/app12105095
[37] Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI) (pp. 234-241). Springer. https://doi.org/10.1007/978-3-319-24574-4_28 DOI: https://doi.org/10.1007/978-3-319-24574-4_28
[38] Yu, F., & Koltun, V. (2016). Multi-scale context aggregation by dilated convolutions. In International Conference on Learning Representations (ICLR). https://arxiv.org/abs/1511.07122
[39] Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7132-7141). IEEE. https://doi.org/10.1109/CVPR.2018.00745 DOI: https://doi.org/10.1109/CVPR.2018.00745
[40] Park, J., Woo, S., Lee, J.-Y., & Kweon, I. S. (2018). BAM: Bottleneck attention module. In British Machine Vision Conference (BMVC). https://arxiv.org/abs/1807.06514
[41] Badrinarayanan, V., Kendall, A., & Cipolla, R. (2017). SegNet: A deep convolutional encoder–decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481-2495. https://doi.org/10.1109/TPAMI.2016.2644615 DOI: https://doi.org/10.1109/TPAMI.2016.2644615
[42] Zhou, Z., Siddiquee, M. M. R., Tajbakhsh, N., & Liang, J. (2020). UNet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE Transactions on Medical Imaging, 39(6), 1856-1867. https://doi.org/10.1109/TMI.2019.2959609 DOI: https://doi.org/10.1109/TMI.2019.2959609
[43] Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder–decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 801-818). Springer. https://doi.org/10.1007/978-3-030-01234-2_49 DOI: https://doi.org/10.1007/978-3-030-01234-2_49
[44] Li, R., Zheng, S., Zhang, C., Duan, C., & Wang, L. (2022). A2-FPN for semantic segmentation of fine-resolution remotely sensed images. International Journal of Remote Sensing, 43(3), 1131-1155. https://doi.org/10.1080/01431161.2022.2030071 DOI: https://doi.org/10.1080/01431161.2022.2030071
[45] Wang, L., Li, R., Zhang, C., Fang, S., Duan, C., Meng, X., & Atkinson, P. M. (2022). UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery. ISPRS Journal of Photogrammetry and Remote Sensing, 190, 196-214. https://doi.org/10.1016/j.isprsjprs.2022.06.008 DOI: https://doi.org/10.1016/j.isprsjprs.2022.06.008
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Computing and Electronic Information Management

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








