A Lightweight Feature Transformation Network with Progressive Training for Rainy Object Detection
DOI:
https://doi.org/10.54097/v3s89j21Keywords:
Adverse weather object detection, Knowledge distillation, Joint trainingAbstract
Rainy weather causes occlusions and blur in images, severely degrading the performance of object detectors deployed in scenarios such as autonomous driving. Existing methods mostly adopt pixel-level deraining preprocessing or adversarial domain adaptation. The former introduces additional computational overhead and error accumulation, while the latter suffers from training instability and tends to forget clear-weather knowledge. This paper proposes a lightweight feature transformation module embedded in the detector backbone, which maps "rainy features" to "rain-free features" in the feature space, allowing the subsequent detection network to adapt to rainy conditions without any modification. To achieve both lightweight design and high accuracy, a knowledge distillation framework based on intermediate feature maps is introduced, where a large-capacity teacher network guides a small student network to learn the feature mapping. Meanwhile, a three-step progressive training strategy is designed to sequentially train the detector, distill the transformation module, and fine-tune the overall network, ensuring training stability and avoiding catastrophic forgetting. Extensive experiments show that our method has essentially reached the rain-free baseline level, while significantly reducing computational cost compared to preprocessing approaches. This work provides a new perspective for the efficient and robust deployment of vision systems under adverse weather conditions.
Downloads
References
[1] Ren, S., He, K., Girshick, R., & Sun, J. (2017). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6), 1137-1149. https://doi.org/10.1109/TPAMI.2016.2577031
[2] Tian, Z., Shen, C., Chen, H., & He, T. (2019). FCOS: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 9627-9636). IEEE. https://doi.org/10.1109/ICCV.2019.00972
[3] Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 779-788). IEEE. https://doi.org/10.1109/CVPR.2016.91
[4] Redmon, J., & Farhadi, A. (2017). YOLO9000: Better, faster, stronger. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6517-6525). IEEE. https://doi.org/10.1109/CVPR.2017.690
[5] Redmon, J., & Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv, arXiv:1804.02767. https://doi.org/10.48550/arXiv.1804.02767
[6] Bochkovskiy, A., Wang, C. Y., & Liao, H. Y. M. (2020). YOLOv4: Optimal speed and accuracy of object detection. arXiv, arXiv:2004.10934. https://doi.org/10.48550/arXiv.2004.10934
[7] Li, C. Y., Li, L. L., Jiang, H. L., Weng, K., Geng, Y., Li, L., Ke, Z., Li, Q., Cheng, M., Nie, W., Li, Y., Zhang, B., Liang, Y., Zhou, L., Xu, X., Chu, X., Wei, X., & Wei, X. (2022). YOLOv6: A single-stage object detection framework for industrial applications. arXiv, arXiv:2209.02976. https://doi.org/10.48550/arXiv.2209.02976
[8] Wang, C. Y., Bochkovskiy, A., & Liao, H. Y. M. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv, arXiv:2207.02696. https://doi.org/10.48550/arXiv.2207.02696
[9] Wang, A., Chen, H., Liu, L., Chen, K., Lin, Z., Han, J., & Ding, G. (2024). YOLOv10: Real-time end-to-end object detection. arXiv, arXiv:2405.14458. https://doi.org/10.48550/arXiv.2405.14458
[10] Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer Vision – ECCV 2014 (Vol. 8693, pp. 740-755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48
[11] Everingham, M., Van Gool, L., Williams, C. K., Winn, J., & Zisserman, A. (2010). The Pascal Visual Object Classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303-308. https://doi.org/10.1007/s11263-009-0275-4
[12] Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., & Bengio, Y. (2013). An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv, arXiv:1312.6211. https://doi.org/10.48550/arXiv.1312.6211
[13] French, R. M. (1999). Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences, 3(4), 128-135. https://doi.org/10.1016/S1364-6613(99)01294-2
[14] Wang, K., Wang, T., Qu, J., Jiang, H., Li, Q., & Chang, L. (2022). An end-to-end cascaded image deraining and object detection neural network. IEEE Robotics and Automation Letters, 7(4), 9541-9548. https://doi.org/10.1109/LRA.2022.3191937
[15] Wang, T., Wang, K., & Li, Q. (2023). Denseformer for single image deraining. International Journal of Applied Mathematics and Computer Science, 33(4), 651-661. https://doi.org/10.34768/amcs-2023-0044
[16] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770-778). IEEE. https://doi.org/10.1109/CVPR.2016.90
[17] Romero, A., Ballas, N., Kahou, S. E., Chassang, G., Gatta, C., & Bengio, Y. (2014). FitNets: Hints for thin deep nets. arXiv, arXiv:1412.6550. https://doi.org/10.48550/arXiv.1412.6550
[18] Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer Vision – ECCV 2014 (Vol. 8693, pp. 740-755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48
[19] Zheng, S., Lu, C., Wu, Y., & Yang, Y. (2022). SAPNet: Segmentation-aware progressive network for perceptual contrastive deraining. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 52-62). IEEE. https://doi.org/10.1109/WACV51458.2022.00012
[20] Chen, X., Huang, Y., & Xu, L. (2020). Multi-scale attentive residual dense network for single image rain removal. In Proceedings of the Asian Conference on Computer Vision (ACCV). Springer.
[21] Wang, C., Xing, X., Wu, Y., Su, Z., & Chen, Y. (2020). DCSFN: Deep cross-scale fusion network for single image rain removal. In Proceedings of the 28th ACM International Conference on Multimedia (pp. 1643-1651). ACM. https://doi.org/10.1145/3394171.3413562
[22] Chen, L., Lu, X., Zhang, J., Chu, X., & Chen, C. (2021). HINet: Half-instance normalization network for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 182-192). IEEE. https://doi.org/10.1109/CVPR46437.2021.00025
[23] Zhao, H., Yap, K. H., Kot, A. C., & Duan, L. (2020). JDNet: A joint-learning distilled network for mobile visual food recognition. IEEE Journal of Selected Topics in Signal Processing, 14(4), 665-675. https://doi.org/10.1109/JSTSP.2020.2980852
[24] Zamir, S. W., Arora, A., Khan, S., Hayat, M., Khan, F. S., & Yang, M. H. (2021). Multi-stage progressive image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 14821-14831). IEEE. https://doi.org/10.1109/CVPR46437.2021.01458
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Computing and Electronic Information Management

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








