Journal of Shandong University(Engineering Science) ›› 2026, Vol. 56 ›› Issue (4): 84-93.doi: 10.6040/j.issn.1672-3961.0.2025.114

• Machine Learning & Data Mining • Previous Articles    

CBAM-U-Net: segmentation of potato pollen images based on U-Net and self-attention mechanism

Shen Yajie1, Xia Lu1, Li Jie1, Tang Mingjing1,2*   

  1. Shen Yajie1, Xia Lu1, Li Jie1, Tang Mingjing1, 2*(1. School of Information Science and Technology, Yunnan Normal University, Kunming 650500, Yunnan, China;
    2. Yunnan Provincial Key Laboratory of Potato Biology, Yunnan Normal University, Kunming 650500, Yunnan, China
  • Published:2026-08-12

Abstract: To address the low efficiency, subjectivity, and poor adaptability to complex backgrounds of traditional manual observation and threshold-based methods in potato pollen microscopy image segmentation and counting, a pollen segmentation method integrating the convolutional block attention module(CBAM)and U-Net(CBAM-U-Net)was proposed. Built on the U-Net framework, the CBAM was introduced to improve the extraction of key features and edge details in pollen regions, while median filtering and histogram equalization were applied for image preprocessing. Experiments on 44 high-resolution potato pollen images and more than 5 000 microscopic images showed that the proposed method achieved accurate pollen segmentation, improved segmentation efficiency by about five times over traditional methods, and demonstrated good robustness and generalization ability on large-scale datasets. The proposed method could improve the automation and accuracy of potato pollen microscopy image analysis and provide technical support for microscopic image analysis in crop breeding, cell biology, and precision agriculture.

Key words: U-Net, CBAM, potato pollen, high throughput segmentation, threshold segmentation

CLC Number: 

  • TP391.9
[1] Singh A, Ganapathysubramanian B, Singh A K, et al. Machine learning for high-throughput stress phenotyping in plants[J]. Trends in Plant Science, 2016, 21(2): 110-124.
[2] Lin E, Lane H Y. Machine learning and systems genomics approaches for multi-omics data[J]. Biomarker Research, 2017, 5(1): 2.
[3] Tester M, Langridge P. Breeding technologies to increase crop production in a changing world[J]. Science, 2010, 327(5967): 818-822.
[4] Ribaut J M, De Vicente M C, Delannay X. Molecular breeding in developing countries: challenges and perspectives[J]. Current Opinion in Plant Biology, 2010, 13(2): 213-218.
[5] Boavida L C, Vieira A M, Becker J D, et al. Gametophyte interaction and sexual reproduction: how plants make a zygote[J]. The International Journal of Developmental Biology, 2005, 49(5/6): 615-632.
[6] Fetter K C, Eberhardt S, Barclay R S, et al. StomataCounter: a neural network for automatic stomata identification and counting[J]. New Phytologist, 2019, 223(3): 1671-1681.
[7] Krizhevsky A, Sutskever I, Hinton G E. ImageNet classification with deep convolutional neural networks[J]. Communications of the ACM, 2017, 60(6): 84-90.
[8] Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation[C] //Medical Image Computing and Computer-Assisted Intervention: MICCAI 2015. Cham, Switzerland: Springer, 2015: 234-241.
[9] Colmer J, O'neill C M, Wells R, et al. SeedGerm: a cost-effective phenotyping platform for automated seed imaging and machine-learning based phenotypic analysis of crop seed germination[J]. New Phytologist, 2020, 228(2): 778-793.
[10] Tello J, Montemayor M I, Forneck A, et al. A new image-based tool for the high throughput phenotyping of pollen viability: evaluation of inter- and intra-cultivar diversity in grapevine[J]. Plant Methods, 2018, 14: 3.
[11] Maraci M A, Bridge C P, Napolitano R, et al. A framework for analysis of linear ultrasound videos to detect fetal presentation and heartbeat[J]. Medical Image Analysis, 2017, 37: 22-36.
[12] Zhao Z X, Chen K X, Yamane S. CBAM-Unet++: easier to find the target with the attention module "CBAM"[C] //2021 IEEE 10th Global Conference on Consumer Electronics(GCCE). Kyoto, Japan: IEEE, 2021: 655-657.
[13] Taud H, Mas J F. Multilayer perceptron(MLP)[M] // Camacho Olmedo M T, Paegelow M, Mas J F, et al. Geomatic approaches for modeling land change scenarios. Cham, Switzerland: Springer, 2017: 451-455.
[14] 周逸凡, 张灵维, 周正东, 等. 基于注意力机制和深度学习的群体语言想象脑电信号分类[J]. 浙江大学学报(工学版), 2024, 58(12): 2540-2546. Zhou Yifan, Zhang Lingwei, Zhou Zhengdong, et al. Classification of group speech imagined EEG signals based on attention mechanism and deep learning[J]. Journal of Zhejiang University(Engineering Science), 2024, 58(12): 2540-2546.
[15] 于贺婷, 刘思萌, 文峰. 基于CBAM注意力机制的智能交通信号控制[J]. 沈阳理工大学学报, 2024, 43(5): 34-40. Yu Heting, Liu Simeng, Wen Feng. Intelligent traffic control technology based on CBAM attention mechanism[J]. Journal of Shenyang Ligong University, 2024, 43(5): 34-40.
[16] 冯庆贺. 面向图像检索的底层视觉与深度卷积特征提取方法研究[D]. 沈阳: 东北大学, 2020: 13-15. Feng Qinghe. Research on low-level vision and deep convolutional feature extraction methods for image retrieval[D]. Shenyang: Northeastern University, 2020: 13-15.
[17] 王军, 张霁云, 程勇. 基于边缘特征和注意力机制的图像语义分割[J]. 计算机系统应用, 2024, 33(7): 63-73. Wang Jun, Zhang Jiyun, Cheng Yong. Image semantic segmentation based on edge features and attention mechanism[J]. Computer Systems & Applications, 2024, 33(7): 63-73.
[18] He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition[C] //2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Las Vegas, USA: IEEE, 2016: 770-778.
[19] He K M, Zhang X Y, Ren S Q, et al. Identity mappings in deep residual networks[C] //Computer Vision-ECCV 2016. Cham, Switzerland: Springer, 2016: 630-645.
[20] 李顺勇, 胥瑞, 李师毅. 加入跳跃连接的深度嵌入K-means聚类[J]. 计算机系统应用, 2024, 33(1): 11-21. Li Shunyong, Xu Rui, Li Shiyi. Deep embedded K-means clustering with skip connections[J]. Computer Systems & Applications, 2024, 33(1): 11-21.
[21] 皮磊, 朱磊, 郑翔, 等. 基于改进Wave-U-Net跳跃连接的盲源分离算法[J]. 信号处理, 2022, 38(4): 835-843. Pi Lei, Zhu Lei, Zheng Xiang, et al. Blind source separation algorithm based on improved Wave-U-Net skip connection[J]. Journal of Signal Processing, 2022, 38(4): 835-843.
[22] Creswell A, Bharath A A. Denoising adversarial autoencoders[J]. IEEE Transactions on Neural Networks and Learning Systems, 2019, 30(4): 968-984.
[23] 史加荣, 王丹, 尚凡华, 等. 随机梯度下降算法研究进展[J]. 自动化学报, 2021, 47(9): 2103-2119. Shi Jiarong, Wang Dan, Shang Fanhua, et al. Research advances on stochastic gradient descent algorithms[J]. Acta Automatica Sinica, 2021, 47(9): 2103-2119.
[24] 赵高长, 张磊, 武风波. 改进的中值滤波算法在图像去噪中的应用[J]. 应用光学, 2011, 32(4): 678-682. Zhao Gaochang, Zhang Lei, Wu Fengbo. Application of improved median filtering algorithm to image denoising[J]. Journal of Applied Optics, 2011, 32(4): 678-682.
[25] 毛本清, 金小梅. 自适应直方图均衡化算法在图像增强处理的应用[J]. 河北北方学院学报(自然科学版), 2010, 26(5): 64-68. Mao Benqing, Jin Xiaomei. Application of self-adaptive histogram equalization algorithm to image enhancement processing[J]. Journal of Hebei North University(Natural Science Edition), 2010, 26(5): 64-68.
[26] Van Valen D A, Kudo T, Lane K M, et al. Deep learning automates the quantitative analysis of individual cells in live-cell imaging experiments[J]. PLoS Computational Biology, 2016, 12(11): e1005177.
[27] Liu X M, Zhao D B, Xiong R Q, et al. Image interpolation via regularized local linear regression[J]. IEEE Transactions on Image Processing, 2011, 20(12): 3455-3469.
[28] 彭程, 李帅, 苗艳龙, 等. 基于三维点云的番茄植株茎叶分割与表型特征提取[J]. 农业工程学报, 2022, 38(9): 187-194. Peng Cheng, Li Shuai, Miao Yanlong, et al. Stem-leaf segmentation and phenotypic trait extraction of tomatoes using three-dimensional point cloud[J]. Transactions of the Chinese Society of Agricultural Engineering, 2022, 38(9): 187-194.
[29] Ranefall P, Wählby C. Global gray-level thresholding based on object size[J]. Cytometry Part A, 2016, 89(4): 385-390.
[30] Song J, Jiao W, Lankowicz K, et al. A two-stage adaptive thresholding segmentation for noisy low-contrast images[J]. Ecological Informatics, 2022, 69: 101632.
[31] Ye J, Xu G. Geometric flow approach for region-based image segmentation[J]. IEEE Transactions on Image Processing, 2012, 21(12): 4735-4745.
[32] Phornphatcharaphong W, Eua-Anant N. Edge-based color image segmentation using particle motion in a vector image field derived from local color distance images[J]. Journal of Imaging, 2020, 6(7): 72.
[33] Xing J W, Yang P, Qing G L T. Robust 2D Otsu's algorithm for uneven illumination image seg-mentation[J]. Computational Intelligence and Neuro-science, 2020, 2020(1): 5047976.
[34] Huang S Y, Hsu W L, Hsu R J, et al. Fully convolutional network for the semantic segmentation of medical images: a survey[J]. Diagnostics, 2022, 12(11): 2765.
[35] Badrinarayanan V, Kendall A, Cipolla R. SegNet: a deep convolutional encoder-decoder architecture for image segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(12): 2481-2495.
[36] Stringer C, Wang T, Michaelos M, et al. Cellpose: a generalist algorithm for cellular segmentation[J]. Nature Methods, 2021, 18(1): 100-106.
[37] Stringer C, Pachitariu M. Cellpose3: one-click image restoration for improved cellular segmentation[J]. Nature Methods, 2025, 22(3): 592-599.
[38] 田萱, 王亮, 丁琪. 基于深度学习的图像语义分割方法综述[J]. 软件学报, 2019, 30(2): 440-468. Tian Xuan, Wang Liang, Ding Qi. Review of image semantic segmentation based on deep learning[J]. Journal of Software, 2019, 30(2): 440-468.
[39] 李刚森. 基于深度学习的细胞核图像分割方法研究[D]. 黑龙江: 哈尔滨工业大学, 2018: 13-15. Li Gangsen. Methodology research of nucleus image segmentation based on deep learning[D]. Heilongjiang: Harbin Institute of Technology, 2018: 13-15.
[40] Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C] //2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018: 7132-7141.
[41] Shan T, Yan J Y. SCA-Net: a spatial and channel attention network for medical image segmentation[J]. IEEE Access, 2021, 9: 160926-160937.
[42] Hou Q B, Zhou D Q, Feng J S. Coordinate attention for efficient mobile network design[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2021: 13708-13717.
[43] Misra D, Nalamada T, Arasanipalai A U, et al. Rotate to attend: convolutional triplet attention module[C] //2021 IEEE Winter Conference on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2021: 3138-3147.
[44] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C] //Proceedings of the 31st Conference on Neural Information Processing Systems(NIPS 2017). Long Beach, USA: Curran Associates, Inc., 2017: 5998-6008.
[45] 赵海丽, 包大泱, 张从豪, 等. 基于改进SDU-YOLOv8的军事飞机目标检测算法[J]. 兵工学报, 2026, 47(1): 250294. Zhao Haili, Bao Dayang, Zhang Conghao, et al. Military aircraft object detection algorithm based on improved SDU-YOLOv8[J]. Acta Armamentarii, 2026, 47(1): 250294.
[1] LI Erchao, ZHANG Zhizhao. Online dynamic demand vehicle routing planning [J]. Journal of Shandong University(Engineering Science), 2024, 54(5): 62-73.
[2] Si YANG, Sitong LI, Jindong ZHANG, Yu BAI. Improvement of bandwidth model for high speed optical communicationlaser and its optimization by parallel computing [J]. Journal of Shandong University(Engineering Science), 2019, 49(1): 17-22.
[3] HUANG Yanhui, FAN Yangyu, SU Xuhui. 3D facial expression tracking using a monocular RGB camera [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2017, 47(4): 7-13.
[4] SHI Wen-Hua, LIU Wei-Dong, SUN Yong-Fu. Research of 1/3 dam breach simulation and personnel evacuation scenario based on digital elevation model DEM in a quake lake [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2009, 39(5): 144-148.
[5] LI Fangjia, GAO Shangce, TANG Zheng*, Ishii Masahiro, Yamashita Kazuya. 3D similar pattern generation of snow crystals with cellular automata [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2009, 39(1): 102-105.
[6] CHEN Cheng-jun,ZHOU Yi-qi,YANG Hong-juan . Study on an approach of transformation and representation based on the SolidWorks model to the virtual assembly model [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2008, 38(1): 61-65 .
[7] SONG Qing,LI Xiao-lei,ZHANG Cheng-jin . Optimization of a postal express mail network based on bottleneck analysis [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2007, 37(5): 29-33 .
[8] ZHAO Wei-hua,WANG Yong,WANG Xian-lun . Space continuous path kinematic simulation of a robot based on VRML [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2007, 37(2): 8-11 .
[9] KONG Xiang-zhen,LIU Yan-jun,WANG Yong,ZHAO Xiu-hua . Compensation and simulation for the deadband of the pneumatic proportional valve [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2006, 36(1): 99-102 .
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!