Journal of Shandong University(Engineering Science) ›› 2026, Vol. 56 ›› Issue (4): 27-37.doi: 10.6040/j.issn.1672-3961.0.2025.109

• Machine Learning & Data Mining • Previous Articles    

Pcapose: semi-supervised animal pose estimation method based on pseudo-labels and consistency training

Zhu Zhaoli1, Zhang Jikai1*, Zeng Xianghao1, Xie Chenjie1, Li Jianbin2   

  1. Zhu Zhaoli1, Zhang Jikai1*, Zeng Xianghao1, Xie Chenjie1, Li Jianbin2(1. School of Digital and Intelligent Industry, Inner Mongolia University of Science and Technology, Baotou 014010, Inner Mongolia, China;
    2. Institute of Animal Science and Veterinary Medicine, Shandong Academy of Agricultural Sciences, Jinan 250100, Shandong, China
  • Published:2026-08-12

Abstract: To address the core challenges such as the difficulty in obtaining large scale labeled data, high costs, and data scarcity in animal pose estimation, a semi-supervised animal pose estimation method based on pseudo-labels and consistency training(Pcapose)was proposed to enhance the model's performance and generalization ability under limited labeled data conditions. An initial model was trained based on the initial labeled data. By jointly using a large amount of unlabeled data, a triple mechanism was constructed to improve the quality of pseudo-labels and the robustness of the model. A confidence driven pseudo-label screening strategy was adopted to select samples with low loss and high confidence to expand the training set. A multi-view consistency detection mechanism was introduced to integrate perturbation information from geometry, illumination, and feature space for re-evaluating and screening samples with high loss and low confidence. A teacher-student consistency framework was built to ensure the teacher model provided stable and accurate pseudo-labels. The experimental results on the AP-10K and Grévy's Zebra datasets showed that Pcapose significantly outperformed existing semi-supervised methods in terms of key point detection accuracy and model robustness, demonstrating superior performance in data scarce scenarios.

Key words: animal pose estimation, semi-supervised learning, pseudo-label screening, multi-view consistency, teacher-student consistency, coordinate classification, data scarcity

CLC Number: 

  • TP391.41
[1] Cao J K, Tang H Y, Fang H S, et al. Cross-domain adaptation for animal pose estimation[C] //2019 IEEE/CVF International Conference on Computer Vision(ICCV). Seoul: IEEE, 2019: 9497-9506.
[2] Li C, Lee G H. From synthetic to real: unsupervised domain adaptation for animal pose estimation[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2021: 1482-1491.
[3] Yu H, Xu Y F, Zhang J, et al. AP-10K: a benchmark for animal pose estimation in the wild[EB/OL].(2021-11-01)[2025-03-22]. https://arxiv.org/abs/2108.12617
[4] Shooter M, Malleson C, Hilton A. SyDog: a synthetic dog dataset for improved 2D pose estimation[EB/OL].(2021-07-31)[2025-03-22]. https://arxiv.org/abs/2108.00249
[5] Lee D H. Pseudo-label: the simple and efficient semi-supervised learning method for deep neural networks[C] //Proceedings of the 2013 ICML Workshop: Challenges in Representation Learning. Atlanta, USA: [s.n.] , 2013: 896.
[6] Scudder H. Probability of error of some adaptive pattern-recognition machines[J]. IEEE Transactions on Information Theory, 1965, 11(3): 363-371.
[7] Tarvainen A, Valpola H. Mean teachers are better role models: weight-averaged consistency targets improve semi-supervised deep learning results[EB/OL].(2015-04-16)[2025-03-22]. https://arxiv.org/abs/1703.01780
[8] Xie Q Z, Dai Z H, Hovy E, et al. Unsupervised data augmentation for consistency training[EB/OL].(2020-11-05)[2025-03-22]. https://arxiv.org/abs/1904.12848
[9] Nassar I, Hayat M, Abbasnejad E, et al. ProtoCon: pseudo-label refinement via online clustering and prototypical consistency for efficient semi-supervised learning[EB/OL].(2023-03-22)[2025-03-22]. https://arxiv.org/abs/2303.13556
[10] Yu Z R, Wang M C, Chen Y B, et al. Denoising and selecting pseudo-heatmaps for semi-supervised human pose estimation[C] //2024 IEEE/CVF Winter Confe-rence on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2024: 6268-6277.
[11] Li C, Lee G H. ScarceNet: animal pose estimation with scarce annotations[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Vancouver, Canada: IEEE, 2023: 17174-17183.
[12] Deng J H, Li W, Chen Y H, et al. Unbiased mean teacher for cross-domain object detection[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2021: 4089-4099.
[13] Cao S C, Joshi D, Gui L Y, et al. Contrastive mean teacher for domain adaptive object detectors[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Vancouver, Canada: IEEE, 2023: 23839-23848.
[14] 郭敏, 张熙涵, 李阳. 融合注意力的教师互一致性半监督医学图像分割[J]. 计算机工程, 2024, 50(9): 313-323. Guo Min, Zhang Xihan, Li Yang. Integrated attentional teacher mutual consistency semi-supervised medical image segmentation[J]. Computer Engineering, 2024, 50(9): 313-323.
[15] Jiang T, Lu P, Zhang L, et al. RTMPose: real-time multi-person pose estimation based on MMPose[EB/OL].(2023-07-03)[2025-03-22]. https://arxiv.org/abs/2303.07399
[16] Wang J D, Sun K, Cheng T H, et al. Deep high-resolution representation learning for visual recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(10): 3349-3364.
[17] Biggs B, Boyne O, Charles J, et al. Who left the dogs out? 3D animal reconstruction with expectation maximization in the loop[C] //Computer Vision-ECCV 2020: 16th European Conference. Glasgow, UK: Springer, 2020: 195-211.
[18] Li C, Lee G H. Coarse-to-fine animal pose and shape estimation[C] //Proceedings of the 35th International Conference on Neural Information Processing Systems. Red Hook, USA: ACM, 2021: 11757-11768.
[19] Lin T Y, Maire M, Belongie S, et al. Microsoft COCO: common objects in context[C] //Computer Vision-ECCV 2014: 13th European Conference. Zurich, Switzerland: Springer, 2014: 740-755.
[20] Rao J Y, Zhao B N, Wang Y. Probabilistic prompt distribution learning for animal pose estimation[C] //2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2025: 29438-29447.
[21] Xu T Y, Rao J Y, Song X N, et al. Learning structure-supporting dependencies via keypoint interactive Transformer for general mammal pose estimation[J]. International Journal of Computer Vision, 2025, 133(7): 3858-3876.
[22] Zhang W J, Liu D N, Cai W D, et al. Cross-view consistency regularisation for knowledge distillation[C] //Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024: 2011-2020.
[23] Sohn K, Berthelot D, Carlini N, et al. FixMatch: simplifying semi-supervised learning with consistency and confidence[EB/OL].(2020-11-25)[2025-03-22]. https://arxiv.org/abs/2001.07685
[24] Arazo E, Ortego D, Albert P, et al. Pseudo-labeling and confirmation bias in deep semi-supervised learning[C] //2020 International joint conference on neural networks(IJCNN). Glasgow, UK: IEEE, 2020: 9207304.
[25] 李萍, 张雪英, 王夙喆, 等. 基于半监督多尺度一致性学习的医学影像分割[J]. 计算机工程, 2025, 51(10): 295-307. Li Ping, Zhang Xueying, Wang Suzhe, et al. Medical image segmentation based on semi-supervised multi-scale consistency learning[J]. Computer Engineering, 2025, 51(10): 295-307.
[26] Sindhwani V, Niyogi P, Belkin M. A co-regularization approach to semi-supervised learning with multiple views[C] //Proceedings of ICML Workshop on Learning with Multiple Views. Bonn, Germany: [s.n.] , 2005: 74-79.
[27] 景攀峰, 梁宇栋, 李超伟, 等. 基于师生学习的半监督图像去雾算法[J]. 计算机应用, 2025, 45(9): 2975-2983. Jing Panfeng, Liang Yudong, Li Chaowei, et al. Semi-supervised image dehazing algorithm based on teacher-student learning [J]. Journal of Computer Applications, 2025, 45(9): 2975-2983.
[28] Li Y J, Yang S, Liu P D, et al. SimCC: a simple coordinate classification perspective for human pose estimation[EB/OL].(2022-07-05)[2025-03-22]. https://arxiv.org/abs/2107.03332
[29] Liu Y T, Wen Q, Chen H X, et al. Crowd counting via cross-stage refinement networks[J]. IEEE Transactions on Image Processing, 2020, 29: 6800-6812.
[30] Rodríguez P, Gonfaus J M, Cucurull G, et al. Attend and rectify: a gated attention mechanism for fine-grained recovery[C] //Computer Vision-ECCV 2018. Munich, Germany: Springer, 2018: 357-372.
[31] Graving J M, Chae D, Naik H, et al. Fast and robust animal pose estimation[EB/OL].(2019-04-26)[2025-03-22]. https://www.biorxiv.org/content/10.1101/620245v1
[32] Xie R C, Wang C Y, Zeng W J, et al. An empirical study of the collapsing problem in semi-supervised 2D human pose estimation[C] //2021 IEEE/CVF International Conference on Computer Vision(ICCV). Montreal, Canada: IEEE, 2022: 11220-11229.
[33] Schmutz H, Humbert O, Mattei P A. Don't fear the unlabelled: safe semi-supervised learning via simple debiasing[EB/OL].(2023-03-03)[2025-03-22]. https://arxiv.org/abs/2203.07512
[34] Scherer S, Schön R, Lienhart R. Pseudo-label noise suppression techniques for semi-supervised semantic segmentation[EB/OL].(2022-10-19)[2025-03-22]. https://arxiv.org/abs/2210.10426
[35] Nguyen K B. Debiasing, calibrating, and improving semi-supervised learning performance via simple ensemble projector[C] //2024 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2024: 2430-2439.
[36] Shen J Q, Jiang Y N, Luo J W, et al. MPE-HRNetL: a lightweight high-resolution network for multispecies animal pose estimation[J]. Sensors, 2024, 24(21): 6882.
[37] He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition[C] //2016 IEEE Confe-rence on Computer Vision and Pattern Recognition(CVPR). Las Vegas, USA: IEEE, 2016: 770-778.
[1] ZHU Hengdong, MA Yingcang, DAI Xuezhen. Adaptive semi-supervised neighborhood clustering algorithm [J]. Journal of Shandong University(Engineering Science), 2021, 51(4): 24-34.
[2] KONG Chao1,2, ZHANG Huaxiang1,2*, LIU Li1,2. A semi-supervised image retrieval algorithm based onfeature fusion of the region of interest [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2014, 44(3): 22-28.
[3] LI Ya-lin1,2, ZHANG Hua-xiang1,2*, FENG Xin-ying1,2. A new multi-label learning algorithm based on semi-supervised learning [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2013, 43(2): 18-22.
[4] XIA Zhan-guo, WAN Ling, CAI Shi-yu, SUN Peng-hui. A semi-supervised clustering algorithm oriented to intrusion detection [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2012, 42(6): 1-7.
[5] XIE Huo-sheng, LIU Min. An ensemble co-training algorithm based on active learning [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2012, 42(3): 1-5.
[6] WEI Wei, ZHANG Yanning. Pose estimation based on semi-supervised latent Dirichlet allocation [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2011, 41(3): 17-22.
[7] SU Hong-lu, LI Fan-zhang*. Semi-supervised image retrieval based on diversity and invariant features [J]. JOURNAL OF SHANDONG UNIVERSITY (ENGINEERING SCIENCE), 2010, 40(5): 150-153.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!