您的位置:山东大学 -> 科技期刊社 -> 《山东大学学报(工学版)》

山东大学学报 (工学版) ›› 2026, Vol. 56 ›› Issue (4): 27-37.doi: 10.6040/j.issn.1672-3961.0.2025.109

• 机器学习与数据挖掘 • 上一篇    

Pcapose:基于伪标签和一致性训练的半监督动物姿态估计方法

朱钊利1,张继凯1*,曾翔皓1,解辰杰1,李建斌2   

  1. 1.内蒙古科技大学数智产业学院, 内蒙古 包头 014010;2.山东省农业科学院畜牧兽医研究所, 山东 济南 250100
  • 发布日期:2026-08-12
  • 作者简介:朱钊利(1999— ),男,山东枣庄人,硕士研究生,主要研究方向为姿态估计. E-mail:614601867@qq.com. *通信作者简介:张继凯(1988— ),男,河北石家庄人,副教授,硕士生导师,博士,主要研究方向为计算机视觉. E-mail:jkzhang0314@imust.edu.cn
  • 基金资助:
    内蒙古自然科学基金资助项目(2024LHMS06007)

Pcapose: semi-supervised animal pose estimation method based on pseudo-labels and consistency training

Zhu Zhaoli1, Zhang Jikai1*, Zeng Xianghao1, Xie Chenjie1, Li Jianbin2   

  1. Zhu Zhaoli1, Zhang Jikai1*, Zeng Xianghao1, Xie Chenjie1, Li Jianbin2(1. School of Digital and Intelligent Industry, Inner Mongolia University of Science and Technology, Baotou 014010, Inner Mongolia, China;
    2. Institute of Animal Science and Veterinary Medicine, Shandong Academy of Agricultural Sciences, Jinan 250100, Shandong, China
  • Published:2026-08-12

摘要: 为应对动物姿态估计中大规模标注数据获取困难、成本高昂及数据稀缺等核心挑战,提出一种基于伪标签和一致性训练的半监督动物姿态估计方法(semi-supervised animal pose estimation method based on pseudo-labels and consistency training, Pcapose),提升模型在有限标注数据下的性能与泛化能力。基于初始标注数据训练初始模型,通过协同利用大量未标注数据构建三重机制,提升伪标签质量与模型鲁棒性。采用置信度驱动的伪标签筛选策略,筛选低损失、高置信度的样本扩充训练集;引入多视图一致性检测机制,融合几何、光照和特征空间的扰动信息,对高损失、低置信度的样本进行再评估与筛选;构建师生一致性框架,确保教师模型提供稳定、准确的伪标签。在AP-10K和Grévy's Zebra数据集上的试验结果表明,Pcapose在关键点检测精度和模型鲁棒性方面显著优于现有半监督方法,在数据稀缺场景下展现出优越性能。

关键词: 动物姿态估计, 半监督学习, 伪标签筛选, 多视图一致性, 师生一致性, 坐标分类, 数据稀缺

Abstract: To address the core challenges such as the difficulty in obtaining large scale labeled data, high costs, and data scarcity in animal pose estimation, a semi-supervised animal pose estimation method based on pseudo-labels and consistency training(Pcapose)was proposed to enhance the model's performance and generalization ability under limited labeled data conditions. An initial model was trained based on the initial labeled data. By jointly using a large amount of unlabeled data, a triple mechanism was constructed to improve the quality of pseudo-labels and the robustness of the model. A confidence driven pseudo-label screening strategy was adopted to select samples with low loss and high confidence to expand the training set. A multi-view consistency detection mechanism was introduced to integrate perturbation information from geometry, illumination, and feature space for re-evaluating and screening samples with high loss and low confidence. A teacher-student consistency framework was built to ensure the teacher model provided stable and accurate pseudo-labels. The experimental results on the AP-10K and Grévy's Zebra datasets showed that Pcapose significantly outperformed existing semi-supervised methods in terms of key point detection accuracy and model robustness, demonstrating superior performance in data scarce scenarios.

Key words: animal pose estimation, semi-supervised learning, pseudo-label screening, multi-view consistency, teacher-student consistency, coordinate classification, data scarcity

中图分类号: 

  • TP391.41
[1] Cao J K, Tang H Y, Fang H S, et al. Cross-domain adaptation for animal pose estimation[C] //2019 IEEE/CVF International Conference on Computer Vision(ICCV). Seoul: IEEE, 2019: 9497-9506.
[2] Li C, Lee G H. From synthetic to real: unsupervised domain adaptation for animal pose estimation[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2021: 1482-1491.
[3] Yu H, Xu Y F, Zhang J, et al. AP-10K: a benchmark for animal pose estimation in the wild[EB/OL].(2021-11-01)[2025-03-22]. https://arxiv.org/abs/2108.12617
[4] Shooter M, Malleson C, Hilton A. SyDog: a synthetic dog dataset for improved 2D pose estimation[EB/OL].(2021-07-31)[2025-03-22]. https://arxiv.org/abs/2108.00249
[5] Lee D H. Pseudo-label: the simple and efficient semi-supervised learning method for deep neural networks[C] //Proceedings of the 2013 ICML Workshop: Challenges in Representation Learning. Atlanta, USA: [s.n.] , 2013: 896.
[6] Scudder H. Probability of error of some adaptive pattern-recognition machines[J]. IEEE Transactions on Information Theory, 1965, 11(3): 363-371.
[7] Tarvainen A, Valpola H. Mean teachers are better role models: weight-averaged consistency targets improve semi-supervised deep learning results[EB/OL].(2015-04-16)[2025-03-22]. https://arxiv.org/abs/1703.01780
[8] Xie Q Z, Dai Z H, Hovy E, et al. Unsupervised data augmentation for consistency training[EB/OL].(2020-11-05)[2025-03-22]. https://arxiv.org/abs/1904.12848
[9] Nassar I, Hayat M, Abbasnejad E, et al. ProtoCon: pseudo-label refinement via online clustering and prototypical consistency for efficient semi-supervised learning[EB/OL].(2023-03-22)[2025-03-22]. https://arxiv.org/abs/2303.13556
[10] Yu Z R, Wang M C, Chen Y B, et al. Denoising and selecting pseudo-heatmaps for semi-supervised human pose estimation[C] //2024 IEEE/CVF Winter Confe-rence on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2024: 6268-6277.
[11] Li C, Lee G H. ScarceNet: animal pose estimation with scarce annotations[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Vancouver, Canada: IEEE, 2023: 17174-17183.
[12] Deng J H, Li W, Chen Y H, et al. Unbiased mean teacher for cross-domain object detection[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2021: 4089-4099.
[13] Cao S C, Joshi D, Gui L Y, et al. Contrastive mean teacher for domain adaptive object detectors[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Vancouver, Canada: IEEE, 2023: 23839-23848.
[14] 郭敏, 张熙涵, 李阳. 融合注意力的教师互一致性半监督医学图像分割[J]. 计算机工程, 2024, 50(9): 313-323. Guo Min, Zhang Xihan, Li Yang. Integrated attentional teacher mutual consistency semi-supervised medical image segmentation[J]. Computer Engineering, 2024, 50(9): 313-323.
[15] Jiang T, Lu P, Zhang L, et al. RTMPose: real-time multi-person pose estimation based on MMPose[EB/OL].(2023-07-03)[2025-03-22]. https://arxiv.org/abs/2303.07399
[16] Wang J D, Sun K, Cheng T H, et al. Deep high-resolution representation learning for visual recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(10): 3349-3364.
[17] Biggs B, Boyne O, Charles J, et al. Who left the dogs out? 3D animal reconstruction with expectation maximization in the loop[C] //Computer Vision-ECCV 2020: 16th European Conference. Glasgow, UK: Springer, 2020: 195-211.
[18] Li C, Lee G H. Coarse-to-fine animal pose and shape estimation[C] //Proceedings of the 35th International Conference on Neural Information Processing Systems. Red Hook, USA: ACM, 2021: 11757-11768.
[19] Lin T Y, Maire M, Belongie S, et al. Microsoft COCO: common objects in context[C] //Computer Vision-ECCV 2014: 13th European Conference. Zurich, Switzerland: Springer, 2014: 740-755.
[20] Rao J Y, Zhao B N, Wang Y. Probabilistic prompt distribution learning for animal pose estimation[C] //2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2025: 29438-29447.
[21] Xu T Y, Rao J Y, Song X N, et al. Learning structure-supporting dependencies via keypoint interactive Transformer for general mammal pose estimation[J]. International Journal of Computer Vision, 2025, 133(7): 3858-3876.
[22] Zhang W J, Liu D N, Cai W D, et al. Cross-view consistency regularisation for knowledge distillation[C] //Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024: 2011-2020.
[23] Sohn K, Berthelot D, Carlini N, et al. FixMatch: simplifying semi-supervised learning with consistency and confidence[EB/OL].(2020-11-25)[2025-03-22]. https://arxiv.org/abs/2001.07685
[24] Arazo E, Ortego D, Albert P, et al. Pseudo-labeling and confirmation bias in deep semi-supervised learning[C] //2020 International joint conference on neural networks(IJCNN). Glasgow, UK: IEEE, 2020: 9207304.
[25] 李萍, 张雪英, 王夙喆, 等. 基于半监督多尺度一致性学习的医学影像分割[J]. 计算机工程, 2025, 51(10): 295-307. Li Ping, Zhang Xueying, Wang Suzhe, et al. Medical image segmentation based on semi-supervised multi-scale consistency learning[J]. Computer Engineering, 2025, 51(10): 295-307.
[26] Sindhwani V, Niyogi P, Belkin M. A co-regularization approach to semi-supervised learning with multiple views[C] //Proceedings of ICML Workshop on Learning with Multiple Views. Bonn, Germany: [s.n.] , 2005: 74-79.
[27] 景攀峰, 梁宇栋, 李超伟, 等. 基于师生学习的半监督图像去雾算法[J]. 计算机应用, 2025, 45(9): 2975-2983. Jing Panfeng, Liang Yudong, Li Chaowei, et al. Semi-supervised image dehazing algorithm based on teacher-student learning [J]. Journal of Computer Applications, 2025, 45(9): 2975-2983.
[28] Li Y J, Yang S, Liu P D, et al. SimCC: a simple coordinate classification perspective for human pose estimation[EB/OL].(2022-07-05)[2025-03-22]. https://arxiv.org/abs/2107.03332
[29] Liu Y T, Wen Q, Chen H X, et al. Crowd counting via cross-stage refinement networks[J]. IEEE Transactions on Image Processing, 2020, 29: 6800-6812.
[30] Rodríguez P, Gonfaus J M, Cucurull G, et al. Attend and rectify: a gated attention mechanism for fine-grained recovery[C] //Computer Vision-ECCV 2018. Munich, Germany: Springer, 2018: 357-372.
[31] Graving J M, Chae D, Naik H, et al. Fast and robust animal pose estimation[EB/OL].(2019-04-26)[2025-03-22]. https://www.biorxiv.org/content/10.1101/620245v1
[32] Xie R C, Wang C Y, Zeng W J, et al. An empirical study of the collapsing problem in semi-supervised 2D human pose estimation[C] //2021 IEEE/CVF International Conference on Computer Vision(ICCV). Montreal, Canada: IEEE, 2022: 11220-11229.
[33] Schmutz H, Humbert O, Mattei P A. Don't fear the unlabelled: safe semi-supervised learning via simple debiasing[EB/OL].(2023-03-03)[2025-03-22]. https://arxiv.org/abs/2203.07512
[34] Scherer S, Schön R, Lienhart R. Pseudo-label noise suppression techniques for semi-supervised semantic segmentation[EB/OL].(2022-10-19)[2025-03-22]. https://arxiv.org/abs/2210.10426
[35] Nguyen K B. Debiasing, calibrating, and improving semi-supervised learning performance via simple ensemble projector[C] //2024 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2024: 2430-2439.
[36] Shen J Q, Jiang Y N, Luo J W, et al. MPE-HRNetL: a lightweight high-resolution network for multispecies animal pose estimation[J]. Sensors, 2024, 24(21): 6882.
[37] He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition[C] //2016 IEEE Confe-rence on Computer Vision and Pattern Recognition(CVPR). Las Vegas, USA: IEEE, 2016: 770-778.
[1] 尹旭,刘兆英,张婷,李玉鑑. 基于弱监督和半监督学习的红外舰船分割方法[J]. 山东大学学报 (工学版), 2022, 52(2): 99-106.
[2] 朱恒东, 马盈仓, 代雪珍. 自适应半监督邻域聚类算法[J]. 山东大学学报 (工学版), 2021, 51(4): 24-34.
[3] 孔超1,2,张化祥1,2*,刘丽1,2. 基于兴趣区域特征融合的半监督图像检索算法[J]. 山东大学学报(工学版), 2014, 44(3): 22-28.
[4] 李雅林1,2,张化祥1,2*,冯新营1,2. 一种新的基于半监督的多标记学习算法[J]. 山东大学学报(工学版), 2013, 43(2): 18-22.
[5] 夏战国,万玲,蔡世玉,孙鹏辉. 一种面向入侵检测的半监督聚类算法[J]. 山东大学学报(工学版), 2012, 42(6): 1-7.
[6] 谢伙生,刘敏. 一种基于主动学习的集成协同训练算法[J]. 山东大学学报(工学版), 2012, 42(3): 1-5.
[7] 魏巍,张艳宁. 基于半监督隐含狄利克雷分配的人脸姿态判别方法[J]. 山东大学学报(工学版), 2011, 41(3): 17-22.
[8] 宿洪禄,李凡长*. 基于相异性和不变特征的半监督图像检索[J]. 山东大学学报(工学版), 2010, 40(5): 150-153.
[9] 崔宝今 林鸿飞 张霄. 基于半监督学习的蛋白质关系抽取研究[J]. 山东大学学报(工学版), 2009, 39(3): 16-21.
[10] 周广通,尹义龙,郭文鹃,任春晓. 基于协同训练的指纹图像分割算法[J]. 山东大学学报(工学版), 2009, 39(1): 22-26.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!