山东大学学报 (工学版) ›› 2026, Vol. 56 ›› Issue (4): 27-37.doi: 10.6040/j.issn.1672-3961.0.2025.109
• 机器学习与数据挖掘 • 上一篇
朱钊利1,张继凯1*,曾翔皓1,解辰杰1,李建斌2
Zhu Zhaoli1, Zhang Jikai1*, Zeng Xianghao1, Xie Chenjie1, Li Jianbin2
摘要: 为应对动物姿态估计中大规模标注数据获取困难、成本高昂及数据稀缺等核心挑战,提出一种基于伪标签和一致性训练的半监督动物姿态估计方法(semi-supervised animal pose estimation method based on pseudo-labels and consistency training, Pcapose),提升模型在有限标注数据下的性能与泛化能力。基于初始标注数据训练初始模型,通过协同利用大量未标注数据构建三重机制,提升伪标签质量与模型鲁棒性。采用置信度驱动的伪标签筛选策略,筛选低损失、高置信度的样本扩充训练集;引入多视图一致性检测机制,融合几何、光照和特征空间的扰动信息,对高损失、低置信度的样本进行再评估与筛选;构建师生一致性框架,确保教师模型提供稳定、准确的伪标签。在AP-10K和Grévy's Zebra数据集上的试验结果表明,Pcapose在关键点检测精度和模型鲁棒性方面显著优于现有半监督方法,在数据稀缺场景下展现出优越性能。
中图分类号:
| [1] Cao J K, Tang H Y, Fang H S, et al. Cross-domain adaptation for animal pose estimation[C] //2019 IEEE/CVF International Conference on Computer Vision(ICCV). Seoul: IEEE, 2019: 9497-9506. [2] Li C, Lee G H. From synthetic to real: unsupervised domain adaptation for animal pose estimation[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2021: 1482-1491. [3] Yu H, Xu Y F, Zhang J, et al. AP-10K: a benchmark for animal pose estimation in the wild[EB/OL].(2021-11-01)[2025-03-22]. https://arxiv.org/abs/2108.12617 [4] Shooter M, Malleson C, Hilton A. SyDog: a synthetic dog dataset for improved 2D pose estimation[EB/OL].(2021-07-31)[2025-03-22]. https://arxiv.org/abs/2108.00249 [5] Lee D H. Pseudo-label: the simple and efficient semi-supervised learning method for deep neural networks[C] //Proceedings of the 2013 ICML Workshop: Challenges in Representation Learning. Atlanta, USA: [s.n.] , 2013: 896. [6] Scudder H. Probability of error of some adaptive pattern-recognition machines[J]. IEEE Transactions on Information Theory, 1965, 11(3): 363-371. [7] Tarvainen A, Valpola H. Mean teachers are better role models: weight-averaged consistency targets improve semi-supervised deep learning results[EB/OL].(2015-04-16)[2025-03-22]. https://arxiv.org/abs/1703.01780 [8] Xie Q Z, Dai Z H, Hovy E, et al. Unsupervised data augmentation for consistency training[EB/OL].(2020-11-05)[2025-03-22]. https://arxiv.org/abs/1904.12848 [9] Nassar I, Hayat M, Abbasnejad E, et al. ProtoCon: pseudo-label refinement via online clustering and prototypical consistency for efficient semi-supervised learning[EB/OL].(2023-03-22)[2025-03-22]. https://arxiv.org/abs/2303.13556 [10] Yu Z R, Wang M C, Chen Y B, et al. Denoising and selecting pseudo-heatmaps for semi-supervised human pose estimation[C] //2024 IEEE/CVF Winter Confe-rence on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2024: 6268-6277. [11] Li C, Lee G H. ScarceNet: animal pose estimation with scarce annotations[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Vancouver, Canada: IEEE, 2023: 17174-17183. [12] Deng J H, Li W, Chen Y H, et al. Unbiased mean teacher for cross-domain object detection[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2021: 4089-4099. [13] Cao S C, Joshi D, Gui L Y, et al. Contrastive mean teacher for domain adaptive object detectors[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Vancouver, Canada: IEEE, 2023: 23839-23848. [14] 郭敏, 张熙涵, 李阳. 融合注意力的教师互一致性半监督医学图像分割[J]. 计算机工程, 2024, 50(9): 313-323. Guo Min, Zhang Xihan, Li Yang. Integrated attentional teacher mutual consistency semi-supervised medical image segmentation[J]. Computer Engineering, 2024, 50(9): 313-323. [15] Jiang T, Lu P, Zhang L, et al. RTMPose: real-time multi-person pose estimation based on MMPose[EB/OL].(2023-07-03)[2025-03-22]. https://arxiv.org/abs/2303.07399 [16] Wang J D, Sun K, Cheng T H, et al. Deep high-resolution representation learning for visual recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021, 43(10): 3349-3364. [17] Biggs B, Boyne O, Charles J, et al. Who left the dogs out? 3D animal reconstruction with expectation maximization in the loop[C] //Computer Vision-ECCV 2020: 16th European Conference. Glasgow, UK: Springer, 2020: 195-211. [18] Li C, Lee G H. Coarse-to-fine animal pose and shape estimation[C] //Proceedings of the 35th International Conference on Neural Information Processing Systems. Red Hook, USA: ACM, 2021: 11757-11768. [19] Lin T Y, Maire M, Belongie S, et al. Microsoft COCO: common objects in context[C] //Computer Vision-ECCV 2014: 13th European Conference. Zurich, Switzerland: Springer, 2014: 740-755. [20] Rao J Y, Zhao B N, Wang Y. Probabilistic prompt distribution learning for animal pose estimation[C] //2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Nashville, USA: IEEE, 2025: 29438-29447. [21] Xu T Y, Rao J Y, Song X N, et al. Learning structure-supporting dependencies via keypoint interactive Transformer for general mammal pose estimation[J]. International Journal of Computer Vision, 2025, 133(7): 3858-3876. [22] Zhang W J, Liu D N, Cai W D, et al. Cross-view consistency regularisation for knowledge distillation[C] //Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024: 2011-2020. [23] Sohn K, Berthelot D, Carlini N, et al. FixMatch: simplifying semi-supervised learning with consistency and confidence[EB/OL].(2020-11-25)[2025-03-22]. https://arxiv.org/abs/2001.07685 [24] Arazo E, Ortego D, Albert P, et al. Pseudo-labeling and confirmation bias in deep semi-supervised learning[C] //2020 International joint conference on neural networks(IJCNN). Glasgow, UK: IEEE, 2020: 9207304. [25] 李萍, 张雪英, 王夙喆, 等. 基于半监督多尺度一致性学习的医学影像分割[J]. 计算机工程, 2025, 51(10): 295-307. Li Ping, Zhang Xueying, Wang Suzhe, et al. Medical image segmentation based on semi-supervised multi-scale consistency learning[J]. Computer Engineering, 2025, 51(10): 295-307. [26] Sindhwani V, Niyogi P, Belkin M. A co-regularization approach to semi-supervised learning with multiple views[C] //Proceedings of ICML Workshop on Learning with Multiple Views. Bonn, Germany: [s.n.] , 2005: 74-79. [27] 景攀峰, 梁宇栋, 李超伟, 等. 基于师生学习的半监督图像去雾算法[J]. 计算机应用, 2025, 45(9): 2975-2983. Jing Panfeng, Liang Yudong, Li Chaowei, et al. Semi-supervised image dehazing algorithm based on teacher-student learning [J]. Journal of Computer Applications, 2025, 45(9): 2975-2983. [28] Li Y J, Yang S, Liu P D, et al. SimCC: a simple coordinate classification perspective for human pose estimation[EB/OL].(2022-07-05)[2025-03-22]. https://arxiv.org/abs/2107.03332 [29] Liu Y T, Wen Q, Chen H X, et al. Crowd counting via cross-stage refinement networks[J]. IEEE Transactions on Image Processing, 2020, 29: 6800-6812. [30] Rodríguez P, Gonfaus J M, Cucurull G, et al. Attend and rectify: a gated attention mechanism for fine-grained recovery[C] //Computer Vision-ECCV 2018. Munich, Germany: Springer, 2018: 357-372. [31] Graving J M, Chae D, Naik H, et al. Fast and robust animal pose estimation[EB/OL].(2019-04-26)[2025-03-22]. https://www.biorxiv.org/content/10.1101/620245v1 [32] Xie R C, Wang C Y, Zeng W J, et al. An empirical study of the collapsing problem in semi-supervised 2D human pose estimation[C] //2021 IEEE/CVF International Conference on Computer Vision(ICCV). Montreal, Canada: IEEE, 2022: 11220-11229. [33] Schmutz H, Humbert O, Mattei P A. Don't fear the unlabelled: safe semi-supervised learning via simple debiasing[EB/OL].(2023-03-03)[2025-03-22]. https://arxiv.org/abs/2203.07512 [34] Scherer S, Schön R, Lienhart R. Pseudo-label noise suppression techniques for semi-supervised semantic segmentation[EB/OL].(2022-10-19)[2025-03-22]. https://arxiv.org/abs/2210.10426 [35] Nguyen K B. Debiasing, calibrating, and improving semi-supervised learning performance via simple ensemble projector[C] //2024 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2024: 2430-2439. [36] Shen J Q, Jiang Y N, Luo J W, et al. MPE-HRNetL: a lightweight high-resolution network for multispecies animal pose estimation[J]. Sensors, 2024, 24(21): 6882. [37] He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition[C] //2016 IEEE Confe-rence on Computer Vision and Pattern Recognition(CVPR). Las Vegas, USA: IEEE, 2016: 770-778. |
| [1] | 尹旭,刘兆英,张婷,李玉鑑. 基于弱监督和半监督学习的红外舰船分割方法[J]. 山东大学学报 (工学版), 2022, 52(2): 99-106. |
| [2] | 朱恒东, 马盈仓, 代雪珍. 自适应半监督邻域聚类算法[J]. 山东大学学报 (工学版), 2021, 51(4): 24-34. |
| [3] | 孔超1,2,张化祥1,2*,刘丽1,2. 基于兴趣区域特征融合的半监督图像检索算法[J]. 山东大学学报(工学版), 2014, 44(3): 22-28. |
| [4] | 李雅林1,2,张化祥1,2*,冯新营1,2. 一种新的基于半监督的多标记学习算法[J]. 山东大学学报(工学版), 2013, 43(2): 18-22. |
| [5] | 夏战国,万玲,蔡世玉,孙鹏辉. 一种面向入侵检测的半监督聚类算法[J]. 山东大学学报(工学版), 2012, 42(6): 1-7. |
| [6] | 谢伙生,刘敏. 一种基于主动学习的集成协同训练算法[J]. 山东大学学报(工学版), 2012, 42(3): 1-5. |
| [7] | 魏巍,张艳宁. 基于半监督隐含狄利克雷分配的人脸姿态判别方法[J]. 山东大学学报(工学版), 2011, 41(3): 17-22. |
| [8] | 宿洪禄,李凡长*. 基于相异性和不变特征的半监督图像检索[J]. 山东大学学报(工学版), 2010, 40(5): 150-153. |
| [9] | 崔宝今 林鸿飞 张霄. 基于半监督学习的蛋白质关系抽取研究[J]. 山东大学学报(工学版), 2009, 39(3): 16-21. |
| [10] | 周广通,尹义龙,郭文鹃,任春晓. 基于协同训练的指纹图像分割算法[J]. 山东大学学报(工学版), 2009, 39(1): 22-26. |
|
||