您的位置:山东大学 -> 科技期刊社 -> 《山东大学学报(工学版)》

山东大学学报 (工学版) ›› 2026, Vol. 56 ›› Issue (4): 52-64.doi: 10.6040/j.issn.1672-3961.0.2025.127

• 机器学习与数据挖掘 • 上一篇    

基于空间-频域信息引导的图像修复算法

侍淑娟1,叶海良1*,曹飞龙2   

  1. 1.中国计量大学理学院, 浙江 杭州 310018;2.浙江师范大学数学与交叉科学研究院, 浙江 杭州 310012
  • 发布日期:2026-08-12
  • 作者简介:侍淑娟(1999— ),女,云南曲靖人,硕士研究生,主要研究方向为深度学习与图像处理. E-mail:atopossj@163.com. *通信作者简介:叶海良(1990— ),男,浙江绍兴人,副教授,硕士生导师,博士,主要研究方向为深度学习与图像处理. E-mail:yehl@cjlu.edu.cn
  • 基金资助:
    国家自然科学基金面上资助项目(62176244,62536006)

A spatial-frequency domain information guided algorithm for image inpainting

Shi Shujuan1, Ye Hailiang1*, Cao Feilong2   

  1. Shi Shujuan1, Ye Hailiang1*, Cao Feilong2(1. College of Sciences, China Jiliang University, Hangzhou 310018, Zhejiang, China;
    2. Institute of Mathematics and Cross-disciplinary Science, Zhejiang Normal University, Hangzhou 310012, Zhejiang, China
  • Published:2026-08-12

摘要: 针对现有图像修复方法难以协同优化空间与频域特征,在上采样过程中难以保持特征全局一致性的问题,提出一种基于空间-频域信息引导的图像修复算法。通过下采样、联合特征提取及上采样3个核心阶段,实现空间与频域信息深度融合。在下采样阶段,设计频域信息引导的下采样模块,结合频域特征有效保留图像的重要结构信息;在联合特征提取阶段,设计空间-频域联合特征提取模块,采用双分支并行架构,分别提取多尺度局部空间特征和基于快速傅里叶变换的全局频域特征,通过特征融合实现局部细节与全局结构的协同表征;在上采样阶段,提出频域信息引导的上采样模块,结合亚像素卷积与双线性插值的优势,引入频域特征增强全局结构一致性,有效平衡修复结果的精细度与自然度。在CelebA-HQ、Places2和Paris StreetView数据集上的试验结果表明,所提算法在峰值信噪比、结构相似性和学习感知图像块相似性指标上优于许多现有方法,有效提升图像修复的纹理连贯性与视觉真实性。

关键词: 深度学习, 图像修复, 特征提取, 快速傅里叶变换, 频域信息

Abstract: Aiming at the issue of jointly optimizing spatial and frequency domain features and preserving global structural consistency during upsampling in image inpainting, a spatial-frequency domain information guided algorithm for image inpainting was proposed. Through three core stages of downsampling, joint feature extraction, and upsampling, the deep integration of spatial and frequency domain information was realized. In the downsampling stage, a frequency domain information-guided downsampling module was designed to preserve key structural information using frequency features. In the joint feature extraction stage, a spatial-frequency domain joint feature extraction module was introduced, employing a dual-branch parallel architecture to separately extract multi-scale local spatial features and global frequency domain features based on fast Fourier transform, realizing a synergistic representation of local details and global structures through feature fusion. In the upsampling stage, a frequency domain information-guided upsampling module was introduced, combining the advantages of sub-pixel convolution and bilinear interpolation while introducing frequency domain features to enhance the consistency of global structures, effectively balancing the delicacy and naturalness of the inpainting results. Experimental results on the CelebA-HQ, Places2, and Paris StreetView datasets demonstrated that the proposed method outperformed existing approaches in terms of peak signal-to-noise ratio, structural similarity index measure, and learned perceptual image patch similarity, effectively enhancing the texture coherence and visual authenticity of image inpainting.

Key words: deep learning, image inpainting, feature extraction, fast Fourier transform, frequency domain information

中图分类号: 

  • TP183
[1] Quan W Z, Chen J X, Liu Y L, et al. Deep learning-based image and video inpainting: a survey[J]. International Journal of Computer Vision, 2024, 132(7): 2367-2400.
[2] 李月龙, 高云, 闫家良, 等. 基于深度神经网络的图像缺损修复方法综述[J]. 计算机学报, 2021, 44(11):2295-2316. Li Yuelong, Gao Yun, Yan Jialiang, et al. Image inpainting methods based on deep neural networks: a review[J]. Chinese Journal of Computers, 2021, 44(11): 2295-2316.
[3] 王真言, 蒋胜丞, 宋奇鸿, 等. 基于Transformer的文物图像修复方法[J]. 计算机研究与发展, 2024, 61(3): 748-761. Wang Zhenyan, Jiang Shengcheng, Song Qihong, et al. Transformer-based image restoration method for cultural relics[J]. Journal of Computer Research and Develop-ment, 2024, 61(3): 748-761.
[4] Oh S W, Lee S, Lee J Y, et al. Onion-peel networks for deep video completion[C] //2019 IEEE/CVF Inter-national Conference on Computer Vision(ICCV). Seoul: IEEE, 2019: 4402-4411.
[5] Wan Z Y, Zhang B, Chen D, et al. Old photo restoration via deep latent space translation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(2): 2071-2087.
[6] Liu Y, Sun P, Wergeles N, et al. A survey and performance evaluation of deep learning methods for small object detection[J]. Expert Systems with Applications, 2021, 172: 114602.
[7] 王相海, 孙丽, 万宇, 等. 非局域样本填充和自适应曲率驱动模型的遥感图像修复算法[J]. 模式识别与人工智能, 2016, 29(8): 735-743. Wang Xianghai, Sun Li, Wan Yu, et al. Remote sensing image inpainting based on non-local sample filling and adaptive curvature driven diffusions model[J]. Pattern Recognition and Artificial Intelligence, 2016, 29(8): 735-743.
[8] Ghorai M, Samanta S, Mandal S, et al. Multiple pyramids based image inpainting using local patch statistics and steering kernel feature[J]. IEEE Transactions on Image Processing, 2019, 28(11): 5495-5509.
[9] He L T, Wang Y L. Iterative support detection-based split Bregman method for wavelet frame-based image inpainting[J]. IEEE Transactions on Image Processing, 2014, 23(12): 5470-5485.
[10] Liang X, Ren X, Zhang Z D, et al. Texture repairing by unified low rank optimization[J]. Journal of Computer Science and Technology, 2016, 31(3): 525-546.
[11] Wang Y F, Guo D S, Zhao H R, et al. Image inpainting via multi-scale adaptive priors[J]. Pattern Recognition, 2025, 162: 111410.
[12] Huang W L, Deng Y, Hui S Q, et al. Sparse self-attention Transformer for image inpainting[J]. Pattern Recognition, 2024, 145: 109897.
[13] Yu J H, Lin Z, Yang J M, et al. Free-form image inpainting with gated convolution[C] //2019 IEEE/CVF International Conference on Computer Vision(ICCV). Seoul: IEEE, 2019: 4471-4480.
[14] Deng Y, Hui S Q, Zhou S P, et al. Learning contextual Transformer network for image inpainting[C] //Pro-ceedings of the 29th ACM International Conference on Multimedia. [S.l.] : ACM, 2021: 2529-2538.
[15] Wang J, Wang C, Huang Q M, et al. Image inpainting based on multi-frequency probabilistic inference model[C] //Proceedings of the 28th ACM International Conference on Multimedia. Seattle, USA: ACM, 2020: 1-9.
[16] Suvorov R, Logacheva E, Mashikhin A, et al. Resolution-robust large mask inpainting with Fourier convolutions[C] //2022 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2022: 3172-3182.
[17] Lu Z Y, Jiang J J, Huang J Q, et al. GLaMa: joint spatial and frequency loss for general image inpainting[C] //2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops(CVPRW). New Orleans, USA: IEEE, 2022: 1300-1309.
[18] Chu T Y, Chen J F, Sun J K, et al. Rethinking fast Fourier convolution in image inpainting[C] //2023 IEEE/CVF International Conference on Computer Vision(ICCV). Paris, France: IEEE, 2024: 23138-23148.
[19] Pathak D, Krähenbühl P, Donahue J, et al. Context encoders: feature learning by inpainting[C] //2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Las Vegas, USA: IEEE, 2016: 2536-2544.
[20] Liu G L, Reda F A, Shih K J, et al. Image inpainting for irregular holes using partial convolutions[C] //Computer Vision-ECCV 2018. Munich, Germany: Springer, 2018: 89-105.
[21] Yu J H, Lin Z, Yang J M, et al. Generative image inpainting with contextual attention[C] //2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018: 5505-5514.
[22] Liu H Y, Jiang B, Song Y B, et al. Rethinking image inpainting via a mutual encoder-decoder with feature equalizations[C] //Computer Vision-ECCV 2020. Glasgow, UK: Springer, 2020: 725-741.
[23] Guo X F, Yang H Y, Huang D. Image inpainting via conditional texture and structure dual generation[C] //2021 IEEE/CVF International Conference on Computer Vision(ICCV). Montreal, Canada: IEEE, 2021: 14114-14123.
[24] 邵新茹, 叶海良, 杨冰, 等. 基于三阶段生成网络的图像修复[J]. 模式识别与人工智能, 2022, 35(12): 1047-1063. Shao Xinru, Ye Hailiang, Yang Bing, et al. Image inpainting with a three-stage generative network[J]. Pattern Recognition and Artificial Intelligence, 2022, 35(12): 1047-1063.
[25] Liu W H, Cun X D, Pun C M, et al. CoordFill: efficient high-resolution image inpainting via parameterized coordinate querying[C] //Proceedings of the AAAI Conference on Artificial Intelligence. Washington, DC, USA: AAAI, 2023: 1746-1754.
[26] Ko K, Kim C S. Continuously masked Transformer for image inpainting[C] //2023 IEEE/CVF International Conference on Computer Vision(ICCV). Paris, France: IEEE, 2024: 13123-13132.
[27] Yu Y C, Zhan F N, Lu S J, et al. WaveFill: a wavelet-based generation network for image inpainting[C] //2021 IEEE/CVF International Conference on Computer Vision(ICCV). Montreal, Canada: IEEE, 2022: 14094-14103.
[28] Li B, Zheng B W, Li H D, et al. Detail-enhanced image inpainting based on discrete wavelet transforms[J]. Signal Processing, 2021, 189:108278.
[29] Jain J, Zhou Y Q, Yu N, et al. Keys to better image inpainting: structure and texture go hand in hand[C] //2023 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2023: 208-217.
[30] Cai X H, Lai Q X, Wang Y W, et al. Poly kernel inception network for remote sensing detection[C] //2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Seattle, USA: IEEE, 2024: 27706-27716.
[31] Woo S, Park J, Lee J Y, et al. CBAM: convolutional block attention module[C] //Computer Vision-ECCV 2018. Munich, Germany: Springer, 2018: 3-19.
[32] Szegedy C, Vanhoucke V, Ioffe S, et al. Rethinking the inception architecture for computer vision[C] //2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Las Vegas, USA: IEEE, 2016: 2818-2826.
[33] Yu W H, Zhou P, Yan S C, et al. InceptionNeXt: when inception meets ConvNeXt[C] //2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Seattle, USA: IEEE, 2024: 5672-5683.
[34] Karras T, Aila T, Laine S, et al. Progressive growing of GANs for improved quality, stability, and variation[PP/OL]. V3.(2018-02-26)[2025-07-10]. https://arxiv.org/abs/1710.10196
[35] Zhou B L, Lapedriza A, Khosla A, et al. Places: a 10 million image database for scene recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(6): 1452-1464.
[36] Doersch C, Singh S, Gupta A, et al. What makes Paris look like Paris?[J]. ACM Transactions on Graphics, 2012, 31(4):101.
[37] Kingma D P, Ba J. Adam: a method for stochastic optimization[PP/OL]. V9.(2017-01-30)[2025-07-10]. https://arxiv.org/abs/1412.6980
[38] Li J Y, Wang N, Zhang L F, et al. Recurrent feature reasoning for image inpainting[C] //2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Seattle, USA: IEEE, 2020: 7757-7765.
[39] Zuo Z W, Zhao L, Li A L, et al. Generative image inpainting with segmentation confusion adversarial training and contrastive learning[C] //Proceedings of the AAAI Conference on Artificial Intelligence. Washington, DC, USA: AAAI, 2023: 3888-3896.
[40] Verma S, Sharma A, Sheshadri R, et al. GraphFill: deep image inpainting using graphs[C] //2024 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV). Waikoloa, USA: IEEE, 2024: 4984-4994.
[41] Li Z, Zhang Y N, Du Y F, et al. STNet: structure and texture-guided network for image inpainting[J]. Pattern Recognition, 2024, 156: 110786.
[42] Cao F L, Xu Q J, Ye H L. Adaptive prior and long-range dependency-based learners for image inpainting[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(11): 10742-10755.
[1] 王倩,张瑞敏,李明津,孟宪静,耿蕾蕾. 基于频域图卷积网络的时空序列预测[J]. 山东大学学报 (工学版), 2026, 56(3): 84-92.
[2] 王新建,景志滨,孟凡成,石建国,张敏昊,张一帆,王庆华,朱彦恺. 基于D-Mamba模型的超短期火电机组发电负荷预测[J]. 山东大学学报 (工学版), 2026, 56(1): 169-178.
[3] 李常刚,李宝亮,曹永吉,王佳颖. 人工智能在电力系统潮流计算中的应用综述及展望[J]. 山东大学学报 (工学版), 2025, 55(5): 1-17.
[4] 周群颖,隋家成,张继,王洪元. 基于自监督卷积和无参数注意力机制的工业品表面缺陷检测[J]. 山东大学学报 (工学版), 2025, 55(4): 40-47.
[5] 周遵富,张乾,石计亮,岳诗琴. 基于纹理和结构交互的人脸图像修复[J]. 山东大学学报 (工学版), 2025, 55(4): 18-28.
[6] 董明书,陈俐企,马川义,张珠皓,孙仁娟,管延华,庄培芝. 沥青路面内部裂缝雷达图像智能判识算法研究[J]. 山东大学学报 (工学版), 2025, 55(3): 72-79.
[7] 薛冰冰,王勇,杨维浩,王川,于迪,王旭. 基于ETC收费数据的高速公路交通流数据修复及实时预测[J]. 山东大学学报 (工学版), 2025, 55(3): 58-71.
[8] 常新功,苏敏惠,周志刚. 基于进化集成的图神经网络解释方法[J]. 山东大学学报 (工学版), 2024, 54(4): 1-12.
[9] 索大翔,李波. 基于Gromov-Wasserstein最优传输的输电线路小目标检测方法[J]. 山东大学学报 (工学版), 2024, 54(3): 22-29.
[10] 聂秀山,巩蕊,董飞,郭杰,马玉玲. 短视频场景分类方法综述[J]. 山东大学学报 (工学版), 2024, 54(3): 1-11.
[11] 宋辉,张轶哲,张功萱,孟元. 基于类权重和最小化预测熵的测试时集成方法[J]. 山东大学学报 (工学版), 2024, 54(3): 36-43.
[12] 刘新,刘冬兰,付婷,王勇,常英贤,姚洪磊,罗昕,王睿,张昊. 基于联邦学习的时间序列预测算法[J]. 山东大学学报 (工学版), 2024, 54(3): 55-63.
[13] 高泽文,王建,魏本征. 基于混合偏移轴向自注意力机制的脑胶质瘤分割算法[J]. 山东大学学报 (工学版), 2024, 54(2): 80-89.
[14] 李璐,张志军,范钰敏,王星,袁卫华. 面向冷启动用户的元学习与图转移学习序列推荐[J]. 山东大学学报 (工学版), 2024, 54(2): 69-79.
[15] 陈成,董永权,贾瑞,刘源. 基于交互序列特征相关性的可解释知识追踪[J]. 山东大学学报 (工学版), 2024, 54(1): 100-108.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!