您的位置:山东大学 -> 科技期刊社 -> 《山东大学学报(工学版)》

山东大学学报 (工学版) ›› 2026, Vol. 56 ›› Issue (4): 10-16.doi: 10.6040/j.issn.1672-3961.0.2025.137

• 机器学习与数据挖掘 • 上一篇    

基于多通道BERT预训练的合同文本检测方法

刘彦北1,2,赵敏全3*,谭松泰3,周李良3, 董新然2,4   

  1. 1.天津工业大学生命科学学院, 天津 300387;2.天津市光电检测技术与系统重点实验室(天津工业大学), 天津 300387;3.南方电网数字平台科技(广东)有限公司, 广东 深圳 518000;4.天津工业大学电子与信息工程学院, 天津 300387
  • 发布日期:2026-08-12
  • 作者简介:刘彦北(1985— ),男,河南安阳人,副教授,硕士生导师,博士,主要研究方向为深度学习、大模型等. E-mail:liuyanbei@tiangong.edu.cn. *通信作者简介:赵敏全(1979— ),男,湖南邵阳人,高级工程师,主要研究方向为数据挖掘、机器学习等. E-mail:1837404139@qq.com
  • 基金资助:
    京津冀基础研究合作资助项目(H2021202008)

Contract text detection based on multi-channel BERT pretraining

Liu Yanbei1,2, Zhao Minquan3*, Tan Songtai3, Zhou Liliang3, Dong Xinran2,4   

  1. Liu Yanbei1, 2, Zhao Minquan3*, Tan Songtai3, Zhou Liliang3, Dong Xinran2, 4(1. School of Life Sciences, Tiangong University, Tianjin 300387, China;
    2. Tianjin Key Laboratory of Optoelectronic Detection Technology and Systems, Tiangong University, Tianjin 300387, China;
    3. Southern Power Grid Digital Platform Technology(Guangdong)Co., Ltd., Shenzhen 518000, Guangdong, China;
    4. School of Electronic and Information Engineering, Tiangong University, Tianjin 300387, China
  • Published:2026-08-12

摘要: 合同风险检测有助于预防项目风险和保障权益,提高合同审查效率。然而,现有方法在合同文本的复杂语义理解和多尺度特征提取方面仍存在不足,难以兼顾全局语义关联与局部关键信息的捕捉。本研究提出一种基于多通道双向编码器表示模型(bidirectional encoder representations from transformers, BERT)的预训练方法MC-BERT(multi-channel BERT),用于合同文本检测。利用BERT预训练模型提取文本的全局特征,并设计多通道卷积模块提取文本的多尺度局部特征,联合注意力机制自适应学习全局与局部特征的权重分布,以精确提取文本的关键信息。在公开的合同文本数据集上,MC-BERT的合同风险判别准确率比现有的流行方法提升4.6%,验证所提模型的有效性。所提出模型融合BERT预训练、多通道卷积模块以及注意力机制的优点,为文本检测领域提供了新的研究思路和技术支持。

关键词: 合同风险检测, 文本检测, BERT预训练模型, 注意力机制, 多通道

Abstract: Contract risk detection helped prevent project risks and protect rights, improving the efficiency of contract review. However, existing methods still faced limitations in understanding the complex semantics of contract texts and extracting multi-scale features, which made it challenging to balance global semantic relations with the capture of local key information. This study proposed a novel multi-channel BERT-based pretraining method for contract text detection. The BERT pretraining model was utilized to extract the global features of the text, while a multi-channel convolution module was designed to capture multi-scale local features. An integrated attention mechanism adaptively learned the weight distribution between global and local features, enabling the precise extraction of key information from the text. On a publicly available contract text dataset, the proposed algorithm improved the contract risk classification accuracy by 4.6% compared to existing popular methods, validating the effectiveness of the proposed model. The model integrated the advantages of BERT pretraining, multi-channel convolution modules, and attention mechanisms, offering new research ideas and technical support for the field of text detection, with broad application prospects and significant practical value.

Key words: contract risk detection, text detection, BERT pre-training, attention mechanism, multi-channel

中图分类号: 

  • TP311.13
[1] 朱璋颖, 陆亦恬, 唐祝寿, 等. 基于隐私政策条款和机器学习的应用分类[J]. 通信技术, 2020, 53(11): 2749-2757. Zhu Zhangying, Lu Yitian, Tang Zhushou, et al. Application classification based on privacy policy terms and machine learning[J]. Communications Technology, 2020, 53(11): 2749-2757.
[2] 高恒源. 基于深度学习的合同风险检测方法及应用[D]. 保定: 河北大学, 2024: 1-3. Gao Hengyuan. Contract risk detection method and application based on deep learning[D]. Baoding: Hebei University, 2024: 1-3.
[3] 曾寒毓. 合同风险预判技术研究及应用[D]. 绵阳: 西南科技大学, 2021: 2-5. Zeng Hanyu. Research and application of contract risk prediction technology[D]. Mianyang: Southwest University of Science and Technology, 2021: 2-5.
[4] Mikolov T, Chen K, Corrado G, et al. Efficient estimation of word representations in vector space[PP/OL]. V2.(2013-09-07)[2025-03-05]. https://arxiv.org/abs/1301
[5] Guntamukkala N, Dara R, Grewal G. A machine-learning based approach for measuring the completeness of online privacy policies[C] //2015 IEEE 14th International Conference on Machine Learning and Applications. 2015, Miami, USA: IEEE, 2016: 289-294.
[6] Lippi M, Palka P, Contissa G, et al. CLAUDETTE: an automated detector of potentially unfair clauses in online terms of service[J]. Artificial Intelligence and Law, 2019, 27: 117-139.
[7] Hendrycks D, Burns C, Chen A, et al. CUAD: an expert-annotated NLP dataset for legal contract review[PP/OL] V2.(2021-11-08)[2025-03-05].https://doi.org/10.48550/arXiv.2103.06268
[8] Torre D, Abualhaija S, Sabetzadeh M, et al. An AI-assisted approach for checking the completeness of privacy policies against GDPR[C] //2020 IEEE 28th International Requirements Engineering Conference. 2020, Zurich, Switzerland: IEEE, 2020: 136-146.
[9] 周红, 王书钰, 黄文路. 基于NLP技术的建设工程合同风险智能检测框架研究[J]. 建筑经济, 2021, 42(6): 94-98. Zhou Hong, Wang Shuyu, Huang Wenlu. Research on intelligent detection framework of construction contract risks based on NLP[J]. Construction Economy, 2021, 42(6): 94-98.
[10] 王书钰. 基于NLP和深度学习的建设工程合同条款缺漏风险智能化检测研究[D]. 厦门: 厦门大学, 2021: 11-13. Wang Shuyu. Research on intelligent detection of missing risk of construction project contract clauses based on NLP and deep learning[D]. Xiamen:Xiamen University, 2021: 11-13.
[11] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C] //NeurIPS 2017. Long Beach, USA:Curran Associates, Inc, 2017: 5998-6008.
[12] 王金政, 杨颖, 余本功. 基于多头协同注意力机制的客户投诉文本分类模型[J]. 数据分析与知识发现, 2023, 7(1): 128-137. Wang Jinzheng, Yang Ying, Yu Bengong. Classifying customer complaints based on multi-head co-attention mechanism[J]. Data Analysis and Knowledge Discovery, 2023, 7(1): 128-137.
[13] 李浩君, 王耀东, 汪旭辉. 中文短文本情感分类: 融入位置感知强化的Transformer-TextCNN模型研究[J]. 计算机工程与应用, 2025, 61(11): 216-226. Li Haojun, Wang Yaodong, Wang Xuhui. Chinese short text sentiment classification: research on Transformer-TextCNN model with location-aware enhancement[J]. Computer Engineering and Applications, 2025, 61(11): 216-226.
[14] Kim Y. Convolutional neural networks for sentence classification[PP/OL]. V2.(2014-09-03)[2025-03-05].https://doi.org/10.48550/arXiv.1408.5882
[15] 国家市场监督管理总局. 全国合同示范文本库[EB/OL].(2022-06-17)[2025-03-05]. https://htsfwb.samr.gov.cn/
[16] Koroteev, Mikhail V. BERT: a review of applications in natural language processing and understanding[PP/OL].(2021-03-22)[2025-03-05]. https://arxiv.org/abs/2103.11943
[17] Cui Y, Che W, Liu T, et al. Revisiting pre-trained models for Chinese natural language processing[PP/OL]. V2.(2020-11-02)[2025-03-05]. https://doi.org/10.48550/arXiv.2004.13922
[18] Yan H, Yi B S, Li H X, et al. Sentiment knowledge-induced neural network for aspect-level sentiment analysis[J]. Neural Computing and Applications, 2022, 34(24): 22275-22286.
[19] Hong Y Z, Yu X G, He N, et al. FASPell: a fast, adaptable, simple, powerful Chinese spell checker based on DAE-decoder paradigm[C] //Proceedings of the 5th Workshop on Noisy User-generated Text(W-NUT 2019). Hong Kong, China:Association for Computational Linguistics, 2019: 160-169.
[20] Cheng X Y, Xu W D, Chen K L, et al. SpellGCN: incorporating phonological and visual similarities into language models for Chinese spelling check[PP/OL]. V2.(2020-03-13)[2025-03-05]. https://arxiv.org/abs/2004.14166
[21] Zhang Riqing, Pang Chao, Zhang Chuanqiang, et al. Correcting Chinese spelling errors with phonetic pre-training[C] //Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. [S.l.] : Association for Computational Linguistics, 2021: 2250- 2261.
[22] Xu Hengda, Li Zhongli, Zhou Qingyu, et al. Read, listen, and see: leveraging multimodal information helps Chinese spell checking[C] //Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021. [S.l.] : Association for Computational Linguistics, 2021: 716-728.
[23] Sun Y, Wang S H, Li Y K, et al. ERNIE 2.0: a continual pre-training framework for language understanding[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(5): 8968-8975.
[24] Loshchilov I, Hutter F. Decoupled weight decay regularization[PP/OL]. V3.(2019-01-04)[2025-03-05]. https://arxiv.org/abs/1711.05101
[25] Van der maaten L. Accelerating t-SNE using tree-based algorithms[J]. J Mach Learn Res, 2014, 15: 3221-3245.
[1] 刘志刚,冯涵霖,周元核,宿健珩,张岩. 面向跨分辨率身份匹配的行人重识别[J]. 山东大学学报 (工学版), 2026, 56(4): 1-9.
[2] 任红伟,孟菲,王继凯,田威杨,魏明召,程之恒,杜聪,吴建清. 基于深度学习的路基病害智能检测方法[J]. 山东大学学报 (工学版), 2026, 56(3): 93-105.
[3] 王倩,张瑞敏,李明津,孟宪静,耿蕾蕾. 基于频域图卷积网络的时空序列预测[J]. 山东大学学报 (工学版), 2026, 56(3): 84-92.
[4] 郑哲明, 孔玲玲, 何印. 基于双图结构的时空图卷积网络短期风电功率预测模型[J]. 山东大学学报 (工学版), 2026, 56(2): 130-138.
[5] 刘飞宇,张静,王亦楠. 基于渐近式特征融合的轻量化SAR舰船检测算法[J]. 山东大学学报 (工学版), 2026, 56(2): 52-59.
[6] 赵峰,刘瑞,王英,陈小强,葛磊蛟,马爱平. 多尺度融合与动态自校正旋转的吊弦检测算法[J]. 山东大学学报 (工学版), 2026, 56(2): 1-10.
[7] 王禹鸥,苑迎春,何振学,何晨. 融合多特征和多头自注意力机制的高校学业命名实体识别[J]. 山东大学学报 (工学版), 2025, 55(6): 35-44.
[8] 周群颖,隋家成,张继,王洪元. 基于自监督卷积和无参数注意力机制的工业品表面缺陷检测[J]. 山东大学学报 (工学版), 2025, 55(4): 40-47.
[9] 李丰,文益民. 融合多尺度视觉和文本语义特征的图像描述生成算法[J]. 山东大学学报 (工学版), 2025, 55(3): 80-87.
[10] 王禹鸥,苑迎春,何振学,王克俭. 改进RoBERTa、多实例学习和双重注意力机制的关系抽取方法[J]. 山东大学学报 (工学版), 2025, 55(2): 78-87.
[11] 邹正标,刘毅志,廖祝华,赵肄江. 动态交通流量预测的时空注意力图卷积网络[J]. 山东大学学报 (工学版), 2024, 54(5): 50-61.
[12] 李家春,李博文,常建波. 一种高效且轻量的RGB单帧人脸反欺诈模型[J]. 山东大学学报 (工学版), 2023, 53(6): 1-7.
[13] 王碧瑶,韩毅,崔航滨,刘毅超,任铭然,高维勇,陈姝廷,刘嘉巍,崔洋. 基于图像的道路语义分割检测方法[J]. 山东大学学报 (工学版), 2023, 53(5): 37-47.
[14] 宋佳芮,陈艳平,王凯,黄瑞章,秦永彬. 基于Affix-Attention的命名实体识别语义补充方法[J]. 山东大学学报 (工学版), 2023, 53(2): 70-76.
[15] 刘方旭,王建,魏本征. 基于多空间注意力的小儿肺炎辅助诊断算法[J]. 山东大学学报 (工学版), 2023, 53(2): 135-142.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!