融合场景多模态先验与稀疏注意力的文本图像超分辨率
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP391.41

基金项目:

自然资源部西南山地自然资源遥感监测工程技术创新中心开放课题(RSMNRSCM-2024-001);新疆生产建设兵团重点研发项目(2024AB064)


Fusing scene multimodal prior and sparse attention for text image super-resolution
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    受复杂背景、模糊、扭曲及变形等因素的影响,从低分辨率文本图像中恢复高分辨率图像极具挑战性.现有方法多依赖递归神经网络提取文本上下文信息,在捕捉长距离依赖及有效运用语义信息方面存在局限.为解决上述问题,本文提出一种融合场景多模态先验与稀疏注意力的文本图像超分辨率方法.首先,创新性地提出场景多模态先验分支,借助先进的内容解析单元和轮廓感知单元,充分挖掘并利用文本识别信息与视觉信息.其次,基于稀疏注意力的超分辨率增强模块从文本行提取上下文信息,并利用多头注意力机制的全局可见性构建字符间相关性,缓解处理长文本序列时的性能衰退.最后,引入结合梯度轮廓和文本结构感知的联合损失函数,显著增强模型提取文本轮廓及处理变形文本方面的能力.实验结果表明,相较于基线模型TATT,本文方法在TextZoom测试集的识别准确率平均提升4.3个百分点,平均峰值信噪比和结构相似性指数指标分别达到21.4 dB与0.790 9,提升了真实场景文本图像超分辨率的性能.

    Abstract:

    Recovering High-Resolution (HR) images from Low-Resolution (LR) text images is extremely challenging due to factors such as complex backgrounds,blurring,warping and distortion. Existing methods,which predominantly rely on recurrent neural networks to extract textual context,often struggle with capturing long-distance dependencies and leveraging semantic information effectively. To address the above problems,this paper proposes a novel text image super-resolution approach that integrates scene multimodal priors and sparse attention. First,we innovatively introduce a Scene Multimodal Prior Branch (SMPB),which leverages an advanced context parsing unit and a contour perception unit to fully explore and utilize both textual and visual information. Second,a Sparse Attention-Based Super-Resolution (SABSR) enhancement module is designed to extract contextual information from text lines. By utilizing the global receptive field of the multi-head attention mechanism,it effectively constructs inter-character correlations,thereby alleviating the performance degradation commonly encountered with long text sequences. Finally,the model's capability in extracting text contours and processing deformed text is significantly enhanced by a joint loss function that combines gradient contour awareness with text structure perception. Experimental results show that compared to the baseline model Text ATTention network(TATT),our method improves recognition accuracy on the TextZoom test set by 4.3 percentage points on average. Furthermore,it achieves average PSNR and SSIM values of 21.4 dB and 0.790 9,respectively,demonstrating superior performance for text image super-resolution in real-world scenarios.

    参考文献
    相似文献
    引证文献
引用本文

周颖,易尧华,余长慧,饶杨莉,王颖洁.融合场景多模态先验与稀疏注意力的文本图像超分辨率[J].南京信息工程大学学报(自然科学版),2026,(3):310-320
ZHOU Ying, YI Yaohua, YU Changhui, RAO Yangli, WANG Yingjie. Fusing scene multimodal prior and sparse attention for text image super-resolution[J]. Journal of Nanjing University of Information Science & Technology, 2026,(3):310-320

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-04-13
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-06-06
  • 出版日期:
文章二维码

地址:江苏省南京市宁六路219号    邮编:210044

联系电话:025-58731025    E-mail:nxdxb@nuist.edu.cn

南京信息工程大学学报 ® 2026 版权所有  技术支持:北京勤云科技发展有限公司