基于改进扩散模型结合条件控制的文本图像生成算法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP391.41

基金项目:

国家自然科学基金(11861003);辽宁省教育厅高等学校基本科研项目(LJKZ0157)


Text-to-image generation based on improved diffusion model combined with conditional control
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对现有的文本图像生成方法存在图像保真度低、图像生成操作难度大、仅适用于特定的任务场景等问题,提出一种新型的基于扩散模型的文本生成图像方法.该方法将扩散模型作为主要网络,设计一种新型结构的残差块,有效提升模型生成性能;通过添加注意力模块CBAM来改进噪声估计网络,增强了模型对图像关键信息的提取能力,进一步提高了生成图像质量;结合条件控制网络,有效地实现了特定姿势的文本图像生成.与KNN-Diffusion、CogView2、textStyleGAN、SimpleDiffusion等方法在数据集CelebA-HQ上做了定性、定量分析以及消融实验,根据评价指标以及生成结果显示,本文方法能够有效提高文本生成图像的质量,FID平均下降36.4%,Inception Score(IS)和结构相似性指数(SSIM)分别平均提高11.4%和3.9%,验证了本文算法的有效性.同时,本文模型结合了ControlNet网络,实现了定向动作的文本图像生成.

    Abstract:

    A novel text-to-image generation method based on diffusion model is proposed to address the problems of low image fidelity,complex generation operation,and narrow applicability to specific task scenarios in existing text-to-image generation methods.This approach takes a diffusion model as the backbone network and designs a novel residual block structure to enhance generation performance.Additionally,a CBAM (Convolutional Block Attention Module) is integrated into the noise estimation network to improve the model's ability to extract key image information,thereby improving output quality.By combining conditional control networks,the approach achieves precise text-to-image generation with user-specific poses.Qualitative and quantitative analyses,along with ablation experiments,were conducted on the CelebA HQ dataset against methods such as KNN-Diffusion,CogView2,textStyleGAN,and Simple diffusion.Evaluation metrics and generation results demonstrate that,the proposed method effectively improves generation quality,with an average decrease of 36.4% in FID (the Fréchet Inception Distance),average increases of 11.4% in IS (Inception Score) and 3.9% in SSIM (Structural Similarity).These results validate the effectiveness of the proposed approach.Furthermore,by integrating the ControlNet framework,the model enables text-to-image generation with controllable directional poses.

    参考文献
    相似文献
    引证文献
引用本文

杜洪波,薛皓元,朱立军.基于改进扩散模型结合条件控制的文本图像生成算法[J].南京信息工程大学学报(自然科学版),2025,17(5):611-623
DU Hongbo, XUE Haoyuan, ZHU Lijun. Text-to-image generation based on improved diffusion model combined with conditional control[J]. Journal of Nanjing University of Information Science & Technology, 2025,17(5):611-623

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2024-06-19
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2025-10-18
  • 出版日期: 2025-09-28
文章二维码

地址:江苏省南京市宁六路219号    邮编:210044

联系电话:025-58731025    E-mail:nxdxb@nuist.edu.cn

南京信息工程大学学报 ® 2026 版权所有  技术支持:北京勤云科技发展有限公司