Can Diffusion Model Achieve Better Performance in Text Generation? Bridging the Gap between Training and Inference! - 专知论文

会员服务 ·

0

推断 · Better · Performer · MoDELS · Processing（编程语言） ·

2023 年 5 月 8 日

Can Diffusion Model Achieve Better Performance in Text Generation? Bridging the Gap between Training and Inference!

翻译：暂无翻译

Zecheng Tang,Pinzheng Wang,Keyan Zhou,Juntao Li,Ziqiang Cao,Min Zhang

Diffusion models have been successfully adapted to text generation tasks by mapping the discrete text into the continuous space. However, there exist nonnegligible gaps between training and inference, owing to the absence of the forward process during inference. Thus, the model only predicts based on the previously generated reverse noise rather than the noise computed by the forward process. Besides, the widely-used downsampling strategy in speeding up the inference will cause the mismatch of diffusion trajectories between training and inference. To understand and mitigate the above two types of training-inference discrepancies, we launch a thorough preliminary study. Based on our observations, we propose two simple yet effective methods to bridge the gaps mentioned above, named Distance Penalty and Adaptive Decay Sampling. Extensive experiments on \textbf{6} generation tasks confirm the superiority of our methods, which can achieve $100\times \rightarrow 200\times$ speedup with better performance.

翻译：暂无翻译

0

相关内容

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

Fe掺杂CuGaS2中间带薄膜材料的制备及光电特性

国家自然科学基金

0+阅读 · 2014年12月31日

拓扑绝缘体电子结构磁性调控的第一性原理研究

国家自然科学基金

0+阅读 · 2014年12月31日

调节性T细胞（Tregs）参与非结核分枝杆菌慢性感染分子免疫调节机制的研究

国家自然科学基金

0+阅读 · 2013年12月31日

石墨烯－有机半导体界面结构及结构与性能关系

国家自然科学基金

0+阅读 · 2012年12月31日

拓扑半金属Sb薄膜的分子束外延生长、能带结构调控和原位同步辐射ARPES研究

国家自然科学基金

0+阅读 · 2012年12月31日

A Natural Bias for Language Generation Models

Arxiv

0+阅读 · 2023年6月23日

Generative Multimodal Entity Linking

Arxiv

0+阅读 · 2023年6月22日

Combining multi-spectral data with statistical and deep-learning models for improved exoplanet detection in direct imaging at high contrast

Arxiv

0+阅读 · 2023年6月21日

Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse Training

Arxiv

0+阅读 · 2023年6月21日

Neural Architecture Search without Training

Neural Architecture Search without Training

Arxiv

10+阅读 · 2021年6月11日

VIP会员

文章信息

相关主题

Processing（编程语言）

相关VIP内容

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

智能体工程（Agent Engineering）

《全球地缘政治环境中的反无人机系统互操作性》252页

专业软件开发者不靠“氛围编程”（Vibe Coding），而靠“控制”：2025 年 AI Agent 在编程中的应用研究

基于大语言模型的智能体化软件问题解决：综述

相关资讯

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

直播 | Interpretable and Trustworthy Graph Geometric Deep Learning

图与推荐

2+阅读 · 2022年11月2日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

相关论文

A Natural Bias for Language Generation Models

Arxiv

0+阅读 · 2023年6月23日

Generative Multimodal Entity Linking

Arxiv

0+阅读 · 2023年6月22日

Combining multi-spectral data with statistical and deep-learning models for improved exoplanet detection in direct imaging at high contrast

Arxiv

0+阅读 · 2023年6月21日

Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse Training

Arxiv

0+阅读 · 2023年6月21日

Neural Architecture Search without Training

Neural Architecture Search without Training

Arxiv

10+阅读 · 2021年6月11日

相关基金

Fe掺杂CuGaS2中间带薄膜材料的制备及光电特性

国家自然科学基金

0+阅读 · 2014年12月31日

拓扑绝缘体电子结构磁性调控的第一性原理研究

国家自然科学基金

0+阅读 · 2014年12月31日

调节性T细胞（Tregs）参与非结核分枝杆菌慢性感染分子免疫调节机制的研究

国家自然科学基金

0+阅读 · 2013年12月31日

石墨烯－有机半导体界面结构及结构与性能关系

国家自然科学基金

0+阅读 · 2012年12月31日

拓扑半金属Sb薄膜的分子束外延生长、能带结构调控和原位同步辐射ARPES研究

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员