根基 SMILES:化学反应预测的严密代表 (Root-aligned SMILES: A Tight Representation for Chemical Reaction Prediction) - 专知论文

会员服务 ·

0

Learning · 可约的 · 知识 (knowledge) · Performer · state-of-the-art ·

2022 年 8 月 12 日

Root-aligned SMILES: A Tight Representation for Chemical Reaction Prediction

翻译：根基 SMILES:化学反应预测的严密代表

Zipeng Zhong,Jie Song,Zunlei Feng,Tiantao Liu,Lingxiang Jia,Shaolun Yao,Min Wu,Tingjun Hou,Mingli Song

from arxiv, Chemical Science 2022. Main paper: 16 pages, 5 figures, and 6 tables; supplementary information: 8 pages, 5 figures and 3 tables. Code repository: https://github.com/otori-bird/retrosynthesis

Chemical reaction prediction, involving forward synthesis and retrosynthesis prediction, is a fundamental problem in organic synthesis. A popular computational paradigm formulates synthesis prediction as a sequence-to-sequence translation problem, where the typical SMILES is adopted for molecule representations. However, the general-purpose SMILES neglects the characteristics of chemical reactions, where the molecular graph topology is largely unaltered from reactants to products, resulting in the suboptimal performance of SMILES if straightforwardly applied. In this article, we propose the root-aligned SMILES (R-SMILES), which specifies a tightly aligned one-to-one mapping between the product and the reactant SMILES for more efficient synthesis prediction. Due to the strict one-to-one mapping and reduced edit distance, the computational model is largely relieved from learning the complex syntax and dedicated to learning the chemical knowledge for reactions. We compare the proposed R-SMILES with various state-of-the-art baselines and show that it significantly outperforms them all, demonstrating the superiority of the proposed method.

翻译：化学反应预测,包括前期合成和反转合成预测,是有机合成的一个根本问题。流行的计算模式将合成预测作为一种序列到顺序的翻译问题,对分子表示采用典型的SMILES。然而,通用SMILES忽略了化学反应的特性,分子图示表层基本上没有从反应剂向产品转变,因此如果直接应用的话,SMILES的性能不尽人意。我们在本篇文章中提议了根整齐的SMILES(R-SMILES),它规定了产品与反应剂SMILES之间的一对一的精确匹配绘图,以便更高效的合成预测。由于严格的一对一的绘图和缩短编辑距离,计算模型基本上从学习复杂的语法和专门学习反应的化学知识中解脱脱去。我们把拟议的R-SMILES与各种最新基准进行比较,并表明它大大超越了它们,显示了拟议方法的优越性。

0

相关内容

Learning

Bioinformatics | MultiGran-SMILES:用于分子性质预测的多粒度SMILES学习

Bioinformatics | MultiGran-SMILES:用于分子性质预测的多粒度SMILES学习

专知会员服务

12+阅读 · 2022年9月25日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

因果图，Causal Graphs，52页ppt

因果图，Causal Graphs，52页ppt

专知会员服务

252+阅读 · 2020年4月19日

【论文推荐】一种用于逆合成预测的图到图框架，A Graph to Graphs Framework for Retrosynthesis Prediction

【论文推荐】一种用于逆合成预测的图到图框架，A Graph to Graphs Framework for Retrosynthesis Prediction

专知会员服务

12+阅读 · 2020年4月1日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

表征天然丰度酵母细胞色素c多构象的液体14N NMR方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

含能分子团簇负离子的红外光解离光谱研究

国家自然科学基金

0+阅读 · 2013年12月31日

聚咔唑聚芴盘状高分子液晶的合成与光电性能

国家自然科学基金

0+阅读 · 2012年12月31日

pH响应离子液体的溶液化学研究

国家自然科学基金

0+阅读 · 2012年12月31日

新型核苷的分子设计、合成和抗乙肝病毒的活性研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于Smiles重排串联反应的并三环杂环体系的构建

国家自然科学基金

0+阅读 · 2011年12月31日

新型氮唑类抗真菌化合物的设计合成及构效关系研究

国家自然科学基金

0+阅读 · 2010年12月31日

Dyrk1A调控CaMKⅡ#948;的可变剪接及其在心脏重构过程中的作用

国家自然科学基金

0+阅读 · 2009年12月31日

硫化氢对心肌细胞内钙通道的调控作用及其分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

TR3相互作用新蛋白机理研究

国家自然科学基金

1+阅读 · 2008年12月31日

Expander Graph Propagation

Arxiv

0+阅读 · 2022年10月6日

Fault-tolerant Coding for Entanglement-Assisted Communication

Arxiv

0+阅读 · 2022年10月6日

Structured Multi-task Learning for Molecular Property Prediction

Arxiv

0+阅读 · 2022年10月6日

Particle clustering in turbulence: Prediction of spatial and statistical properties with deep learning

Arxiv

0+阅读 · 2022年10月5日

HYPRO: A Hybridly Normalized Probabilistic Model for Long-Horizon Prediction of Event Sequences

Arxiv

0+阅读 · 2022年10月4日

Mastering Spatial Graph Prediction of Road Networks

Arxiv

0+阅读 · 2022年10月3日

Gradient Gating for Deep Multi-Rate Learning on Graphs

Arxiv

0+阅读 · 2022年10月2日

Citation Trajectory Prediction via Publication Influence Representation Using Temporal Knowledge Graph

Arxiv

0+阅读 · 2022年10月2日

Metro: Memory-Enhanced Transformer for Retrosynthetic Planning via Reaction Tree

Arxiv

0+阅读 · 2022年9月30日

Deep Reinforcement Learning for List-wise Recommendations

Arxiv

13+阅读 · 2018年1月5日

VIP会员

文章信息

相关主题

知识 (knowledge)

state-of-the-art

相关VIP内容

Bioinformatics | MultiGran-SMILES:用于分子性质预测的多粒度SMILES学习

Bioinformatics | MultiGran-SMILES:用于分子性质预测的多粒度SMILES学习

专知会员服务

12+阅读 · 2022年9月25日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

因果图，Causal Graphs，52页ppt

因果图，Causal Graphs，52页ppt

专知会员服务

252+阅读 · 2020年4月19日

【论文推荐】一种用于逆合成预测的图到图框架，A Graph to Graphs Framework for Retrosynthesis Prediction

【论文推荐】一种用于逆合成预测的图到图框架，A Graph to Graphs Framework for Retrosynthesis Prediction

专知会员服务

12+阅读 · 2020年4月1日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【普林斯顿博士论文】在线学习：优化、控制与学习理论

不确定环境下无人机三维路径规划研究 | 221页

【NeurIPS2025】《LeapFactual：基于条件流匹配的可靠视觉反事实解释》

大语言模型将如何改变军事指挥结构

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Expander Graph Propagation

Arxiv

0+阅读 · 2022年10月6日

Fault-tolerant Coding for Entanglement-Assisted Communication

Arxiv

0+阅读 · 2022年10月6日

Structured Multi-task Learning for Molecular Property Prediction

Arxiv

0+阅读 · 2022年10月6日

Particle clustering in turbulence: Prediction of spatial and statistical properties with deep learning

Arxiv

0+阅读 · 2022年10月5日

HYPRO: A Hybridly Normalized Probabilistic Model for Long-Horizon Prediction of Event Sequences

Arxiv

0+阅读 · 2022年10月4日

Mastering Spatial Graph Prediction of Road Networks

Arxiv

0+阅读 · 2022年10月3日

Gradient Gating for Deep Multi-Rate Learning on Graphs

Arxiv

0+阅读 · 2022年10月2日

Citation Trajectory Prediction via Publication Influence Representation Using Temporal Knowledge Graph

Arxiv

0+阅读 · 2022年10月2日

Metro: Memory-Enhanced Transformer for Retrosynthetic Planning via Reaction Tree

Arxiv

0+阅读 · 2022年9月30日

Deep Reinforcement Learning for List-wise Recommendations

Arxiv

13+阅读 · 2018年1月5日

相关基金

表征天然丰度酵母细胞色素c多构象的液体14N NMR方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

含能分子团簇负离子的红外光解离光谱研究

国家自然科学基金

0+阅读 · 2013年12月31日

聚咔唑聚芴盘状高分子液晶的合成与光电性能

国家自然科学基金

0+阅读 · 2012年12月31日

pH响应离子液体的溶液化学研究

国家自然科学基金

0+阅读 · 2012年12月31日

新型核苷的分子设计、合成和抗乙肝病毒的活性研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于Smiles重排串联反应的并三环杂环体系的构建

国家自然科学基金

0+阅读 · 2011年12月31日

新型氮唑类抗真菌化合物的设计合成及构效关系研究

国家自然科学基金

0+阅读 · 2010年12月31日

Dyrk1A调控CaMKⅡ#948;的可变剪接及其在心脏重构过程中的作用

国家自然科学基金

0+阅读 · 2009年12月31日

硫化氢对心肌细胞内钙通道的调控作用及其分子机制

国家自然科学基金

0+阅读 · 2009年12月31日

TR3相互作用新蛋白机理研究

国家自然科学基金

1+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员