CHA2: 向反分子设计进取的心电图 (CHA2: CHemistry Aware Convex Hull Autoencoder Towards Inverse Molecular Design)

Optimizing molecular design and discovering novel chemical structures to meet certain objectives, such as quantitative estimates of the drug-likeness score (QEDs), is NP-hard due to the vast combinatorial design space of discrete molecular structures, which makes it near impossible to explore the entire search space comprehensively to exploit de novo structures with properties of interest. To address this challenge, reducing the intractable search space into a lower-dimensional latent volume helps examine molecular candidates more feasibly via inverse design. Autoencoders are suitable deep learning techniques, equipped with an encoder that reduces the discrete molecular structure into a latent space and a decoder that inverts the search space back to the molecular design. The continuous property of the latent space, which characterizes the discrete chemical structures, provides a flexible representation for inverse design in order to discover novel molecules. However, exploring this latent space requires certain insights to generate new structures. We propose using a convex hall surrounding the top molecules in terms of high QEDs to ensnare a tight subspace in the latent representation as an efficient way to reveal novel molecules with high QEDs. We demonstrate the effectiveness of our suggested method by using the QM9 as a training dataset along with the Self- Referencing Embedded Strings (SELFIES) representation to calibrate the autoencoder in order to carry out the Inverse molecular design that leads to unfold novel chemical structure.

翻译：优化分子设计和发现新的化学结构以实现某些目标,例如对药物类比分(QEDs)的定量估计,由于离散分子结构的庞大组合设计空间,分子设计空间几乎不可能全面探索整个搜索空间,以全面利用具有相关属性的新结构。为了应对这一挑战,将棘手的搜索空间缩小为低维潜积体积,有助于通过反向设计对分子候选分子进行更易变的检查。自动编码器是合适的深层次学习技术,配有将离散分子结构降低到潜藏空间的编码器,以及将搜索空间反向分子设计的一个解码器。隐蔽空间的连续特性几乎无法全面探索整个搜索空间,从而利用离散化学结构的特性来全面探索新分子。然而,探索这一隐蔽空间需要一定的洞察力才能产生新的结构。我们提议使用高QED的螺旋门环环绕着顶部分子进入一个紧密的子空间,在潜伏层结构中使搜索空间反向分子结构转变。

相关内容

自编码器

关注 140

自动编码器是一种人工神经网络，用于以无监督的方式学习有效的数据编码。自动编码器的目的是通过训练网络忽略信号“噪声”来学习一组数据的表示（编码），通常用于降维。与简化方面一起，学习了重构方面，在此，自动编码器尝试从简化编码中生成尽可能接近其原始输入的表示形式，从而得到其名称。基本模型存在几种变体，其目的是迫使学习的输入表示形式具有有用的属性。自动编码器可有效地解决许多应用问题，从面部识别到获取单词的语义。

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

33页PPT【AI+天气预测】，AI and Machine learning for weather predictions

专知会员服务

35+阅读 · 2022年3月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日