GEN: Highly Efficient SMILES Explorer Using Autodidactic Generative Examination Networks

Ruud van Deursen; Peter Ertl; Igor Tetko; Guillaume Godin

doi:10.26434/chemrxiv.9796874.v1

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

GEN: Highly Efficient SMILES Explorer Using Autodidactic Generative Examination Networks

12 September 2019, Version 1

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

Recurrent neural networks have been widely used to generate millions of de novo molecules in a known chemical space. These deep generative models are typically setup with LSTM or GRU units and trained with canonical SMILES. In this study, we introduce a new robust architecture, Generative Examination Network GEN, based on bidirectional RNNs with concatenated sub-models to learn and generate molecular SMILES within a trained target space. GENs autonomously learn the target space in a few epochs while being subjected to an independent online examination to measure the quality of the generated set. Here we have used online statistical quality control (SQC) on the percentage of valid molecular SMILES as examination measure to select the earliest available stable model weights. Very high levels of valid SMILES (95-98%) can be generated using multiple parallel encoding layers in combination with SMILES augmentation using unrestricted SMILES randomization. Our architecture combines an excellent novelty rate (85-90%) while generating SMILES with strong conservation of the property space (95-99%). Our flexible examination mechanism is open to other quality criteria.

Keywords

SMILES string representation

assessment measures

quality control mechanisms

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Now Published

GEN: highly efficient SMILES explorer using autodidactic generative examination networks

Ruud van Deursen, Peter Ertl, Igor V. Tetko, Guillaume Godin journal article

Journal of Cheminformatics , Volume 12, Issue 1

Online publication date: Apr 10, 2020

Version History

Sep 12, 2019 Version 1

Metrics

2,860

549

Views

Downloads

License

The content is available under CC BY NC ND 4.0

DOI

10.26434/chemrxiv.9796874.v1

Author’s competing interest statement

No conflict of interest

GEN: Highly Efficient SMILES Explorer Using Autodidactic Generative Examination Networks

Authors

Abstract

Keywords

Comments

Now Published

Version History

Metrics

License

DOI

Author’s competing interest statement

Share