Size-Extensive Molecular Machine Learning with Global Descriptors

Hyunwook Jung; Sina Stocker; Christian Kunkel; Harald Oberhofer; Byungchan Han; Karsten Reuter; Johannes T. Margraf

doi:10.26434/chemrxiv.10002020.v1

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

Size-Extensive Molecular Machine Learning with Global Descriptors

22 October 2019, Version 1

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

Machine learning (ML) models are increasingly used to predict molecular prop- erties in a high-throughput setting at a much lower computational cost than con- ventional electronic structure calculations. Such ML models require descriptors that encode the molecular structure in a vector. These descriptors are generally designed to respect the symmetries and invariances of the target property. However, size- extensivity is usually not guaranteed for so-called global descriptors. In this contri- bution, we show how extensivity can be build into ML models with global descriptors such as the Many-Body Tensor Representation. Properties of extensive and non- extensive models for the atomization energy are systematically explored by training on small molecules and testing on small, medium and large molecules. Our result shows that the non-extensive model is only useful in the size-range of its training set, whereas the extensive models provide reasonable predictions across large size differences. Remaining sources of error for the extensive models are discussed.

Keywords

Machine learning

Kernel ridge regression

Many-body tensor representation

Size-extensivity

Atomization energy

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Now Published

Size‐Extensive Molecular Machine Learning with Global Representations

Hyunwook Jung, Sina Stocker, Christian Kunkel, Harald Oberhofer, Byungchan Han, Karsten Reuter, Johannes T. Margraf journal article

ChemSystemsChem , Volume 2, Issue 4

Online publication date: Feb 04, 2020

Version History

Oct 22, 2019 Version 1

Metrics

3,795

880

Views

Downloads

License

The content is available under CC BY NC ND 4.0

DOI

10.26434/chemrxiv.10002020.v1

Author’s competing interest statement

no conflict of interest

Size-Extensive Molecular Machine Learning with Global Descriptors

Authors

Abstract

Keywords

Comments

Now Published

Version History

Metrics

License

DOI

Author’s competing interest statement

Share