Linear Graphlet Models for Accurate and Interpretable Cheminformatics

Michael Tynes; Michael G Taylor; Jan Janssen; Daniel J Burrill; Danny Perez; Ping Yang; Nicholas Lubbers

doi:10.26434/chemrxiv-2024-r81c8-v2

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

Linear Graphlet Models for Accurate and Interpretable Cheminformatics

05 August 2024, Version 2

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

Advances in machine learning have given rise to a plurality of data-driven methods for predicting chemical properties from molecular structure. For many decades, the cheminformatics field has relied heavily on structural fingerprinting, while in recent years much focus has shifted toward leveraging highly parameterized deep neural networks which usually maximize accuracy. Beyond accuracy, to be useful and trustworthy in scientific applications, machine learning techniques often need intuitive explanations for model predictions and uncertainty quantification techniques so a practitioner might know when a model is appropriate to apply to new data. Here we revisit graphlet histogram fingerprints and introduce several new elements. We show that linear models built on graphlet fingerprints attain accuracy that is competitive with the state of the art while retaining an explainability advantage over black-box approaches. We show how to produce precise explanations of predictions by exploiting the relationships between molecular graphlets and show that these explanations are consistent with chemical intuition, experimental measurements, and theoretical calculations. Finally, we show how to use the presence of unseen fragments in new molecules to adjust predictions and quantify uncertainty.

Keywords

graph

graphlet

interpretability

uncertainty quantification

machine learning

molecular graph

QSAR

Supplementary materials

Title

Description

Actions

Title

Supplementary Information

Description

Supplementary Information including figures, tables, and other details referenced in the text.

Actions

Supplementary weblinks

Title

Description

Actions

Title

Minervachem

Description

A python library for cheminformatics and machine learning

Actions

View

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Version History

Aug 05, 2024 Version 2

Feb 26, 2024 Version 1

Version Notes

Updated to address peer review comments.

Metrics

1,264

681

Views

Downloads

Citations

License

The content is available under CC BY NC ND 4.0

DOI

10.26434/chemrxiv-2024-r81c8-v2

Funding

United States Department of Energy

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

Linear Graphlet Models for Accurate and Interpretable Cheminformatics

Authors

Abstract

Keywords

Supplementary materials

Supplementary weblinks

Comments

Version History

Version Notes

Metrics

License

DOI

Funding

Author’s competing interest statement

Ethics

Share