Benchmarking molecular feature attribution methods with activity cliffs

José Jiménez Luna; Miha Skalic; Nils Weskamp

doi:10.26434/chemrxiv-2021-pp88m

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

Benchmarking molecular feature attribution methods with activity cliffs

17 September 2021, Version 1

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

Feature attribution techniques are popular choices within the explainable artificial intelligence toolbox, as they can help elucidate which parts of the provided inputs used by an underlying supervised-learning method are considered relevant for a specific prediction. In the context of molecular design, these approaches typically involve the coloring of molecular graphs, whose presentation to medicinal chemists can be useful for making a decision of which compounds to synthesize or prioritize. The consistency of the highlighted moieties alongside expert background knowledge is expected to contribute to the understanding of machine-learning models in drug design. Quantitative evaluation of such coloring approaches, however, has so far been limited to substructure identification tasks. We here present an approach that is based on maximum common substructure algorithms applied to experimentally-determined activity cliffs. Using the proposed benchmark, we found that molecule coloring approaches in conjunction with classical machine-learning models tend to outperform more modern, deep-learning-based alternatives. However, none of the tested feature attribution methods sufficiently and consistently generalized when confronted with unseen examples.

Keywords

graph neural networks

Supplementary materials

Title

Description

Actions

Title

Supporting data

Description

Influence of variables such as molecular similarity between training and benchmark sets, training set size, and out-of-fold performance on color agreement for all model combinations.

Actions

Supplementary weblinks

Title

Description

Actions

Title

GitHub Repository

Description

Accompanying code for replication of the results.

Actions

View

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Now Published

Benchmarking Molecular Feature Attribution Methods with Activity Cliffs

José Jiménez-Luna, Miha Skalic, Nils Weskamp journal article

Journal of Chemical Information and Modeling , Volume 62, Issue 2

Online publication date: Jan 12, 2022

Version History

Sep 17, 2021 Version 1

Metrics

3,321

1,441

Views

Downloads

Citations

License

The content is available under CC BY NC 4.0

DOI

10.26434/chemrxiv-2021-pp88m

Funding

Boehringer Ingelheim

ETH RETHINK

SNF

205321_182176

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

Benchmarking molecular feature attribution methods with activity cliffs

Authors

Abstract

Keywords

Supplementary materials

Supplementary weblinks

Comments

Now Published

Version History

Metrics

License

DOI

Funding

Author’s competing interest statement

Ethics

Share