Learning on Compressed Molecular Representations

Jan Weinreich; Daniel Probst

doi:10.26434/chemrxiv-2023-v1s2s-v3

Biological and Medicinal Chemistry

Search within Biological and Medicinal Chemistry

Learning on Compressed Molecular Representations

07 February 2024, Version 3

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small molecules or single-letter amino acid sequences for proteins. Here, we extend the previously introduced approach with support for regression and multitask classification and subsequently apply it to the prediction of molecular properties and protein-ligand binding affinities. We further propose converting numerical descriptors into string representations, enabling the integration of text input with domain-informed numerical descriptors. Finally, we show that the method can achieve performance competitive with chemical fingerprint- and GNN-based methodologies in general, and perform better than comparable methods on quantum chemistry and protein-ligand binding affinity prediction tasks.

Keywords

machine learning

artificial intelligence

gzip

compression

data set

Supplementary weblinks

Title

Description

Actions

Title

GitHub Repostiroy

Description

The GitHub repository containing all the code and data described in the manuscript.

Actions

View

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Version History

Feb 07, 2024 Version 3

Sep 29, 2023 Version 2

Jul 20, 2023 Version 1

Version Notes

Overall changes due to resubmission. - Title change - Abstract change - Ran additional benchmarks - Added ECFP-based kNN as control in benchmarks

Metrics

3,244

1,685

Views

Downloads

Citations

License

The content is available under CC BY 4.0

DOI

10.26434/chemrxiv-2023-v1s2s-v3

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

Learning on Compressed Molecular Representations

Authors

Abstract

Keywords

Supplementary weblinks

Comments

Version History

Version Notes

Metrics

License

DOI

Author’s competing interest statement

Ethics

Share