MASSA Algorithm: automated rational sampling of training and test subsets for QSAR modelling

Gabriel Corrêa Veríssimo; Simone Queiroz Panteleão; Jadson Castro Gertrudes; Thales Kronenberger; Kathia Maria Honório; Vinicius Gonçalves Maltarollo

doi:10.26434/chemrxiv-2022-dct7l-v2

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

MASSA Algorithm: automated rational sampling of training and test subsets for QSAR modelling

15 May 2023, Version 2

This is not the most recent version. There is a

newer version

of this content available

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

QSAR models capable of predicting biological, toxicity, and pharmacokinetic properties were widely used to search lead bioactive molecules in chemical databases. The dataset’s preparation to build these models has a strong influence on the quality of the generated models, and sampling requires that the original dataset be divided into training (for model training) and test (for statistical evaluation) sets. This sampling can be done randomly or rationally, but the rational division is superior. In this paper, we present MASSA, a Python tool that can be used to automatically sample datasets by exploring the biological, physicochemical, and structural spaces of molecules using PCA, HCA, and K-modes. The proposed algorithm is very useful when the variables used for QSAR are not available or to construct multiple QSAR models with the same training and test sets, producing models with lower variability and better values for validation metrics. These results were obtained even when the descriptors used in the QSAR/QSPR were different from those used in the separation of training and test sets, indicating that this tool can be used to build models for more than one QSAR/QSPR technique. Finally, this tool also generates useful graphical representations that can provide insights into the data.

Keywords

QSAR

Training and test splitting

Sampling

Hierarchical clustering analysis (HCA)

K-modes

Python

Supplementary materials

Title

Description

Actions

Title

Supplementary Material

Description

Similarity maps for the seven datasets and their training-test distributions from MASSA, random, and referential (from the original study) approaches.

Actions

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Version History

Sep 29, 2023 Version 3

May 15, 2023 Version 2

Aug 16, 2022 Version 1

Version Notes

Revised version after peer reviewing and rejection from IEEE/ACM Transactions on Computational Biology and Bioinformatics. Major changes comprise grammatical corrections; explanations about the implementation of the elbow method distance calculations; differentiation between K-mean and K-modes; comparison between employment randomly sampled datasets, and rational (stratifies) one exemplifying with a worst-case scenario.

Metrics

1,220

760

Views

Downloads

License

The content is available under CC BY NC ND 4.0

DOI

10.26434/chemrxiv-2022-dct7l-v2

Funding

CNPq

FAPEMIG

FAPESP

CAPES

Pró-Reitoria de Pesquisa of the Universidade Federal de Minas Gerais

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

MASSA Algorithm: automated rational sampling of training and test subsets for QSAR modelling

Authors

Abstract

Keywords

Supplementary materials

Comments

Version History

Version Notes

Metrics

License

DOI

Funding

Author’s competing interest statement

Ethics

Share