MASSA Algorithm: automated rational sampling of training and test subsets for QSAR modelling

Gabriel Corrêa Veríssimo; Simone Queiroz Pantaleão; Philipe de Oliveira Fernandes; Jadson Castro Gertrudes; Thales Kronenberger; Kathia Maria Honório; Vinicius Gonçalves Maltarollo

doi:10.26434/chemrxiv-2022-dct7l-v3

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

MASSA Algorithm: automated rational sampling of training and test subsets for QSAR modelling

29 September 2023, Version 3

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

QSAR models capable of predicting biological, toxicity, and pharmacokinetic properties were widely used to search lead bioactive molecules in chemical databases. The dataset’s preparation to build these models has a strong influence on the quality of the generated models, and sampling requires that the original dataset be divided into training (for model training) and test (for statistical evaluation) sets. This sampling can be done randomly or rationally, but the rational division is superior. In this paper, we present MASSA, a Python tool that can be used to automatically sample datasets by exploring the biological, physicochemical, and structural spaces of molecules using PCA, HCA, and K-modes. The proposed algorithm is very useful when the variables used for QSAR are not available or to construct multiple QSAR models with the same training and test sets, producing models with lower variability and better values for validation metrics. These results were obtained even when the descriptors used in the QSAR/QSPR were different from those used in the separation of training and test sets, indicating that this tool can be used to build models for more than one QSAR/QSPR technique. Finally, this tool also generates useful graphical representations that can provide insights into the data.

Keywords

QSAR

Training and test splitting

Sampling

Hierarchical clustering analysis (HCA)

K-modes

Python

Supplementary materials

Title

Description

Actions

Title

Supplementary Material

Description

Similarity maps for the seven datasets and their training-test distributions from MASSA, random, and referential (from the original study) approaches.

Actions

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Now Published

MASSA Algorithm: an automated rational sampling of training and test subsets for QSAR modeling

Gabriel Corrêa Veríssimo, Simone Queiroz Pantaleão, Philipe de Olveira Fernandes, Jadson Castro Gertrudes, Thales Kronenberger, Kathia Maria Honorio, Vinícius Gonçalves Maltarollo journal article

Journal of Computer-Aided Molecular Design , Volume 37, Issue 12

Online publication date: Oct 07, 2023

Version History

Sep 29, 2023 Version 3

May 15, 2023 Version 2

Aug 16, 2022 Version 1

Version Notes

Comparison with other algorithms (Kennard-Stone, SPXY, and Sphere Exclusion) and modeling with Random Forest as representative of the machine learning method were added to this version. A new author was added (Philipe de Oliveira Fernandes) due to the execution of the applicability domain of RF models and the help on the analysis/writing.

Metrics

1,212

759

Views

Downloads

Citations

License

The content is available under CC BY NC ND 4.0

DOI

10.26434/chemrxiv-2022-dct7l-v3

Funding

CNPq

FAPEMIG

FAPESP

CAPES

Pró-Reitoria de Pesquisa of the Universidade Federal de Minas Gerais

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

MASSA Algorithm: automated rational sampling of training and test subsets for QSAR modelling

Authors

Abstract

Keywords

Supplementary materials

Comments

Now Published

Version History

Version Notes

Metrics

License

DOI

Funding

Author’s competing interest statement

Ethics

Share