Theoretical and Computational Chemistry

Discovery of Highly Polymorphic Organic Materials: A New Machine Learning Approach

Abstract

Polymorphism is the capacity of a molecule to adopt different conformations or molecular packing arrangements in the solid state. This is a key property to control during pharmaceutical manufacturing because it can impact a range of properties including stability and solubility. In this study, a novel approach based on machine learning classification methods is used to predict the likelihood for an organic compound to crystallise in multiple forms. A training dataset of drug-like molecules was curated from the Cambridge Structural Database (CSD) and filtered according to entries in the Drug Bank database. The number of separate forms in the CSD for each molecule was recorded. A metaclassifier was trained using this dataset to predict the expected number of crystalline forms from the compound descriptors. This approach was used to estimate the number of crystallographic forms for an external validation dataset. These results suggest this novel methodology can be used to predict the extent of polymorphism of new drugs or not-yet experimentally screened molecules. This promising method complements expensive ab initio methods for crystal structure prediction and as integral to experimental physical form screening, may identify systems that with unexplored potential.

Content

Thumbnail image of Machine learning-based approach to predict the polymorphability of organic compounds v11 chemxiv.pdf
download asset Machine learning-based approach to predict the polymorphability of organic compounds v11 chemxiv.pdf 1 MB [opens in a new tab]

Supplementary material

Thumbnail image of output of classifiers (Autosaved).xlsx
download asset output of classifiers (Autosaved).xlsx 0.09 MB [opens in a new tab]
output of classifiers (Autosaved)
Thumbnail image of SI1- Predictive models of polymorphism.docx
download asset SI1- Predictive models of polymorphism.docx 0.07 MB [opens in a new tab]
SI1- Predictive models of polymorphism
Thumbnail image of SI2- Experimental solvents screening.xlsx
download asset SI2- Experimental solvents screening.xlsx 0.03 MB [opens in a new tab]
SI2- Experimental solvents screening
Thumbnail image of Solvent.docx
download asset Solvent.docx 0.03 MB [opens in a new tab]
Solvent