Basis for accurate protein pKa prediction with
machine learning

Zhitao Cai; Tengzi Liu; Qiaoling Lin; Jiahao He; Xiaowei Lei; Fangfang Luo; Yandong Huang

doi:10.26434/chemrxiv-2023-g4qpb

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

Basis for accurate protein pKa prediction with machine learning

30 January 2023, Version 1

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

pH regulates protein structures and the resulting functions in many biological processes via protonation and deprotonation of ionizable side chains where the titration equilibra is determined by pKa. To accelerate pH-dependent molecular mechanism research in life science or industrial protein and drug designs, fast and accurate pKa prediction is crucial. Here we present a theoretical pK data set PHMD549, which was successfully applied to four distinct machine learning methods, including DeepKa that was proposed in our previous work. To reach a valid comparison, EXP67S was selected as the test set. Encouragingly, DeepKa was improved significantly and outperforms other state-of-the-art methods, except for the constant-pH molecular dynamics, which was utilized to create PHMD549. More importantly, DeepKa reproduced experimental pKa orders of acidic dyads in five enzyme catalytic sites. Apart from structural proteins, DeepKa was found applicable to intrinsically disordered peptides. Further, in combination with solvent exposures, it's revealed that DeepKa offers the most accurate prediction under the challenging circumstance that hydrogen bonding or salt bridge interaction is partly compensated by desolvation for a buried side chain. Finally, our benchmark data qualify PHMD549 and EXP67S as the basis for future developments of protein pKa prediction tools driven by artificial intelligence. In addition, DeepKa built on PHMD549 has been proved an efficient protein pKa predictor and thus can be applied immediately to, for example, pKa database construction, protein design, drug discovery and so on.

Keywords

molecular dynamics

deep learning

pKa prediction

Supplementary materials

Title

Description

Actions

Title

Supporting Information

Description

Supplemental figures and tables including statistics of ionizable residues in proteins, explanation of Henderson-Hasselbalch equation, pKa convergence plot by CpHMD simulations, statistics of pKa database PHMD549, overlaps of four test sets, structures and pKa's of two HEWL proteins, residue type-specific assessment of DeepKa, pKa values of acids in NUPR1 protein, RMSD estimation for PKAI+ and pKa shift distributions and solvent accessibility calculations for PDB and AlphaFold structures.

Actions

Supplementary weblinks

Title

Description

Actions

Title

Code and data

Description

The code of DeepKa and the relevant data, including the pKa database PHMD549 and test set EXP67S.

Actions

View

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Now Published

Basis for Accurate Protein pKa Prediction with Machine Learning

Zhitao Cai, Tengzi Liu, Qiaoling Lin, Jiahao He, Xiaowei Lei, Fangfang Luo, Yandong Huang journal article

Journal of Chemical Information and Modeling , Volume 63, Issue 10

Online publication date: May 05, 2023

Version History

Jan 30, 2023 Version 1

Metrics

1,152

428

Views

Downloads

Citations

License

The content is available under CC BY NC 4.0

DOI

10.26434/chemrxiv-2023-g4qpb

Funding

National Natural Science Foundation of China

11804114

National Key R&D Program of China

2018YFD0901004

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

Basis for accurate protein pKa prediction with machine learning

Authors

Abstract

Keywords

Supplementary materials

Supplementary weblinks

Comments

Now Published

Version History

Metrics

License

DOI

Funding

Author’s competing interest statement

Ethics

Share