Improving Compound-Protein Interaction Prediction by Self-Training with Augmenting Negative Samples

Takuto Koyama; Shigeyuki Matsumoto; Hiroaki Iwata; Ryosuke Kojima; Yasushi Okuno

doi:10.26434/chemrxiv-2023-jsrq5

Biological and Medicinal Chemistry

Search within Biological and Medicinal Chemistry

Improving Compound-Protein Interaction Prediction by Self-Training with Augmenting Negative Samples

22 February 2023, Version 1

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

Identifying compound-protein interactions (CPIs) is crucial for drug discovery. Because experimentally validating CPIs is often time-consuming and costly, computational approaches are expected to facilitate the process. Rapid growths of available CPI databases have accelerated the development of many machine learning methods for CPI predictions. However, their performance, particularly their generalizability against external data, often suffers from a data imbalance attributed to the lack of experimentally validated inactive (negative) samples. In this study, we developed a self-training method for augmenting both credible and informative negative samples to improve the performance of models impaired by data imbalances. The constructed model demonstrated a higher performance than those constructed with other conventional methods for solving data imbalances, and the improvement was prominent for external datasets. Moreover, examination of the prediction score thresholds for pseudo-labeling during self-training revealed that augmenting the samples with ambiguous prediction scores is beneficial for constructing a model with high generalizability. The present study provides guidelines for improving CPI predictions on real-world data, thus facilitating drug discovery.

Keywords

compound-protein interactions

model generalizability

self-training

Supplementary materials

Title

Description

Actions

Title

Supporting Information

Description

Complementary figures and additional details

Actions

Supplementary weblinks

Title

Description

Actions

Title

Data and Source Code

Description

Data and source code are provided

Actions

View

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Now Published

Improving Compound–Protein Interaction Prediction by Self-Training with Augmenting Negative Samples

Takuto Koyama, Shigeyuki Matsumoto, Hiroaki Iwata, Ryosuke Kojima, Yasushi Okuno journal article

Journal of Chemical Information and Modeling

Online publication date: Jul 17, 2023

Version History

Feb 22, 2023 Version 1

Metrics

943

346

Views

Downloads

Citations

License

The content is available under CC BY 4.0

DOI

10.26434/chemrxiv-2023-jsrq5

Funding

Japan Agency for Medical Research and Development

JP22nk0101111

Japan Society for the Promotion of Science

JP20K12063

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

Improving Compound-Protein Interaction Prediction by Self-Training with Augmenting Negative Samples

Authors

Abstract

Keywords

Supplementary materials

Supplementary weblinks

Comments

Now Published

Version History

Metrics

License

DOI

Funding

Author’s competing interest statement

Ethics

Share