Papers
Topics
Authors
Recent
Search
2000 character limit reached

Protein Multimer Structure Prediction via Prompt Learning

Published 29 Feb 2024 in cs.CE | (2402.18813v1)

Abstract: Understanding the 3D structures of protein multimers is crucial, as they play a vital role in regulating various cellular processes. It has been empirically confirmed that the multimer structure prediction~(MSP) can be well handled in a step-wise assembly fashion using provided dimer structures and predicted protein-protein interactions~(PPIs). However, due to the biological gap in the formation of dimers and larger multimers, directly applying PPI prediction techniques can often cause a \textit{poor generalization} to the MSP task. To address this challenge, we aim to extend the PPI knowledge to multimers of different scales~(i.e., chain numbers). Specifically, we propose \textbf{\textsc{PromptMSP}}, a pre-training and \textbf{Prompt} tuning framework for \textbf{M}ultimer \textbf{S}tructure \textbf{P}rediction. First, we tailor the source and target tasks for effective PPI knowledge learning and efficient inference, respectively. We design PPI-inspired prompt learning to narrow the gaps of two task formats and generalize the PPI knowledge to multimers of different scales. We provide a meta-learning strategy to learn a reliable initialization of the prompt model, enabling our prompting framework to effectively adapt to limited data for large-scale multimers. Empirically, we achieve both significant accuracy (RMSD and TM-Score) and efficiency improvements compared to advanced MSP models. The code, data and checkpoints are released at \url{https://github.com/zqgao22/PromptMSP}.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (52)
  1. Rl-mlzerd: Multimeric protein docking using reinforcement learning. Frontiers in Molecular Biosciences, 9:969394, 2022.
  2. The protein data bank. Nucleic acids research, 28(1):235–242, 2000.
  3. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  4. Predicting the structure of large protein complexes using alphafold and monte carlo tree search. Nature communications, 13(1):6028, 2022.
  5. Alphafold encodes the principles to identify high affinity peptide binders. BioRxiv, pp.  2022–03, 2022.
  6. Ranking peptide binders by affinity with alphafold. Angewandte Chemie, 135(7):e202213362, 2023.
  7. Multifaceted protein–protein interaction prediction based on siamese residual rcnn. Bioinformatics, 35(14):i305–i314, 2019.
  8. Wiener graph deconvolutional network improves graph self-supervised learning. In AAAI, pp.  7131–7139, 2023.
  9. Flexible protein-protein docking with a multi-track iterative transformer. bioRxiv, 2023.
  10. Structural analysis of protein complexes by cryo electron microscopy. Bacterial Protein Secretion Systems: Methods and Protocols, pp.  377–413, 2017.
  11. Multi-lzerd: multiple protein docking for asymmetric complexes. Proteins: Structure, Function, and Bioinformatics, 80(7):1818–1833, 2012.
  12. Protein complex prediction with alphafold-multimer. biorxiv, pp.  2021–10, 2021.
  13. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning, pp.  1126–1135. PMLR, 2017.
  14. Independent se (3)-equivariant models for end-to-end rigid protein docking. arXiv preprint arXiv:2111.07786, 2021.
  15. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723, 2020.
  16. Hierarchical graph learning for protein–protein interaction. Nature Communications, 14(1):1093, 2023a.
  17. Handling missing data via max-entropy regularized graph autoencoder. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pp.  7651–7659, 2023b.
  18. Improved docking of protein models by a combination of alphafold2 and cluspro. Biorxiv, pp.  2021–09, 2021.
  19. Bottom-up structural proteomics: cryoem of protein complexes enriched from the cellular milieu. Nature methods, 17(1):79–85, 2020.
  20. Protein structure determination by x-ray crystallography. Bioinformatics: Data, Sequence Analysis and Evolution, pp.  63–87, 2008.
  21. Prediction of multimolecular assemblies by multiple docking. Journal of molecular biology, 349(2):435–447, 2005.
  22. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  23. Diffdock-pp: Rigid protein-protein docking with diffusion models. arXiv preprint arXiv:2304.03889, 2023.
  24. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907, 2016.
  25. Network-based prediction of protein interactions. Nature communications, 10(1):1240, 2019.
  26. Semi-supervised graph classification: A hierarchical graph perspective. In The World Wide Web Conference, pp.  972–982, 2019.
  27. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021.
  28. A survey of graph meets large language model: Progress and future directions. arXiv preprint arXiv:2311.12399, 2023.
  29. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130, 2023.
  30. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35, 2023.
  31. Improving generalization in equivariant graph neural networks with physical inductive biases. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=3oTPsORaDH.
  32. Learning to predict reciprocity and triadic closure in social networks. ACM Transactions on Knowledge Discovery from Data (TKDD), 7(2):1–25, 2013.
  33. xtrimodock: Rigid protein docking via cross-modal representation learning and spectral algorithm. bioRxiv, pp.  2023–02, 2023.
  34. Protein x-ray crystallography and drug discovery. Molecules, 25(5):1030, 2020.
  35. Recent advances in natural language processing via large pre-trained language models: A survey. ACM Computing Surveys, 2021.
  36. Do vision transformers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34:12116–12128, 2021.
  37. Exploiting cloze questions for few shot text classification and natural language inference. arXiv preprint arXiv:2001.07676, 2020.
  38. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  39. Using strong triadic closure to characterize ties in social networks. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pp.  1466–1475, 2014.
  40. All in one: Multi-task prompting for graph neural networks. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining (KDD’23), pp.  2120–2131, 2023a.
  41. Graph prompt learning: A comprehensive survey and beyond. arXiv preprint arXiv:2311.16534, 2023b.
  42. Rethinking graph neural networks for anomaly detection. In International Conference on Machine Learning, pp.  21076–21089. PMLR, 2022.
  43. Gadbench: Revisiting and benchmarking supervised graph anomaly detection. In Thirty-seventh Conference on Neural Information Processing Systems, 2023.
  44. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  45. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  46. Learning harmonic molecular representations on riemannian manifold. arXiv preprint arXiv:2303.15520, 2023.
  47. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826, 2018.
  48. The hdock server for integrated protein–protein docking. Nature protocols, 15(5):1829–1852, 2020.
  49. Normalized l3-based link prediction in protein–protein interaction networks. BMC bioinformatics, 24(1):59, 2023.
  50. Differentiable prompt makes pre-trained language models better few-shot learners. arXiv preprint arXiv:2108.13161, 2021.
  51. Scoring function for automated assessment of protein structure template quality. Proteins: Structure, Function, and Bioinformatics, 57(4):702–710, 2004.
  52. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.  16816–16825, 2022.
Citations (7)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We found no open problems mentioned in this paper.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 12 likes about this paper.