Skip to content
ai-proteomics:-from-protein-identification-to-virtual-cells-–-nature-methods

AI proteomics: from protein identification to virtual cells – Nature Methods

References

  1. Aebersold, R. & Mann, M. Mass spectrometry-based proteomics. Nature 422, 198–207 (2003).

    Article  PubMed  Google Scholar 

  2. Guo, T., Steen, J. A. & Mann, M. Mass-spectrometry-based proteomics: from single cells to clinical applications. Nature 638, 901–911 (2025). This review provides an up-to-date overview of MS-based proteomics, tracing its evolution from large-scale quantification to ultrasensitive single-cell and spatial analysis and discussing its translation into clinical applications such as biomarker discovery and diagnostics.

    Article  PubMed  Google Scholar 

  3. Lindsay, R. K., Buchanan, B. G., Feigenbaum, E. A. & Lederberg, J. DENDRAL: a case study of the first expert system for scientific hypothesis formation. Artif. Intell. 61, 209–261 (1993).

    Article  Google Scholar 

  4. Mann, M., Kumar, C., Zeng, W.-F. & Strauss, M. T. Artificial intelligence for proteomics and biomarker discovery. Cell Syst. 12, 759–770 (2021).

    Article  PubMed  Google Scholar 

  5. Kalhor, M., Lapin, J., Picciani, M. & Wilhelm, M. Rescoring peptide spectrum matches: boosting proteomics performance by integrating peptide property predictors into peptide identification. Mol. Cell. Proteomics 23, 100798 (2024).

    Article  PubMed  PubMed Central  Google Scholar 

  6. Demichev, V., Messner, C. B., Vernardis, S. I., Lilley, K. S. & Ralser, M. DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat. Methods 17, 41–44 (2020). A widely adopted tool based on a neural network for improving peptide and protein identification in MS-based proteomics.

    Article  PubMed  Google Scholar 

  7. Bittremieux, W. et al. Deep learning methods for de novo peptide sequencing. Mass Spectrom. Rev. 45, 507–526 (2026).

    Article  PubMed  Google Scholar 

  8. Xiao, Q. et al. High-throughput proteomics and AI for cancer biomarker discovery. Adv. Drug Deliv. Rev. 176, 113844 (2021).

  9. Kustatscher, G. et al. Understudied proteins: opportunities and challenges for functional proteomics. Nat. Methods 19, 774–779 (2022).

    Article  PubMed  Google Scholar 

  10. Lee, M. Recent advances in deep learning for protein-protein interaction analysis: a comprehensive review. Molecules 28, 5169 (2023).

    Article  PubMed  PubMed Central  Google Scholar 

  11. Qian, L. et al. AI-empowered perturbation proteomics for complex biological systems. Cell Genomics 4, 100691 (2024).

    Article  PubMed  PubMed Central  Google Scholar 

  12. Mani, D. R. et al. Cancer proteogenomics: current impact and future prospects. Nat. Rev. Cancer 22, 298–313 (2022).

    Article  PubMed  PubMed Central  Google Scholar 

  13. Savage, S. R. et al. Pan-cancer proteogenomics expands the landscape of therapeutic targets. Cell 187, 4389-4407.e15 (2024).

  14. Bunne, C. et al. How to build the virtual cell with artificial intelligence: priorities and opportunities. Cell 187, 7045–7063 (2024). This article presents a framework for building virtual cells with AI, combining multi-scale biological state representations and virtual instruments to simulate cellular dynamics and enable predictive modeling.

    Article  PubMed  PubMed Central  Google Scholar 

  15. Qian, L., Dong, Z. & Guo, T. Grow AI virtual cells: three data pillars and closed-loop learning. Cell Res. 35, 319–321 (2025).

    Article  PubMed  PubMed Central  Google Scholar 

  16. Perez-Riverol, Y. et al. The PRIDE database at 20 years: 2025 update. Nucleic Acids Res. 53, D543–D553 (2025).

    Article  PubMed  PubMed Central  Google Scholar 

  17. Doerr, A. Proteomics data reuse with MassIVE-KB. Nat. Methods 16, 26 (2019).

    Article  PubMed  Google Scholar 

  18. Panagakis, Y. et al. Tensor methods in computer vision and deep learning. Proc. IEEE 109, 863–890 (2021).

  19. Deng, J. et al. ImageNet: a large-scale hierarchical image database. In IEEE Conf. Computer Vision and Pattern Recognition 248–255 (IEEE, 2009).

  20. Berman, H. M. et al. The Protein Data Bank. Nucleic Acids Res. 28, 235–242 (2000).

    Article  PubMed  PubMed Central  Google Scholar 

  21. Omenn, G. S. et al. The 2024 report on the human proteome from the HUPO Human Proteome Project. J. Proteome Res. 23, 5296–5311 (2024).

  22. Deutsch, E. W. et al. Proteomics standards iInitiative at twenty years: current activities and future work. J. Proteome Res. 22, 287–301 (2023).

  23. Li, Y. et al. Proteogenomic data and resources for pan-cancer analysis. Cancer Cell 41, 1397–1406 (2023).

  24. Rodriguez, H., Zenklusen, J. C., Staudt, L. M., Doroshow, J. H. & Lowy, D. R. The next horizon in precision oncology: proteogenomics to inform cancer diagnosis and treatment. Cell 184, 1661–1670 (2021).

    Article  PubMed  PubMed Central  Google Scholar 

  25. Tully, B. et al. Addressing the challenges of high-throughput cancer tissue proteomics for clinical application: ProCan. Proteomics 19, 1900109 (2019).

  26. He, F. et al. π-HuB: the proteomic navigator of the human body. Nature 636, 322–331 (2024).

  27. Bittremieux, W., May, D. H., Bilmes, J. & Noble, W. S. A learned embedding for efficient joint analysis of millions of mass spectra. Nat. Methods 19, 675–678 (2022).

    Article  PubMed  PubMed Central  Google Scholar 

  28. Pan, Q., Shai, O., Lee, L. J., Frey, B. J. & Blencowe, B. J. Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing. Nat. Genet. 40, 1413–1415 (2008).

  29. Phan, L. et al. The evolution of dbSNP: 25 years of impact in genomic research. Nucleic Acids Res. 53, D925–D931 (2025).

  30. Chung, C.-R. et al. dbPTM 2025 update: comprehensive integration of PTMs and proteomic data for advanced insights into cancer research. Nucleic Acids Res. 53, D377–D386 (2025).

  31. Wen, B. et al. Assessment of false discovery rate control in tandem mass spectrometry analysis using entrapment. Nat. Methods 22, 1454–1463 (2025). This article establishes a rigorous theoretical framework for evaluating FDR control in proteomics using entrapment experiments.

  32. Tran, N. H. et al. NovoBoard: a comprehensive framework for evaluating the false discovery rate and accuracy of de novo peptide sequencing. Mol. Cell. Proteomics 23, 100849 (2024).

    Article  PubMed  PubMed Central  Google Scholar 

  33. Sanh, V., Debut, L., Chaumond, J. & Wolf, T. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. Preprint at arXiv https://doi.org/10.48550/arXiv.1910.01108 (2020).

  34. Shazeer, N. et al. Outrageously large neural networks: the sparsely-gated mixture-of-experts layer. Preprint at arXiv https://doi.org/10.48550/arXiv.1701.06538 (2017).

  35. von Mering, C. et al. Comparative assessment of large-scale data sets of protein–protein interactions. Nature 417, 399–403 (2002).

  36. Garlick, J. M. & Mapp, A. K. Selective modulation of dynamic protein complexes. Cell Chem. Biol. 27, 986–997 (2020).

    Article  PubMed  PubMed Central  Google Scholar 

  37. Huttlin, E. L. et al. Dual proteome-scale networks reveal cell-specific remodeling of the human interactome. Cell 184, 3022–3040.e28 (2021).

  38. Bludau, I. et al. Complex-centric proteome profiling by SEC-SWATH-MS for the parallel detection of hundreds of protein complexes. Nat. Protoc. 15, 2341–2386 (2020).

  39. Xu, Y., Fan, X. & Hu, Y. In vivo interactome profiling by enzyme‐catalyzed proximity labeling. Cell Biosci. 11, 27 (2021).

  40. Bajpai, A. K. et al. Systematic comparison of the protein-protein interaction databases from a user’s perspective. J. Biomed. Informatics 103, 103380 (2020).

  41. Wu, S., Zhang, S., Liu, C.-M., Fernie, A. R. & Yan, S. Recent advances in mass spectrometry-based protein interactome studies. Mol. Cell. Proteomics 24, 100887 (2025).

  42. Feng, S. et al. Hypergraph models of biological networks to identify genes critical to pathogenic viral response. BMC Bioinformatics 22, 287 (2021).

  43. Method of the Year 2024: spatial proteomics. Nat. Methods 21, 2195–2196 (2024).

  44. Goltsev, Y. et al. Deep profiling of mouse splenic architecture with CODEX multiplexed imaging. Cell 174, 968–981.e15 (2018).

  45. Angelo, M. et al. Multiplexed ion beam imaging of human breast tumors. Nat. Med. 20, 436–442 (2014).

  46. Mund, A. et al. Deep Visual Proteomics defines single-cell identity and heterogeneity. Nat. Biotechnol. 40, 1231–1240 (2022). This article introduces Deep Visual Proteomics, an AI-driven framework that integrates imaging and ultrasensitive mass spectrometry to map protein expression at single-cell resolution, advancing spatial and functional proteomics.

  47. Rosenberger, F. A. et al. Deep Visual Proteomics maps proteotoxicity in a genetic liver disease. Nature 642, 484–491 (2025).

  48. Nordmann, T. M. et al. Spatial proteomics identifies JAKi as treatment for a lethal skin disease. Nature 635, 1001–1009 (2024).

  49. Xu, Y. et al. Multimodal single cell-resolved spatial proteomics reveal pancreatic tumor heterogeneity. Nat. Commun. 15, 10100 (2024).

  50. Bury, A. G. et al. A subcellular cookie cutter for spatial genomics in human tissue. Anal. Bioanal. Chem. 414, 5483–5492 (2022).

  51. Dong, Z. et al. Spatial proteomics of single cells and organelles on tissue slides using filter-aided expansion proteomics. Nat. Commun. 15, 9378 (2024).

  52. Qin, R., Ma, J., He, F. & Qin, W. In-depth and high-throughput spatial proteomics for whole-tissue slice profiling by deep learning-facilitated sparse sampling strategy. Cell Discov. 11, 21 (2025).

    Article  PubMed  PubMed Central  Google Scholar 

  53. Hu, B. et al. High-resolution spatially resolved proteomics of complex tissues based on microfluidics and transfer learning. Cell 188, 734–748.e22 (2025). This study combines microfluidics with transfer learning to achieve high-resolution, spatially resolved proteomics of complex tissues, providing an AI-driven approach that enables precise and high-throughput mapping of tissue microenvironments and advances spatial proteomics applications.

  54. Mitchell, D. C. et al. A proteome-wide atlas of drug mechanism of action. Nat. Biotechnol. 41, 845–857 (2023).

  55. Zecha, J. et al. Decrypting drug actions and protein modifications by dose- and time-resolved proteomics. Science 380, 93–101 (2023).

    Article  PubMed  PubMed Central  Google Scholar 

  56. Eckert, S. et al. Decrypting the molecular basis of cellular drug phenotypes by dose-resolved expression proteomics. Nat. Biotechnol. 43, 406–415 (2025).

  57. Zhao, W. et al. Large-scale characterization of drug responses of clinically relevant proteins in cancer cell lines. Cancer Cell 38, 829–843.e4 (2020).

  58. Yuan, B. et al. CellBox: interpretable machine learning for perturbation biology with application to the design of cancer combination therapy. Cell Syst. 12, 128–140.e4 (2021). This study presents CellBox, an interpretable machine learning framework that models cellular responses to perturbations, enabling AI-driven prediction and rational design of effective cancer combination therapies and thereby advancing predictive and mechanistic proteomics in perturbation biology.

  59. Vogel, C. & Marcotte, E. M. Insights into the regulation of protein abundance from proteomic and transcriptomic analyses. Nat. Rev. Genet. 13, 227–232 (2012).

    Article  PubMed  PubMed Central  Google Scholar 

  60. Liu, Y., Beyer, A. & Aebersold, R. On the dependency of cellular protein levels on mRNA abundance. Cell 165, 535–550 (2016).

    Article  PubMed  Google Scholar 

  61. Chung, H. et al. Joint single-cell measurements of nuclear proteins and RNA in vivo. Nat. Methods 18, 1204–1212 (2021).

  62. Zitnik, M., Agrawal, M. & Leskovec, J. Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics 34, i457–i466 (2018).

    Article  PubMed  PubMed Central  Google Scholar 

  63. Cao, Z.-J. & Gao, G. Multi-omics single-cell data integration and regulatory inference with graph-linked embedding. Nat. Biotechnol. 40, 1458–1466 (2022).

  64. Gayoso, A. et al. Joint probabilistic modeling of single-cell multi-omic data with totalVI. Nat. Methods 18, 272–282 (2021).

  65. Gao, Y., Feng, Y., Ji, S. & Ji, R. HGNN+: general hypergraph neural networks. IEEE Trans. Pattern Anal. Mach. Intell. 45, 3181–3199 (2023).

  66. Xu, H. et al. A whole-slide foundation model for digital pathology from real-world data. Nature 630, 181–188 (2024).

  67. Senior, A. W. et al. Improved protein structure prediction using potentials from deep learning. Nature 577, 706–710 (2020).

  68. Cui, H. et al. Towards multimodal foundation models in molecular cell biology. Nature 640, 623–633 (2025).

  69. Roohani, Y. H. et al. Virtual Cell Challenge: toward a Turing test for the virtual cell. Cell 188, 3370–3374 (2025).

  70. Karr, J. R. et al. A whole-cell computational model predicts phenotype from genotype. Cell 150, 389–401 (2012).

  71. Macklin, D. N. et al. Simultaneous cross-evaluation of heterogeneous E. coli datasets via mechanistic simulation. Science 369, eaav3751 (2020).

  72. Ye, C. et al. Comprehensive understanding of Saccharomyces cerevisiae phenotypes with whole-cell model WM_S288C. Biotechnol. Bioeng. 117, 1562–1574 (2020).

  73. Österlund, T., Nookaew, I., Bordel, S. & Nielsen, J. Mapping condition-dependent regulation of metabolism in yeast through genome-scale modeling. BMC Syst. Biol. 7, 36 (2013).

  74. Rood, J. E. et al. The Human Cell Atlas from a cell census to a unified foundation model. Nature 637, 1065–1071 (2025).

  75. Theodoris, C. V. et al. Transfer learning enables predictions in network biology. Nature 618, 616–624 (2023).

  76. Hao, M. et al. Large-scale foundation model on single-cell transcriptomics. Nat. Methods 21, 1481–1491 (2024).

    Article  PubMed  Google Scholar 

  77. Sun, R. et al. A perturbation proteomics-based foundation model for virtual cell construction. Preprint at bioRxiv https://doi.org/10.1101/2025.02.07.637070 (2025).

  78. Adduri, A. K. et al. Predicting cellular responses to perturbation across diverse contexts with State. Preprint at bioRxiv https://doi.org/10.1101/2025.06.26.661135 (2025).

  79. Desiere, F. et al. The PeptideAtlas project. Nucleic Acids Res. 34, D655–D658 (2006).

  80. Choi, M. et al. MassIVE.quant: a community resource of quantitative mass spectrometry–based proteomics datasets. Nat. Methods 17, 981–984 (2020).

  81. Dai, C. et al. quantms: a cloud-based pipeline for quantitative proteomics enables the reanalysis of public proteomics data. Nat. Methods 21, 1603–1607 (2024). This study introduces quantms, a cloud-based, automated pipeline for quantitative proteomics that facilitates large-scale reanalysis of public datasets, advancing AI-driven proteomics by providing standardized, accessible and reproducible data resources for machine learning applications.

  82. Liu, Z. et al. DIA-BERT: pre-trained end-to-end transformer models for enhanced DIA proteomics data analysis. Nat. Commun. 16, 3530 (2025). This study develops DIA-BERT, a pretrained end-to-end transformer model that enhances DIA proteomics data analysis, demonstrating how large language model architectures can be adapted to improve peptide identification, quantification and overall AI-driven interpretation of MS data.

  83. Gao, H. et al. iDIA-QC: AI-empowered data-independent acquisition mass spectrometry-based quality control. Nat. Commun. 16, 892 (2025).

    Article  PubMed  PubMed Central  Google Scholar 

  84. Jun, A. et al. MassNet: billion-scale AI-friendly mass spectral corpus enables robust de novo peptide sequencing. Preprint at bioRxiv https://doi.org/10.1101/2025.06.20.660691 (2025).

  85. Rehfeldt, T. G. et al. ProteomicsML: an online platform for community-curated data sets and tutorials for machine learning in proteomics. J. Proteome Res. 22, 632–636 (2023).

    Article  PubMed  PubMed Central  Google Scholar 

  86. Zolg, D. P. et al. Building ProteomeTools based on a complete synthetic human proteome. Nat. Methods 14, 259–262 (2017).

  87. Marx, H. et al. A large synthetic peptide and phosphopeptide reference library for mass spectrometry–based proteomics. Nat. Biotechnol. 31, 557–564 (2013).

    Article  PubMed  Google Scholar 

  88. Kryshtafovych, A., Schwede, T., Topf, M., Fidelis, K. & Moult, J. Critical assessment of methods of protein structure prediction (CASP)—round XIV. Proteins 89, 1607–1617 (2021).

  89. Mann, S. P., Treit, P. V., Geyer, P. E., Omenn, G. S. & Mann, M. Ethical principles, constraints, and opportunities in clinical proteomics. Mol. Cell. Proteomics 20, 100046 (2021).

    Article  PubMed  PubMed Central  Google Scholar 

  90. Bandeira, N., Deutsch, E. W., Kohlbacher, O., Martens, L. & Vizcaíno, J. A. Data management of sensitive human proteomics data: current practices, recommendations, and perspectives for the future. Mol. Cell. Proteomics 20, 100071 (2021).

    Article  PubMed  PubMed Central  Google Scholar 

  91. Cai, Z. et al. Federated deep learning enables cancer subtyping by proteomics. Cancer Discov 15, 1803–1818 (2025). This study applies federated deep learning to proteomics data for cancer subtyping, demonstrating how privacy-preserving AI frameworks can integrate decentralized datasets to enhance model generalization, data security and precision in AI-driven clinical proteomics.

    Article  PubMed  PubMed Central  Google Scholar 

  92. Vašíček, J. et al. ProHap enables human proteomic database generation accounting for population diversity. Nat. Methods 22, 273–277 (2025).

    Article  PubMed  Google Scholar 

  93. Qian, L. et al. Towards the construction of a virtual yeast. Nature 655, 59–70 (2026).

    Article  PubMed  Google Scholar 

Download references

Acknowledgements

T.G., Y.S., J.A., Z.L., R.S., L.Q., Y.C., Z.D., Y.D. and H.G. acknowledge the National Key R&D Program of China (grant no. 2022YFF0608403), the National Natural Science Foundation of China (Key Joint Research Program) (grant no. U24A20476), National Natural Science Foundation of China (Major Research Plan) (grant no. 92259201), National Natural Science Foundation of China (Young Scientist Fund) (grant no. 32401239), Zhejiang Provincial Natural Science Foundation of China (grant no. LQ24C050002) and National Key R&D Program of China (grant no. 2021YFA1301600). B.Z. acknowledges the Robert and Janice McNair Foundation. M.M. acknowledges the Max Planck Society for the Advancement of Science. W.B. acknowledges the Research Foundation–Flanders (G087625N and G0AHY25N). C.L. was supported by an Australian Research Council (ARC) Future Fellowship (FT240100798) and a National Health and Medical Research Council of Australia (NHMRC) Ideas Grant (2024/GNT2037597). C.C. and F.H. were supported by the National Key Research and Development Program of China (2024YFA1210400 and 2021YFA1301603) and the National Natural Science Foundation of China (32088101). J.A.V. and Y.P.-R. acknowledge BBSRC grants BB/X001911/1, BB/V018779/1, BB/Y513829/1, Wellcome grant 223745/Z/21/Z and EMBL core funding. S.H.P. was supported by NIGMS/National Institutes of Health award R01GM147653. H.H. acknowledges BBSRC grant BB/X002179/1. V.D. was supported by BMBF grant 161L0221. M.L. acknowledges the National Key R&D Program of China (no. 2022YFA1304603), the National Key Research and Development Program of China (2024YFA1306400) and Canadian NSERC grant OGP0046506. Y.W. was supported by direct national funding from the Chinese Ministry of Technology to Pengcheng Laboratory, Research and Development Program of Guangzhou Laboratory (SRPG22-001). G.S.O. acknowledges National Institutes of Health grants P30ES017885-11-S1 and U24CA271037. C.M.O. was supported by the Canada Research Chairs program (950-01-126) and the Canadian Institutes for Health Research Foundation Grant program (FDN-148408). E.W.D. acknowledges National Institutes of Health grants R01 GM087221 and R24 GM148372. L.C. was supported by the Natural Science Foundation of China (T2341007, T2350003, 12131020, 42450084, 42450135, 12326614, 12426310, 12301620 and 42450192), Shenzhen Medical Research Fund (E250200621,E250200620), Tianfu Jincheng Laboratory (TFJCPI20260001) and Zhejiang Province Vanguard Goose-Leading Initiative (2025C01114). We acknowledge the support of the π-Hub project.

Author information

Authors and Affiliations

  1. Affiliated Hangzhou First People’s Hospital, State Key Laboratory of Medical Proteomics, School of Medicine, School of Future Biomedicine, Westlake University, Hangzhou, China

    Yingying Sun, Jun A, Zhiwei Liu, Rui Sun, Liujia Qian, Yi Chen, Zhen Dong, Huanhuan Gao, Yamin Deng & Tiannan Guo

  2. Westlake Center for Intelligent Proteomics, Westlake Laboratory of Life Sciences and Biomedicine, Hangzhou, China

    Yingying Sun, Jun A, Zhiwei Liu, Rui Sun, Liujia Qian, Yi Chen, Zhen Dong & Tiannan Guo

  3. Biology Department, Brigham Young University, Provo, UT, USA

    Samuel H. Payne

  4. Department of Computer Science, University of Antwerp, Antwerp, Belgium

    Wout Bittremieux

  5. Department of Biochemistry, Charité Universitätsmedizin Berlin, Berlin, Germany

    Markus Ralser & Vadim Demichev

  6. Department of Biochemistry and Molecular Biology, Immunity and Cancer Programs, Biomedicine Discovery Institute, Monash University, Clayton, Victoria, Australia

    Chen Li & Jiangning Song

  7. Department of Medicine, School of Clinical Sciences, Monash University, Clayton, Victoria, Australia

    Chen Li

  8. European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge, UK

    Yasset Perez-Riverol, Juan Antonio Vizcaíno & Henning Hermjakob

  9. Harvard Medical School, Ludwig Center at Harvard, Boston, MA, USA

    Asif Khan

  10. Harvard Medical School, Broad Institute, Ludwig Center at Harvard, DF/HCC Cancer Center, Boston, MA, USA

    Chris Sander

  11. Department of Biology, Institute of Molecular Systems Biology, ETH Zürich, Zurich, Switzerland

    Ruedi Aebersold

  12. Bruker Ltd., Milton, Ontario, Canada

    Jonathan R. Krieger

  13. AI for Life Sciences Lab, Tencent, Shenzhen, China

    Jianhua Yao

  14. State Key Laboratory of Medical Proteomics, AI for Science Institute, Beijing, China

    Wen Han & Linfeng Zhang

  15. State Key Laboratory of Medical Proteomics, National Center for Protein Sciences (Beijing), Beijing, China

    Yunping Zhu, Cheng Chang, Ping Xu & Fuchu He

  16. Thermo Fisher Scientific GmbH, Bremen, Germany

    Yue Xuan

  17. Informatics and Predictive Sciences Research, Bristol Myers Squibb, Cambridge, MA, USA

    Benjamin Boyang Sun

  18. Department of Chemistry, Fudan University, Shanghai, China

    Liang Qiao

  19. Department of Computer Science, Luddy School of Informatics, Computing and Engineering, Indiana University, Bloomington, IN, USA

    Haixu Tang

  20. ProCan®, Children’s Medical Research Institute, Faculty of Medicine and Health, The University of Sydney, Westmead, New South Wales, Australia

    Qing Zhong

  21. Center for Computational Mass Spectrometry, Dept. Computer Science and Engineering, Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California, San Diego (UCSD), La Jolla, CA, USA

    Nuno Bandeira

  22. Peng Cheng Laboratory, Shenzhen, China

    Ming Li & Yu Wang

  23. Bioinformatics Solutions Inc, Waterloo, Ontario, Canada

    Ming Li

  24. University of Waterloo, Waterloo, Ontario, Canada

    Ming Li

  25. AI for Science Institute, Center for Machine Learning Research, School of Mathematical Sciences, Peking University, Beijing, P. R. China

    Weinan E

  26. Research Institute of Intelligent Complex Systems, Fudan University, Shanghai, China

    Siqi Sun

  27. School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China

    Yuedong Yang

  28. Departments of Computational Medicine & Bioinformatics, Internal Medicine, Human Genetics & Genomics, and Environmental Health, University of Michigan, Ann Arbor, MI, USA

    Gilbert S. Omenn

  29. Department of Materials Science and Engineering, Westlake University, Hangzhou, China

    Yue Zhang, Jiaxing Huang & Kaicheng Yu

  30. State Key Laboratory of Mathematical Science, Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China

    Yan Fu

  31. Deming Department of Medicine, School of Medicine, Tulane University, New Orleans, LA, USA

    Xiaowen Liu

  32. Department of Oral Biological and Medical Sciences, Centre for Blood Research, University of British Columbia, Vancouver, British Columbia, Canada

    Christopher M. Overall & Peter A. Bell

  33. Institute for Systems Biology (ISB), Seattle, WA, USA

    Eric W. Deutsch

  34. School of Mathematical Sciences and School of AI, Shanghai Jiao Tong University, Shanghai, China

    Luonan Chen

  35. Computational Systems Biochemistry Research Group, Max Planck Institute of Biochemistry, Martinsried, Germany

    Jürgen Cox

  36. International Academy of Phronesis Medicine (Guang Dong), Guangzhou International Bio Island, Guangzhou, China

    Fuchu He

  37. School of Big Data, Anhui University, Hefei, China

    Huilin Jin

  38. School of Biological Science and Medical Engineering & School of Engineering Medicine, Beihang University, Beijing, China

    Chao Liu

  39. Westlake University High-Performance Computing Center, Westlake University, Hangzhou, China

    Nan Li

  40. School of Computer Science and Engineering, Beihang University, Beijing, China

    Zhongzhi Luan

  41. School of Communication and Information Engineering, Institute of Smart City, Shanghai University, Shanghai, China

    Wanggen Wan

  42. Bristol Myers Squibb, Cambridge, MA, USA

    Tai Wang

  43. Eye Hospital and Institute for Advanced Study on Eye Health and Diseases, Institute for clinical Data Science, Wenzhou Medical University, Wenzhou, China

    Kang Zhang

  44. Department of Computer Science, Sichuan University, Chengdu, China

    Le Zhang

  45. Department of Proteomics and Signal Transduction, Max Planck Institute of Biochemistry, Martinsried, Germany

    Matthias Mann

  46. Lester and Sue Smith Breast Center, Baylor College of Medicine, Houston, TX, USA

    Bing Zhang

Authors

  1. Yingying Sun
  2. Jun A
  3. Zhiwei Liu
  4. Rui Sun
  5. Liujia Qian
  6. Samuel H. Payne
  7. Wout Bittremieux
  8. Markus Ralser
  9. Chen Li
  10. Yi Chen
  11. Zhen Dong
  12. Yasset Perez-Riverol
  13. Asif Khan
  14. Chris Sander
  15. Ruedi Aebersold
  16. Juan Antonio Vizcaíno
  17. Jonathan R. Krieger
  18. Jianhua Yao
  19. Wen Han
  20. Linfeng Zhang
  21. Yunping Zhu
  22. Yue Xuan
  23. Benjamin Boyang Sun
  24. Liang Qiao
  25. Henning Hermjakob
  26. Haixu Tang
  27. Huanhuan Gao
  28. Yamin Deng
  29. Qing Zhong
  30. Cheng Chang
  31. Nuno Bandeira
  32. Ming Li
  33. Weinan E
  34. Siqi Sun
  35. Yuedong Yang
  36. Gilbert S. Omenn
  37. Yue Zhang
  38. Ping Xu
  39. Yan Fu
  40. Xiaowen Liu
  41. Christopher M. Overall
  42. Yu Wang
  43. Eric W. Deutsch
  44. Luonan Chen
  45. Jürgen Cox
  46. Vadim Demichev
  47. Fuchu He
  48. Jiaxing Huang
  49. Huilin Jin
  50. Chao Liu
  51. Nan Li
  52. Zhongzhi Luan
  53. Jiangning Song
  54. Kaicheng Yu
  55. Wanggen Wan
  56. Tai Wang
  57. Kang Zhang
  58. Le Zhang
  59. Peter A. Bell
  60. Matthias Mann
  61. Bing Zhang
  62. Tiannan Guo

Contributions

T.G., B.Z. and M.M. supervised the work. Y.S., J.A., Z.L., R.S., L.Q., S.H.P., Z.D., Y.C. and Y.D. wrote the original draft and created the figures. Y.S., J.A., Z.L., R.S., L.Q., S.H.P., Z.D., Y.C., Y.D., W.B., M.R., C.L., Y.P.-R., A.K., C.S., R.A., J.A.V., J.R.K., J.Y., H.W., L.Z., Y.Z., Y.X., B.B.S., L.Q., H.H., H.T., H.G., Q.Z., C.C., N.B., M.L., W.E., S.S., Y.Y., G.S.O., Y.Z., P.X., Y.F., X.L. and C.M.O., Y.W., E.W.D., L.C., J.C., V. D., F. H., J.H., H.J., C.L., N.L., Z.L., J.S., K.Y., W.W., T.W., K.Z., L.Z. and P.A.B. reviewed the draft and provided feedback on the manuscript.

Corresponding authors

Correspondence to Matthias Mann, Bing Zhang or Tiannan Guo.

Ethics declarations

Competing interests

T.G. is the founder of Westlake Omics Inc. B.Z. received research funding from AstraZeneca and consulting fees from Inotiv. M.M. is an indirect investor in Evosep. X.L. has a project contract with Bioinformatics Solutions Inc. J.R.K. is an employee of Bruker Ltd. Milton, Canada. Y.X. is an employee of Thermo Fisher Scientific. B.B.S. is currently a full-time employee of Bristol Myers Squibb. V.D. holds shares in Aptila Biotech. M.R. is the founder and shareholder of Eliptica Ltd. C.S. is a scientific advisor for Cytoreason Ltd. The remaining authors declare no competing interests.

Peer review

Peer review information

Nature Methods thanks the anonymous reviewers for their contribution to the peer review of this work. Primary Handling Editor: Arunima Singh, in collaboration with the Nature Methods team.

Additional information

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary information

Supplementary Table 1 (download XLSX )

The table summarizes software and tools that apply traditional machine learning or deep learning models across the six key areas of MS-based proteomics discussed in the Perspective. It includes their corresponding subsections, modules, publication year, software/outcome name, model, model type, specific tasks, last corresponding author, first author and journal article link.

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Sun, Y., A, J., Liu, Z. et al. AI proteomics: from protein identification to virtual cells. Nat Methods (2026). https://doi.org/10.1038/s41592-026-03085-y

Download citation

  • Received:

  • Accepted:

  • Published:

  • Version of record:

  • DOI: https://doi.org/10.1038/s41592-026-03085-y

colind88

Back To Top