π-HuB Research Highlights | π-HelixNovo2 Delivers Accessible, High-Performance De Novo Peptide Sequencing for Global Proteomics

14

07

2026

16

07

2026

The π-HuB Secretariat is glad to share a practical computational proteomics advance jointly completed by Prof. Cheng Chang’s team at the National Center for Protein Sciences, Beijing and Prof. Yu Wang’s group from Pengcheng Laboratory. Their method paper π-HelixNovo2: Making Accurate Online De Novo Peptide Sequencing Available to All was published online in Genomics, Proteomics & Bioinformatics (GPB) in June 2026. This work addresses three common limitations of existing de novo peptide sequencing algorithms and builds an AI model paired with a cloud computing platform to lower operation barriers for proteomics researchers.

 

Core Research Challenges

De novo peptide sequencing is a standard method for identifying unannotated novel peptides and proteins from tandem mass spectrometry data. Transformer-based tools including Casanovo and the earlier π-HelixNovo have improved analysis efficiency, yet three practical issues remain unsolved.

1. Incomplete MS2 spectral signals lead to lost b/y fragment ions, and single-direction decoding accumulates sequence errors without effective correction.

2. Generative AI models produce uncertain prediction results, lacking simple and stable filtering rules to distinguish reliable peptide-spectrum matches from false identifications.

3. Most mainstream tools require coding skills and dedicated local computing equipment, which are hard to access for many wet-lab researchers.

They updated the π-HelixNovo framework and developed the π-HelixNovo2 system together with supporting filtering and cloud deployment workflows to solve these problems.

 

Core Technical Innovations of π-HelixNovo2

The model adopts two improved strategies and a standardized screening workflow to raise identification accuracy:

1. Complementary spectrum encoding: Based on the mass matching rule of b/y ion pairs, complementary ion signals are calculated to supplement missing fragment peaks. Separate encoders process original and complementary spectra to extract complete mass spectrum features and reduce interference from incomplete fragmentation.

2. Shared-encoder bidirectional ensemble decoding: Two sets of decoders generate peptide sequences from N-terminal and C-terminal directions respectively, and multiple decoder groups work together to stabilize prediction results. Comparative tests show complementary spectrum brings an average 11.8% rise in peptide recall, bidirectional decoding adds a 15.4% improvement, and combining both delivers a total 25.7% recall increase compared with baseline Transformer models.

3. Dual-constraint peptide filtering pipeline: Two fixed thresholds are set to screen credible results: peptide confidence score reaches 1, and mass deviation between precursor and predicted peptide is less than 0.1 Da. A hybrid analysis workflow combining database search and de novo sequencing is also proposed to detect novel peptides for both large and small mass spectrometry datasets.

 

Comprehensive Performance Validation & Practical Application

The team compared π-HelixNovo2 with widely used algorithms on datasets covering multiple species, antibodies, multi-enzyme digests and non-enzymatic samples as well as gut metaproteome data.

1. On the nine-species cross-validation dataset, π-HelixNovo2 gains 9.2%–19.4% higher peptide recall than other tools. Tests on three independent datasets achieve recall growth of 11.9%, 13.7% and 17.3% respectively.

2. Models trained on larger datasets show better generalization; the version trained on 30 million MassIVE-KB PSM data outperforms π-HelixNovo and Casanovo by 7.2%–19.3% on antibody datasets. Targeted fine-tuning improves analysis of non-tryptic peptides obviously: peptide recall of multi-enzyme samples rises by 10.5%–58.2%, while recall for non-enzymatic peptides increases over 65%.

3. In gut metaproteome analysis, π-HelixNovo2 reanalyzed 9.5 million unassigned MS2 spectra and detected around twice the number of high-confidence PSMs and microbial peptides as π-HelixNovo. After screening sequences originating from gut microbes, thousands of unique microbial peptides were obtained, which helps improve taxonomic differentiation for research on host-microbe interactions.

 

Open Visual Online Computing Platform for Global Scientists

To cut computing resource costs for researchers, the team deployed π-HelixNovo2 on China Computing NET and Pengcheng Cloud Brain II supercomputer and built a visual public platform (https://openi.pcl.ac.cn/OpenI/pi-HelixNovo-NPU). The platform provides free GPU/NPU resources and pre-installed operating environments. Users can upload data, run model training, fine-tuning and inference through simple mouse operations. Multi-node parallel computing shortens processing time for large-scale mass spectrometry datasets. Trial accounts, source codes, model weights and English operation tutorials are all open access for the whole proteomics community.

 

Relevance for π-HuB

This open AI analysis pipeline offers usable computational support for π-HuB’s human proteome mapping work. It can coordinate with existing separation technologies, unify consistent analysis standards across different labs and facilitate processing of clinical and multi-omics cohort data, while providing reference algorithm designs for follow-up tool development under the initiative.

 

We congratulate Professor Cheng Chang, Professor Yu Wang, first author Dr. Tingpeng Yang and all participating team members on this practical methodological study. The π-HuB Research Highlights channel keeps receiving high-impact research sharing from Council Members and global partners to promote academic communication. We look forward to exchanges on proteomics computing tools at the π-HuB Summit and Council Meeting in September 2026.

 

Reference

Yang T, Ling T, Sun B, et al. π-HelixNovo2: Making Accurate Online De Novo Peptide Sequencing Available to All. Genomics, Proteomics & Bioinformatics. Published online June, 2026. doi:10.1093/gpbjnl/qzag049/8716247


You are applying

Laboratory staff of biological mass spectrometry platform

Full Name*

E-mail*

Resume uploading

The size cannot exceed 5M and supports word, pdf, and html

Other

Our page uses cookies

We use cookies to personalize and enhance your browsing experience on our website. By clicking "Accept All", you agree to use cookies. You can read our Cookie Policy for more information.