Menu

Home

home (1)
glycolab Dinko Soic

NovoGlyco: Comprehensive Glycoproteomics Platform

Overview

NovoGlyco is a fully untargeted glycoproteomics platform for identifying and characterizing prokaryotic protein glycosylation directly from shotgun proteomics data. NovoGlyco integrates de novo discovery of oxonium ions, sequence tag database matching, and mass offset binning to extract glycan mass, composition and attachment types. Results are provided with an interactive visualization framework that supports exploration of glycan features and glycoproteins.

NovoGlyco works in tandem with Oxonium Browser for discovering diagnostic sugar oxonium ions in MS/MS data.

The platform is designed for prokaryotic systems where glycan compositions are often novel or poorly characterised, traditional database search approaches fail due to unknown modifications.

Available as a standalone executable, a Docker version, and Python source code at SourceForge: https://sourceforge.net/projects/novoglycox/files/

How It Works

┌───────────────┐     ┌───────────────┐     ┌───────────────┐     ┌───────────────┐
│  Input Files  │───▶│ SAGE Database │────▶│   Spectral    │───▶│  Oxonium Ion  │
│  Processing   │     │    Search     │     │  Processing   │     │  Annotation   │
└───────────────┘     └───────────────┘     └───────────────┘     └───────────────┘
                                                                          │
                                                                          ▼
┌───────────────┐     ┌───────────────┐     ┌───────────────┐     ┌───────────────┐
│  Interactive  │◀────│ Glycopeptide  │◀───│ Sequence Tag  │◀───│ DirectTag De  │
│   Dashboard   │     │  Validation   │     │   Matching    │     │Novo Sequencing│
└───────────────┘     └───────────────┘     └───────────────┘     └───────────────┘

The pipeline:

  1. SAGE database search identifies unmodified peptides and creates a focused protein database
  2. Spectral processing decharges all unmatched MS2 spectra and calculates precursor offsets
  3. Oxonium ion annotation flags which diagnostic sugar fragment ions are present per spectrum
  4. De novo sequencing (DirecTag) generates sequence tags from all unmatched spectra
  5. Tag matching identifies peptides via dictionary-based lookup, optionally validated by Y0 ion presence
  6. Mass delta and offset binning reveals glycan masses and monosaccharide compositions
  7. Interactive dashboard presents glycoprotein candidates with oxonium evidence for exploration

Key Features

Glycan Database-Independent Search
All unmatched spectra undergo tag matching regardless of oxonium ion content, without requiring a predefined glycan database. This catches glycopeptides with novel or unexpected glycan compositions, including those where oxonium ions are absent or below threshold. Y0 ion validation is enabled by default for increased confidence but can be disabled to maximise sensitivity (see Analysis Parameters). The oxonium-filtered view is available as a confidence layer on top of the untargeted results.

Glycan Composition Analysis
Three complementary mass metrics — mass deltas, precursor offsets, and peptide offsets — provide independent evidence for glycan mass and monosaccharide composition. Precursor offset analysis works even when peptide identification fails, enabling glycan exploration across all oxonium-positive spectra. See Key Metrics for details.

Protein-Centric Dashboard
An interactive Plotly Dash interface presents results as ranked glycoprotein candidates with collapsible peptide details. Each protein shows PSM count, unique peptides, median tag count, and oxonium evidence as coloured dots. Oxonium filter checkboxes allow progressive refinement from untargeted results to high-confidence glycopeptide candidates. See Interactive Dashboard for a full guide.

Oxonium Ion Co-occurrence
A heatmap reveals which monosaccharides co-occur on the same spectra, supporting glycan structure reconstruction directly from shotgun data.

Getting Started

For the Python and Docker versions, installation instructions, input file requirements, Docker commands, and parameter configuration are described in the README files included in the respective ZIP packages. No installation is required for the standalone executable.

Wiki Pages

Page Description
Key Metrics Explanation of mass delta, precursor offsets, peptide offsets, tag counts, and other analytical metrics
Interactive Dashboard Guide to the dashboard visualisations, glycoprotein candidate table, oxonium filter, and Excel export
Analysis Parameters Complete parameter reference with defaults, descriptions, and recommended combinations
System Architecture Technical documentation of the pipeline architecture, module descriptions, and data flow

Citation

If you use this software in your research, please cite:

Šoić D. and Pabst M. NovoGlyco: mapping protein glycosylation in prokaryotes. bioRxiv. 2026.

Contacts

Dinko Šoić (soic@imsb.biol.ethz.ch)
Martin Pabst (m.pabst@tudelft.nl)


Auth0 Logo