Structure-Aware Inference from Incomplete Mass Spectra: From De Novo Sequencing to PTM Discovery

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

University of Waterloo

Abstract

Machine-learning systems are often developed under a convenient assumption: the set of possible answers is known before inference begins. Molecular discovery violates this assumption. A tandem mass spectrum is an incomplete and noisy record of a molecule, while the databases and molecular vocabularies used to train computational models contain only previously described molecules. A model forced to choose among known answers may therefore confuse missing evidence with a negative result or assign an unexpected molecule to the nearest familiar category. This dissertation argues that computational inference should preserve both what a measurement supports and what it leaves unresolved until additional evidence justifies a more specific molecular claim. Rather than treating a database mismatch as the end of analysis, the methods developed here retain informative relationships among measured fragments, use them to formulate explicit hypotheses, and keep the evidence that generates a hypothesis distinct from the evidence used to accept it. Experiments on peptide and glycopeptide inference demonstrate why this distinction matters. When peptide fragments are missing, postponing uncertain local sequence decisions prevents one ambiguous region from corrupting the remainder of the sequence. For branched glycans, relationships between measured fragments and partial molecular compositions recover evidence that cannot be represented by a single linear ordering. The accompanying evaluation work further shows that apparent accuracy and estimated error rates depend on the data distribution, scoring method, and construction of negative controls. Expanding the space of possible molecular explanations must therefore be accompanied by evidence appropriate to the claim being made. These studies culminate in RNovA, which extends de novo sequencing from recognizing predefined modifications to discovering post-translational modifications beyond the training data and reference vocabulary. Without retraining for each new modification, RNovA can detect unexpected mass changes, reconstruct the corresponding modified peptides, and turn computational outputs into experimentally testable molecular hypotheses. Controlled experiments recovered modification types absent from training, while selected discoveries in biological samples were supported using synthetic reference peptides. Together, these results show that artificial intelligence need not treat an existing database as the boundary of molecular discovery. By preserving structural evidence in the measurement and making unresolved information explicit, it can propose and test molecular explanations that were not enumerated in advance.

Description

Keywords

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By