Structure-Aware Inference from Incomplete Mass Spectra: From De Novo Sequencing to PTM Discovery

dc.contributor.authorMao, Zeping
dc.date.accessioned2026-08-17T14:15:40Z
dc.date.issued2026-08-17
dc.date.submitted2026-08-12
dc.description.abstractMachine-learning systems are often developed under a convenient assumption: the set of possible answers is known before inference begins. Molecular discovery violates this assumption. A tandem mass spectrum is an incomplete and noisy record of a molecule, while the databases and molecular vocabularies used to train computational models contain only previously described molecules. A model forced to choose among known answers may therefore confuse missing evidence with a negative result or assign an unexpected molecule to the nearest familiar category. This dissertation argues that computational inference should preserve both what a measurement supports and what it leaves unresolved until additional evidence justifies a more specific molecular claim. Rather than treating a database mismatch as the end of analysis, the methods developed here retain informative relationships among measured fragments, use them to formulate explicit hypotheses, and keep the evidence that generates a hypothesis distinct from the evidence used to accept it. Experiments on peptide and glycopeptide inference demonstrate why this distinction matters. When peptide fragments are missing, postponing uncertain local sequence decisions prevents one ambiguous region from corrupting the remainder of the sequence. For branched glycans, relationships between measured fragments and partial molecular compositions recover evidence that cannot be represented by a single linear ordering. The accompanying evaluation work further shows that apparent accuracy and estimated error rates depend on the data distribution, scoring method, and construction of negative controls. Expanding the space of possible molecular explanations must therefore be accompanied by evidence appropriate to the claim being made. These studies culminate in RNovA, which extends de novo sequencing from recognizing predefined modifications to discovering post-translational modifications beyond the training data and reference vocabulary. Without retraining for each new modification, RNovA can detect unexpected mass changes, reconstruct the corresponding modified peptides, and turn computational outputs into experimentally testable molecular hypotheses. Controlled experiments recovered modification types absent from training, while selected discoveries in biological samples were supported using synthetic reference peptides. Together, these results show that artificial intelligence need not treat an existing database as the boundary of molecular discovery. By preserving structural evidence in the measurement and making unresolved information explicit, it can propose and test molecular explanations that were not enumerated in advance.
dc.identifier.urihttps://hdl.handle.net/10012/23979
dc.language.isoen
dc.pendingfalse
dc.publisherUniversity of Waterlooen
dc.titleStructure-Aware Inference from Incomplete Mass Spectra: From De Novo Sequencing to PTM Discovery
dc.typeDoctoral Thesis
uws-etd.degreeDoctor of Philosophy
uws-etd.degree.departmentDavid R. Cheriton School of Computer Science
uws-etd.degree.disciplineComputer Science
uws-etd.degree.grantorUniversity of Waterlooen
uws-etd.embargo.terms0
uws.contributor.advisorLi, Ming
uws.contributor.affiliation1Faculty of Mathematics
uws.peerReviewStatusUnrevieweden
uws.published.cityWaterlooen
uws.published.countryCanadaen
uws.published.provinceOntarioen
uws.scholarLevelGraduateen
uws.typeOfResourceTexten

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Mao_Zeping.pdf
Size:
10.42 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
6.4 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections