CA2303758C

Improved methods of identifying peptides and proteins by mass spectrometry

Abstract

A method of identifying a protein, polypeptide or peptide by means of mass spectrometry and especially by tandem mass spectrometry is disclosed. The method preferably models the fragmentation of a peptide or protein in a tandem mass spectrometer to facilitate comparison with an experimentally determined spectrum. A fragmentation model is used which takes account of all possible fragmentation pathways which a particular sequence of amino acids may undergo. A peptide or protein may be identified by comparing an experimentally determined mass spectrum with spectra predicted using such a fragmentation model from a library of known peptides or proteins. Alternatively, a de novo method of determining the amino acid sequence of an unknown peptide using such a fragmentation model may be used.

CA2303758C, drawing sheet 1
Sheet 1 of 3

Term

Term ended

Expired 6 April 2020, 6.5 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

31 claims: 4 independent, 27 dependent

  1. 1
    CA 02303758 2003-05-28 20208-1777 CLAIMS:1. A method of identifying a most probable amino acid sequence(s) which would account for a fragmentation mass spectrum of a protein or peptide, said method comprising the 5 steps of: a) producing the fragmentation mass spectrum (D) from said protein or peptide;b) providing a plurality of trial amino acid sequences;c) for each of a plurality of fragmentation routes (f) which 10 together represent the possible ways that a trial amino acid sequence (S) will fragment, calculating: i) the probability P(D given f) of the fragmentation mass spectrum (D), assuming a particular fragmentation route ii) the probability P(t given S) cf having the fragmentation route (f) from each of said trial sequences (S) ;d) calculating a likelihood factor P(D given S) that each trial sequence (S) is the true amino acid sequence of said peptide or protein by summing probabilistically over said plurality of 2C fragmentation routes, such that: P(D given S) = f P(D given f).P(f given S) e) selecting one or more of said trial sequences (S) which have 25' the highest likelihood factors calculated in step d) as being the most probable amino acid sequence(sj of said protein or peptide . CA 02303758 2003-05-28 20208-1777
  2. 2
    The method as claimed Ln claim 1 wherein said plurality of fragmentation routes represents all the possible ways that said trial sequence might fragment.
  3. 5
    The method as claimed in any one of claims 1 to 4 10 wherein at least one cf said fragmentation routes is described by using one or more Markov chains to represent a series of ions generated by successive losses of amino acid residues, wherein the probability of an ion m said series being observed is influenced by the probability of a preceding ion in the 15 series being observed.
  4. 28
    Apparatus for identifying the most probable amino acid sequence is) in an unknown protein or peptide, the apparatus comprising:a mass spectrometer for generating a fragmentation mass spectrum (D) from, a protein or peptide, and data processinç} means comprising (a) means for calculating, for each of a plurality of fragmentation routes (f;which together represent the possible ways that a trial amino acid sequence (S) wilL fragment, (i) the probability P (D given f) of the fragmentation mass spectrum (D), assuming a particular fragmentation route (f), and (ii) the probability P (f given S) of having the fragmentation route (f) from each of said trial sequences (S);(b) means for calculating a likelihood factor P (D given S) that each trial sequence (3) is the true amino acid sequence of said protein or peptide by summing probabilistically over said plurality of fragmentation routes, such that: P(D given S') = , P(D given f ).P(f given S) , whereby the most probable ammo acid sequence (s) of said protein or peptide corresponds to said trial amino acid sequence(s) (S) having the highest likelihood factor(s).