Method and apparatus for identifying proteins in mixtures
25 claims: 7 independent, 18 dependent
- 1混合物中のタンパク質を同定する方法であって、 混合物を消化してペプチド混合物を得る工程と、 消化された混合物をLC/MSシステムに適用する工程と、 LC/MSシステムの質量分析計部分において低エネルギーモードを適用する工程と、 LC/MSシステムの質量分析計部分において高エネルギーモードを適用してペプチド混合物に由来するペプチド前駆体をフラグメント化する工程と、 低エネルギーモード中にペプチド前駆体から取得されたペプチド前駆体のデータから、質量情報および保持時間情報を取得する工程と、 高エネルギーモード中にフラグメント化されたペプチド前駆体から取得されたフラグメントデータから、質量情報および保持時間情報を取得する工程と、 タンパク質のデータベースから標的ペプチドを選択する工程と、 前記データベースから選択された 標的ペプチドと関連する標的前駆体に関連する質量を決定する工程と、 標的前駆体に相当するYイオンおよびBイオン各々に関連する質量を決定する工程と、 標的前駆体に関連する質量、およびYイオンおよびBイオンに関連する質量を、ペプチドの前駆体データおよびフラグメントデータに関連する質量情報と比較する工程と、 比較から得られる質量の一致を検出する工程と、 質量の一致の保持時間を決定する工程と、 検出された質量の一致および検出された質量の一致の保持時間に基づいて、標的ペプチドがペプチド混合物中に存在するか否かを同定する工程とを含む方法。
- 2クロマトグラフピーク幅の間に、高エネルギーモードおよび低エネルギーモード各々が複数回適用されるように、高エネルギーモードおよび低エネルギーモードの適用を十分な頻度で交番するプロトコルに従って、高エネルギーモードと低エネルギーモードとを切り換える工程をさらに含む、請求項1に記載の方法。
- 3検出クロマトグラムを生成する工程をさらに含む、請求項1に記載の方法。
- 4検出クロマトグラムを生成するために、質量の一致に相当する検出ガウスのピークを個々の保持時間値域に加える工程をさらに含む、請求項3に記載の方法。
- 5検出閾値を決定する工程と、 特定の保持インターバル中の質量の一致の計数が、検出閾値を超えるか否かに基づいて、ペプチドが存在するかどうかを決定する工程とをさらに含む、請求項1に記載の方法。
- 6検出閾値に関連するノイズフロアを推定する工程と、 計数およびノイズフロアの推定に基づいて、ペプチド同定の有意性を分析する工程とをさらに含む、請求項5に記載の方法。
- 7低エネルギーモードが、LC/MSシステムの質量分析計部分の衝突セルにおいて低電圧を印加することによって適用され、高エネルギーモードが、LC/MSシステムの質量分析計部分の衝突セルにおいて高電圧を印加することによって適用される、請求項1に記載の方法。
- 8データベースとして、ペプチド混合物中で見付けられる可能性があるタンパク質を有する特化されたデータベースを選択する工程をさらに含む、請求項1に記載の方法。
- 9ピーク特性を用いて、特定のクロマトグラフピークが標的前駆体に関連しているかどうかを決定する工程をさらに含む、請求項1に記載の方法。
- 10ピーク特性がピーク形状である、請求項9に記載の方法。
- 11ピーク形状が、ピークの頂点、上昇勾配変曲点、および下降勾配変曲点の分析を含む、請求項10に記載の方法。
- 12混合物中のタンパク質を同定するシステムであって、 混合物を消化してペプチド混合物を得る手段と、 混合物を適用して混合物を分離する液体クロマトグラフと、 液体クロマトグラフによって出力された分離された混合物中の質量を測定するための質量分析計とを備え、該質量分析計が、ペプチドをフラグメント化する衝突セルと、フラグメント化されたペプチドおよびフラグメント化されていないペプチドに関連する質量情報を決定する質量分析装置とを備え、前記システムがさらに、 標的ペプチドが選択されるタンパク質のデータベースと、 データベースおよび質量分析計出力に結合され、分析コンピュータで実行されるソフトウェアを有する分析コンピュータとを備え、前記ソフトウェアは、分析コンピュータに、 前記データベースから選択された 標的ペプチドと関連する標的前駆体に関連する質量を決定させ、標的前駆体に対応するYイオンおよびBイオン各々に関する質量を決定させ、標的前駆体ならびにYイオンおよびBイオンに関連する質量を、フラグメント化されたおよびフラグメント化されていないペプチドに関連する質量情報と比較させ、比較から得られる質量の一致を検出させ、質量の一致の保持時間を決定させ、検出された質量の一致および検出された質量の一致の保持時間に基づいて、標的ペプチドが混合物中にあるか否かを同定させる、システム。
- 13高電圧を衝突セルに印加してペプチドをフラグメント化し、低電圧を印加してフラグメン化されていないペプチドについて質量分析器からデータを提供する、請求項12に記載のシステム。
- 14衝突セルは、クロマトグラフピーク幅の間に、高エネルギーモードおよび低エネルギーモード各々が複数回適用されるように、高エネルギーモードおよび低エネルギーモードの適用を十分な頻度で交番するプロトコルに従って、高エネルギーモードと低エネルギーモードとを切り換えるように構成された、請求項12に記載のシステム。
- 15コンピュータソフトウェアが、コンピュータに、さらに検出クロマトグラムを生成させる、請求項12に記載のシステム。
- 16コンピュータソフトウェアが、コンピュータに、さらに質量の一致に相当する検出ガウスのピークを個々の保持時間値域に加えて、検出クロマトグラムを生成させる、請求項15に記載のシステム。
- 17コンピュータソフトウェアが、コンピュータにさらに、 検出閾値を決定させ、 特定の保持インターバル中の質量の一致の計数が、検出閾値を超えるか否かに基づいて、ペプチドが存在するかどうかを決定させる、請求項12に記載のシステム。
- 18コンピュータソフトウェアが、コンピュータにさらに、 検出閾値に関連するノイズフロアを推定させ、 計数およびノイズフロアの推定に基づいて、ペプチド同定の有意性を分析させる、請求項17に記載のシステム。
- 19低エネルギーモードが、LC/MSシステムの質量分析計部分の衝突セルにおいて低電圧を印加することによって適用され、高エネルギーモードが、LC/MSシステムの質量分析計部分の衝突セルにおいて高電圧を印加することによって適用される、請求項1 4 に記載の システム 。
- 20データベースが、混合物中で見付けられる可能性があるタンパク質を有する特化されたデータベースである、請求項12に記載のシステム。
- 21コンピュータソフトウェアが、コンピュータに、ピーク特性を用いて特定のクロマトグラフピークが標的前駆体に関連しているかどうかを決定させる、請求項12に記載のシステム。
- 22コンピュータソフトウェアが、コンピュータに、ピーク形状を用いて特定のクロマトグラフピークが標的前駆体に関連しているかどうかを決定させる、請求項21に記載のシステム。
- 23コンピュータソフトウェアが、コンピュータに、ピークの頂点、上昇勾配変曲点、および下降勾配変曲点を分析させる、請求項22に記載のシステム。
- 24混合物中の前駆体を同定する方法であって、 液体クロマトグラフ部分および質量分析計部分を有するLC/MSシステムに混合物を適用する工程と、 LC/MSシステムの質量分析計部分において低エネルギーモードを適用する工程と、 LC/MSシステムの質量分析計部分において高エネルギーモードを適用して混合物に由来する前駆体をフラグメント化する工程と、 低エネルギーモード中に前駆体から取得された前駆体のデータから質量情報および保持時間情報を取得する工程と、 高エネルギーモード中にフラグメント化された前駆体から取得されたフラグメントデータから質量情報および保持時間情報を取得する工程と、 データベースから標的前駆体を選択する工程と、 前記データベースから選択された 標的前駆体に関連する質量を決定する工程と、 前記データベースから選択された 標的前駆体に相当するフラグメントイオンに関連する質量を決定する工程と、 前記データベースから選択された 標的前駆体に関連する質量およびフラグメントイオンに関連する質量を、前駆体データおよびフラグメントデータに関連する質量情報と比較する工程と、 比較から得られる質量の一致を検出する工程と、 質量の一致の保持時間を決定する工程と、 検出された質量の一致および検出された質量の一致の保持時間に基づいて、標的前駆体が混合物中に存在するか否かを同定する工程とを含む方法。
- 25混合物中の前駆体を同定する方法であって、 液体クロマトグラフ部分および質量分析計部分を有するLC/MSシステムに混合物を適用する工程と、 LC/MSシステムの質量分析計部分において 前駆体およびフラグメントの混合物を収集する 固定エネルギーモードを適用する工程と、 固定エネルギーモード中に取得された前駆体およびフラグメントのデータから、質量情報および保持時間情報を取得する工程と、 データベースから標的前駆体を選択する工程と、 前記データベースから選択された 標的前駆体に関連する質量を決定する工程と、 前記データベースから選択された 標的前駆体に相当するフラグメントイオンに関連する質量を決定する工程と、 前記データベースから選択された 標的前駆体に関連する質量およびフラグメントイオンに関連する質量を、前駆体データおよびフラグメントデータに関連する質量情報と比較する工程と、 比較から得られる質量の一致を検出する工程と、 質量の一致の保持時間を決定する工程と、 検出された質量の一致および検出された質量の一致の保持時間に基づいて、標的前駆体が混合物中に存在するか否かを同定する工程とを含む方法。
Independent claims25
181 paragraphs, as filed
This application claims the benefit of US Patent Provisional Application No. 60/572532, filed May 20, 2004, which is incorporated herein by reference in its entirety.
This application is co-pending application No. PCT / US05 with Agent Document No. WAA-393, entitled "Systems and Methods for Grouping Precursor and Fragment Ions Using Selected Ion Chromatograms". Related to / issue.
The present invention generally relates to proteomics. More specifically, the present invention uses liquid chromatography combined with mass spectrometry to identify and quantify peptides in proteins and complex mixtures, as well as to produce precursors and fragment ions in mass spectrometers. It relates to identifying and quantifying molecules in a mixture. The invention also relates to tracking the retention time of peptides in a mixed complex using liquid chromatography in combination with mass spectrometry. More importantly, the present invention provides a method for peptide identification that does not require the presence of a mass of precursor ions. This allows the method to include both chemically modified and post-translationally modified peptides, allelic differences, point mutant-containing peptides, as well as sequences stored in the referenced database. It is possible to identify any modification of.
Proteomics generally relates to studies involving complex mixtures of proteins. The field of proteomics includes studying and cataloging proteins in biological systems. Proteomics studies typically focus on protein identification, determination of relative abundance differences between different conditions, or both. Identifying and quantifying proteins in complex biological samples is a fundamental problem in proteomics.
Liquid chromatography (LC / MS) combined with mass spectrometry has become a fundamental tool in proteomics research. Separation of untreated proteins or their proteolyzed peptide products by liquid chromatography (LC) and then analysis by mass spectrometry (MS) underlies many common proteomics methodologies. Methods of measuring changes in protein expression levels are of great interest as they can form the basis for finding biomarkers and clinical diagnostics.
Traditional proteomics studies typically digest the protein of interest first to produce a particular set of proteolytic peptides, rather than studying the untreated protein directly. The resulting peptide is then characterized during proteomics analysis. A common enzyme used for such digestion is trypsin. In trypsin digestion, the proteins present in the complex mixture are cleaved to produce peptides as determined by the cleavage specificity of the proteolytic enzyme. From the observed peptide identity and concentration, algorithms known in the art can estimate the identity and concentration of the parent protein.
In LC / MS analysis, this peptide digestion is separated and analyzed by online mass spectrometry (MS) following online liquid chromatography (LC) separation. Ideally, the mass of a single peptide measured with sufficient accuracy is sufficient to uniquely identify the peptide. However, in practice, the mass accuracy achieved is typically greater than or equal to about 10 ppm. In general, such mass accuracy is not sufficient to uniquely identify a peptide based solely on mass measurements. For example, with a mass accuracy of 10 ppm, about 10 peptide sequences are identified by a typical database search. Lower search restrictions on mass accuracy, chemical or post-translational modifications, H<sub>2</sub>O or NH<sub>3</sub>This number of sequences will increase significantly, considering the loss of, point mutations, etc. Sequence repositories typically contain translated DNA sequences annotated to substrates known by homology. Therefore, if a peptide sequence is modified by either deletion or substitution, it would be erroneous to tentatively identify the peptide by precursor mass alone.
In addition, the two peptides may have the same amino acid composition but different sequences. Mass accuracy alone is not sufficient to distinguish peptides that differ in sequence rather than composition. A fragmentation method for breaking a peptide into fragment ions is known. These fragments may correspond to the subsequences of the original peptide, but other types of fragment ions may be observed. The fragment mass of this data can be used to confirm or estimate the precursor sequence.
In the case of peptide precursors, subsequences can result from fragmentation at the single peptide bond of the precursor. Such fragmentation results in two subsequences. When a fragment containing the peptide C-terminus is ionized, it is called a Y ion, and when a fragment containing a peptide N-terminus is ionized, it is called a B ion.
Known protein identification methods search databases using accurate mass retention time (AMRT) data for precursors and fragments obtained from LC / MS tests. For example, methods of obtaining such data are described in US Pat. No. 6,717,130 to Bateman (Bateman), which is incorporated herein by reference in its entirety. Bateman can obtain such data using high-energy and low-energy switching protocols that are applied as part of the LC / MS analysis of a single infusion of peptide mixture. In such data, the low energy spectrum contains ions mainly from unfragmented precursors and the high energy spectrum mainly contains ions from fragmented precursors.
To identify the presence of proteins in such data, AMRT (which experimentally describes those ions from peptides or fragments) is selected from the low energy data. If trypsin is used for digestion, this AMRT is assumed to be a trypsin precursor. Using this AMRT data, a known method searches a database of peptide masses to find trypsin peptides whose masses are within the mass search window or threshold.
A hit is considered if the theoretical peptide mass from the database is within the mass search window for the mass of a precursor measured in the data. That is, the precursor of the data was hit by the peptide of the database, or otherwise, the peptide of the database was hit by the precursor of the data.
This search provides a hit list of potentially matching peptides from the database. These potentially matching database peptides may or may not be weighted by statistical factors. As a possible result of such a search, potentially matching database peptides are not identified, or one potentially matching database peptide is identified or likely to match 2 One or more database peptides will be identified. The higher the resolution of the MS, the better instrument calibration is assumed, and the smaller the ppm threshold, the less identification errors.
If there is one or more hits to a theoretical peptide in the database, conventional searches use data from high-energy AMRTs to validate possible matching database peptides. High-energy AMRTs are first searched to separate high-energy AMRTs that occur at the same retention time as the low-energy AMRTs being verified. Typically, the isolated high-energy AMRT is an AMRT whose retention time is substantially the same as the low-energy AMRT under verification.
For each peptide in the hitlist database, the algorithm determines the masses of all Y and B ions that can be obtained through collision-induced dissociation of the precursor. In pairs, this separated high-energy AMRT data is retrieved for each of these Y and B ions. Peptide sequences that have the highest number of hits or meet other criteria are returned as the exact hits, i.e. identification of the target precursor. This result can be saved and displayed.
This step can be repeated for each low energy AMRT in the digestion mixture. Further analysis can be performed on the results, including storing the results, displaying the results, quantifying the results, and combining the results with the results of other injections.
During the search, multiple charge states and multiple isotopes can be searched. You may also search for ions, ie AMRTs with reduced charge. In addition, better confidence can be obtained by applying experimentally generated confidence rules to help identify valid hits and using a higher number of high energy hits.
<p> In summary, given the set of data obtained by the LC / MS system, known protein identification methods search a database of theoretical protein sequences to identify proteins in that data. That is, a known protein identification method is started with the data and searches the database. In contrast, the invention described below begins with a database and retrieves data.</p>
<p> In contrast to conventional protein identification methods, embodiments of the invention begin with a theoretical peptide sequence (not always, but generally obtained from a protein database) and correspond to the theoretical peptide sequence. Search the data for evidence of precursors and fragment ions. If a sufficient number of such masses are found in the data with a common retention time, the peptide sequence is identified in the data. If this method finds one or more peptide sequences associated with a given protein in the data, the protein is considered to be identified in the sample.</p><p> An embodiment of the present invention uses a preselected database to retrieve data acquired using an LC / MS system. For example, in one embodiment of the invention, the eluent output by the liquid chromatograph (LC) is introduced into the mass spectrometer (MS) via the ESI interface. The first quadrupole (Ql) of the MS simply acts as an ion guide. An AC voltage is applied to the collision cell. The spectrum is collected in all precursors and all fragments thereof in an alternating fashion as described in Bateman.</p><p> More specifically, embodiments of the present invention collect spectra that alternate uniformly over time between the low-energy and high-energy spectra. There is no selection of MS spectra applied prior to high energy fragmentation. The high energy mode spectrum contains fragment ions of all precursor ions. Due to the high duty cycle of data acquisition in this alternating mode, chromatograph profiles of all detected precursors and fragments are preserved. This data acquisition mode allows the determination or measurement of retention time, as well as the m / z and intensity determination or measurement of all ions found in low and high energy modes.</p><p> The low energy mode corresponds to the conventional LC / MS acquisition. The high energy mode is referred to herein as an alternative ascending energy mode. High energy or rising energy mode LC / MS<sup>E</sup>Corresponds to acquisition. The low energy mode includes a spectrum of primary precursor ions. The high energy mode includes a spectrum of primary fragment ions.</p><p> the term As used herein, the following terms have the specified meaning. Protein: A specific primary sequence of amino acids grouped together as a single polypeptide. Peptide: A specific sequence of amino acids, grouped as a single polypeptide, contained within the primary sequence of a protein. Trypsin peptide: A peptide produced from a protein sequence obtained from enzymatic cleavage of a protein by trypsin. In the following description, the digested peptide will be referred to as a trypsin peptide for convenience. However, it should be understood that embodiments of the present invention apply to other methods for peptide digestion. Precursor Peptide: A trypsin peptide (or other protein cleavage product) produced directly using a protein cleavage protocol. Precursor peptides from the sample are chromatographed and sent to a mass spectrometer. In mass spectrometers, an ion source ionizes these precursor peptides to produce a positively charged proteinized precursor morphology. The mass of such a positively charged protein compounded form can be referred to as the precursor mwHPlus or MH +. In the following, the term "precursor mass" is used, which generally refers to the proteinified mwHPlus or MH + mass of an ionized peptide precursor. Fragments: Multiple types of fragments, MS<sup>E</sup>Can occur in the spectrum. In the case of tryptic peptide precursors, the fragment may comprise a polypeptide ion produced from collision fragmentation of the untreated peptide precursor, the amino acid primary sequence of which was contained within the original precursor peptide. Y and B ions are examples of such peptide fragments. Fragments of trypsin peptide are ammonium ion, phosphate ion (PO)<sub>3</sub>), Such as a functional group, a mass label cleaved from a particular molecule or class of molecules, or water from a precursor (H)<sub>2</sub>O) molecule or ammonia (NH)<sub>3</sub>) The molecule "neutral loss" may also be included.</p><p> Y and B ions: When a peptide is fragmented at a peptide bond and the charge is retained in an N-terminal fragment, the fragment ions are called B ions. When the charge is retained in the C-terminal fragment, the fragment ion is called the Y ion. A more descriptive list of possible fragments and their names can be found in Roepstorff and Fohlman, Biomed Mass Spectrom, 1984, 11 (11): 601, and Johnson et al., Anal. Chem 1987, 59 (21): 2621: 2625. It is provided and is incorporated herein by reference.</p><p> Chromatograph profile: In LC / MS analysis, the intensity vs. time of the chromatograph peak at a single mass corresponding to a single precursor or fragment ion. The mass chromatogram may include a chromatograph profile of one or more such ions.</p><p> Vertex retention time or chromatograph retention time: The point of the chromatograph profile when an entity reaches its maximum intensity during LC / MS analysis.</p><p> Ions: Each peptide appears as a population of ions due to the natural abundance of component isotopes. Ions have a retention time and m / z value. A mass spectrometer (MS) detects only ions. This LC / MS method produces a variety of observed measurements for any of the detected ions. This includes charge-to-mass ratio (m / z) m, retention time, and signal strength of the ions.</p><p> mwHPlus: Neutral monoisotopic mass of peptide + weight of one proton 1.070825amu.</p><p> AMRT: Accurate mass retention time. AMRT is an empirical description of peptides in terms of their mass, retention time, and total strength. When the peptide elutes from the chromatograph column, the peptide elutes over a particular retention time interval and reaches its maximum signal in a single retention time (vertex retention time). After ionization and (possibly) fragmentation, the peptide appears as a set of related ions. The different ions in this set correspond to different isotopic compositions and charges of a common peptide. Each ion in this set of related ions produces a single vertex retention time and peak shape. Since these ions arise from a common peptide, the apex retention time and peak shape of each ion are the same within some measurement tolerance. MS acquisition of each peptide produces multiple ion detections for all isotope and charge states, all sharing the same vertex retention time and peak shape within some measurement tolerance.</p><p> In LC / MS separation, a single peptide (precursor or fragment) produces multiple ion detections and appears as a cluster of ions in multiple charged states. Deconvolution of these ion detections from such clusters indicates that, at a particular retention time, there is a single body of unique monoisotopic mass of the measured signal strength of the charge state, while giving rise to AMRT. Suggest.</p><p> It is not possible to directly infer from AMRT whether the sequence is a precursor, a fragment, or a chemically modified peptide, not to mention what it is. Molecules other than peptides can be described using AMRT.</p><p> Protein Database: In an embodiment of the invention, the user selects or supplies a database of proteins. Alternatively, a default database or other predetermined database may be used. Each protein is described by its primary sequence of amino acids. It is the user's responsibility to choose which database (or database subset) to compare with the data. The user may select a database that is intended to closely match the protein under study. For example, an E. coli database could be compared with data obtained from E. coli cell lysates. Similarly, the human serum database will be compared with the data obtained from human serum. The user may select a subset database. The user may select a superset database, such as all proteins listed in SwissProt. The user may select a database containing simulated proteins, described by a random sequence of amino acids. Such a random database will be used in controlled trials to evaluate and calibrate protein identification systems and search for algorithms. The user may use a database that combines both naturally occurring sequences or artificial sequences.</p><p> From the protein database, the software can deduce from each sequence the sequences and masses of trypsin precursor ions, Y and B ions, and other possible fragment ions that will result from their precursors.</p>
FIG. 1 is a schematic representation of a system for identifying and quantifying proteins in a biological complex mixture according to an embodiment of the invention. Sample 102 is injected into the liquid chromatograph 104 via the injector 106. The pump 108 pumps the sample through the column 110 and decomposes the mixture into its components according to the retention time through the column.
The output from the column is input to the mass spectrometer 112 for analysis. First, the sample is desolvated and ionized by the desolvation / ionization device 114. Desolvation may be any method for desolvation, including, for example, heaters, gases, heaters in combination with gases, or other desolvation methods. The ionization may be any ionization, including, for example, electrospray ionization (ESI), atmospheric pressure chemical ionization (APCI), matrix-assisted laser desorption (MALDI), or other ionization methods. The ions obtained from the ionization are supplied to the collision cell 118 by the voltage gradient applied to the ion guide 116. Collision cells 118 can be used to pass ions (low energy) or fragment ions (high energy). For example, as described in Bateman, an AC voltage is applied to the collision cell 118 to cause fragmentation. Spectrum is collected for low-energy (non-collision) precursors and high-energy (results of collisions) fragments.
The output of the collision cell 118 is input to the mass spectrometer 120. The mass spectrometer 120 may be any mass spectrometer including a quadrupole type, a time-of-flight (TOF) type, an ion trap type, a magnetic field type mass spectrometer, and a combination thereof. The detector 122 detects ions coming out of the mass spectrometer 122. The detector 122 can be integrated with the mass spectrometer 120. For example, in the case of a TOF mass spectrometer, the detector 122 can be a microchannel plate detector that counts the intensity of ions, i.e. counts the number of ions that collide with the detector.
The storage medium 124 provides a permanent recording device for storing ion counts for analysis. For example, the storage medium 124 may be an internal or external computer disk. The analysis computer 126 analyzes the stored data. Data can be analyzed in real time without being stored in the storage medium 124. In real-time analysis, the detector 122 sends the data to be analyzed directly to computer 126 without first storing it in a permanent recording device.
Collision cell 118 performs fragmentation of precursor ions. Fragmentation can be used to determine the primary sequence of a peptide and then identify its original protein.
Collision cell 118 contains a gas such as helium, argon, nitrogen, air, or methane. When a charged peptide interacts with a gas atom, the resulting collision can fragment the peptide by breaking it into one or more characteristic bonds. The most common fragments obtained are described as Y or B ions. Such fragmentation involves a low voltage state (low energy <5V) to obtain the MS spectrum of the peptide precursor and a high voltage state (high energy> 15V) to obtain the MS spectrum of the collision-induced precursor fragment. ), Which can be achieved as an online fragmentation by switching the voltage in the collision cell. High and low voltages are referred to as high and low energies because they use high or low voltage to impart kinetic energy to the ions, respectively.
Various protocols can be used to determine when and how to switch the voltage to obtain such MS / MS acquisition. For example, conventional methods trigger voltage in either target mode or data dependency mode (data dependency analysis, DDA). These methods also include gas phase separation (or preselection) of the bound, target precursor. Acquire a low energy spectrum and verify it in real time with software. When the desired mass reaches a specified intensity value in the low energy spectrum, the voltage in the collision cell is switched to a high energy state. A high energy spectrum is then acquired for the preselected precursor ions. These spectra include fragments of precursor peptides found at low energy. After a sufficiently high energy spectrum has been collected, the data acquisition returns to low energy in order to continue searching for the precursor mass at an intensity suitable for high energy collision analysis.
Although conventional switching methods can be used, embodiments of the present invention preferably use a novel fragmentation protocol in which the voltage is switched in a simple alternating cycle. This switching is done frequently enough so that the plurality of high energy spectra and the plurality of low energy spectra are contained in a single chromatograph peak. Unlike traditional switching protocols, this cycle is independent of the content of the data.
In summary, each sample 102 is injected into the LC / MS system. The LC / MS system produces two sets of spectra: a set of low energy spectra and a set of high energy spectra. The set of low energy spectra mainly contains ions associated with the precursor. The set of high energy spectra mainly contains ions associated with the fragment. These spectra are stored in the storage medium 124. After data acquisition, these spectra can be extracted from the storage medium, displayed, and processed by the post-acquisition algorithm on the analysis computer 126.
The data obtained by the high-low protocol makes it possible to accurately determine the retention time, mass-to-charge ratio, and intensity of all ions and AMRTs collected in low-energy and high-energy modes. In general, different ions are found in two different modes, and then the spectra obtained in each mode are analyzed separately to determine the ion retention time, mass-to-charge ratio, and intensity observed in each mode. It is determined.
Ions from a common precursor, such as those found in one or both modes, share the same retention time and peak shape. The high-low protocol allows for meaningful comparison of ion retention time and peak shape within or between modes. This comparison can then be used to group the ions found in both the low and high energy spectra by their common retention time and peak shape. Figures 2, 8, and 9 and the discussion below show how low-energy and high-energy spectra can be used to find ions with a common retention time and peak shape.
FIG. 2 shows the time during which the spectrum was obtained during elution of the peaks obtained by applying the low energy mode and the high energy mode to the police box according to one embodiment of the present invention. Figure 2 shows that the chromatographic profile and precursor-related ion retention times can be reconstructed into both high-energy and low-energy spectral data.
Peak 202 represents the elution profile of a single precursor chromatograph. The horizontal axis is the elution time. The vertical axis is arbitrary and represents the concentration and chromatograph profile of the precursor that change with time when eluted from the chromatograph column.
The plots 204a (low energy) and 204b (high energy) in FIG. 2 represent the same chromatograph peak 202, with the horizontal axis representing time and the vertical axis representing ion intensity.
Eluted molecules passed through a mass spectrometer produce ions in both low and high energy modes. The ions produced in the low energy mode are primarily the ions of precursor ions in different isotopic and charged states that may be possible. In proteomics studies, precursor ions are peptides produced from enzymatic digestion (typically trypsin digestion) of untreated proteins. In high energy mode, the ions are primarily charged states of fragment ions of different isotopes and their precursors. The high energy mode can also be called an ascending energy mode.
In the peak 202 plot, the alternating density graphs represent the time during which the spectra were collected using low and high energy voltages during the elution of the chromatograph peaks shown. This bar is evenly replaced every hour. Plot 204a exemplifies the time when a low energy voltage was applied to the collision cell, resulting in a low energy spectrum. Plot 204b shows the time when a high energy voltage was applied to the collision cell, resulting in a high energy spectrum. As shown in 204a and 204b in the figure, chromatograph peaks are sampled multiple times in high energy and low energy modes. From these plurality of samples, the exact retention times of all the ions associated with the peak and observed in the high and low energy spectra can be estimated. These exact retention times are obtained by interpolating the intensities sampled by the individual spectra.
A certain holding time t<sub>r</sub>In the case of molecules that elute at, all relevant ions will be observed to elute at exactly the same retention time within some measurement accuracy. This phenomenon is shown in FIGS. 8A to 8B and 9A to 9C. Ions associated with a common precursor have the same chromatographic peak shape. 8A-8B show exemplary spectra extracted at 72.5 minutes containing isotopes with Z = + 2 ions. The six ions 802a-802f (appearing with masses 740amu-742.5amu) shown in FIGS. 8A-8B are isotopes of trypsin peptide from cellotransferrin precursor protein in human serum. The sequence of this peptide is MYLGYEYVTAIR. The interval between the mass-to-charge ratio (m / z) of this ion is 0.5amu, indicating that the charge of the ion is Z = 2. At Z = 2,<sup>12</sup>The m / z value of the C monoisotopic ion is 739.96amu.
FIG. 8B shows a series of mass chromatograms corresponding to each of the six ions shown in FIG. 8A. As shown by vertical line 804, the retention time of each ion has the same value of 72.53 minutes, indicating that the chromatograph vertices of each ion go into the same single scan. Vertical line 804 with arrows shows that the retention times of each of the six chromatograph peaks 812a-812f are the same, and the ions found in the top plot are associated with a common precursor. It is confirmed that it is.
FIG. 9A is a plot of an exemplary high energy spectrum showing a large number of spectral peaks 910-916. FIG. 9C is a series of plots showing chromatograph profiles corresponding to some of the peaks appearing in FIG. 9A. Spectral peaks 910, 911, 912, and 913 have chromatographic peaks that correspond to those depicted in FIG. 9C, such as peaks 920, 921, 922, and 923. The chromatograph peaks of these four spectral peaks have substantially the same peak shape and retention time, consistent with the hypothesis that they are from the same peptide. Spectral peaks 914, 915, and 916 correspond to chromatograph peaks 924, 925, and 926. The retention times of peaks 924, 925, and 926 are substantially the same as their peak shape, consistent with the hypothesis that spectral peaks 914, 915, and 916 are from the same peptide.
FIG. 9B depicts only spectral peaks having the same retention time as the retention time of peak 910. Therefore, only peaks 910, 911, 912, and 913 are redrawn as in the vertical lines 930, 931, 932, and 933 in Figure 9B. Since the retention times are out of sync, all other peaks in Figure 9A are excluded and not shown in Figure 9B. This example demonstrates that chromatographic information can be used to select related spectrum peaks or exclude unrelated peaks. As described below, the PDS algorithm of the embodiments of the present invention relies on matching retention times between the precursor and its fragments as seen in both low and high energy spectra.
3A to 3C are flowcharts of the peptide identification method according to the embodiment of the present invention. For ease of reference, the method of the invention is referred to as a peptide data retrieval (PDS) algorithm. Given the peptide sequence that corresponds to a protein in an accurate mass database, embodiments of the invention search the data to determine if the peptide sequence is present. No previous knowledge of peptide retention time is required. However, knowing the range of time the peptide elutes can help reduce the number of data AMRTs retrieved. This reduction speeds up calculations and reduces false positive detection rates. This also allows lower detection thresholds (described below) to be applied to the detection chromatogram.
Ions or AMRTs must be obtained from both low-energy and high-energy data before retrieving the data. FIG. 3A is a flowchart of a method for acquiring ions or AMRT according to an embodiment of the present invention. The method in Figure 3A is applied to low-energy and high-energy data to obtain the required ions or AMRT.
In step 302, the collected spectrum is read from the hard drive. In step 304, ions are detected from the spectrum. For example, by a peak finding algorithm applied to spectral chromatograms and mass chromatograms, such as described in Bateman (US Pat. No. 6,717,130), which is incorporated herein by reference, or "liquid chromatography / mass analysis data." PCT / US 05/04180 ("International Patent Application No. 4180", co-pending International Patent Application No. 4180, filed February 11, 2005, entitled "Instruments and Methods for Identifying Peaks and Generating Spectrums and Chromatograms" ), The ions can be detected by the two-dimensional convolution method. In one embodiment of the invention, the properties of the measured ion are its retention time, mass-to-charge ratio (m / z), and intensity. Step 304 stores a list of these ionic properties in a table. From this table, a list of these ions and their properties is entered in method 306. Method 306 determines the exact mass retention time (AMRT), writes the AMRT parameters to the table, and stores these parameters.
In LC / MS tests, peptides appear as a set of ions, with each ion corresponding to a peptide in a different isotope and charge state. AMRT is a set of ions produced by a peptide. The properties of AMRT are determined from its set of ions containing AMRT.
One AMRT corresponds to a set of ions from the ion list obtained in step 304. Therefore, method 306 decomposes the ion list into a set of ions. Here each set is AMRT. The characteristics of ARMT are determined from such a set. AMRT is described by four parameters: retention time, mwHPlus, intensity, and partial charge state. AMRT consists of a set of at least two or more ions. Two or more ions are required to establish that charge, and therefore to establish ARMT's mwHPlus.
AMRT retention time and mwHPlus are the minimum mass, that is, the retention time and mwHPlus of monoisotopic ions in the set. The intensity of AMRT is the sum of the intensities of the ions in the set. It is also possible to derive a figure of merit known as the partial charge state for each AMRT. This partial charge state is the sum of the charges of each ion weighted by the partial intensity of that ion relative to the AMRT intensity. Ions that are not bundled into a set are treated as a single ion described by their retention time, m / z, and intensity. A single ion can effectively be considered AMRT if the charge for that ion is assumed or assigned by convention.
The peptide mwHPlus is neutral and is the monoisotopic mass of the peptide plus the mass of one proton, [M + H], or mwHPlus, or MH.<sup>+</sup>Is called. When all the atoms are in their lowest mass, the most abundant isotope state, the monoisotopic mass M is the mass of the peptide.
PDS and EDA algorithms (described below), such as those applied to peptides, take AMRT retention time, mwHPlus, and intensity as inputs. This list of single ions is stored with the AMRT list. A single ion can effectively be considered AMRT if the charge for that ion is assumed or assigned by convention. Therefore, a single ion can optionally be included with the ARMT list as an input to the PDS and EDA algorithms. Alternatively, the PDS and EDA algorithms (as described below) can be applied only to that ion, as is obtained in step 304, skipping step 306.
The algorithm performed in step 306 utilizes the known properties of the peptide mass spectrometric properties. Peptides are known to appear in the mass spectrum as a set of ions at different values m / z. Examples of such spectra are incorporated herein by reference by Barbara Seliger Larsen (edit), Charles N. McEwen (edit), Marcel Dekker, Mass Spectrometry of Biological Materials, 2nd Edition (March 1998). 1st), pp. 34-46.
Peptide ions can have a possible mass-to-charge ratio of m / z = [M + Z × H + N × 1.003] / Z = [M + N × 1.00335] / Z + H. In the formula, M is the monoisotopic mass of the neutral peptide, H = 1.00728amu, the mass of the proton due to the charge of the peptide, Z is the charge of the peptide, and N is the number of isotopes of the peptide. Is (integer) and 1.00335amu is<sup>13</sup>With C isotope<sup>12</sup>The mass difference with the C isotope. This mass difference is an approximation of the actual mass difference that occurs between isotopes of the same peptide.
The value for N = 0 corresponds to this monoisotopic state. The monoisotopic mass of the peptide for N = 0 and Z = 1 is then [M + H].
FIG. 8A shows part of a mass spectrum containing a set of 6 ions 802a-802f associated with a single peptide. Such sets are commonly referred to as clusters, or ion clusters. The m / z spacing between the ions in this cluster is 0.5 amu, which is evidence that the ions are in different isotopic states of peptides with a charge of Z = 2. The smallest mass ion is the monoisotope, which appears at m / z = 739.96. The mwHPlus of this peptide is estimated to be 739.96 × 2-1.00739 = 147.91amu. If this peptide appears at charge Z = 1, its monoisotope appears at m / z of 1478.91amu, and if its peptide appears at Z = 3, its monoisotope is (1478.91 + 2) / 3 Appears at = 493.64amu. The mass spectrum of a peptide can be represented as consisting of clusters of one or more ions. Each cluster corresponds to an ion of the same charge. Different clusters correspond to ions in different charged states.
In the LC / MS data generated here, AMRT appears as a set of ions (where each ion is represented by retention time, m / z, and intensity). From the ion list obtained in step 304, it is easy to estimate the set of ions (where each set corresponds to AMRT and each such AMRT is presumed to correspond to a peptide). .. Since each AMRT is assumed to be derived from a peptide, rules can be used to reach such a set. For example, ions from peptides (as Bateman describes) should occur with a common retention time, and the mass-to-charge relationship between such sets of ions must conform to the above rules. Method 306 then collects a set of ions whose properties meet such a rule. It is presumed that each of such sets corresponds to AMRT and each of such AMRTs corresponds to peptides.
The method at 306 is then applied to the ion list obtained from the low energy spectrum to obtain a low energy AMRT. It is also possible to obtain a high energy AMRT by separately applying the method to the high energy ion list obtained from the high energy spectrum.
In summary, step 306 determines that AMRT is a set of ions in which each ion in the set is presumed to originate from a single common peptide. The peptide can be a precursor or a fragment. Step 306 utilizes the known properties of the peptide spectrum as described and described above, as well as the common retention time of such ions, to estimate the set of ions corresponding to AMRT. Step 306 then calculates the AMRT parameters (holding time, mwHPlus, intensity, and partial charge state) from each ion set, stores the parameters, and records the ions containing each ion set.
To estimate their charge states Z and mwHPlus, a known algorithm deconvolves the peptide spectrum in which the ions were found. Examples of such algorithms are Karl R. Clauser, Peter Baker, and Alma L. Burlingame, "Role of Accurate Mass Measurement (+/- 10ppm) in Protein Identification Strategies Employing MS or MS / MS and Database Searching," Anal. .Chem. 1999, No. 71, pp. 2871-2882. Another such algorithm is Zhongqi Zhang and Alan G. Marshall's "A Universal Algorithm for Fast and Automated Charge State Deconvolution of Electrospray Mass-to-Charge Ratio." Spectra , J. Am Soc. Mass Spectrom. 1998, No. 9, pp. 224-233. Each of these is incorporated herein by reference.
However, these known algorithms work on a single spectrum only. Thus, such an algorithm can determine the charge state and mwHPlus of each peptide observed in the spectrum, but not the exact retention time for such peptides.
The method used in step 306 to obtain AMRT from the ion list is novel and is described in FIGS. 17A-17C. Figure 17A summarizes the method. Step 1704 takes the ion list 1702 and determines the set of ion clusters. Each ion cluster contains ions with the same charge Z and retention time. The ion of this cluster with the lowest mass is the monoisotopic ion of the peptide. A list of this cluster is stored in the list at 1706 and this list is entered in step 1708.
Step 1708 validates this list to determine which clusters have the same retention time and mwHPlus but different charge states. If several clusters with the same retention time and mwHPlus as well as different charge states appear, 1708 estimates that these clusters should be from the same peptide. The clusters are grouped into a single set, which combined set is the AMRT for a single peptide. If a cluster with a unique retention time and mwHPlus emerges, it is presumed that the corresponding peptide produced the only cluster, and 1708 determines that one cluster is AMRT. Step 1710 stores the ions that were not grouped into AMRT and clusters. It is this combined list that is input to the PDS and EDA algorithms.
FIG. 17B shows how step 1704 identifies an ion cluster from the ion list 1702. A nested iterative loop reduces the two search parameters Zm and Nm. The initial values of these search parameters are Zmax and Nmax. In each path, Zm is the charge of the cluster and Nm is the minimum number of ions required to be present in the cluster. For low energy ions, the initial parameters are Zmax = 6 and Nmax = 8. For high energy ions, these initial parameters are Zmax = 3 and Nmax = 8.
Step 1736 finds all pairs of ions with the same retention time and separated at m / z by 1.00335 / Zm across all ions in the list. This value 1.00335amu is<sup>13</sup>With C isotope<sup>12</sup>The mass difference with the C isotope. This mass difference is an approximation of the actual mass difference that occurs between isotopes of the same peptide. Using a mass threshold of 20 ppm (as described below), this single approximation is sufficient to determine if a pair of ions are isotopes of a common peptide.
Step 1738 clusters the ion pairs. Thus, if ion 7 is paired with ion 10 and ion 10 is paired with ion 15, ions 7, 10, and 15 form a cluster. If ion 15 is paired with another ion, the cluster is expanded by that one additional ion. Ions are considered to be paired only if they are not labeled. Initially, the ions are unlabeled and all ions are considered. As described below, the ions are labeled in the next step. Step 1738 determines all clusters that may meet the pairing requirements.
The retention time requirement applied in step 1736 is determined by the retention time window and the m / z requirement applied in step 1736 is determined by the ppm window. The retention window is 20% of the chromatographic peak width (FWHM) and +/- 0.1 minutes for the chromatographic peak width of 0.5 minutes (FWHM). The ppm window at TOF with a resolution of 15000 is +/- 20ppm. That is, the ions are paired only if their retention time difference is within this window and their mass difference from the m / z model above is within that ppm window.
In step 1738, the set of ions is recorded as a cluster and labeled if two additional conditions are met. The number of ions in the cluster must be greater than or equal to Nm and the intensity ratio of the ions N = 1 and N = 0 must be within the range of values expected for such ions. Let r be the intensity ratio of N = 1 ion to N = 0 monoisotopic ion. The ion intensity distribution of the cluster is well known and is described in the references cited above. In the method described herein, the nominal intensity ratio r is estimated by r0 = (mwHPlus / 20) × 0.0107. In the formula (mwHPlus / 20) is an approximate value of the number of carbon atoms of the peptide, 0.0107 is<sup>12</sup>C atom pair<sup>13</sup>Approximate abundance of C atoms. The acceptable range for ions with an intensity ratio of N = 1 to N = 0 is 40% or r0 × 1.4 and r0 / 1.4. Therefore, it is necessary that r> r0 / 1.4 and r <r0 × 1.4.
If these two rules are not met, steps 1740 and 1742 do not label the ions and consider them in subsequent iterations. These rules are applied to take into account and remove accidental pairs of ions from unrelated peptides.
Step 1744 gets the cluster parameters for the clusters acquired in step 1738 and received in steps 1740 and 1742. The cluster parameters are retention time, mwHPlus, intensity, and charge. The retention time and mwHPlus for a cluster is the minimum mass retention time and mwHPlus for that cluster. The strength of this cluster is the sum of the strengths of the ions in the cluster. The charge of the cluster is a Zm parameter. If ion clusters are received, step 1744 labels the ions so that they are no longer considered during the next iteration. Step 1746 stores the cluster parameters containing the ions that form the cluster for the received cluster.
The next iteration reduces the charge parameter Zm. Therefore, in the second iteration, clusters with charges Zm = Zmax-1 and ions greater than or equal to Nm = Nmax are found. Continue the iteration and decrease Zm until Zm = 1. When Zm = 1, all charged state clusters containing ions above Nm = Nmax are identified. After reaching Zm = 1, the next iteration resets Zm = Zmax and decrements Nmax by 1 so that Nm = Nmax-1. This nested iteration proceeds until Nm = 2. Therefore, the iteration progresses from the maximum value to the minimum value for Nm of the outer loop and from the maximum value to the minimum value for Zm of the inner loop.
Step 1750 stores all clusters (cluster parameters and associated ions) found in the original ion list, and all ions not found to be present in the cluster.
The operation of step 1708 is shown in FIG. 17C. 1750 looping over the list of ion clusters. The incremental variable nc refers to the number of clusters and is initialized to 1. Step 1776 finds all clusters with the same retention time and mwHPlus as cluster nc. If step 1778 determines that no other such cluster exists, step 1780 records that the cluster nc is AMRT and stores its parameters in step 1786. Therefore, for a given retention time, if there is only one cluster with a given mwHPlus value, that cluster is considered to be the AMRT of the peptide. That is, there are peptides that appear as a single cluster. The AMRT is added to the AMRT list. AMRT parameters are the same as cluster parameters.
If step 1778 determines that one or more clusters have the same retention time and mwHPlus as cluster nc, step 1782 records that this set of clusters is AMRT. Step 1784 combines these clusters into a single AMRT, obtains their parameters, and 1786 accumulates the results. That is, if there are multiple clusters with the same mwHPlus, it is inferred that the peptide is present in the data with multiple ions appearing in these different charge and isotopic states. The AMRT parameter is the retention time of the strongest cluster, the mwHPlus of the strongest cluster, the intensity is the sum of the intensities of all clusters, and the partial charge state is the charge of the cluster weighted by the partial intensity of each cluster. Is the total of.
The retention time requirement applied in step 1778 is determined by the retention time window and the m / z requirement applied in step 1778 is determined by the ppm window. These parameters are acquired and applied in the same manner as in the case of ion pair determination above.
When all clusters are looped over, the loop ends and all results are stored. The final result is an AMRT list that includes all AMRT parameters and ions associated with those AMTSs, as well as ions that were not part of the cluster. It is this final list that is input to the PDS and EDA algorithms.
Other inputs to the PDS and EDA algorithms are sequences of target peptide precursors and their fragment sequences. FIG. 3B is a flow chart of a method of selecting a target precursor peptide using a selected database of protein sequences. In step 310, a suitable database is selected. The database preferably contains protein sequences that correspond to all proteins present or likely to be present in the sample. In the database, each protein is described by its primary sequence of amino acids. From such sequences, it is possible to predict peptides resulting from digestion protocols as well as other properties such as hydrophobicity and charge state. For example, trypsin digestion cleaves sequences at the known amino acids K and R. Based on these cleavage products, the corresponding masses obtained from the Y and B ion fragments as well as the collision fragmentation can be predicted. Therefore, this database provides a model of mass and other physical attributes that can occur in low-energy and high-energy spectra.
During protein identification, the AMRT or ions found in the data are compared to the masses contained in the database to provide reliable identification of the peptides present in the obtained LC / MS data. Ideally, all database peptides in the data are identified without error.
In step 312, in silico digestion is performed on one or more of the protein sequences in the database to produce precursor peptides in that database. In silico digestion is a synthetic digestion based on known digestive properties such as those mentioned above. In step 314, the exact mass of the precursor peptide is determined by looking at the amino acid sequences that make up the precursor peptide. The exact mass and sequence corresponding to the precursor peptide is stored for later use.
FIG. 3C is a flowchart of a method for identifying a peptide in a mixture according to an embodiment of the present invention. The method begins by selecting a precursor peptide (target precursor) from the database. Using the peptides in the selected database, the masses of the Y and B ion fragments corresponding to the selected peptides are determined or obtained from the database. In this way, the list of masses is put together. This list of masses includes the masses corresponding to each of the Y and B ions (possibly excluding the lowest masses Y and B ions), as well as the inherent mass associated with the unfragmented precursor itself. ..
Other precursor masses and fragment masses, such as those corresponding to chemically modified precursor peptides, may be considered. Examples of such modifications are by glycosylation or phosphorylation. Other fragment masses, such as fragmentation at peptide bonds other than Y or B bonds, may be considered.
Then, for each mass in this list, LC / MS and LC / MS<sup>E</sup>Search for both low-energy and high-energy AMRTs from the data. A matching mass, or hit, occurs when the mass from the database (precursor or fragment) is within the mass search window for the mass measured in the data (at low or high energy). All AMRTs hit are recorded along with their mass, retention time, and intensity. Matching masses, or hits, are accumulated for each of the multiple retention time ranges. Ranges with accumulation above the detection threshold are considered to be associated with the target precursor.
As described below, the method decisively uses the AMRT retention time match found at low and high energies. Also, as described below, the method can identify AMRTs in data that are related but not identical to peptides in the database. Such identification can be made when the mass associated with the peptide fragment from the database significantly overlaps the AMRT found in the data with substantially the same retention time.
With reference to FIG. 3C, in step 350, the precursor peptide (eg, trypsin peptide) is selected from in silico digested peptides. This peptide, alternatively referred to as the target sequence or target precursor, is described by its mass (mwHPlus) and its mass of Y and B ions in step 352 (mwHPlus). Any peptide in the database can be selected as the target precursor for step 350.
In step 352, a list of the target sequence and the exact masses of its Y and B ions is determined. In step 354, all in the data in the list of high and low energies with mwHPlus within the search tolerance (eg 20 ppm) using the target, precursor sequence, and masses of its Y and B ions. Search for AMRT. In step 356, AMRTs that match the masses in the mass list from the database to within the search tolerance are recorded, labeled, or identified. Search tolerances can be user-specified or can be automatically determined from the data by known statistical means. The automatic method for determining the mwHPlus tolerance is described below.
Ideally, the low energy spectrum contains only precursor ions. In fact, the precursor ions can be fragmented at the ion source so that the low energy spectrum can contain the fragment ions of the precursor. Such ions are called insource fragments and generally appear at attenuated intensities.
Ideally, the high energy spectrum contains only fragment ions. However, in practice, collision fragmentation of precursor ions may not be completed, so that the high energy spectrum may include precursor ions. Generally, such precursor ions appear in high energy mode intensities attenuated relative to their intensities in low energy mode.
Thus, the mass of the precursor or fragment can appear in either low-energy or high-energy data, or both. If the mass of the list appears in the low energy AMRT data, the PDS algorithm records, labels, or identifies it as appearing in that data for that mode. If a mass in the list appears in high-energy AMRT data, the PDS algorithm records, labels, or identifies it as appearing in that data for that mode. Therefore, the present invention utilizes all ions derived from a common precursor molecule, regardless of the mode in which such ions were generated or detected.
In step 358, a detection chromatogram is formed. Assuming that sequences with detectable levels of ions are present in the data, all such ions will be labeled during the search performed in step 356. However, many other ions that do not correspond to that sequence are also labeled. The detection chromatogram shows the number of ions (both low and high energy) labeled within the retention interval, and for each retention time, the vertical signal is the number of labels observed during the retention interval. Is. The effect of false positive labeling is to generate baseline noise.
According to one embodiment of the invention, the detection chromatogram is a simple histogram. The histogram is a series of ranges, the center of each range corresponds to the holding time, and the width of the range corresponds to the holding time interval. The histogram is formed by each hit with a simple one-up count, and the range corresponds to the specific retention time interval that included the retention time of the hit mass.
According to the second embodiment of the present invention, the detection chromatogram is derived using the accumulated Gaussian peaks. In a second embodiment of the invention, each hit AMRT is represented by a Gaussian peak in the detection chromatogram. FIG. 14 is a flowchart of a method for generating a detection chromatogram according to the second embodiment.
In step 1402, the detection peak width is established. The detection peak width is the width of the Gaussian peak added to the detection chromatogram for each hit. The width of the added Gaussian peak (hereinafter referred to as the detected Gaussian peak or the detected Gaussian peak) is set to the specified part of the FWHM of the chromatographic peak in the data. According to one embodiment of the invention, this portion is 10%. Therefore, if the FWHM peak width of a typical chromatograph peak is 0.5 minutes, then the Gaussian FWHM of detection is 0.05 minutes.
The time range of the detection chromatogram corresponds to the time range of separation. If the FWHM peak width of a typical chromatograph peak is 0.5 minutes, for example, the sample duration of the detection chromatogram is chosen to be about 1% of that width, ie 0.005 minutes. In step 1404, the detection chromatogram is initialized. The initial values for all points in the detection chromatogram are set to zero. The list of hit AMRTs (found in step 350) is traversed (looped over).
In step 1406, a detection Gaussian peak corresponding to a hit is added. This is done by analyzing all the hit low and high energy AMRTs. For each of the hit low-energy and high-energy AMRTs, a single detection gauss of unit height (with a width of detection peak width) is added to the detection chromatogram at the respective retention time of the AMRT or ion.
If two AMRTs with different masses elute at the same time, their detected gauss will reach a peak with a peak height of 2. If N AMRTs with different masses elute at the same time, their detection Gauss reaches a peak with a peak height N.
The width of the Gaussian detection corresponds to the standard error used to measure the retention time of the peak. A method for determining the standard error of retention time measurement is described below.
Returning to FIG. 3, step 362 identifies the maximum of the detection chromatogram. Determine the peptide detection threshold. The peptide detection threshold determines whether a peptide has been identified. The method by which the detection threshold can be determined is specified below. For example, the peptide detection threshold may be selected to be 4 AMRTs. Therefore, a peptide is considered identified if at least 4 AMRTs are present in the same retention window. AMRTs detected in both low and high energy spectra can help with this count.
That is, in step 362, if (A) the number of AMRT thresholds or more is found, (B) the relevant retention time of AMRT is within +/- 0.05 minutes, and (C) the mwHPlus value of AMRT. If all of the above are within 20 ppm of the molecular weight of the fragment and the molecular weight of the precursor in the selected peptide database, then the target peptide of the database is determined to be present in the data. In one embodiment of the invention, the detection chromatogram is constructed such that if (A) is true, the mass contributing to the maximum must also satisfy (B) and (C). The PDS algorithm for this point, if present, meets this condition and identifies AMRTs that suggest that the precursor selected by this (target precursor) is present.
Any maximal value above the detection threshold suggests whether the selected peptide is present in the data or some peptide is present in the data closely related to the selected peptide. When used in this context, the term "closely related" refers to database peptides and retention time t.<sub>r</sub>Means that there is a significant sequence match with the peptides found in the data at. Note that certain AMRTs with a precursor molecular weight (mwHPlus) can make such a detection as to whether they have been found. Therefore, one or more retention times can be found.
FIG. 15 is a flowchart of a method for identifying sequences that can be used in step 362 and their retention times according to an embodiment of the present invention. Upon completion of the method shown in FIG. 15, in addition to the target precursors found in the low energy LC / MS data, the sequences of the low energy associated (but not identical) sequences of the target precursors. Precursors found in LC / MS data are also identified.
In step 1502, the detection threshold is established. The detection threshold can be determined from all the maximums found in the detection chromatogram. Each maximum of the detection chromatogram has a value. From these values, the median is obtained. The detection threshold is typically set to about 4 times the median value. This detection threshold corresponds to the maximum number of fragments that can simply happen to fall within the detection peak width. Typical values for the detection threshold vary from 5 to 10 fragment ions per 0.05 minute detection peak width.
In step 1504, all peaks in the detection chromatogram above the threshold are recorded. If the target peptide (or some peptide with a sequence associated with the target peak) is not present in the data, the peak will not be detected in the detection chromatogram. On the other hand, if a target peptide and / or a peptide having a sequence associated with the target is present in sufficient concentration, there may be one or more maxima that exceed its detection threshold.
In step 1506, the retention time of the detection peak above the threshold is acquired as the retention time of the peptide. When a target peptide (or sequence-related target peptide) is detected, its retention time value is t.<sub>d</sub>Is. The height of the detection chromatogram gives an approximate number of ions detected for the target peptide (or sequence-related target peptide), and the location at the maximum time is for the target peptide (or sequence-related target peptide). Retention time when eluted from the chromatograph column.
Returning to FIG. 3, step 364 collects low-energy and high-energy AMRTs for each identification. These AMRTs are collected according to the following rules: That is, the value t from the detection chromatogram<sub>d</sub>Is present in the hit list, assuming that it is the elution time of the peptide, and the detection range, i.e. +/- 0.05 min t in the experiment<sub>d</sub>All AMRTs within are recorded, labeled, or collected. Therefore, these collected AMRTs are then subject to two conditions, (A) the associated retention time of AMRTs is t.<sub>d</sub>Satisfy that it is at +/- 0.05 minutes and (B) the mwHPlus value of AMRT is within 20 ppm of the molecular weight of the fragment and the molecular weight of the precursor in the selected peptide database.
The number of AMRTs collected by this rule is t<sub>d</sub>It will be close to the height of the detection chromatogram in, but not necessarily the same. The detection Gaussian peaks may not exactly match, as the peptide-related AMRTs may have slightly different retention times due to measurement errors. If the AMRT retention times do not have the same value, the height of the detection peak in the detection chromatogram that exceeds the detection threshold can be non-integer. However, the number of AMRTs collected by the above rules must obviously be integer values.
In step 366, this collected AMRT is stored. If desired, the spectrum of the collected AMRT can be displayed. In step 368, the search is repeated for the next precursor peptide, if necessary, by returning to step 350. If this search is not repeated, further analysis can be performed in step 370. Such further analysis can combine the results with the results from other injections or display the results while quantifying the identified peptides.
Combining the results from other injections can consist of comparing the retention times when the same peptide appears in two or more injections. Combining the results from other injections can consist of comparing the intensity of the corresponding AMRT found in two or more injections. These injections may be repeated injections of the same mixture or injections of two samples made under different conditions.
Retention times and intensities from repeated infusions can be compared to match for further confirmation that the peptide was identified correctly. If this injection is from a different sample (or condition), the retention times can be compared to match for further confirmation that the peptide was identified correctly, and the intensities can be compared or rationalized to two conditions. Changes in the amount of peptide in the sample between can be clarified.
Rules can be applied to the list of peptide identifications to estimate which proteins are present in the original sample.
FIG. 16 is a flow chart for collecting AMRT and ions associated with each of the identified sequences. In step 1602, a retention time window is established. Generally, the retention time window is set equal to the detected peak width, which is +/- 0.05 minutes in the above example. In step 1604, the retention time is detected and the retention time t<sub>d</sub>All labeled low-energy and high-energy AMRTs within the retention time window threshold centered above are collected. The detection peak provides a retention time for the peptide to elute. Retention time, detection retention time t<sub>d</sub>Ions within the retention time window threshold centered on are the ions detected for the peptide. This ion collection includes all ions or AMRT hits at low and high energies. These ions may or may not contain a mass corresponding to the target precursor.
The result is the ions found in the retention time window centered at the apex of the detection peak. These ions have a mass corresponding to the mass of the peptide fragment and generally include, but not always, the mass of the target precursor. In step 1606, this result is stored in storage. Further, in step 1606, this result can be displayed to the user.
FIG. 4A is an exemplary plot showing all AMRTs found at high energy. FIG. 4B is an exemplary plot showing all AMRTs found at low energy. In FIGS. 4A and 4B, the vertical axis is AMRT mwHPlus and the horizontal axis is retention time. FIG. 4C is an exemplary detection chromatogram derived from the data shown in FIGS. 4A and 4B, showing the number of hits per retention time corresponding to the precursor mass and fragment mass for a given peptide sequence. There is. Figure 4C clearly shows that the peak of hit occurrence is close to about 78 minutes. This peak contains precursor peptides found in the data. The distribution observable in Figure 4C shows the noise background where the importance of the peak of hit occurrence can be determined.
FIG. 5A is an exemplary plot showing only AMRT hits in the high energy plot of FIG. 4A, which corresponds to the mass of the precursor and the mass of the fragment for a given peptide sequence. FIG. 5B is an exemplary plot showing only AMRT hits in the low energy plot of FIG. 4B, which corresponds to the mass of the precursor and the mass of the fragment for a given peptide sequence. FIG. 5C is a histogram plot of the hits of the high and low energy plots of FIGS. 5A and 5B, which correspond to the mass of the precursor and the mass of the fragment for a given peptide sequence. FIG. 5C is an exemplary detection chromatogram similar to FIG. 4C, with the addition of a threshold of 502. Threshold 502 indicates the number of hits required to indicate the presence of the peptide. A clear precursor peptide 506 associated with the sequence was identified with a retention time of approximately 78 minutes, where more than 40 hits were counted. Peptide 504 associated with a possible sequence is identified with a retention time of 43 minutes. The hit of low-energy AMRT against the fragment mass supports insource fragmentation. Insource fragmentation of the precursor ion gives the fragment ion observed in the low energy spectrum.
Peak properties can be used to further assist peptide identification. One such property is the peak shape. All chromatograph peaks associated with the same precursor peptide must have the same peak shape and peak width. However, chromatograph peaks associated with different peptides may not have the same peak shape and width. Thus, the two peptides can elute at the same chromatograph retention time but have different peak shapes and / or widths. Therefore, peak geometries can be used to eliminate concurrency that could otherwise be misidentified. This peak shape property can be used to reduce the threshold, i.e., the number of ions required to match at the same retention time, to suggest the presence of the target peptide. Similarly, the peak shape can confirm the relationship between AMRTs.
All chromatograph peaks associated with the same precursor peptide must essentially have the same peak shape and width, but such peak shape or width changes occur due to measurement errors. May be observed. Another source of change is interference from other peaks that are not related to the precursor.
As shown in FIGS. 6A-6B, the peaks have several times in which the peak width and peak shape can be determined in comparison. These include vertex time (holding time), ascending inflection point time, and descending inflection point time. The inflection point can be obtained from the time of zero intersection of the quadratic derivative of the peak shape. The quadratic derivative can be determined by the Savitzky-Golay filter or the associated polynomial filter. FIG. 6A shows an exemplary chromatograph peak 602. FIG. 6B shows a plot of the second derivative of chromatograph peak 602. It shows the times of the vertices 604 and the inflection points 606a and 606b of the second derivative trace. These times can be compared to the peak shape and width. For reference, the dotted lines indicate the peak points of the top plot that correspond to the time seen in the bottom plot.
The time difference between the inflection point of the downhill slope and the inflection point of the uphill slope represents the peak width. For Gaussian chromatograph peaks, this width is twice the Gaussian standard deviation. The ratio of peak heights of ascending and descending inflection points is an additional measure of peak asymmetry or peak shape. The magnitude of the time difference between the apex time and the time of the ascending and descending inflection points is another measure of peak width. These time ratios are a measure of peak shape or asymmetry.
Considering the peak shape, additional processing of the detection chromatogram can be performed. As mentioned above, the maximum of the detection chromatogram is found. Compare the shapes and widths of the peaks that match during that maximum. If its width or shape is an outlier, the peak is rejected.
FIG. 7 is a flowchart of a method for comparing peak shapes and peak widths according to an embodiment of the present invention. In step 702, the peak retention times are compared. In step 704, the inflection points of the peaks are compared. The peak inflection point is the time between the inflection points. For example, the time between inflection points 606a and 606b in FIGS. 6A-6B is the inflection width of peak 602. In step 706, the magnitude of the difference between the apex time and the inflection time of the ascending gradient is compared. In step 708, the magnitude of the difference between the apex time and the inflection time of the descending gradient is compared.
In step 710, this time is analyzed to determine if they fall within the time threshold. In one embodiment of the invention, the time threshold for each comparison is a detection width of 0.05 minutes. Therefore, all time comparisons must be within 0.05 minutes to consider the peak as corresponding to the same peptide. This threshold may be user-specified or statistically determined. User-specified thresholds or statistically determined thresholds can be part of absolute time or peak width.
Suitable peaks used to compare peak shape and size are clusters of peptide-related ions.<sup>12</sup>C Monoisotopic peak. this<sup>12</sup>The C monoisotope is the lowest mass peak where all isotopes are in their most abundant state. Other peaks of the peptide cluster of ions can be used as well. In addition, the mean retention time and ascending and descending inflection times can be used to compare peak shapes and widths.
In a second embodiment of the invention, instead of or in addition to searching for AMRT, precursors, fragments, and ions corresponding to their isotopes are searched for. One advantage of using ions is that for low intensity peptides, the peptides may appear as single ions. For example, when using AMRT, at least two ions are required to detect AMRT and establish its charge state.
This ion-based search is in many ways similar to the AMRT search described above. FIG. 10 is a flowchart of a method for identifying a peptide in a complex mixture using ions according to an embodiment of the present invention. A summary of the steps in the ion-based PDS algorithm is as follows.
In step 1002, low-energy and high-energy ions are obtained from a single injection of the peptide mixture. Peptide mixtures are generally obtained from digests of protein samples. Low energy and high energy data are acquired using the high voltage / low voltage switching technology as described above. In step 1004, a database of proteins is selected. In step 1006, a list of peptides corresponding to the proteins in the database is obtained using the rules for peptide digestion.
In step 1008, the target precursor is selected from a database (eg, trypsin peptide). This peptide is described by its mass (mwHPlus) and its mass of Y and B ions (mwHPlus). In step 1009, data is retrieved for the mass corresponding to the mass of the selected precursor. Given these masses, step 1010 records all ions in the high and low energy list data that have masses within the search tolerance (eg, within 20 ppm). When using ions as search criteria, the search must be greater than multiple charge states and isotopic numbers. The charge state is generally limited to 1 to 3 for high energy fragments. The charge state is generally limited to 1-6 for low energies.
In a preferred embodiment of this ion-based search, the charge states of all ions obtained from the low energy spectrum are assumed to be Z = 2, and the charge states of all ions obtained from the high energy spectrum are Z. It is assumed that = 1. These allocations are made because Z = 2 is the most commonly observed low energy change of the peptide and Z = 1 is the most commonly observed high energy change of the peptide. The number of isotopes of all ions is assumed to be N = 0. That is, all ions are assumed to be in their monoisotopic state. These charge states and isotopic allocations are required to determine the mwHPlus value for each ion. It is these mwHPlus values that are then compared to the masses obtained from the database in step 1009, as described below.
In step 1012, the retention time of each hit ion is recorded. The output at this stage is a list of ions from high-energy and low-energy data hit with some mass from the peptides in the database.
In step 1016 as described above, a synthetic detection chromatogram is produced. In one embodiment of the invention, each hit ion is represented by a Gaussian peak as follows. The time range of the chromatogram corresponds to the time range of separation. If the chromatogram has a FWHM peak width of 0.5 minutes, for example, in one embodiment of the invention, the sample duration of the detection chromatogram is selected to be about 1% of that width, ie 0.005 minutes. The initial values of all points on the detection chromatogram are set to 0. The list of hit ions is crossed (looped over). For each input in the list, a Gaussian peak (the peak of the detected Gaussian) is added to the detection chromatogram. The width of the Gaussian peak is set to the specified portion of the FWHM of the chromatograph peak of the data. In one embodiment of the invention, this portion is 10%. Therefore, the Gaussian FWHM of detection is 0.05 minutes. The Gaussian peak width corresponds to the standard error used to measure the peak retention time.
If two ions with different masses elute at the same time, their detected gauss will reach a new peak with a peak height of 2. If N ions with different masses elute at the same time, their detected gauss will reach a new peak with a peak height N.
Steps 1018 to 1026 are the same as steps 362 to 370 in FIG. 3C above. The maximum of the detection chromatogram is found in step 1018. The threshold is determined. The method by which the detection threshold can be determined is specified as follows. Possible values for the threshold are four or more ions present in the same retention window (obtained by summing the low energy AMRT with the high energy AMRT).
All maxima above the detection threshold suggest that certain peptides are present in the data, either identical to or closely related to the peptides in the database. The maximum determines the retention time of detection. Retention times with a significant number of ions suggest that the target peptide is present in the data. Given this retention time, all low-energy and high-energy ions within this retention time threshold and within the mass threshold are selected in step 1020. Additional thresholds can be applied to determine the number of ions that meet these requirements. In step 1022, the group of ions above the threshold is recorded. These groups suggest peptide identification. In step 1024, if there are additional precursors, the search is repeated for the next precursor peptide by returning to step 1008. In step 1026, further analysis of the results may be performed.
To include any relevant ions as a hit when generating the detection chromatogram,<sup>12</sup>C ions may impose the additional requirement that they be found for each of the charge and isotope clusters. That is,<sup>13</sup>If C is seen, in the same charge state<sup>12</sup>Not counted unless C is seen.
The PDS algorithm described above has a number of advantages over conventional peptide identification methods. One of the prior art problems is that it assumes that the low energy AMRT (or ion) is only the precursor peptide. However, fragmentation can occur at low energy (insource fragmentation) as part of the ionization and focusing steps. Therefore, in conventional systems, the search can be initiated by AMRT, which is not actually a precursor. Such a search either does not result in a hit to the target, or a hit to the target results in an erroneous false identification. Such false identifications are called false positives.
However, using embodiments of the present invention, for example, as can be seen in FIG. 5B, insource fragments that appear at low energies are detected. Since embodiments of the present invention detect AMRTs in low-energy and high-energy data, all low-energy AMRTs that are in fact fragments (not trypsin precursors) are identified.
Another advantage of embodiments of the present invention is that the search can be performed and the peptide sequence can be identified without detecting the precursor mass. That is, peptides from the database may have a precursor molecular weight (mwHPlus) M. In the conventional method, the search is initiated by finding a peptide having a molecular weight (mwHPlus) M in the database. If the database does not contain peptides with this molecular weight, conventional systems do not perform identification.
However, the peptide mixture may contain peptides that are related but not identical to the peptides obtained from the database. For example, the peptide may be present in a sample that is chemically associated with the peptide in the database. In this case, the Y and / or B ions may be present, but their precursor mwHPlus may not be present. Using embodiments of the invention, ions that are overly abundant (exceeding the detection threshold) at a common retention time support the presence of molecules in the sample that are closely related to the sequence of the target precursor. Examples of steps that can result in such situations are modification of the primary sequence of the protein, or post-translational modification of the protein, or that one or the other end of the peptide may be modified or clipped after digestion. There is.
An example of the primary sequence of a modified protein is a single nucleotide polymorphism (SNP), which is a single nucleotide polymorphism difference in the DNA sequence. SNPs can occur with a frequency of 1 change / 100 bases. In some organisms, SNPs can give rise to trypsin peptides that differ by a single amino acid from the theoretical sequence derived from the protein database. This single amino acid substitution is sufficient to alter the mass of the precursor in relation to the theoretical mass of the unmodified sequence. This substitution leaves the rest of the peptide sequence intact. In particular, up to the point of amino acid substitution, the series of Y and B ions of the modified peptide is identical to the series of unmodified sequences.
Therefore, peptides in the sample mixture may result in having substantially the same sequence as the peptides in the database with mwHPlus mass M. However, changes in the sequence or chemical composition of a peptide in a sample generally change its mass precursor. Therefore, a precursor of mass M will not be present in the data, but a significant number of Y and B ions in the sequence corresponding to the mass derived from the database will be present in the sample data. Ions corresponding to these sub-sequence ions appear in the data with substantially the same retention time. Therefore, the accumulation of hits does not require the theoretical mass of precursors in the database to be present in the data.
The retention time of the modified sequence is generally different from the retention time of the target sequence (if any). The modified and target sequences are derived from two different peptide molecules. Each of the two molecules is retained differently in chromatographic separation. Therefore, if both the modified peptide and the (unmodified) target peptide are present in the sample, the detection chromatogram for the target sequence will reveal two detection peaks, one for each peptide. The retention time of these detection peaks reflects the retention time of each molecule.
It is noteworthy for the present invention that the mass tolerance of hits can reflect the inherent mass accuracy of the data. That is, it is not necessary to extend the mass tolerance and consider sequence modifications that may affect the precursor mass. For modified sequences, the present invention uses narrow mass tolerances that reflect the inherent mass accuracy of the data, eliminating the theoretical mass of the precursor. However, such a narrow mass tolerance further makes it possible to hit the Y and B ions present in the data corresponding to the theoretical masses of Y and B obtained from the database.
Thus, when sufficient hits accumulate during the retention time of elution of the modified peptide, a match with only the Y and B ions produces sufficient hits and the modified peptide is present in the data. The modified peptide is detected because it can be detected. The present invention can detect modified peptide sequences by using the theoretical fragment mass of unmodified peptides.
Therefore, the search can be performed and the peptide can be found in the sample data associated with the peptide in the database. In addition, the sample data can provide partial sequence information from the hit ions. Thus, the monoisotope of the database peptide may not be present in the LC / MS or LC / MS-E data, but in embodiments of the invention the database peptide is in the sample in a potentially modified form Can be further identified to be present in.
Given the detection of the modified peptide, subsequent examination of the mass in the data eluting during its retention time can reveal the exact sequence and mass of the modified peptide. For example, the ab initio sequencing algorithm can be applied to the mass of the data to determine the peptide sequence.
FIG. 11 shows peptides from a database that were found to be in the data at about 87 minutes. The highest value of mwHPlus seen at this retention time is consistent with its precursor, mwHPlus. Figure 11 appears to have a significant number of hits (more than expected for noise when it exceeds the threshold), but there are no precursors in the database, at about 49 minutes. An example of a peptide is also shown. Specifically, at about 49 minutes, about 5 ions are observed with a common retention time of 49 minutes. No ion with mwHPlus associated with the peptide is found. This can be an example of a peptide that is present in the sample and is chemically related to the peptide in the database. It should be noted that the two peptides elute at different retention times, suggesting different chemical compositions.
Thus, as shown by the plot in FIG. 11, the peptide mixture may contain peptides that are related to but not identical to the peptides obtained from the database.
In summary, embodiments of the present invention provide a number of advantages over conventional systems. These include retrieving data using a database with known peptide mass. Embodiments of the invention mistakenly identify the low energy fragment as a low energy (trypsin) precursor, but instead the peptide fragment in the low energy data (prior art erroneously assumes that each low energy AMRT is a precursor). (Estimate) can be identified.
Embodiments of the present invention can be configured to perform an archive search of data. Peptides should be able to be identified as long as the mass resolution of the MS is high enough, the retention time resolution is high enough, and the data are available using the high / low switching protocol. The intensities and retention times of these peptides can then be measured from the hit ions in the data.
Further, embodiments of the present invention allow for global retrieval of data. For example, high energy data can be archived worldwide and can be retrospectively retrieved using this algorithm.
Embodiments of the invention can detect peptides in data that share a sequence with peptides in a database, even if the shared sequences are not identical or the sharing is not complete. For example, the presence of subsequences of Y and B ions is typically sufficient to identify AMRT as related to the peptide from the database.
Also, embodiments of the present invention detect peptides in chemically modified data that have the same sequence as the peptides in the database. The presence of subsequences of the Y and B ions is sufficient to identify AMRT as related to the peptide from the database.
Embodiments of the present invention use retention time matching in high / low switching MS analysis. Improving chromatographic resolution directly leads to improved ability to identify peptides. For example, the chromatographic peak width is reduced (the resolution is increased), so that the detection threshold can be reduced. Embodiments of the invention further use peak shape and width matching to tune high / low MS analysis.
Embodiments of the invention can use retention time specificity to further identify ions that are chemical modifications of a peptide that has a peptide that has lost a precursor, eg, a neutral water molecule such as an ammonia molecule. .. After the peptide has been identified as present in the data, all ions associated with that peptide can be identified. As a result, these lower levels of peptides are not erroneously identified as precursors or Y or B fragments.
Embodiments of the present invention provide measurement of background noise. As a result, the significance or reliability of the identification can be calculated based on the statistical behavior of the data. Rather than a histogram, the number of hits in the retention interval, the detection chromatogram, can only count the longest contiguous sequence of hits in the retention interval. That is, when Y2, Y4, Y5, Y6, Y7, Y10 are hit, only these four ions (Y4, Y5, Y6, Y7) form a contiguous sequence of Y ions. The detection histogram will contain 4 ions at retention time. The significance of the detection threshold can be evaluated by Monte Carlo means or by obtaining the statistical characteristics of background hits from the data. In addition, embodiments of the present invention can be configured to assess the significance of hits using the distribution of peptides in the database.
Embodiments of the present invention can also be configured to use intensity and sequence rules in database retrieval. For example, if the candidate precursor AMRT is found in both high and low energy spectra, the intensity of the precursor in the high energy spectrum can be compared to the intensity of the precursor found in the low energy spectrum. If AMRT is a precursor, its intensity at high energy must be less than that at low energy. If the intensity at high energies is measured to exceed the intensity at low energies, the AMRT found at low energies and high energies is actually perhaps yet another precursor fragment. In this case, the PDS algorithm can be configured to eliminate low energy AMRT as a possible precursor, thereby eliminating possible false positive identifications.
As an example of another rule, if the intensity of fragment AMRT is determined to be out of proportion to the intensity of other fragments ARMT that may be associated with that precursor, then with a given precursor. Hits are deleted.
As an example of another rule, for a particular precursor sequence, the amino acid composition may indicate that the particular Y or B ion of that precursor should be effectively ionized. If the relative or absolute intensities of the fragment AMRT do not match such a model of ionization efficiency, the discrepancy may provide a basis for eliminating sequence identification. On the other hand, if the relative or absolute intensities of the fragment AMRT are consistent with a model of such ionization efficiency, this may support sequence identification.
The problem with analyzing AMRT is the dynamic range of their intensity. AMRT of a set of data can occur with intensity over a large dynamic range. The dynamic range of this intensity can be greater than 1: 1000. Advances in MS technology can extend this dynamic range to over 1: 10,000. The large dynamic range results from two effects. First, changes in peptide composition result in changes in ionization efficiency, so a given protein in a sample can produce AMRT with a dynamic range of 1: 100 or greater. As a result, some AMRTs can be ionized more effectively than others. Second, different proteins with a large dynamic range of concentration can occur in the sample.
Therefore, low intensity AMRTs can be generated from multiple sources. For example, low-intensity AMRT can result from high concentrations of proteins with weakly ionized peptides, or from low concentrations of proteins with effective ionization of peptides.
Therefore, high intensity dynamic range is a potential source of false positive identification. AMRT or ions from high concentrations of weakly ionized proteins can be misinterpreted as resulting from highly ionized peptides from low concentrations of proteins.
A method of dealing with this complexity caused by dynamic range is disclosed. The method, called Electronic Depletion Algorithm (EDA), identifies and removes all ions associated with high-concentration proteins before low-concentration proteins are analyzed. Removal of AMRTs associated with high concentrations of protein removes all high-intensity AMRTs in the sample, as well as all low-intensity ions in the sample coming from high-concentration proteins. Therefore, low-intensity AMRT from high-concentration proteins is not confused with AMRT from low-concentration proteins that are effectively ionized.
Elimination of low-intensity AMRT and ions associated with high concentrations of protein reduces an important source of false positives. For example, changes in low intensity AMRT may be misinterpreted as evidence of biomarkers of low concentrations of protein. In fact, in this case, it is a misidentified, weakly ionized fragment of a high concentration of protein.
FIG. 12 is a flowchart of a method of executing EDA according to an embodiment of the present invention. In step 1202, all AMRTs of the data are identified using the PDS described above. In step 1203, the peptide corresponding to AMRT is identified. The identified AMRTs are stored in step 1204 as intensities from the highest to the lowest intensities and are stored in the list.
In step 1206, the strongest peptide remaining on the list (the top item on the list) is selected. In step 1208, PDS is used to label each of the precursors, Y and B ions associated with the peptide with the peptide. In step 1210, each neutral loss AMRT associated with the precursor, Y ion, and B ion labeled in step 1208 is labeled with a peptide. In step 1212, all labeled AMRTs are removed from the data. In step 1214, if there is no AMRT with an intensity higher than a predetermined intensity threshold determined, the method ends in step 1216. On the other hand, if there is another AMRT with an intensity higher than the intensity threshold, the treatment continues in step 1206 with the remaining peptide having the highest intensity.
Embodiments of the invention can be configured to measure retention time and standard error of mwHPlus. The retention time measurement error associated with this method is the AMRT and ion error common to a single peptide. Therefore, this error is not an elution error. This error is due solely to the limitations used to measure the retention time.
The measurement error of mwHPlus related to the embodiments of the present invention is an error related to a database of accurately mass-measured proteins and peptides. This error is due to the limitations used to measure m / z with a mass spectrometer. There are two sources of this m / z error: statistical noise and calibration error.
These errors are used to set the thresholds used in determining which AMRT constitutes a hit and, in some cases, in constructing the detection chromatogram of the PDS method described above, so that the retention time error and The mwHPlus error must be estimated.
The retention time error and mwHPlus error are measured as the standard deviation of the error distribution. Given the standard deviation, the threshold used in determining which AMRT constitutes a hit and having a significant number of hits in the detection chromatogram is several times the standard deviation. A typical value may be 3 for a false positive rate of 3 sigma or 1/1000, or may specify an official false positive rate of 6 sigma or 1/1000000.
According to one embodiment of the invention, ions that are common at low and high energies are identified to determine retention time error and mwHPlus error. For example, trypsin precursors that appear at low energies also appear at high energies, despite their reduced intensity. The nominal difference in retention time between ions is necessarily zero. The nominal difference in mwHPlus is zero.
The observed difference in retention time between these ions is a measure of the retention time error. The standard deviation of this error, as determined by combining the errors from many such pairs, is the basis for measuring the standard error of retention time.
Based on the above, FIG. 13 is a method of determining the retention time error and the mwHPlus error according to one embodiment of the present invention. In step 1302, all low energy AMRTS (or ions) are looped. Use large thresholds for mass and retention time (eg, 50 ppm and 0.5 minutes, respectively) to find all matching high-energy AMRTs (or ions) in step 1304. In step 1306, the error between the retention time and the match with mwHPlus is analyzed. In step 1308, outliers are removed using standard techniques such as median filtering. In step 1310, the standard deviation for the resulting distribution is calculated. In step 1312, the standard deviation of mwHPlus is multiplied by statistical factors (3-6) to establish a molecular weight threshold that can be used in determining which AMRT hits in the PDS algorithm. In step 1314, the standard deviation of retention time is multiplied by statistical factors (3-6) to establish a retention time threshold that can be used in the step of generating the detection chromatogram of the PDS algorithm.
The detection threshold used in determining whether a peptide is present in the data can be determined in the following manner. If N or more AMRTs (or ions) are found, the peptides in the database are considered to have been detected in the data. This number N depends on the complexity of the sample. The more complex the sample, the larger N should be. By examining the background of the hits, N can be determined experimentally from the data. N can be established using standard histograms or other statistical techniques. In fact, values of N from 4 to 6 were found to be acceptable detection thresholds for AMRT found in the high / low data.
The results obtained from the present invention are a list of peptides identified in the sample. Each such peptide in the list contains the peptide sequence from the database, and the measured retention time of the precursor, the measured and theoretical mass, the measured intensity of the precursor, as well as found in the data. It also includes the measured retention time, intensity and mass of fragment ions, which are also associated with the precursor. It is expected that these results will be highly useful in the field of proteomics. For example, four uses of such proteomics are identification of proteins in samples, tracking of peptide retention times between samples, and quantification of peptides and proteins between samples.
Peptides identified by this method provide evidence that the original protein sequence has developed in the sample. For example, the list of proteins corresponding to the trypsin peptides identified by the present invention is one way to identify the proteins present in a sample. From the observed peptide identification and concentration, algorithms known in the art can also estimate the identification and concentration of the parent protein.
Entrants on this list may contain false positives. That is, if the peptide is identified incorrectly, the protein may be identified incorrectly. False positives can be reduced or eliminated using one of several means known in the prior art. To identify a protein, it can be required to identify two or more peptides for that protein by the methods of the invention. The user may specify that such peptides be detected when N> 1. Alternatively, the user may specify that a protein sequence in the smallest percentage range (by amino acids) is achieved by detecting one or more peptides for the protein. The user may also request that the protein be identified on each of the LC / MS analyzes of the sample repeated several times. Other rules that may be known in the art can also be applied to reduce false positive identification of proteins.
The present invention can be applied to data obtained from multiple samples analyzed on the same or different LC / MS systems. If peptide sequences from the database are found in more than one such sample or analysis, the retention times of that peptide from each sample or analysis can be compared. Therefore, the present invention provides a means of tracking the retention time of a peptide from injection to injection. Peptides can also be detected in data obtained by different instruments using different LC separation methodologies. Therefore, the retention time of the peptide can also be compared between such different instruments and LC separation.
Concordance of retention times of the same sequence as seen in multiple injections with the same or different instruments can be verified. Verification of such matches can be used to detect and eliminate false positive identifications.
When considering the strength of the same peptide, as seen in repeated injections of a given mixture, or when injecting samples under different conditions, the prior art method is applied to calibrate the strength. Changes in peptide and protein expression can be determined. For example, if the same peptide sequence is detected between injections of two samples, the ratio of the intensity of the corresponding precursors can also be calculated. If these peptides are from standard proteins for calibration, the ratio allows for relative or absolute concentration calibration of each infusion. If the peptide is from a protein that is endogenous to the sample, the ratio can be used to determine changes in the expression level of the peptide or the original protein.
The methods of the invention can be applied to mixtures other than mixtures of peptides. Given a database containing (1) any mixture of molecules, and (2) the masses of those molecules and their fragments, the method can be used to identify the molecules in a sample.
For the method applied, the sample is analyzed by the described LC / MS system. Fragments of precursor molecules by the above methods, as well as the theoretical masses of precursors and fragments, are known. If these conditions are met, the precursor can be identified by the method of the present invention. In the discussion of the above method, the precursor refers to the mass of ions found at low energy. Fragments refer to the mass of precursor fragments found at high or possibly low energy. The method then identifies the original molecule to be separated on the column.
Thus, for example, metabolic tests can benefit from the present invention. The digestive process is not required for metabolic molecules. All that is required for the present invention is a list of exact theoretical masses that correspond to the precursors and their associated fragments. Using such a list, the method can detect the presence of that set or subset of mass and the retention time when the original precursor molecule elutes from the chromatographic column.
In one preferred embodiment, the PDS and EAD algorithms are applied to the data acquired from the spectra collected by alternating high and low energy modes. However, both of these algorithms can be applied to spectra collected in only one mode, i.e. in a fixed energy mode. Thus, for example, these algorithms can be applied to only one of the low energy spectrum or the high energy spectrum. That is, in principle, both modes may be collected, but peptides can be identified using PDS or EDA algorithms that apply only to a single mode. As long as some fragmentation of the precursor ions is done in single energy mode only, these algorithms will detect the presence of precursors and / or fragments thereof in the data. The insource fragmentation apparent in Figures 5C and 11C indicates that the PDS or EDA algorithm may be applied only to low energy data.
Therefore, the requirement that data be collected in two modes is preferred, but not necessarily necessary for the application of PDS or EDA algorithms.
It should also be noted that the PDS or EDA algorithm can be applied to spectral data intentionally acquired with only a single fixed energy. In fact, it would also be advantageous to adjust the voltage (or voltage step) in a fixed energy mode to correspond to a voltage in the middle of the voltage typically used for low or high energy acquisition. The purpose would be to collect a spectrum containing the optimal mixture of precursors and fragments. Such acquisitions may also be useful for peptide identification using PDS or EDA algorithms.
The above disclosure of preferred embodiments of the present invention is for illustration and explanation. It is not intended to be inclusive or limited to the exact form in which the invention is disclosed. Considering the above disclosure, a number of modifications and modifications of the embodiments described herein will be apparent to those skilled in the art. The scope of the present invention is defined only by the appended claims and their equivalents.
Also, in describing a representative embodiment of the invention, the specification may indicate the methods and / or processes of the invention as a particular sequence of steps. However, to the extent that the method or process does not comply with the particular process sequence described herein, the method or process should not be limited to the particular process sequence described. As one of ordinary skill in the art will understand, the sequence of other steps is possible. Therefore, the particular sequence of steps described herein should not be construed as limiting the scope of the claims. Also, the scope of claims directed to the methods and / or processes of the present invention should not be limited to the performance of those steps in the written order, which will be varied by those skilled in the art. It can be easily understood that at the same time it is still in the spirit and scope of the present invention.
<figref num="1">FIG. 5 is a schematic showing a system for identifying and quantifying proteins in a biological complex according to an embodiment of the invention.</figref><figref num="2">It is a graph which shows the time when the exemplary spectrum was acquired. These spectra are obtained as a result of applying the alternating low energy mode and high energy mode according to one embodiment of the present invention.</figref><figref num="3A">It is a flowchart which shows the peptide identification method by one Embodiment of this invention.</figref><figref num="3B">It is a flowchart which shows the peptide identification method by one Embodiment of this invention.</figref><figref num="3C">It is a flowchart which shows the peptide identification method by one Embodiment of this invention.</figref><figref num="4A">It is an exemplary plot showing all AMRTs seen at high energy.</figref><figref num="4B">It is an exemplary plot showing all AMRTs seen at low energy.</figref><figref num="4C">Illustrative detection chromatograms derived from the data shown in FIGS. 4A and 4B, showing the number of hits per retention time for a theoretical given peptide from a database. FIG. 4C shows an example of a precursor peptide sequence in a database that appears to have a significant number of hits at about 78 minutes.</figref><figref num="5A">FIG. 4A is an exemplary plot showing only AMRT hits in the high energy plot of FIG. 4A.</figref><figref num="5B">It is an exemplary plot showing only AMRT hits in the low energy plot of FIG. 4B.</figref><figref num="5C">Histogram plots showing hits of the high and low energy plots of FIGS. 5A and 5B. FIG. 5C shows an example of a precursor peptide sequence in a database that appears to have a significant number of hits at about 78 minutes.</figref><figref num="6A">It shows how the second derivative zero intersection of the chromatograph peak is obtained.</figref><figref num="6B">It shows how the second derivative zero intersection of the chromatograph peak is obtained.</figref><figref num="7">It is a flowchart which shows the method of comparing the peak shape and the peak width by one Embodiment of this invention.</figref><figref num="8A">An exemplary spectrum of a peptide is shown.</figref><figref num="8B">It is a graph which shows the series of mass chromatograms corresponding to each of 6 ions shown in FIG. 8A.</figref><figref num="9A">It is a plot of an exemplary high energy spectrum showing a large number of spectral peaks.</figref><figref num="9B">It is a plot which shows only the spectral peak which has the same retention time as one retention time of the peak shown in FIG. 9A.</figref><figref num="9C">A series of plots showing chromatograph profiles corresponding to most of the peaks appearing in Figure 9A.</figref><figref num="10">FIG. 5 is a flowchart showing a method for identifying a peptide in a complex mixture using ions according to an embodiment of the present invention.</figref><figref num="11">It is a plot showing an example of a peptide at about 49 minutes, which appears to have a significant number of hits but no precursors in the database.</figref><figref num="12">It is a flowchart which shows the method of executing the electronic depletion algorithm (EDA) by one Embodiment of this invention.</figref><figref num="13">It is a flowchart which shows the method of determining the holding time error and mwHPlus error by one Embodiment of this invention.</figref><figref num="14">It is a flowchart which shows the method of generating the detection chromatogram which added the peak of the detection Gauss.</figref><figref num="15">It is a flowchart which shows the method of sequence identification and identification of their retention time by one Embodiment of this invention.</figref><figref num="16">It is a flowchart which shows the method of collecting AMRT and an ion associated with each sequence identification by one Embodiment of this invention.</figref><figref num="17A">It is a flowchart which shows the method of determining AMRT from an ion list by one Embodiment of this invention.</figref><figref num="17B">It is a flowchart which shows the method of determining AMRT from an ion list by one Embodiment of this invention.</figref><figref num="17C">It is a flowchart which shows the method of determining AMRT from an ion list by one Embodiment of this invention.</figref>
29 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29
Every citation, both waysCites: the store holds 4 of 5
| Document | Relation | Office |
|---|---|---|
| JP08334493A | Cites | Japan |
| JP2002110081A | Cites | Japan |
| WO2004019035A2 | Cites | World Intellectual Property Organization (WIPO) |
| JP2004503792A | Cites | Japan |
| Anna Pelander et al.,“Toxicological Screening with Formula-Based Metabolite Identification by Liquid Chromatography/Time,Anal. Chem.,2003年,Vol.75, No.21,pp.5710-5718 | Non-patent | – |
| Edward H.Kerns et al.,“Buspirone metabolite structure profile using a standard liquid chromatographic-mass spectrometric,Journal of Chromatography B,1997年,Vol.698, No.1/2,pp.133-145 | Non-patent | – |
19 members in 6 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 57253204 | United States of America | P | |
| 57253204 | United States of America | P | |
| 60572532 | United States of America | – | |
| 2005017742 | United States of America | W | |
| 2005017742 | United States of America | W | |
| 2004572532 | – | – | – |
| 2005017742 | – | – | – |
| US20040572532P | – | – | – |
| WO2005US17742 | – | – | – |
Members19
| Document | Office | Kind | |
|---|---|---|---|
| WO2005114930A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2005114930A3 | World Intellectual Property Organization (WIPO) | A3 | |
| GB0625396D0 | United Kingdom | D0 | |
| EP1756852A2 | European Patent Office (EPO) | A2 | |
| GB2430794A | United Kingdom | A | |
| DE112005001166T5 | Germany | T5 | |
| JP2007538260A | Japan | A | |
| US2008070314A1 | United States of America | A1 | |
| GB2430794B | United Kingdom | B | |
| EP1756852A4 | European Patent Office (EPO) | A4 | |
| US7851742B2 | United States of America | B2 | |
| US2011260049A1 | United States of America | A1 | |
| US8193485B2 | United States of America | B2 | |
| JP5008564B2This record | Japan | B2 | |
| US2012267522A1 | United States of America | A1 | |
| US8373115B2 | United States of America | B2 | |
| US2013282293A1 | United States of America | A1 | |
| EP1756852B1 | European Patent Office (EPO) | B1 | |
| DE112005001166B4 | Germany | B4 |
24 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A712A711 | A711 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5008564
- Publication, DOCDB
- 5008564
- Publication, EPODOC
- JP5008564B
- Application
- 2007527477
- Application, DOCDB
- 2007527477
- Application, EPODOC
- JP20070527477
Titles2
- Japanese
- 混合物中のタンパク質を同定する方法および装置
- English
- Methods and equipment for identifying proteins in mixtures
Classification
- CPC, 9
- G01N33/6848
- G01N2030/8831
- H01J49/0036
- H01J49/0431
- G16B15/00
- G01N33/6818
- H01J49/00
- H01J49/0027
- H01J49/426
- IPC, 6
- G01N27 62
- G01N30 72
- G01N30 88
- G01N33 68
- G16B15 00
- H04L13 18
