EP2236628A2

Reagents, methods and libraries for bead-based sequencing

Abstract

The present invention provides methods for determining a nucleic acid sequence by performing successive cycles of duplex extension along a single stranded template. The cycles comprise steps of extension, ligation, and, preferably, cleavage. In certain embodiments the methods make use of extension probes containing phosphorothiolate linkages and employ agents appropriate to cleave such linkages. In certain embodiments the methods make use of extension probes containing an abasic residue or a damaged base and employ agents appropriate to cleave linkages between a nucleoside and an abasic residue and/or agents appropriate to remove a damaged base from a nucleic acid. The invention provides methods of determining information about a sequence using at least two distinguishably labeled probe families. In certain embodiments the methods acquire less than 2 bits of information from each of a plurality of nucleotides in the template in each cycle. In certain embodiments the sequencing reactions are performed on templates attached to beads, which are immobilized in or on a semi-solid support. The invention further provides sets of labeled extension probes containing phosphorothiolate linkages or trigger residues that are suitable for use in the method. In addition, the invention includes performing multiple sequencing reactions on a single template by removing initializing oligonucleotides and extended strands and performing subsequent reactions using different initializing oligonucleotides. The invention further provides efficient methods for preparing templates, particularly for performing sequencing multiple different templates in parallel. The invention also provides methods for performing ligation and cleavage. The invention also provides new libraries of nucleic acid fragments containing paired tags, and methods of preparing microparticles having multiple different templates (e.g., containing paired tags) attached thereto and of sequencing the templates individually. The invention also provides automated sequencing systems, flow cells, image processing methods, and computer-readable media that store computer-executable instructions (e.g., to perform the image-processing methods) and/or sequence information. In certain embodiments the sequence information is stored in a database.

EP2236628A2, drawing sheet 1
Sheet 1 of 75

Term

Term ended

Projected expiry passed 1 February 2026, 0.6 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

15 claims: 1 independent, 14 dependent

  1. 1
    A method for of distinguishing a polymorphism on a template sequence from a sequencing error, the method comprising the steps of:(I) performing a sequencing reaction on a probe-template complex that has a duplex portion formed by hybridization between the template sequence and a probe, wherein the sequencing reaction comprises: (a) contacting the complex with a collection of oligonucleotide probes, the collection comprising at least two probe families, wherein: (i) a probe family is a group of one or more different probes that are identically labeled, and (ii) the probe families are labeled distinguishably from each other;(b) extending the duplex portion by ligating an extendable terminus of the duplex portion with a labeled probe from the collection that has hybridized to the template;(c) detecting the label of the ligated probe to identify its probe family;and (d) repeating steps (a) to (c) for a desired number of cycles to obtain an ordered series of probe family names of successively ligated probes, (II) optionally performing at least one more sequencing reaction;and (III) generating an ordered list of probe family names from at least one ordered series of probe family names from a sequencing reaction;(IV) comparing the ordered list of probe family names derived from the template sequence in step (III) with at least one ordered list generated from a reference sequence, and (V) ascribing a difference between the ordered lists to: i. a sequencing error if the difference consists of a single isolated probe family name, or ii. a polymorphism if the difference includes two or more adjacent probe family names.