High throughput nucleic acid sequencing by expansion
Abstract
A method for sequencing a target nucleic acid, comprising: a) providing a daughter strand produced by a template-directed synthesis, the daughter strand comprising a plurality of subunits coupled in a sequence corresponding to a contiguous nucleotide sequence of all or a portion of the target nucleic acid, in which the individual subunits comprise an anchor, at least one probe or nucleobase residue, and at least one selectively cleavable link; b) cleaving the at least one selectively cleavable link to give an Xpandomer of a length greater than the plurality of the daughter strand subunits, the Xpandomer comprising the anchors and indicator elements for analyzing the genetic information in a sequence corresponding to the sequence of contiguous nucleotides of all or a portion of the target nucleic acid; and c) detect the indicator elements of the Xpandomer.

Term
1.7 yearsto projected expiry
Projected expiry 19 June 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
8 claims: 1 independent, 7 dependent
- 1ES 2 559 313 T3 REIVINDICACIONES 1. Un método para la secuenciación de un ácido nucleico diana, que comprende:a) proporcionar una hebra hija producida por una síntesis dirigida por molde, comprendiendo la hebra hija una pluralidad de subunidades acopladas en una secuencia correspondiente a una secuencia de nucleótidos contigua de toda o una porción del ácido nucleico diana, en la que las subunidades individuales comprenden un anclaje, al menos una sonda o residuo de nucleobase, y al menos un enlace selectivamente escindible;b) escindir el al menos un enlace selectivamente escindible para dar un Xpandómero de una longitud mayor que la pluralidad de las subunidades de la hebra hija, comprendiendo el Xpandómero los anclajes y elementos indicadores para analizar la información genética en una secuencia correspondiente a la secuencia de nucleótidos contigua de toda o una porción del ácido nucleico diana;y c) detectar los elementos indicadores del Xpandómero.
- 2El método de la reivindicación 1, en el que los elementos indicadores para analizar la información genética están asociados con:los anclajes del Xpandómero;la hebra hija antes de la escisión del al menos un enlace selectivamente escindible;o el Xpandómero después de la escisión del al menos un enlace selectivamente escindible.
- 3El método de la reivindicación 1, en el que el Xpandómero comprende además toda o una porción de la al menos una sonda o residuo de nucleobase.
- 4El método de la reivindicación 3, en el que los elementos indicadores para analizar la información genética son o están asociados con la al menos una sonda o residuo de nucleobase.
- 5El método de la reivindicación 1, en el que el al menos un enlace selectivamente escindible es:un enlace covalente;un enlace intra-anclaje;un enlace entre o dentro de las sondas o residuos de nucleobase de la hebra hija;o un enlace entre las sondas o residuos de nucleobase de la hebra hija y un molde diana.
- 6El método de la reivindicación 1, en el que el Xpandómero comprende la siguiente estructura:en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;y α representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;α representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos ES 2 559 313 T3 contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;K en la que T representa el anclaje;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;en la que T representa el anclaje;N representa un residuo de nucleobase;ES 2 559 313 T3 κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;r « en la que T representa el anclaje;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;en la que T representa el anclaje;N representa un residuo de nucleobase;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;en la que T representa el anclaje;N representa un residuo de nucleobase;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;χ 1 representa un enlace con el anclaje de una subunidad adyacente;y χ 2 representa un enlace inter-anclaje;o en la que ES 2 559 313 T3 T representa el anclaje;n 1 y n 2 representa una primera porción y una segunda porción, respectivamente, de un residuo de nucleobase;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;y a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies comprende información de secuencias de la secuencia de nucleótidos contigua de una porción del ácido nucleico diana.
- 7El método de la reivindicación 6, en el que, antes de la escisión del al menos un enlace selectivamente escindible, la hebra hija comprende un dúplex de molde-hebra hija que tiene la siguiente estructura:en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;P 1 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 1 es complementaria;P 2 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 2 es complementaria;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;y a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;P 1 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 1 es complementaria;P 2 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 2 es complementaria;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;ES 2 559 313 T3 en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;P 1 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 1 es complementaria;P 2 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 2 es complementaria;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;γΤί P 1 ~ P 2 P 1 - P 2 en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;P 1 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 1 es complementaria;P 2 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 2 es complementaria;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;-pi-p 2 -P’-P 2 81 ES 2 559 313 T3 en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;P 1 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 1 es complementaria;P 2 representa una secuencia de nucleótidos contigua de al menos un residuo de nucleótido de la hebra molde a la que P 2 es complementaria;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a tres;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;NN* en la que T representa el anclaje;N representa un residuo de nucleobase;N' representa un residuo de nucleótido de la hebra molde al que N es complementario;~ representa el al menos un enlace selectivamente escindible;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;en la que T representa el anclaje;N representa un residuo de nucleobase;N' representa un residuo de nucleótido de la hebra molde al que N es complementario;~ representa el al menos un enlace selectivamente escindible;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;ES 2 559 313 T3 -Ν’en la que T representa el anclaje;N representa un residuo de nucleobase;N' representa un residuo de nucleótido de la hebra molde al que N es complementario;~ representa el al menos un enlace selectivamente escindible;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;y χ representa un enlace con el anclaje de una subunidad adyacente;en la que T representa el anclaje;N representa un residuo de nucleobase;N' representa un residuo de nucleótido de la hebra molde al que N es complementario;~ representa el al menos un enlace selectivamente escindible;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana;χ 1 representa un enlace con el anclaje de una subunidad adyacente;y χ 2 representa un enlace inter-anclaje;o N Ν' en la que T representa el anclaje;N representa un residuo de nucleobase;N' representa un residuo de nucleótido de la hebra molde al que N es complementario;V representa un sitio de escisión interno del residuo de nucleobase;κ representa la subunidad k n en una cadena de m subunidades, en la que m es un número entero superior a diez;y a representa una especie de un motivo de subunidad seleccionado de una biblioteca de motivos de subunidad, ES 2 559 313 T3 en la que cada una de las especies es complementaria a la secuencia de nucleótidos contigua de una porción del ácido nucleico diana.
- 8El método de la reivindicación 6, en la que la hebra hija se forma a partir de una pluralidad de construcciones de sustrato de oligómero que tienen la siguiente estructura:rTn en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;y R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;s representa un primer grupo conector;δ representa un segundo grupo conector;y “----” representa una reticulación intra-anclaje escindible;en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;ε representa un primer grupo conector;δ representa un segundo grupo conector;y “----” representa una reticulación intra-anclaje escindible;ES 2 559 313 T3 en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;ε representa un primer grupo conector;y δ representa un segundo grupo conector;en la que T representa el anclaje;P 1 representa un primer resto de sonda;P 2 representa un segundo resto de sonda;~ representa el al menos un enlace selectivamente escindible;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;ε representa un primer grupo conector;y δ representa un segundo grupo conector;R 1 -N-R 2 en la que T representa el anclaje;N representa un residuo de nucleobase;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;ε representa un primer grupo conector;δ representa un segundo grupo conector;y “----” representa una reticulación intra-anclaje escindible;ES 2 559 313 T3 en la que T representa el anclaje;N representa un residuo de nucleobase;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;~ representa el al menos un enlace selectivamente escindible;ε representa un primer grupo conector;δ representa un segundo grupo conector;y “----” representa una reticulación intra-anclaje escindible;en la que T representa el anclaje;N representa un residuo de nucleobase;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;ε representa un primer grupo conector;δ representa un segundo grupo conector;y “----” representa una reticulación intra-anclaje escindible;en la que T representa el anclaje;N representa un residuo de nucleobase;R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija;ε1 y ε2 representan los primeros grupos conectores iguales o diferentes;δ1 y δ2 representan los segundos grupos conectores iguales o diferentes;y “----” representa una reticulación intra-anclaje escindible;o -ΤΗ 1 - Ñ - R 2 en la que T representa el anclaje;N representa un residuo de nucleobase;V representa un sitio de escisión interno del residuo de nucleobase;y R 1 y R 2 representan los mismos grupos terminales o diferentes para la síntesis dirigida por molde de la hebra hija.
Independent claims8
778 paragraphs in 51 sections, as filed
ES 2 559 313 T3
DESCRIPTION
High-throughput Nucleic Acid Sequencing by Expansion
Background
Technical field
The present invention is generally related to nucleic acid sequencing, as well as related methods.
Description of Related Art
Nucleic acid sequences encode the information necessary for living things to function and reproduce, and are essentially a blueprint for life. Determining such sequences is therefore a useful tool in pure research on how and where organisms live, as well as in applied science such as drug development. In medicine, sequencing tools can be used to diagnose and develop treatments for various pathologies, including cancer, heart disease, autoimmune disorders, multiple sclerosis, or obesity. In industry, sequencing can be used to design improved synthetic organisms or enzymatic processes. In biology, such tools can be used to study ecosystem health, for example, and thus have a wide range of utility.
A unique DNA sequence of the individual provides valuable information regarding their susceptibility to certain diseases. The sequence will provide patients the opportunity to screen for early detection and to receive preventive treatment. In addition, given a patient's individual project, clinicians will be able to administer personalized therapy to maximize drug efficacy and to minimize the risk of an adverse drug response. Similarly, project determination of pathogens can lead to new infectious disease treatments and more robust pathogen monitoring. Whole genome DNA sequencing will provide the basis for modern medicine.
DNA sequencing is the process of determining the order of the chemical constituents of a given DNA polymer. These chemical constituents, which are called nucleotides, exist in DNA in four common forms: deoxyadenosine (A), deoxyguanosine (G), deoxycytidine (C), and deoxythymidine (T). Sequencing a diploid human genome requires determining the sequential ord of approximately 6 billion nucleotides.
Currently, most DNA sequencing is done using the chain termination method developed by Frederick Sanger. This technique, called Sanger sequencing, uses sequence-specific termination of DNA synthesis and fluorescently modified nucleotide reporter substrates to derive sequence information. This method sequences a target nucleic acid strand, or read length, up to 1000 bases in length using a modified polymerase chain reaction. In this modified reaction sequencing is randomly interrupted at selected base types (A, C, G, or T) and the lengths of the interrupted sequences are determined by capillary gel electrophoresis. The length then determines what type of base is located at that length. Many overlapping read lengths are produced and their sequences are overlapped using data processing to determine the most reliable fit of the data. This process of producing sequence read lengths is very time consuming and expensive and is now being replaced by newer methods that are more efficient.
Sanger's method was used to provide most of the sequence data in the Human Genome Project that generated the first complete sequence of the human genome. This project took 10 years and nearly $ 3 billion to complete. Given these significant performance and cost limitations, it is clear that DNA sequencing technologies will need to be drastically improved in order to achieve the established goals proposed by the scientific community. To this end, several second-generation technologies, which far exceeded the basic performance and cost limitations of Sanger sequencing, are gaining a growing share of the sequencing market. Yet these "sequencing by synthesis" methods fall short of achieving the performance, cost and quality goals required by markets such as whole genome sequencing for personalized medicine.
For example, 454 Life Sciences is producing instruments (eg, the genome sequencer) that can process 100 million bases in 7.5 hours with an average read length of 200 nucleotides. Their approach uses a variation of the polymerase chain reaction ("PCR") to produce a homogeneous colony of target nucleic acid, hundreds of bases in length, on the surface of a bead. This process is called emulsion PCR. Hundreds of thousands of such pearls are then arranged on a "picotiter plate". The plate is then prepared for further sequencing whereby each type of nucleic acid base is sequentially washed onto the plate. The target beads incorporating the base produce a pyrophosphate by-product that can be used to catalyze a reaction that produces light that is then detected with a camera.
Illumina Inc. has a similar process that uses reversible termination nucleotides and fluorescent labels to perform nucleic acid sequencing. The average read length for the Illumina 1G analyzer is
ES 2 559 313 T3 less than 40 nucleotides. Rather than using emulsion PCR to amplify sequence targets, Illumina has an approach to amplifying colonies by PCR on a matrix surface. Both the 454 and Illumina approaches use complicated polymerase amplification to increase signal intensity, perform base measurements during the rate-limiting sequence extension cycle, and have limited read lengths due to incorporation errors that they degrade the measurement signal with respect to noise proportionally to the read length.
Applied Biosystems uses reversible termination ligation instead of synthetic sequencing to read DNA. Like the 454 genome sequencer, the technology uses bead-based emulsion PCR to amplify the sample. Since most beads do not carry PCR products, the researchers then use an enrichment step to select DNA-coated beads. The biotin-coated beads are spread and immobilized on a streptavidin-coated glass slide matrix. The immobilized beads are then run through an 8-mer probe hybridization process (each labeled with four different fluorescent dyes), ligation, and cleavage (between bases 5<sup>to</sup> and 6<sup>to</sup> to create a site for the next round of ligation). Each probe interrogates two bases, at positions 4 and 5, using a 2-base encoding system, which is recorded by a camera. Similar to Illumina's approach, the average read length for the Applied Biosystems Solid platform is less than 40 nucleotides.
Other approaches are being developed to avoid the time and expense of the polymerase amplification step by measuring individual DNA molecules directly. Visigen Biotechnologies, Inc. is measuring fluorescently labeled bases as they are sequenced by incorporating a second fluorophore into an engineered DNA polymerase and using Forster resonance energy transfer (FRET) for nucleotide identification. This technique faces the challenges of separating the base signals that are separated by less than one nanometer and by a polymerase incorporation action that will have very large statistical variation.
A process that is developed by LingVitae sequence cDNA inserted into immobilized plasmid vectors. The process uses a class IIS restriction enzyme to cleave the target nucleic acid and bind an oligomer to the target. Typically, one or two nucleotides at the 5 'or 3' terminal overhang generated by the restriction enzyme determine which of a library of oligomers in the ligation mixture will be added to the sticky cut end of the target. Each oligomer contains "signal" sequences that uniquely identify the nucleotide (s) it replaces. The cleavage and ligation process is then repeated. The new molecule is then sequenced using specific labels for the various oligomers. The product of this process is called a "designer polymer" and always consists of a nucleic acid longer than the one it replaces (for example, a target dinucleotide sequence is substituted with an "extended" polynucleotide sequence of no less than 100 base pairs). An advantage of this process is that the strand of the duplex product can be amplified if desired. A disadvantage is that the process is necessarily cyclical and the continuity of the mold would be lost if multiple simultaneous restriction cuts were made.
US Patent No. 7,060,440 to Kless describes a sequencing process that involves incorporating oligomers by polymerization with a polymerase. A modification of the Sanger method, with end-terminated oligomers as substrates, is used to construct sequencing ladders by gel electrophoresis or capillary chromatography. Although the coupling of oligomers by end-ligation is well known, the use of a polymerase to couple oligomers in a template-directed process was used as a new advantage.
Polymerization techniques are expected to grow in potency as modified polymerases (and ligases) become available through genetic engineering and bioprospecting, and methods of removing exonuclease activity by polymerase modification are already known. For example, published US Patent Application 2007/0048748 to Williams describes the use of mutant polymerases to incorporate dye-labeled nucleotides and other modified nucleotides. Substrates for these polymerases also include γ-phosphate labeled nucleotides. Both high speed of incorporation and reduction in error rate were found with chimeric and mutant polymerases.
Furthermore, a great effort has been made by both academic and industrial teams to sequence native DNA using non-synthetic methods. For example, Agilent Technologies, Inc., along with university collaborators, are developing a single-molecule detection method that coils DNA through a nanopore to make measurements as it passes through. As with Visigen and LingVitae, this method must overcome the problem of efficiently and accurately obtaining distinct signals from individual nucleobases separated by sub-nanometric dimensions, as well as the problem of developing reproducible pore sizes of similar size. As such, direct sequencing of DNA by detection of its constituent parts now has to be accomplished in a high throughput process due to the small size of nucleotides in the strand (approximately 4 Angstroms center to center) and the limitations of the resolution of signal with respect to noise and corresponding signals in this respect. Direct detection is further complicated by the inherent secondary structure of DNA, which is not easily elongated in a perfectly linear polymer.
Although significant advances have been made in the field of DNA sequencing, new and improved methods remain a need in the art. The present invention meets these needs and
ES 2 559 313 T3 provides other related advantages.
Document US 2002/028458 A1 discloses a method of sequencing all or part of a target nucleic acid molecule that determines the sequence of a portion of said nucleic acid molecule, in addition to information regarding the position of said portion, and in particular provides a new sequencing method that involves augmenting one or more bases of said bases to aid in identification.
WO 00/79257 A discloses a method of evaluating a polymer molecule that includes linearly connected monomer residues. Such a method includes providing a polymer molecule in a liquid, contacting the liquid with an insulating solid state membrane having a detector capable of detecting characteristics of the polymer molecule, and causing the polymer molecule to pass through a limited region of the solid state membrane so that monomers of the polymer molecule cross the boundary in sequential order, whereby the polymer molecule interacts linearly with the detector and adequate data are obtained to determine the characteristics of the polymer molecule.
WO 2006/076650 A2 discloses compositions, methods and kits for selectively amplifying and detecting target sequences. In some embodiments, a circularizable probe and / or probe pair are disclosed to selectively amplify target sequences.
Short summary
Broadly speaking, corresponding methods and devices and products are disclosed that overcome spatial resolution challenges by existing high-throughput nucleic acid sequencing techniques. This is accomplished by encoding the nucleic acid information into an extended length surrogate polymer that is easier to detect. The surrogate polymer (referred to herein as an "Xpandomer") is formed by template-directed synthesis that preserves the original genetic information of the target nucleic acid, while also increasing the linear separation of individual elements from the sequence data.
In one embodiment, a method for sequencing a target nucleic acid is disclosed, comprising: a) providing a daughter strand produced by template directed synthesis, the daughter strand comprising a plurality of subunits coupled in a sequence corresponding to a contiguous nucleotide sequence of all or a portion of the target nucleic acid, wherein the individual subunits comprise an anchor, at least one probe or nucleobase residue, and at least one selectively cleavable bond; b) cleaving the at least one selectively cleavable bond to give an Xpandomer of a length greater than the plurality of subunits of the daughter strand, the Xpandomer comprising the anchors and reporter elements to analyze the genetic information in a sequence corresponding to the sequence of contiguous nucleotides of all or a portion of the target nucleic acid; and c) detecting the indicator elements of the Xpandomer.
In more specific embodiments, the reporter elements for analyzing genetic information can be associated with the Xpandomer anchors, with the daughter strand before cleavage of the at least one selectively cleavable bond, or with the Xpandomer after cleavage of the at least one bond. selectively cleavable. The Xpandomer may further comprise all or a portion of the at least one probe or nucleobase residue, and the reporter elements for analyzing the genetic information may be associated with the at least one probe or nucleobase residue or may themselves be the probe or residues. nucleobase. In addition, the selectively cleavable bond can be a covalent bond, an intra-anchor bond, a bond between or within the probes or nucleobase residues of the daughter strand, and / or a bond between the probes or nucleobase residues of the strand. daughter and a target mold.
In other embodiments, the Xpandomers have the following structures (I) to (X):
Structure (I):
<img file="ES2559313T3_D0001.tif" />
in which
T represents the anchor;
P<sup>1</sup> represents a first probe moiety;
P<sup>2</sup> represents a second probe moiety;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than three; already represents one species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid.
ES 2 559 313 T3
Structure (II):
<img file="ES2559313T3_D0002.tif" />
in which
T represents the anchor;
P<sup>1</sup> represents a first probe moiety;
P<sup>2</sup> represents a second probe moiety;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than three;
α represents a species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid; and χ represents a bond with the anchor of an adjacent subunit.
Structure (III):
<img file="ES2559313T3_D0003.tif" />
in which
T represents the anchor;
P<sup>1</sup> represents a first probe moiety;
P<sup>2</sup> represents a second probe moiety;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than three;
α represents a species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid; and χ represents a bond with the anchor of an adjacent subunit.
Structure (IV):
<img file="ES2559313T3_D0004.tif" />
in which
T represents the anchor;
P<sup>1</sup> represents a first probe moiety;
P<sup>2</sup> represents a second probe moiety;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than three;
α represents a species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid; Y
ES 2 559 313 T3 * represents a bond with the anchor of an adjacent subunit.
Structure (V):
in which
T represents the anchor;
χ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than three;
a represents a species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid; and χ represents a bond with the anchor of an adjacent subunit.
Structure (VI):
κ in which
T represents the anchor;
N represents a nucleobase residue;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than ten;
a represents a species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid; and χ represents a bond with the anchor of an adjacent subunit.
Structure (VII):
in which
T represents the anchor;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than ten;
a represents a species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid; and χ represents a bond with the anchor of an adjacent subunit.
ES 2 559 313 T3
Structure (VIII):
<img file="ES2559313T3_D0005.tif" />
in which
T represents the anchor;
N represents a nucleobase residue;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than ten;
a, represents a species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid; and χ represents a bond with the anchor of an adjacent subunit.
Structure (IX):
r -i (xx<sup>1</sup> x<sup>2</sup> t
LNJ<sub>K</sub> in which
T represents the anchor;
N represents a nucleobase residue;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than ten;
a represents a species of a subunit motif selected from a library of subunit motifs, wherein each of the species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid;
χ represents a bond with the anchor of an adjacent subunit; and χ<sup>2</sup> represents an inter-anchor bond.
Structure (X):
<img file="ES2559313T3_D0006.tif" />
in which
T represents the anchor;
n<sup>1</sup> and n<sup>2</sup> represents a first portion and a second portion, respectively, of a nucleobase residue;
κ represents subunit k<sup>n</sup> in a chain of m subunits, where m is an integer greater than ten; already represents one species of a subunit motif selected from a library of subunit motifs, wherein each species comprises sequence information from the contiguous nucleotide sequence of a portion of the target nucleic acid.
In addition, oligomer substrate constructs are disclosed for use in template-directed synthesis for sequencing a target nucleic acid.
The oligomer substrate constructs comprise a first probe moiety attached to a second probe moiety, each of the first and second probe moieties having a terminal group suitable for targeted synthesis.
ES 2 559 313 T3 per template, and an anchor having a first end and a second end with at least the first end of the anchor attached to at least one of the first and second probe moieties, wherein the oligomer substrate construct when used in template directed synthesis is capable of forming a daughter strand comprising a limited Xpandomer and having a plurality of subunits coupled in a sequence corresponding to the contiguous nucleotide sequence of all or a portion of the target nucleic acid, wherein the individual subunits comprise an anchor, the first and second probe moieties, and at least one selectively cleavable bond.
In addition, oligomer substrate constructs are disclosed for use in template-directed synthesis for sequencing a target nucleic acid. The oligomer substrate constructs comprise a nucleobase residue with end groups suitable for template-directed synthesis, and an anchor having a first end and a second end with at least the first end of the anchor attached to the nucleobase residue, wherein the monomer substrate construct when used in template directed synthesis is capable of forming a daughter strand comprising a limited Xpandomer and having a plurality of subunits coupled in a sequence corresponding to the contiguous nucleotide sequence of all or a portion of the target nucleic acid, wherein the individual subunits comprise an anchor, the nucleobase residue, and at least one selectively cleavable bond.
Furthermore, template-daughter strand duplexes comprising a daughter strand duplexed with a template strand are disclosed, in addition to methods of forming the same from the template strand and the oligomer or oligomer substrate constructs.
These and other aspects of the invention will become apparent upon reference to the accompanying drawings and upon detailed description. To this end, various references are set forth herein that describe in more detail certain procedures, compounds, and / or compositions.
Brief description of the drawings
In the figures, identical reference numerals identify similar elements. The sizes and relative positions of the elements in the figures are not necessarily drawn to scale and some of these elements are arbitrarily enlarged and positioned to improve the readability of the figure. Furthermore, the particular shapes of the elements as they are drawn are not intended to express any information regarding the actual shape of the particular elements, and have only been selected to facilitate recognition in the figures.
Figures 1A and 1B illustrate the limited spacing between nucleobases that must be resolved in order to determine the nucleotide sequence in a nucleic acid target.
Figures 2A to 2D schematically illustrate various representative substrate structures useful in the invention.
Figures 3A, 3B and 3C are schemes illustrating simplified steps for synthesizing an Xpandomer of a target nucleic acid.
Figure 4 is a simple model illustrating a FRET nanopore device for sequencing an Xpandomer.
Figure 5 is a graph with channels for red, green and blue fluorescence emissions, and illustrates how analog signals can be decoded into digital information that corresponds to the genetic information of sequences encoded in an Xpandomer. The attached table (Figure 6) shows how the data is decoded. Employing three multi-state fluorophores, the base sequence can be read with high resolution digitally from a single molecule in real time as the Xpandomer is wound through the nanopore.
Figure 6 is a look-up table from which the data in Figure 5 is derived.
Figures 7A-E are ligation product gels.
Figure 8 is an overview of oligomeric Xpandomers.
Figure 9 is an overview of monomeric Xpandomers.
Figures 10A to 10E depict Class I Xpandomers, intermediates and precursors in symbolic and graphical form. These precursors are called Xprobes if they are monophosphates and Xmers if they are triphosphates.
Figure 11 is a condensed scheme of a method of synthesis of an Xpandomer by solution ligation using an end-terminated hairpin primer and class I substrate constructs.
Figure 12 is a condensed scheme of a method of synthesis of an Xpandomer by solution ligation using a double-ended hairpin primer and class I substrate constructs.
Figure 13 is a condensed scheme of a method of synthesis of an Xpandomer by ligation on an immobilized template, without primers, using class I substrate constructs.
Figure 14 is a condensed scheme of a method of synthesis of an Xpandomer by cyclic step ligation using reversibly terminated class I substrate constructs on hybridized templates for immobilized primers.
Figure 15 is a condensed scheme of a method of synthesis of an Xpandomer by promiscuous assembly and chemical coupling using primerless class I substrate constructs.
Figure 16 is a condensed scheme of a method of synthesis of an Xpandomer by polymerization in
ES 2 559 313 T3 dissolution using a hairpin primer and class I triphosphate substrate constructs.
Figure 17 is a condensed scheme of a method of synthesis of an Xpandomer on an immobilized template using class I triphosphate substrate constructs and a polymerase.
Figures 18A to 18E depict a class II Xpandomer, Xpandomer intermediate, and substrate construction in graphic and symbolic language.
Figures 19A to 19E depict a class III Xpandomer, Xpandomer intermediate and substrate construction in symbolic and graphic language.
Figure 20 is a condensed scheme of a method of synthesis of an Xpandomer on an immobilized template using class II substrate constructs combining primerless chemical coupling and hybridization.
Figure 21 is a condensed scheme of a method of synthesis of an Xpandomer on an immobilized template using a primer, class II substrate constructs and a ligase.
Figures 22A through 22E depict a class IV Xpandomer, Xpandomer intermediate and substrate construction in symbolic and graphical form.
Figures 23A through 23E depict a class V Xpandomer, Xpandomer intermediate and substrate construction in symbolic and graphical form.
Figure 24 is a condensed scheme of a method of synthesis of an Xpandomer by solution polymerization using an adaptamer primer and class V triphosphate substrate constructs.
Figure 25 illustrates structures of deoxyadenosine (A), deoxycytosine (C), deoxyguanosine (G), and deoxythymidine (T).
Figures 26A and 26B illustrate derivatized nucleotides with functional groups.
Figures 27A and 27B illustrate probe members incorporating derivatized nucleobases.
Figures 28A through 28D illustrate in more detail class I-IV substrates of the invention, showing here examples of selectively cleavable bond cleavage sites in the probe backbone and indicating loop end bonds connecting the cleavage sites. .
Figures 29A to 29D illustrate class I-IV substrates of the invention in more detail, showing here examples of selectively cleavable bond cleavage sites in the probe backbone and indicating loop end bonds connecting the cleavage sites. .
Figure 30 illustrates a method of assembling a "probe-loop" construct, such as an Xprobe or an Xmer.
Figure 31 illustrates a method of assembling a class I substrate construction, in which the loop contains indicator constructs.
Figures 32A to 32C illustrate the use of PEG as a polymeric anchor.
Figures 33A to 33D illustrate poly-lysine as a polymeric anchor and dendrimeric constructs derived from poly-lysine scaffolds.
Figures 34a through 34C illustrate selected anchor loop closure methods integrated with the segment indicator construction assembly.
Figures 35A and 35B illustrate a method of synthesis of a single reporter segment by randomized polymerization of precursor blocks.
Figures 36A through 36I illustrate indicator constructions.
Figure 37 is a table showing the chemical composition and methods for assembling indicator constructs and their corresponding indicator codes.
Figures 38A through 38F are suitable adapters for use in end functionalization of target nucleic acids.
Figure 39 is an adapter cassette for introducing a terminal ANH functional group onto a dsDNA template.
Figure 40 is a scheme for immobilizing and preparing a template for the synthesis of an Xpandomer.
Figures 41A through 41E illustrate selected physical stretching methods.
Figures 42A through 42C illustrate selected electrostretching methods.
Figures 43A through 43D illustrate methods, reagents and adapters for stretching in gel matrices.
Figures 44A through 44C describe the construction and use of "drag marks."
Figure 45 depicts a promiscuous hybridization / ligation-based method for synthesizing an Xpandomer.
Figures 46A and B describe nucleobases used to fill gaps.
Figures 47A and B describe simulations of the appearance of voids.
Figures 48A and B describe simulations of the appearance of voids.
Figure 49 illustrates how the gaps are filled with 2mers and 3mers.
Figures 50A and B depict gap filling simulations using combinations of 2-mer and 3-mer.
Figure 51 illustrates the use of 2mer and 3mer adjuvants to break down secondary structure.
Figure 52 depicts bases useful as adjuvants.
Figure 53 depicts nucleotide substitutions used to reduce secondary structure.
Figure 54 depicts a magnetic bead transport nanopore detection model.
Figure 55 illustrates a conventional nanopore detection method.
Figure 56 illustrates a transverse electrode nanopore detection method.
Figure 57 illustrates a microscopic detection method.
Figure 58 illustrates detection by electron microscopy.
ES 2 559 313 T3
Figure 59 illustrates detection using atomic force microscopy.
Figures 60A to 60E depict a class VI Xpandomer, Xpandomer intermediate, and substrate construction in graphic and symbolic language. These precursors are called RT-NTP.
Figure 61 is a condensed scheme of a method of synthesis of an Xpandomer on an immobilized template using reversibly terminated class VI triphosphate substrate constructs and a polymerase.
Figures 62A to 62E depict a class VII Xpandomer, Xpandomer intermediate and substrate construction in graphic and symbolic language. These precursors are called RT-NTP.
Figures 63A to 63E depict a class VIII Xpandomer, Xpandomer intermediate, and substrate construction in graphic and symbolic language. These precursors are called RT-NTP.
Figures 64A to 64E depict a class IX Xpandomer, Xpandomer intermediate and substrate construction in symbolic and graphic language. These precursors are called RT-NTP.
Figure 65 is a condensed scheme of a method of synthesis of an Xpandomer on an immobilized template using class IX triphosphate substrate constructs and a polymerase.
Figures 66A to 66E depict a class X Xpandomer, Xpandomer intermediate and substrate construction in graphic and symbolic language. These precursors are called XNTP.
Figure 67 is a condensed scheme of a method of synthesis of an Xpandomer by solution polymerization using a hairpin primer and class X triphosphate substrate constructs.
Detailed description
Certain specific details are set forth in the following description in order to provide a thorough understanding of various embodiments. However, one skilled in the art will understand that the invention can be practiced without these details. In other cases, well-known structures have not been shown or described in detail to avoid unnecessarily illegible descriptions of the embodiments. Unless the context requires otherwise, throughout the specification and claims that follow, the word "comprise" and variations thereof, such as "comprises" and "comprising", are to be construed in an open inclusive sense, that is, as "including, but not limited to." Furthermore, the headings herein are provided for convenience only and do not construe the scope or meaning of the claimed invention.
Reference throughout this specification to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the occurrences of the phrase "in one embodiment" at various places throughout this specification are not all necessarily with reference to the same embodiment. Furthermore, the particular features, structures, or characteristics can be combined in any suitable way in one or more embodiments. Thus, as used in this specification and the appended claims, the singular forms "a", "an", "the" and "the" include plural referents, unless the content clearly dictates otherwise. It should also be noted that the term "or" is generally used in its sense that it includes "and / or", unless the content clearly dictates otherwise.
Definitions
As used herein, and unless the context dictates otherwise, the following terms have the meanings specified below.
"Nucleobase" is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. A nucleobase can exist naturally or synthetically. Non-limiting examples of nucleobases are adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purines substituted at position 8 with methyl or bromine, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7 -deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6-diaminopurine, N6-ethane2,6-diaminopurine, 5-methylcytosine, 5-alkynyl (C3-C6) -cytosine, 5-fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylaloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine and the non-naturally occurring nucleobases described in US Patent Nos. 5,432,272 and 6,150 .510 and PCT applications WO 92/002258, WO 93/10820, WO 94/22892 and WO 94/24144, and Fasman ("Practical Handbook of Biochemistry and Molecular Biology", pp. 385-394, 1989, CRC Press, Boca Mouse, LO).
"Nucleobase residue" includes nucleotides, nucleosides, fragments thereof, and related molecules that have the property of binding to a complementary nucleotide. Deoxynucleotides and ribonucleotides, and their various analogues, are contemplated within the scope of this definition. Nucleobase residues can be members of oligomers and probes. "Nucleobase" and "nucleobase residue" may be used interchangeably herein and are generally synonymous, unless the context dictates otherwise.
"Polynucleotides", also called nucleic acids, are covalently linked series of nucleotides in which the 3 'position of the pentose of one nucleotide is linked by a phosphodiester group to the 5' position of the next. DNA (deoxyribonucleic acid) and RNA (ribonucleic acid) are biologically produced polynucleotides in which nucleotide residues are linked in a specific sequence by phosphodiester bonds. As used herein, the terms "polynucleotide" or "oligonucleotide" encompass any compound of
ES 2 559 313 T3 polymer having a linear nucleotide backbone. Oligonucleotides, also called oligomers, are generally shorter chain polynucleotides.
"Complementary" generally refers to the duplexing of specific nucleotides to form Watson-Crick canonical base pairs, as understood by those skilled in the art. However, complementary as referred to herein also includes base pairing of nucleotide analogs, including, but not limited to, 2'-deoxyinosine and 5-nitroindole-2'-deoxyribo, which are capable of pairing of Universal bases with nucleotides A, T, G or C and blocked nucleic acids, which enhance the thermal stability of the duplexes. One skilled in the art will recognize that the stringency of hybridization is a determinant in the degree of matching or mismatch of the duplex formed by hybridization.
"Nucleic acid" is a polynucleotide or an oligonucleotide. A nucleic acid molecule can be deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination of both. Nucleic acids are generally referred to as "target nucleic acids" or "target sequence" if they are targeted for sequencing. Nucleic acids can be mixtures or pools of molecules targeted for sequencing.
"Probe" is a short strand of nucleobase residues, generally referring to two or more contiguous nucleobase residues that are generally single-stranded and complementary to a target nucleic acid sequence. As integrated into "Substrate Members" and "Substrate Constructs", probes can be up to 20 nucleobase residues in length. Probes can include modified nucleobase residues and modified intra-nucleobase linkages in any combination. Probe backbones can be linked together by any of several types of covalent bonds, including, but not limited to, ester bond, phosphodiester, phosphoramide, phosphonate, phosphorothioate, phosphorothiolate, amide, and any combination thereof. The probe can also have 5 'and 3' end linkages including, but not limited to, the following moieties: monophosphate, triphosphate, hydroxyl, hydrogen, ester, ether, glycol, amine, amide, and thioester.
"Selective hybridization" refers to specific complementary binding. Polynucleotides, oligonucleotides, probes, nucleobase residues, and fragments thereof selectively hybridize to target nucleic acid strands, under hybridization and wash conditions that minimize non-specific binding. As is known in the art, high stringency conditions can be used to achieve selective hybridization conditions that favor perfect mating. Conditions for hybridization such as salt concentration, temperature, detergents, PEG, and GC neutralizing agents such as betaine can be varied to increase the stringency of hybridization, i.e., the requirement for exact C-matches for base-pairing with G , and A for base pairing with T or U, along with a contiguous strand of a duplex nucleic acid.
"Template directed synthesis", "template directed assembly", "template directed hybridization", "template directed binding" and any other template directed process refer to a process by which nucleobase residues or probes are attached. selectively to a complementary target nucleic acid, and are incorporated into a nascent daughter strand. A daughter strand produced by template-directed synthesis is complementary to the single-stranded target from which it is synthesized. It should be noted that the corresponding sequence of a target strand can be deduced from the sequence of its daughter strand, if this is known. "Template-directed polymerization" and "template-directed ligation" are special cases of template-directed synthesis whereby the resulting daughter strand is polymerized or ligated, respectively.
"Contiguous" indicates that a sequence continues without a break or absent nucleobase. The contiguous nucleotide sequence of the template strand is said to be complementary to the contiguous sequence of the daughter strand.
"Substrates" or "substrate members" are oligomers, probes, or nucleobase residues that have binding specificity to the target template. Substrates are generally combined with anchors to form substrate constructions. Substrates of the substrate constructs that form the primary backbone of the daughter strand are also substrates or substrate members of the daughter strand.
"Substrate constructs" are reagents for template-directed synthesis of daughter strands, and are generally provided in the form of libraries. Substrate constructs generally contain a substrate member for complementary binding to a target template and both an anchor member and anchor binding sites to which an anchor can bind. Substrate constructions are provided in a variety of forms adapted to the invention. Substrate constructs include both "oligomeric substrate constructs" (also called "probe substrate constructs") and "monomeric substrate constructs" (also called "nucleobase substrate constructs").
"Subunit motif" or "motif" refers to a repeating subunit of a polymer backbone, the subunit having an overall shape characteristic of repeat subunits, but also having species-specific elements that encode genetic information. The complementary nucleobase residue motifs are represented in substrate construct libraries according to the number of possible combinations of the binding nucleobase elements in the basic complementary sequence in each motif. If the nucleobase binding elements are four (for example, A, C, G, and T), the number of possible motifs of combinations of
ES 2 559 313 T3 four elements is 4<sup>x</sup>, where x is the number of nucleobase residues in the motif. However, other degenerate base-pairing-based motifs, substituting uracil for thymidine in ribonucleobase residues or other sets of nucleobase residues, can lead to larger libraries (or smaller libraries) of motif-bearing substrate constructs. The motifs are also represented by species-specific reporter constructs, such as the indicators that constitute a reporter anchor. There is generally a one-to-one correlation between the reporter construction motif that identifies a particular substrate species and the binding complementarity and specificity of the motif.
"Xpandomer intermediate" is an intermediate (also referred to herein as a "daughter strand") assembled from substrate constructs, and is formed by template-directed assembly of substrate constructs using a target nucleic acid template. Optionally, other bonds are formed between confined substrate constructs which may include polymerization or ligation of the substrates, anchor-to-anchor bonds, or anchor-to-substrate bonds. The Xpandomer intermediate contains two structures; specifically, the limited Xpandomer and the primary backbone. The limited Xpandomer comprises all of the anchors in the daughter strand, but may comprise all, a portion, or none of the substrate as required by the method. The primary backbone comprises all confined substrates. Under the process step in which the primary backbone is fragmented or dissociated, the limited Xpandomer is no longer limited and is the Xpandomer product that spreads as the anchors stretch. "Duplex daughter strand" refers to an Xpandomer intermediate that hybridizes or duplexes with the target template.
"Primary backbone" refers to a contiguous or segmented backbone of daughter strand substrates. A commonly encountered primary backbone is the ribosyl 5'-3 'phosphodiester backbone of a native polynucleotide. However, the primary backbone of a daughter strand may contain nucleobase analogs and oligomer analogs not linked by phosphodiester linkages or linked by a mixture of phosphodiester linkages and other backbone linkages, including, but not limited to, the following linkages : phosphorothioate, phosphorothiolate, phosphonate, phosphoramidate, and peptide nucleic acid "PNA" linkages including phosphono-PNA, serine-PNA, hydroxyproline-PNA, and combinations thereof. If the daughter strand is in its duplex form (ie, duplex daughter strand), and the substrates are not covalently bound between the subunits, the substrates are nevertheless contiguous and form the primary backbone of the daughter strand.
"Limited Xpandomer" is an Xpandomer in a configuration before it has been expanded. The limited Xpandomer comprises all anchor members of the daughter strand. It is limited from expanding by at least one bond or anchor bond that is attached to the primary backbone. During the expansion process, the primary backbone of the daughter strand fragments or dissociates to transform the limited Xpandomer into an Xpandomer.
"Limited Xpandomer backbone" refers to the limited Xpandomer backbone. It is a covalent synthetic backbone co-assembled together with the primary backbone in the formation of the daughter strand. In some cases, both skeletons may not be discrete, but both may have the same substrate or portions of the substrate in their composition. The limited Xpandomer backbone always comprises anchors, while the primary backbone does not comprise anchor members.
"Xpandomer" or "Xpandomer product" is a synthetic molecular construct produced by expansion of a limited Xpandomer, which is itself synthesized by mold-directed assembly of substrate constructs. The Xpandomer is elongated relative to the target template from which it was produced. It is composed of a concatenation of subunits, each subunit a motif, each motif a member of a library, comprising sequence information, an anchor, and optionally a portion, or all, of the substrate, all of which are derived from the formative substrate construct. . The Xpandomer is designed to expand to be larger than the target template, thus reducing the linear density of sequence information from the target template along its length. In addition, the Xpandomer optionally provides a platform to increase the size and abundance of indicators, which in turn improves signal over noise for detection. Better linear information density and stronger signals increase resolution and reduce sensitivity requirements for detecting and decoding the template strand sequence.
"Selectively cleavable bond" refers to a bond that can be cleaved under controlled conditions such as, for example, conditions for the selective cleavage of a phosphorothiolate bond, a photocleavable bond, a phosphoramide bond, a 3'-OBD-ribofuranosyl-2 bond ', a thioether bond, a selenoether bond, a sulfoxide bond, a disulfide bond, a deoxyribosyl-5'-3' phosphodiester bond, or a ribosyl-5'-3 'phosphodiester bond, in addition to other cleavable bonds known in the art. A selectively cleavable bond may be an intraanchor bond or between or within a probe or nucleobase residue or it may be the bond formed by hybridization between a probe and a template strand. Selectively cleavable bonds are not limited to covalent bonds, and can be non-covalent bonds or associations, such as those based on hydrogen bonds, hydrophobic bonds, ionic bonds, pi bond ring stacking interactions, Van der Waals interactions, and the like.
"Rest" is one of two or more parts into which something can be divided, such as, for example, the various parts of an anchor, a molecule or a probe.
ES 2 559 313 T3 "Anchor" or "anchor member" refers to a polymer or molecular construct having a generally linear dimension and with a terminal moiety at each of the two opposite ends. An anchor is attached to a substrate with a bond in at least one terminal moiety to form a substrate construction. The terminal moieties of the anchor can be connected to cleavable bonds with the substrate or cleavable intra-anchor bonds that serve to limit the anchor in a "limited configuration". After synthesizing the daughter strand, each terminal moiety has a terminal bond that is directly or indirectly coupled to other anchors. The coupled anchors comprise the limited Xpandomer which further comprises the daughter strand. The anchors have a "limited configuration" and an "expanded configuration". The limited configuration is found in substrate constructs and in the daughter strand. The limited anchor configuration is the precursor to the expanded configuration, as found in Xpandomer products. The transition from the limited configuration to the expanded configuration results in the cleavage of selectively cleavable bonds that may be within the primary backbone of the daughter strand or intra-anchor bonds. An anchor in a limited configuration is also used when an anchor is added to form the daughter strand after assembly of the "primary backbone". The anchors may optionally comprise one or more reporters or reporter constructs along their length that may encode substrate sequence information. The anchor provides a means of expanding the length of the Xpandomer and thus reducing the linear density of sequence information.
"Anchor constructs" are anchors or anchor precursors composed of one or more anchor segments or other architectural components for assembling anchors such as indicator constructs, or indicator precursors, including polymers, graft copolymers, block copolymers, ligands of affinity, oligomers, haptens, aptamers, dendrimers, linking groups or affinity linking group (eg, biotin).
"Anchor element" or "anchor segment" is a polymer having a generally linear dimension with two terminal ends, in which the ends form terminal bonds to concatenate the anchor elements. The anchoring elements can be segments of anchoring constructions. Such polymers can include, but are not limited to: polyethylene glycols, polyglycols, polypyridines, polyisocyanides, polyisocyanates, poly (triarylmethyl) methacrylates, polyaldehydes, polypyrrolinones, polyureas, polyglycol phosphodiesters, polyacrylates, polymethacrylates, polycarbonylamides, polyvinyl esters, polycarbonates, polyamines, polystyrene, polycarbonates, polyamines, polyvinyl esters, polycarbonates, polyamines, polyvinyl esters, polyvinylphosphonates, polyacetamides, polysaccharides, polyhyaluranates, polyamides, polyimides, polyesters, polyethylenes, polypropylenes, polystyrenes, polycarbonates, polyterephthalates, polysilanes, polyurethanes, polyethers, polyamino acids, polyglycines, polyrolines, N-substituted polylysine, polypeptides, N-substituted peptides in the side chain, N-substituted poly-glycine, peptide-substituted peptides, side chain carboxyl, homopeptides, oligonucleotides, ribonucleic acid oligonucleotides, deoxynucleic acid oligonucleotides, modified oligonucleotides to prevent Watson-Crick base pairing, oligonucleotide analogs, polycytidylic acid, polyadenylic acid, polyuridylic acid, polythymidine, polyphosphate, polynucleotides, polyribonucleotides, polyethylene glycol phosphodiesters, polyethylene glycol aninucleotide analogs, peptide-threinucleotide analogs glycol-polynucleotide analogs, morpholino-polynucleotide analogs, blocked nucleotide oligomer analogs, polypeptide analogs, branched polymers, comb polymers, star polymers, dendritic polymers, random, gradient, and block copolymers, anionic polymers, cationic polymers, stem-loop-forming polymers, rigid segments, and flexible segments.
"Peptide nucleic acid" or "PNA" is a nucleic acid analog having nucleobase residues suitable for hybridization to a nucleic acid, but with a backbone comprising amino acids or derivatives or analogs thereof.
"Phosphono-peptide nucleic acid" or "pPNA" is a peptide nucleic acid in which the backbone comprises amino acid analogs, such as N- (2-hydroxyethyl) phosphonoglycine or N- (2-aminoethyl) phosphonoglycine, and the bonds between Nucleobase units are via phosphonoester or phosphonoamide linkages.
"Serine nucleic acid" or "SerNA" is a peptide nucleic acid in which the backbone comprises serine residues. Such residues can be linked by amide or ester linkages.
"Hydroxyproline nucleic acid" or "HypNA" is a peptide nucleic acid in which the backbone comprises 4-hydroxyproline residues. Such residues can be linked by amide or ester linkages.
"Reporter element" is a signaling element, molecular complex, compound, molecule or atom that also comprises an associated "reporter detection characteristic". Other reporter elements include, but are not limited to, FRET resonant donor or acceptor, dye, quantum dot, bead, dendrimer, over-converting fluorophore, magnet particle, electron scatterer (eg, boron), mass, bead. gold, magnetic resonance, ionizable group, polar group, hydrophobic group. Still others are fluorescent labels, such as, but not limited to, Ethidium Bromide, SYBR Green, Texas Red, Acridine Orange, Pyrene, 4-Nitro-1,8naphthalimide, TOTO-1, YOYO-1, Cyanine 3 ( Cy3), cyanine 5 (Cy5), phycoerythrin, phycocyanin, allophycocyanin, FITC, rhodamine, 5 (6) -carboxyfluorescein, fluorescent proteins, DOXYL (N-oxyl-4,4-dimethyloxazolidine), PROXYL (N-oxyl2,2, 5,5-tetramethylpyrrolidine), TEMPO (N-oxyl-2,2,6,6-tetramethylpiperidine), dinitrophenyl, acridines, coumarins, Cy3 and Cy5 (Biological Detection Systems, Inc.), Erythrosine, Coumaric Acid, Umbelliferone, Texas Red-Rhodamine, Tetramethylhodamine, Rox, 7-Nitrobenzo-1-Oxa-1-Diazole (NBD), Oxazole, Thiazole, Pyrene, Fluorescein or lantamides; also
ES 2 559 313 T3
3 14 35 125 32 131 radioisotopes (such as P, H, C, S, I, P or I), ethidium, europium, ruthenium and samarium or other radioisotopes; or mass labels, such as, for example, C5-modified pyrimidines or N7-modified purines, where the mass modifying groups can be, for example, halogen, ether or polyether, alkyl, ester or polyester, or of the general type XR, where X is a linking group and R is a mass modifying group, chemiluminescent labels, spin tags, enzymes (such as peroxidases, alkaline phosphatases, beta-galactosidases and oxidases), antibody fragments, and affinity ligands (such as an oligomer, hapten, and aptamer). The association of the indicator element with the anchor can be covalent or non-covalent, and direct or indirect. Representative covalent associations include connector and zero connector bonds. Links to the anchor backbone or to an element attached to the anchor such as a dendrimer or side chain are included. Representative non-covalent bonds include hydrogen bonds, hydrophobic bonds, ionic bonds, pi bond ring stacking, Van der Waals interactions, and the like. Ligands, for example, are associated by specific affinity binding with binding sites on the reporter element. Direct association can take place at the time of anchor synthesis, after anchor synthesis, and before or after Xpandomer synthesis.
An "indicator" is made up of one or more indicator elements. Indicators include what are known as "labels" and "marks." The probe or nucleobase residue of the Xpandomer can be considered an indicator. The indicators are used to analyze the genetic information of the target nucleic acid.
"Reporter construction" comprises one or more reporters that can produce a detectable signal (s), wherein the detectable signal (s) generally contain sequence information. This signal information is called the "flag code" and is subsequently decoded into genetic sequence data. An indicator construct can also comprise anchor segments or other architectural components including polymers, graft copolymers, block copolymers, affinity ligands, oligomers, haptens, aptamers, dendrimers, linking groups or affinity linking group (e.g. , biotin).
"Reporter detection characteristic" referred to "signal" describes all possible measurable or detectable elements, properties or characteristics used to communicate the genetic information of a reporter sequences directly or indirectly to a measuring device. These include, but are not limited to, fluorescence, multi-wavelength fluorescence, emission spectrum fluorescence quenching, FRET, emission, absorbance, reflectance, dye emission, quantum dot emission, bead imaging, molecular complexes, magnetic susceptibility, electron scattering, ionic mass, magnetic resonance, molecular complex dimension, molecular complex impedance, molecular charge, induced dipole, impedance, molecular mass, quantum state, charge capacity, magnetic spin state, inducible polarity, nuclear decay, resonance, or complementarity.
"Indicator code" is the genetic information of a measured signal of an indicator construct. The flag code is decoded to provide sequence-specific genetic information data.
"Xsonda" is an expandable oligomeric substrate construct. Each Xsonda has a probe member and an anchor member. The anchor member generally has one or more indicator constructions. Xprobes with 5'-monophosphate modifications are compatible with enzymatic ligation-based methods for the synthesis of Xpandomers. Xprobes with 5 'and 3' linker modifications are compatible with chemical ligation-based methods for the synthesis of Xpandomers.
"Xmero" is an expandable oligomeric substrate construct. Each Xmer has an oligomeric substrate member and an anchor member, the anchor member generally having one or more reporter constructs. Xmers are 5'-triphosphates compatible with polymerase-based methods to synthesize Xpandomers.
"RT-NTP" is a 5 'triphosphate modified expandable nucleotide substrate construct ("monomeric substrate") compatible with template-dependent enzymatic polymerization. An RT-NTP has a modified deoxyribonucleotide triphosphate ("DNTP"), ribonucleotide triphosphate ("RNTP"), or a functionally equivalent analog substrate, collectively referred to as the nucleotide triphosphate substrate ("NTPS"). An RT-NTP has two distinct functional components; specifically, a nucleobase 5'-triphosphate and an anchor or anchor precursor. After daughter strand formation, the anchor binds between each nucleotide at positions that allow controlled RT expansion. In one class of RT-NTP (eg class IX), the anchor is attached after polymerization of RT-NTP. In some cases, RT-NTP has a reversible end terminator and an anchor that selectively crosslinks directly with adjacent anchors. Each anchor can be uniquely encoded with reporters that specifically identify the nucleotide to which it is attached.
"XNTP" is an expandable 5'-triphosphate modified nucleotide substrate compatible with template-dependent enzymatic polymerization. An XNTP has two distinct functional components; specifically, a 5'-triphosphate nucleobase and an anchor or anchor precursor that is attached within each nucleotide at positions that allow expansion of Re controlled by intra-nucleotide cleavage.
ES 2 559 313 T3 "Processive" refers to a coupling process substrate that is generally continuous and proceeds with directionality. Although not bound by theory, both ligases and polymerases, for example, exhibit processive behavior if substrates are added to a nascent daughter strand incrementally without interruption. The hybridization and ligation steps, or hybridization and polymerization, are not considered independent steps if the net effect is the processive growth of the nascent daughter strand. Some, but not all, primer-dependent processes are processive.
"Promiscuous" refers to a substrate coupling process that proceeds from multiple points on a template at once, and is not primer dependent, and indicates that strand expansion occurs in parallel (simultaneously) from more than one point of origin.
"Single base extension" refers to a cyclic staging process in which monomeric substrates are added one at a time. Generally, the coupling reaction is limited to proceeding beyond a single substrate extension in any one step by the use of reversible blocking groups.
"Single probe extension" refers to a cyclic staging process in which oligomeric substrates are added one at a time. Generally, the coupling reaction is limited to proceeding beyond a single substrate extension in any one step by the use of reversible blocking groups.
"Corresponds to" or "corresponding" is used herein to refer to a contiguous single-stranded sequence of a probe, oligonucleotide, oligonucleotide analog, or daughter strand that is complementary, and thus "corresponds to" all or a portion of a sequence. of target nucleic acids. The complementary sequence of a probe can be said to correspond to its target. Unless otherwise stated, both the complementary sequence of the probe and the complementary sequence of the target are individually contiguous sequences.
"Nuclease resistant" refers to is a bond that is resistant to a nuclease enzyme under conditions in which a phosphodiester bond of DNA or RNA will generally cleave. Nuclease enzymes include, but are not limited to, DNase I, exonuclease III, mung bean nuclease, RNase I, and RNase H. One of skill in this field can easily assess the relative nuclease resistance of a given bond.
"Ligase" is an enzyme generally for binding 3'-OH 5'-monophosphate nucleotides, oligomers, and their analogues. Ligaases include, but are not limited to, NAD-dependent ligaments<sup>+</sup> including tRNA ligase, Taq DNA ligase, Thermus filiformis DNA ligase, Escherichia coli DNA ligase, Tth DNA ligase, Thermus scotoductus DNA ligase, thermostable ligase, Ampligase thermostable DNA ligase, VanC-like ligase, 9 ° N DNA ligase, DNA Tsp ligase, and novel ligases discovered by bioprospecting. Ligaases also include, but are not limited to, ATP-dependent ligases including RNA ligase T4, DNA ligase T4, DNA ligase T7, DNA ligase Pfu, DNA ligase I, DNA ligase III, DNA ligase IV, and novel ligases discovered by bioprospecting. These ligases include non-mutant isoforms, mutants, and genetically engineered variants.
"Polymerase" is an enzyme generally for binding 3'-OH 5'-triphosphate nucleotides, oligomers, and their analogues. Polymerases include, but are not limited to, DNA-dependent DNA polymerases, DNA-dependent aRn polymerases, RNA-dependent DNA polymerases, RNA-dependent DNA polymerases, T7 DNA polymerase, T3 DNA polymerase, T4 DNA polymerase, T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, Klenow fragment DNA polymerase I, Thermophilus aquaticus DNA polymerase, Tth DNA polymerase, VentR® DNA polymerase (New England Biolabs), Deep VentR® DNA polymerase (New England Biolabs), Bst DNA polymerase large fragment, Stoeffel fragment, 9 ° N DNA polymerase, 9 ° N DNA polymerase, Pfu DNA polymerase, Tfl DNA polymerase, Tth DNA polymerase, RepliPHI Phi29 polymerase, Tli DNA polymerase, Eukaryotic DNA polymerase beta, telomerase, Therminator ™ polymerase (New England Biolabs), KOD HiFi ™ DNA polymerase (Novagen), KOD1 DNA polymerase, Q-beta replicase, terminal transferase, AMV reverse transcriptase, M-MLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered by bioprospecting, and polymerases cited in US 2007/0048748, US 6,329,178, US 6,602,695 and US 6,395,524. These polymerases include non-mutant isoforms, mutants, and genetically engineered variants.
"Encode" or "analyze" are verbs with reference to transferring from one format to another, and refers to transferring the genetic information from the base sequence of the target template into an array of reporters.
"Extragenetic" refers to any structure on the daughter strand that is not part of the primary backbone; for example, an extragenetic reporter is not the nucleobase itself found in the primary backbone.
"Hetero-copolymer" is a material formed by combining different units (eg, monomer subunit species) into chains of a "copolymer". Hetero-copolymers are formed from discrete "subunit" constructs. A "subunit" is a region of a polymer composed of a well-defined motif, in which each motif is a species and carries genetic information. The term hetero-copolymer is also used herein to describe a polymer in which all of the blocks are building blocks of repeating motifs, each motif having species-specific elements. The daughter strand and the Xpandomer are both hetero-copolymers, whereby each subunit motif encodes 1 or more bases of the target template sequence and the entire target sequence is further defined by the motif sequence.
ES 2 559 313 T3 "Solid support" is a solid material that has a surface for the attachment of molecules, compounds, cells, or other entities. The surface of a solid support can be flat or non-flat. A solid support can be porous or non-porous. A solid support can be a chip or matrix that comprises a surface, and that can comprise glass, silicon, nylon, polymers, plastics, ceramics, or metals. A solid support can also be a membrane, such as a nylon, nitrocellulose, or polymeric membrane, or a plate or disk, and comprise glass, ceramics, metals, or plastics, such as, for example, polystyrene, polypropylene, polycarbonate or polyalomer. A solid support can also be a bead, resin or particle of any shape. Such particles or beads can comprise any suitable material, such as glass or ceramic, and / or one or more polymers, such as, for example, nylon, polytetrafluoroethylene, TEFLON ™, polystyrene, polyacrylamide, Sepaharose, agarose, cellulose, cellulose derivatives , or dextran, and / or may comprise metals, particularly paramagnetic metals, such as iron.
"Reversibly blocking" or "terminator" refers to a chemical group that when attached to a second chemical group in a moiety prevents the second chemical group from entering particular chemical reactions. A wide range of protecting groups are known in synthetic organic and bioorganic chemistry that are suitable for particular chemical groups and are compatible with particular chemical processes, which means that they will protect particular groups during those processes and can be subsequently removed or modified (see, eg Metzker et al. Nucleic Acids Res., 22 (20): 4259, 1994).
"Linker" is a molecule or moiety that joins two molecules or moieties, and provides separation between the two molecules or moieties so that they are capable of functioning in their intended form. For example, a linker may comprise a diamine hydrocarbon chain that is covalently attached via a reactive group on one end to an oligonucleotide analog molecule and via a reactive group on the other end to a solid support, such as, for example, a pearl surface. Coupling of linkers to nucleotides and substrate constructs of interest can be accomplished using coupling reagents that are known in the art (see, for example, Efimov et al., Nucleic Acids Res. 27: 4416-4426, 1999). The methods of derivatization and coupling of organic molecules are well known in the arts of organic and bioorganic chemistry. A connector can also be cleavable or reversible.
Overview
In general terms, methods and devices and corresponding products for replicating single molecule target nucleic acids are described. Such methods use "Xpandomers" that allow sequencing of the target nucleic acid with high throughput and accuracy. An Xpandomer encodes (analyzes) the nucleotide sequence data of the target nucleic acid in a linearly expanded format, thus improving spatial resolution, optionally with amplification of signal intensity. These processes are referred to herein as "expansion sequencing" or "SBX".
As shown in Figure 1A, native duplex nucleic acids have an extremely compact linear data density; approximately 3.4 A center-to-center spacing between sequential stacked bases (2) of each strand of the double helix (1), and thus it is tremendously difficult to directly image or sequence with any accuracy and speed. When the double-stranded form is denatured to single-stranded polynucleotides (3,4), the resulting base-to-base separation distances are similar, but the problem is compounded by domains of the secondary structure.
As shown in Figure 1B, Xpandomer (5), here illustrated as a concatenation of short oligomers (6,7) held together by extragenetic T anchors (8,9), is a synthetic or "surrogate" substitution for the target. nucleic acid to be sequenced. The bases complementary to the template are incorporated into the Xpandomer, but the regularly spaced anchors serve to increase the distance between the short oligomers (here each shown with four nucleobases represented by circles). The Xpandomer is prepared by a process in which a synthetic duplex intermediate is first formed by replicating a template strand. The daughter strand is unique in that it has both a linear backbone made up of the oligomers and a limited Xpandomer backbone comprised of folded anchors. The anchors are then opened or "expanded" to transform the product into a chain of elongated anchors. Figuratively, the daughter strand can be visualized as having two overlapping skeletons: one linear (primary skeleton) and the other with "accordion" folds (Xpandomer limited). Selective cleavage of bonds in the daughter strand allows the accordion folds to expand to produce an Xpandomer product. This process will be explained below in more detail, but it should be noted that the choice of four nucleobases per oligomer and anchor data as shown in Figure 1B is for the purpose of illustration only, and should not be construed to limit the invention in any way. .
The separation distance "D" between neighboring oligomers in the Xpandomer is now a process dependent variant and is determined by the length of the anchor T. As will be shown, the length of the anchor T is designed in the substrate constructions, the structural elements from which the Xpandomer is prepared. The separation distance D can be selected to be, for example, greater than 0.5 nm, or greater than 2 nm, or greater than 5 nm, or greater than 10 nm, or greater than 50 nm. As the separation distance increases, the process of discrimination or "resolution" of the individual oligomers becomes progressively easier. This would also be true if, instead of oligomers, individual nucleobases from another species of Xpandomer
ES 2 559 313 T3 will be strung together on an anchor chain.
Referring back to Figure 1A, native DNA is replicated by a semi-conservative replication process; Each new DNA molecule is a "duplex" of a template strand (3) and a native daughter strand (4). The sequence information is passed from the template to the native daughter strand by a process of "template-directed synthesis" that preserves the genetic information inherent in the base pair sequence. The native daughter strand in turn receives a template for a native daughter strand of the next generation, etc. Xpandomers are formed by a similar template-directed synthesis process, which can be an enzymatic or chemical coupling process. However, unlike native DNA, once formed, Xpandomers cannot replicate by a semi-conservative biological replication process and are not suitable for amplification by processes such as PCR. The Xpandomer product is designed to limit unwanted secondary structure.
Figures 2A to 2D show representative Xpandomer class I (20,21,22,23) substrates. These are the structural elements from which Xpandomers are synthesized. Other Xpandomer substrates (ten classes are disclosed herein) are discussed in later sections. The Xpandomer substrate constructs shown here have two functional components; specifically, a probe member (10) and an "anchor" member (11) in a loop configuration. The loop forms the elongated "T" anchor of the final product. For convenience of explanation only, the probe member is again represented with four nucleobase residues (14,15,16,17) as shown in Figure 2B.
These substrate constructs can be end-modified with R groups, for example, as a 5'-monophosphate, 3'-OH suitable for use with a ligase (referred to herein as "Xsonda") or as a 5'- triphosphate, 3'-OH suitable for use with a polymerase (referred to herein as an "Xmer"). Other R groups can be of use in various protocols. In the first example shown in Figure 3B, the present inventors present the use of Xprobes in the synthesis of an Xpandomer from a template strand of a target nucleic acid by a ligase-dependent process.
The four nucleobase residues (14,15,16,17) of the probe member (10) are selected to be complementary to a contiguous four nucleotide sequence of the template. Each "probe" is thus designed to hybridize to the template in a complementary four nucleotide sequence. By supplying a library of many such probe sequences a contiguous complementary replica of the template can be formed. This daughter strand is called an "Xpandomer intermediate". Xpandomer intermediates are in duplex or single chain forms.
The anchor loop is attached to the probe member (10) at the second and third nucleobase residues (15,16). The second and third nucleobase residues (15,16) are also linked together by a "selectively cleavable bond" (25) represented by a "V". Cleavage of this cleavable bond allows the anchor loop to expand. The linearized anchor can be said to "connect" the selectively cleavable binding site of the primary polynucleotide backbone of a daughter strand. The cleavage of these bonds breaks the primary backbone and forms the longest Xpandomer.
Selective cleavage of the selectively cleavable linkages (25) can be done in a variety of ways including, but not limited to, chemical cleavage of phosphorothiolate linkages, ribonuclease digestion of ribosyl 5'-3 'phosphodiester linkages, cleavage of photo-cleavable linkages , and the like, as discussed in greater detail below.
The substrate construction (20) shown in Figure 2A has a single anchor segment, represented here by an ellipse (26), for attachment of indicator elements. This segment is flanked with spacer anchor segments (12,13), all of which together form the anchor construction. In this regard, one to many dendrimer (s), polymer (s), branched polymer (s) or combinations may be used, for example, to construct the anchor segment. For the substrate construction (21) of Figure 2B, the anchoring construction is composed of three anchoring segments for joining indicator elements (27,28,29), each of which is flanked with an anchoring segment spacer. The combination of reporter elements together form a "reporter construct" to produce a unique digital reporter code (for identification of probe sequences). These reporter elements include, but are not limited to, fluorophores, FRET tags, beads, ligands, aptamers, peptides, haptens, oligomers, polynucleotides, dendrimers, stem-loop structures, affinity tags, mass tags, and the like. The anchor loop (11) of the substrate construction (22) in Figure 2C is "bare". The genetic information encoded in this construct is not encoded on the anchor, but is associated with the probe (10), for example, in the form of labeled nucleotides. The substrate construct (23) of Figure 2D illustrates the general principle: as indicated by the asterisk (*), the probe sequence information is encoded or "parsed" in the substrate construct in a more modified way. easily detected in a sequencing protocol. Because the sequence data is better physically resolved after cleavage of the selectively cleavable bond (25) to form the linearly elongated Xpandomer polymer, the asterisk (*) represents any form of encoded genetic information for which this is a benefit. . The bioinformatic element or elements (*) of the substrate construction, whatever their shape, may be directly detectable or they may be precursors to which the detectable elements are added in a post-assembly marking step. In some cases, the
ES 2 559 313 T3 genetic information is encoded in a molecular property of the substrate construct itself, eg, a multi-state mass tag. In other cases, the genetic information is encoded by one or more fluorophores from the donor pairs: FRET acceptor, or a nanomolecular barcode, or a ligand or combination of ligands, or in the form of some other labeling technique extracted from The matter. Various embodiments will be discussed below in more detail.
The anchor generally serves several functions: (1) to link sequentially, directly or indirectly, to adjacent anchors that form the Xpandomer intermediate; (2) stretching and expanding to form an elongated chain of anchors upon cleavage of selected bonds in the primary backbone or within the anchor (see Figure 1B); and / or (3) provide a molecular construct to incorporate reporter elements, also called "tags" or "tags," that encode the sequence information of the nucleobase residue of its associated substrate. The anchor can be designed to optimize coding function by adjusting spatial gaps, abundance, informational density, and signal intensity of its constituent reporter elements. A wide range of reporter properties is useful for amplifying the signal intensity of the encoded genetic information within the substrate construct. The literature directed at reporters, molecular barcodes, affinity binding, molecular marking, and other reporter elements, is well known to a person skilled in the art.
It can be seen that if each substrate of a substrate construct contains x nucleobases, then a library representing all possible sequential combinations of x nucleobases would contain 4<sup>x</sup> probes (when nucleobases A, T, C or G are selected). If other bases are used, fewer or more combinations may be needed. These substrate libraries are designed so that each substrate construct contains (1) a probe (or at least one nucleobase residue) complementary to any one of the possible target sequences of the nucleic acid to be sequenced and (2) a construct A unique reporter encoding the identity of the target sequence to which that particular probe (or nucleobase) is complementary. A library of probes containing two nucleobases would have 16 unique members; a probe library containing three nucleobases would have 64 unique members, and so on. A representative library would have all four individual nucleobases, but configured to accommodate an anchoring medium.
The synthesis of an Xpandomer is illustrated in Figures 3A to 3C. The substrate represented here is an Xprobe and the method can be described as hybridization with primer-dependent processive ligation in free solution.
Many well-known molecular biology protocols, such as protocols for fragmenting target DNA and ligating end adapters, can be adapted for use in sequencing methods and are used herein to prepare target DNA (30) for sequencing. Here, the present inventors illustrate, in broad terms, what would be familiar to those skilled in the art, processes for polishing the ends of fragments and ligation of blunt ends of adapters (31,32) designed for use with sequencing primers. . These actions are shown in stage I of Figure 3A. In steps II and III, the target nucleic acid is denatured and hybridized with suitable primers (33) complementary to the adapters.
In Figure 3B, the primed template strand from step III is contacted with a library of substrate (36) and ligase (L) constructs, and in step IV conditions are adjusted to favor hybridization, followed by ligation. in a free 3'-OH of a primer-template duplex. Optionally, in step V, the ligase dissociates, and in steps VI and VII it can be recognized that the hybridization and ligation process produces extension by cumulative addition of substrates (37,38) to the end of the primer. Although priming can occur from adapters at both ends of a single stranded template, growth of a nascent Xpandomer daughter strand is shown here to proceed from a single primer, for simplicity only. The extension of the daughter strand is represented in stages VI and VII, which are repeated continuously (incrementally, without interruption). These reactions occur in free solution and proceed until a sufficient amount of product has been synthesized. In step VIII the formation of a completed Xpandomer intermediate (39) is shown.
Relatively long lengths of the contiguous nucleotide sequence can be efficiently replicated in this way to form Xpandomer intermediates. It can be seen that continuous read lengths ("contigs") corresponding to long template strand fragments can be achieved with this technology. It will be apparent to one of ordinary skill in the art that billions of these single molecule SBX reactions can be done simultaneously in an efficient batch process in a single tube. Subsequently, the random sequencing products of these syntheses can be sequenced.
The next steps of the SBX process are depicted in Figure 3C. Step IX shows denaturation of the duplex Xpandomer intermediate, followed by cleavage of selectively cleavable bonds in the backbone, with the selectively cleavable bonds designed so that the anchor loops "open", forming the Xpandomer product linearly. elongated (34). Such selective cleavage can be accomplished by any number of techniques known to one of ordinary skill in the art, including, but not limited to, cleavage of phosphorothiolate with metal cations as disclosed by Mag et al. ("Synthesis and selective cleavage of an oligodeoxynucleotide containing a bridged internucleotide 5'-phosphorothioate linkage", Nucleic Acids Research 19 (7): 1437-1441, 1991), acid-catalyzed cleavage of phosphoramidate as disclosed by Mag et al. ("Synthesis and selective cleavage of oligodeoxyribonucleotides containing non-chiral internucleotide phosphoramidate linkages",
ES 2 559 313 T3
Nucleic Acids Research 17 (15): 5973-5988, 1989), selective nuclease cleavage of phosphodiester bonds as disclosed by Gut et al. ("A novel procedure for efficient genotyping of single nucleotide polymorphisms", Nucleic Acids Research 28 (5): E13, 2000) and separately by Eckstein et al. ("Inhibition of restriction endonuclease hydrolysis by phosphorothioate-containing DNA", Nucleic Acids Research, 25; 17 (22): 9495, 1989) and selective cleavage of the photo-cleavable linker modified phosphodiester backbone as disclosed by Sauer et al. ("MALDI mass spectrometry analysis of single nucleotide polymorphisms by photocleavage and charge-tagging", Nucleic Acids Research 31,11 e63, 2003), Vallone et al. ("Genotyping SNPs using a UV-photocleavable oligonucleotide in MALDITOF MS", Methods Mol. Bio. 297: 169-78, 2005) and Ordoukhanian et al. ("Design and synthesis of a versatile photocleavable DNA building block, application to phototriggered hybridization", J. Am. Chem. Soc. 117, 9570-9571, 1995).
Basic process refinements, such as the washing steps and stringent setting, are well within the expertise of an experienced molecular biologist. Variations to this process include, for example, immobilization and analysis of the target strands, stretching and other techniques to reduce the secondary structure during Xpandomer synthesis, post-expansion labeling, end-functionalization, and alternatives to ligase to bind the substrates. will be discussed in the materials that follow.
The synthesis of Xpandomers is done to facilitate nucleic acid detection and sequencing, and is applicable to nucleic acids of all types. The process is a method of "expanding" or "lengthening" the length of backbone elements (or subunits) that encode the sequence information (expanded with respect to small nucleotide-to-nucleotide distances of native nucleic acids) and optionally also serves to increase signal intensity (relative to the almost indistinguishable low intensity signals observed for native nucleotides). As such, reporter elements incorporated into the expanded synthetic backbone of an Xpandomer can be detected and processed using a variety of detection methods, including detection methods well known in the art (e.g., a CCD camera, a force microscope atomic, or a regulated mass spectrometer), as well as by methods such as a massively parallel nanopore sensor array, or a combination of methods. Detection techniques are selected based on the optimal signal with respect to noise, performance, cost, and the like.
Returning to Figure 4, a simple model of a sensing technology is shown; specifically, a nanopore (40) with a FRET donor (42) in a membrane (44), which is excited by light of wavelength λ · ι. As the Xpandomer product (41) elongates and is transported through the nanopore (40) in the direction of the arrow (45), serial bursts of emission of wavelength λ2 of excited fluorophores are detected, in the vicinity of the pore. The emission wavelengths (λ<sub>2</sub>) are temporarily separated based on the length of the anchor and the speed of the Xpandomer passing through the nanopore. By capturing these analog signals and processing them digitally, the sequence information can be read directly from the Xpandomer. It should be noted that in this detection method, the nanopore and the membrane can have many paths through which the Xpandomer can translocate. FRET detection requires that there be at least one excited FRET donor along each path. In contrast, a Coulter counter-based nanopore may only have additional translocation hole at the cost of the signal relative to noise.
In the nanopore-based detection technique of Figure 4, representing Xprobe chains of the structure shown in Figure 2B, the anchor constructs contain multi-element reporter constructs, as indicated by the box-like reporter members (27, 28,29) arranged along the anchor. Relevant nanopore sequencing technology is disclosed, for example, by Branton et al. in US Patent No. 6,627,067 and by Lee et al. (Lee, JW and A Meller. 2007. Rapid sequencing by direct nanoscale reading of nucleotide bases in individual DNA chains. In "New High Throughput Technologies for DNA Sequencing and Genomics, 2", Elsevier).
Figure 5 demonstrates how multi-element reporter constructs, herein comprising FRET acceptor fluorophores, appear to a detector positioned at the FRET gate. It can be seen in this multi-channel representation of broadcasts that analog signals are generated at generally regular intervals of time and can be analyzed as a type of digital code (here called an indicator code and, for this example, an Xsonda ID) that reveals the identity and order of the Xsonda subunits and thus the genetic sequence of the illustrated Xpandomer. Various combinations of reporters can be used to create a library of reporter codes sequentially encoding any 4-base combination of A, T, G, or C of the described Xprobe. In this example, combinations of three fluorophores are used to produce twenty-two reporter codes. In this way, it is observed that the sequence ACTG is followed by GCCG; followed by AAAT. Vertically arranged dotted lines separate the fluorimetry data and the corresponding Xpandomer subunits (shown schematically). An interpretive algorithm immediately below the representation shows how regularly separated analog signals are transformed into a readable genetic sequence.
The Xsonda substrate construct illustrated in Figure 5, which uses a multi-element anchor construct composed of three reporter-labeled segments, each flanked with spacer anchor segments, to encode substrate sequence identity , is further elaborated. The first anchor segment is indicator code # 1 (reading from left to right), and it reads as a high signal in the red channel. The second anchor segment is indicator code # 9, and it reads as a high green signal and
ES 2 559 313 T3 a low red signal. The third anchor segment is indicator code # 8, and it reads as a low blue signal and a low red signal. Indicator code No. 1 is used as a clock or timing signal; Flag code # 9 encodes the first probe residue "AC"; Reporter code # 8 encodes the second "TG" residue of the probe. Taken together, the sequence reporter code of "1-9-8" corresponds to a particular species of Xsonda (Xsonda ID 117), which in turn corresponds by design to the sequence fragment "ACTG". Three Xsonda IDs encode the entire contiguous sequence shown in the representation, "ACTGGCCGAAAT". The fluorophore emissions, the table for decoding reporter codes and sequence fragments, and the corresponding physical representations of the reporter constructs, are separated by the dotted lines of the figures according to structural subunits of the Xpandomer so that it can be easily observed. how sequence information is decoded and digitized.
Figure 6 is a table of fluorophore labels from which the example of Figure 5 was prepared. This more generally illustrates the use of multi-state indicator code combinations to analyze information in the form of detectable signals. Fluorophores having twenty-two possible emission states are used to form the reporter constructs of this example. Three fluorophore tags per oligomer are more than adequate to encode all possible 4-mer combinations of A, T, C, and G. Increasing the length of the anchor improves the resolution between emissions of fluorophore tags, benefiting the accuracy of the matching step. detection, a principle that is generally applicable.
Useful reporters with anchor constructs of this type are of many types, not simply fluorophores, and can be measured using a corresponding wide range of high-throughput and accurate detection technologies, technologies that might not otherwise be useful for sequencing native nucleic acids due to at limited resolution. Massively parallel state-of-the-art detection methods, such as nanopore sensor arrays, are facilitated by the more measurable characteristics of Xpandomers. Inefficiencies in sequencing detection processes can be reduced by pre-purifying batches of Xpandomers to remove incomplete or short reaction products. Methods for end-modification of synthesized Xpandomers are provided that can be used for both purification and as a means of facilitating presentation of Xpandomers to the detector. Furthermore, the reading process is not limited by limitation to hooded, uncapped, nucleotide extension, marking, or other simultaneous processing methods.
Figure 7A depicts a partial duplex template designed with a twenty base 5 'overhang nucleotide to demonstrate processive substrate ligation and primer initiated template directed ligation in free solution. Figure 7B is a photograph of a gel demonstrating ligation of substrates using the primer-template format described in Figure 7A. For this example, dinucleotide oligomeric substrates of the sequence 5 'phosphate CA 3' hybridize to the template in the presence of a primer and T4 DNA ligase. The non-duplexed end overhang (if any) is then digested with nuclease and the ligation products are separated on a 20% acrylamide gel. Ligation produces product polymers containing demonstrably bound subunits. As indicated by the banding pattern, the ligase-positive reactions running in lanes 1, 3, 5, 7, and 9, containing progressively longer templates (4, 8, 12, 16, and 20 bases, respectively), clearly demonstrate the sequential ligation of 2mer substrates (high lengths of exonuclease-protected duplexes). Lanes 2, 4, 6, 8 and 10 are negative controls that do not contain ligase and show complete exonuclease digestion of unligated products.
Figure 7C is a second gel showing template directed ligation of substrates. Four progressively longer positive control templates were tested, again duplexed with an extension primer (4, 8, 12, and 16 template bases, respectively). Again, oligomeric dinucleotide substrates of the sequence 5 'phosphate CA 3' hybridize to the template in the presence of a primer and T4 DNA ligase. The non-duplexed end overhang (if any) is then digested with nuclease and the ligation products are separated on a 20% acrylamide gel. It is observed that the oligomeric substrates (again 2mer) bind to the template in lanes 1, 2, 3 and 4, but not in lanes 5 and 6, where the template strands contain a mismatch with the 5 'dinucleotide ( phosphate) CA 3 '(lane 5-5' CGcG 3 'template; lane 6-5' GGGG 3 'template).
The gel results shown in Figure 7D demonstrate multiple template-directed ligation of a bis (amino) modified tetranucleotide probe. The aliphatic amino modifiers were of the binding and composition described in Figure 26. For this example, a 5 '(phosphate) C (amino) A (amino) CA 3' sequence tetranucleotide oligomeric substrate hybridized to a variety of progressively longer complementary templates (duplexed with an extension primer) in the presence of a primer and T4 DNA ligase. The nucleotide overhang from the end if duplexing (if any) was then nuclease digested and the ligation products separated on a 20% acrylamide gel. Ligation produces product polymers containing demonstrably bound subunits. Lanes 1 and 2 represent 16mer and 20mer size controls. Lanes 3, 4, 5, 6, 7, 8, and 9 show ligation products for progressively longer complementary templates (4, 6, 8, 12, 16, 18, and 20 template bases, respectively). Multiple tetramer ligations are observed for longer template reactions (lanes 6-9). Lane 10 shows essentially complete ligase inhibition due to template-probe mismatch (template - 5 'CGCG 3').
ES 2 559 313 T3
The gel results shown in Figure 7E demonstrate multiple template-directed ligation of a bis (amino) modified hexanucleotide probe. The aliphatic amino modifiers were of the binding and composition described in Figure 26. For this example, a hexanucleotide oligomeric substrate of the sequence 5 '(phosphate) CA (amino) C (amino) ACA 3' hybridized to a variety of progressively longer complementary templates (duplexed with an extension primer) in the presence of a primer and t4 DNA ligase. The nucleotide overhang from the end if duplexing (if any) was then nuclease digested and the ligation products separated on a 20% acrylamide gel. Ligation produces product polymers containing demonstrably bound subunits. Lanes 1 and 2 represent 16mer and 20mer size controls. Lanes 3, 4, 5, 6, 7, 8, and 9 show ligation products for progressively longer complementary templates (4, 6, 8, 12, 16, 18, and 20 template bases, respectively). Multiple tetramer ligations are observed for longer template reactions (lanes 5-9). Lane 10 shows almost complete ligase inhibition due to probe template mismatch (template - 5 'CGCGCG 3').
Substrates include both probe members (ie, oligomers and template-specific binding members for assembling the Xpandomer intermediate) and monomers (ie, individual nucleobase members as template-specific binding members). The present inventors call the first substrates "probe type" and the second substrates "monomer type". As illustrated in Figure 8, probe-like Xpandomers have five basic subgenera, while Figure 9 illustrates five basic subgenera of monomer-like Xpandomers. The tables in Figures 8 and 9 include three columns: the first describing substrate constructs, the second Xpandomer intermediates, and the third the characteristic Xpandomer products of the subgenus (per row). The tables are provided here as an overview, with methods doing and using the same things that are disclosed in greater detail herein below. In Figures 8 and 9, "P" refers to a probe member, "T" to an anchor member (or loop anchor or anchor arm precursor), "N" to a monomer (a single nucleobase or nucleobase residue) and "R" to a terminal group.
More specifically, the table in Figure 8 uses the following nomenclature:
P is a probe substrate member and is composed of P<sup>1</sup>-P<sup>2</sup>, in which P<sup>1</sup> is a first probe residue and P<sup>2</sup> is a second probe moiety;
T is an anchor;
The brackets indicate a subunit of the daughter strand, in which each subunit is a subunit motif that has a species-specific probe member, further in which said probe members of said subunit motifs are sequentially complementary to the sequence of Corresponding contiguous nucleotides of the template strand, here denoted P<sup>1</sup> - P<sup>2</sup>, and form a primary backbone of the Xpandomer intermediate, and wherein the anchoring members, optionally in combination with the probe moieties, form a limited Xpandomer backbone. The cleavage of one or more selectively cleavable bonds within the Xpandomer intermediate allows expansion of the subunits to produce an Xpandomer product, the subunits of which are also indicated by square brackets;
α indicates a species of subunit motif selected from a library of subunit motifs;
s is a first connecting group attached to a first end or residue of a probe or anchor member; under controlled conditions, s is capable of selectively reacting with, directly or via cross-linkers, the δ connecting group of a confined end of an adjacent subunit to form covalent or equivalently durable bonds;
δ is a second connecting group attached to a first end or residue of a probe or anchor member; under controlled conditions, δ is capable of selectively reacting with, directly or via cross-linkers, the ε connecting group from a confined end of an adjacent subunit to form covalent or equivalently durable bonds;
χ represents a bond with an adjacent subunit and is the bond product of the reaction of the connecting groups δ and ε;
~ indicates a selectively cleavable bond, which may be the same or different when multiple selectively cleavable bonds are present;
R<sup>1</sup> includes, but is not limited to, hydroxyl, hydrogen, triphosphate, monophosphate, ester, ether, glycol, amine, amide, and thioester;
R<sup>2</sup> includes, but is not limited to, hydroxyl, hydrogen, triphosphate, monophosphate, ester, ether, glycol, amine, amide, and thioester; and κ indicates subunit k<sup>n</sup> in a chain of m subunits, where κ = 1, 2, ... am, where m> 3, and generally m> 20, and preferably m> 50, and more preferably m> 1000.
More specifically, and in the context of the table in Figure 9, the following nomenclature is used:
N is a nucleobase residue;
T is an anchor;
The brackets indicate a subunit of the daughter strand, in which each subunit is a subunit motif that has a species-specific nucleobase residue, additionally in which said nucleobase residues of said subunit motifs are sequentially complementary to the sequence of
ES 2 559 313 T3 corresponding contiguous nucleotides of the template strand, denoted herein N, and form a primary backbone of the Xpandomer intermediate, and wherein the anchor members, optionally in combination with the nucleobase residues, form a backbone of Limited Xpandomer. Cleavage of one or more selectively cleavable bonds within the Xpandomer intermediate allows expansion of the subunits to produce an Xpandomer product, the subunits of which are also indicated by square brackets; n<sup>1</sup> is a first portion of a nucleobase residue;
n<sup>2</sup> is a second portion of a nucleobase residue;
ε is a first connecting group attached to a first end or residue of a probe or anchor member; under controlled conditions, ε is capable of selectively reacting with, directly or via cross-linkers, the δ connecting group of a confined end of an adjacent subunit to form covalent or equivalently durable bonds;
δ is a second connecting group attached to a first end or residue of a probe or anchor member; under controlled conditions, δ is capable of selectively reacting with, directly or via cross-linkers, the ε connecting group from a confined end of an adjacent subunit to form covalent or equivalently durable bonds;
χ represents a bond with an adjacent subunit and is the bond product of the reaction of bonding groups δ and ε;
χ<sup>1</sup> is the bond product of the reaction of bonding groups δ<sup>1</sup> and ε<sup>1</sup>; χ<sup>2</sup> is the bond product of the reaction of bonding groups δ<sup>2</sup> and ε<sup>2</sup>;
~ indicates a selectively cleavable bond, which may be the same or different when multiple selectively cleavable bonds are present;
R<sup>1</sup> includes, but is not limited to, hydroxyl, hydrogen, triphosphate, monophosphate, ester, ether, glycol, amine, amide, and thioester;
R<sup>2</sup> includes, but is not limited to, hydroxyl, hydrogen, triphosphate, monophosphate, ester, ether, glycol, amine, amide, and thioester; and κ indicates subunit k<sup>n</sup> in a chain of m subunits, in which κ = 1, 2, ... am, in which m> 10, and generally m> 50, and usually m> 500 or> 5,000.
Oligomeric constructions
Xpandomer precursors and constructs can be divided into two categories based on the substrate (oligomeric or monomeric) used for template-directed assembly. The structure, precursors and synthesis methods of the Xpandomer for those based on the oligomer substrates are discussed below.
Substrate constructs are reactive precursors for the Xpandomer and generally have an anchoring member and a substrate. The substrate treated herein is an oligomer or probe substrate, generally made up of a plurality of nucleobase residues. Generating combinatorial-type libraries of two to twenty nucleobase residues per probe, generally 2 to 10 and usually 2, 3, 4, 5 or 6 nucleobase residues per probe, useful probe libraries are generated as reagents in the synthesis of precursors of Xpandomers (substrate constructs).
The probe is generally described as having two probe moieties, P<sup>1</sup> And p<sup>2</sup>. These probe residues are generally represented in the figures as dinucleotides, but in general P<sup>1</sup> And p<sup>2</sup> each have at least one nucleobase residue. In the example of a probe with two nucleobase residues, the probe residues P<sup>1</sup> And p<sup>2 </sup>they would be individual nucleobase residues. The number of nucleobase residues for each is chosen, appropriately, for the Xpandomer synthesis method and may not be the same in P<sup>1</sup> And p<sup>2</sup>.
For substrate constructs in which ε and δ linker groups are used to create intersubunit bonds, a wide range of suitable commercially available chemistries (Pierce, Thermo Fisher Scientific, USA) can be adapted for this purpose. Common linker chemistries include, for example, NHS esters with amines, maleimides with sulfhydryls, imidoesters with amines, EDC with carboxyls for reactions with amines, pyridyl disulfides with sulfhydryls, and the like. Other embodiments involve the use of functional groups such as hydrazide (HZ) and 4-formylbenzoate (4FB) which can then be further reacted to form bonds. More specifically, a wide range of crosslinkers (hetero- and homo-bifunctional) are widely available (Pierce) including, but not limited to, sulfo-SMCC (sulfosuccinimidyl 4- [N-maleimidomethyl] cyclohexane-1-carboxylate), SIA (iodoacetate N-succinimidyl), sulfo-EMCS ([N-emaleimidocaproyloxy] sulfosuccinimide ester), sulfo-GMBS (N- [g-maleimidobutyryloxy] sulfosuccinimide ester), AMAS (N- (a-maleimidoacetoxy) succinimide ester), BMPS (N EMCA acid (Ne-maleimidocaproic) -ester of [βmaleimidopropyloxy] succinimide), EDC (1-ethyl-3- [3-dimethylaminopropyl] carbodiimide hydrochloride), SANPAH (Nsuccinimidyl-6- [4'-azido-2 '-nitrophenylamino] hexanoate), SADP (N-succinimidyl (4-azidophenyl) -1,3'-dithiopropionate), PMPI (N- [p-Maleimidophenyl] isocyanate, BMPH (N- [p-maleimidopropionic acid hydrazide], trifluoroacetic acid salt)), EMCH ([Ne-maleimidocaproic acid hydrazide], trifluoroacetic acid salt), SANH (succinimidyl 4-hydrazinonicotinateacetonehydrazone), SHTH (succinimidyl 4-hydrazidoterephthalate hydrochloride) and C6-SFB (C6-succinimidyl 4-formylbenzoate). Therefore, the method disclosed by Letsinger et al. ("Phosphorothioate oligonucleotides having modified internucleoside linkages", US Patent No. 6,242,589) can be adapted to form phosphorothiolate linkages.
ES 2 559 313 T3
In addition, well-established protection / deprotection chemistries are widely available for common linker moieties (Benoiton, "Chemistry of Peptide Synthesis", CRC Press, 2005). Amino protection includes, but is not limited to, 9-fluorenylmethyl carbamate (Fmoc-NRR '), t-butyl carbamate (Boc-NRR'), benzyl carbamate (Z-NRR ', Cbz-NRR') , acetamide-trifluoroacetamide, phthalimide, benzylamine (Bn-NRR '), triphenylmethylamine (TrNRR') and benzylidene-p-toluenesulfonamide (Ts-NRR '). Carboxyl protection includes, but is not limited to, methyl ester, t-butyl ester, benzyl ester, st-butyl ester, and 2-alkyl-1,3-oxazoline. Carbonyl include, but are not limited to, 1,3-dimethylacetal dioxane, and 1,3-dithiano N, N-dimethylhydrazone. Hydroxyl protection includes, but is not limited to, methoxymethyl ether (MOM-OR), tetrahydropyranyl ether (THP-OR), t-butyl ether, allyl ether, benzyl ether (Bn-OR), t-butyldimethylsilyl ether (TbDMS- OR), t-butyldiphenylsilyl ether (TBDPS-OR), acetic acid ester, pivalic acid ester and benzoic acid ester.
Although the anchor is frequently represented as a reporter construct with three reporter groups, various reporter configurations can be displayed on the anchor, and may comprise individual indicators that identify probe constituents, individual indicators that identify probe species, molecular barcodes that identify probe species, or the anchor may be bare polymer (lacking reporters). In the case of the bare polymer, the indicators can be the probe itself, or they can be on a second anchor attached to the probe. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
As discussed above, Figure 8 provides an overview of oligomeric constructs of the invention, distinguishing five classes: classes I, II, III, IV, and V. These classes apply to both Xprobes and Xmeros. Each class will be covered individually below.
Class I oligomeric constructions
Returning to Figure 10, class I oligomeric constructs are described in more detail. Figures 10A to 10C employ adapted notation to show these molecules as substrates and as heterocopolymer products of the SBX process. The figures are read from left to right, showing first the probe-substrate construct (Xpandomer oligomeric precursor), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product prepared for sequencing.
As shown in Figure 10A, a class I substrate construct has an oligomeric probe member (-P<sup>1</sup>~ P<sup>2</sup>-) (100) and an anchoring member, T (99). The anchor is attached by two terminal bonds (108,109) to P probe moieties<sup>1</sup> And p<sup>2</sup>. These limitations prevent the anchor from elongating or expanding and thus in a limited configuration. Under mold-directed assembly, the substrates duplex with the target template so that the substrates are confined.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-Oh, they would find use in a ligation protocol, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with terminal linker groups for chemical coupling or with non-linker groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
The tilde (~) in Figure 10A and 10B indicates a selectively cleavable bond that separates two residues from the probe member. The terminal bonds of the anchor are attached to the two residues of the probe member that are separated by the selectively cleavable bond. The anchor links the first probe moiety to the second probe moiety, forming a loop connecting the selectively cleavable bond. When the probe member is intact (uncleaved), the probe member can bind with high fidelity to the template sequence and the anchor loops in the "limited configuration." When this link is split, the anchor loop can be opened and the anchor is in the "expanded configuration."
Substrate constructs are reagents used for the template-dependent assembly of a daughter strand, which is an intermediate composition for producing Xpandomers. Figure 10B shows the daughter strand of the duplex, a hetero-copolymer with repeating subunits (shown in brackets). The primary skeleton of the daughter strand (-P<sup>1</sup>~ P<sup>2</sup>-) and the target template strand (-P<sup>1</sup>- P<sup>2</sup>-) as a duplex (95). Each subunit of the daughter strand is a repeat motif composed of a probe member and an anchor member, T (99), the anchor member in limited configuration. The motifs have species-specific variability, indicated here by the superscript "a". Each particular subunit on the daughter strand is selected from a library of motifs by a template-directed process and its probe binds to a corresponding sequence of complementary nucleotides on the template strand. In this way, the nucleobase residue sequence of the probes forms a contiguous complementary copy of the female target template.
The daughter strand is composed of an Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors (99) become your
ES 2 559 313 T3 "expanded configuration", the limited Xpandomer is converted to the Xpandomer product.
It can be seen that the daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer" backbone. The primary backbone is composed of contiguously confined probe substrates. The "limited Xpandomer backbone" prevents selectively cleavable bond between P probe residues<sup>1</sup> And p<sup>2</sup> and is formed by linked backbone residues, each backbone residue being a linear linkage of P<sup>1</sup> to the anchor at P<sup>2</sup>, and in which P<sup>2 </sup>can be additionally linked with the P<sup>1</sup> of the next skeleton remainder. It can be seen that the limited Xpandomer backbone connects or loops over the selectively cleavable bonds of the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone fragments.
Figure 10C is a representation of the Xpandomer class I product after dissociation from the template strand and after cleavage of the selectively cleavable bonds from the primary backbone. Methods for template strand dissociation include heat denaturation, or selective digestion with a nuclease, or chemical degradation. The Xpandomer product strand contains a plurality of κ subunits, in which κ indicates the κ subunit<sup>11</sup> in a chain of m subunits constituting the daughter strand, where κ = 1.2, 3 am, where m> 3, and generally m> 20, and preferably m> 50, and more preferably m> 1000. Each subunit is made up of an anchor (99) and P probe residues<sup>1</sup> And p<sup>2</sup>. The anchoring member T, now in "expanded configuration", is seen stretched to its length between the cleaved probe residues P<sup>1</sup> And p<sup>2</sup>, which remain covalently bound to adjacent subunits. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 10D shows the substrate construction of Figure 10A as a molecular model, in which the probe member (100) is arbitrarily represented with four nucleobase residues (101,102,103,104), two of which (102,103) bind to the anchor ( 99) by terminal links (108,109). Between the two terminal bonds of the anchor is a selectively cleavable bond, shown as "V" (110) on the probe member (100). This bond binds P probe residues<sup>1</sup> And p<sup>2</sup> referred to in Figure 10A. The anchor loop shown here has three indicators (105,106,107), which can also be species-specific for motifs.
Figure 10E shows the product Xpandomer after cleavage of the selectively cleavable bonds in the substrate. Cleavage results in limited Xpandomer expansion and is indicated by "E" (dark arrows). Residues (110a, 110b) of the selectively cleavable bond mark the cleavage event. The subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 10C.
In the Xpandomer product (Figure 10E) the primary backbone is now fragmented and not covalently intact because the probe members have cleaved, separating each P<sup>1</sup> (92) and P<sup>2</sup> (94). During the cleavage process, the limited Xpandomer is released to become the Xpandomer product. The Xpandomer includes each concatenated subunit in the sequence. Bound within each subunit are the remainder of the P probe<sup>1</sup>, anchor and rest of probe P<sup>2</sup>. The Xpandomer anchor members (99) that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information from the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Figure 11 depicts a condensed scheme of a method of preparation of one embodiment of a class I Xpandomer; The method illustrates the preparation and use of substrates and products shown in Figures 10D and 10E. The method can be performed in free solution and is described using a ligase (L) to covalently couple confined Xprobes. Methods of relieving the secondary structure in the mold are covered in a later section. Tailored conditions for hybridization and ligation are well known in the art and such conditions can easily be optimized by one skilled in the art.
Ligaases include, but are not limited to, NAD-dependent ligaments<sup>+</sup> including tRNA ligase, Taq DNA ligase, Thermus filiformis DNA ligase, Escherichia coli DNA ligase, Tth DNA ligase, Thermus scotoductus DNA ligase, thermostable ligase, Ampligase thermostable DNA ligase, VanC type ligase, 9 ° N DNA ligase, DNA Tsp ligase, and novel ligases discovered by bioprospecting. Ligaases also include, but are not limited to, ATP-dependent ligases including RNA ligase T4, DNA ligase T4, DNA ligase T7, DNA ligase Pfu, DNA ligase I, DNA ligase III, DNA ligase IV, and novel ligases discovered by bioprospecting. These ligases include non-mutant isoforms, mutants, and genetically engineered variants.
Referring to Figure 11, and in preparation for synthesis, a target nucleic acid (110) is provided and the ends are polished in preparation for blunt end ligation of adapters. Step I shows the ligation of hairpin primers (120) with the target nucleic acid. The free 5 'end of the primers is blocked with a movable blocking group (119). The primers will prime both strands of the target nucleic acid. Adapters are generally added in excess. The locking groups on the hot ends of the
ES 2 559 313 T3 primers are removed in step II, and the two template strands are separated by denaturation. In step III, the primed single-stranded template (111) is contacted with a library of substrate constructs (as represented by construct (112) for the purpose of illustration) and with ligase, L, under conditions permissible for hybridization. of complementary probe substrate (113) and ligation at the reactive end of the primer, as shown in step IV. Generally, hybridization and ligation are performed at a temperature above the melting temperature of the substrate to reduce non-specific side reactions. Each substrate construction in this example contains an anchor that exhibits three indicators. Each probe substrate has a selectively cleavable bond (indicated by a "V") between the two anchor binding sites. In step V, a second substrate construct (114) is added by template directed ligation and hybridization, etc. In step VI the formation of a fully extended Xpandomer intermediate (117) is demonstrated. This intermediate can be denatured from the template strand and selectively cleaved at the cleavage sites shown, thus forming a product Xpandomer suitable for sequencing. In some embodiments, denaturation is not required and the template strand can be digested instead.
Figure 12 is a condensed scheme of a second method, here for the preparation of another embodiment of a class I Xpandomer. In preparation for synthesis, a target nucleic acid (120) is provided and the ends are polished in preparation for blunt end ligation of adapters (121,122). Step I shows the ligation of doubly blocked hairpin primer precursors to the target nucleic acid. One end of the duplex hairpin primers is blocked with movable blocking groups (125a, 125b, 125c, 125d) provided to prevent ligation and concatenation of the template or adapter strands. Adapters are generally added in excess. The blocking groups are removed in step II, and the two strands of the template are separated by denaturation. In stage III, the hairpin primers self-hybridize, forming priming sites (126,127) for the subsequent ligation of substrate constructs, which can advance bidirectionally, that is, in both a 3 'to 5' and a 5 direction. 'to 3'. In step IV, the primed templates are contacted with a library of substrate constructs (128) under conditions permissible for complementary probe substrate hybridization and ligation. The ligation proceeds incrementally (ie, spreading the growing ends with apparent processivity) through a process of hybridization of complementary probe substrates and ligation at the ends of the nascent daughter strands. Each substrate construction in this example contains an anchor loop that exhibits reporter groups. In step V the formation of a completed Xpandomer intermediate (129) is depicted. Optionally, the template strand can be removed by nuclease digestion, releasing the Xpandomer. The intermediate product can be selectively cleaved at the cleavage sites shown, thus forming a product Xpandomer suitable for sequencing. Product Xpandomers are formed in free solution.
A method based on immobilized template strands is shown in Figure 13. Here, the template strands are anchored to a bead (or other solid phase support) by an adapter (131). The template is shown in contact with substrate constructs (132), and in step I, conditions are adapted so that hybridization occurs. It can be seen that "islands" of hybridized confined substrate constructs are formed. In step II, the addition of ligase, L, causes ligation of the confined substrate constructs, thus forming multiple contiguous sequences of gapped-separated ligated intermediates. In stage III, the conditions are adjusted to favor the dissociation of hybridized material of low molecular weight or unpaired, and in stage IV, the reactions of stages II to III are repeated one or more times to favor the formation of products of longer extension. This primerless process is referred to herein as "promiscuous ligation." The ligation can be extended bi-directionally and spliced intersections can be sealed with ligase, thus filling in the gaps. In step V, after optimization of the desired product lengths, the immobilized duplexes are washed to remove unreacted substrate and ligase. Then, in step VI, the daughter strands (shown here as a single-stranded Xpandomer intermediate) (138,139) are dissociated from the template. Selective cleavage of selectively cleavable bonds from the intermediate results in the formation of the Xpandomer product (not shown). In this embodiment, the immobilized mold can be reused. Once the Xpandomer products are sequenced, the contigs can be assembled by well-known algorithms to overlap and align the data to build a consensus sequence.
Referring to Figure 14, a method of using immobilized primers is shown. End-tailored templates (or random template sequences, depending on the nature of the immobilized primers) (142) hybridize to the immobilized primers (140) in step I. In step II, the immobilized templates (143) are contacted with a library of substrate constructs, the members of which are shown as (144), and conditions are set for template-directed hybridization. In this example, the 3 'OH ends of the probe member (R group) substrate construct have been substituted (146) to further reversibly block extension. In step III, the confined ends of the adjacent substrate and primer construct, or free end of the nascent growing daughter strand, are ligated and the 3 'OH end of the nascent daughter strand is activated by removing the blocking R group ( 146). As indicated in steps IV and V, this stepped cyclic addition process can be repeated multiple times. Typically, a wash step is used to remove unreacted substrates between each extension step.
The process is thus analogous to what is called "single base cyclic extension", but would be more appropriately called "single probe cyclic extension" here. Although the ligase, L is shown, the process can be performed with a ligase, polymerase, or by any suitable chemical coupling protocol to bind
ES 2 559 313 T3 oligomers in template directed synthesis. Chemical coupling can occur spontaneously at the confined ends of the hybridized probes, or a condensing agent can be added at the beginning of stage III and each resulting stage V of the cycle. The end-blocking group R is configured such that free initiated polymerization cannot occur on the mold or in solution. Step VI shows the formation of a complete Xpandomer intermediate (149); no more substrate can be added. This intermediate product can be dissociated from the template, the single-stranded product is then cleaved to open the backbone as previously described.
This method can be adapted for selective sequencing of particular targets in a mixture of nucleic acids, and for analyzing sequencing methods on sequencing matrices, for example, by non-random selection of immobilized primers. Alternatively, universal or random primers can be used as shown.
Figure 15 depicts a promiscuous hybridization method on an immobilized template (150) (step I), in which the library substrate constructs (152) are modified with a chemical functional group that is selectively reactive (156), depicted like an open triangle, with a confined probe. A detail of the chemical functional group of the substrate constructs is shown in the expanded portion shown by the hatched circle (Figure 15a). At a certain hybridization density, the coupling is initiated as shown in step II, producing high molecular weight Xpandomer intermediates linked by the crosslinked product (157), represented as a filled triangle, of the coupling reaction. A detail of the cross-linked probes is shown in the expanded portion shown in the hatched circle (Figure 15b) in the product from step II. This process can be accompanied by steps for the selective cleavage and removal of low molecular weight products and any possible mismatched products. The coupling chemistries for this promiscuous chemical coupling method are known to those of skill in the art and include, for example, the techniques disclosed in US Pat. 6,951,720 to Burgin et al.
In another embodiment, polymerase-based methods for assembling product Xpandomers are disclosed. Substrate triphosphates (Xmers) are generally the appropriate substrate for reactions involving a polymerase. The selection of a suitable polymerase is part of an optimization process of the experimental protocol. As shown in Figure 16 for illustration, and although it is not intended to be limiting, a reaction mixture containing a template (160) and a primer (161) is contacted with a library of substrate constructs (162) and a polymerase (P), under conditions optimized for template-directed polymerization. In step I, the polymerase begins to processively add dinucleotide Xmers (anchors with two reporters) to the template strand. This process continues in stages II and III. Each probe subunit added is a particular species selected by specific binding to the next adjacent oligomer of the template so as to form a continuous complementary copy of the template. Although not bound by theory, it is believed that the polymerase helps to ensure that the incoming probe species added with the nascent strand are specifically complementary to the next available contiguous segment of the template. Loeb and Patel describe mutant DNA polymerases with high activity and improved fidelity (US Patent No. 6,329,178). Williams, for example, in US patent application 2007/0048748 has shown that polymerases can be modified for high speed of incorporation and reduction in error rate, clearly linking error rate not with hybridization accuracy. , but with the processivity of polymerase. Step III produces a completed Xpandomer intermediate (168). The single-stranded Xpandomer intermediate is then treated by a procedure that may involve denaturation of the template strand (not shown). The primary backbone of the daughter strand is selectively cleaved to expand the anchors, thus forming an Xpandomer product suitable for use in a sequencing protocol, as previously explained.
As shown in Figure 17, polymerase-driven template-directed synthesis of an Xpandomer can be accomplished by alternative techniques. Here, an immobilized primer (170) to which a processed template strand (171) is hybridized in step I. In step II, the polymerase, P, is processively coupled specifically with complementary substrate constructs (175) from a library of such constructs (represented by 174) in the reaction mixture. Conditions and reagent solutions are adjusted to promote processive polymerase activity. As shown here, hybridization in stage II and polymerization in stage III are separate activities, but the activities of the polymerase need not be isolated in that way. In step IV, the incremental processive addition of complementary substrate constructs continues cyclically (continuously without interruption), producing the fully charged Xpandomer intermediate (177) as depicted resulting from step IV. The Xpandomer intermediate can be dissociated and expanded in preparation for use in a sequencing protocol as previously described. Note that this method also lends itself to sequencing methods analyzed by selection of suitable immobilized primers. In addition, the methods for stretching the mold to release the secondary structure are easily adapted to this method, and are discussed in later sections.
Polymerases include, but are not limited to, DNA-dependent DNA polymerases, DNA-dependent RNA polymerases, RNA-dependent DNA polymerases, RNA-dependent RNA polymerases, T7 DNA polymerase, T3 DNA polymerase, T4 DNA polymerase, T7 RNA polymerase, RNA T3 polymerase, SP6 RNA polymerase, DNA polymerase I, Klenow fragment, Thermophilus aquaticus DNA polymerase, Tth DNA polymerase, VentR® DNA polymerase (New England Biolabs), Deep VentR® DNA polymerase (New England Biolabs), large fragment
ES 2 559 313 T3 DNA polymerase Bst, Stoeffel fragment, DNA polymerase 9 ° N, DNA polymerase 9 ° N, DNA polymerase Pfu, DNA polymerase Tfl, DNA polymerase Tth, RepliPHI polymerase Phi29, DNA polymerase TIi, DNA polymerase beta eukaryotic, telomerase, Therminator ™ polymerase (New England Biolabs DNA polymerase), KOD HiFi ™ (Novagen), KOD1 DNA polymerase, Q-beta replicase, terminal transferase, AMV reverse transcriptase, MMLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered by bioprospecting, and polymerases cited in US 2007/0048748, US 6,329,178, US 6,602,695 and US 6,395,524. These polymerases include non-mutant isoforms, mutants, and genetically engineered variants.
Class II and III oligomeric constructions
Referring to Figures 18A through 18E, they describe class II oligomeric constructs in more detail, which (along with isomeric class III oligomeric constructs) can be both Xprobes and Xmers.
Figures 18A to 18C are read from left to right, showing first the probe-substrate construct (Xpandomer oligomeric precursor), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product prepared for sequencing.
As shown in Figure 18A, a class II substrate construct has an oligomeric probe member (-P<sup>1</sup>-P<sup>2</sup>-) (180) and an anchoring member, T (181). The anchor is attached by a single terminal bond (184) from a first terminal moiety to the P probe moiety<sup>2</sup>. At the distal end of the anchor (186), a second terminal moiety has a connecting group δ and is positioned close to R<sup>2</sup>. The second terminal residue also has a cleavable intra- anchor crosslinking (187) to limit it to this location. The cleavable cross-link (187) is indicated by a dotted line, which may indicate, for example, a disulfide bond. These limitations prevent the anchor from elongating or expanding and thus it is in a limited configuration. A second connecting group ε is positioned near the distal end (189) of the probe member near R<sup>1</sup>. Under mold-directed assembly, the substrates duplex with the target template so that the substrates are confined. Under controlled conditions, the δ and ε linker groups of the confined substrates bind to form a χ bond between adjacent substrate constructs (shown in Figures 18B and 18C). These linking groups are positioned on the substrate construct to limit these linking reactions with adjacent confined substrate constructs. The substrate construction does not preferentially bond to itself. Suitable binding and protection / deprotection chemistries for δ, ε, and χ are detailed in the description of the general oligomeric construct.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-Oh, they would find use in a ligation protocol, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with the terminal connecting groups for chemical coupling or with non-connecting groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
Substrate constructs are reagents used for the template-dependent assembly of a daughter strand, an intermediate composition to produce Xpandomers. Figure 18B shows the daughter strand of the duplex, a heterocopolymer with repeating subunits / shown in brackets). Primary skeleton of the daughter strand (~ P<sup>1</sup>-P<sup>2</sup>~) and target template strand (-p<sup>1</sup>-p<sup>2</sup>-) as a duplex (185). Each subunit of the daughter strand is a repeat motif comprising a probe member and an anchor member. The motifs have species-specific variability, indicated here by the superscript a. Each particular subunit on the daughter strand is selected from a library of motifs by a template-directed process and its probe binds to a corresponding sequence of complementary nucleotides on the template strand. In this way, the nucleobase residue sequence of the probes forms a contiguous complementary copy of the target template strand.
Each tilde (~) indicates a selectively cleavable bond. The internal bond between the remains P<sup>1</sup> And p<sup>2</sup> of a probe member are not selectively cleavable linkages, but inter-probe linkages (between subunits) are necessarily selectively cleavable as required to expand the anchors and the Xpandomer. In one embodiment, no direct link is formed between the separate subunit probes, thus eliminating the need for subsequent selective cleavage.
The daughter strand is composed of the Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors are converted to their "expanded configuration", the limited Xpandomer is converted to the Xpandomer product. The anchors are limited by the χ bonds formed by the connection with the probe members of adjacent subunits and, optionally, the intra-anchor bonds if they are still present. The χ bond joins the anchoring member of a first subunit with the confined end of an adjacent second subunit and is formed by bonding the positioned connecting groups, δ of the first subunit and ε of the second subunit.
It can be seen that the daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer" backbone. The primary backbone is composed of the contiguously confined probe substrates. The skeleton
"Limited Xpandomer ES 2 559 313 T3" prevents selectively cleavable bond between subunit substrates and is formed by χ-linked backbone residues, each backbone residue being a linear anchor bond, to P<sup>2</sup>, to P<sup>1</sup>, linking each link χ P<sup>1</sup> to the anchoring of the next skeleton rest. It can be seen that the limited Xpandomer backbone connects or loops over the selectively cleavable bonds of the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone fragments.
In Figure 18B, the connecting groups δ and ε have been cross-linked and now form an intra-subunit χ bond. After forming the χ bond, the intra-anchor bond can be broken, although it is shown here intact (dotted line on substrate). Generally, the formation of the χ bond depends on the proximity of the δ connecting group on the first subunit and the position of the ε connecting group of a confined second subunit, such that they are positioned and contacted during or after mold-directed assembly. of substrate constructions.
In other embodiments, crosslinking depends only on hybridization to the template to bring the two linker groups together. In still other embodiments, the χ bond is preceded by enzymatic coupling of the P probe members along the primary backbone, with formation of phosphodiester bonds between adjacent probes. In the structure shown here, the primary backbone of the daughter strand has been formed, and the inter-substrate bonds are represented by a tilde (~) to indicate that they are selectively cleavable. After dissociating or degrading the target template strand, cleaving the selectively cleavable bonds (including intra-anchor bonds), the limited Xpandomer is released and converted to the Xpandomer product.
Figure 18C is a representation of the Xpandomer class II product after dissociation from the template strand and after cleavage of selectively cleavable bonds (including those in the primary backbone and, if not already cleaved, intra-bonds). -anchorage). Methods for template strand dissociation include heat denaturation, or selective digestion with a nuclease, or chemical degradation. The Xpandomer product strand contains a plurality of κ subunits, where κ indicates the k subunit<sup>n</sup> in a chain of m subunits constituting the daughter strand, where κ = 1, 2, 3 am, where m> 3, and generally m> 20, and preferably m> 50, and more preferably m> 1000. Each subunit is made up of an anchor, and P probe residues<sup>1</sup> And p<sup>2</sup>. The anchor, T (181), is seen in its expanded configuration and stretches to its length between P<sup>2</sup> And p<sup>1</sup> of adjacent subunits. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 18D shows the substrate construction of Figure 18A as a molecular model, in which the probe member (180), represented with four nucleobase residues (81,82,83,84), is attached to the anchor (181) by a bond of the first terminal moiety of the anchor (184). An intra-anchor link (85) of a second terminal moiety is at the distal end of the anchor. A connecting group (δ) (86) is also arranged on the second terminal moiety and the corresponding second connecting group (ε) (87) is anchored with the end of the probe opposite the connecting group (δ). The anchor loop shown here has three indicators (78,79,80), which may also be species-specific for motifs.
Figure 18E shows the substrate construction after incorporation into the product Xpandomer. The subunits are cleaved and expanded and linked by χ bonds (88), formed by linking the connecting groups δ and ε referred to in Figure 18A. A subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 18C. "E" again indicates expansion.
In the Xpandomer product (Figure 18E) the primary backbone has been fragmented and is not covalently contiguous because any direct bonds between adjacent subunit probes have been cleaved. Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information of the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor is represented as a reporter construct with three reporter groups, various reporter constructs may be present on the anchor, and may comprise individual reporters that identify probe constituents, individual reporters that identify probe species, molecular barcodes that identify the probe species, or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
The class III oligomeric constructs, illustrated in Figures 19A to 19E, are non-isomers with the class II constructs discussed above. No additional description is included because the description of
ES 2 559 313 T3 class II is adequate to understand this class.
This class can serve to emphasize that all classes can be reflected in the mirror image application (i.e., exchange of the R groups<sup>1</sup> and R<sup>2</sup>). Furthermore, this serves to illustrate that the classes described are not intended to be comprehensive, but rather reflect some of the many possible arrangements encompassed by the present invention.
Figure 20 represents a condensed scheme of a method of preparation of a first embodiment of a class II Xpandomer; The method illustrates the preparation and use of substrates and products shown in Figures 18D and 18E. The method is done with solid phase chemistries. Methods of relieving the secondary structure in the mold are covered in a later section. Suitable conditions adapted for hybridization and chemical coupling are well known in the art and conditions can easily be optimized by one skilled in the art.
Step I of Figure 20 shows a reaction mixture containing an immobilized template (200) and a library of substrate reagents (201). Substrate constructs are observed to specifically bind to template in template directed hybridization. The conditions are adjusted to optimize the complementarity and fidelity of the union. As shown in the figure insert (see Figure 20a), each confined substrate construction brings the δ (202) functional group into proximity on the distal aspect of the anchor, shown here attached to the anchor stem by intra-crosslinking. - anchor (203), represented by the adjacent triangles, and the functional group ε (204) of the confined probe member.
In stage II, a crosslinking reaction occurs between hybridized proximally confined ends of the probe members involving the two functional groups δ and ε, thus forming an intersubunit probe anchor χ bond (205), represented as an open oval , as shown in the insertion of the figure (see Figure 20b). Hybridization occurs in parallel at various sites on the template, promiscuously, and chemical coupling can occur in a cycle of hybridization (stage III), stringent fusion and / or washing (stage IV), and chemical coupling (stage V). The cycle can be repeated to increase the number of contiguous subunits assembled to form the Xpandomer intermediate. Step VI illustrates a completed Xpandomer intermediate with two contiguous product strands of varying length. A similar method can be used with class III Xpandomers.
Figure 21 illustrates a method of processive ligation of class II substrates on an immobilized template. Step I shows a primer (210) that hybridizes to template (212), the tailored primer with a chemically reactive functional group ε shown in the figure insert as (214) (see Figure 21a). A reaction mixture containing class II substrates (216) is then added to step II. As shown in the figure insert (Figure 21a), these substrate constructs have δ (217) and ε (214) reactivity at opposite ends of the probe-anchor member. A first substrate construct is observed to specifically bind to the template in template directed hybridization. The conditions are adjusted to optimize the complementarity and fidelity of the union. A ligase is then used to covalently link the first probe to the primer (step II).
In steps III and IV, the process of hybridization and ligation of substrate constructs continues in order to construct the Xpandomer intermediate product shown formed in step IV. After this, in step V, crosslinking is performed between the δ (217) and ε (214) groups (see Figure 21b), producing a χ bond as represented in Figure 21c as (219). As shown in the figure inserts (Figure 21b and 21c), the functional group δ (217) on the anchor is limited by an intra-anchor crosslinking (211), represented by the adjacent triangles, until the enlace bond is formed . The completed Xpandomer intermediate is optionally dissociated from the template strand and cleaved to form an Xpandomer product suitable for sequencing. A similar method can be used with class III Xpandomers. This method can also be adapted for use with a polymerase by substituting triphosphate substrate constructs.
Class IV and V oligomeric constructions
With reference to Figures 22A to 22E, class IV oligomeric constructs are described in more detail.
Figures 22A to 22C are read from left to right, showing first the probe-substrate construction (Xsonda precursors or Xpandomer Xmer), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product prepared for sequencing.
Figure 22A shows a class IV substrate construct having oligomeric probe member (229) with P probe moieties.<sup>1</sup> And p<sup>2</sup> that are attached to the anchor, T (220). Anchor T is attached to P<sup>1</sup> And p<sup>2</sup> by appropriate bonding to the first and second terminal moieties of the anchor, respectively. The connecting groups ε of the first terminal residue and δ of the second terminal residue are positioned near the ends R<sup>1</sup> and R<sup>2</sup> of the probe, respectively (in an alternative embodiment, the positions of the functional groups may be reversed). Under controlled conditions, the functional groups δ (222) and s (221) will react to form a χ bond as shown in Figure 22B. These linking groups are positioned in the substrate construct to limit these linking reactions with adjacent confined substrate constructs. Substrate construction
ES 2 559 313 T3 preferentially does not bind to itself. Suitable binding and protection / deprotection chemistries for δ, ε, and χ are detailed in the oligomeric constructs overview.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-Oh, they would find use in a ligation protocol, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with the terminal connecting groups for chemical coupling or with non-connecting groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
Substrate constructs are reagents used for the template-directed assembly of a daughter strand, an intermediate composition to produce Xpandomers. Figure 22B shows the duplex daughter strand, a heterocopolymer with repeating subunits (shown in brackets). Primary skeleton of the daughter strand (-P<sup>1</sup>~ P<sup>2</sup>-) and target template strand (-P<sup>1</sup>-P<sup>2</sup>-) as a duplex (228). Each subunit of the daughter strand is a repeat motif comprising a probe member and an anchor member. The motifs have species-specific variability, indicated here by the superscript a. Each particular subunit on the daughter strand is selected from a library of motifs by a template-directed process and its probe binds to a corresponding sequence of complementary nucleotides on the template strand. In this way, the nucleobase residue sequence of the probes forms a contiguous complementary copy of the target template strand.
The tilde (~) indicates a selectively cleavable bond. The internal bond between the remains P<sup>1</sup> And p<sup>2</sup> of a probe member is necessarily selectively cleavable as required to expand the anchors and the Xpandomer. In one embodiment, no direct link is formed between the separate subunit probes.
The daughter strand is composed of an Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors are converted to their "expanded configuration", as shown in Figure 22C, the limited Xpandomer is converted to the Xpandomer product. The anchors are limited by the χ bonds formed by connection to the anchoring member of adjacent subunits and by the probe bonds. The χ bond joins the anchor member of a first subunit with the anchor of an adjacent second subunit and is formed by bonding the positioned connecting groups, δ of the first subunit and ε of the second subunit.
It can be seen that the daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer backbone" backbone. The primary backbone is composed of the contiguously confined probe substrates. The "limited Xpandomer backbone" is the linear bonding of the anchors in each subunit linked together by the χ bonds that bypass the subunit probe substrates. The χ bond results from a reaction of the ε functional group of a first subunit with the δ functional group of a confined second subunit. It can be seen that the limited Xpandomer backbone connects or loops over the selectively cleavable bonds of the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone fragments.
In Figure 22B, the connecting groups δ and ε have been cross-linked and now form an intra-subunit χ bond. Generally, the formation of the χ bond depends on the placement of the connecting group δ, in the first subunit, and the connecting group ε, of a second confined subunit, such that they come into contact during or after the mold-directed assembly of constructs. substrate.
In other embodiments, crosslinking of the χ bond depends only on hybridization to the template to bring the two linking groups together. In still other embodiments, the formation of the χ bond is preceded by enzymatic coupling of the P probe members along the primary backbone with phosphodiester linkages between adjacent probes. In the structure shown in Figure 22B, the primary backbone of the daughter strand has been formed, and the bond between the probe moieties is represented by a check mark (~) to indicate that it is selectively cleavable. After dissociating or degrading the target template strand, cleaving the selectively cleavable bonds, the limited Xpandomer is released and converted to the Xpandomer product as shown in Figure 22C.
In this regard, Figure 22C is a representation of the class IV Xpandomer product after dissociation from the template strand and after cleavage of the selectively cleavable bonds from the primary backbone. Methods for template strand dissociation include heat denaturation, or selective digestion with a nuclease, or chemical degradation. The Xpandomer product strand contains a plurality of κ subunits, where κ indicates the k subunit<sup>n</sup> in a chain of m subunits constituting the daughter strand, wherein m> 3, and generally m> 20, and preferably m> 50, and more preferably m> 1000. Each subunit is made up of an anchor (220), and lateral probe residues P<sup>1</sup> And p<sup>2</sup>. The anchor, T, is seen in its expanded configuration and stretches to its length between adjacent subunits. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
ES 2 559 313 T3
Figure 22D shows the substrate construction of Figure 22A as a molecular model, in which the probe member, depicted with four nucleobase residues (open circles), is attached to the anchor by a bond from the first terminal residue of the anchor. The connecting group (221), shown as ε in Figure 22A, is also from the first terminal moiety of the anchor. A connecting group (222), shown as δ in Figure 22A, is disposed on a second terminal moiety at the distal end of anchor (220). The anchor loop shown here has three indicators (800,801,802), which can also be species-specific for motifs. A selectively cleavable bond, shown as a "V" (225), is located within the probe member (229).
Figure 22E shows the substrate construction after incorporation into the product Xpandomer. The subunits are excised, shown as lines (225a, 225b), and are expanded and linked by χ bonds (223,224), formed by linking the connecting groups δ and ε referred to in Figure 22A. A subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 22C.
In the Xpandomer product of Figure 22E, the primary backbone has been fragmented and is not covalently contiguous because any direct bonds between adjacent subunit probes have been cleaved. Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information of the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor is represented as a reporter construct with three reporter groups, various reporter constructs may be present on the anchor, and may comprise individual reporters that identify probe constituents, individual reporters that identify probe species, molecular barcodes that identify the probe species, or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
Class V substrate constructs are similar to class IV constructs, with the primary difference being the position of the cleavable linkers. Figures 23A to 23C are read from left to right, showing first the probe-substrate construction (Xsonda precursors or Xpandomer Xmer), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product prepared for sequencing.
Figure 23A illustrates a class V substrate construct having first and second terminal T anchor residues (239) linked with two selectively cleavable terminal bonds (234, 238) (represented as two "-verticals). These cleavable bonds then bind to the first and second probe moieties, P<sup>1</sup> And p<sup>2</sup>, of an oligomeric probe member (235). The connecting groups ε (230) and δ (231) of said first and second terminal residues are positioned near the ends R<sup>1</sup> and R<sup>2</sup> of the probe (again, the positions of these functional groups may be reversed). Under controlled conditions, the functional groups δ and ε are reacted to form a χ bond. These linking groups are positioned in the substrate construct to limit these linking reactions with adjacent confined substrate constructs. The substrate construction preferentially does not bond with itself. Suitable binding and protection / deprotection chemistries for δ, ε, and χ are detailed in the oligomeric constructs overview.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-OH, would find use in a ligation protocol as found in Xprobes, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol as found in Xmeros. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with the terminal connecting groups for chemical coupling or with non-connecting groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
Substrate constructs are reagents used for the template-directed assembly of a daughter strand, an intermediate composition to produce Xpandomers. Figure 23B shows the duplex daughter strand, a heterocopolymer with repeating subunits (shown in brackets). Primary skeleton of the daughter strand (-P<sup>1</sup>-P<sup>2</sup>-) and target template strand (-P<sup>1</sup>-P<sup>2</sup>-) as a duplex (236). Each subunit of the daughter strand is a repeat motif comprising a probe member and an anchor member. The motifs have species-specific variability, indicated here by the superscript a. Each particular subunit on the daughter strand is selected from a library of motifs by a template-directed process and its probe binds to a corresponding sequence of complementary nucleotides on the template strand. In this way, the nucleobase residue sequence of the probes forms a contiguous complementary copy of the target template strand.
The tilde (~) indicates a selectively cleavable bond. Links connect remains P<sup>1</sup> And p<sup>2</sup> of a probe member with the anchor and are necessarily selectively cleavable as required to expand the anchors and the
ES 2 559 313 T3
Xpandomer. In one embodiment, no direct link is formed between the separate subunit probes.
The daughter strand is composed of an Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors are converted to their "expanded configuration", as shown in Figure 23C, the limited Xpandomer is converted to the Xpandomer product. The anchors are limited by the χ bonds formed by connection with the anchoring members of adjacent subunits and by the selectively cleavable bonds (234,238). The χ bond joins the anchor member of a first subunit with the anchor of an adjacent second subunit and is formed by bonding the positioned connecting groups, δ of the first subunit and ε of the second subunit.
It can be seen that the daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer backbone" backbone. The primary backbone is composed of the contiguously confined probe substrates. The "limited Xpandomer backbone" is the linear bonding of the anchors in each subunit linked together by the χ bonds that bypass the subunit probe substrates. The χ bond results from a reaction of the ε functional group of a first subunit with the δ functional group of a confined second subunit. It can be seen that the limited Xpandomer backbone connects or loops over the selectively cleavable bonds that connect to the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone dissociates or is otherwise fragmented.
In Figure 23B, the connecting groups δ and ε have been cross-linked and now form an intra-subunit χ bond. Generally, the formation of the χ bond depends on the placement of the connecting group δ, in the first subunit, and the connecting group ε, of a second confined subunit, such that they come into contact during or after the mold-directed assembly of constructs. substrate.
In some protocols, the cross-linking reaction relies only on template hybridization to bring the two reactive groups together. In other protocols, the binding is preceded by the enzymatic coupling of the probe members, with formation of phosphodiester bonds between adjacent probes. In the structure shown in Figure 23B, the primary backbone of the daughter strand has been formed. The anchor, now attached to adjacent subunits by χ bonds, and comprises the bounded Xpandomer backbone. Upon cleavage of the selectively cleavable bonds (~), the bounded Xpandomer is separated from the primary backbone to become the Xpandomer product, and its now unconstrained anchors expand linearly to their full length as shown in Figure 23C.
In this regard, Figure 23C is a representation of the class V Xpandomer product after cleavage of the selectively cleavable bonds that dissociates the primary backbone. The Xpandomer product strand contains a plurality of κ subunits, where κ indicates the k subunit<sup>n</sup> in a chain of m subunits constituting the daughter strand, wherein m> 3, and generally m> 20, and preferably m> 50, and more preferably m> 1000. Each subunit is formed of an anchor, T (239), as seen in its expanded configuration and stretches to its length between adjacent subunits. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 23D shows the substrate construction of Figure 23A as a molecular model, in which the probe member (235), represented with four nucleobase residues (open circles), is attached to the first and second terminal residues of the anchor by two cleavable bonds (234,238). The connecting group (232), shown as ε in Figure 23A, is from the first terminal residue of the anchor and the connecting group (233), shown as δ in Figure 23A, is from the second terminal residue of the anchor. The anchor loop shown here has three indicators (237a, 237b, 237c), which may also be species-specific for motifs.
Figure 23E shows the substrate construction after incorporation into the product Xpandomer. The subunits are cleaved (234a, 234b, 238a, 238b) and expanded and linked by χ bonds (249,248), formed by linking the connecting groups δ and ε referred to in Figure 23A. A subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 23C.
In the Xpandomer product of Figure 23E, the primary backbone (235) has been cleaved (dissociated). Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information from the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor is represented as a reporter construct with three reporter groups, various reporter constructs may be present on the anchor, and may comprise individual reporters that identify probe constituents, individual reporters that identify probe species, molecular barcodes that identify the probe species, or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
ES 2 559 313 T3
The preparation and use of a class V Xpandomer is illustrated in Figure 24. In the setup for synthesis in step I, a single-stranded template (245) is contacted and hybridized with sequencing primer (246). The first assembly (247) is then contacted with a library of class V substrate constructs and a polymerase (step II). In step III, the substrates have been processively added in a mold-directed polymerization. In step IV, polymerization of the primary backbone of the daughter strand is complete and the reactive functional groups on the side arms of the confined anchor crosslink, forming the χ anchor-to-anchor bonds. Finally, in step V, the cleavable bonds in the stems of the anchor loops are severed, releasing the synthetic anchoring skeleton from the oligomeric daughter strand and template. This Xpandomer (249) is thus entirely constructed of anchor linkages and is shown to spontaneously expand as it moves away from the rest of the synthetic intermediate. Here, the genetic information corresponding to the target polynucleotide sequence is encoded in the contiguous subunits of the anchors.
Preparation and use of Xmeros
Class I embodiments include Xprobes and Xmeros. Xprobes are monophosphates, although Xmers are triphosphates. "Xmers" are expandable oligonucleotide triphosphate substrate constructs that can be polymerized in an enzyme-dependent template-directed synthesis of an Xpandomer. Like Xprobes, Xmero substrate constructs have a characteristic "probe-loop" shape as illustrated in Figures 10A and 10C, in which R<sup>1</sup> is 5'-triphosphate and R<sup>2</sup> is 3'-OH. Note that the substrate constructs are oligonucleobase triphosphates or oligomer-like triphosphates, but the probe members (i.e., the oligomer) have been modified with an anchor construct and a selectively cleavable bond between the terminal bonds of the anchor as shown. in Figure 10D, the function of which is further illustrated in Figure 10E.
DNA and RNA polymerases can incorporate dinucleotide triphosphate, trinucleotide and tetranucleotide oligonucleotides with a level of efficiency and fidelity in a primer-dependent processive process as disclosed in US Pat. 7,060,440 to Kless. Anchor-modified oligonucleotide triphosphates of length n (n = 2, 3, 4, or more) can be used as substrates for polymerase-based incorporation into Xpandomers. Suitable enzymes for use in the methods shown in Figures 16 and 17 include, for example, DNA-dependent DNA polymerases, DNA-dependent RNA polymerases, RNA-dependent DNA polymerases, RNA-dependent RNA polymerases, T7 DNA polymerase, DNA polymerase. T3, T4 DNA polymerase, T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, DNA polymerase I, Klenow fragment, Thermophilus aquaticus DNA polymerase, Tth DNA polymerase, VentR® DNA polymerase (New England Biolabs), Deep VentR® DNA polymerase (New England Biolabs), Bst large DNA polymerase fragment, Stoeffel fragment, 9 ° N DNA polymerase, 9 ° N DNA polymerase, Pfu DNA Polymerase, DNA Polymerase Tfl, Tth DNA polymerase, RepliPHI Phi29 polymerase, TIi DNA polymerase, Eukaryotic DNA polymerase beta, telomerase, Therminator ™ polymerase (New England Biolabs), KOD HiFi ™ DNA polymerase (Novagen), KOD1 DNA polymerase, Q-beta replicase, transferase terminal, AMV reverse transcriptase, M-MLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered by bioprospecting, and the polymerases cited in US 2007/0048748, US 6329178, US 6602695 and US 6395524. These Polymerases include non-mutant isoforms, mutants, and genetically engineered variants.
Polymerization of Xmers is a method of synthesis of Xpandomers and is illustrated in Figure 16, for example, in which the 2mer substrate is provided as a triphosphate. Because Xmers are processively polymerized, the extension, crosslinking, end activation and high stringency wash steps normally associated with cyclic sequencing by synthetic methods are optionally eliminated with this approach. Thus, the reaction can be carried out in solution. The synthesis of Xpandomers with Xmers can also be performed with immobilized templates, as illustrated in Figure 17, in which an Xmer 4mer triphosphate is processively polymerized in a primer-dependent template-directed synthesis.
A variety of methods can be employed for the robust synthesis of Xmers of 5'-triphosphate. As described by Burgess and Cook ("Synthesis of Nucleoside Triphosphates", Chem. Rev. 100 (6): 2047-2060, 2000), these methods include (but are not limited to) reactions using nucleoside phosphoramidites, synthesis by nucleophilic attack of pyrophosphate on activated nucleoside monophosphates, synthesis by nucleophilic attack of phosphate on activated nucleoside pyrophosphate, synthesis by nucleophilic attack of diphosphate on activated syntone phosphate, synthesis involving nucleoside-derived activated phosphites or phosphoramidites, synthesis involving the direct displacement of 5'-O- leaving groups by triphosphate nucleophiles, and biocatalytic methods. A representative method for producing polymerase compatible dinucleotide substrates uses N-methylimidazole to activate the 5'-monophosphate group; subsequent reaction with pyrophosphate (tributylammonium salt) produces triphosphate (Abramova et al., "A easy and effective synthesis of dinucleotide 5'-triphosphates", Bioorganic and Med Chem 15, 6549-6555, 2007).
As discussed in more detail below, the Xmero anchor construction is related in design, composition and bonding to that of the anchor used for the Xprobes. In many embodiments, the genetic information is encoded on the anchor, and thus each anchor in each substrate construct is a species-specific anchor. The encoded information about the anchor is encoded with a reporter code that digitizes the genetic information. For example, the five-bit binary encoding in the anchors would produce 32 codes of
ES 2 559 313 T3 unique sequence (2<sup>5</sup>). This strategy can be used to encode only the 16 combinations of two nucleobase residues per probe member of a 2mer library, regardless of the orientation of the anchor. Similar to the Xprobe encoding, a variety of functionalization and labeling strategies for Xmers can be considered, including (but not limited to): functionalized dendrimers, polymers, branched polymers, nanoparticles, and nanocrystals as part of the anchor scaffold, in addition of indicator chemistries and indicator signals - to be detected with the appropriate detection technology. Base-specific labels can be introduced (by attachment to the anchor) both before and after polymerization of the X-mer, by covalent bonding or by affinity targeting.
Design and synthesis of Xsondas and Xmeros
An overview of synthetic and cleavage strategies are presented below, starting with the probe oligomers with selectively cleavable linkages, followed by the anchor and reporter anchor constructs.
An objective of an X-probe- or X-mer-based SBX method is to assemble a replica of the target nucleic acid as completely and efficiently as possible by template-directed synthesis, generally a process or combination of selected hybridization, ligation, polymerization, or cross-linking processes. chemistry of suitable precursor compositions, referred to herein as "substrates." The Xprobes and Xmeros substrates are supplied as reagent libraries (eg, as parts of kits for sequencing) for this purpose. Libraries are generally combinatorial in nature and contain probe members selected to specifically bind to any or all of the complementary sequences, such as would be found in a target polynucleotide. The number of probes required in a library for this purpose is a function of the size of the probe. Each probe can be considered to be a sequence fragment, and a sufficient variety of probe members must be present to form a contiguous copy of the contiguous sequence of complementary sequence fragments of the target polynucleotide. For probes in which each oligomer is a dimer, there are 16 possible species combinations of A, T, C and G. For probes in which each oligomer is a trimer, then there are 64 possible species combinations of A, T, C and G, etc. When random genomic fragments are sequenced, all of those species are likely to be required in a library of reagents.
Xprobes and Xmers are oligomeric substrate constructs that are divided into five different functional classes. The oligomeric substrate constructs have two distinct functional components: a modified oligonucleobase or "probe" member, and an anchor member ("T"). The probe is attached to the anchor member by a "probe-loop" construction, wherein the anchor loop is a precursor to the linearized anchor member of the final product Xpandomer. Each T anchor can be encoded with reporters (commonly referred to as "tags" or "tags"), or combinations thereof, that uniquely identify the probe sequence to which it is attached. In this way, the sequence information of the assembled Xpandomer is more easily detected.
The oligomer is the probe portion of the Xprobe. The probe is a modified oligonucleobase having a chain of x deoxyribonucleotides, ribonucleotides, or more generally nucleobase residues (where x can be 2, 3, 4, 5, 6, or more). In these discussions, a probe 2, 3, 4, 5, or 6 nucleobase residues in length can be referred to as a 2mer, 3mer, 4mer, 5mer, or 6mer, respectively.
Substrate construct reagents can be synthesized with a 5'-3 'phosphodiester oligonucleotide backbone, whose oligomer has nucleotides A, T, G, and C (structures shown in the table of Figure 25), or other nucleic acid analogs. hybridizable such as those having a peptide backbone, phosphono-peptide backbone, serine backbone, hydroxyproline backbone, mixed peptide-phosphono-peptide backbone, mixed peptide-hydroxyproline backbone, mixed hydroxyproline-phosphono-peptide backbone, mixed serine-phosphono-peptide backbone, treose backbone, glycol backbone, morpholino backbone, and the like, as known in the art. Deoxyribonucleic acid oligomers and ribonucleic acid oligomers, and mixed oligomers of the two, can also be used as probes. Other bases may also be substituted, such as uracil for thymidine, and inosine for the degenerate base. Fragmented nucleobase residues that have complementarity can also be used.
A more complete recitation of degenerate and unstable bases known in the art includes, but is not limited to, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer of xanthine and hypoxanthine, 8-azapurine, 8-substituted purines with methyl or bromine, 9-oxo-N<sup>6</sup>-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N<sup>4</sup>-ethanocytosine, 2,6-diaminopurine, N<sup>6</sup>-ethane-2,6-diaminopurine, 5-methylcytosine, 5-alkynyl (C3-C6) -cytosine, 5-fluorouracil, 5-bromouracil, thiouracil, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, pseudoisocytosine, isoguanine, 7,8-dimethylaloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine, and the nucleobases described in US Pat. Nos. 5,432,272 and 6,150,510, PCT published WO 92/002258, WO 93/10820, WO 94/22892 and Wo 94/22144, and in Fasman, Practical Handbook of Biochemistry and Molecular Biology, pp. 385-394, CRC Press, Boca Raton, IA, 1989.
As is known in the art, oligomers can be designed to include nucleotide modifiers. In some embodiments, these serve as attachment points for the anchor member or members. Purine derivatives
ES 2 559 313 T3 and pyrimidine suitable for the synthesis of derivatized oligomers are well known in the art. Two such representative modified bases are shown in Figures 26A and 26B, in which a 5-amino modified cytosine derivative and an 8-amino modified guanine residue are depicted.
As illustrated in Figures 27A and 27B, taking a 4mer oligomer as an example (here illustrated as 5'monophosphate), any two of the four base positions on the oligomer can be modified to create binding sites by known chemistry. The modified nucleotides at probe residues 2 and 3 (on opposite sides of a selectively cleavable bond, represented as "V") are illustrated in Figure 27A. This figure illustrates a 4mer oligomer with amino linkers attached to the oligomer's cytosine and guanosine. Figure 27B illustrates a 4mer oligomer with benzaldehyde functional groups with the cytosine and guanosine of the oligomer. The details are illustrative of methods well known in the art. For simplicity, most of the illustrations provided herein will assume 4-mer, unless otherwise indicated, but it is understood that other libraries of substrate constructs or combinations of libraries may be employed in the practice of the present invention.
Cleavage
Generally, the Xsondas and Xmero substrate constructs have selectively cleavable linkages that allow controlled expansion of the anchor. As previously cited, such selective cleavage can be achieved by any number of techniques known to one of ordinary skill in the art, including, but not limited to, cleavage of the phosphorothiolate backbone with metal cations, acid cleavage of backbone modifications of phosphoramidate, selective nuclease cleavage of standard phosphodiester bonds using nuclease resistant phosphorothioate modifications for backbone protection, photoscission of nitrobenzyl modified skeleton linkers, and reduction of disulfide bonds.
Modification of substrate probes to include selectively cleavable linkages is illustrated in Figures 28A through 28D and Figures 29A through 29D. Figure 28A shows an example of an Xsonda dimer with a 2'-OH ribosyl group susceptible to cleavage by ribonuclease H in a DNA / RNA duplex Xpandomer intermediate. The bond is thus selectively cleavable, provided the other nucleotide (s) in the Xprobe are resistant to cleavage by RNase (e.g. 2'-o-methyl-pentose, 2'-deoxyribose nucleobases , "blocked" LNA nucleobases, and glycol- or peptide-linked nucleobases). These cleavage sites have other uses, for example, a ribonucleotide in the penultimate 5'-nucleobase and an adapter provides a cleavable linker between the Xpandomer and an immobilized support.
Figure 28B shows a phosphodiester-linked Xprobe that couples two nucleotides. In this figure, in addition to Figures 28A, 28C and 28D, the anchors for connecting the selectively cleavable linkage of the probe are indicated at (282) and (284). This linkage is selectively cleavable with mung bean nuclease, S1 nuclease, DNase I, or other DNases, for example, if other linkages that join the subunit anchors together are nuclease resistant. Synthesis of a 2mer library, for example, with a standard phosphate bond between the anchor attachment points and the phosphorothioate bond (s) at the position (s) of the nucleotide backbone that goes (n ) to remain intact, provides the desired cleavage pattern. Figure 28C is an Xprobe dimer held together by a 3'-phosphorothiolate bond, and in Figure 28D it is held together by a 5'-phosphorothiolate bond. These linkages are selectively cleavable by chemical attack, for example, with iodoethanol as described by Gish et al. ("DNA and RNA sequence determination based on phosphorothioate chemistry", Science 240 (4858): 15201522, 1988) or by cleavage with divalent metal cations as described by Vyle et al. ("Sequence- and strand-specific cleavage in oligodeoxyribonucleotides and DNA containing 3'-thiothymidine." Biochemistry 31 (11): 3012-8, 1992). Other backbone cleavage options include, but are not limited to, UV-induced photoredox cleavage (such as by adaptation of nitrobenzyl photo-cleavage groups) as described by Vallone et al. ("Genotyping sNps using a UV-photocleavable oligonucleotide in MALDI-tOf MS", Methods Mol. Bio. 297: 169-78, 2005), acid cleavage of phosphoramidate bonds as described by Obika et al. (“Acid-Mediated Cleavage of Oligonucleotide P3 '^ N5' Phosphoramidates Triggered by Sequence-Specific Triplex Formation”, Nucleosides, Nucleotides and Nucleic Acids 26 (8,9): 893-896, 2007) and periodate-catalyzed cleavage of modifications of the 3'-OBD-ribofuranosyl-2'-deoxy backbone as disclosed by Nauwelaerts et al. ("Cleavage of DNA without loss of genetic information by incorporation of a disaccharide nucleoside", Nucleic Acids Research 31 (23): 6758-6769, 2003).
As with Xprobes, cleavage of the poly-Xmer backbone to produce an Xpandomer takes place in a variety of ways. As shown in Figure 29A, for example, an Xmer containing an RNase digestible ribonucleotide base can be selectively cleaved at that position as long as the other nucleotide (s) in the Xmer are resistant to cleavage. by RNase (eg, 2'-O-methylpentose and 2'-deoxyribose nucleobases, "blocked" LNA nucleobases, and glycol- or peptide-linked nucleobases). For the Xmer described in Figure 29A, the 5 'base is a cytidine standard 2'-hydroxyl ribonucleotide and the 3' base is a 2 'RNase guanine resistant deoxyribonucleotide. The Xmero design allows selective RNase cleavage of the Xmero backbone to expand an Xpandomer. Alternatively, as shown in Figure 29B, DNase can be used to cleave all non-phosphorothioate protected backbone linkages. Consequently, a 2mer library, for example, with a standard phosphate bond between the anchor attachment points and the phosphorothioate bond (s) at the nucleotide backbone position (s) to remain intact ,
ES 2 559 313 T3 provides the desired cleavage pattern. Figure 29C is a dimer of Xmer held together by a 3'-phosphorothiolate bond, and Figure 29D is held together by a 5'-phosphorothiolate bond. These bonds are selectively cleavable by chemical attack, for example, with iodoethanol or by cleavage with divalent metal cations as previously mentioned. Other backbone cleavage options include (but are not limited to) UV-induced photoredox cleavage (such as by adaptation of nitrobenzyl photo-cleavage groups) and acid cleavage of phosphoramidate bonds, both of which have been cited above in Figure 28. In Figures 29A to 29D, the anchors for connecting the selectively cleavable linkage of the probe are indicated at (292) and (294).
Turning now to Figure 30, in a first general embodiment of a scheme for the synthesis of the "probe-loop" class I substrate construct, two nucleobase residues (circles) at the second and third positions on the probe are modified to create L1 and L2 attachment points for the two ends L1 'and L2' of the anchor. The anchor is shown here as separately pre-assembled and is attached to the probe member in a synthetic step (arrow). Intra-anchor disulfide bonds (represented by the two triangles) can be used in the assembly and use of these substrate constructs. The introduction of a reducing agent to the Xpandomer product will selectively break the disulfide bridges that hold the anchor together, thus allowing expansion of the Xpandomer backbone. Photo-cleavable bonds are also useful in the folding anchors during assembly with subsequent release and unfolding after exposure to light.
In other embodiments, the phosphodiester backbone of the substrate can be modified to create attachment points for anchoring as disclosed by Cook et al. ("Oligonucleotides with novel, cationic backbone substituents: aminoethylphosphonates", Nucleic Acids Research 22 (24): 5416-5424, 1994), Agrawal et al. (“Site specific functionalization of oligonucleotides for attaching two different reporter groups”, Nucleic Acids Research 18 (18): 5419-5423, 1990), De Mesmaeker et al., (“Amide backbone modifications for antisense oligonucleotides carrying potential intercalating substituents: Influence on the thermodynamic stability of the corresponding duplexes with RNA and DNA-complements ”, Bioorganic & Medicinal Chemistry Letters 7 (14): 1869-1874, 1997), Shaw et al. (Boranophosphates as mimics of natural phosphodiesters in dNa ”, Curr Med Chem. 8 (10): 1147-55, 2001), Cook et al. (US Patent No. 5,378,825) and Agrawal ("Functionalization of Oligonucleotides with Amino Groups and Attachment of Amino Specific Reporter Groups", Methods in Molecular Biology Vol. 26, 1994). The nucleobase residues that make up the probe member can be substituted with nucleobase analogs to alter the functionality of Xprobes. For example, blocked nucleic acids ("LNAs") can be used to increase the stability of the probe duplex. If chemical coupling of Xprobes (rather than enzymatic ligation) is envisaged, the 5 'and 3' probe ends can be further derivatized to allow chemical crosslinking.
Design, composition and synthesis of indicator constructions
In one embodiment, the anchors are encoded with "reporter constructs" that uniquely identify the sequence of nucleobase residues (or "probe" of Xprobes, Xmeros, and other oligomer substrates of Figure 8) or nucleobase (as in XNTP, RT-NTP and monomeric substrates of Figure 9) with which it is anchored. Reporters are reporters or combinations of reporters generally associated with anchors that serve to "analyze" or "encode" the sequence information inherent in the substrates and inherent in the order in which the substrates are incorporated into the Xpandomer. In some embodiments, the anchor is just a spacer and the indicators are, or are associated with, the substrate.
Figure 31 depicts a substrate anchor assembly method similar to that of Figure 30, but the pre-assembled anchor includes indicator groups (shown as the three rectangular portions of the anchor) and is called an "indicator construction." Anchor and indicator constructs can be prepared by a variety of polymer chemistries, and their use and synthesis are discussed in more detail here.
In the practice of the present invention, anchors can serve a variety of functions, for example: (1) as an anchor to link sequentially, directly or indirectly, adjacent anchors along the nucleobase backbone, (2) as a spacer to stretch or expand so that an elongated chain of anchored subunits, called an Xpandomer, is formed after excision of the skeleton, and / or (3) optionally comprises reporter constructs or reporter precursors that encode the nucleobase or oligomeric sequence information of the individual substrate construct to which the anchor is associated.
Indicator constructs are physical manifestations of indicator codes, which are bioinformational and digital in nature. The reporter codes analyze or encode the genetic information associated with the sequence fragment of the probe or nucleobase with which the reporter construct and the anchor are linked. The indicator constructs are designed to optimize the detectability of the indicator code by adjusting spatial gaps, abundance and signal intensity of the constituent indicators. Reporter constructs can incorporate a wide range of structural and signal elements including, but not limited to, polymers, dendrimers, beads, aptamers, ligands, and oligomers. These indicator constructs are made by a variety of polymer chemistries and are discussed further below.
In one embodiment, the reporter constructs are attached to the probe or nucleobase by a polymer anchor. Anchors can be constructed from one or more durable water- or solvent-soluble polymers including, but not limited to, the following segment (s): polyethylene glycols, polyglycols, polypyridines,
ES 2 559 313 T3 polyisocyanides, polyisocyanates, poly (triarylmethyl) methacrylates, polyaldehydes, polypyrrolinones, polyureas, polyglycol phosphodiesters, polyacrylates, polymethacrylates, polyacrylamides, polyvinyl esters, polystyrenes, polybamyroides, polyvinyl esters, polystyrenes, polybamyroides, polyvinyl esters, polystyrenes, polybamyroides, polyvinyl esters, polystyrenes, polybamyridones , polyacetamides, polysaccharides, polyhyaluranates, polyamides, polyimides, polyesters, polyethylenes, polypropylenes, polystyrenes, polycarbonates, polyterephthalates, polysilanes, polyurethanes, polyethers, polyamino acids, polyglycines, polyrolines, N-substituted polylysine, polypeptides, N-substituted peptides in the side chain, poly-N-substituted glycine, peptoids, carboxyl substituted peptides lateral, homopeptides, oligonucleotides, ribonucleic acid oligonucleotides, deoxynucleic acid oligonucleotides, Oligonucleotides modified to prevent Watson-Crick base pairing, oligonucleotide analogs, polycytidylic acid, polyadenylic acid, polyuridylic acid, polythymidine, polyphosphate, polynucleotides, polyribonucleotides, polyethylene glycol phosphodiesters, polyethylene glycol aninucleotide analogs, peptide-threinucleotide analogs glycol-polynucleotide analogs, morpholino-polynucleotide analogs, blocked nucleotide oligomer analogs, polypeptide analogs, branched polymers, comb polymers, star polymers, dendritic polymers, random, gradient, and block copolymers, anionic polymers, cationic polymers, stem-loop polymers, rigid segments, and flexible segments. Such polymers can be circularized at points of attachment on a substrate construction as described in, for example, Figures 30 and Figure 31.
The anchor is generally resistant to entanglement or folds so that it is compact. Polyethylene glycol (PEG), polyethylene oxide (PEO), methoxypolyethylene glycol (mPEG), and a wide variety of similarly constructed PEG derivatives (PEG) are widely available polymers that can be used in the practice of the present invention. Modified PEGs are available with a variety of bifunctional and heterobifunctional end crosslinkers and are synthesized in a wide range of lengths. PEGs are generally soluble in water, methanol, benzene, dichloromethane, and many common organic solvents. PEGs are generally flexible polymers that do not normally interact non-specifically with biological chemicals.
Figure 32A illustrates the repeating structure of a PEG polymer. Figure 32B shows an Xprobe or Xmero with a bare PEG anchor secured to the probe backbone by, for example, amine terminated linkers (not shown) using standard linker chemistries. Figure 32C shows the same substrate construction after cleavage of the probe backbone at a selectively cleavable ("V") bond, with the PEG polymer flexibly accommodating elongation of the Xpandomer. In some embodiments, the PEG polymer segments are piecewise assembled over the anchor to provide expansion length or to minimize steric issues, such as, for example, in the anchor stems near the terminal end links that connect the anchoring arms with the substrate.
Other polymers that can be used as anchors, and provide "scaffolding" for indicators, include, for example, polyglycine, polyproline, polyhydroxyproline, polycysteine, poly-serine, polyaspartic acid, polyglutamic acid, and Similar. Side chain functionalities can be used to construct functional group rich scaffolds for added signal capacity or complexity.
Figure 33A shows the structure of poly-lysine. In the anchor constructs embodiments described in Figures 33B to 33D, the poly-lysine anchor segments create a scaffold for indicator attachment. In Figure 33B, the ε-amino groups of the lysine side chains (indicated by arrows) provide functionality for the attachment of pluralities of reporter elements to a substrate construct, amplifying the reporter code. Figure 33C illustrates a star dendrimer attached to a substrate construct) with poly-lysine side chains (arrows).
Figure 33D illustrates dendrimer oligomer charge that can be detected by adding labeled complementary oligomers in a post-assembly labeling and "signal amplification" step. This provides a useful method of preparing a universal anchor by attaching an unlabeled dendrimer complex with multiple oligomeric reporter groups to a probe, and then treating the probe-bound dendrimer with a selection of one or two complementary labeled probes, a dendrimer is obtained. " painted ”specific to the individual substrate species. Different probe / anchor constructions can be painted with different complementary labeled probes.
In another embodiment of this approach, the backbone of the reporter system comprises eight unique oligonucleotides that are spatially encoded in a binary way using two distinguishable fluorescent reporters, each excited by the same FRET donor. Prior to or after coupling of the anchor constructs to their respective substrate construct, the anchor constructs are sequence-encoded by hybridizing the appropriate mixture of fluorescent reporter elements to create the appropriate probe-specific binary code. Variations of this approach are employed using coded, unlabeled dendrimers, polymers, branched polymers, or beads as the backbone of the reporter system. The oligonucleotides can again be used for binary coding of the reporter construct. An advantage with this approach is that the signal intensity is significantly amplified and that the coding is not dependent after a single hybridization event, both of which decrease the possibility of measurement and / or coding error.
ES 2 559 313 T3
Still another embodiment replaces the above-described oligonucleotide coding strategy with affinity-linked heterospecific ligands to produce, for example, a similarly binary encoded reporter construct. Using a 9-bit binary coding strategy, this non-universal reporter construct, in its simplest form, employs only a single coupling chemistry to simultaneously mark all anchors.
Given the flexibility of the SBX approach, a wide variety of indicators are used to produce unique measurable signals. Each anchor is uniquely encoded by one or many different indicator segments. The scaffold with which the reporter moieties are attached can be constructed using a wide variety of existing structural features including, but not limited to, dendrimers, beads, polymers, and nanoparticles. Depending on the coding scheme, one or many distinctly separate indicator scaffolds may be used for the indicator code of each anchor. Any number of options are available for the direct and indirect binding of reporter moieties to the reporter scaffold, including (but not limited to): indicator encoding of chemically reactive polymer (s) integrated into anchor constructions; indicator encoding of chemically reactive surface groups on dendrimer (s) integrated into the anchor backbone; and indicator coding of chemically reactive surface groups on bead (s) integrated into the anchor. In this context, a "bead" is broadly considered to denote any crystalline, polymeric, latex, or composite particle or microsphere. For all three examples, the abundance of indicator can be significantly increased by linking, with the indicator scaffolds, polymers that are loaded with multiples of indicators. These polymers can be as simple as a 100 residue poly-lysine or more advanced, such as labeled oligomeric probes.
Small size anchor constructions can also be used. For example, anchor constructions can be elongated in a post-processing step using targeted methods to insert spacer units, thus reducing the size of the indicator anchor.
Reduction of the size and mass of the substrate construction can also be achieved by using unmarked anchors. By removing bulky reporters (and reporter scaffolding such as dendrimers, which for some coding embodiments comprise more than 90% of the mass of the anchor), hybridization and / or coupling kinetics can be enhanced. Post-assembly anchor marking can then be employed. Reporters are attached to one or more linker chemistries that are distributed throughout the anchor constructs using spatial or combinatorial strategies to encode base sequence information. A simple binary coding scheme can use only a reactive binding chemistry for the post-assembly labeling of the Xpandomer. More complicated labeling schemes, which may require hundreds of unique linkers, use an oligonucleotide-based strategy for labeling Xpandomers. Another embodiment of post-labeled Xprobes or Xmers is to use the resulting nucleotide sequences derived from P<sup>1</sup> And p<sup>2</sup> (see Figure 10) that follow after cleavage and expansion of the Xpandomer for reporter binding by hybridization of a library of labeled probes. Similarly, other labeling and / or detection techniques can directly identify the more spatially resolved nucleotides.
Anchors and reporter constructs are employed that bind to substrate constructs with a desired level of accuracy, since miscoupling leads to inefficiencies in detection, and can also lead to polymer termination or indicator code rearrangement ( for example, if an asymmetric indicator code is used). The fidelity of the SBX process can be linked to the synthesis purity of the substrate constructs. Following purification of the anchor / reporter construct to enrich the full-length product, the construct can be directly coupled to a heterobifunctional (directional reporter encoding) or homobifunctional (symmetric or oriented reporter encoding) oligonucleotide probe. As with all polymer synthesis methods, purification (size, affinity, HPLC, electrophoresis, etc.) is used after the synthesis of substrate constructs and assembly are completed to ensure the high purity of expandable probe constructs of full length.
Synthesis of Class I Substrate Constructs Presenting Indicators
The synthesis of class I substrate constructs with reporter or reporter precursors presented on the anchors can be carried out in a variety of ways. A stepped process can be used to assemble a hairpin anchor polymer that is connected near the tie ends of the anchor of the substrate construct by a disulfide bridge. This orients the reactive ends so that docking of the anchor to the probe is highly favored. C6 amino modifier phosphoramidites are commercially available for all four nucleotides (Glen Research, USA) and are used to bind the anchor to form, for example, the completed substrate construct. Alternatively, benzaldehyde modified nucleotide linker chemistry may be employed. Size and / or affinity purification is useful for enriching correctly assembled substrate constructs.
Heterobifunctionalization of probes can advantageously be done on a solid support matrix, as is usual for the synthesis of oligos and peptides, or in solution with appropriate purification methods. A wide variety of ready-to-use heterobifunctional and homobifunctional crosslinking reagents are available to modify, for example, amine, carboxyl, thiol and hydroxyl moieties, and to produce a variety of chemical compounds.
ES 2 559 313 T3 strong and selective binding. Since C6 amino modifiers are available for all four deoxyribonucleotides, and can be made available for all four ribonucleotides, the functionalization strategies described here use ready-to-use amine-based crosslinking methods in conjunction with well-established amine protection / deprotection chemistries. . However, given the wide variety of phosphoramidite and crosslinking chemistries known in the art, methods not described herein can also be considered and produce equivalent products.
The need for probe heterobifunctionalization can be eliminated if the indicator coding strategy produces digitally symmetric coding or uses directional reference points (parity bits) to identify code orientation. In this case, two internal amine probe modifiers are sufficient, as any mating orientation of the reporter constructs on the anchor would produce unique probe specific sequence identification.
One or many polymer segments can be assembled sequentially using covalent linkages catalyzed by chemicals (eg, cross-linkers) or enzymes (eg, nucleic acid probe hybridization and ligation) to form an end-functionalized circular anchor. Given the current state of the art for polymer synthesis methods, the chemical crosslink synthesis approach constitutes a representative embodiment. As is customary for many polymer synthesis methods, a solid support matrix can be used as a scaffold for synthesis. Polymer segments can be assembled one at a time based on end functionalization, as mixed pair segments having different end functionalizations, or as linked pairs - two polymer segments with different homobifunctional end moieties (eg, hydrazide and amine) paired by disulfide bridges.
Labeling chemistries, which include both linker and indicator element moieties, are developed and optimized based on the high signal performance and stability, low polymer cross-reactivity and crosslinking, and the structural (reinforcement) rigidity these chemistries confer to the Xpandomer backbone, which may be important for sample preparation and detection as discussed below.
In Figure 31 discussed above, the complete hairpin anchor / reporter construct is assembled independently of the oligonucleotide probe and then attached by homobifunctional or heterobifunctional linker chemistries to the probe member. In an alternative embodiment, as shown in Figure 34, the anchors are stepwise circularized by construction on immobilized probe sequences. The reporter and anchor construct are synthesized with heterobifunctional (directional) or homobifunctional (symmetric or oriented) anchor segments directly linked to an oligonucleotide probe. The probe sequence shown in Figure 34A is a 4mer and includes P probe residues.<sup>1</sup> And p<sup>2</sup> (the second and third circles) separated by a selectively cleavable bond ("V"). Solid state synthesis techniques are used in the synthesis of the indicator construct. This synthesis is integrated with closure of the anchor loop. In step I of Figure 34A, a first anchor segment (341) with first reporter group (342) is added using specific functional group chemistry indicated by L1 and L1 '(the L1 linker on one of the probe moieties is blocks, as represented by the small rectangle). In step II, a second anchor segment (344) with second indicator group (345) is added using specific functional group chemistry indicated by L2 'and L2'. In step III, a third anchor segment (346) with indicator group (347) is added using specific functional group chemistry indicated by L2 and L1 '. In stage IV, and after the elimination of the blocking group of the L1 site on the P residue<sup>2</sup> of the probe (again represented by the small rectangle), the loop is closed after coupling of L1 'and L1.
In yet another embodiment, as illustrated in Figure 34B, the heterobifunctional linker chemistry can be used again to ensure that the anchor is directionally positioned on the probe (although this is not necessary for all coding strategies). A 4mer probe is again represented by the P probe residues<sup>1</sup> And p<sup>2</sup> (the second and third circles) separated by a selectively cleavable bond ("V"). In stage I, two anchors (341,344) with indicator segments (342,345) are brought into contact with functional groups L1 and L2 on P<sup>1</sup> And p<sup>2</sup>; the chemistries are specific to each anchor. The anchors at this stage can be stabilized with intra-anchor bonds (represented by the contiguous triangles). In stage II, a third anchor (346) with indicator segment (347) is used to "cap" the anchors after removing blocking groups (shown as the small rectangles) on the predecessor segments. The hooded segment can also be stabilized with an intra-segment bond (again represented as contiguous triangles).
Turning now to Figure 34C, an embodiment based on the addition of P probe residues is shown<sup>1</sup> And p<sup>2</sup> to separate terminal links of a preformed anchor. In stage I, the preformed anchor is first reacted with P<sup>1</sup> contacting L1 with L1 '. In stage II, the anchor is then reacted with P<sup>2 </sup>contacting L2 with L2 '. The anchor can be stabilized by an intra-anchor bond (represented by contiguous triangles), which puts P<sup>1</sup> in proximity to P<sup>2</sup>. The two probe moieties are then ligated in step III to form the selectively cleavable ("V") bond between P<sup>1</sup> And p<sup>2</sup> (the second and third circles). The ligation of P<sup>1</sup> And p<sup>2</sup> it can optionally be facilitated by duplexing said probe moieties with a complementary template.
Another embodiment for the synthesis of an indicator building segment is disclosed in Figure 35A. Using solid state chemical methods, a cleavable linker (351) is first anchored onto a solid substrate (350). A first anchor segment with a reversible bond (352) is reacted with the connector in step I, and
ES 2 559 313 T3 then in step II is reacted with a combinatorial library of monomers M1 and M2, shown in this example as a 4: 1 stoichiometric mixture of the respective monomers. Random copolymer synthesis is done in this way to produce unique anchor segment compositions. If peptide or amino acid monomers are used, this can be done with mixed anhydride chemistry, for example, by producing variable length random copolymer peptide anchors (Semkin et al., "Synthesis of peptides on a resin by the mixed anhydride method "Chemistry of Natural Compounds 3 (3): 182-183, 1968; Merrifield et al.," Solid Phase Peptide Synthesis. I. The Synthesis of a Tetrapeptide. "J. Am. Chem. Soc. 85 (14): 2149-2154, 1963). In step III, an L1-locked end linker (356) (represented as a small rectangle) is then added to the peptide segment. Following cleavage of the segment from the solid support (not shown), the anchor segment can be incorporated into a reporter construct using a variety of methods, some of which are described in Figure 34.
An alternative to this embodiment is demonstrated in Figure 35B. A reporter construct with randomly incorporated peptide fragments is synthesized as before in steps I, II and III, but in step III the end connecting group (358) is provided with a heterofunctional L2 linker. Unlike the previous example, the monomers are provided in equal proportions. After cleavage from the solid support (not shown), the indicator construct is available for further incorporation into anchors from the previous examples and is a heterobifunctional linker. The reporter constructs serve to encode the genetic information of the probe sequence fragment, as described in more detail below.
Polypeptides are useful reporter constructs and also serve as anchors. Random, periodic, alternating and block copolymers can be used in conjunction with homopolymer, for example, for the composition of the anchor segment and anchor construction. The polypeptide segments can be end-functionalized using succinimidyl-containing heterobifunctional crosslinkers for the conversion of amine to, for example, a hydrazide or 4-formylbenzoate (4FB). Polypeptides can be produced both by chemical synthesis and by cloning and overexpression in biological systems (bacteria, yeast, baculovirus / insect, mammal). Depending on the desired length of the segment, the anchor segments can range from N> 2 (short segments) to N> 1000 (long segments). Amine side group protection chemistries can be used as appropriate. As an alternative to using ready-to-use cross-linkers, polypeptides can be chemically synthesized with directly attached hydrazide, 4FB, and NHS moieties.
End-functionalized polypeptide segments using maleimide-containing heterobifunctional crosslinkers can also be used for the conversion of thiol to a hydrazide or 4-formylbenzoate (4FB), in the synthesis of anchors and reporter constructs. Similarly end-functionalized polypeptide segments using EDC crosslinkers can be employed for the conversion of carboxyl to a hydrazide (HZ) or 4-formylbenzoate.
The synthesis of the five classes of oligomeric substrate constructs having modified anchors with reporter constructs is accomplished using the above synthetic methods. Similarly, the synthesis of monomeric substrates having modified anchors with reporter constructs can also be achieved by the above synthetic mechanisms. The chemistries for these variants are generally applicable to the genera of Xpandomer species shown in Figures 8 and 9. The present inventors now return to coding strategies and rules for transporting genetic information in Xpandomers with reporter constructs.
Indicator constructs and indicator code strategies
An "indicator code" is a digital representation of a particular signal or signal sequence that is embedded in the indicators of a particular indicator construct. While the "indicator construct" is a physical manifestation of extragenetic information, the indicator code is its digital equivalent.
Digital encoding requires flag codes to follow certain rules. For example, at least 256 indicator codes are required to identify all possible combinations of a library of Xprobes 4meras. By having more flag codes than are possible, flag construction combinations are advantageous because additional states can be used for other purposes such as marking gaps, providing positional information, or identifying parity errors or high order errors.
Several strategies can be considered to physically represent an indicator code. The anchor can be divided into one or many encodable segments, each of which can be labeled both before and after Xpandomer assembly. Variable signal levels (tag amount), lengths (signal duration), and segment shapes of tagged anchor constructs can be used to increase encoding options. Encoding can also be expanded using multiplexable tags. For example, using a mass tag labeling approach, a large library of spectrally distinct tags can be used to encode only a single reporter segment; 14 distinct mass labels used in combinations of three labels clustered on a single anchor construct segment would create 364 unique 3-mass spectra. For multi-segmented anchoring, post-assembly marking of the Xpandomer of most or all of the anchoring skeleton may have the added benefit of increasing the stiffness of the Xpandomer, possibly making it easier to handle for detection and improving stability.
ES 2 559 313 T3
Figure 36A (and also Figure 2A) illustrates an anchor with a single indicator segment of an Xprobe or an Xmer. This approach takes advantage of highly multiplexable reporter labeling, such as mass spectrometer labels, to produce a large library of spectrally distinct outputs. A cleavable mass label is a molecule or molecular complex of cleavable indicators that can be readily ionized to a minimal number of ionization states to produce accurate mass spectra. When carefully controlled, a mass spectrometer can detect as few as a few hundred such mass mark indicators. This indicator code example does not need indicator positional information about the anchor skeleton to determine the status of the code (although positional information is required to distinguish one indicator code from the next). This feature simplifies anchor construction and possibly shortens anchor length requirements.
In one embodiment, using cleavable mass tag labels, the elongated Xpandomer can be presented for detection by a nanopore ion source (electrospray ionization, atmospheric pressure chemical ionization, photoionization) or by surface deposition (nanopeine, nanochannel, flux laminar, electrophoretic and the like) followed by ionization by desorption with laser, ion beam, or electron beam sources (matrix-assisted laser desorption ionization "MALDI", electrospray desorption ionization "DESI", silica desorption ionization "DIOS", secondary ion mass spectrometry "SIMS").
An example of 9-bit anchor constructs with 2 detection states ("1" and "0"), producing 512 code identities, is described in Figure 36B. In one embodiment, the anchor constructs consist of the "1" and "0" segments that produce two levels of electrical impedance as measured in a Coulter type nanopore detector. In a second embodiment, the anchor constructions consist of electrically conductive segments "1" and non-conductive segments "0". In a third embodiment, the anchor constructs consist of fluorescent "1" segments and non-fluorescent "0" segments. A plurality of different indicator elements can be considered for this type of encoding. For either of these approaches, the probe binding and anchor assembly chemistry may be identical - only the composition of the reporter segment would need to be changed. With this simple format, marking can be done both before and after Xpandomer assembly. Post-labeling is desired, as the unlabeled indicator segment is significantly less massive and as such is useful at a much higher concentration than a fully labeled indicator construct. Depending on the strategy, the polymer segments can be: (1) encoded by conjugatable or reactive surface chemistries (eg, poly-lysine, poly-glutamic acid), (2) non-reactive (eg, PEG, low polymers reactivity), or (3) a mixture of both reactive and non-reactive polymers. Reactive groups include, but are not limited to, primary amines (-NH3), carboxyls (COOH), thiols (-SH), hydroxyls (-OH), aldehydes (-RCOH), and hydrazide moieties (-RNN). Labeling of reactive segments, which may include deprotection of reactive groups, can be done directly on the substrate constructs, after formation of the Xpandomer intermediate, after cleavage of the backbone to produce the Xpandomer, or at any other time. in the SBX process as appropriate to produce the best results.
Additional levels may be possible, as shown in Figure 36C. For example, a binary directional coding strategy (non-symmetric indicator construct) requires at least eight indicator coding segments and a ninth segment, a codable hood segment (for anchoring loop closure and as a possible construct reference point core substrate) to produce the 256 codes minimally required for a four-base substrate (512 codes if the cap segment is coded).
Returning to Figure 36D, it is shown that a 7-bit flag construct, each flag with three detection states, produces 2187 detectable code identities. The use of flexible polymer spacers can be used for steric reasons.
A rigid equivalent of the previous 7-bit flag construction is shown in Figure 36E.
Figure 36F depicts an example of rigid 4 segment anchor constructs with eight total distinct states or tag combinations per segment. Using seven of the mark combinations to mark three of the anchor segment constructs will produce 343 unique code identities and leave 1 segment available for identification of subunit boundaries, parity, or other functional purpose. This embodiment may use a mixed marking approach where 1 to 3 different markings can be incorporated on each segment to produce at least eight unique combinations per segment as shown in Figure 37, in which combinations of different indicator types are contemplated. and indicator construction chemistry. Mass marking and fluorescent marking options, among others, can be used as described. The labeling of Xpandomer post-assembly assembly anchor constructs is driven by the abundance and identity of three crosslinker residues as described in Figure 37. Anchor segment lengths can be on the order of 100-1000 nm for limited diffraction measurements or <100 nm for near-field measurements, if these detection technologies are to be used. Shorter anchors can be used for other detection methods.
ES 2 559 313 T3
Rigid 3-segment anchor constructions with 22 total different brand combinations per segment are depicted in Figure 36G. The use of 21 of the tag combinations described in Figure 37 to tag two of the anchor segment constructs will yield 441 unique anchor construct identities and leave one segment available for identification of Xprobe subunit boundaries. This embodiment uses a mixed tagging approach where one to three different tags can be incorporated into each segment to produce up to 22 unique combinations per segment (Figure 37). Mass marking and fluorescent marking options, among others, can be used for this embodiment. As depicted in Figure 37, the marking of Xpandomer post-assembly anchor constructs is driven by the abundance and identity of three chemical moieties.
By designing anchor / indicator segments in which the abundance of reactive groups and the spatial dimension (radial distance from the polymer backbone) can be varied, coding levels can be achieved in which at least three total code states are possible: High "2", Medium "1" and Low "0" (Figure 36H). On the other hand, a level two, three mark coding strategy (i.e. 21 states per indicator) would only require two indicator coding segments to produce 441 codes and could use an additional segment for code orientation (Figure 36I) .
Flag coding can be designed to reduce errors inherent with flag construction and associated with detection technologies using various approaches. In the case of Coulter-type nanopore detection, the rate at which an Xpandomer passes through the pore and the modulation of the current it produces can depend on many factors including, but not limited to, the state of charge of the portion of the Xpandomer within the pore, the electrolyte concentrations, the surface charge states of the nanopore, the applied potential, the friction effects that limit the movement of Xpandomers, and the relative dimensions of both the Xpandomer and the nanopore. If the speed is not predictable, the current modulation decoding cannot use time (and constant speed) to solve the indicator measurement assignments. An encoding embodiment solves this question using 3-state encoding. The indicator signal is the impedance that causes a mark to the conductivity of electrolytes through the nanopore. By providing three possible impedance levels for an indicator, one bit of information is transition-encoded with the next flag. By design, this transition is always a change to one of the other two states. If all three states are marked A, B, and C, then a sequence of flags never has 2 A, 2 B, or 2 C together. In this way, the information is encoded in level transitions and is therefore independent of speed through the nanopore. One coding scheme is to assign all transitions A to B, B to C and C to A a value of "0" while transitions B to A, C to B and A to C are assigned the value "1". For example, the ABACBCA detected sequence is decoded to 0,1,1,1,0,0.
Although precise timing cannot reliably resolve sequential tags, it may be sufficient to differentiate tag sequence spacing on single anchor constructs from next sequential anchor. Additional spacers anchor lengths at either end of the reporter tag sequence (at the substrate construct attachment points) can provide a large time gap outlining anchor construct codes over time.
In cases where such proper timing is insufficient, a frame shift error may occur. Frame offset errors result when the listener reads a series of marks from multiple anchor constructs codes (frames), but does not correctly delineate the start of the code (frame start). This produces bad codes. One embodiment to solve this is to add more bits to the code than are needed to identify the corresponding base sequences (which are typically 1 to 4 bases in length). For example, eight bases are required to uniquely identify a 4-base sequence. Each 2-bit pair describes a unique base. By adding a parity bit to each base, the anchor construct code increases to 12 bits. A high parity error rate (close to 50%) would indicate a frame shift error resulting in the absence of a state change. In addition to an absence of state change, another type of error that can occur in this nanopore detection technology is that of an incorrectly read state. Single indicator errors can be isolated to the particular base using the parity bit and the value "unknown base" can then be assigned.
Xpandomers labeled using both electrical impedance, electrical conduction, and fluorescent segments can be measured in solution using a variety of nanopore, nanopeine, or nanochannel formats. Alternatively, Xpandomers can be surface deposited as spatially distinct linear polymers using nano comb, nanochannel, laminar flow, or electrophoretic deposition methods, among others. As with the dissolution approach, direct detection of the elongated Xpandomer at the surface can be done by measuring the signal characteristics of labeled anchor segments. Depending on the composition of the deposition material (conduction, insulation) and the underlying substrate (conduction, insulation, fluorescent), a variety of detection technologies can be considered for mark detection.
Additional SBX methods
Various SBX methods using class IV substrate oligomeric constructs were illustrated in the previous figures (Figures 11-17, 20, 21, 24). The present inventors now consider optional methods that
ES 2 559 313 T3 complement these protocols.
Endpoints functionalization
Figure 38 illustrates the preparation and use of target end adapters. Figure 38A illustrates end-functionalized complementary oligonucleotides duplexed to form an adapter with a bifunctional conjugate end ("L1" and "L2") and an enzymatically ligatable end (5'-phosphate and 3'-OH). These adapters can also be designed with additional functionality. For example, as shown in Figure 38B, end-functionalized adapters can be synthesized with nested backbone cross-linkers ("L3") and a cleavable bond ("V"). For simplicity, the cleavable bond V may have the same cleavable bond chemistry (or enzymology) used to release or expand the Xpandomer, although other cleavable linkers can be used if it is desired to differentiate between independent cleavage steps. Figures 38C and 38D show steps for the construction of a multifunctional adapter of Figure 38B. As shown, a magnetic bead with a surface-anchored oligonucleotide complementary to each adapter strand (two different bead mixtures) can be used to assemble differentially modified oligonucleotide segments. Once assembled, the segments can be enzymatically linked to covalently link the hybridized segments. With this approach, each segment can be individually modified in a way not available for standard oligonucleotide synthesis.
Manipulation of the Xpandomer can be helpful for efficient sample presentation and detection. For example, terminal affinity tags can be used to selectively modify one end of the Xpandomer to allow electrophoretic elongation. Attachment of a bulky neutral charge modifier to either the 3 'or 5' end (not both) produces an electrophoretic drag on the Xpandomer that causes the unmodified end to elongate as it travels with the detector. End modifiers, including, but not limited to, microbeads, nanoparticles, nanocrystals, polymers, branched polymers, proteins, and dendrimers, can be used to influence the structure (elongation), position, and rate at which the Xpandomer is presented. to the detector by conferring unique differentiating properties to its extremes such as charge (+ / - / neutral), buoyancy (+ / - / neutral), hydrophobicity and paramagnetism, to name a few. In the examples provided, the end modifications produce a drag load that allows the Xpandomer to elongate; however, the opposite strategy can also be employed in which end modification is used to pull the Xpandomer into and through the detector. With this approach, pulling the end modifier facilitates elongation of the Xpandomer. The end adapter on the template strand may also optionally contain one or more nucleic acids that will be used to synthesize the frame registration and validation signals in the finished Xpandomer (see Figure 54).
The incorporation of an affinity modifier can be done both before, during, and after the synthesis of Xpandomers (hybridization, ligation, washing, cleavage). For example, terminal affinity-tagged primers that are complementary to the adapter sequence can be pre-loaded to the ssDNA target under highly specific conditions prior to Xpandomer synthesis. The primer and its affinity tag can be incorporated into the full-length Xpandomer and can be used to selectively modify its end. A possibly more elegant approach is to enzymatically incorporate end modifiers. Terminal transferase (TdT), for example, is a template independent polymerase that catalyzes the addition of deoxynucleotides to the 3 'hydroxyl end of single or double stranded DNA molecules. TdT has been shown to add modified nucleotides (biotin) to the 3 'end (Igloi et al., "Enzymatic addition of fluorescein- or biotin-riboUTP to oligonucleotides results in primers suitable for DNA sequencing and PCR", BioTechniques 15, 486-497 , 1993). A wide range of enzymes are suitable for this purpose, including (but not limited to) RNA ligases, DNA ligases, and DNA polymerases.
Fork adapters are illustrated in Figures 38E and 38F. Hairpin adapters find use in ligation of blunt ends to give self-priming template strands. An advantage of this approach is that the daughter strand remains covalently coupled to the template strand and re-hybridizes more quickly after a fusion to remove unbound material and low molecular weight fragments. As shown in Figure 38F, these adapters may contain built-in pre-formed linker functionalities for downstream purification or manipulation, and may also contain cleavage sites for more efficient harvesting of Xpandomer daughter strands.
Target template preparation and analysis
To perform long continuous DNA whole genome sequencing, the use of the Xsonda-based SBX methods assumes that the DNA is prepared in a manageable way for hybridization, ligation, gap filling if required, expansion and measurement. The Xpandomer assembly process can be improved, in some embodiments, by surface immobilization of the DNA target in order to (1) reduce the complexity and cross-effects of hybridization, (2) improve washing, (3) allow for target manipulation (elongation) in order to facilitate enhanced hybridization, and / or (4) if nanopore sensor is used, provide a continuous interface with the detection process. As described in detail below, methods for target template preparation, analysis, and surface binding are expected to improve data quality and sequence assembly.
ES 2 559 313 T3
Most whole genome sequencing methods require fragmentation of the target genome into more manageable chunks. The longest chromosome in the human genome (chromosome 1) is ~ 227 Mb and the smallest chromosome (chromosome 22) is ~ 36 Mb. For most of the highly processive and continuous embodiments described herein, 36 Mb of sequencing continuous is too long to be sequenced with high efficiency. However, to take advantage of the inherent ability of the long read length of SBX, DNA fragment lengths of> 1 kb are targeted. Accordingly, various strategies can be employed to perform genome fragmentation and prepare a DNA target set compatible with the SBX method as disclosed herein.
One embodiment involves fragmenting the entire genome into 1-10 Kb chunks (average 5 Kb). This can be done by both restriction enzymes and hydrodynamic / mechanical shear. The fragments are then blunt-ended and repaired in preparation for blunt-ended ligation to a sequence adapter ("SA") or sequence adapter-dendrimer ("SAD") construct. The non-blunt end of SA or the SAD construct is designed to be incapable of ligation, so as to prevent assembly of multimers of such adapters. Any SA or SAD-linked targets that are present can be affinity purified away from free adapters. At this time, if the capture dendrimer is not introduced with the SA construct, then this can be done (using an effective excess dendrimer), followed by purification to isolate the target-SAD complex. Once purified, the complex is ready for binding to the test surface. For this purpose, it is desired that only a single target complex binds per reaction site. Alternatively, the capture dendrimer can instead be associated with the test surface, in which case the purified SA-target complex can be directly bound to the dendrimer that is already localized and covalently bound to the surface. A similar method of assembling DNA target on surfaces has been disclosed by Hong et al. ("DNA microarrays on nanoscale-controlled surface", Nucleic Acids Research, 33 (12): e106, 2005).
Another approach, called the "analysis" method, involves coarsely fragmenting the genome into 0.5-5 Mb chunks (using rare cutting restriction enzymes or by hydrodynamic / mechanical shear), followed by capturing such fragments to a microarray. modified (or cleavage surface) composed of gene / loci-specific oligonucleotide capture probes. Custom microarrays and oligonucleotide capture probe sets are widely available from various commercial sources (Arraylt, Euorfins-Operon, Affymetric Inc.). This additional analysis may provide an advantage for back end sequence assembly. Capture of these large fragments involves both partial and complete target denaturation in order to allow the capture probes to bind or duplex with specific target DNA. To reduce non-specific hybridization it may be necessary to load the capture matrix under dilute conditions so as to prevent cross-hybridization between templates.
Each set of cleavage capture probes, which can be displayed over a large area, is designed to provide ~ 3 Mb linear genome resolution. To provide efficient genome analysis, each set of individual capture probes can be comprised of 3 to 5 gene / loci specific oligonucleotides, linearly spaced on the genome by 0.5-1.5 Mb each. Capture probes are selected to be completely unique to the target fragment, thus providing both specificity and redundancy to the method. Given the 3 Mb resolution of a 3 Gb genome, this approach requires a capture array composed of specific probes of approximately 1000 genes / loci. Sets of specific microarray-ready oligonucleotide capture probes for human gene targets are widely available ready-to-use (Operon Biotechnologies, Huntsville AL, USA).
Once the non-specific binding events have been eliminated, each capture matrix can then proceed as independent reactions through the remainder of the genome preparation process in the same manner as discussed above for the non-assay method. The primary difference being that the target sample can now be analyzed positionally on an SBX test surface or as individual solution-based reactions, thus reducing the complexity of post-assembly of the data acquisition sequence.
Another embodiment is to have Xsonda-unanchored target hybridization and Xpandomer production. In this case, the SBX test is carried out in free solution. This approach may use one or a combination of physical manipulations such as using electrophoretic, magnetic, drag tags, or terminal positive / negative buoyancy functionalities under static or laminar flow conditions, as a means of elongating the target DNA before and during probe hybridization and ligation, eg, and, to elongate or expand the cleaved Xpandomer prior to detection. Free solution synthesis of Xpandomers, without immobilization, can be done using polymerases and ligases (with and without primers) and can also be done using chemical ligation methods. Both triphosphate substrate constructs and monophosphate substrate constructs can be used. Simultaneous synthesis of multiple and mixed nucleic acid target Xpandomers is envisioned. Generally, substrate triphosphate constructs are capable of continuous processive polymerization in solution and can be adapted to single tube protocols for sequencing a single molecule massively in parallel in free solution, for example.
ES 2 559 313 T3
Surface assembly of nucleic acid targets
One method of preparing target nucleic acids for sequencing uses end-functionalized double-stranded DNA target as shown in Figure 39. For this example, each adapter is typically provided with an ANH (reactive hydrazide) group useful for further processing. , for example, using 4-hydrazinonicotinate-acetone-succinimidyl hydrazone (SANH) bonding chemistry and an amine end-modified adapter oligonucleotide duplex. Similarly, an amine reactive SANh can be used to create reactive hydrazide moieties. SANH is readily conjugated with aldehydes such as 4-formylbenzoate (4-FB) to form stable covalent hydrazone bonds. Amine reactive C6-succinimidyl 4-formylbenzoate (C6 SFB) can be used to create a reactive benzaldehyde moiety. The end-functionalized complementary oligonucleotides are duplexed to form a bifunctional conjugatable end (SANH and amine) and an enzymatically ligatable end (3'-OH and 5'-phosphate). Ligation of the adapter creates an end-functionalized dsDNA target that can cross-link to surface-anchored aldehyde groups. Figure 39 illustrates end-functionalized dsDNA target using SANH and amine end-modified adapter oligonucleotide duplexes.
For many of the SBX methods described, nucleic acid targets can be covalently anchored to a flat coated solid support (stainless steel, silicon, glass, gold, polymer). As depicted in Figure 40, target binding sites can be produced by derivatization of 4FB from a SAM monolayer. The method uses the adapter ANH of Figure 39, and in stage I of Figure 40, the ANH is reacted with the 4FB heads of the monolayer and the template is denatured (stage II). The other strand of the template, which is also end-labeled, can be captured in a separate reaction. In step III, the free amine at the 3 'end of the single chain template is then reacted with a bead, for example here shown as a buoyancy bead, so that the target template can be stretched. These capture complexes can be assembled either by random self-assembly of a stoichiometrically balanced mixture of end-functionalized polymers (e.g., thiol-PEG-hydrazide for target binding; thiol-PEG-methoxy for hooded auto-monolayer). assembled "SAM") or by spatially resolved reactive bond point stamping using lithographic techniques. The stamped lithographic method can produce consistently separate target binding sites, although this is difficult to do for single molecule binding, although the random self-assembly method would probably produce more variable target binding separation, but has a high percentage of binding. of individual molecules. A wide range of monofunctional, bifunctional, and heterobifunctional crosslinkers are commercially available from a variety of sources. Compatible monofunctional, bifunctional, and heterobifunctional polymers of crosslinker (polyethylene glycol, poly-I-lysine) are also available from a wide variety of commercial sources.
The DNA target density of 1 trillion targets over a 100 cm surface<sup>2</sup> would require an average per target area of 10 um<sup>2</sup>. Target separation in this range provides sufficient target separation to prevent significant cross-reaction of target nucleic acids 5000 bases in length anchored to beads (bead diameter 100-1000 nm; 5 Kb dsDNA = 1700 um). The target area can be easily expanded if the cross-reactivity of the target and / or bead (> 1 bead / target) is determined to be unacceptably high.
Target elongation using beads or nanoparticles
As shown in Figure 40, a bead or nanoparticle (400) anchored on the free end of target DNA can be used to elongate and maintain the target in its single-stranded conformation during, for example, hybridization of Xsonda libraries. Retention of the target in an elongated conformation significantly reduces the frequency and stability of the target intramolecular secondary structures that form at the lower temperatures. Reducing or even eliminating secondary structure influences promotes efficient assembly of high fidelity substrate constructions.
A variety of approaches can be employed to apply elongation force to the single-stranded target. For example, paramagnetic beads / nanoparticles / polymers allow the use of the magnetic field to deliver controlled directional force to the target anchored to the surface. In addition, by controlling the direction of the magnetic field lines, paramagnetic beads / particles can be used to guide and hold the Xprobe-loaded target on the substrate surface during the washing steps. By sequestering the targets along the surface, loss of target due to shear forces can be minimized. As an alternative to elongation of the magnetic field, positive and negative buoyancy beads / particles that are both more and less dense than water, respectively, can be used to provide elongation force. All beads / particles are surface coated as necessary to minimize non-specific interactions (bead aggregation, probe binding) and functionalized to allow covalent crosslinking with adapter-modified targets.
Figures 41A through 41D illustrate representative bead-based target elongation strategies. Target elongation using paramagnetic beads / end-anchored particles attracted to an external magnetic field (B) is illustrated in Figure 41A. Target sequestration on the substrate surface (to reduce target shear) using paramagnetic beads / end-anchored particles attracted to an external magnetic field is illustrated in Figure 41B. Figures 41C and 41D illustrate target elongation using
ES 2 559 313 T3 end-anchored negative buoyancy beads / particles (higher density than water) and positive buoyancy beads / particles (lower density than water), respectively. Free dissolution methods for target elongation using end-anchored moieties, for example, using both a positive and negative buoyancy bead to functionalize opposite ends of a target, provide an elegant alternative to reducing the secondary structure of targets.
The use of these elongation strategies in the preparation of Xpandomers is illustrated in Figure 41E. Here, in step I, the immobilized template (416) is stretched using a floating bead duplexed to the target by the adaptamer primer (417). The template is then contacted in step II with substrate constructs and these are then bound to the single-stranded target, producing a double-stranded Xpandomer intermediate (411) with characteristic probe-loop construction and two backbones, one via the backbone. primary of the polynucleotide and the other by the limited Xpandomer backbone. In step III, the template strand is denatured, and in step IV the single-stranded Xpandomer intermediate is cleaved at selectively cleavable bonds in the primary backbone, causing looping and elongation of the fully extended substitute backbone Xpandomer product. (419).
Elongation of DNA targets using polymer end modifications
An alternative to the methods described in Figure 40 uses long functionalized polymers (rather than beads) covalently linked at the free ends of surface-anchored DNA targets to elongate the target. Electro-stretching and winding of polymer end modifications through the porous substrate followed by capture (sequestration) of polymer within the substrate produces fully elongated single-stranded target DNA significantly free of secondary structure. Figure 42A illustrates this method, showing the winding of the target strands through pores in a substrate. Porous substrates include, but are not limited to, gel matrix, porous aluminum oxide, and porous membranes. Capture of fully elongated polymer can be achieved by controlled chemical functionalization in which crosslinking or binding of polymer to porous substrate can be selectively initiated after complete elongation of the target. A variety of crosslinking or attachment strategies are available. For example, crosslinking of carboxyl functionalized elongated polymer to amine functionalized porous substrate can be accomplished with the introduction of the crosslinking agent 1-ethyl-3- [3-dimethylaminopropyl] carbodiimide hydrochloride (EDC).
Another embodiment illustrated in Figure 42B, is to electrostretch the polymer towards a functionalized surface that crosslinks or binds to the polymer. In this case, the polymer is not necessarily wound through the substrate, but instead cross-links or binds to functional groups on the surface. For this approach, substrates include, but are not limited to, gel matrix, porous aluminum oxide, porous membranes, and non-porous conductive surfaces such as gold. Both porous and non-porous substrates can be functionalized using a variety of ready-to-use crosslinking (hydrazide-aldehyde; oro-thiol) or binding (biotin-streptavidin) chemistries.
Alternatively, an enclosed electrical aperture system can be employed that uses variable electric field regions (high, low, ground) to capture charged polymer segments. The method, illustrated in Figure 42C, produces a type of locking Faraday cage that can be used to hold the polymer and target DNA in their elongated position. This approach requires efficient partitioning or confinement of electric fields, and has the advantage of allowing adjustable target elongation (throughout SBX synthesis) by modifying aperture parameters. Outside of the Faraday cage, the influences of electric fields on the substrate constructions are minimal.
In a related method, an elastic polymer is used to deliver a constant stretching force to the less elastic target nucleic acid. This method can be used as a supplement to any of the elongation methods described herein, including magnetic bead, buoyancy deformation, density (gravity) deformation, or electrostretching. The elastic force stored within the polymer provides a cushioned elongation force more consistent with the target strand.
Electro-straightening gel matrix based target substrate
An alternative method for the presentation of single-stranded target DNA (versus solid support binding) is shown in Figure 43A in which the target is covalently bound to a sieve gel matrix (430). In Figure 43b (inset), it can be seen that the substrate constructions associate with the anchored molds, which have been stretched straight. Electric fields can be used for elongation of targets and presentation of substrate constructs. This approach has advantages over other methods for DNA target array in terms of target density and target elongation (electro-straightening). For example, crosslinking of DNA to acrylamide by acrylate-modified end adapters (attached to oligonucleotides) is done routinely. Acrydite is an available phosphoramidite (Matrix Technologies, Inc., Hudson, NH, USA) that has been widely used to incorporate modifications of the 5 'end of methacryl to oligonucleotides: the double bond in the Acrydite group reacts with activated double bonds of acrylamide (Kenney et al., "Mutation typing using electrophoresis and gel-immobilized Acrydite probes", Biotechniques 25 (3): 516-21, 1998). Amine-functionalized agarose is also available ready-to-use (G Biosciences, St. Louis, MO,
ES 2 559 313 T3
USA) and can similarly be used to cross-link DNA using ready-made amine-reactive cross-linking chemistries (Spagna et al., “Stabilization of a β-glucosidase from Aspergillus niger by binding to an amine agarose gel”, J. of Mol. Catalysis B: Enzymatic 11 (2,3): 63-69, 2000).
As with previously described solid support ligation methods, DNA fragments are usefully functionalized with end adapters to produce, for example, a modification of 5'-Acrydite. To ensure uniformity of targets, dsDNA targets are denatured (and maintained in a denatured state) while crosslinking with the sieving matrix. The target can be denatured using a variety of existing techniques including, but not limited to, urea, alkaline pH, and thermal fusion. Once it has been cross-linked with a gel matrix, the denatured target DNA can be straightened by applying an appropriate electric field (similar to electrophoresis). Diffusion, electrophoresis, and neutralization of denaturants within the gel matrix along with close temperature control, with the addition of hybridization-compatible solutions, produces an environment conducive to hybridization of substrate constructs. For example, Xprobes can be continuously presented to electrophoresis anchored targets. This forming allows for the recycling of probes and provides good control with respect to the flow rate of the X-probes.
The target density of a series of 5 mm x 50 mm x 20 mm gel matrices produces a reagent volume of 5.0 x 10<sup>12</sup> p.m<sup>3</sup>, which is substantially higher than what can be presented on a flat surface. Figure 43C shows the chemical structures for acylamide, bis-acrylamide, and oligonucleotide end modification with Acrydite. Figure 43D illustrates a covalently attached DNA target polyacrylamide gel matrix.
Enzymatic processing of substrate constructs can be done within the gel matrix. Enzyme assays using polymerases (PCR) and ligases are routinely used for the modification of oligonucleotides covalently coupled to acrylamide gel pads (Proudnikov et al., "Immobilization of DNA in Polyacrylamide Gel for the Manufacture of DNA and DNA-Oligonucleotide Microchips" , Analytical Biochemistry 259 (1): 34-41, 1998). Ligase, for example, is delivered to Xsonda targets assembled within the gel matrix either by passive diffusion, electrophoresis, or a combination of both. A wide range of matrix densities and pore sizes are considered, however, passive diffusion benefits from matrix relative density. low and large pores. Active delivery of ligase by electrophoresis can be achieved by binding charge modifiers (polymers and dendrimers) to the enzyme to enhance its migration through the matrix.
A variation of the gel matrix method is to electrophoresis end-modified single-stranded target polynucleotides through a gel matrix. End modifications can be used to create drag on one end of the target. A cationic end modification can be made at the other end of the target so that it is carried through the gel. A temperature zone or gradient can also be induced in the gel so as to ramp or modulate the stringency of hybridization as the targets progress through the gel matrix. Enzymatic processing can be done within the gel matrix or after the hybridized target-substrate complexes exit the gel matrix. The process can be repeated as necessary to produce the desired average Xpandomer length.
Figure 44 illustrates the use of drag marks for post-synthesis manipulation of Xpandomers. The drag mark (represented as a diamond) serves as the dissolution equivalent of a stretching and transport technique (Meagher et al., “Free-solution electrophoresis of DNA modified with drag-tags at both ends”, Electrophoresis 27 (9) : 1702-12, 2006). In Figure 44A, the drag marks are linked by linker chemistry to functionalized adapters on the 5 'ends of the single stranded template (L1' of the drag mark combines with L1 of the template). In the step shown, the processive addition of substrate constructs is shown already in progress. In Figure 44B, the carry-over tag is added by hybridizing with a complementary adapter. In Figure 44C, the Xpandomer is first treated so that a connector is attached to a 3 'adapter (shown as the small square with L1'), and then the connector (L1) is used in stage II to join the drag mark to it. The terminal transferase, which extends a 3'-OH end of single- or double-stranded DNA, polymerizes linker-modified nucleotide triphosphates such as biotinylated nucleotides. Other enzymes can also be used to add modified bases or oligomers. In Figure 44C, a drag mark is added to the 3 'end of the Xpandomer intermediate. Drag marks can include, but are not limited to, nanoparticles, beads, polymers, branched polymers, and dendrimers.
Gaps and gap filling with hybridization and ligation of substrate constructs
Multiple variants in ligation methods can be employed to address gap management and gap filling. In one embodiment, cyclic step ligation is used, in which probes are sequentially assembled from a surface anchored primer duplexed with the target polynucleotide. In another embodiment, hybridization and "promiscuous" ligation of substrate constructs occur simultaneously across the entire target DNA sequence, generally without a primer.
For both the cyclic and promiscuous ligation approaches, it is possible that, for example, some portion of the 4-mer substrate constructs (eg, 25%) and some portion of the 5-mer substrate constructs (eg, 20%) hybridizes adjacently so as to allow ligation. Of these duplexes
ES 2 559 313 T3 adjacent, some percentage may contain mismatched sequences that will produce both measurement error (if linked) and failure to link. Most cases of ligation failure (due to mismatch) is inconsequential, as unligated probes are removed prior to the next round of hybridization. The hybridization / wash / ligation / wash cycle can be repeated several hundred times, if desired (6 minutes / cycle = 220 cycles / day and 720 cycles / 3 days; 10 minutes / cycle = 144 cycles / day and 432 cycles / 3 days).
In the cyclic process (illustrated with Xprobes in Figure 14), Xprobes are sequentially assembled from one end of a surface anchored primer duplexed with the target DNA; one Xsonda per cycle. If this is the case, the target read length can be calculated using the following assumptions: 25% adjacent hybridization (4-mer) with 20% perfect probe duplex fidelity would yield only 20 sequential 4-mer ligations (400 x 25% x 20%) after 400 cycles. Using these assumptions, a 400 cycle assay (1.5 days at 6 '/ cycle) would yield an average product length of 80 nucleotides.
In the promiscuous process (illustrated with class I Xprobes in Figure 45), if Xprobe ligation and hybridization reactions are allowed to occur spontaneously and simultaneously along the target DNA template, replication of a template can be achieved. much longer DNA using far fewer cycles. After each cycle of hybridization and washing, ligation can be used to connect the remaining duplexes that are both adjacent and 100% correct. Another more stringent wash can be used to remove smaller ligation products (all 8mers and less; some 12mers), in addition to all unbound 4mers. With ligation reactions occurring throughout the 1,000-10,000 nucleotide target sequence, the processivity of the assay increases dramatically. As the promiscuous approach to Xprobe hybridization and ligation is independent of the length of the target DNA template, most of the target template can replicate in a fraction of the cycles required for the serial method. Longer cycle times can be used to compensate for kinetic limitations if such an adjustment increases the fidelity and / or number of hybridization / ligation reactions.
Figure 45 illustrates a sequential progression of the promiscuous process. Step I illustrates hybridization of 4mer Xprobes at multiple loci along the DNA target. Stage II illustrates the hybridization and ligation of adjacent Xprobes; the bound Xprobes stabilize and thus continue to hybridize even though the unligated Xprobes melt. Stage III and Stage IV illustrate another cycle of thermal hybridization followed by ligation and thermal fusion of unligated Xprobes. Each cycle preferentially extends existing linked X-probe chains. As illustrated in step IV, after repeated cycling, the DNA target is saturated with X-probes that leave gaps, shorter than a probe length, in which no duplexing has occurred. To complete the Xpandomer, the Xprobes are linked through the gaps by enzymatic or chemical means as illustrated in step V.
Standard gap filling
As illustrated with Xprobes in step V of Figure 45, after the cycle of hybridization and ligation process is complete, the sequence gaps can be filled in to produce a continuous Xpandomer. Gaps along the target DNA backbone can be filled using well established DNA polymerase / ligase based gap filling processes (Stewart et al., "A quantitative assay for assessing allelic proportions by iterative gap ligation", Nucleic Acids Research 26 (4): 961-966, 1998). These gaps occur when strings of adjacent Xprobes meet each other and have a gap length of 1, 2 or 3 nucleotides between them (assuming a class 1 4mer Xprobe as illustrated in Figure 45). Gap filling can also be done by chemical crosslinking (Burgin et al.). After the gap filling is complete, the target DNA complement, which is composed primarily of Xprobes ligated with periodic 1, 2, or 3 nucleotide charges (purification, cleavage, end modification, reporter labeling) can be processed as appropriate to produce a measurable Xpandomer. As the SBX test products are prepared and purified in batch processing prior to the detection step, detection is efficient and is not limited by any simultaneous biochemical processing.
In order to differentiate gaps from the nucleotide-specific reporter signal, modified deoxynucleotide triphosphates (which are both already labeled and capable of being labeled after assay) can be used to identify the sequence of gaps. Suitable deoxynucleotide triphosphates are illustrated in Figures 46A (linker-modified bases) and 46B (biotin-modified bases).
The length and frequency of gaps depend on several synthesis variables, including number of cycles, stringency of hybridization, library strategy (random sequencing vs. sublibrary with stoichiometric adjustment), target template length (100b -1 Mb) and reaction density (0.1-10 B). Gap frequency and gap length can be significantly reduced using maximum stringency conditions, provided the conditions are compatible with the specified test time interval. Target length and reaction density are also important factors with respect to both alignment fidelity and gap filling probability.
The promiscuous hybridization process can use thermal cycling methods to improve the stringency of hybridization and to increase the frequency of alignments of adjacent substrate constructs. Hybridization, ligation, and thermal fusion continue under routine accurate thermal cycling until most of the target sequence is probe-duplexed. Loosely bound non-specific probe-target duplexes can
ES 2 559 313 T3 be removed by a simple washing step, again under precise thermal control. Enzymatic or chemical ligation can be performed to link any adjacently hybridized oligomeric construct along the target DNA, generating longer and more stabilized sequences. Enzymatic ligation has the added benefit of providing a fidelity cross-check of the addition duplex, since the probe-mismatched targets are not efficiently ligated. Using a second wash under precise thermal cycling conditions, unligated substrate constructs can fuse the DNA backbone; longer intermediate products will remain. Longer linked sequences grow at multiple loci along the target DNA until replication is nearly complete. The hybridization / wash / ligation / wash process is repeated zero to many times until most of the target template has replicated.
Temperature cycling coupled with stabilization of the substrate construct due to base stacking can be used to enrich for adjacent hybridization events. Thermal cycling conditions, similar in concept to "temperature ramp down PCR", can be used to remove less stable non-adjacent duplexes. For example, by repetitively cycling between the annealing temperature and the upper melting temperature ("Tm") calculated for the oligos library (as determined by single probe fusion - ie, no base stacking stabilized) events can be positively selected. hybridization of adjacent probes. With sufficient enrichment efficiency, the number of cycles can be significantly reduced. This can be done under single-bath or multi-bath conditions.
With a small library size (256 4mer substrate constructs, for example), the library can be partitioned into sub-libraries based on an even narrower melting temperature range. This strategy can significantly benefit from the method if Tm bias cannot be controlled by modification of bases, hybridization dissolution adjuvants, and the like. Additionally, the probe stoichiometry can be adjusted (up or down) to compensate for any residual bias that may still exist in the library / sub-libraries.
Computer modeling was performed to evaluate the statistics of the occurrence of gaps and the lengths of consecutively connected substrate constructs that hybridize on a target DNA template. The model simulates a thermal hybridization / ligation cycle with a complete library of 256 4mer Xprobes. The model results presented here are based on the following model process:
i) The hybridization / ligation step simulates random 4-mer Xprobes that meet a randomly sequenced DNA template at random positions and that hybridize if they mate. No Xprobe can overlap any nucleotides. Hybridization continues until all locations on the target hybridize more than 3 nucleotides in length.
ii) The thermal fusion stage is simulated by eliminating all Xprobes that are in shorter chains than M Xprobes in length (where m = 2, 3, 4, 5 ...). The longest chains remain on the DNA template. A "String" is defined here as multiple consecutive Xprobes with no gaps between them.
iii) Repeat the cycle defined by i) and ii) so that Xprobes are constructed from the existing multi-Xprobe chain loci. Cycling stops when there is no change between 2 consecutive cycles.
Figure 47 and Figure 48 each have two graphs illustrating the statistics of running the model 100 times on a random DNA template. For processing purposes, the DNA target was chosen to be 300 nucleotides in length (but the statistic discussed here is not changed if the length of the template is 5000 nucleotides). Figure 47 shows the results for M = 3, where shorter strands of M = 3 "melt" the DNA template. Figure 47A shows the statistics of the chain lengths in the last cycle. The distribution appears to be almost exponential, ranging from 12mer (M = 3) to some 160mers long with an average of ~ 28 nucleotides in length. Figure 47B shows that an interval of 5 to 20 cycles is needed to complete the DNA template using this strand "grow" method and that 90% of the runs are completed after 12 cycles. Figures 48A and 48B show the same type of information for the case where M = 4, a more stringent thermal fusion. In this case, the average chain length increases to ~ 40 nucleotides, but requires ~ 18 cycles to complete 90% of the runs. Longer chain lengths reduce the likelihood of having a gap as you would expect. The mean probability of not filling a nucleotide position for these data of M = 4 has a mean of 0.039 (approximately 1 in 25).
Gap filling with 3-pole and 2-pole substrate constructions
An extension of the basic 4mer hybridization embodiment is to use 3mer and 2mer substrate constructs to fill in the 3-base and 2-base sequence gaps, respectively. Figure 49 shows an example of gap filling by sequential or simultaneous addition of shorter Xprobes. If the majority of the target template is duplexed successfully with Xprobes after 4mer hybridization, most of the remaining gaps are both 1, 2, and 3 bases in length. The average size of the Xpandomer can be significantly increased if a moderate to high percentage of the 2- and 3-base gaps can be filled and ligated with smaller X-probes. As the hybridization of these smaller probes is done after most of the target is duplexed, the temperature stringency can be lowered to allow duplexing of 3-mer and then 2-mer, respectively. The duplex stability of the smaller Xprobes increases due to base stacking stabilization. Following ligation, the remaining gaps are filled using a polymerase to insert nucleotides, followed by
ES 2 559 313 T3 new by ligation to bind all remaining adjacent 5'-phosphates and 3'-hydroxyls. Polymerase-incorporated linker modified nucleotides allow for the marking of gaps, both before and after incorporation. These markings may or may not identify the base. Based on statistical models, summarized in the diagrams of Figure 50, the full-length Xpandomer polymers have more than 95% coverage of the target sequence (greater than 16 bases). Depending on the model, the use of 3-mer / 2-mer probe gap filling lowers the percentage of unlabeled bases by a factor of 6 (assuming high-efficiency incorporation), which, for a 5,000-base target sequence, would yield 5,000 sequence bases mainly contiguous with less than 50 individual base holes without discriminating across the entire dyne replicon. As such, 16X coverage greater than 99% of the unidentified gap sequences is identified.
Gap filling with these smaller 3mer and 2mer Xprobes, for example, extends the length of the average Xpandomer three times. The diagrams summarizing the statistical modeling results as shown in Figures 50A and 50B indicate that the average length 3mer / 2mer filled gap Xpandomers, under the conditions described, are in the range of 130 bases of sequence. Furthermore, with 5000 gigabases of Xpandomer sequence synthesized in this way, which is the equivalent of converting only 2.5 ng of target sequence to Xpandomer, size selection of the 10% longest Xpandomer fragments from that population would yield 500 gigabases. sequence with average lengths in the range of 381 bases, although selecting the longest 2% size of Xpandomer fragments would yield 100 gigabases of sequence with average lengths in the 554 base range. These fragment lengths can be achieved without the need to fill individual base gaps with polymerase. If hybridization of 4mer Xprobes is done without any gap filling, statistical models indicate that the average lengths for the longest 10% and 2% of the fragments would be 106 and 148 bases, respectively.
Addition of adjuvants to reduce the target secondary structure
Long surface anchored single-stranded target DNA can present a challenging hybridization target due to intramolecular secondary structures. The addition of a bead anchored at the ends, as described in Figure 41E, reduces the appearance of these intramolecular formations, and destabilizes them when they occur. However, it may be found necessary to further decrease the formation of secondary structures, which can effectively block hybridization of substrate constructs, with the addition of adjuvants.
The addition of unlabeled 2-mer and / or 3-mer oligonucleotide probes to, for example, a hybridization mixture of X-probes can serve this purpose. To eliminate the possibility of incorporation into the Xpandomer, these probes are synthesized to be non-ligable (no 5'-phosphate or 3'-hydroxyl). Due to their small size, 2mer and / or 3mer adjuvants are unlikely to form stable duplexes at the hybridization temperatures of 4mer Xprobes; however, their presence in the reaction mixture at moderate to high concentrations reduces the frequency and stability of the target secondary structure by weakly and transiently blocking access to intramolecular nucleotide sequences that may otherwise be duplexed. Figure 51 is an illustration of how the addition of 2mer or 3mer adjuvants inhibits the formation of secondary structures (for simplicity only, 4mer Xprobes were not shown in the figure). Coupled with the elongation force provided by the bead anchor, adjuvants significantly decrease the frequency of secondary structure formation.
These dinucleotide and / or trinucleotide adjuvants can be composed of standard nucleotides, modified nucleotides, universal nucleotides (5-nitroindole, 3-nitropyrrole, and deoxyinosine), or any combination thereof, to create all necessary sequence combinations. Figure 52 shows common universal base substitutions that were used for this purpose.
Another approach to reducing the target secondary structure is to replicate target DNA to produce a synthetic cDNA target that has reduced secondary structure stability compared to native DNA. Figure 53 lists some nucleotide analogs, as illustrated in US application 2005/0032053 ("Nucleic acid molecules with reduced secondary structure"), that can be incorporated into the synthetic cDNA target described. These analogs, which include (but are not limited to) N4-ethyldeoxycytidine, 2-aminoadenosine-5'-monophosphate, 2-thiouridine-monophosphate, inosine-monophosphate, pyrrolopyrimidine-monophosphate, and 2-thiocytidine-monophosphate, have been shown to cDNA. This approach can be used in conjunction with target elongation and hybridization aids to reduce secondary structure.
Detection and measurement
As previously mentioned, the Xpandomer can be labeled and measured by any number of techniques. The massive data output potential of the SBX method matches well with nanopore-based sensor arrays or equivalent technologies. In one embodiment, the nanopore array can serve as an ionization source on the leading end of a mass spectrometer, where the reporter codes on the Xpandomer are cleavable mass spectroscopy marks. Other embodiments involve the use of a nanopore sensor such as electrical impedance / conductance or FRET.
ES 2 559 313 T3
One detection embodiment uses a nanopore array as depicted in Figure 54. In this embodiment, the Xpandomer is assembled using previously described methods, except that the DNA target is not anchored to an immobilized solid support, but rather is anchored to a magnetic bead that has its anchor probe wound through a nanopore substrate. Furthermore, a multitude of FRET donor fluorophores, for two excitation wavelengths, are anchored with the nanopore input (shown as small squares). FRET acceptor fluorophores constitute the indicators incorporated on the Xpandomer. After the bound Xprobes are cleaved from the Xpandomer, the Xpandomer is stretched by a force that can be electrostatic, magnetic, gravitational, and / or mechanical and can be facilitated by the bead used to extend the original DNA target (shown attached to the top of the Xpandomer). In the final measurement stage, the Xpandomer is pulled through the nanopore by applying a magnetic force on the magnetic bead shown below the nanopore structure. The donor fluorophores at the entrance to the nanopore are excited with a light source (λ1) and as the reporters pass proximal with the donor fluorophores, they are excited and emit their distinctive fluorescence (λ<sub>2</sub>) that decodes the associated nucleotide sequence.
Figure 54 also illustrates one embodiment of encoding the sequence in the indicator codes. On each anchor are 4 reporter sites, each loaded with a combination of two types of FRET acceptor fluorophore. This provides four measurable states on each site using relative intensity level. If the fluorescence of acceptor fluorophores were red and green, the encoding of these states with the 4 nucleotides can be: A: green> red, C: only red, G: only green and T: red> green.
In another embodiment, the Xpandomer is labeled with mass label indicators that are measured using a mass spectrometer. The Xpandomer is dosed into a narrow capillary that feeds the indicators sequentially into an electrospray ionizer. To allow mass spectrometer measurement of discrete mass mark indicators, the mass marks can be photo-cleaved from the indicator scaffold just prior to, during, or after ionization of the mass mark. Mass spectrometer based on magnetic sector, quadrupole, or time of flight ("TOF") can be used for the detection of mass marks. The instrument only needs to distinguish a limited number of separate mass marks. This can be used to improve the sensitivity and performance of the instrument. Therefore, using an instrument that has multiple channels to perform ionization and detection in parallel increases performance by orders of magnitude.
One detection embodiment of the mass spectrometer approach is a multi-channel TOF mass detector that can read> 100 channels simultaneously. A suitable instrument would use a multi-channel ionization source that feeds Xpandomer on multiple channels at a concentration and rate that maximizes channel usage, thus maximizing the rate of good quality data output. Such a source of ionization requires having adequate spacing of the mass mark indicator segments. Dough marks can be photo-cleaved as they emerge from a nanopore and become ionized. The dispersion requirements are not high, so a short flight tube is all that is required. Extremely high measurement output is possible with a multi-channel mass spectrometer detection approach. For example, an array of 100 nanopore ion channels that reads at the rate of 10,000 reporter codes per second (with 4 nucleotide measurements per reporter code) would achieve an instrument throughput of> 4 Mbase / second.
In another embodiment, the nanopore is used in a similar way to a Coulter counter (Figure 55). In this implementation, the charge density of the Xpandomer is designed to be similar to that of native DNA. The indicators are designed to produce 3 levels of impedance as measured in, for example, the nanopore detector. Xpandomers occur in a free solution that has high electrolyte concentrations, such as 1M KCl. The nanopore is 2 to 15 nm in diameter and 4 to 30 nm in length. To achieve good resolution of anchor construction marks, they are chosen to be close to the same length or longer. The diameter of the marks can be different for the 3 levels. The Xpandomer segment near the anchor construct links has no indicators (eg PEG), this segment will have a particular impedance level. One of the three indicator levels can be equivalent. Different impedance flags can be produced, involving both impedance level and temporal response, by varying segment lengths, charge density, and molecular density. For example, to achieve 3 different impedance levels, the labeled anchor construct segments can be chemically encoded to couple to one of three different polymer types, each with a different length and charge density.
When the Xpandomer passes through the nanopore detector, the current is modulated depending on what type of label is present. The amount of charged polymer that resides in the nanopore affects both the current of electrolyte species and the rate of translocation.
Polymer-based detection by nanopores is demonstrated in US Patent Nos. 6,465,193 and 7,060,507, for example, and as expected it is shown that the physical parameters of a polymer modulate the electrical output of a polymer. nanopore.
In another embodiment of a nanopore-based detection apparatus (Figure 56), the side electrodes attached to the nanopore are used to measure the impedance or conductivity from side to side in the nanopore, although a voltage is applied across the film solid support. This has the advantage of separating the translocation function from the impedance function. As the Xpandomer is transported through the nanopore, again the
ES 2 559 313 T3 current modulation. Micropipetting and microfluidic techniques are employed, along with drag marks, magnetic beads, electrophoretic stretching techniques, etc., in order to transport the Xpandomer through the nanopore. For example, end-labeled free solution electrophoresis, also called ELFSE, is a method of breaking the charge balance with respect to friction of free trailing DNA that can be used for electrophoresis of Xpandomers in free solution (Slater et al., "End-labeled free-solution electrophoresis of DNA", Electrophoresis 26: 331-350, 2005).
The methods of anchoring, stretching, marking and measuring large DNA fragments are well established (Schwartz et al., "A single-molecule barcoding system using nanoslits for DNA analysis", PNAS, 104 (8): 2673-2678, 2007 ; and Blanch et al., "Electrokinetic Stretching of Tethered DNA", Biophysical Journal 85: 2539-2546, 2003). However, resolution of individual nucleobases for the purposes of whole genome sequencing of native nucleic acids is beyond the capabilities of these techniques. Various "single molecule" Xpandomer detection methods are depicted in Figure 57-59. A microscope and any of a variety of direct imaging techniques that take advantage of the highest spatial resolution of the Xpandomer structure that can be conceived are depicted in Figure 57. An Xpandomer molecule is laid flat on a generally flat surface and scanned along its length. Examples include light microscopy and fluorescence microscopy. The use of high resolution microscopy is limited in most cases to> 100 nm resolution which is still possible with long anchors (100 nm per indicator, for example). Higher resolution (requiring anchors <100 nm per reporter, for example) can be achieved using techniques such as total internal reflection fluorescence microscopy (TIRF: Starr, TE et.al., Biophys. J. 80, 1575-84, 2001 ), zero-mode waveguides (Levine MJ et al., Science 299, 682-85, 2003) or near-field optical microscopy (NSOM: de Lang et al., J. Cell Sci. 114, 4153-60 2001) or confocal laser scanning microscopy. In the case of localized FRET, excitation interactions can be localized to <10 nm, leading to anchor lengths of ~ 10 nm / reporter.
The detection and analysis of large DNA molecules by electron microscopy is well established (Montoliu et al., "Visualization of large DNA molecules by electron microscopy with polyamines: application to the analysis of yeast endogenous and artificial chromosomes", J. Mol. Bio 246 (4): 486-92, 1995), however, accurate and high-throughput polynucleotide sequencing using these methods has proven difficult. Transmission electron microscopy (TEM) and scanning (SEM) for the detection of an Xpandomer are conceptually indicated in Figure 58. Here a focused electron beam is used to sweep an Xpandomer, which is again generally flat on a surface. Aspects of the structure of the Xpandomer serve to decode genetic information about the skeleton. Specimen fixation and spray coating techniques, which allow imaging of individual features and the size of molecule atoms, can be used to enhance magnification.
Also of interest are nanoelectrode aperture electron tunnel conductance spectroscopy, in which a tunnel electron beam is modulated between two nanoelectrode tips by transport of the Xpandomer between the tips (Lee et al., "Nanoelectrode-Gated Detection of Individual Molecules with Potential for Rapid DNA Sequencing ”, Solid State Phenomena 121-123: 1379-1386, 2007). The Xpandomer disturbs the tunnel current by its scanning-conduction effect, which can be amplified on native DNA by using suitable reporters. This technique has the advantage that specimen fixation and the vacuum requirement are avoided, and in theory, massively parallel arrays of electrode gates can be employed to read many Xpandomers in parallel.
Atomic force microscopy is conceptually illustrated in Figure 59. In a simple embodiment, a nanotube mounted on a sensitive cantilever is swept across a surface and the attractive and repulsive forces between the probe and the sample surface are translated into a topological picture of the surface being swept. This technique can achieve very high resolution, but has relatively slow scanning speeds (M. Miles, Science 277, 1845-1847 (1997)). Scanning Tunnel Electron Microscopy (STM) is a related technology for surface imaging; The probe, however, does not touch the surface, but rather a tunnel current is measured between the surface and the probe. Here, the Xpandomer can be laid flat on a surface and physically swept up with the probe tip, much like a phonograph needle on a recording.
Sequence assembly
The published human genome reference sequence (or other reference sequence) can be used as an alignment tool to aid in the assembly of massive amounts of SBX produced sequence data. Despite the likely inclusion of small, positionally identified sequence gaps, the long read length capabilities described for Xpandomer-based SBXs simplify and improve the fidelity of whole genome sequence assembly. As discussed above, the process can be further simplified by dividing contiguous fragments into dimensionally confined locations on the test reaction surface. In this embodiment, the analysis method can dramatically reduce assembly time and error.
ES 2 559 313 T3
Monomeric constructions
Figure 9 provides an overview of monomeric constructs of the invention. A total of five classes are distinguished including four classes of RT-NTP (VI, VII, VIII and IX) and one class of XNTP (X). Each class will be covered individually below.
Class VI to X monomeric constructs are distinguished from class I to V oligomeric constructs in that they use a single nucleobase residue as a substrate. In the following description, N refers to any nucleobase residue, but is typically a nucleotide triphosphate or analog herein. It has attachment points on an anchor (also described herein), for example, to the heterocyclic rings of the base, with the ribose group, or with the α-phosphate of the nucleobase residue. As described, the primary method of template-directed synthesis uses polymerase, but any method that can perform template-directed synthesis is appropriate, including chemical and enzymatic ligation methods.
For substrate constructs where s and δ linker groups are used to create inter-subunit bonds, a wide range of suitable commercially available chemistries (Pierce, Thermo Fisher Scientific, USA) can be adapted for this purpose. Common linker chemistries include, for example, NHS esters with amines, maleimides with sulfhydryls, imidoesters with amines, EDC with carboxyls for reactions with amines, pyridyldisulfides with sulfhydryls, and the like. Other embodiments involve the use of functional groups such as hydrazide (HZ) and 4-formylbenzoate (4FB), which can then be further reacted to form bonds. More specifically, a wide range of crosslinkers (hetero- and homo-bifunctional) are widely available (Pierce), including, but not limited to, sulfo-SMCC (sulfosuccinimidyl 4- [N-maleimidomethyl] cyclohexane-1-carboxylate). ), SIA (N-succinimidyl iodoacetate), sulfo-EMCS ([N-emaleimidocaproyloxy] sulfosuccinimide ester), sulfo-GMBS (N- [g-maleimidobutyryloxy] sulfosuccinimide ester), AMAS (N- (a- maleimidoacetoxy) succinimide), BMPS (N EMCA acid (Ne-maleimidocaproic) -ester of [βmaleimidopropyloxy] succinimide), EDC (1-ethyl-3- [3-dimethylaminopropyl] carbodiimide hydrochloride), SANPAH (Nsuccinimidyl-6- [4'-azido-2 '-nitrophenylamino] hexanoate), SADP (N-succinimidyl (4-azidophenyl) -1,3'-dithiopropionate), PMPI (N- [p-Maleimidophenyl] isocyanate, BMPH (N- [p-maleimidopropionic acid hydrazide], trifluoroacetic acid salt)), EMCH ([Ne-maleimidocaproic acid hydrazide], trifluoroacetic acid salt), SANH (succinimidyl 4-hydrazinonicotinateacetonehydrazone), SHTH (succinimidyl 4-hydrazidoterephthalate hydrochloride) and C6-SFB (C6-succinimidyl 4-formylbenzoate). Therefore, the method disclosed by Letsinger et al. ("Phosphorothioate oligonucleotides having modified internucleoside linkages", US Patent No. 6,242,589) can be adapted to form phosphorothiolate linkages.
In addition, well-established protection / deprotection chemistries are widely available for common linker moieties (Benoiton, "Chemistry of Peptide Synthesis", CRC Press, 2005). Amino protection includes, but is not limited to, 9-fluorenylmethyl carbamate (Fmoc-NRR '), t-butyl carbamate (Boc-NRR'), benzyl carbamate (Z-NRR ', Cbz-NRR') , acetamide-trifluoroacetamide, phthalimide, benzylamine (Bn-NRR '), triphenylmethylamine (TrNRR') and benzylidene-p-toluenesulfonamide (Ts-NRR '). Carboxyl protection includes, but is not limited to, methyl ester, t-butyl ester, benzyl ester, st-butyl ester, and 2-alkyl-1,3-oxazoline. Carbonyl include, but are not limited to, 1,3-dimethylacetal dioxane, and 1,3-dithiano N, N-dimethylhydrazone. Hydroxyl protection includes, but is not limited to, methoxymethyl ether (MOM-OR), tetrahydropyranyl ether (THP-OR), t-butyl ether, allyl ether, benzyl ether (Bn-OR), t-butyldimethylsilyl ether (TbDMS- OR), t-butyldiphenylsilyl ether (TBDPS-OR), acetic acid ester, pivalic acid ester and benzoic acid ester.
Herein, the anchor is often depicted as an indicator construct with three indicator groups, various indicator constructs may be present on the anchor, and may comprise a single indicator to identify the substrate, multiple indicators to identify the substrate, or the anchor can be bare polymer (which has no indicators). Note that flags can be used for detection timing, error correction, redundancy, or other functions. In the case of the bare polymer, the indicators can be the substrate itself, or they can be on a second anchor attached to the substrate. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
Indicator coding strategies are disclosed above, and discussed further below. For example, the two-bit binary encoding of each monomer would produce four unique code sequences (11,10, 01, 00), which can be used to identify each sequence base (adenine "A", cytosine "C", guanine " G ", thymidine" T "), assuming that the substrate coupling is directional. If it has no address, then a third bit provides unambiguous coding. Alternatively, a 4-state multiplexed unique indicator construct provides a unique indicator code for each sequence base. A variety of functionalization and labeling strategies can be considered for anchor constructs, including, for example, functionalized dendrimers, polymers, branched polymers, nanoparticles, and nanocrystals as part of an indicator scaffold, in addition to indicator chemistries with a characteristic of detection to be detected with appropriate detection technology including, for example, fluorescence, FRET emitters or excitators, charge density, size or length. Specific marks of base can be incorporated into the anchor as marked substrates both before and after assembly of the Xpandomer. Once the Xpandomer has been fully released and elongated, the indicator codes can be detected and analyzed using a variety of
ES 2 559 313 T3 of detection methods.
Substrate libraries suitable as monomeric substrates include (but are not limited to) modified ATP, GTP, CTP, TTP, and UTP.
Class VI monomeric constructions
Figure 60 describes class VI monomeric substrate constructs (a type of RT-NTP) in more detail. Figures 60A to 60C are read from left to right, showing first the monomeric substrate construction (Xpandomer precursor having a single nucleobase residue), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product ready for sequencing.
As shown in Figure 60A, class VI monomeric substrate constructs have an anchor, T (600), linked by a bond (601) of a first terminal residue to a nucleobase residue of the substrate, N. A connecting group , ε, is arranged in the first terminal residue (603) of the anchor close to R<sup>1</sup>. At the distal end (604) of the anchor, a second terminal moiety with a second connecting group, δ, is positioned close to R<sup>2</sup>. The second terminal moiety of the anchor is secured to the first terminal moiety in proximity to the nucleobase by a selectively cleavable intra-anchor cross-link (or by other limitation). Intra-anchor cleavable crosslinking (605) is denoted herein by dotted line, which may indicate, for example, a disulfide bond or a photocleavable linker. This limitation prevents the anchor from elongating or expanding and is said to be in its "limited configuration." Under mold-directed assembly, the substrates duplex with the target template so that the substrates are confined. Under controlled conditions, the placed connecting groups δ and ε of the confined substrates bind to form a bond between adjacent substrate constructs. The connecting groups δ and ε of a monomeric substrate construct do not form an intra-substrate bond due to positioning limitations. Suitable bonding and protection / deprotection chemistries for δ, ε, and χ are detailed in the monomeric constructs overview.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-Oh, they would find use in a ligation protocol, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with the terminal connecting groups for chemical coupling or with non-connecting groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
During assembly, the monomeric substrate construct is first polymerized onto the extensible end of the nascent daughter strand by a mold-directed polymerization process using a single-stranded mold as a guide. Generally, this process starts from a primer and proceeds in the 5 'to 3' direction. Generally, a DNA polymerase or other polymerase is used to form the daughter strand, and the conditions are selected such that a complementary copy of the template strand is obtained. Subsequently, the δ connecting group, which is now positioned with the ε connecting group of the adjacent subunit anchor, is selectively crosslinked to form a χ bond, which is an inter-subunit inter-anchor bond. The χ bonds link the anchors in a continuous chain, forming an intermediate product called the "duplex daughter strand", as shown in Figure 60B. After the χ bond is formed, the intra-anchor bond can be broken.
The duplex daughter strand (Figure 60B) is a hetero-copolymer with subunits shown in brackets. The primary backbone (~ N ~) k, template strand (-N '-) k, and anchor (T) are shown as a duplexed daughter strand, where κ indicates a plurality of repeat subunits. Each subunit of the daughter strand is a repeat "motif" and the motifs have species-specific variability, indicated here by the superscript a. The daughter strand is formed from species of the monomeric substrate construct selected by a template-directed process from a library of motif species, with the monomer substrate of each species of substrate construct binding to a corresponding complementary nucleotide on the strand. target mold. In this way, the nucleobase residue sequence (ie, primary backbone) of the daughter strand is a contiguous complementary copy of the target template strand.
Each tilde (~) indicates a selectively cleavable bond shown here as inter-substrate bonds. These are selectively cleavable to release and expand the anchors (and the Xpandomer) without degrading the Xpandomer itself.
The daughter strand is composed of an Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors are converted to their "expanded configuration", the limited Xpandomer is converted to the Xpandomer product. The anchors are limited by inter-subunit χ bonds, substrate bonding, and optionally intra-anchor bonds if they are still present. The χ bond joins the first terminal moiety of the anchor of a first subunit with the second terminal moiety of the anchor at the confined end of a second subunit and is formed by linking the positioned connecting groups, ε of the first subunit and δ of the second subunit.
ES 2 559 313 T3
It can be seen that the daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer" backbone. The primary backbone is composed of the contiguously confined and polymerized monomeric substrates. The "limited Xpandomer backbone" avoids selectively cleavable bond between monomer substrates and is formed by χ-linked backbone moieties, each backbone moiety being an anchor. It can be seen that the limited Xpandomer backbone connects via the selectively cleavable bonds of the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone fragments.
The anchor enlace bonding (crosslinking of the δ and ε linking groups) is generally preceded by enzymatic coupling of the monomer substrates to form the primary backbone, with, for example, phosphodiester bonds between adjacent bases. In the structure shown here, the primary backbone of the daughter strand has been formed, and the inter-substrate, are represented by a tilde (~) to indicate that they are selectively cleavable. After dissociating or degrading the target template strand, cleaving the selectively cleavable bonds (including intraanchor bonds), the limited Xpandomer is released and converted to the Xpandomer product. Methods for template strand dissociation include heat denaturation, or selective digestion with a nuclease, or chemical degradation. A selection cleavage method uses nuclease digestion in which, for example, phosphodiester linkages from the primary backbone are digested by a nuclease and anchor-to-anchor linkages are nuclease resistant.
Figure 60C is a representation of the Xpandomer class VI product after dissociation from the template strand and after cleavage of selectively cleavable bonds (including those in the primary backbone and, if not already cleaved, intra-bonds). -anchorage). The Xpandomer product strand contains a plurality of κ subunits, where κ indicates the k subunit<sup>n</sup> in a chain of m subunits that make up the daughter strand, where κ = 1.2, 3 am, where m> 10, usually m> 50, and usually m> 500 or> 5,000. Each subunit is formed of an anchor in its expanded configuration and stretches to its length between the χ bonds of adjacent subunits. The lateral substrate is attached to the anchor in each subunit. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 60D shows the substrate construction of Figure 60A as a molecular model, in which the monomer substrate member, represented with a nucleobase residue (606), is attached to the anchor by a bond (607) of the first terminal residue of the anchor. Also disposed on the first terminal moiety is a connecting group (608), shown as ε in Figure 60A. A second connecting group (609), shown as δ in Figure 60A, is disposed on the second terminal moiety at the distal end of the anchor. A selectively cleavable intra-anchor bond (602), represented by the contiguous triangles, is shown to limit the anchor by linking the first and second terminal moieties. The connecting groups ε and δ are positioned not to interact and to align preferentially near the sides of R<sup>1</sup> and R<sup>2</sup> of the substrate, respectively. The anchor loop (600) shown here has three indicators (590,591,592), which may also be species-specific for motifs.
Figure 60E shows the substrate construction after incorporation into the product Xpandomer. The subunits are cleaved and expanded and linked by χ bonds (580,581), represented as an open oval, formed by linking the connecting groups δ and ε referred to in Figure 60A. A subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 60C.
In the Xpandomer product of Figure 60E, the primary backbone has been fragmented and is not covalently contiguous because any direct bonds between adjacent subunit substrates have been cleaved. Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information of the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor is represented as an indicator construct with three indicator groups, various indicator constructs may be present on the anchor, and may comprise individual indicators that identify monomer or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
Turning now to Figure 61, a single base extension (SBE) method with monomeric substrate constructs is shown. The end-tailored target templates (or random target template sequences, depending on the nature of the immobilized primers) are first hybridized to the immobilized primers. Prior to step I, the immobilized templates (611) are contacted with a library of monomeric substrate constructs, one member of which is shown for illustration (612), and polymerase (P). Conditions are set for mold directed polymerization. In this example, the 5 'ends of the first monomeric substrate (R<sup>2</sup> of Figures 60A and 60D) are polymerized to initiate the nascent daughter strand. The substrate has
ES 2 559 313 T3 substituted at the 3 'ends (R<sup>1</sup> Figures 60A and 60D), to reversibly lock the further extension. This is shown in more detail in enlarged Figure 61a (dotted lines), in which the monomeric substrate is shown confining the primer. This substrate is a 5'-triphosphate, and a phosphodiester bond is formed with the primer by the action of the polymerase. It should be noted that the reactive functional group (δ in Figure 60A) shown in Figure 61c is normally deactivated (by reaction or other means) in the first added substrate (as shown in Figure 61a by the rectangle).
In step I of Figure 61, the blocking group is removed (no rectangle) to allow the addition of another monomer (extension). Methods to reversibly block the 3 'end include the use of Pd to catalyze the removal of an allyl group to regenerate a viable 3'-hydroxyl end, or the use of 3'-O- (2-nitrobenzyl) terminated nucleotides, wherein the active hydroxyls can be regenerated by exposure to a UV source to cleavage the terminal residue as described by Ju et al. (“Four-color DNA sequencing by synthesis using cleavable fluorescent nucleotide reversible terminators”, PNAS 26; 103 (52): 19635-19640, 2006, and “Four-color DNA sequencing by synthesis on a chip using photocleavable fluorescent nucleotides”, PNAS 26; 102 (17): 5926-31, 2005). The regenerated 3'-OH is shown in more detail in Figure 61b in enlarged view. In this view, the confined ends of the primer and adjacent substrate construct are polymerized and the 3'OH end of the nascent daughter strand is activated by removing the blocking group. In step II, the template is contacted with a library of monomeric substrate constructs, in which the δ functional groups are reactive under controlled conditions, and a monomer substrate is polymerized with the nascent daughter strand. This polymerization places the ε (608) and δ (609) groups of the confined substrate constructs as shown in Figure 61c. In step III, under controlled conditions, these groups react to form a χ (580) bond, as shown in more detail in Figure 61d.
As indicated, cycling through stages II, III, and IV extends the nascent daughter each time by additional substrate (construct). Typically, a wash step can be used to remove unreacted reagents between steps. The process is thus analogous to what is called in the literature, "single-base extension" cyclical. The process is shown with polymerase, P, but can be adapted for a suitable ligase or chemical ligation protocol to link substrate constructs in template directed synthesis. Step V shows the daughter strand intermediate for Xpandomer (605). This intermediate product can be dissociated from the template and primer, for example, with a nuclease that attacks the primary backbone of the daughter strand, thus relieving limited anchors and releasing the Xpandomer product.
As with all SBE methods, efficient inter-cycle washing is helpful in reducing undesirable side reactions. To further facilitate the incorporation of individual bases through template regions with high secondary structure, the extension temperature can be varied through each extension cycle and / or additives or adjuvants, such as betaine, TMACL, or PEG, can also be added to neutralize the effects of the secondary structure (as is done in conventional polymerase extension protocols). And, if necessary, the stoichiometry of the substrate construct species can be varied to compensate for the reaction bias favoring certain bases, such as C or G.
An alternative method of producing class VI Xpandomers is to do processive polymerization based on polymerase. DNA and RNA polymerases, in addition to any similarly functioning enzymes, were shown to catalyze the precise polymerization of RT-NTP, absent reversibly terminal blocking R groups, can be considered for this approach.
A wide range of crosslinking chemistries are known in the art, and are useful for the formation of χ bonds. These include the use of NHS esters with amines, maleimides with sulfhydryls, imidoesters with amines, EDC with sulfhydryls and carboxyls for reactions with amines, pyridyldisulfides with sulfhydryls, etc. Other embodiments involve the use of functional groups such as hydrazide (HZ) and 4-formylbenzoate (4FB), which can then be further reacted to link subunits. In one option, two different bonding chemistries, ε1 / δ1 and ε2 / δ2 (also referred to as L1 / L1 'and L2 / L2' elsewhere in this document) that react to form χ1 and χ2 bonds, respectively, and can be used to differentially functionalize two sets of RT-NTP substrate constructs. For example, if an SBE cycle is performed with a set of RT-NTP functionalized with ε1 / δ2, the next round of SBE would use the set of ε2 / δ1, causing a reaction of placed δ2 / ε2 to form χ1. The orderly activation of the crosslinking pairs is useful to minimize crosslinking errors and unwanted side reactions.
Class VII monomeric constructions
Class VII molecules are analogous to those of class VI previously described. The primary difference is that the cleavable bond is between the anchor and the substrate rather than between the substrates. In Figure 62, class VII monomeric substrate constructs (a type of RT-NTP) are disclosed in more detail. Figures 62A to 62C are read from left to right, showing first the monomeric substrate construction (Xpandomer precursor having a single nucleobase residue), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product ready for sequencing.
ES 2 559 313 T3
As shown in Figure 62A, class VII monomeric substrate constructs have an anchor, T (620), attached by a selectively cleavable linker (621) to a first terminal residue to a substrate nucleobase residue, N. Other connecting group, ε (622), is arranged in the first terminal residue of the anchor close to R<sup>1</sup>. At the distal end of the T anchor a second terminal moiety with a second connecting group, δ (623), is positioned close to R<sup>2</sup>. The second terminal moiety of the anchor is secured to the first terminal moiety in proximity to the nucleobase by a selectively cleavable intra-anchor cross-link (624) (or by another limitation). Intra-anchor cleavable crosslinking is denoted herein by dotted line, which may indicate, for example, a disulfide bond or a photocleavable linker. This limitation prevents the anchor from elongating or expanding and is in a "limited configuration". Under mold-directed assembly, the substrates duplex with the target template so that the substrates are confined. Under controlled conditions, the placed connecting groups δ and ε of the confined substrates bind to form a bond between adjacent substrate constructs. The connecting groups δ and ε of a monomeric substrate construct do not form an intra-substrate bond due to positioning limitations. Suitable bonding and protection / deprotection chemistries for δ, ε, and χ are detailed in the monomeric constructs overview.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-Oh, they would find use in a ligation protocol, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with the terminal connecting groups for chemical coupling or with non-connecting groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
During assembly, the monomeric substrate construct is first polymerized onto the extensible end of the nascent daughter strand by a mold-directed polymerization process using a single-stranded mold as a guide. Generally, this process starts from a primer and proceeds in the 5 'to 3' direction. Generally, a DNA polymerase or other polymerase is used to form the daughter strand, and the conditions are selected such that a complementary copy of the template strand is obtained. Subsequently, the δ connecting group, which is now positioned with the ε connecting group of the adjacent subunit anchor, is caused to cross-link and form a χ bond, which is an inter-subunit inter-anchor bond. The χ bonds join the anchors in a continuous chain, forming an intermediate product called the "duplex daughter strand", as shown in Figure 62B. After the χ bond is formed, the intra-anchor bond can be broken.
The daughter strand of the duplex (Figure 62B) is a hetero-copolymer with subunits shown in brackets. The primary backbone (-N-) k, template strand (-N '-) k, and anchor (T) are shown as a duplexed daughter strand, where κ indicates a plurality of repeating subunits. Each subunit of the daughter strand is a repeat "motif" and the motifs have species-specific variability, indicated here by the superscript a. The daughter strand is formed from species of the monomeric substrate construct selected by a template-directed process from a library of motif species , with the monomer substrate of each species of substrate construct binding to a corresponding complementary nucleotide on the target template strand. In this way, the nucleobase residue sequence (ie, primary backbone) of the daughter strand is a contiguous complementary copy of the target template strand.
The tilde (~) indicates a selectively cleavable bond shown here as the anchor to the substrate linker. These are selectively cleavable to release and expand the anchors (and the Xpandomer) without degrading the Xpandomer itself.
The daughter strand is composed of an Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors are converted to their "expanded configuration", the limited Xpandomer is converted to the Xpandomer product. The anchors are limited by the inter-subunit χ bonds, the cleavable bond to the substrate, and optionally, the intra-anchor bonds if they are still present. The χ bond joins the first terminal moiety of the anchor of a first subunit with the second terminal moiety of the anchor at the confined end of a second subunit and is formed by linking the positioned connecting groups, ε of the first subunit and δ of the second subunit.
It can be seen that the daughter strand has two backbones, a "primary backbone" and a "limited Xpandomer backbone". The primary backbone is composed of the contiguously confined and polymerized monomeric substrates. The limited Xpandomer backbone avoids the selectively cleavable bond that connects to the substrate and is formed by χ-linked backbone moieties, each backbone moiety being an anchor. It can be seen that the limited Xpandomer backbone connects via the selectively cleavable bonds connected to the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone dissociates or fragments.
Anchor χ bonding (crosslinking of the δ and ε linking groups) is generally preceded by enzymatic coupling of the monomer substrates to form the primary backbone with, for example, phosphodiester bonds between adjacent bases. In the structure shown here, the primary skeleton of the daughter strand has formed, and the
ES 2 559 313 T3 inter-substrate, are represented by a tilde (~) to indicate that they are selectively cleavable. After dissociating or degrading the target template strand, cleaving the selectively cleavable bonds (including intraanchor bonds), the limited Xpandomer is released and converted to the Xpandomer product. Methods for template strand dissociation include heat denaturation, or selective digestion with a nuclease, or chemical degradation. A selection cleavage method uses nuclease digestion in which, for example, phosphodiester linkages from the primary backbone are digested by a nuclease and anchor-to-anchor linkages are nuclease resistant.
Figure 62C is a representation of the Xpandomer class VII product after dissociation from the template strand and after cleavage of selectively cleavable bonds (including those attached to the primary backbone and, if not already cleaved, intra-bonds). -anchorage). The strand of the Xpandomer product contains a plurality of κ subunits, in which the k subunit indicates<sup>n</sup> in a chain of m subunits constituting the daughter strand, where κ = 1, 2, 3 am, where m> 10, usually m> 50, and usually m> 500 or> 5,000. Each subunit is formed from an anchor 620 in its expanded configuration and stretches to its length between the χ bonds of adjacent subunits. The primary skeleton has been completely removed. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 62D shows the substrate construction of Figure 62A as a molecular model, in which the monomer substrate member, represented with a nucleobase residue (626), linked by a selectively cleavable bond (625) to the first terminal residue of the anchorage. Also disposed on the first terminal moiety is a connecting group (629), shown as ε in Figure 62A. A second connecting group (628), shown as δ in Figure 62A, is disposed on the second terminal moiety at the distal end of the anchor. A selectively cleavable intra-anchor bond (627) is shown, represented by the contiguous triangles, which limits the anchor by linking the first and second terminal moieties. The connecting groups ε (629) and δ (628) are positioned not to interact and to align preferentially near the sides of R<sup>1</sup> and R<sup>2</sup> of the substrate, respectively. The anchor loop shown here has three indicators (900,901,902), which may also be species-specific for motifs.
Figure 62E shows the substrate construction after incorporation into the product Xpandomer. The subunits are cleaved and expanded and linked by χ (910) bonds, represented as an open oval, formed by linking the connecting groups δ and ε referred to in Figure 62A. A subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 62C.
In the Xpandomer product of Figure 62E, the primary backbone has dissociated or fragmented and is separated from the Xpandomer. Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information of the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor is represented as an indicator construct with three indicator groups, various indicator constructs may be present on the anchor, and may comprise individual indicators that identify monomer or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
Class VIII monomeric constructions
Class VIII molecules are analogs of class VI previously described. The primary difference is that the connecting group, ε, is connected directly to the substrate rather than to the anchor. In Figure 63, the present inventors describe class VIII monomeric substrate constructs (a type of RT-NTP) in more detail. Figures 63A to 63C are read from left to right, showing first the monomeric substrate construction (Xpandomer precursor having a single nucleobase residue), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product ready for sequencing.
As shown in Figure 63A, class VII monomeric substrate constructs have an anchor, T (630), linked by a linkage (631) of a first terminal residue to a nucleobase residue of the substrate, N. At the end distal of the anchor (632), a second terminal moiety with a second connecting group, δ, is preferentially positioned close to R<sup>2</sup>. The second terminal moiety of the anchor is secured to the first terminal moiety in proximity to the nucleobase by a selectively cleavable intra-anchor crosslinking (or by other limitation). Intra-anchor cleavable crosslinking (633) is denoted herein by dotted line, which may indicate, for example, a disulfide bond or a photocleavable linker. This limitation prevents the anchor from elongating or expanding and is said to be in its "limited configuration." A connecting group, ε (635), is attached to the monomer substrate preferentially close to R<sup>1</sup>. Under mold-directed assembly, substrates form a duplex with the mold
ES 2 559 313 T3 target so that the substrates are confined. Under controlled conditions, the δ and ε connecting groups of the confined substrates are positioned and linked to form a bond between adjacent substrate constructs. The connecting groups δ and ε of a monomeric substrate construct do not form an intra-substrate bond due to positioning limitations. Suitable bonding and protection / deprotection chemistries for δ, ε, and χ are detailed in the monomeric constructs overview.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-Oh, they would find use in a ligation protocol, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with the terminal connecting groups for chemical coupling or with non-connecting groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
During assembly, the monomeric substrate construct is first polymerized onto the extensible end of the nascent daughter strand by a mold-directed polymerization process using a single-stranded mold as a guide. Generally, this process starts from a primer and proceeds in the 5 'to 3' direction. Generally, a DNA polymerase or other polymerase is used to form the daughter strand, and the conditions are selected such that a complementary copy of the template strand is obtained. Subsequently, the δ connecting group, which is now positioned with the ε connecting group of the adjacent subunit anchor, is caused to cross-link and form a χ bond, which is an inter-subunit bond. The χ bonds provide a second inter-subunit bond (polymerized inter-substrate bonds are the first) and form an intermediate product called the "duplex daughter strand", as shown in Figure 63B.
The duplex daughter strand (Figure 63B) is a hetero-copolymer with subunits shown in brackets. The primary backbone (~ N ~) k, template strand (-N '-) k, and anchor (T) are shown as a duplexed daughter strand, where κ indicates a plurality of repeat subunits. Each subunit of the daughter strand is a repeat "motif" and the motifs have species-specific variability, indicated here by the superscript a. The daughter strand is formed from species of the monomeric substrate construct selected by a template-directed process from a library of motif species, with the monomer substrate of each species of substrate construct binding to a corresponding complementary nucleotide on the strand. target mold. In this way, the nucleobase residue sequence (ie, primary backbone) of the daughter strand is a contiguous complementary copy of the target template strand.
Each tilde (~) indicates a selectively cleavable bond shown here as inter-substrate bonds. These are necessarily selectively cleavable to release and expand the anchors (and the Xpandomer) without degrading the Xpandomer itself.
The daughter strand is composed of an Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors are converted to their "expanded configuration", the limited Xpandomer is converted to the Xpandomer product. Anchors are limited by inter-subunit χ linkages, substrate binding, and optionally intra-anchor linkages if they are still present. The χ bond joins the substrate of a first subunit with the anchor of the second terminal moiety at the confined end of a second subunit and is formed by linking the positioned connecting groups, ε of the first subunit and δ of the second subunit.
It can be seen that the daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer" backbone. The primary backbone is composed of the contiguously confined and polymerized monomeric substrates. The "limited Xpandomer backbone" avoids selectively cleavable bond between monomer substrates and is formed by χ-bonded backbone moieties, each skeleton moiety being a substrate-linked anchor which then binds to the next moiety anchor skeleton with a χ bond. It can be seen that the limited Xpandomer backbone connects via the selectively cleavable bonds of the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone fragments.
Anchor χ bonding (crosslinking of the δ and ε linking groups) is generally preceded by enzymatic coupling of the monomer substrates to form the primary backbone, with, for example, phosphodiester bonds between adjacent bases. In the structure shown here, the primary backbone of the daughter strand has been formed, and the inter-substrate bonds are represented by a tilde (~) to indicate that they are selectively cleavable. After dissociating or degrading the target template strand, cleaving the selectively cleavable bonds (including intra-anchor bonds), the limited Xpandomer is released and converted to the Xpandomer product. Methods for template strand dissociation include heat denaturation, or selective digestion with a nuclease, or chemical degradation. A selection cleavage method uses nuclease digestion in which, for example, phosphodiester linkages from the primary backbone are digested by a nuclease and anchor-to-anchor linkages are nuclease resistant.
ES 2 559 313 T3
Figure 63C is a representation of the Xpandomer class VIII product after dissociation from the template strand and after cleavage of selectively cleavable bonds (including those in the primary backbone and, if not already cleaved, intra-bonds). -anchorage). The Xpandomer product strand contains a plurality of κ subunits, in which the K subunit indicates<sup>n</sup> in a chain of m subunits constituting the daughter strand, where κ = 1, 2, 3 am, where m> 10, usually m> 50, and usually m> 500 or> 5,000. Each subunit is formed of an anchor in its expanded configuration and stretches to its length between the χ bonds of adjacent subunits. The lateral substrate is attached to the anchor in each subunit. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 63D shows the substrate construction of Figure 63A as a molecular model, in which the monomer substrate member, represented with a nucleobase residue (634), is attached to the anchor by a bond (631) of the first terminal residue of the anchor. A second connecting group (639), shown as δ in Figure 63A, is disposed on the second terminal moiety at the distal end of the anchor. A selectively cleavable intra-anchor bond (633) is shown, represented by the contiguous triangles, which limits the anchor by linking the first and second terminal moieties. Also attached to the substrate is a connecting group (638), shown as ε in Figure 63A. The connecting groups ε (638) and δ (639) are positioned not to interact and to align preferentially near the sides of R<sup>1</sup> and R<sup>2</sup>of the substrate, respectively. The anchor loop shown here has three indicators (900,901,902), which may also be species-specific for motifs. The cleavable intra-anchor crosslink (633) is shown in Figure 63D and Figure 63E which is positioned closer to the substrate than δ (639). In Figures 60D and Figure 60E, the positions are switched. This positioning can be anywhere and in one embodiment both connector functions can be on a single multifunctional group.
Figure 63E shows the substrate construction after incorporation into the product Xpandomer. The subunits are cleaved and expanded and linked by χ (970,971) bonds, represented as an open oval, formed by linking the connecting groups δ and ε referred to in Figure 63A. A subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 63C.
In the Xpandomer product of Figure 63E, the primary backbone has been fragmented and is not covalently contiguous because any direct bonds between adjacent subunit substrates have been cleaved. Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information of the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor (630) is depicted as an indicator construct with three indicator groups, various indicator constructs may be present on the anchor, and may comprise individual indicators that identify monomer or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
Class IX monomeric constructions
One class IX substrate construction is distinguished from the other RT-NTP in that it has two anchor attachment points to which a free anchor is attached after the primary backbone is assembled. Figure 64 describes class IX monomeric substrate constructs (akin to RT-NTP) in more detail. Figures 64A to 64C are read from left to right, showing first the monomeric substrate construction (Xpandomer precursor having a single nucleobase residue), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product ready for sequencing.
As shown in Figure 64A, the class IX monomeric substrate construct has a substrate nucleobase residue, N, with two anchor binding sites, the connecting groups δ · ι and δ<sub>2</sub>. Also shown is a free anchor, T (640), with the ε1 and ε2 connecting groups of a first and second terminal moiety of the anchor. The first and second terminal moieties of the anchor limit free anchoring by a selectively cleavable intra-anchor crosslinking (647) and serve to position the connecting groups ε1 and ε<sub>2</sub>. The cleavable crosslinking is indicated herein by dotted line, it may indicate, for example, a disulfide bond or a photocleavable linker. This limitation prevents the anchor from elongating or expanding and is said to be in its "limited configuration." The connecting groups, δ and δ<sub>2</sub>, are attached to the monomer substrate oriented near R<sup>1</sup> and R<sup>2</sup> respectively. Under template-directed synthesis, the substrates form a duplex with the target template and the connecting group of a substrate construct, for example δ1 and the connecting group of the confined substrate construct, for example δ<sub>2</sub>, are placed. Under controlled conditions, these placed connectors are brought into contact with the ε connectors placed at the end of the free anchor. Two selective bonding reactions, ε1 and δ1, occur to form χ<sup>1</sup> and ε2 and δ2 to form χ<sup>2</sup>; the adjacent substrates are now connected by the limited anchor. Bonding and chemistry
ES 2 559 313 T3 protection / deprotection suitable for δ, ε and χ are detailed in the general description of monomeric constructs.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-phosphate and R<sup>2</sup> = 3'-Oh, they would find use in a ligation protocol, and R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. Optionally, R<sup>2</sup> can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> they can be configured with the terminal connecting groups for chemical coupling or with non-connecting groups for a hybridization-only protocol. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
During assembly, the monomeric substrate construction (without anchor) is polymerized onto the extensible end of the nascent daughter strand by a mold-directed polymerization process using a single-stranded mold as a guide. Generally, this process starts from a primer and proceeds in the 5 'to 3' direction. Generally, a DNA polymerase or other polymerase is used to form the daughter strand, and the conditions are selected such that a complementary copy of the template strand is obtained. The connecting group 81 is now positioned with the connecting group δ2 of the adjacent monomer substrate. After polymerization of the primary skeleton, free anchors crosslink to form χ bonds<sup>1</sup> and links χ<sup>2</sup> between the two anchor ends and two of the adjacent substrates.
In one embodiment, the free anchors have no sequence information and are called "naked." In this case, single linker chemistry can be used so that connecting groups 61 and δ2 are the same and connecting groups ε1 and ε2 are the same with a type of bonding.
In the embodiment where the free anchor comprises base-type information (as in the indicators), there are species-specific free anchors. Generally, there are four types of base that would require four types of free anchor with the corresponding base information. Various methods can be used to bind the free anchor species to their correct base. In one method, four heterospecific bond chemistries are used to further differentiate the linker pair δ1 and ε1, now expressed as δ1<sub>α</sub> and ε-ia in which a is one of four types and form bond types χ<sup>1α</sup>. and δ1<sub>α</sub> of a type a will bind only to each other. In this method, the pair of connectors δ2 and ε2 are caused to bond only after Ó1 is formed.<sub>to</sub> and ε ^. In a second method, different selectively unprotectable protecting groups, in which each protecting group is associated with a type of linker, are used to selectively block the linking groups. In a first cycle, there is no protection in a base type and its associated free anchor type is ε1 bonded to δ1 to form a χ bond.<sup>1</sup>. In a second cycle, one type of protecting group is removed from one type of base and its associated free anchor type is ε1 bonded to δ1 to form a χ bond.<sup>1</sup>. This last cycle is repeated for the next two base types and after completion, the linker pair δ2 and ε2 are caused to bind. Washing steps are included between each stage to reduce bonding errors. Without loss of generality, the remaining description will be in terms of links χ<sup>1 </sup>and links χ<sup>2</sup>. Note that after the anchor binds and the χ bonds are formed, the intra-anchor bond can be broken.
As shown in Figure 64B, the links χ<sup>1</sup> they provide inter-subunit bonding (in addition to polymerized inter-substrate bonds) and form an intermediate product called the "duplex daughter strand." The primary backbone (~ N ~) k, template strand (-N '-) k, and anchor (T) are shown as a duplexed daughter strand, where κ indicates a plurality of repeating subunits. Each subunit of the daughter strand is a repeat "motif" and the motifs have species-specific variability, indicated here by the superscript a. The daughter strand is formed from species of the monomeric substrate construct selected by a template-directed process from a library of motif species, with the monomer substrate of each species of substrate construct binding to a corresponding complementary nucleotide on the strand. target mold. In this way, the nucleobase residue sequence (ie, primary backbone) of the daughter strand is a contiguous complementary copy of the target template strand.
Each tilde (~) indicates a selectively cleavable bond shown here as inter-substrate bonds. These are necessarily selectively cleavable to release and expand the anchors (and the Xpandomer) without degrading the Xpandomer itself.
The daughter strand is composed of an Xpandomer precursor called the "limited Xpandomer" which is additionally composed of anchors in the "limited configuration". When the anchors are converted to their "expanded configuration", the limited Xpandomer is converted to the Xpandomer product. The anchors are limited by the χ bonds to adjacent substrates and, optionally, the intra-anchor bonds (if they are still present).
It can be seen that the daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer" backbone. The primary backbone is composed of the contiguously confined and polymerized monomeric substrates. The "limited Xpandomer backbone" avoids selectively cleavable bond between monomer substrates and is formed by restos-linked backbone residue residues<sup>1</sup>, each skeleton remnant being a
ES 2 559 313 T3 anchor bound to a substrate (by a bond χ<sup>2</sup>) which then binds to the next backbone remainder anchor with a bond enlace<sup>1</sup>. It can be seen that the limited Xpandomer backbone connects via the selectively cleavable bonds of the primary backbone, and will remain covalently intact when these selectively cleavable bonds are cleaved and the primary backbone fragments.
Anchor χ bonding (crosslinking of the δ and ε linking groups) is generally preceded by enzymatic coupling of the monomer substrates to form the primary backbone, with, for example, phosphodiester bonds between adjacent bases. In the structure shown here, the primary backbone of the daughter strand has been formed, and the inter-substrate bonds are represented by a tilde (~) to indicate that they are selectively cleavable. After dissociating or degrading the target template strand, cleaving the selectively cleavable bonds (including intra-anchor bonds), the limited Xpandomer is released and converted to the Xpandomer product. Methods for template strand dissociation include heat denaturation, or selective digestion with a nuclease, or chemical degradation. A selection cleavage method uses nuclease digestion in which, for example, phosphodiester linkages from the primary backbone are digested by a nuclease and anchor-to-anchor linkages are nuclease resistant.
Figure 64C is a representation of the Xpandomer class IX product after dissociation from the template strand and after cleavage of selectively cleavable bonds (including those in the primary backbone and, if not already cleaved, intra-bonds). -anchorage). The Xpandomer product strand contains a plurality of κ subunits, in which the K subunit indicates<sup>n</sup> in a chain of m subunits that make up the daughter strand, where κ = 1.2, 3 am, where m> 10, usually m> 50, and usually m> 500 or> 5,000. Each subunit is made up of an anchor (640) in its expanded configuration bonded to a monomer substrate and with the χ bond<sup>1</sup> with the next adjacent subunit. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 64D shows the substrate construction of Figure 64A as a molecular model, in which the monomer substrate member, represented with a nucleobase residue (641), has two bonds (642,643), shown in Figure 64a as δ Ι and δ<sub>2</sub>, which will be points of attachment for free anchoring. The free anchor is shown with two connecting groups (644,645), shown as ε1 and ε2 in Figure 64A, on the first and second terminal residues of the anchor. A selectively cleavable intra-anchor linkage (647) is shown limiting anchoring by linking the first and second terminal residues. Linker groups are positioned to promote crosslinking of anchor ends between monomer substrates and prevent crosslinking through a monomer substrate.
The anchor loop shown here has three indicators (900,901,902), which can also be species-specific for motifs, but require a method to correctly bind to the correct bases in the primary backbone.
Figure 64E shows the substrate construction after incorporation into the product Xpandomer. Subunits split and expand and are linked by bonds χ<sup>1</sup> inter-anchors (930,931), and links χ<sup>2</sup> interchanges (932,933). Each subunit is an anchor linked to a monomer substrate and connected on the following bond χ<sup>1</sup>. A subunit is indicated by dotted lines vertically grouping the repeating subunit, as represented by the brackets in the accompanying Figure 64C.
In the Xpandomer product of Figure 64E, the primary backbone has been fragmented and is not covalently contiguous because any direct bonds between adjacent subunit substrates have been cleaved. Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information of the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor is represented as an indicator construct with three indicator groups, various indicator constructs may be present on the anchor, and may comprise individual indicators that identify monomer or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
Figure 65 demonstrates an Xpandomer synthesis method using class IX RT-NTP substrate constructs. A target template (650) is first selected and hybridized with an immobilized primer. In step I, the primer is extended by template directed synthesis of a daughter strand. This process continues in stage II, and an enlarged view (dotted arrow) of the daughter strand is shown in Figure 65a. Illustrated are template strand, primer, and polymerized class IX nucleobase substrate constructs (without anchors), each substrate construct with chemical functionalities, represented as the lock and key symbols, for the chemical addition of anchor reagents. The functionalities are selected so that each monomer has a binding site of the
ES 2 559 313 T3 specific base anchor and a universal binding site. Hairpin anchors with four species-specific connectors are introduced in stage III. The anchors are attached to the primary skeleton in a base-specific mode according to the base-specific connectors as shown in Figure 65b, an enlarged view (dotted arrow). Here, black and white circles indicate the universal chemical binding chemistry and the diamond and fork shapes indicate the base specific binding chemistries. In step IV and Figure 65c, an enlarged view (dotted arrow), after all base specific bonding is complete, causes the universal connectors to form bonds. Note that the universal connectors on both anchors and the primary backbone are in close proximity to their base specific connector counterpart to avoid linking errors. The chemical bonding of the anchoring reagents is shown to be complete, and a limited Xpandomer has formed on the template. The Xpandomer is then released (not shown) by dissociation from the template, and cleaving the selectively cleavable bonds (primary backbone and intra-anchor bonds).
Class X monomeric constructions
Class X substrate constructs, also called XNTP, differ from RT-NTP in that the anchor is contained within each monomer substrate to form an intra-substrate anchor, with each XNTP substrate having a selectively cleavable bond within the substrate that , once cleaved, allows expansion of the limited anchor. In Figure 66, the present inventors describe class X monomeric substrate constructs in more detail. Figures 66A to 66C are read from left to right, showing first the monomeric substrate construction (Xpandomer precursor having a single nucleobase residue), then the daughter strand of the intermediate duplex in the center, and on the right the Xpandomer product ready for sequencing.
As shown in Figure 66A, the class X monomeric substrate construct has a substrate nucleobase residue, N, which has two residues (662,663) separated by a selectively cleavable bond (665), each residue being attached to one end of an anchor (660). The ends of the anchor can be attached to modifications of the linking group on the heterocycle, the ribose group, or the phosphate backbone. The monomer substrate also has an intra-substrate cleavage site positioned within the phosphororibosyl backbone so that cleavage provides for limited anchor expansion. For example, to synthesize a class X ATP monomer, the amino linker on 8 - [(6-amino) hexyl] -amino-ATP or N6- (6-amino) hexyl-ATP can be used as a first point of anchor attachment, and, a mixed backbone linker, such as the non-linker modification (N-1-aminoalkyl) phosphoramidate or (2-aminoethyl) phosphonate, can be used as a second anchor attachment point. Furthermore, a linker backbone modification such as a phosphoramidate (3 'OPN 5') or a phosphorothiolate (3 'OPS 5'), for example, can be used for selective chemical cleavage of the primary backbone.
R<sup>1</sup> and R<sup>2</sup> they are end groups configured as appropriate for the synthesis protocol in which the substrate construct is used. For example, R<sup>1</sup> = 5'-triphosphate and R<sup>2</sup> = 3'-OH for a polymerase I protocol. The R<sup>1</sup> 5 'triphosphate can include mixed backbone modifications, such as an aminoethyl phosphonate or 3'-OPS-5' phosphorothiolate, to allow anchor bonding and backbone cleavage, respectively. Optionally, R<sup>2 </sup>can be configured with a reversible blocking group for cyclic addition of a single substrate. Alternatively, R<sup>1</sup> and R<sup>2</sup> They can be configured with terminal connecting groups for chemical coupling. R<sup>1</sup> and R<sup>2</sup> they can be of the general type XR, where X is a linking group and R is a functional group.
During assembly, the monomeric substrate construct is polymerized onto the extensible end of the nascent daughter strand by a mold-directed polymerization process using a single-stranded template as a guide. Generally, this process starts from a primer and proceeds in the 5 'to 3' direction. Generally, a DNA polymerase or other polymerase is used to form the daughter strand, and the conditions are selected such that a complementary copy of the template strand is obtained.
As shown in Figure 66B, the nucleobase residues polymerize one subunit with the next and form an intermediate product called the "duplex daughter strand." The primary backbone (-N-) k, template strand (-N '-) k, and anchor (T) are shown as a duplexed daughter strand, where κ indicates a plurality of repeating subunits. Each subunit of the daughter strand is a repeat "motif" and the motifs have species-specific variability, indicated here by the superscript a. The daughter strand is formed from species of the monomeric substrate construct selected by a template-directed process from a library of motif species, with the monomer substrate of each species of substrate construct binding to a corresponding complementary nucleotide on the strand. target mold. In this way, the nucleobase residue sequence (ie, primary backbone) of the daughter strand is a contiguous complementary copy of the target template strand.
The ("V") shown in Figure 66B above the nucleobase residue indicates a selectively cleavable bond that splits the substrate into the first and second residues. Following cleavage, the first residue (669) of a subunit will remain linked to the second residue (668) of the adjacent subunit and within a subunit, each residue will be connected, one to the other by the anchor. These are necessarily selectively cleavable to release and expand the anchors (and the Xpandomer) without degrading the Xpandomer itself.
The daughter strand has two backbones, a "primary backbone" and the "limited Xpandomer" backbone. The primary backbone is composed of the contiguously confined and polymerized monomeric substrates. The "limited Xpandomer backbone" prevents selectively cleavable bond within the monomer substrate and is formed by
ES 2 559 313 T3 inter-substrate bonds that link the backbone residues, each backbone residue being an anchor linked to two residues of the nucleobase residues still intact. The bounded Xpandomer backbone connects via the selectively cleavable bond within each monomer, and will remain covalently intact when these selectively cleavable bonds are cleaved and the monomers are cleaved into n-portions.<sup>1</sup> and n<sup>2</sup> shown in Figure 66C.
Cleavage is preceded by enzymatic coupling of the monomer substrates to form the primary backbone, with, for example, mixed or phosphodiester backbone linkages between adjacent bases. In the structure shown here, the primary backbone of the daughter strand has been formed. After dissociating or degrading the target template strand and cleaving the selectively cleavable bonds, the limited Xpandomer is released and converted to the Xpandomer product. Methods for dissociation of the template strand include, for example, heat denaturation.
Figure 66C is a representation of the X-class Xpandomer product after dissociation from the template strand and after cleavage of the selectively cleavable bonds. The Xpandomer product strand contains a plurality of κ subunits, where κ indicates the k subunit<sup>n</sup> in a chain of m subunits constituting the daughter strand, where κ = 1, 2, 3 am, where m> 10, usually m> 50, and usually m> 500 or> 5,000. Each subunit is made up of an anchor in its expanded configuration linked to portions n<sup>1</sup> and n<sup>2</sup> of a monomer substrate, and each subunit is linked to the next by monomer polymerization bonds. Each subunit, an a-subunit motif, contains species-specific genetic information established by the template-directed assembly of the Xpandomer intermediate (daughter strand).
Figure 66D shows the substrate construction as a molecular model, in which the nucleobase member (664) is attached to a first and second nucleobase residue, each residue having an anchor binding site (662,663). The anchor (660) comprises indicator groups (900,901,902). A selectively cleavable bond that separates the two nucleobase residues is indicated by a "V" (665).
Figure 66E shows the Xpandomer product. The subunits comprise the expanded anchor (660) attached to the nucleobase portions (669,668), shown as n<sup>1</sup> and n<sup>2</sup> in Figure 66C, each subunit linked by inter-nucleobase bonds. Through the cleavage process, the limited Xpandomer is released to become the Xpandomer product. Anchor members that were previously in the limited configuration are now in the expanded configuration, thus serving to linearly stretch the sequence information from the template target. The expansion of the anchors reduces the linear density of the sequence information throughout the Xpandomer and provides a platform to increase the size and abundance of indicators, which in turn improves the signal over noise for the detection and decoding of the template sequence.
Although the anchor is represented as an indicator construct with three indicator groups, various indicator constructs may be present on the anchor, and may comprise individual indicators that identify monomer or the anchor can be a bare polymer. In some cases, one or more reporter precursors are present on the anchor, and the reporters are affinity bound or covalently bound upon assembly of the Xpandomer product.
Figure 67 shows a method of assembling an Xpandomer with XNTP class X substrate constructs. In the first view a hairpin primer is used to prime a template and the template is contacted with a polymerase and monomer substrates of class X. It is shown in step I that polymerization extends the nascent daughter strand processively by template-directed addition of monomer substrates. The enlarged view (Figure 67a) illustrates this in more detail. The daughter strand complementary to the template strand is shown to be composed of modified nucleobase substrates with internal cleavage site ("V"). In stage II, the process of formation of the Xpandomer intermediate is completed, and in stage III, a process of cleavage, dissociation of the intermediate and expansion of the daughter strand is shown in progress. In the enlarged view (Figure 67b) it is shown that internal cleavage of the nucleobase substrates alleviates the constraint on the anchors, which expand, elongating the Xpandomer backbone.
EXAMPLE 1.
SYNTHESIS OF A “CA” CHIMERICAL XSONDA 2MERA WITH SELECTIVELY SCINDIBLE RIBOSIL-5'-3 'INTERNUCLEOTIDIC LINK
The oligomeric substrate constructs are composed of probe members and anchor members and have a general "probe-loop" construction. The synthesis of the probe member is carried out using well established solid phase oligomer synthesis methods. In these methods, the addition of nucleobases to a nascent probe strand on a resin is carried out with, for example, phosphoramidite chemistry (US Pat. Nos. 4,415,732 and 4,458,066), and milligrams or grams of synthetic oligomer can be economically synthesized using readily available automated synthesizers. Typical solid phase oligonucleotide synthesis involves repetitively performing four steps: deprotection, coupling, capping, and oxidation. However, at least one bond in the probe of a class I Xsonda substrate construct is a selectively cleavable bond, and at least two probe moieties are modified for acceptance of a member.
ES 2 559 313 T3 anchor. The selectively cleavable bond is located between the probe moieties selected for anchor attachment (ie, "between" should not be limited to what it means "between adjacent nucleobase members" because the first and second sites of the anchor attachment they only need to be positioned anywhere on a first and second residue of the probe, respectively, the residues being linked by the selectively cleavable bond). In this example, a ribosyl-5'-3 'internucleotide linkage, which is selectively cleavable by ribonuclease H, is the selectively cleavable linkage and the two anchor attachment points are the first and second nucleobase residues of a 2mer probe.
Synthesis of linker-modified Xprobes is accomplished using commercially available phosphoramidites, for example, from Glen Research (Sterling, VA, USA), BioGenex (San Ramon, CA, USA), Dalton Chemical Laboratories (Toronto, USA). Canada), Thermo Scientific (US), Link Technologies (UK), and others, or can be custom synthesized. Well-established synthesis methods can be used to prepare the disclosed probe in which the 3 'nucleobase, which for this example is a C6-amino deoxyadenosine modifier, is first attached to a universal support using 5'-dimethoxytrityl-N6-benzoyl. -N8- [6- (trifluoroacetylamino) -hex-1-yl] -8-amino-2'deoxyadenosine-3 '- [(2-cyanoethyl) - (N, N-diisopropyl)] - phosphoramidite, followed by the addition of a C6-amino cytidine modifier using 5'-dimethoxytrityl-N-dimethylformamidine-5- [N (trifluoroacetylaminohexyl) -3-acrylimido] -cytidine, 3 '- [(2-cyanoethyl) - (N, N- diisopropyl)] - phosphoramidite, in which 5'-cytidine is a ribonucleotide. The addition of a chemical phosphorylation reagent followed by standard cleavage, deprotection, and purification methods completes the synthesis. The dinucleotide product is a 3 '5'-phosphate (aminoC6-cytosine) (aminoC6-deoxyadenosine) with a centrally cleavable 5', 3 'ribosyl bond and amino linkers on each base.
The Xsonda anchor for this example is constructed from bis-N-succinimidyl- [pentaethylene glycol] ester (Pierce, Rockford IL; Product No. 21581). The linker amines of the modified pCA oligomer are cross-linked with bis (NHS) PEG5 according to the manufacturer's instructions. A product of the molecular weight expected for the circularized PEG-probe construct is obtained.
EXAMPLE 2.
SYNTHESIS OF A “TATA” 4MERA XSONDA WITH SELECTIVELY CLEAVABLE PHOSPHOROTHIOLATE LINK
4mer Xprobes can be synthesized with a phosphorothiolate bond as the selectively cleavable bond. The following example describes the synthesis of a 5 'phosphate (dT) (aminoC6-dA) (dT) (aminoC6-dA) 3' tetranucleotide.
First a 5 'mercapto-deoxythymidine is prepared as described by Mag et al. ("Synthesis and selective cleavage of an oligodeoxynucleotide containing a bridged internucleotide 5'-phosphorothioate linkage", Nucl Acids Res 19: 1437-41, 1991). Thymidine is reacted with two equivalents of p-toluenesulfonyl chloride in pyridine at room temperature, and the resulting 5'-tosylate is isolated by crystallization from ethanol. Tosylate is converted to 5 '- (S-trityl) -mercapto-5'-deoxy-thymidine with five equivalents of sodium tritylthiolate (prepared in situ). The nucleotide 5 '- (S-trityl) -mercapto-thymidine is purified and reacted with 2-cyanoethoxy-bis- (N, N-diisopropylamino-phosphane) in the presence of tetrazole to prepare the structural element of 3'-O -phosphoramidite.
To begin automated synthesis, the C6-amino deoxyadenosine modifier is first attached to a universal support using 5'-dimethoxytrityl-N6-benzoyl-N8- [6- (trifluoroacetylamino) -hex-1-yl] -8-amino- 2'-deoxyadenosine-3 '- [(2cyanoethyl) - (N, N-diisopropyl)] - phosphoramidite, followed by the addition of mercaptothymidine phosphoramidite prepared above. Before adding the next C6 amino dA phosphoramidite, the S-trityl group is first deprotected with 50 mM aqueous silver nitrate and the resin is washed with water. The resin is then normally treated with a reducing agent such as DTT to remove secondary disulfides formed during cleavage. The column is then washed again with water and acetonitrile, and the free thiol is reacted under standard conditions with the C6 amino deoxyadenosine phosphoramidite in the presence of tetrazole, thus forming "ATA" with a connecting phosphorothiolate linkage S3 '^ P5' between terminal deoxyadenosine and 3-mercapto-thymidine. In the next cycle standard deoxythymidine phosphoramidite is added. Finally, the addition of a chemical phosphorylation reagent, followed by conventional cleavage, deprotection, and purification methods, completes the synthesis. The tetranucleotide product is a 3 '5' phosphate (dT) (aminoC6-dA) (dT) (aminoC6-dA).
The phosphorothiolate bond is selectively cleavable, for example, with AgCl, acid or with iodoethanol (Mag et al, "Synthesis and selective cleavage of an oligodeoxynucleotide containing a bridged internucleotide 5'phosphorothioate linkage", Nucleic Acids Research, 19 (7): 1437-1441, 1991). Because the selectively cleavable bond is between the second and third nucleobases, the anchor is designed to connect this bond, and can bind to any two nucleobases (or any two primary backbone attachment points) on either side of the selectively cleavable bond. Methods for linker and zero linker chemistries include, for example, provision of linkers with primary amines used in oligomer synthesis, as described in Example 1. Amine-modified linkers on the "TATA" oligomer are normally protected during oligomer synthesis and are deprotected in the normal course of oligonucleotide synthesis completion.
The Xsonda anchor for this example is then constructed from bis-epoxide activated poly (ethylene glycol) diglycidyl ether (SigmaAldrich, St. Louis MO, Product No. 475696). The epoxide reaction of amines
ES 2 559 313 T3 with activated PEG end groups is performed in dilute solution to minimize any competition concatenation reaction. Similar reactivity is obtained with mixed anhydrides or even acid chlorides, and heterobifunctional linking groups can be employed to guide the attachment of the anchor. The anchored reaction products are separated by preparative HPLC and characterized by mass spectroscopy. A product of approximate molecular weight is obtained for the circularized 4mer PEG-probe construct (approximately 2.5 Kd). With this method a PEG anchor distribution is obtained with approximately Mn = 500. This corresponds to an anchor of approximately 40 Angstroms (at approximately 3.36 A / PEG unit).
EXAMPLE 3.
SYNTHESIS OF A 3MERA “CTA” X-SOUND WITH 5'-3 'LINK SELECTIVELY SCINDABLE PHOSPHODIESTER
Xprobes can also be synthesized with a phosphodiester bond as the selectively cleavable bond. The following example describes the synthesis of a 3 '5' phosphate (aminoC6-dC) (aminoC6-dT) (dA) trinucleotide with a non-connecting phosphorothioate modification. Phosphodiester bonds are attacked by a variety of nucleases. A phosphorothioate bond, with non-connecting sulfur, is used as the nuclease resistant bond in this example.
For the automated synthesis in the 3 'to 5' direction a solid support of deoxyadenosine immobilized in CPG (5'-dimethoxytrityl-N-benzoyl-2'-deoxyadenosine, 3'-succinoyl-long chain alkyl-CPG 500) is used . In the first cycle, the C6-amino-deoxythymidine phosphoramidite modifier (5'-dimethoxytrityl-5- [N- (trifluoroacetylaminohexyl) -3-acrylimido] -2'-deoxyuridine, 3 '- [(2-cyanoethyl) - (N, N -diisopropyl)] - phosphoramidite) is coupled. Before hooded, the immobilized dA is reacted with sulfurizing reagent (Glen Research, Sterling VA; Cat No. 40-4036), also known as Beaucage's reagent, following the manufacturer's protocol. The reagent is generally added through a separate port on the synthesizer. After thiolation, amino-modified C6-deoxycytidine (5'-dimethoxytrityl-N-dimethylformamidine-5- [N- (trifluoroacetylaminohexyl) -3-acrylimido] -2'-deoxycytidine, 3 '- [(2-cyanoethyl) - (N, Ndiisopropyl)] - phosphoramidite) is coupled. Prior to capping, the immobilized dC is reacted with sulfurizing reagent (Glen Research, Sterling VA; Cat No. 40-4036), also known as Beaucage's reagent, following the manufacturer's protocol. The resulting 3mer "CTA" has a phosphorothioate bond between T and A, and a phosphodiester bond between C and T. Addition of a chemical phosphorylation reagent, followed by standard, complete cleavage, deprotection and purification methods synthesis.
The resistance of phosphorothioate bonds to nuclease attack is well characterized, for example, by Matsukura et al ("Phosphorothioate analogs of oligodeoxynucleotides: inhibitors of replication and cytopathic effects of human immunodeficiency virus", PNAS 84: 7706-10, 1987), by Agrawal et al ("Oliogodeoxynucleoside phosphoramidates and phosphorothioates as inhibitors of human immunodeficiency viruses", PNAS 85: 7079-83, 1988), and in US Patent No. 5770713. Both C and T are linker or zero linker modified, with derivatization serving to attach an anchor member.
EXAMPLE 4.
SYNTHESIS OF A 3MERA “ATA” XSONDA WITH 5 'NPO 3' LINK SELECTIVELY SCINDABLE PHOSPHORAMIDATE
The C6-amino deoxyadenosine modifier is first attached to a universal support using 5'-dimethoxytrityl-N6benzoyl-N8- [6- (trifluoroacetylamino) -hex-1-yl] -8-amino-2'-deoxyadenosine-3'- [(2-cyanoethyl) - (N, N-diisopropyl)] phosphoramidite, followed in the next cycle by the addition of 5'-amino-dT blocked with MMT (5'-monomethoxytritylamino-2'-deoxythymidine). After deblocking, the 5'-amino end is reacted with amino modified C6-deoxyadenosine phosphoramidite under standard conditions. The addition of a chemical phosphorylation reagent, followed by standard cleavage, deprotection, and purification methods completes the synthesis. The 5 '(aminoC6-dA) (OPN) (dT) (OPO) (aminoC6-dA) 3' trinucleotide has a phosphoramidate bond between the 5 'aminoC6-dA and the penultimate dT.
This phosphoramidate bond is selectively cleavable under conditions where phosphodiester bonds remain intact by treating the oligomer with 80% acetic acid as described by Mag et al. ("Synthesis and selective cleavage of oligodeoxynucleotides containing non-chiral internucleotide phosphoramidate linkages", Nucl. Acids Res. 17: 5973-5988, 1989). By attaching an anchor to connect the 5'NPO phosphoramidate bond, an Xpandomer containing such dimers can be expanded by selective cleavage of the phosphoramidate bonds from the primary backbone.
EXAMPLE 5.
SYNTHESIS OF AN XSONDA 6MERA “CACCAC” WITH AN INTERNAL PHOTOSESCENDIBLE LINK
After standard 3 'to 5' synthesis with unmodified deoxycytosine and deoxyadenosine phosphoramidites, amine-modified C6 dC is coupled (5'-dimethoxytrityl-N-dimethylformamidine-5- [N- (trifluoroacetylaminohexyl) -3
ES 2 559 313 T3 acrylimido] -2'-deoxycytidine, 3 '- [(2-cyanoethyl) - (N, N-diisopropyl)] - phosphoramidite; Glen Research; Cat No. 10-1019), thus forming a CAC trimer. For the next cycle, a photo-cleavable linker (3- (4,4'-dimethoxytrityl) -1- (2-nitrophenyl) -propan-1-yl - [(2-cyanoethyl) - (N, N-diisopropyl)] - is attached phosphoramidite; Glen Research; Cat No. 10-4920). In the next cycle, a second amino-modified dC is added, followed by two final rounds of standard addition of dA and dC phosphoramidites, respectively. The resulting product, "CAC-pc-CAC" contains amino linkers at the third and fourth base positions, and can be modified by adding an anchor that connects the selectively cleavable bond formed by the photocleavable nitrobenzene construct between the two bases. modified with amino. Selective cleavage of a photo-cleavable linker modified phosphodiester backbone is disclosed, for example, by Sauer et al. ("MALDI mass spectrometry analysis of single nucleotide polymorphisms by photocleavage and charge-tagging", Nucleic Acids Research 31,11 e63, 2003), Vallone et al. ("Genotyping SNPs using a UV-photocleavable oligonucleotide in MALDI-TOF MS", Methods Mol. Bio. 297: 169-78, 2005) and Ordoukhanian et al. ("Design and synthesis of a versatile photocleavable DNA building block, application to phototriggered hybridization", J. Am. Chem. Soc. 117, 9570-9571, 1995).
EXAMPLE 6.
SYNTHESIS OF AN XMERA SUBSTRATE CONSTRUCTION
Xmeras substrate constructs are closely related in design and composition to Xsondas. An Xmera library is synthesized, for example, by 5'-pyrophosphorylation of Xprobes. Established procedures for pyrophosphate treatment of 5'-monophosphates include, for example, Abramova et al. ("A easy and effective synthesis of dinucleotide 5'-triphoshates", Bioorganic Medicinal Chemistry 15: 6549-55, 2007). In this method, the terminal monophosphate of the oligomer is activated as needed for subsequent reaction with pyrophosphate by first reacting the terminal phosphate as a salt of cetyltrimethylammonium with equimolar amounts of triphenylphosphine (Ph3P) and 2,2'-dipyridyl disulfide (PyS ) 2 in DMF / DMSO, using DMAP (4-dimethylaminopyridine) or 1Melm (1-methylimidazole) as a nucleophilic catalyst. The product is precipitated with LiClO4 in acetone and purified by anion exchange chromatography.
A variety of other methods can be considered for the robust synthesis of Xmers of 5'-triphosphate. As described by Burgess and Cook (Chem Rev 100 (6): 2047-2060), these methods include, but are not limited to, reactions using nucleoside phosphoramidites, synthesis by nucleophilic attack of pyrophosphate on activated nucleoside monophosphates, synthesis by nucleophilic attack of phosphate on activated nucleoside pyrophosphate, synthesis by nucleophilic attack of diphosphate on activated syntone phosphate, synthesis involving activated phosphites or phosphoramidites derived from nucleosides, synthesis involving the direct displacement of 5'-O- leaving groups by triphosphate nucleophiles, and biocatalytic methods. A specific method that produced polymerase compatible dinucleotide substrates uses N-methylimidazole to activate the 5'-monophosphate group; subsequent reaction with pyrophosphate (tributylammonium salt) produces triphosphate (Bogachev, 1996). In another procedure, the phosphoramidate trinucleotides have been synthesized by Kayushin (Kayushin AL et al. 1996. A convenient approach to the synthesis of trinucleotide phosphoramidites. Nucl Acids Res 24: 3748-55).
EXAMPLE 7.
SYNTHESIS OF XANDOMERS WITH POLYMERASE PHOSPHORAMIDATE CLEAVAGE
The synthesis of an Xpandomer is carried out using a substrate construct prepared with 5'-triphosphate and 3'-OH ends. US Patent No. 7,060,440 to Kless describes the use of polymerases to polymerize triphosphate oligomers, and the method is adapted herein for the synthesis of Xpandomers. The substrate construct consists of a 2mer probe member "pppCA" with selectively cleavable 5 'NPO inter-nucleotide phosphoramidate linkage and a PEG linker loop construct. An accompanying primer and template strand is synthesized and purified before use. The sequence: "TGTGTGTGTGTGTGTGTGTGATCTACCGTCCGTCCC" is used as a template. The sequence "5'GGGACGGACGGTAGAT" is used as a primer. A 5 'end HEX (5'hexachlorofluorescein) is used on the primer as a label. Primer and template hybridization forms a template with duplex primer and single-stranded protruding nucleotide from the free 3'-OH and 5 'ends of twenty bases in length. Substrate constructs and Sequenase ™ brand recombinant T7 DNA polymerase (US Biochemicals Corp., Cleveland, OH) are then added and polymerization continues for 30 min under conditions adjusted for optimal polymerization. A sample from the polymerization reaction is mixed with gel loading buffer and electrophoresed on a 20% TBE-acrylamide gel (Invitrogen, USA) along with a non-polymerase negative control and a marker of MW to confirm polymerization of Xmeros.
The intermediate Xpandomer product is treated with 80% acetic acid for 5 h at room temperature according to the procedure of Mag et al. ("Synthesis and selective cleavage of oligodeoxyribonucleotides containing non-chiral internucleotide phosphoramidate linkages", Nucl. Acids Res., 17: 5973-88, 1989) to selectively cleave phosphoramidate bonds, which is also confirmed by electrophoresis.
ES 2 559 313 T3
EXAMPLE 8.
SYNTHESIS OF XPANDOMERS WITH PHOSPHOROTHIOATE CLEARANCE BY POLYMERASE
The synthesis of an Xpandomer is performed using a substrate construct prepared with 5'-triphosphate and 3'-OH ends. US Patent No. 7,060,440 to Kless describes the use of polymerases to polymerize triphosphate oligomers. The substrate construct is a "pppCA". The substrate construct is designed with a selectively cleavable inter-nucleotide phosphorothiolate backbone linkage and a PEG 2000 anchor loop construct. A template strand and accompanying primer are synthesized and purified before use. The sequence: "TGTGTGTGTGTGTGTGTGTGATCTACCGTCCGTCCC" is used as the target template. The sequence "5'GGGACGGACGGTAGAT" is used as a primer. A 5 'end HEX (5'hexachloro-fluorescein) is used on the primer as a label. Primer and template hybridization forms a template with duplex primer and single-stranded protruding nucleotide from the free 3'-OH and 5 'ends of twenty bases in length. The substrate and Therminator ™ DNA polymerase constructs (New England Biolabs, USA) are then added with buffer and salts optimized for polymerization. Polymerization continued for 60 min under conditions adjusted for optimal polymerization. A sample from the polymerization reaction is mixed with gel loading buffer and electrophoresed on a 20% TBE-acrylamide gel (Invitrogen, USA) along with a non-polymerase negative control and a marker of MW to confirm polymerization of Xmeros.
The phosphorothiolate bonds of the Xpandomer intermediate are selectively cleavable, for example, with AgCl, acid, or with iodoethanol (Mag et al, "Synthesis and selective cleavage of an oligodeoxynucleotide containing a bridged intemucleotide 5'-phosphorothioate linkage", Nucleic Acids Research , 19 (7): 1437-1441, 1991). Cleavage is confirmed by gel electrophoresis.
EXAMPLE 9.
SYNTHESIS OF CHIMERICAL XPANDOMERS WITH LIGASE
The synthesis of an Xpandomer is performed using a substrate construct prepared with 5'-monophosphate and 3'-OH ends. The substrate construct is a chimeric "5 'p dC rA ~ dC dA 3"' 4mer in which the penultimate 5 'adenosine is a ribonucleotide and the remainder of the substrate is deoxyribonucleotide. The substrate construct is designed with a selectively cleavable inter-nucleotide 5'-3 'phosphodiester ribosyl linkage (as shown by and a PEG 2000 anchor loop construct, in which the anchor is attached to the "C" and Terminal "A" of the 4mer. A template strand and accompanying primer is synthesized and purified before use. The sequence: "TGTGTGTGTGTGTGTGTGTGATCTACCGTCCGTCCC" is used as the target template. The sequence "5'GGGACGGACGGTAGAT" is used as the primer. A 5 'end HEX (5'hexachloro-fluorescein) is used on the primer as a label. Primer and template hybridization forms a template with duplex primer and single-stranded protruding nucleotide from the free 3'-OH and 5 'ends of twenty bases in length. T4 DNA ligase and substrate constructs (Promega Corp, Madison, WI, USA; Cat No. M1801) are then added with optimized temperature, buffer and salts for transient probe ligation and hybridization. The ligation continues for 6 hours. A sample from the ligation reaction is mixed with gel loading buffer and electrophoresed on a 20% TBE-acrylamide gel (Invitrogen, Carlsbad, CA, USA) along with a non-ligase negative control. and a MW marker to confirm X-probe ligation.
In a second step, the Xpandomer intermediate is treated with ribonuclease H to cleave the RNase-labile 5'-3 'phosphodiester bond to produce an Xpandomer product which is also confirmed by gel electrophoresis.
EXAMPLE 10.
PREPARING AN ALPHA-PHOSPHATE CONNECTOR CONSTRUCTION
Anchor linkers can also be attached to phosphorothioate diester S or a phosphoramidate N-amide, as discussed by Agrawal ("Site specific functionalization of oligonucleotides for attaching two different reporter groups", Nuc. Acids Res. 18: 5419-23, 1990). A method of functionalization of two different internucleotide backbone linkages is disclosed: aminohexyl phosphoramidate (N-1 amino alkyl) linker and phosphorothioate. The C6 amine prepared as described by Agrawal is used as a linker for the synthesis of an inter-nucleobase anchor construct of the present invention. Derivatization of N3'-P5 'bonds has also been reported (Sinyakov et al., "Functionalization of the oligonucleotides containing an internucleotide phosphoramidate bond", Russian J Bioorganic Chem, 29: 100-102, 2003).
EXAMPLE 11.
HETEROBIFUNCTIONAL PATHWAYS FOR ANCHORING CONSTRUCTION
The synthesis of modified oligomers containing a selectively cleavable bond are described in Examples 1-5. Here, a 4mer with a C6-amino modified base in the second position and 4-formylbenzoate linker modified base in the third position of the 4mer is prepared by synthetic oligomer chemistry.
ES 2 559 313 T3 standard. The second and third bases are separated by a selectively cleavable bond selected from ribosyl 5'-3 'phosphodiester bond, deoxyribosyl 5'-3' phosphodiester bond, phosphorothiolate bond (5 'OPS 3' or 5 'SPO 3'), phosphoramidate bond (5 'OPN 3' or 5 'NPO 3'), or a photoclearable bond. The amine of the C6 linker is reacted with sulfo-EGS to form an active ester NHS group. The synthesis of heterobifunctional anchor loops is then carried out by reacting the modified probe member with an anchor functionalized with hydrazide and amine terminal end groups. A final circularized product is obtained in the anchor.
EXAMPLE 12.
COLOR MARKED ANCHOR CONSTRUCTION
A single reporter Xsonda substrate construct is prepared. A 2mer is first synthesized with C6 amines on the first and second bases using conventional methods. The first and second bases are separated by a selectively cleavable bond selected from ribosyl 5'-3 'phosphodiester bond, deoxyribosyl 5'-3' phosphodiester bond, phosphorothiolate bond (5 'OPS 3' or 5 'SPO 3'), phosphoramidate bond (5 'OPN 3' or 5 'NPO 3'), or a photocleavable bond. In this example, the anchor is a species-specific end-functionalized PEG 2000 molecule with a single internal linking group, which in this case is a maleimide-functionalized, centralized, localized linker anchor. The indicator is a dye-labeled dendrimer attached to the maleimido linking group on the anchor via a sulfhydryl moiety on the dendrimer. The resulting indicator construct thus contains an indicator on the anchor. A cystamine dendrimer (Dendritic Nanotechnologies, Mt Pleasant, MI, USA; Cat No. DNT-294 G3) with a diameter of 5.4 nm and 16 surface amines per dendrimer medium) is used as an indicator. By attaching specific dyes, or combinations of dyes, to the amine groups on the dendrimer, the substrate construct species is uniquely marked for identification.
EXAMPLE 13.
PEPTIDE MARKED ANCHOR CONSTRUCTION
A single reporter Xsonda substrate construct is prepared. A 2mer is first synthesized with C6 amines on the first and second bases using conventional methods. The first and second bases are separated by a selectively cleavable bond selected from ribosyl 5'-3 'phosphodiester bond, deoxyribosyl 5'-3' phosphodiester bond, phosphorothiolate bond (5 'OPS 3' or 5 'SPO 3'), phosphoramidate bond (5 'OPN 3' or 5 'NPO 3'), or a photocleavable bond. In this example, the anchor is a species-specific end-functionalized PEG 2000 molecule with a single internal linking group, which in this case is a maleimide-functionalized, centralized, localized linker anchor. The indicator is a dye-labeled dendrimer attached to the maleimido linking group on the anchor via a sulfhydryl moiety on the dendrimer. The resulting indicator construct thus contains an indicator on the anchor. A cystamine dendrimer (Dendritic Nanotechnologies, Mt Pleasant, MI, USA; Cat No. DNT-294 G3) with a diameter of 5.4 nm and 16 surface amines per dendrimer medium) is used as an indicator. By attaching specific dyes, or combinations of dyes, to the amine groups on the dendrimer, the substrate construct species is uniquely marked for identification. By binding specific peptides to the dendrimer, which is amine functionalized, the substrate construct species is labeled for further identification. The dimensions and charge of the bound peptides are used as detection characteristics in a detection apparatus.
EXAMPLE 14.
HETEROBIFUNCTIONAL WAY FOR INDICATOR ANCHORING CONSTRUCTIONS
Substrate constructs with multiple indicators are prepared. A library of 2mers with C6-amino modified base in the first position and 4'-formylbenzoate (4FB) modified base in the second position of the 2mer is prepared by standard organic chemistry. The first and second bases are separated by a selectively cleavable bond selected from ribosyl 5'-3 'phosphodiester bond, deoxyribosyl 5'-3' phosphodiester bond, phosphorothiolate bond (5 'OPS 3' or 5 'SPO 3'), phosphoramidate bond (5 'OPN 3' or 5 'NPO 3'), or a photoclearable bond. The amine is reacted with sulfo-EGS to form an active ester NHS group. In a second step, a species-specific bifunctional amine reporter segment (segment 1) is reacted with the active NHS group and a species-specific bifunctional hydrazide reporter segment (segment 4) is reacted with 4FB. The free amine on segment 1 is then reacted with sulfo-EGS to form an active ester NHS. A species-specific heterobifunctional hood consisting of a pair of reporter segments (segments 2 and 3) with amine end groups and 4FB is then reacted with the construct, closing the anchor loop.
The resulting flag construction thus contains four flags about the anchor construction. In this example, the indicators are polyamine dendrons with a cystamine linker to covalently attach to each polymer segment by thioether bonds (Dendritic Nanotechnologies, Mt Pleasant MI, USA; Cat No. DNT-294: G4, 4, 5 nm in diameter, 32 surface amines per dendrimer) and the polymer members are end-functionalized PEG 2000, each with an internal linking group. Each anchor segment comprises a
ES 2 559 313 T3 single indicator. With four directionally coupled segments, 2 are available<sup>4</sup> Possible combinations of indicator code.
EXAMPLE 15.
HETEROBIFUNCTIONAL ROUTE FOR CONSTRUCTION OF INDICATOR WITH POST-SYNTHESIS MARKING
Substrate constructs with multiple indicators are prepared. A 4mer with C6-amino modified base in the first position and 4'-formylbenzoate modified base (4FB) in the second position of the 2mer is prepared by standard organic chemistry. The second and third bases are separated by a selectively cleavable bond selected from ribosyl 5'-3 'phosphodiester bond, deoxyribosyl 5'-3' phosphodiester bond, phosphorothiolate bond (5'O-PS 3 'or 5' SPO 3 '), phosphoramidate bond (5 'OPN 3' or 5 'NPO 3'), or a photocleavable bond. The amine is reacted with sulfo-EGS to form an active ester NHS group. In a second step, a species-specific bifunctional amine reporter segment (segment 1) is reacted with the active NHS group and a species-specific bifunctional hydrazide reporter segment (segment 4) is reacted with 4FB. The free amine in segment 1 is then reacted with sulfo-EGS to form an active ester NHS. Then, a species-specific heterobifunctional hood consisting of a pair of indicator segments (segments 2 and 3) with amine end groups and 4FB is reacted with the construct, closing the anchor loop.
The resulting flag construction thus contains four flags about the anchor construction. This example describes the post-marking of the anchor construction. Covalently attached to each anchor segment is a 16mer oligomer that is used for indicator attachment. Each anchor segment is composed in part of a functionalized PEG molecule. The indicator is dye-labeled polyamine dendrimer with a cystamine linker (Dendritic Nanotechnologies, Mt Pleasant MI, USA; Cat No. DNT-294: G5, 4.5 nm diameter, 64 surface amines per dendrimer medium) for coupling with the 16mer oligomer probe. Upon assembly of the Xpandomer, the dye-labeled dendrimers hybridize to the oligomeric anchor segments. This labeling approach is analogous to the method described by DeMattei et al. ("Designed Dendrimer Syntheses by Self-Assembly of Single-Site, ssDNA Functionalized Dendrons", Nano Letters, 4: 771-77, 2004).
EXAMPLE 16.
INDICATOR ELEMENTS MARKED WITH COLOR
Referring to Examples 14 and 15, the surface amines of the dendrimeric indicator elements are labeled by active ester chemistry. Alexa Fluor 488 (green) and Alexa Fluor 680 (red) are available for one-step binding as active esters of sNHS from Molecular Probes (Eugene OR). The density and ratio of the dyes are varied to produce a distinctive molecular mark on each indicator element.
EXAMPLE 17.
MULTI-STATUS INDICATOR ELEMENTS
Various dye palettes are selected by techniques similar to those used in M-FISH and SKY, as known to those skilled in the art, and conjugated to a dendrimer indicator. Thus, the reporter element of each anchor constitutes a "spectral direction", whereby a single dendrimeric construct with a multiplicity of dye-binding sites can create a plurality of reporter codes. With reference to the anchor construction described in Examples 14 and 15, a 5-state spectral direction produces 625 indicator code combinations.
EXAMPLE 18.
PEG-5000 SPACER ANCHOR
Anchor segments are constructed from a durable aqueous / organic solvent soluble polymer that possesses little to no binding affinity for SBX reactants. Modified PEG 5000 is used for the flexible anchor spacers flanking a poly-lysine 5000 indicator. The free PEG ends are functionalized for binding to the probe member. The polymer is circularized by crosslinking to probe binding sites using heterobifunctional linker chemistry.
EXAMPLE 19.
PREPARING A GROUND MARK INDICATOR COMPOSITION
An anchor is synthesized as follows: Cleavable mass labels are covalently coupled to poly-l-lysine peptide functionalized G5 dendrimers (Dendritic Nanotechnologies, Mt Pleasant MI) using standard amine coupling chemistry. The G5 dendrimers are ~ 5.7 nm in diameter and provide 128 reactive surface groups. A strip of ten G5 dendrimers functionalized with 10,000 polylysine peptides
ES 2 559 313 T3 molecular weight provides approximately 100,000 reporter binding sites together on a ~ 57 nm dendrimer segment. A total of approximately 10,000 mass marks are available for detection on the fully assembled segment, assuming only 10% occupancy of the available binding sites. Using the previously described 3-mass mark coding method (Figure 37), approximately 3,300 copies of each mass mark are available for measurement. Alternatively, a single G9 dendrimer (with 2048 reactive groups) functionalized with 10,000 molecular weight poly-lysine has available approximately 170,000 mass tag binding sites in a 12 nm segment of the reporter construct.
To detect a sequence with mass marks, a method of controlled release of the mass mark indicators is used at the point of measurement by use of photo-excisable connectors. Sequential fragmentation of the anchor is not necessary. The mass label indicators associated with each subunit of the Xpandomeric polymer are measured in one step. For example, a set of 13 mass brand indicators ranging from 350 Daltons to 710 Daltons (that is, a 30 Dalton scale of mass brand indicators) has 286 combinations of three mass brands each. Thus, any one of the 256 different 4mers is associated with only one particular combination of 3 mass brand indicators. The Xpandomer encoded sequence information is easily detected by subunit mass spectroscopy. Because the Xpandomer subunits are spatially well separated, the Xpandomer detection and manipulation technology need not be highly sophisticated.
EXAMPLE 20.
DIRECT ANALYSIS OF UNMARKED SUBSTRATE CONSTRUCTIONS
A library of substrate constructs is synthesized; anchors do not contain indicators. Following the preparation of an Xpandomer product, the individual bases of the Xpandomer product are analyzed by electron tunnel spectroscopy.
EXAMPLE 21.
HYBRIDIZATION ASSISTED ANALYSIS OF UNMARKED SUBSTRATE CONSTRUCTIONS
A library of oligomeric substrate constructs is synthesized; anchors do not contain indicators. Following the preparation of an Xpandomer product, a full set of labeled probes are then hybridized to the Xpandomer product and the duplexed probes are analyzed sequentially.
EXAMPLE 22.
SYNTHESIS OF A DEOXIADENOSINE TRIPHOSPHATE WITH ANCHORAGE OF LYS-CYS-PEG-POLIGLUTAMATOPEG-CYS-COOH
Lysine with a BOC-protected amino side chain is immobilized on a resin and reacted with a cysteine residue using standard peptide synthesis methods. The amino side chain of lysine will be the epsilon reactive functional group of the RT-NTP anchor (see class VI, VII). Cysteine will form a first half of an intra-anchor disulfide bond. The unprotected amine on the cysteine is modified with SANH (PierceThermo Fisher, USA; Cat No. 22400: Bioconjugate Toolkit) to form a hydrazide.
Separately, a spacer segment is prepared from bis-amino PEG 2000 (Creative PEGWorks, Winston Salem NC; Cat No. PSB 330) by functionalization of free amines with C6-SFB (Pierce-Thermo Fisher, USA. .; Cat No. 22400: Bioconjugate Toolkit), forming a bis-4FB PEG spacer segment; the product is purified.
The bis-4FB PEG spacer segment is then coupled to the hydrazide linker on the cysteine, leaving a 4FB group as a terminal reactive group, and the resin is washed.
Separately, a polyglutamate segment (each glutamate derivatized at the gamma-carboxyl with 5 PEO units of methyl-capped PEG) is prepared. The C-terminus is converted to an amine with the coupling agent EDC and diaminohexane. SANH is used to form a dihydrazide-terminated polyglutamate segment, and the product is purified. The dihydrazide-terminated polyglutamate segment is reacted with the terminal 4Fb group on the resin, forming a terminal hydrazide, and the resin is washed.
Separately, a PEG-2000 spacer segment (amine and carboxyl terminated, Creative PEGWorks; PHB-930) is reacted with SFB to generate a 4FB end group. This spacer segment is reacted with the hydrazide group on the resin, forming a carboxyl-terminated chain. The resin is washed off again.
A cysteine residue is coupled to the free carboxyl using standard reactive peptide synthesis. The terminal carboxyl of cysteine is protected with O-benzyl. The resulting product is washed again and then cleaved from the resin. The free carboxyl generated by cleavage is then modified with EDC, aminohexyl and SANH
ES 2 559 313 T3 to form a reactive hydrazide.
Separately, a C6 amine modified deoxyadenosine triphosphate (N6- (6-amino) hexyl-dATP, Jena Bioscience, Jena De; Cat No. NU-835) is treated with SFB (Pierce Bioconjugate Toolkit, Cat No. 22419) to form a functional group 4FB. By combining the modified base with the reactive hydrazide from the preceding steps, an anchor-probe substrate construct is assembled. BOC from the lysine side chain is removed prior to use. Under generally oxidizing conditions, cysteines associate to form an intra-anchor disulfide bond, stabilizing the anchor in a limited compact form.
EXAMPLE 23.
SYNTHESIS OF A RT-NTP TRIPHOSPHATE LIBRARY
The anchored triphosphate nucleotide bases A, T, C and G with intra-anchor -SS- linkage are prepared as described in Example 22, but the charge and physical parameters of the PEGylated glutamate segments used for each base are selected. to provide a different indicator feature.
EXAMPLE 24.
SYNTHESIS OF XPANDOMERS BY SBE USING RT-NTP ANCHORED WITH POLYGLUTAMATE
RT-NTP modified adenosine and guanosine nucleotide triphosphates with intraanchor disulfide bonds are prepared. The bases are further modified so that they are reversibly blocked at the 3 'position. Allyl-based reversible blocking chemistry is as described by Ruparel ("Design and synthesis of a 3'-O-allyl photocleavable fluorescent nucleotide as a reversible terminator for DNA sequencing by synthesis" PNAS, 102: 593237, 2005). Modified base anchors are constructed with delta functional group and epsilon functional group generally as shown in Figure 61. The delta functional group is a carboxyl of a cysteine side of the anchor and the epsilon functional group is an amine side chain of a lysine near the junction of the anchor with the purines. The anchors are further modified so that they contain nucleobase-specific modified polyglutamate segments.
The sequence "TCTCTCTCTCTCTCTCATCTACCGTCCGTCCC" is used as a template. The sequence "5'GGGACGGACGGTAGAT" is used as a primer. A 5 'end HEX (5'hexachloro-fluorescein) is used on the primer as a label. The Xpandomer synthesis method is essentially as described for Figure 61. In a first SBE priming cycle, the modified nucleobase is added with Klenow polymerase under conditions adapted for polymerization and a single base with the nascent daughter strand is added on the 3'-OH end of the primer. Because the substrate construction is locked in the 3 'position, no further polymerization occurs.
The amino side chain linking group (epsilon) on the first RT-NTP added is capped and will remain so during the SBE reaction. The terminal carboxyl group of the anchor is deprotected and the 3'OH on the substrate is unblocked; the complex is washed before the next round of SBE.
In a second cycle of SBE, another nucleobase is polymerized with the nascent Xpandomer intermediate. The χ bond is formed between the free amine of the epsilon linking group on the first nucleobase and the carboxyl linking group on the second nucleobase anchor using EDC and sulfo-NHS as a cross-linking agent (Pierce Cat Nos. 22980 and 24510). The carboxyl on the delta connecting group of the anchor is deprotected and the 3 'OH on the substrate is unblocked; the complex is washed before a next round of SBE.
The SBE cycle can be repeated multiple times, thus forming an Xpandomer intermediate in the limited configuration. Each anchor in the growing chain of χ-linked nucleobases is in the limited Xpandomer configuration.
EXAMPLE 25.
SPLITTING BY NUCLEASA AND TCEP TO FORM X-CLASS XANDOMER
The Xpandomer intermediate of Example 24 is nuclease cleaved, forming an Xpandomer product composed of individual nucleobases linked by anchor segments and χ linkages. The nuclease also degrades the template and any associated primers, releasing the product. The intra-anchor disulfide bonds are cleaved by the addition of a reducing agent (TCEP, Pierce Cat. No. 20490).
The Xpandomer product is filtered and purified to remove truncated tunes and by-products of nuclease digestion. Detection and analysis of linearized Xpandomer can be done using a wide variety of existing and future generation methods.
ES 2 559 313 T3
EXAMPLE 26.
SYNTHESIS OF A DEOXIADENOSINE TRIPHOSPHATE WITH AN INTRA-ANCHORED PHOTOSESCENDABLE CONNECTOR
Glycine was immobilized on a resin and reacted with a cysteine. The amino group of cysteine is then deprotected and reacted with a glutamate, glutamate with a modified side chain with a photolabile linker terminating in an OBenzyl-protected carboxyl, such as a 2-nitroveratrylamine linker adapted from described by Holmes et al. ("Reagents for combinatorial organic synthesis: development of a new O-nitrobenzyl photolabile linker for solid phase synthesis", J Org Chem, 60: 2318-19, 1995). Cysteine will be the "epsilon functional group" of the RT-NTP anchor. The unprotected amine on glutamate is modified with SANH (Pierce-Thermo Fisher, USA; Cat No. 22400: Bioconjugate Toolkit) to form a hydrazide. The glutamate side chain will form a photo-cleavable intra-anchor linker upon synthesis of the anchor.
Separately, a spacer segment is prepared from bis-amino PEG 2000 (Creative PEGWorks, Winston Salem, NC, USA; Cat No. PSB 330) by functionalizing the free amines with C6-SFB (Pierce Bioconjugate Toolkit, Cat No. 22423), forming a bis-4FB PEG spacer segment, and the product is purified. The bis-4FB PEG spacer segment is then coupled to the hydrazide linker on glutamate, leaving a 4FB group as the terminal reactive group, and the resin is washed.
Separately, a polyglutamate segment (with t-butyl protected side chains) is prepared. The C-terminus is converted to an amine with the coupling agent EDC and diaminohexane. SANH is used to form a di-hydrazide terminated polyglutamate segment, and the product is purified. The dihydrazide-terminated polyglutamate segment is reacted with the terminal 4Fb group on the resin, forming a terminal hydrazide on the resin, and the resin is washed.
A PEG-2000 spacer segment (with free amino and FMOC-protected amino ends; Cat No. PHB-0982, Creative PEGWorks) is modified with C6 SFB to form a 4FB and FMOC-amino modified PEG spacer segment. The end of 4FB is reacted with the hydrazide group on the resin, forming a FMOC-amino terminated chain. The resin is washed off again.
A lysine residue is coupled to the free amine of the spacer via a peptide bond. The lysine residue is protected on the side chain by BOC and the alpha-amine of lysine is protected by FMOC. The OBenzyl terminal carboxyl of the photo-cleavable linker and the BOC-protected side chain of lysine are then deprotected and cross-linked with EDC / sulfo-NHS to circularize the anchor.
The resulting product is washed again and then cleaved from the resin. The free glycine carboxyl generated by the cleavage is then modified with EDC, aminohexyl, and SANH to form a reactive hydrazide.
Separately, a C6 amine modified deoxyadenosine triphosphate (N6- (6-amino) hexyl-dATP, Jena Bioscience, Jena De; Cat No. NU-835) is treated with SFB (Pierce Bioconjugate Toolkit, Cat No. 22419) to form a functional group 4FB. By combining the modified base with the reactive hydrazide from the preceding steps, an anchor-probe substrate construction is assembled. The photosispendable intra-anchor connector stabilizes the anchor in a limited compact form. FMOC is then removed and the free terminal amine is reacted with sulfo-EMCS (Pierce; Cat No. 22307) to introduce a terminal maleimido functional group.
EXAMPLE 27.
SYNTHESIS OF XPANDOMERS BY SBE USING PHOTO-ESSENTIAL RT-NTP
Modified RT-NTPs adenosine and guanosine nucleotide triphosphates with photocleavable intra-anchor linkages are prepared. The bases are further modified so that they are reversibly blocked at the 3 'position. Allyl-based reversible blocking chemistry is as described by Ruparel ("Design and synthesis of a 3'O-allyl photocleavable fluorescent nucleotide as a reversible terminator for DNA sequencing by synthesis", PNAS, 102: 5932-37, 2005) . Modified base anchors are constructed with delta functional group and epsilon functional group generally as shown in Figure 61. The delta connecting group is an amine of a lysine terminal side of the anchor and the epsilon connecting group is a sulfhydryl of a cysteine near the anchor attachment point. The anchors are further modified so that they contain species-specific modified polyglutamate segments.
The sequence "TCTCTCTCTCTCTCTCATCTACCGTCCGTCCC" is used as a template. The sequence "5'GGGACGGACGGTAGAT" is used as a primer. A 5 'end HEX (5'hexachloro-fluorescein) is used on the primer as a label. The Xpandomer synthesis method is essentially as described for Figure 61. In a first SBE priming cycle, the modified nucleobase (A) is added with Klenow polymerase under conditions adapted for polymerization and a single base is added with the nascent daughter strand on the 3'-OH end of the primer. Because the substrate construction is blocked at the 3 'position, no further polymerization occurs. The immobilized primer-template complex is then washed to remove unreacted substrate.
ES 2 559 313 T3
The sulfhydryl (epsilon) side chain linking group on the first added RT-NTP is capped and will remain so during the SBE reaction. The terminal amino group of the anchor is deprotected and the 3'OH on the substrate is unblocked; the complex is washed before the next round of SBE.
In a second cycle of SBE, another nucleobase (G) is polymerized with the nascent Xpandomer intermediate. The χ bond is formed between the amino (delta connecting group) on the first nucleobase and the sulfhydryl (epsilon connecting group) on the second base anchor using the GMBS crosslinking reagent (Pierce; Cat No. 22309). The delta amino linking group on the second RT-NTP is deprotected and the 3'-OH of the substrate is unblocked; the complex is washed before the next round of SBE.
The SBE cycle can be repeated multiple times, thus forming an Xpandomer intermediate in the limited configuration. Each anchor in the growing chain of χ-linked nucleobases is in the limited Xpandomer configuration.
EXAMPLE 28.
NUCLEASA AND PHOTOSCISSION TO FORM X-CLASS XANDOMER
The Xpandomer intermediate of Example 27 is nuclease cleaved, forming an Xpandomer product composed of individual nucleobases linked by anchor segments and χ linkages. The nuclease also degrades the template and any associated primers, releasing the product. Intra-anchor photocleavable bonds are cleaved by exposure to UV light.
The Xpandomer product is filtered and purified to remove truncated tunes and by-products of nuclease digestion. Detection and analysis of linearized Xpandomer can be done using a wide variety of existing and future generation methods.
EXAMPLE 29.
SYNTHESIS OF AN IN SITU RT-NTP ANCHOR
Using standard peptide synthesis methods on a solid support, a peptide is prepared having the structure (Resin-C ') - Glu-Cys- (Gly-Ala) 10-Pro-Ser-Gly-Ser-Pro- (Ala -Gly) 10-Cys-Lys. The terminal amine is reacted with SANH (Pierce, Cat No. 22400) to create a hydrazide linker.
Separately, a C6 amine modified deoxyadenosine triphosphate (N6- (6-amino) hexyl-dATP, Jena Bioscience, Jena De; Cat No. NU-835) is treated with SFB (Pierce Bioconjugate Toolkit, Cat No. 22419) to form a functional group 4FB. By combining the modified base with the reactive hydrazide from the preceding steps, an anchor-probe substrate construct is assembled.
The construct is then cleaved from the resin. After deprotection and under generally oxidizing conditions, the cysteines associate to form an intra-anchor disulfide bond, stabilizing the beta hairpin, which contains a terminal free carboxyl (a lateral delta connecting group on the anchor) and a lysine near the point of attachment. bonding of the anchor (the amine side chain and epsilon connecting group).
Disulfide is representative of the intra-anchor stabilization depicted in class II, III, VI, VII and VIII substrate constructs (see Figures 8 and 9), although illustrated herein with more specific reference to class II, III, VI, VII and VIII substrate constructs, although illustrated herein with more specific reference to class II, III, VI, VII and VIII substrate constructs. VI, VII and VIII. The length of the unfolded anchor, assuming a residue CC peptide bond length of 3.8 A, is approximately 10 nm, but assumes a compact shape due to hydrogen bonding at the beta hairpin.
As described by Gellman ("Foldamers, a manifesto", Acc Chem Res 31: 173-80, 1998), a wide variety of polymers, not simply peptides, can be folded into compact forms. Such polymers include oligopyridines, polyisocyanides, polyisocyanates, poly (triarylmethyl) methacrylates, polyaldehydes, polyproline, RNA, oligopyrrolinones, and oligoureas, all of which have exhibited the ability to fold into compact secondary structures and expand under suitable conditions. Thus, the peptide examples presented here are representative of a much larger class of anchoring chemistries, where limitations to unexpanded anchoring can include hydrogen bonding and hydrophobic interactions, eg, in addition to intra-anchor crosslinks.
EXAMPLE 30.
SYNTHESIS OF XPANDOMERS USING XWONDERS
In one embodiment of SBX, a library of 256 X-probe 4-mer Xprobes is presented to an elongated anchored single-stranded DNA target for hybridization. The hybridization step continues under routine precise thermal cycling to promote long Xprobe chains. Loosely bound non-specific probe-target duplexes are removed by a simple washing step, again under precise thermal control. Enzymatic ligation is performed to bind any X-probe strands along the target DNA, followed by a second wash. By repeating the hybridization / wash / ligation / wash cycle, longer linked sequences grow in multiples
ES 2 559 313 T3 loci along the target DNA until replication of the target template is generally complete.
Unfilled gaps along the target DNA are filled using a well-established DNA polymerase and ligase-based gap filling process (Lee, "Ligase Chain Reaction", Biologicals, 24 (3): 197-199, 1996). The nucleotides incorporated into the gaps have a unique reporter code to indicate a gap nucleotide. The completed Xpandomer intermediate, which is composed of the original DNA target with complementary duplexed and ligated Xprobes with occasional 1, 2, or 3 nucleotide hole charges, is cleaved to produce an Xpandomer. The cleavable linker for this example is a backbone modification of the 3 'OP-N 5' substrate. Selective cleavage is catalyzed by the addition of acetic acid at room temperature.
The Xpandomer is filtered and purified to remove truncated products and subsequently elongated to form a linear structure of linked reporter codes. Detection and analysis of the Xpandomer product can be done using a wide variety of existing methods.
EXAMPLE 31.
SYNTHESIS OF XNTP XPANDOMERS USING POLYMERASE
The synthesis of a class X Xpandomer is performed using a modified substrate construct 8 - [(6-amino) hexyl] -amino-deoxyadenosine triphosphate having a mixed backbone consisting of a non-connecting 2-aminoethyl phosphonate and a connecting phosphorothiolate (3 ' OP-S 5 ') in alpha phosphate. An intra-nucleotide anchor is attached to the 2-aminoethyl phosphonate linker and to a C6 amino linker on 8 - [(6-amino) hexyl] -aminodeoxyATP. The sequence "TTTTTTTTTTTTTTTTTTTATCTACCGTCCGTCCC" is used as a template. The sequence "5'GGGACGGACGGTAGAT" is used as a primer. A 5 'end HEX (5'hexachloro-fluorescein) is used on the primer as a label. Primer and template hybridization forms a template with duplex primer and single-stranded protruding nucleotide from the free 3'-OH and 5 'ends of twenty bases in length. The substrate and polymerase constructs are then added and polymerization continues for 60 min under conditions adjusted for optimal polymerization. A sample from the polymerization reaction is mixed with gel loading buffer and electrophoresed on a 20% TBE-acrylamide gel (Invitrogen, USA) along with a non-polymerase negative control and a marker of MW to confirm XNTP polymerization.
The Xpandomer intermediate is treated with a divalent cation (see Mag et al. 1991. "Synthesis and selective cleavage of an oligodeoxynucleotide containing a bridged internucleotide 5'-phosphorothioate linkage", Nucleic Acids Research, 19 (7): 1437-1441) to selectively cleave the phosphorothiolate bonds between the anchor junction and deoxyribose, which is confirmed by electrophoresis.
EXAMPLE 32.
SYNTHESIS OF XNTP XPANDOMERS USING POLYMERASE
The synthesis of a class X Xpandomer is performed using a modified substrate construct N<sup>6</sup>- (6-amino) hexyl-deoxyadenosine triphosphate having a mixed backbone consisting of a non-connecting (N-1-aminoalkyl) phosphoramidate and a connecting phosphorothiolate (3 'OPS 5') on the alpha phosphate. An intranucleotide anchor is attached to the N-1-aminoalkyl group and to a C6 amino on N linker<sup>6</sup>- (6-amino) hexyl-deoxyATP. The sequence "TTTTTTTTTTTTTTTTTTTATCTACCGTCCGTCCC" is used as a template. The sequence "5'GGGACGGACGGTAGAT" is used as a primer. A 5 'end HEX (5'hexachloro-fluorescein) is used on the primer as a label. Primer and template hybridization forms a template with duplex primer and single-stranded protruding nucleotide from the free 3'-OH and 5 'ends of twenty bases in length. The substrate and polymerase constructs are then added and polymerization continues for 60 min under conditions adjusted for optimal polymerization. A sample from the polymerization reaction is mixed with gel loading buffer and electrophoresed on a 20% TBE-acrylamide gel (Invitrogen, USA) along with a non-polymerase negative control and a marker of MW to confirm XNTP polymerization.
The Xpandomer intermediate is treated with iodoethanol (see Gish et al ("DNA and RNA sequence determination based on phosphorothioate chemistry", Science, 240 (4858): 1520-1522, 1988) or by cleavage with divalent metal cations as described by Vyle et al ("Sequence- and strand-specific cleavage in oligodeoxyribonucleotides and DNA containing 3'-thiothymidine". Biochemistry 31 (11): 3012-8, 1992) to selectively cleave the phosphorothiolate bonds between the anchor junction and deoxyribose, which is confirmed by electrophoresis.
EXAMPLE 33.
SYNTHESIS OF XNTP XPANDOMERS USING LIGASA
The synthesis of a class X Xpandomer is performed using a modified substrate construct 8 - [(6-amino) hexyl] -amino-deoxyadenosine monophosphate having a mixed backbone consisting of a non-connecting 2-aminoethyl phosphonate and a connecting phosphorothiolate (3 ' OP-S 5 ') in alpha phosphate. An intra-nucleotide anchor is attached to the 2-aminoethyl phosphonate linker and to a C6 amino linker on the 8 - [(6-amino) hexyl] -amino
ES 2 559 313 T3 deoxyAMP. The sequence "TTTTTTTTTTTTTTTTTTTATCTACCGTCCGTCCC" is used as a template. The sequence "5'GGGACGGACGGTAGAT" is used as a primer. A 5 'end HEX (5'hexachloro-fluorescein) is used on the primer as a label. Primer and template hybridization forms a template with duplex primer and single-stranded protruding nucleotide from the free 3'-OH and 5 'ends of twenty bases in length. The 5 substrate and ligase constructs are then added and ligation continues for 5 hours under conditions set for ligation.
A sample of the ligation reaction is mixed with gel loading buffer and electrophoresed on a 20% TBE-acrylamide gel (Invitrogen, USA) along with a non-polymerase negative control and a marker of MW to confirm XNTP ligation.
The intermediate product of Xpandomer is treated with 80% acetic acid for 5 h at room temperature 10 according to the procedure of Mag et al ("Synthesis and selective cleavage of oligodeoxyribonucleotides containing nonchiral internucleotide phosphoramidate linkages", Nucl. Acids Res. 17: 5973- 88, 1989) to selectively cleave the phosphoramidate bonds between the anchor attachment point and deoxyribose, which is confirmed by electrophoresis.
The various embodiments described above can be combined to provide additional embodiments.
Contents51
81 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81
42 members in 12 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 945031P | United States of America | – | |
| 94503107 | United States of America | P | |
| 981916P | United States of America | – | |
| 98191607 | United States of America | P | |
| 305P | United States of America | – | |
| 30507 | United States of America | P | |
| 2008067507 | United States of America | W |
Members42
| Document | Office | Kind | |
|---|---|---|---|
| AU2008265691A1 | Australia | A1 | |
| CA2691364A1 | Canada | A1 | |
| WO2008157696A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2009035777A1 | United States of America | A1 | |
| WO2008157696A3 | World Intellectual Property Organization (WIPO) | A3 | |
| CA2740973A1 | Canada | A1 | |
| WO2009055617A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2171088A2 | European Patent Office (EPO) | A2 | |
| KR20100044790A | Republic of Korea | A | |
| EP2215259A1 | European Patent Office (EPO) | A1 | |
| JP2010530759A | Japan | A | |
| US2010297644A1 | United States of America | A1 | |
| CN101910410A | China | A | |
| HK1143188A | Hong Kong, China | A | |
| HK1143188A1 | Hong Kong, China | A1 | |
| US7939259B2 | United States of America | B2 | |
| CN102083998A | China | A | |
| US2011237787A1 | United States of America | A1 | |
| US2011251079A1 | United States of America | A1 | |
| US8324360B2 | United States of America | B2 | |
| US8349565B2 | United States of America | B2 | |
| US2013172197A1 | United States of America | A1 | |
| US8592182B2 | United States of America | B2 | |
| AU2008265691B2 | Australia | B2 | |
| JP5745842B2 | Japan | B2 | |
| EP2171088B1 | European Patent Office (EPO) | B1 | |
| EP2952587A1 | European Patent Office (EPO) | A1 | |
| DK2171088T3 | Denmark | T3 | |
| ES2559313T3This record | Spain | T3 | |
| US2016208344A1 | United States of America | A1 | |
| CN102083998B | China | B | |
| KR101685208B1 | Republic of Korea | B1 | |
| HK1217736A | Hong Kong, China | A | |
| HK1217736A1 | Hong Kong, China | A1 | |
| IL202821A | Israel | A | |
| IL224295A | Israel | A | |
| US9920386B2 | United States of America | B2 | |
| US2018334729A1 | United States of America | A1 | |
| CA2691364C | Canada | C | |
| US2022064741A1 | United States of America | A1 | |
| EP2952587B1 | European Patent Office (EPO) | B1 | |
| ES2959127T3 | Spain | T3 |
Numbers
- Publication
- 2559313
- Application
- 8771483
Titles2
- Spanish
- Secuenciación de ácidos nucleicos de alto rendimiento por expansión
- English
- High performance nucleic acid sequencing by expansion
Classification
- CPC, 6
- C12Q1/6806
- C12Q1/6897
- C12Q1/6869
- C12Q1/6876
- C12Q1/68
- C07H21/00
- IPC, 3
- C12Q1 68
- C07H21 02
- C07H21 04