Improved strategies for sequencing complex genomes using high throughput sequencing technologies
Abstract
Method for determining a genomic sequence comprising the steps of: (a) providing a first subset of the genome by digesting the genome with at least a first restriction endonuclease to provide restriction fragments; (b) ligation of at least one adapter with the restriction fragments of the first subset to provide a first set of restriction fragments linked to the adapter; (c) selectively amplifying the first set of restriction fragments bound to the adapter using a first primer combination in which at least one first primer contains a section complementary to the adapter and a portion of the restriction endonuclease recognition sequence and which also contains a first sequence selected at the 3 '' end of the primer sequence, the first sequence selected between 0 and 10 selective nucleotides, to provide a first subset of amplified restriction fragments bound to the adapter; (d) repeat step (c) with at least one second and / or one or more additional combinations of the primer in which the primer contains a different second and / or additional selected sequence at its 3 '' end containing the same number of selective nucleotides, to provide second and / or additional subsets of amplified restriction fragments bound to the adapter; (e) fragment each of the first, second and / or additional subsets of amplified restriction fragments linked to the adapter to generate first, second and / or additional sequencing libraries; (f) determining at least a part of the nucleotide sequences of at least a part of the fragments contained in each of the first, second and / or additional libraries; (g) align the sequence of the fragments in each of the first, second and / or additional libraries to generate amplified restriction fragments linked to the adapter obtained representing dispersed fractions of the genome; (h) repeat steps (a) - (g) for at least a second and / or additional restriction endonucleases; (i) align the condoms obtained in steps (g) and (h) for each of the second and / or additional restriction endonucleases to provide a genome sequence.

Term
Term ended
Projected expiry passed 23 June 2026, 0.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
16 claims: 8 independent, 8 dependent
- 1ES 2 344 802 T3 ES 2 344 802 T3 CLAIMS REIVINDICACIONES 1. Procedure to determine a genomic sequence that comprises the steps of:1. Procedimiento para determinar una secuencia genómica que comprende las etapas de: (a) proporcionar un primer subconjunto del genoma mediante la digestión del genoma con por lo menos una primera endonucleasa de restricción para proporcionar fragmentos de restricción;(a) providing a first subset of the genome by digesting the genome with at least a first restriction endonuclease to provide restriction fragments;(b) performing the ligation of at least one adapter with the restriction fragments of the first subset to provide a first set of restriction fragments linked to the adapter;(b) realizar la ligación de por lo menos un adaptador con los fragmentos de restricción del primer subconjunto para proporcionar un primer conjunto de fragmentos de restricción ligados al adaptador;(c) amplificar selectivamente el primer conjunto de fragmentos de restricción ligados al adaptador utilizando una primera combinación de cebador en la que por lo menos un primer cebador contiene una sección complementaria al adaptador y a una parte de la secuencia de reconocimiento de la endonucleasa de restricción y que contiene además una primera secuencia seleccionada en el extremo 3' de la secuencia del cebador, comprendiendo la primera secuencia seleccionada entre 0 y 10 nucleótidos selectivos, para proporcionar un primer subconjunto de fragmentos de restricción amplificados ligados al adaptador;(c) selectively amplifying the first set of adapter-ligated restriction fragments using a first primer combination in which at least one first primer contains a section complementary to the adapter and a portion of the restriction endonuclease recognition sequence, and further containing a first sequence selected at the 3 'end of the primer sequence, the first selected sequence comprising between 0 and 10 selective nucleotides, to provide a first subset of amplified restriction fragments linked to the adapter;(d) repeat step (c) with at least one second and / or one or more additional primer combinations in which the primer contains a different second and / or additional selected sequence at its 3 'end containing the same number of selective nucleotides, to provide second and / or additional subsets of amplified restriction fragments linked to the adapter;(d) repetir la etapa (c) con por lo menos un segundo y/o una o más combinaciones adicionales del cebador en las que el cebador contiene una distinta secuencia seleccionada segunda y/o adicionales en su extremo 3' que contiene el mismo número de nucleótidos selectivos, para proporcionar unos subconjuntos segundo y/o adicionales de fragmentos de restricción amplificados ligados al adaptador;(e) fragmentar cada uno de los subconjuntos primero, segundo y/o adicionales de fragmentos de restricción amplificados ligados al adaptador para generar unas genotecas de secuenciación primera, segunda y/o adicionales;(e) fragmenting each of the first, second, and / or additional subsets of adapter-ligated amplified restriction fragments to generate first, second, and / or additional sequencing libraries;(f) determinar por lo menos una parte de las secuencias de nucleótidos de por lo menos una parte de los fragmentos contenidos en cada una de las genotecas primera, segunda y/o adicionales;(f) determining at least a part of the nucleotide sequences of at least a part of the fragments contained in each of the first, second and / or additional libraries;(g) aligning the sequence of the fragments in each of the first, second and / or additional libraries to generate contigs of the obtained adapter-linked amplified restriction fragments that represent sparse fractions of the genome;(g) alinear la secuencia de los fragmentos en cada una de las genotecas primera, segunda y/o adicionales para generar cóntigos de los fragmentos de restricción amplificados ligados al adaptador obtenidos que representan fracciones dispersas del genoma;(h) repeating steps (a) - (g) for at least a second and / or additional restriction endonucleases;(h) repetir las etapas (a)-(g) para por lo menos unas endonucleasas de restricción segunda y/o adicionales;(i) align the contigs obtained in steps (g) and (h) for each of the second and / or additional restriction endonucleases to provide a genome sequence. (i) alinear los cóntigos obtenidos en las etapas (g) y (h) para cada una de las endonucleasas de restricción segunda y/o adicionales para proporcionar una secuencia del genoma.
- 3Method according to any of claims 1 or 2, wherein at least one of the first, second and / or additional restriction endonucleases is a frequent cutter. 3. Procedimiento según cualquiera de las reivindicaciones 1 ó 2, en el que por lo menos una de las endonucleasas de restricción primera, segunda y/o adicionales es una cortadora frecuente.
- 8Method according to any of claims 1 to 7, in which the first, second and additional selected sequences have the same number of nucleotides but differ in nucleotide sequences from each other from the selective sequence found at the 3 'end of the primer . 8. Procedimiento según cualquiera de las reivindicaciones 1 a 7, en el que las secuencias seleccionadas primera, segunda y adicionales presentan el mismo número de nucleótidos pero difieren en las secuencias de nucleótidos entre sí de la secuencia selectiva que se encuentra en el extremo 3' del cebador.
- 14Method according to any of claims 1 to 8, in which sequencing comprises the steps of:14. Procedimiento según cualquiera de las reivindicaciones 1 a 8, en el que secuenciación comprende las etapas de: (f1) enlazar los adaptadores de la secuenciación con los fragmentos;(f1) linking the sequencing adapters to the fragments;(f2) aparear los fragmentos de la secuenciación enlazados con el adaptador con las perlas, apareándose cada perla con un fragmento simple, (f3) emulsionar las perlas en microrreactores de agua en aceite, comprendiendo cada microrreactor de agua en aceite una perla simple;(f2) pairing the adapter-linked sequencing fragments with the beads, each bead pairing with a single fragment, (f3) emulsifying the beads in water-in-oil microreactors, each water-in-oil microreactor comprising a single bead;(f4) perform emulsion PCR to amplify adapter-linked fragments on the surface of beads (f5) select / enrich beads containing amplified adapter-linked fragments (f6) load beads into wells, each well comprising one simple pearl;and (f7) generating a pyrophosphate signal. (f4) realizar la PCR en emulsión para amplificar los fragmentos enlazados con el adaptador sobre la superficie de perlas (f5) seleccionar/enriquecer las perlas que contienen fragmentos amplificados enlazados con el adaptador (f6) cargar las perlas en pocillos, comprendiendo cada pocillo una perla simple;y (f7) generar una señal de pirofosfato.
- 15Procedimiento según cualquiera de las reivindicaciones anteriores, en el que la construcción de cóntigos se ve ayudada además por la utilización de secuencias de nucleótidos obtenidas a partir de otras fuentes, que comprenden, pero sin limitarse a las mismas, secuencias del extremo BAC, secuencias aleatorias BAC, secuencias EST o secuencias aleatorias del genoma entero. fifteen. Method according to any of the preceding claims, in which contig construction is further aided by the use of nucleotide sequences obtained from other sources, including, but not limited to, BAC end sequences, random sequences BAC, EST sequences or random sequences of the entire genome.
- 16Method according to any of the preceding claims, wherein the method for reducing the complexity of the mixture is based on indexing linkers, chromatin immunoprecipitation (CHIP) or PCR primers directed against conserved motifs. 16. Procedimiento según cualquiera de las reivindicaciones anteriores, en el que el procedimiento para reducir la complejidad de la mezcla se basa en ligadores de indización, la inmunoprecipitación de la cromatina (CHIP) o cebadores de la PCR dirigidos contra motivos conservados.
Independent claims8
186 paragraphs in 11 sections, as filed
ES 2 344 802 T3
DESCRIPTION
Improved strategies for complex genome sequencing using high-throughput sequencing technologies.
Technical field
The present invention relates to the fields of molecular biology and genetics. The present invention relates to improved strategies for determining the sequence of preferably complex (ie large) genomes, based on the use of sequencing techniques.
Background of the invention
Assembling whole genome random sequences into large genomes (100 Mbp or larger) with draft genome sequences is a complicated matter. Many plants and animals also contain a large number of repeating sequences, further complicating the problem. This computing problem is even greater with the appearance of high-capacity sequencing techniques, such as 454 Life Science techniques. These techniques are no longer based on the Sanger dideoxysequencing method, but mainly on sequencing by synthesis (pyrosequencing), which is easier to perform on a solid surface. Synthesis sequencing provides a large number of sequences, albeit relatively short in length (approximately 100 bp [base pairs]) compared to the relatively long length of between 500 and 1000 bp as is common in the dideoxysequencing method of Sanger.
One of the disadvantages of such short fragments is that assembling contig to determine genomic sequence requires great computing power, making current sequencing procedures relatively expensive and time consuming. Consequently, there is a need for complex sequencing procedures that are inexpensive, reliable and fast, ie large genomes, to enhance the technology that has sometimes been referred to as the "$ 1000 genome", ie a procedure that allows the determination of the entire sequence of a complex genome (particularly human) costing no more than $ 1000. This would allow, among other things, the development of personalized medications.
Summary of the invention
The present inventors have now discovered that said problem can be solved with a different strategy and that techniques of high sequencing capacity can be used efficiently in the assembly of the genome.
The present invention comprises the use of a technique that divides the genome into complementary and reproducible parts by restricting the genome with one or more restriction endonucleases to obtain a set of restriction fragments and subsequently provide a subset of restriction fragments by selective amplification. . The subassembly is sequenced and assembled to a contig. By repeating this step for one or more different sets of restriction endonucleases, different contigs are obtained. These different contigs are used to assemble the draft genome sequence. The present invention does not require any knowledge of sequence and can be applied to genomes of any size and complexity. The present invention can be adjusted for any type and size of genome. The present invention provides faster, more reliable and faster access to any genome of interest and therefore allows to accelerate the analysis of the genome.
Definitions
In the following description and examples, a number of terms are used. In order to provide a clear and consistent understanding of the present specification and claims, including the scope to be provided to said terms, the following definitions are provided. Except as defined otherwise herein, all technical and scientific terms have the same meanings commonly ascribed to them by those of ordinary skill in the art to which the present invention pertains.
Nucleic acid: a nucleic acid according to the present invention can comprise any polymer or oligomer of pyrimidine and purine bases, preferably cytosine, thymine and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry ("Principles of Biochemistry ”), at 793-800 (Worth Pub. 1982)). The present invention contemplates any deoxyribonucleotide, ribonucleotide or peptide component of a nucleic acid, and any chemical variant thereof, such as the methylated, hydroxymethylated or glycosylated forms thereof. The polymers or oligomers can have a heterogeneous or homogeneous composition, and can be isolated from natural sources or can be produced artificially or synthetically. Furthermore, nucleic acids can be DNA or RNA, or a mixture thereof, and can exist permanently or transiently in single or double-stranded form, comprising the homoduplex, heteroduplex and hybrid states.
Complexity reduction: The term complexity reduction is used to indicate a procedure in which the complexity of a nucleic acid sample, such as genomic DNA, is reduced by generating a subset of the sample. This subset can be representative of the entire sample (that is, complex) and is
ES 2 344 802 T3 preferably a reproducible subset. Reproducible means in the present context that when the complexity of the same sample is reduced using the same procedure, the same subset is obtained, or at least one of comparable. The method used in reducing complexity can be any method of reducing complexity known in the art. Examples of complexity reduction procedures include, for example, AFLP® (Keygene NV, The Netherlands; see, for example, EP 0 534 858), procedures described by Dong (see, for example, documents WO 03/012118, WO 00/24939, EP 1 124 990) or Barany (eg US 6 534 293), indexed linkage (Unrau et al., See below), etc. The complexity reduction methods used in the present invention have in common that they are reproducible. Reproducible in the sense that when the complexity of the same sample is reduced in the same way, the same subset of the sample is obtained, contrary to what happens with a more random reduction in complexity such as microdissection or the use of MRNA (cDNA) that represents a part of the transcribed genome in a selected tissue and that its reproducibility depends on the tissue, the isolation period, etc.
Labeling: The term labeling refers to the addition of a label to a nucleic acid sample in order to be able to distinguish it from a second or another nucleic acid sample. Labeling can be accomplished, for example, by adding a sequence identifier during complexity reduction or by any other means known in the art. Such a sequence identifier can, for example, be a unique base sequence with a variable but defined length that is used to identify a specific sample of nucleic acid. Common examples thereof are, for example, ZIP sequences. By using such a tag, the origin of one can be determined with additional processing. In the case of combining processor products from different nucleic acid samples, the different nucleic acid mixtures would have to be identified using different labels.
Tagged library: The term tagged library refers to a tagged nucleic acid library.
Sequencing: The term sequencing refers to the determination of the order of nucleotides (base sequences) in a nucleic acid sample, for example DNA or RNA.
Align and Alignment: The terms "align" and "alignment" refer to the comparison of two or more nucleotide sequences based on the presence of short or long stretches of identical or similar nucleotides. Various nucleotide sequence alignment procedures are known in the art, as will be described in greater detail below. Assembly is sometimes used synonymously.
High-capacity identification techniques: High-capacity identification techniques, often abbreviated as HTS, consist of a procedure of scientific experimentation especially suitable for the fields of biology and chemistry. Using a combination of current robotics and other specialized laboratory hardware, it enables the researcher to effectively identify large numbers of samples simultaneously.
Restriction endonuclease: a restriction endonuclease or restriction enzyme is an enzyme that recognizes specific nucleotide sequences (selected site) in a double-stranded DNA molecule and will separate both strands of the DNA molecule at each selected site.
Restriction fragments: DNA molecules produced by digestion with a restriction endonuclease are referred to as restriction fragments. Any given genome (or nucleic acid, regardless of its origin) will be digested by a particular restriction endonuclease into a described set of restriction fragments. The DNA fragments that are obtained from restriction endonuclease cleavage can continue to be used in various techniques and, for example, can be detected by gel electrophoresis.
Gel electrophoresis: In order to detect restriction fragments, an analytical procedure may be required to fractionate double-stranded DNA molecules based on size. The technique most commonly used to achieve this fractionation is gel (capillary) electrophoresis. The speed at which said DNA fragments move in said gels depends on their molecular weight; thus, the travel distance decreases as the length of the fragment increases. Gel electrophoresis fractionated DNA fragments can be directly visualized by a staining procedure, for example silver staining or a stain using ethidium bromide, if the number of fragments found in the band is small enough. Alternatively, further processing of the DNA fragments can incorporate detectable labels on the fragments, such as fluorophores or radioactive labels.
Ligation: refers to the enzymatic reaction catalyzed by an enzyme ligase in which two molecules of double-stranded DNA are joined together with a covalent bond as a ligation. In general, both DNA strands are joined together with a covalent bond, but it is also possible to avoid the ligation of one of the two strands by chemical or enzymatic modification of one of the ends of the strands. In this case, the covalent bond will occur in only one of the two DNA strands.
Synthetic oligonucleotide: refers to single-stranded DNA molecules that preferably have between about 10 and about 50 bases, which can be chemically synthesized as oligo
ES 2 344 802 T3 synthetic nucleotides. In general, such synthetic DNA molecules are designed to have a unique or intended nucleotide sequence, although it is possible to synthesize families of molecules that have related sequences and that have different nucleotide compositions at specific positions of the nucleotide sequence. . The term "synthetic oligonucleotide" will be used to refer to DNA molecules that have a designed or intended nucleotide sequence.
Adapters: short double-stranded DNA molecules with a limited number of base pairs, for example between about 10 and about 30 base pairs in length, that are designed in such a way that they can be ligated to the ends of restriction fragments. The adapters are generally composed of two synthetic oligonucleotides that have nucleotide sequences that are partially complementary to each other. When two synthetic oligonucleotides are mixed in solution under appropriate conditions, they will pair with each other forming a double-stranded structure. After mating, one end of the adapter molecule is designed such that it is compatible with the end of a restriction fragment and can be ligated to it; the other end of the adapter can be designed in such a way that it cannot be ligated, but this need is not the case (dual ligature adapters).
Adapter-linked restriction fragments: Restriction fragments that have been capped with adapters as a result of ligation.
Primers: In general, the term primers refers to a strand of DNA that can initiate DNA synthesis. DNA polymerase cannot synthesize DNA de novo without primers: it can only extend an existing strand of DNA in a reaction in which the complementary strand is used as a template to direct the order of nucleotides to be assembled. We will refer to the synthetic oligonucleotide molecules that are used in the polymerase chain reaction (PCR) as primers.
DNA Amplification: The term DNA amplification will be commonly used to denote the in vitro synthesis of double-stranded DNA molecules using PCR. It is to be noted that other amplification procedures exist and can be used in the present invention without departing from the essence thereof.
Detailed description of the invention
The present invention provides a method for determining a genomic sequence comprising the steps of:
(a) providing a first subset of the genome by digesting the genome with at least a first restriction endonuclease to provide restriction fragments;
(b) performing the ligation of at least one adapter with the restriction fragments of the first subset to provide a first set of restriction fragments linked to the adapter;
(c) selectively amplifying the first set of adapter-ligated restriction fragments using a first primer combination in which at least one first primer contains a section complementary to the adapter and a portion of the restriction endonuclease recognition sequence, and further containing a first sequence selected at the 3 'end of the primer sequence, the first selected sequence comprising between 1 and 10 selective nucleotides, to provide a first subset of amplified restriction fragments linked to the adapter;
(d) repeat step (c) with at least one second and / or one or more additional primer combinations in which the primer contains a different second and / or additional selected sequence at its 3 'end containing the same number of selective nucleotides, to provide second and / or additional subsets of amplified restriction fragments linked to the adapter;
(e) fragmenting each of the first, second and / or additional subsets of amplified restriction fragments linked to the adapter, optionally followed by selection of the size of the fragments in the optimal size range, to generate first sequencing libraries , second and / or additional, and then optionally mixing the libraries;
(f) determining (at least a part of) the nucleotide sequences of (at least a part of) the fragments contained in each of the first, second and / or additional sequencing libraries;
(g) aligning the sequence of the fragments in each of the first, second, and / or additional libraries to generate contigs of the adapter-linked amplified restriction fragments obtained from the genome subset (s);
(h) repeating steps (a) - (g) for at least a second and / or additional restriction endonucleases;
(i) align the contigs obtained in steps (g) and (h) for each of the second and / or additional restriction endonucleases to provide a genome sequence.
ES 2 344 802 T3
In step (a) of the procedure, the genome of interest is subjected to one or more restriction endonucleases. In certain embodiments, at least two restriction endonucleases are used. In certain embodiments, particularly with large genomes, three or more restriction endonucleases can be used. Digestion of the genome provides a first subset of the genome. Restriction endonucleases can be frequent cutters (i.e. usually 4 and 5 cutters, i.e. restriction endonucleases having a recognition sequence of 4 or 5 nucleotides, respectively) or they can be rare cutters, (i.e. usually cutters 6 and higher (7, 8, ... etc., ie, restriction endonucleases having a recognition sequence of 6 or more nucleotides, respectively)), or combinations thereof. In certain embodiments, a combination of a rare and a frequent cutter is used. In certain embodiments, two rare cutters can be used. Restriction endonucleases can be of any type, including types IIs and IIsa that cut DNA outside of its recognition sequence, both on one and both sides of the recognition sequence.
In step (b) of the procedure, at least one adapter is ligated to the restriction fragments obtained in step (a). Preferably, the adapters are such that the restriction site is not reestablished by ligation of the adapter. It is also possible, for example, in the case of two or more restriction endonucleases to use two or more different adapters. This stage of ligation produces restriction fragments linked to the adapter. The adapters, depending on the restriction endonuclease, may have a blunt end or may contain a protruding end.
In certain embodiments, the adapter can be a set of adapters known as indexing linkers (Unrau et al., 1994, Gene, 145: 163-169).
In step (c), the first set of adapter-ligated restriction fragments is amplified using a first primer pool. The primer combination comprises at least a first primer containing a section complementary to (at least a part of) the adapter and a part of the restriction endonuclease recognition sequence used in genome restriction. Usually, the part of the recognition sequence is that part that remains after restriction of the sequence with the restriction endonuclease. At its 3 'end, the primer contains a first selected sequence. The first selected sequence comprises a previously selected set of 1 to 10 nucleotides, preferably between 1 and 8 selected nucleotides, preferably between 1 and 5, more preferably between 1 and 3. Said primer may have the following illustrative structure (for 2 selective nucleotides (AC)) "5 '- specific region of the adapter - specific region of the restriction sequence - AC-3'". This exemplary first primer contains 2 AC-selective nucleotides that will only amplify fragments linked to the adapter that comprise the complementary sequence TG as the first two nucleotides obtained from the sequence of the restriction fragment. This provides the first subset of amplified restriction fragments linked to the adapter.
The first primer combination may also comprise two selective primers, each carrying a selected sequence at its 3 'end. The primers can be tagged to allow grouping strategies.
Amplification is preferably done using long PCR. In certain embodiments, the use of PCR is preferred.
In step (d), the selective amplification is repeated with the second and additional primer combinations. At least one of the primers in each of the additional primer combinations contains a different selected sequence at its 3 'end. The selection of the selected sequences is carried out in such a way that, given the number of selected nucleotides, all possible permutations of the selective nucleotides are used. In the example above this means AT, AG, AA, CA, CT, CG, CA, etc. In practice this means that all adapter-linked restriction fragments in the genome subset (ie, in the set of restriction fragments obtained using one or more restriction endonucleases) have been amplified.
In a preferred embodiment of the present invention, the reduction of the complexity of the genome by selective amplification is carried out using the AFLP technique.<sup>®</sup> (Keygene NV, The Netherlands; see for example EP 0 534 858 and Vos et al. (1995) AFLP: a new technique for DNA fingerprinting, Nucleic Acids Research, vol. 23, n. 21, 4407-4414).
AFLP consists of a procedure for the amplification of a selective restriction fragment. AFLP does not require any prior sequence information and can be performed on any starting DNA. In general, the AFLP comprises the stages of:
(a) digesting a nucleic acid, in particular DNA, with one or more specific restriction endonucleases, to fragment the DNA into the corresponding series of restriction fragments;
(b) ligating the restriction fragments thus obtained with an adapter of a double-stranded synthetic oligonucleotide, one end thereof being compatible with one or both ends of the restriction fragments, to thereby produce restriction fragments of the initial DNA linked to the adapter, preferably labeled;
ES 2 344 802 T3 (c) contacting the adapter-linked, preferably tagged, restriction fragments under hybridization conditions with one or more oligonucleotide primers containing selective nucleotides at their 3 'end;
(d) amplifying the adapter-ligated restriction fragment, preferably labeled hybridized with the primers by PCR or a similar technique in such a way as to cause additional elongation of the primers hybridized along the restriction fragments of the initial DNA with the that the primers hybridize; and (e) detecting, identifying or recovering the amplified or elongated DNA fragment thus obtained.
AFLP thus provides a reproducible subset of adapter-linked fragments. A useful variant of the AFLP technique does not use selective nucleotides (ie + 0 / + 0 primers) and is sometimes referred to as linker PCR. This also provides a very adequate reduction in complexity, particularly for smaller genomes.
For a further description of AFLP, its advantages, its embodiments, as well as the techniques, enzymes, adapters, primers and additional compounds and tools used therein, reference is made to US 6,045,994, EP-BO 534 858, EP 976835 and EP 974672, WO 01/88189 and Vos et al. Nucleic Acids Research, 1995, 23,4407-4414.
Thus, in a preferred embodiment of the method of the present invention, the complexity of the genome is reduced by the following steps:
(a) digesting the nucleic acid sample with at least one restriction endonuclease to fragment the same into restriction fragments;
(b) ligating the obtained restriction fragments with at least one adapter of a double-stranded synthetic oligonucleotide having an end compatible with one or both ends of the restriction fragments to produce restriction fragments linked to the adapter;
(c) contacting said adapter-linked restriction fragments with one or more oligonucleotide primers under hybridization conditions; and (d) amplifying said adapted restriction fragments by elongating one or more oligonucleotide primers, (e) at least one of the one or more oligonucleotide primers comprising a nucleotide sequence presenting a nucleotide sequence as the terminal part. of the chains at the ends of said adapted restriction fragments, comprising the nucleotides involved in the formation of the selected sequence for said restriction endonuclease and comprising at least a part of the nucleotides present in the adapters, in which, optionally, at least one of said primers comprises at its 3 'end a selected sequence comprising at least one nucleotide arranged immediately adjacent to the nucleotides involved in the formation of the selected sequence for said restriction endonuclease.
AFLP constitutes a highly reproducible procedure for reducing complexity and is therefore particularly suitable as a method according to the present invention.
To date, in the sequencing art, the use of such selective amplification in the determination of the sequence of entire genomes and, in particular, complex genomes has not been described or proposed. The AFLP technique is known in the art as a fingerprinting technique and has not yet been identified as a solution to aid in the sequencing of complex genomes. In particular, the use of a set of primer combinations that cover all or most of the nucleotide permutations for a given number of selective nucleotides (for example, 16 primer combinations in the case of two selective nucleotides) provides a reliable method. and fast to provide complementary and reproducible subsets of a genome that can be sequenced. In certain embodiments, the primers used to reduce complexity contain one or more thioate linkages to increase their selectivity and / or performance.
In certain alternative embodiments, the reduction in complexity comprises the CHIP method. Other methods suitable for complexity reduction are chromatin immunoprecipitation (ChiP). This means that DNA is isolated from the nucleus, while proteins such as transcription factors bind to DNA. With the Chip technique, an antibody against the protein is used first, resulting in an Ab-protein-DNA complex. By purifying said complex and precipitating it, the DNA to which said protein binds is selected. Subsequently, the DNA can be used in the construction and sequencing of the library. That is, this is a method to perform a complexity reduction in a non-random way directed to specific functional areas.
ES 2 344 802 T3 science; in the present example, specific transcription factors. Alternative embodiments may utilize the design of the PCR primers directed against conserved motifs such as SSRs regions, NBS (nucleotide binding regions), promoter / enhancer sequences, telomere consensus sequences, MADS gene sequences, families of genes from the adenosine triphosphatase gene family and other gene families.
In step (e), the first second and additional sequencing libraries are generated for each subset of amplified restriction fragments linked to the adapter. Libraries are typically generated by fragmentation of adapter-ligated amplified restriction fragments. Fragmentation can be performed by physical techniques, ie, cutting, sonication, or other fragmentation procedures. In step (f), at least a part, but preferably all, of the nucleotide sequence of at least a part of, but preferably all of the fragments contained in the libraries is determined.
Sequencing can in principle be performed using any means known in the art, such as the dideoxyribonucleotide chain termination procedure. It is preferred, however, that the sequencing is performed using high-throughput sequencing procedures, such as the procedures described in WO 03/004690, WO 03/054142, WO 2004/069849, WO 2004/070005, WO 2004 / 070007 and WO 2005/003375 (all in the name of 454 Life Sciences), by Seo et al. (2004) Proc. Nati. Acad. Sci. USA 101: 5488-93, and the techniques of Helios, Solexa, US Genomics. It is preferred that sequencing is performed using the apparatus and / or procedure described in WO 03/004690, WO 03/054142, WO 2004/069849, WO 2004/070005, WO 2004/070007 and WO 2005/003375 (all them on behalf of 454 Life Sciences). The described technique allows sequencing of 40 million bases in a simple series and is 100 times faster and cheaper than alternative techniques. The sequencing technique consists of approximately 5 steps: 1) DNA fragmentation and ligation of a specific adapter to create a single-stranded DNA (ssDNA) library; 2) pairing ssDNA with beads, emulsifying the beads in water-in-oil microreactors and performing emulsion PCR to amplify individual ssDNA molecules on the beads; 3) selection / enrichment of the beads containing amplified ssDNA molecules on their surface 4) sedimentation of the DNA carrying them on a PicoTiterPlate®; and 5) simultaneous sequencing in 100,000 wells by generating an optical pyrophosphate signal. The procedure will be explained in more detail later.
In a preferred embodiment, sequencing comprises the steps of:
(a) pairing the matched fragments to the beads, each bead mating with a single matched fragment;
(b) emulsifying the beads in water-in-oil microreactors, each water-in-oil microreactor comprising a single bead;
(c) loading the beads into wells, each well comprising a single bead; and generating a pyrophosphate signal.
In the first step (a), the sequencing adapters are linked to fragments in the pool of the pool. Said sequencing adapter comprises at least a "key" region for bead pairing, a sequencing primer region and a PCR primer region. In this way the adapted fragments are obtained.
In a first step, the matched fragments are paired with the beads, each bead pairing with a single matched fragment. To the pool of matched fragments, excess beads are added to ensure the matching of a single matched fragment per bead for most beads (Poisson distribution).
In a subsequent step, the beads are emulsified in water-in-oil microreactors, each water-in-oil microreactor comprising a single bead. The PCR reagents are present in the water-in-oil microreactors allowing a PCR reaction to occur in the microreactors. Subsequently, the microreactors are disrupted and the beads comprising DNA (DNA positive beads) are enriched.
In a later step, the beads are loaded into wells, each well comprising a single bead. The wells are preferably part of a PicoTiter<sup>TM</sup>Plate that allows the simultaneous sequencing of a large number of fragments.
Following the addition of enzyme-bearing beads, the sequence of the fragments is determined using pyrosequencing. In successive stages, the PicoTiter<sup>TM</sup>Plate and beads as well as their enzymes are subjected to different deoxyribonucleotides in the presence of conventional sequencing reagents and, by incorporating a deoxyribonucleotide, an optical signal is generated and recorded. Incorporation of the correct nucleotide will generate a pyrosequencing signal that can be detected.
Pyrosequencing itself is known in the art and is described, among others, at www.biotagebio.com; www.pyrosecuencing.com/ technical section. The technique is further applied in, for example, WO 03/004690, WO 03/054142, WO 2004/069849, WO 2004/070005, WO 2004/070007 and WO 2005/003375 (all in the name of 454 Life Sciences ).
ES 2 344 802 T3
In step (g) of the method of the present invention, the determined sequences of the fragments of the first, second and / or additional libraries are aligned. The alignment provides fragment contigs in subsets of the amplified restriction fragments linked to the adapter. In this way, for each amplified restriction fragment linked to the adapter, a contig is generated from the sequenced fragments, that is, the contig of an amplified restriction fragment linked to the adapter, is constructed from the alignment of the sequence of the various fragments obtained from the fragmentation of step (e). By constructing contigs from the sequences that represent the scattered restriction fragments of a small part of the genome, the problems of contig construction that are a consequence of the abundant repeated sequences are considerably reduced, obtaining draft sequence of the genome of a higher quality. A containing fewer errors due to false joins of repeated sequences. In addition, the assembly process will be less complex from a computer point of view and, therefore, faster to carry out. By aligning the sequences of the various libraries, contigs of each restriction fragment from the set of restriction fragments can be constructed for each primer combination. This results in a set of contigs, each corresponding to a particular restriction fragment. As a result, each restriction fragment obtained from restricting the genome with at least one restriction endonuclease now has a certain sequence (contig). The process of the present invention is illustrated in Figures 1 and 2.
Sequence alignment procedures for comparative purposes are well known in the art. Various alignment algorithms and programs are described in: Smith & Waterman (1981) Adv. Appl. Math. 2: 482; Needleman & Wunsch (1970) J. MoI. Biol. 48: 443; Pearson & Lipman (1988) Proc. Natl. Acad. Sci. USA 85: 2444; Higgins & Sharp (1988) Gene 73: 237-244; Higgins & Sharp (1989) CABIOS 5: 151-153; Corpet et al. (1988) Nucl. Acids Res. 16: 10881-90; Huang et al. (1992) Computer Appl. in the Biosci. 8: 155-65; and Pearson et al. (1994) Meth. MoI. Biol. 24: 307-31, which are incorporated herein by reference. Altschul et al. (1994) Nature Genet. 6: 119-29 (incorporated herein by reference) presents a detailed discussion of sequence alignment procedures and analogy calculations.
NCBI's Basic Local Alignment Search (BLAST) program (Altschul et al., 1990) is available from a variety of sources, including the National Center for Biological Information (NCBI, Bethesda, Md.) And on the Internet, for use in conjunction with blastp, blastn, blastx, tblastn, and tblastx sequence analysis programs. Available at <http://www.ncbi.nlm.nih.gov/BLAST/>. A description of how to determine sequence identity using this program is available at <http://www.ncbi.nlm.nih.gov/BLAST/blast_help.html>. A further application can be found in microsatellite extraction (see Varshney et al. (2005) Trends in Biotechn. 23 (1): 48-55.
Usually, the alignment is performed on the sequence data that has been extracted for the adapters / primers and / or identifiers but with reconstructed restriction enzyme recognition sequences, that is, using only the sequence data from the fragments. originating from the nucleic acid sample. Typically, the sequence data obtained is used to identify the origin of the fragment (that is, from which sample), the sequences obtained from the adapter and / or identifier are removed from the data, and alignment is performed on this extracted set. .
In step (h), the entire procedure is repeated at least once with one or more different restriction endonucleases, that is, a restriction endonuclease preferably containing a different recognition site than the first endonuclease, to provide a set of second or even additional restriction fragments that are subsequently bound to the adapter, are selectively amplified using combinations of the primer with a selected sequence that is independently selected, that is, it has no relationship whatsoever with the selected sequence, (both in the number and in the type of nucleotides) with which they have been used with the same objective with the first restriction endonuclease. For example, the first subset can be obtained by restriction with MseI / PstI and selectively amplified using a primer selective for the MseI residues of the recognition site and carrying 2 selective nucleotides at its 3 'end. The second subset can be obtained by digestion of EcoRI / HindIII and selective amplification with a primer selective for the EcoRI residues of the selective 1 nucleotide recognition site.
Therefore, a second (and / or additional) contig set for all restriction fragments can be obtained in this way, in a similar way as described hereinabove. This is necessary, since for a given restriction endonuclease, the fractions of the genome that are being sequenced are complementary, they do not overlap. Contigs obtained with different enzyme combinations overlap and, therefore, allow the generation of a contig from them and, thus, allow the formation of a genomic sequence (draft).
In step (i) of the procedure, the contigs obtained from the previous steps of the procedure for each fragment are aligned to form a genome sequence.
In certain embodiments, contig construction of both restriction fragments and genomic sequence can be contributed to by using genome nucleotide sequences derived from other sources, including but not limited to the same, BAC end sequences, BAC random sequences, EST sequences or whole genome random sequences.
ES 2 344 802 T3
The method of the present invention is independent of the DNA source, that is, it can be applied to all organisms since it does not require prior information on the sequence. Through the appropriate selection of selective enzymes, adapters, primers and the number of nucleotides, the scalable technique is presented that can be applied to genomes of all sizes and complexities. Furthermore, the fractions of the genome that are obtained with the different combinations of the primer and / or with the selective primers that differ from each other in the specific selective nucleotide sequences of the 3 'end, are complementary. This means that when, for a given number of selective nucleotides, all permutations are used (1 selective nucleotide equals 4 variants (A, C, T, G), 2 selective nucleotides 16, 3 selective nucleotides 64 variants, etc.), the combined restriction fragments constitute the restricted genome.
Brief description of the figures
Figure 1: Starting from genomic DNA, a restriction endonuclease pool digestion (enzyme pool 1, EC1) is performed, resulting in a set of restriction fragments. Restriction fragments are linked with adapters after which the adapter-linked restriction fragments are amplified with a first selective primer combination (PC1) to obtain n fragments. Each fragment is cleaved for high-throughput sequencing and subjected to sequencing and alignment to generate restriction fragment contigs. In this way, the n sequence of all or most adapter-linked restriction fragments is determined for a primer pool.
Figure 2: For each possible primer combination of a given enzyme combination (EC1), the steps of fragmentation of the adapter-linked amplified restriction fragments, sequencing, alignment and contig construction are repeated. This means that when, for example, selective amplification is performed with primers each carrying a selective nucleotide at their 3 'end (i.e. + 1 / + 1 primers), 16 primer combinations (PC1 ... PCm) cover all permutations and with 16 primer combinations all adapter-linked restriction fragments have been amplified and subsequently sequenced. From contigs generated with EC1, that is, from EC1 / PC1 ... EC1 / PCm, an assembly will cover a large part of the genome, but will need to be fixed in order to provide a genomic sequence. For this purpose, a second enzyme combination is used (and, if necessary, a third and a fourth, etc.). The steps of Figures 1 and 2 are repeated with Enzyme Combination 2 (EC2), ie, restriction, ligation with the adapter, etc. Selective amplification is performed with a set of selective primers that can usually be different (sequence and selectivity) from the primers used with EC1. The adapter-linked restriction fragments are amplified again with all possible permutations of a set of selective primers, obtaining distinct and complementary subsets. Fragmentation of each subset of selectively amplified adapter-linked restriction fragments, and subsequent high-throughput sequencing, contig construction, etc. leads to a second assembly, which again covers a large part of the genome. From these two assemblies (and the optional third, fourth, etc. enzyme combination), which overlap each other in large areas, the draft sequence of the genome under investigation is generated.
Figure 3. Computer predicted 662 base pair sequence of contig 606, containing overlapping restriction fragments EcoRI / HindIII + C / + CT and BamHI / XbaI + C + G.
Figure 4. Computer-predicted contig 606 observed sequence counts based on AFLP sequencing of EcoRI / HindIII + C / + CT (r1_9_35974-36087) and BamHI / XbaI + C / + G (r2_9_36138-36200 fragment libraries) ). It is to be noted that the 42 base pair overlap between (r1_9_35974-36087) and (r2_9_36138-36200) is completely covered by the sequences obtained from both fragment libraries.
The present invention can be illustrated by the following examples which are not intended to limit the present invention in any way, but are presented for illustrative purposes only.
Whole Genome Sequencing Using Long PCR
Stage 1
DNA is restricted using two 6 cutters A and B (eg EcoRI and HindIII). This generates three types of restriction fragments: AA (25%), BB (25%) and AB (50%) with an average length of approximately 3-4 kb, depending on the GC content of the genome of interest, and the selection of restriction endonucleases. After ligation of the adapters, long PCR is performed with + X / + Y primers (ie, one of the primers contains X selective nucleotides and the other Y), up to a sequence of 1 Mb per primer combination. In the case that X = 2 and Y = 3, this is repeated for all 1024 primer combinations. In the case that X = 1 and Y = 2, this is repeated for all 64 primer combinations.
A: 2700 Mbp maize genome size: AB type fragments = 1350 Mbp. By selective amplification + 2 / + 3 (1024 primer combinations) the amplification product of each of the primer combinations contains, on average, a sequence of 1350/1024 = 1.32 Mbp. With an average length of approximately 3000 base pairs per AB fragment, this produces 132,000/3000 = 440 AB fragments.
ES 2 344 802 T3
B: 130 Mbp from Arabidopsis: 65 Mbp type AB fragments. With X = 1 and Y = 2, each primer combination (PC) contains a sequence of approximately 1 Mbp. With an average length of 3000 base pairs per fragment, that is, 1000000/3000 = 330 AB fragments.
Stage 2
Library construction by cutting each set of adapter-ligated amplified restriction fragments and sequencing using emulsion PCR in conjunction with 454 Technologies pyrosequencing as described elsewhere herein, for each primer combination (1024 or 64-fold ). Sequencing using this technology provides 40 Mbp sequence data per library, which means that each library has a 40-fold repeat level in the sequence. By varying the number of nucleotides, a different number of AB fragments are amplified and a different level of repeats is obtained. This can be determined in its implementation.
Stage 3
Assembly of sequences by sequence library (per PC)
Assembly is performed to generate contigs of all AB fragments that have been PC-amplified. With this, approximately between 300 and 500 contigs are produced per PC, the mean length of which will correspond to the mean length of the AB fragment. Sequencing of all PCs results in contig numbers ranging from several tens of thousands to several hundred thousands (Arabidopsis 21000, maize 450000).
Stage 4
Repeat steps 1 to 3 for at least one other enzyme combination (EC), eg AC. This is essential since all EC AB PCs provide only complementary contigs and no overlapping contigs, thus covering only 50% of the genome. Genome coverage can also be increased by processing all AA and BB fragments. By using additional ECs, overlap between the AB and AC contigs is achieved and genome coverage is increased.
Stage 5
Assembly of all contigs of primer combinations AB (optionally also AA, BB) and A-C to a genomic sequence (draft)
One of the advantages of such a procedure lies in the fact that one of the problems of genome assembly and the possibility of wrong contig formation due to the fact that the presence of multiple repeated sequences is minimized by the formation of small sparse contig (i.e. , not adjacent) of 1-10 kb in a 1-5 Mbp genome fraction instead of the entire genome. Contig with much longer lengths can be tagged in the early stages as a false bond product and discarded. An additional advantage is that the assembly is less complex from a computer point of view due to the initial assembly (step 3), fewer sequences are involved than when the entire genomic sequence is to be assembled in one step. An additional advantage is that the selective amplification process allows the entire procedure to be graduated to a genome of any size and is universally applicable.
Whole Genome Sequencing Using One Rare and One Frequent Cutter
Stage 1
As in the previous case, with a cutter 6 (EcoRI) and a cutter 4 (MseI). The average length of the fragment is approximately 250 base pairs. The AB fragments represent approximately 8-15% of the genome. Compared to restriction enzyme digestion using two cutter restriction enzymes 6, an average of about 1 less selective nucleotide is needed to obtain a sequence complexity level per PC of about 1-5 Mbp.
A: 2700 Mbp maize genome size: AB-type fragments = 270 Mbp, (10%) by selective amplification + 2 / + 2 (256 primer combinations) contains the amplification product of each of the combinations of the primer, on average, 270/256 sequence = 1.05 Mbp. With an average length per AB fragment of approximately 250 base pairs, 1050000/250 = 4200 AB fragments / contigs are obtained.
B: 130 Mbp of Arabidopsis: 13 Mbp AB-type fragments (10%). With X = 1 and Y = 1, each primer combination (PC) contains approximately a 1 Mbp sequence. With an average length of 250 base pairs per fragment, that is, 1000000/3000 = 4000 AB fragments.
ES 2 344 802 T3
Stage 2
As in the previous case. To avoid systematic error of fragments that are too short, size selection can be used to eliminate fragments smaller than 100-150 base pairs.
Stage 3
Assembly of sequences by sequence library (per PC)
Assembly is performed to generate contigs of all AB fragments that have been PC-amplified. This produces several thousand contigs per PC whose average length will correspond to the length of the AB fragment (250 base pairs). Sequencing of all PCs results in contig numbers ranging from several tens of thousands to approximately one million (Arabidopsis 64000, maize 1000000).
Stage 4
Repetition of steps 1-3 with different CEs (AC, BC, CC, CD, AD, etc.). This is necessary since the AB enzyme combination PCs do not cover more than 8-15% of the genome and, as in the previous case, the contigs of the PCs do not overlap.
Stage 5
As in the previous case.
Whole Genome Sequencing Using a Restriction Endonuclease
Stage 1
Digestion of DNA with a restriction endonuclease A (EcoRI, for example). Restriction fragments of approximately 3-4 kb, depending on the GC content and the choice of enzyme. The mixture is bound to the adapter and the long PCR is carried out (see above) with selective primers that reduce the amount of sequence per PC to approximately 1 Mb. In the case of X = 2 and Y = 3 it is repeated for all 1024 PCs. In the case (X, Y) = (+ 1 / + 2) it is repeated for all 64 PCs.
A: 2700 Mbp maize genome size: AA-type fragments = 2700 Mbp, by selective amplification + 2 / + 3 (1024 primer combinations) contains the amplification product of each of the primer combinations, in average, 2700/1024 sequence = 2.64 Mb. With an average length per AA fragment of approximately 3000 base pairs, 2640000/300 = 880 AA fragments / contigs are obtained.
B: 130 Mbp Arabidopsis: 130 Mb AA-like fragments. With X = 1 and Y = 2, each primer combination (PC) contains approximately a 2 Mb sequence. With an average length of 3000 base pairs per fragment i.e. 2,000,000/3,000 = 660 AA fragments.
Stage 2
Library construction by cutting each set of adapter-ligated amplified restriction fragments and sequencing using emulsion PCR in conjunction with 454 Technologies pyrosequencing as described elsewhere herein, for each primer combination (1024 or 64-fold ). Sequencing using this technology provides 40 Mbp sequence data per library, which means that each library has a 20-fold repeat level in the sequence.
Stage 3
Assembly of sequences by sequence library (per PC)
Assembly is performed to generate contigs of all AA fragments that have been PC amplified. This theoretically produces between 600 and 900 contigs per PC whose mean length will correspond to the length of the AA fragment (3000 base pairs). Sequencing of all PCs results in contig numbers ranging from several tens of thousands to approximately a few hundred thousand (Arabidopsis 42000, maize 900000).
ES 2 344 802 T3
Stage 4
Repeat steps 1-3 with at least one other CS (BB). This is necessary since the AB enzyme combination PCs do not cover more than 8-15% of the genome and, as in the previous case, the contigs of the PCs do not overlap.
Stage 5
As in the previous case.
Example 1
This example describes the ability to use high-throughput sequencing of AFLP fragments obtained from 2 combinations of restriction enzymes to determine the genomic sequence of a complex plant genome.
In the present example, the following steps were carried out:
A) Computer prediction of restriction fragments by AFLP of the Arabidopsis genomic sequence (Genbank), using the RECOMB program, described in WO 0044937 (Keygene N. V)
The entire genomic sequence of the Arabidopsis genome (Colombia ecotype) was downloaded from Genbank. Computer prediction of AFLP + 1 / + 1 fragments was performed for the restriction enzyme combination BamHI / XbaI using the selective nucleotides + C and + G, respectively, using RECOMB. Similarly, prediction of AFLP + 1 / + 2 fragments was performed for the combination of restriction enzymes EcoRI / HindIII using the selective nucleotides + C and + CT. Collection of AFLP fragments obtained from two computer digests resulted in various sequences of AFLP fragments (out of approximately 14) that overlapped between the BamHI / XbaI and EcoRI / HindIII enzyme combinations. One of the overlapping restriction fragments forms a contig called contig "606", which has a total length of 662 base pairs. The sequence of said contig is represented in figure 3.
The predicted EcoRI / HindIII AFLP + C / + CT fragment in said contig is 218 base pairs in length, which is 32.9% of the total contig length of 606 base pairs. The BamHI / XbaI AFLP + C / + G fragment is 486 base pairs long, which is equivalent to 73.4% of the total contig length. Both fragments overlap by 42 base pairs, as indicated in Figure 3.
B) Preparation and amplification of the AFLP template
The templates of the genomic DNA of the Arabidopsis Colombia ecotype and AFLP templates were prepared for the combinations of restriction enzymes EcoRI / HindIII and BamHI / XbaI based on the protocols described by Zabeau & Vos, 1993: Selective restriction fragment amplification; a general method for DNA fingerprinting (EP 0534858-A1, B1; US patent 6045994) and Vos et al. (Vos, P., Hogers, R., Bleeker, M., Reijans, M., van de Lee, T., Homes, M., Frijters, A., Pot, J., Peleman, J., Kuiper, M. et al. (1995) AFLP: a new technique for DNA fingerprinting. Nucl. Acids Res., 21, 4407-4414).
The following sequences (5'-3 ') of the adapter were used:
BamHI:
91M35: CTCGTAGACTGCGTACC [SEQ ID 1]
93U01: GATCGGTACGCAGTC [SEQ ID 2]
XbaI:
90K02: CTCGTAGACTGCGTACA [SEQ ID 3]
92A16: CTAGTGTACGCAGTCT [SEQ ID 4]
ES 2 344 802 T3
EcoRI:
91M35: CTCGTAGACTGCGTACC [SEQ ID 5]
91M36: AATTGGTACGCAGTCTAC [SEQ ID 6]
HindIII:
91M35: CTCGTAGACTGCGTACC [SEQ ID 7]
91M37: AGCTGGTACGCAGTCTAC [SEQ ID 8]
Selective (+ 1 / + 1) amplifications (E / H and B / X) were performed using the following phosphorothioate (5'-3 ') primers:
BamHI + C-thio: 96R22thio: GACTGCGTACCGATCsCsC [SEQ ID 9]
Xbal + G: 96X03thio: GACTGCGTACACTAGsAsG [SEQ ID 10]
EcoRI + C-thio: 93T14thio: GACTGCGTACCAATTsCsC [SEQ ID 11]
HindIII + C-thio: 95H18thio GACTGCGTACCAGCTTsCsT [SEQ ID 12] with an "s" indicating the position of the phosphorothioate linkages of the oligonucleotides.
The mixtures of the reactions by AFLP had the following composition:
¿Ul 1/10 in AFLP template diluted with MQ ¿ul herculase II PCR buffer
0.5 μ dNTP (20 mM)
1.5 μ \ AFLP primer 1 (50 ng / jul)
1.5 μl AFLP 2 primer (50 ng / jul) μl Herculase® II Fusion DNA polymerase
30.5 μ! by MQ
The conditions of the PCR cycle were as follows:
Initial denaturation 94 ° C 2 min
Denaturation 94 ° C 10 sec
Mating 56 ° C 30 sec 10 cycles
Elongation 68 ° C 2 min
Denaturation 94 ° C 15 sec
Mating 56 ° C 30 sec 20 cycles
Elongation 68 ° C 2 min * * Setting: 20 sec per cycle
ES 2 344 802 T3
After AFLP amplification, reaction products were purified using Qiagen columns following the manufacturers' protocols.
C) Preparation of the 454 sequence library
Two 454 sequence libraries were prepared using purified BamHI / XbaI AFLP fragments and EcoRI / HindIII AFLP fragments as starting DNA, respectively, as described by Margulies et al., Beginning with nebulization (fragmentation) of the purified AFLP reaction products. A simple run of the 454 sequence was performed using a GS20 sequencing instrument (Roche Molecular Diagnostics), applying each of the fragment libraries of the two AFLP enzyme combinations to one half of a GS20 PicoTiterPlate.
D) Data processing
After completing the sequence run, the original data was processed using the GS20 RUNASSEMBLY program. Data obtained from the EcoRI / HindIII and BamHI / XbaI AFLP fragment libraries were processed separately and together, obtaining contigs from the overlapping sequence reads.
The contig obtained from RUNASSEMBLY were then mapped against the reference genome (contig 606 predicted in step a) above using RUNMAPPING, in order to determine to what extent the BamHI / XbaI + C / + predicted fragments had been sequenced. G and EcoRI / HindIII + C / + CT AFLP contained in contig 606. The coverage percentages obtained from the respective libraries are represented in the following table.
TABLE
Summary of Arabidopsis Contig 606 Sequence Coverage Summary
<td></td><td>% Expected</td><td>Observed Number of contigs at contig 606</td><td>% Coverage of the sequence observed in contig 606</td><td>Length of overlap (base pairs)</td><td>Coverage overlap (base pairs)</td><td>Coverage overlap (%)</td>
<td>EcoRI / HindIII</td><td> 32.9</td><td> 3</td><td> 35,2</td><td> 42</td><td> 42</td><td> 100</td>
<td>BamHI / Xbal</td><td> 71,4</td><td> 2</td><td> 59,0</td><td> 42</td><td> 42</td><td> 100</td>
<td>EcoRI / HindIII + BamHI / Xbal</td><td> 100</td><td> 4</td><td> 84,8</td><td> 42</td><td> 42</td><td> 100</td>
The resulting contig sequences are depicted in Figure 4.
These results demonstrate the feasibility of determining the genomic sequence of complex plant genomes by digesting total genomic DNA with multiple combinations of restriction enzymes by AFLP, followed by assembly of contig by fragment library and subsequent assembly of contig in sequence. genomics of the plant.
Contents11
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
70 members in 11 offices
Priority claims16
| Document | Office | Kind | Date |
|---|---|---|---|
| 69305305 | United States of America | P | |
| 69305305 | United States of America | P | |
| 71489705 | United States of America | P | |
| 71489705 | United States of America | P | |
| 06075104 | European Patent Office (EPO) | A | |
| 06075104 | European Patent Office (EPO) | A | |
| 75903406 | United States of America | P | |
| 75903406 | United States of America | P | |
| 06075104 | – | – | – |
| 06757809693053P | – | – | – |
| 714897P | – | – | – |
| 759034P | – | – | – |
| EP20060075104 | – | – | – |
| US20050693053P | – | – | – |
| US20050714897P | – | – | – |
| US20060759034P | – | – | – |
Members70
| Document | Office | Kind | |
|---|---|---|---|
| AU2006259990A1 | Australia | A1 | |
| CA2613248A1 | Canada | A1 | |
| WO2006137733A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2006137734A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1910562A1 | European Patent Office (EPO) | A1 | |
| EP1910563A1 | European Patent Office (EPO) | A1 | |
| CN101278058A | China | A | |
| JP2008546404A | Japan | A | |
| JP2008546405A | Japan | A | |
| US2009036323A1 | United States of America | A1 | |
| US2009142758A1 | United States of America | A1 | |
| CN101641449A | China | A | |
| EP1910563B1 | European Patent Office (EPO) | B1 | |
| AT465274T | Austria | T | |
| ATE465274T1 | Austria | T1 | |
| DE602006013831D1 | Germany | D1 | |
| ES2344802T3This record | Spain | T3 | |
| EP1910562B1 | European Patent Office (EPO) | B1 | |
| AT491045T | Austria | T | |
| ATE491045T1 | Austria | T1 | |
| DE602006018744D1 | Germany | D1 | |
| AU2006259990B2 | Australia | B2 | |
| EP2292788A1 | European Patent Office (EPO) | A1 | |
| DK1910562T3 | Denmark | T3 | |
| EP2302070A2 | European Patent Office (EPO) | A2 | |
| ES2357549T3 | Spain | T3 | |
| EP2302070A3 | European Patent Office (EPO) | A3 | |
| EP2292788B1 | European Patent Office (EPO) | B1 | |
| AT557105T | Austria | T | |
| ATE557105T1 | Austria | T1 | |
| DK2292788T3 | Denmark | T3 | |
| EP2302070B1 | European Patent Office (EPO) | B1 | |
| ES2387878T3 | Spain | T3 | |
| DK2302070T3 | Denmark | T3 | |
| US2012316073A1 | United States of America | A1 | |
| ES2393318T3 | Spain | T3 | |
| CN102925561A | China | A | |
| JP2013063095A | Japan | A | |
| JP5220597B2 | Japan | B2 | |
| CN101641449B | China | B | |
| US8685889B2 | United States of America | B2 | |
| US8785353B2 | United States of America | B2 | |
| US2014213462A1 | United States of America | A1 | |
| US2014274744A1 | United States of America | A1 | |
| US9023768B2 | United States of America | B2 | |
| US2015232924A1 | United States of America | A1 | |
| CN102925561B | China | B | |
| CN105039313A | China | A | |
| JP5823994B2 | Japan | B2 | |
| US2016060686A1 | United States of America | A1 | |
| US2016145689A1 | United States of America | A1 | |
| US9376716B2 | United States of America | B2 | |
| US9447459B2 | United States of America | B2 | |
| US9453256B2 | United States of America | B2 | |
| US9493820B2 | United States of America | B2 | |
| US2017137872A1 | United States of America | A1 | |
| US2017206314A1 | United States of America | A1 | |
| US2018025111A1 | United States of America | A1 | |
| US9896721B2 | United States of America | B2 | |
| US9898576B2 | United States of America | B2 | |
| US9898577B2 | United States of America | B2 | |
| US2018137241A1 | United States of America | A1 | |
| US2018247017A1 | United States of America | A1 | |
| US10095832B2 | United States of America | B2 | |
| CN105039313B | China | B | |
| US10235494B2 | United States of America | B2 | |
| US2019147977A1 | United States of America | A1 | |
| US2020185056A1 | United States of America | A1 | |
| US10978175B2 | United States of America | B2 | |
| US2021202035A1 | United States of America | A1 |
Numbers
- Publication, DOCDB
- 2344802
- Publication, EPODOC
- ES2344802T
- Application
- 6757809
- Application, DOCDB
- 06757809
- Application, EPODOC
- ES20060757809T
Titles2
- Spanish
- ESTRATEGIAS MEJORADAS PARA LA SECUENCIACION DE GENOMAS COMPLEJOS UTILIZANDO TECNOLOGIAS DE SECUENCIACION DE ALTO RENDIMIENTO.
- English
- IMPROVED STRATEGIES FOR THE SEQUENCING OF COMPLEX GENOMES USING HIGH PERFORMANCE SEQUENCING TECHNOLOGIES.
Classification
- CPC, 2
- C12Q1/6827
- C12Q1/6869
- IPC, 1
- C12Q1 68