Strategies for high throughput identification and detection of polymorphisms
Abstract
Use in a complexity reduction method of an adapter that carries a 3 'protruding end of T in the reduction of mixed labeling of an amplified DNA sample and / or in the reduction or prevention of concatamer formation of DNA fragments of a DNA sample comprising amplified restriction fragments bearing a 3 'protruding end of Obtained of a complexity reduction.

Term
Term ended
Projected expiry passed 23 June 2026, 0.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
6 claims: 2 independent, 4 dependent
- 1ES 2 393 318 T3 ES 2 393 318 T3 CLAIMS REIVINDICACIONES 1. Use, in a method of reducing complexity, of an adapter bearing a 3 'protruding end of T in the reduction of mixed labeling of an amplified DNA sample and / or in the reduction or prevention of the formation of concatamers of fragments of DNA from a DNA sample comprising amplified restriction fragments bearing a 3 'overhang of A obtained from a reduction in complexity. 1. Utilización en un método de reducción de la complejidad, de un adaptador que porta un extremo protuberante 3' de T en la reducción del etiquetado mixto de una muestra de ADN amplificado y/o en la reducción o prevención de la formación de concatámeros de fragmentos de ADN de una muestra de ADN que comprende fragmentos de restricción amplificados que portan un extremo protuberante 3' de A obtenido de una reducción de complejidad.
Independent claims2
435 paragraphs in 27 sections, as filed
ES 2 393 318 T3
DESCRIPTION
Strategies for the identification and detection of high-throughput polymorphisms.
Technical field
The present invention relates to the fields of molecular biology and genetics. The invention relates to the rapid identification of multiple polymorphisms in a nucleic acid sample. The identified polymorphisms can be used for the development of high-throughput screening systems for polymorphisms in test samples.
Background of the invention
The exploration of genomic DNA has been the desire of the scientific community, particularly the medical community, for a long time. Genomic DNA is the key to the identification, diagnosis, and treatment of diseases such as cancer and Alzheimer's disease. In addition to disease identification and treatment, genomic DNA scanning could provide significant advantages in plant and animal husbandry efforts, providing answers to food and nutrition problems around the world.
It is known that many diseases are associated with specific genetic components, in particular with polymorphisms in specific genes. Identifying polymorphisms in large samples, such as genomes, is currently a laborious and time-consuming task. However, this identification is of great value for areas such as biomedical research, pharmaceutical product development, tissue typing, genotyping, and population studies.
Summary description of the invention
The present invention provides a method for efficiently identifying and reliably detecting polymorphisms in a complex, eg large sample of nucleic acids (eg DNA or RNA) in a rapid and inexpensive manner using a combination of high throughput methods.
Such integration of high-throughput methods together provide a platform that is particularly suitable for the rapid and reliable identification and detection of polymorphisms in highly complex nucleic acid samples, where conventional identification and mapping of polymorphisms would be laborious and time-consuming.
One of the things the present inventors have found is a solution to identify polymorphisms, preferably single nucleotide polymorphisms, albeit similarly (micro) satellites and / or indels, particularly in large genomes. The method is unique in its applicability to both large and small genomes, although it provides particular advantages in large genomes, particularly polyploid species.
For identifying SNPs (and subsequently detecting the identified SNPs) several possibilities are available in the art. In a first option, the entire genome can be sequenced, and this can be carried out in several individuals. It is a largely theoretical exercise, being cumbersome and expensive, and despite the rapid development of technology, it is simply not feasible to apply to every organism, especially those with larger genomes. The second option is to use available (chunked) sequence information, such as EST libraries. This allows the generation of PCR primers, resequencing and comparison between individuals. Again the above requires initial sequence information that is not available or is available only in a limited amount. In addition, separate PCR assays must be developed for each region, which is a huge addition to development time and costs.
The third option is to limit yourself to part of each individual's genome. The difficulty is that the provided part of the genome must be the same in different individuals in order to provide a comparable result for the successful identification of SNPs. The present inventors have now solved this dilemma by integrating highly reproducible methods to select part of the genome by high throughput sequencing for the identification of polymorphisms integrated with sample preparation and high throughput identification platforms. The present invention accelerates the polymorphism identification procedure and uses the same elements in the subsequent procedure to exploit the discovered polymorphisms, allowing effective and reliable high-throughput genotyping.
Applications further contemplated in the present invention include screening-enriched microsatellite libraries, performing AFLP-transcript profiling cDNA (northern digital), complex genome sequencing, EST library sequencing (in full-length cDNA or in AFLP-cDNA), microRNA scanning (sequencing of small insert libraries), Bacterial artificial chromosome (BAC) sequencing (contig), AFLP / AFLP-cDNA in an analysis approach of
ES 2 393 318 T3 segregating clusters, routine detection of AFLP fragments, eg for marker assisted backcrossing (MABC), etc.
Definitions
In the description and examples later a series of expressions is used. In order to provide a clear and consistent understanding of the specification and claims, including the scope to be provided for such terms, the following definitions are provided. Unless defined otherwise herein, all technical and scientific terms used have the same meanings commonly understood by one of ordinary skill in the art to which the present invention belongs.
Polymorphism: Polymorphisms refer to the presence of two or more variants of a nucleotide sequence in a population. A polymorphism can comprise one or more base changes, an insertion, a repeat, or a deletion. A polymorphism includes, for example, a single sequence repeat (SSR) and a single nucleotide polymorphism (SNP), which is a variation that occurs in the event that a single nucleotide is altered: adenine (A), thymine ( T), cytosine (C) or guanine (G). A variation must generally occur in at least 1% of the population for it to be considered a SNP. SNPs make up 90% of all human genetic variations and occur every 100 to 300 bases throughout the human genome. Two out of three SNPs substitute cytosine (C) for thymine (T). Variations in the DNA sequences of, for example, humans or plants, can affect how they cope with diseases, bacteria, viruses, chemical compounds, drugs, etc.
Nucleic acid: a nucleic acid according to the present invention can include any pyrimidine or purine-based polymer or oligomer, preferably cytosine, thymine, and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, pages 793-800, Worth Publ. 1982 ). The present invention contemplates any deoxyribonucleotide, ribonucleotide or peptide-nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated or glycosylated forms of said bases, and the like. The polymers or oligomers can be of heterogeneous or homogeneous composition, and can be isolated from natural sources or produced artificially or synthetically. In addition, nucleic acids can be DNA or RNA, or a mixture thereof, and can exist permanently or transiently in single-stranded or double-stranded forms, including homoduplex, heteroduplex, and hybrid states.
Complexity reduction: the term complexity reduction is used to refer to a method in which the complexity of a nucleic acid sample, such as genomic DNA, is reduced by generating a subset of the sample. This subset may be representative of the entire sample (ie the complex sample) and is preferably a reproducible subset. The term "reproducible" in the present context refers to the fact that, by reducing the complexity of the same sample using the same method, the same or at least a comparable subset is obtained. The method used for the complexity reduction can be any method of reducing complexity known in the art. Examples of complexity reduction methods include, for example, AFLP® (Keygene NV, The Netherlands; see, for example, EP Patent No. 0 534 858), methods described by Dong (see, for example , WO 03/012118 and 00/24939), the indexed junction (Unrau et al., see below), etc. The complexity reduction methods used in the present invention have in common that they are reproducible. The term reproducible is used in the sense that the complexity of the same sample is reduced in the same way, the same subset of the sample is obtained, and not in the sense of a more random reduction in complexity, such as microdissection or the use of mRNA (cDNA), which represents a part of the transcribed genome in a selected tissue and its reproducibility depends on the selection of the tissue, the time of isolation, etc.
Labeling: The term labeling refers to the addition of a label to a nucleic acid sample in order to distinguish it from a second or subsequent nucleic acid samples. Labeling can be carried out, for example, by adding a sequence identifier during complexity reduction or by any other means known in the art. Said sequence identifier can be, for example, a single sequence of bases in length variable but defined, used only to identify a specific sample of nucleic acids. Typical examples thereof are, for example, ZIP sequences. By using such a label, the origin of a sample can be determined after further processing. In the event that processed products originating from different nucleic acid samples are combined, the different nucleic acid samples must be identified using different labels.
Tagged Library: The term tagged refers to a library of tagged nucleic acids.
Sequencing: The term sequencing refers to determining the order of nucleotides (base sequences) in a nucleic acid sample, for example DNA or RNA.
Align and Alignment: The term align and alignment refers to the comparison between two or more nucleotide sequences based on the presence of short or long segments of identical or similar nucleotides. Various methods for aligning nucleotide sequences are known in the art, as further explained
ES 2 393 318 T3 later.
Detection probes: The term detection probes is used to refer to probes designed to detect a specific sequence of nucleotides, in particular sequences that contain one or more polymorphisms.
High Throughput Screening - High throughput screening, often abbreviated HTS, is a method for scientific experimentation especially relevant to the fields of biology and chemistry. Using a combination of modern robotics and other specialized laboratory equipment, it enables the researcher to efficiently screen large quantities of samples simultaneously.
Test sample nucleic acids: The term "test sample nucleic acids" is used to denote a nucleic acid sample that is investigated for polymorphisms using the method of the present invention.
Restriction endonuclease: A restriction endonuclease or restriction enzyme is an enzyme that recognizes a specific nucleotide sequence (target site) in a double-stranded DNA molecule, and cuts both strands of the DNA molecule at all target sites.
Restriction fragments: DNA molecules produced by digestion with a restriction endonuclease are called restriction fragments. Any given genome (or nucleic acid, regardless of its origin) is digested by a particular restriction endonuclease into a discrete set of restriction fragments. DAN fragments that result from restriction endonuclease cleavage can further be used in a variety of techniques and can be detected by, for example, gel electrophoresis.
Gel electrophoresis: In order to detect restriction fragments, an analytical method may be necessary to fractionate double-stranded DNA molecules based on size. The most commonly used technique to achieve this fractionation is gel (capillary) electrophoresis. The rate at which DNA fragments move in such gels depends on their molecular weight; in this way, the distances traveled are reduced as the length of the fragment increases. The DNA fragments fractionated by gel electrophoresis can be directly visualized by a staining procedure, for example silver staining or staining using ethidium bromide, in case the number of fragments included in the standard is sufficiently small. Alternatively, further processing of the DNA fragments can incorporate detectable labels on the fragments, such as fluorophores or radioactive labels.
Ligation: The enzymatic reaction catalyzed by a ligase enzyme in which two double-stranded DNA molecules are covalently linked together is called ligation. In general, both DNA strands are covalently linked to each other, although it is also possible to avoid ligation of one of the two strands by chemical or enzymatic modification of one of the ends of the strands. In this case, covalent bonding will occur in only one of the two DNA strands.
Synthetic oligonucleotide: Single-stranded DNA molecules that preferably have between about 10 and about 50 bases, that can be chemically synthesized, are called synthetic oligonucleotides. In general, these synthetic DNA molecules are designed to have a unique or desired nucleotide sequence, although it is possible to synthesize families of molecules that have related sequences and that have different nucleotide compositions at specific positions within the nucleotide sequence. The term "synthetic oligonucleotide" is used to refer to DNA molecules that have a designed or desired nucleotide sequence.
Adapters: short double-stranded DNA molecules with a limited number of base pairs, eg, between about 10 and about 30 base pairs in length, that are designed so that they can bind to the ends of restriction fragments. Adapters are generally composed of two synthetic oligonucleotides that have nucleotide sequences that are partially complementary to each other. By mixing the two synthetic oligonucleotides in solution under appropriate conditions, they pair with each other forming a double-stranded structure. After hybridization, one end of the adapter molecule is designed so that it is compatible with the end of a restriction fragment and can be ligated thereto; the other end of the adapter can be designed in such a way that it cannot be bonded, although this is not necessarily the case (doubly bonded adapters).
Adapter-linked restriction fragments: Restriction fragments to which adapter caps have been added.
Primers: In general, the term primers refers to a strand of DNA that can prime DNA synthesis. DNA polymerase cannot synthesize DNA de novo without primers: it can only extend an existing DNA strand in a reaction in which the complementary strand is used as a template to direct the order of nucleotides to be assembled. Reference is made to the synthetic oligonucleotide molecules that are used
ES 2 393 318 T3 in a polymerase chain reaction (PCR) as primers.
DNA amplification: The term DNA amplification is typically used to refer to the in vitro synthesis of double-stranded DNA molecules using PCR. It is noted that other amplification methods exist and that they can be used in the present invention without departing from the essence thereof.
Detailed description of the invention
The present invention provides a method for identifying one or more polymorphisms, said method comprising the steps of:
a) provide a first sample of nucleic acids of interest,
b) carry out a complexity reduction of the first nucleic acid sample of interest, providing a first library of the first nucleic acid sample,
c) carry out consecutively or simultaneously steps a) and b) with a second or subsequent sample of nucleic acids of interest, obtaining a second or subsequent library of the second or subsequent sample of nucleic acids of interest,
d) sequencing at least a part of the first library and the second or subsequent libraries,
e) aligning the sequences obtained in step d),
f) determining one or more polymorphisms between the first nucleic acid sample and the second or subsequent nucleic acid sample in the alignment of step e),
g) use the polymorphism or polymorphisms determined in step f) to design one or more detection probes,
h) provide a test sample of nucleic acids of interest,
i) carrying out the complexity reduction of step b) on the nucleic acid test sample of interest, providing a test library of the nucleic acid test sample,
j) subjecting the test library to high throughput screening to identify the presence, absence or quantity of polymorphisms determined in step f) using the detection probes designed in step g).
In step a), a first sample of nucleic acids of interest is provided. Said first nucleic acid sample of interest is preferably a complex nucleic acid sample, such as total genomic DNA or a cDNA library. It is preferred that the complex nucleic acid sample is total genomic DNA.
In step b), a reduction of the complexity of the first nucleic acid sample of interest is carried out, providing a first library of the first nucleic acid sample.
In one embodiment of the invention, the step of reducing the complexity of the nucleic acid sample comprises enzymatically cutting the nucleic acid sample into restriction fragments, separating the restriction fragments, and selecting a particular group of restriction fragments. Optionally, the selected fragments are then ligated to adapter sequences containing templates / PCR primer binding sequences.
In a complexity reduction embodiment, a type II endonuclease is used to digest the nucleic acid sample and the restriction fragments are selectively ligated to adapter sequences. The adapter sequences may contain various nucleotides at the overhang to be ligated and only the adapter with the corresponding set of nucleotides at the overhang is ligated to the fragment and subsequently amplified. This technology is described in the art as indexing connectors. Examples of this principle can be found in, among others, Unrau P. and Deugau KV, Gene 145: 163-169, 1994.
In another embodiment, the complexity reduction method uses two restriction endonucleases having different target sites and frequencies and two different adapter sequences.
In another embodiment of the invention, the complexity reduction step comprises performing an arbitrarily primed PCR on the sample.
In yet another embodiment of the invention, the complexity reduction step comprises removing repeat sequences by denaturation and rehybridization of DNA and subsequent removal of double-stranded duplexes.
In another embodiment of the invention, the complexity reduction step comprises hybridizing the nucleic acid sample with a magnetic bead that binds to an oligonucleotide probe containing a desired sequence. This embodiment may further comprise exposing the hybridized sample to a single-stranded DNA nuclease to remove the single-stranded DNA and ligating an adapter sequence containing a class II restriction enzyme to release the magnetic bead. This embodiment may or may not comprise amplification of the isolated DNA sequence. Furthermore, the adapter sequence may or may not be used as a template for the oligonucleotide PCR primer. In this embodiment, the adapter sequence may or may not contain a sequence tag or identifier.
ES 2 393 318 T3
In another embodiment, the complexity reduction method comprises exposing the DNA sample to a mismatch binding protein and digesting the sample with a 3 'to 5' exonuclease and then a single stranded nuclease. This embodiment may or may not include the use of a magnetic bead attached to the mismatched binding protein.
In another embodiment of the present invention, complexity reduction comprises the CHIP method as described hereinafter, or the design of PCR primers directed against conserved motifs, such as SSRs, NBS regions (nucleotide binding regions ), promoter / enhancer sequences, telomere consensus sequences, MADS box genes, ATPase gene families, and other gene families.
In step c), steps a) and b) are carried out consecutively or simultaneously with a second or subsequent sample of nucleic acids of interest, obtaining a second or subsequent library of the second or subsequent sample of nucleic acids of interest. Said second or subsequent nucleic acid sample of interest is preferably also a complex nucleic acid sample, such as total genomic DNA. It is preferred that the complex nucleic acid sample is total genomic DNA. It is also preferred that said second or subsequent nucleic acid sample is related to the first nucleic acid sample. The first nucleic acid sample and the second or subsequent nucleic acid can be, for example, different lines of a plant, such as different lines of the pepper plant, or different varieties. Steps a) and b) can be carried out for merely a second nucleic acid sample of interest, although they can also be carried out additionally for a third, fourth, fifth, etc. nucleic acid sample of interest.
It should be noted that the method according to the present invention will be most useful in carrying out complexity reduction using the same method and under substantially the same conditions, preferably identical, for the first nucleic acid sample and the second or subsequent nucleic acid samples. Under such conditions, similar (comparable) fractions of the (complex) nucleic acid samples are obtained.
In step d), at least a part of the first library and the second or subsequent libraries are sequenced. It is preferred that the amount of overlap of the sequenced fragments of the first library and second or subsequent libraries is at least 50%, more preferably at least 60%, still more preferably at least 70%, still more preferably of at least 80%, still more preferably of at least 90% and still more preferably of at least 95%.
Sequencing can be carried out, in principle, by any means known in the art, such as the dideoxy chain termination method. However, it is preferred that sequencing is carried out using high-throughput sequencing methods, such as the methods disclosed in WO 03/004690, 03/054142, 2004/069849, n 2004/070005, 2004/070007 and 2005/003375 (all in the name of 454 Corporation), in Seo et al., Proc. Natl. Acad. Sci. USA 101: 5488-93, 2004, and the techniques of Helios, Solexa, US Genomics, etc. It is more preferred that the sequencing is carried out using the apparatus and / or method disclosed in WO 03/004690, 03/054142, 2004/069849, 2004/070005, no. 2004/070007 and no 2005/003375 (all in the name of 454 Corporation). The technology described allows the sequencing of 40 million bases in a single operation and is 100 times faster and cheaper than competing technology. Sequencing technology generally consists of 4 steps: 1) DNA fragmentation and ligation of specific adapters to a single stranded DNA library (ssDNA), 2) hybridization of ssDNA to beads and emulsification of the beads in water microreactors in oil, 3) deposition of the DNA-bearing beads on a PicoTiterPlate® plate, and 4) simultaneous sequencing in 100,000 wells by generating a pyrophosphate light signal. The method is explained in more detail later.
In step e), the sequences obtained in step d) are aligned providing an alignment. Sequence alignment methods for comparison purposes are well known in the art. Various alignment algorithms and programs are described in: Smith and Waterman (1981) Adv. Appl. Math. 2: 482; Needleman and Wunsch, J. Mol. Biol. 48: 443, 1970; Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85: 2444, 1988; Higgins and Sharp, Gene 73: 237-244, 1988; Higgins and Sharp, CABIOS 5: 151-153, 1989; Corpet et al., Nucl. Acids Res. 16: 10881-90, 1988; Huang et al., Computer Appl. in the Biosci. 8: 155-65, 1992; and Pearson et al., Meth. Mol. Biol. 24: 307-31, 1994. Altschul et al., Nature Genet. 6: 119-29, 1994, present a detailed discussion of sequence alignment methods and homology calculations.
NCBI's Basic Local Alignment Search Tool (BLAST) (Altschul et al., 1990) is available from various sources, including the National Center for Biological Information (NCBI, Bethesda, Md.) And on the Internet, for use with the blastp, blastn, blastx, tblastn, and tblastx sequence analysis programs. They can be accessed at <http://www.ncbi.nlm.nih.gov/BLAST/>. A description of how to determine sequence identity using such a program is available at <http: //www.ncbi.nlm.nih.gov/BLAST/blast_help.html>. An additional application could be microsatellite exploration (see Varshney et al., Trends in Biotechn. 23 (1): 48-55, 2005).
Typically the alignment is carried out on sequence data that has been clipped for the
ES 2 393 318 T3 adapters / primers and / or identifiers, that is, using only the sequence data of the fragments originating from the nucleic acid sample.
Typically, the sequence data obtained is used to identify the origin of the fragment (i.e., the provenance sample), the adapter and / or identifier derived sequences are removed from the data, and alignment is performed on this trimmed set. .
In step f), one or more polymorphisms between the first nucleic acid sample and the second or later nucleic acid sample in the alignment are determined. The alignment can be performed so that the sequences derived from the first nucleic acid sample and the second or subsequent nucleic acid samples can be compared. The differences that reflect polymorphisms can then be identified.
In step g), the polymorphism (s) determined in step g) are used to design detection probes, for example for detection by hybridization on DNA chips or on a bead-based detection platform. Detection probes are designed so that a polymorphism is reflected therein. In the case of single nucleotide polymorphisms (SNPs), detection probes typically contain the variant alleles of SNPs in the central position so as to maximize allele discrimination. Such probes can advantageously be used to screen test samples exhibiting a certain polymorphism. Probes can be synthesized using any method known in the art. Probes are typically designed to be suitable for high throughput screening methods.
In step h), a nucleic acid test sample of interest is provided. The nucleic acid test sample can be any sample, although preferably it is another line or variety that must be mapped to identify polymorphisms. A collection of test samples representing the germ plasm of the organisms studied is commonly used in order to experimentally validate that the polymorphism (SN) is genuine and detectable, and to calculate the allele frequencies of the observed alleles. Optionally, samples from a genetic mapping population are included in the validation stage in order to also determine the position in the genetic map of the polymorphism.
In step i), the complexity reduction of step b) is carried out on the nucleic acid test sample of interest, providing a test library of the nucleic acid test sample. It is highly preferred that throughout the method according to the present invention, the same method for complexity reduction is used, using substantially the same, preferably identical conditions, thus covering a similar fraction of the sample. However, it is not necessary to obtain a tagged test library, although a tag may be present on the fragments in the test library.
In step j), the test library is subjected to high-throughput screening to identify the presence, absence or quantity of the polymorphisms determined in step f) using the detection probes designed in step g). The person skilled in the art knows several methods for high throughput screening using probes. It is preferred that one or more probes designed using the information obtained in step g) are immobilized on a matrix, such as a DNA chip, and that said matrix is subsequently contacted with the test library under hybridization conditions. DNA fragments in the test library that are complementary to one or more probes on the array will hybridize under such conditions to such probes, and can be detected in this way. Other high-throughput screening methods are also contemplated within the scope of the present invention, such as immobilization of the test library obtained in step j) and contacting said immobilized test library with the probes designed in step h) under hybridization conditions.
Affymetrix, among others, provides another high-throughput sequencing screening technique that uses chip-based detection of SNPs and bead technology provided by Illumina.
In an advantageous embodiment, step b) in the method according to the present invention further comprises the step of labeling the library to obtain a labeled library, and said method further comprises step c1) of combining the first labeled library and a second or subsequent ones tagged libraries to get a combined library.
It is preferred that labeling is carried out during the complexity reduction step to reduce the number of steps required to obtain the first labeled library from the first nucleic acid sample. Such simultaneous labeling can be achieved by, for example, AFLP, using adapters that comprise a unique identifier (nucleotide) for each sample.
The labeling aims to distinguish between samples of different origin, for example obtained from different plant lines, in the case that libraries of two or more nucleic acid samples are combined to obtain a combination library. Thus, different labels are preferably used to prepare the labeled libraries of the first nucleic acid sample and the second or subsequent acid samples.
ES 2 393 318 T3 nucleic. In the case where, for example, five nucleic acid samples are used, it is intended to obtain five differently labeled libraries, the five different labels indicating the respective original samples.
The tag can be any tag known in the art to distinguish nucleic acid samples, but is preferably a short identifier sequence. Said identifier sequence can be, for example, a unique sequence of bases of variable length used to indicate the origin of the library obtained by reduction of complexity.
In a preferred embodiment, the labeling of the first library and the second or subsequent libraries is carried out using different labels. As discussed above, it is preferred that each nucleic acid sample library is identified with its own label. The nucleic acid test sample does not need to be labeled.
In a preferred embodiment of the invention, complexity reduction is carried out by means of AFLP<sup>® </sup>(Keygene NV, The Netherlands, see, for example, EP Patent No. 0 534 858 and Vos et al., (1995). AFLP: a new technique for DNA fingerprinting, Nucleic Acids Research 23 (21): 4407-4414 , nineteen ninety five).
AFLP is a method for the selective amplification of restriction fragments. AFLP does not require prior sequence information and can be carried out on any starting DNA. In general, the AFLP comprises the stages of:
(a) digestion of a nucleic acid, in particular a DNA or cDNA, with one or more specific restriction endonucleases, to fragment the DNA into a corresponding series of restriction fragments, (b) ligation of the restriction fragments obtained therefrom manner with a synthetic double-stranded oligonucleotide adapter, one end of which is compatible with one or both ends of the restriction fragments, thereby producing adapter-linked, preferably tagged, restriction fragments of the starting DNA, (c) contacting the adapter-linked, preferably tagged restriction fragments, under hybridization conditions with at least one oligonucleotide primer that contains at least one selective nucleotide at its 3 'end, (d) amplification of restriction fragments linked with adapters, preferably labeled, hybridized to the primers by PCR or a similar technique such that additional elongation of the hybridized primers is caused along the restriction fragments of the starting DNA to which the primers were hybridized, and (e) detecting, identifying or recovering the amplified or elongated DNA fragment obtained in this way.
AFLP thus provides a reproducible subset of adapter-ligated fragments. Other suitable methods for complexity reduction are chromatin immunoprecipitation (ChiP). The above refers to the isolation of nuclear DNA, whereas proteins such as transcription factors bind to DNA. With ChiP an antibody against the protein is first used, resulting in an Ab-DNA protein complex. Through the purification of this complex and its precipitation, the DNA to which said protein binds is selected. The DNA can then be used for library construction and sequencing. That is, it is a method to carry out a complexity reduction in a non-random manner directed to functional areas; in the present example, specific transcription factors.
A useful variant of AFLP technology uses non-selective nucleotides (ie + 0 / + 0 primers) and is sometimes referred to as linker PCR. It also provides a very adequate complexity reduction.
For a further description of AFLP, its advantages, embodiments, as well as the techniques, enzymes, adapters, primers, and additional compounds and tools used therein, reference is made to US Patents No. 6,045,994, EP No. B -0 534 858, EP No. 976835 and EP No. 974672, WO No. 01/88189, and Vos et al., Nucleic Acids Research 23: 4407-4414, 1995.
Thus, in a preferred embodiment of the method of the present invention, complexity reduction is carried out by:
- digestion of the nucleic acid sample with at least one restriction endonuclease to fragment it into restriction fragments,
ligation of the restriction fragments obtained with at least one synthetic double-stranded oligonucleotide adapter having an end compatible with one or both ends of the restriction fragments to produce adapter-ligated restriction fragments,
- contacting said adapter-linked restriction fragments with one or more oligonucleotide primers under hybridization conditions, and
- amplification of said adapter-ligated restriction fragments by elongation of one or more of the oligonucleotide primers,
ES 2 393 318 T3 in which at least one of the oligonucleotide primer (s) includes a nucleotide sequence having the same nucleotide sequence as the terminal parts of the strands at the ends of said adapter-linked restriction fragments, including the nucleotides involved in the formation of the target sequence for said restriction endonuclease, and including at least part of the nucleotides present in the adapters, wherein, optionally, at least one of said primers includes at its 3 'end a selected sequence comprising at least one nucleotide located immediately adjacent to the nucleotides involved in the formation of the target sequence for said restriction endonuclease.
AFLP is a highly reproducible method for reducing complexity and is therefore particularly suitable for the method according to the present invention.
In a preferred embodiment of the method according to the present invention, the adapter or primer comprises a tag. This is particularly the case for the identification of polymorphisms, where it is important to distinguish between sequences derived from separate libraries. Incorporation of an oligonucleotide tag into an adapter or primer is very convenient because no additional steps are necessary to tag a library.
In another embodiment, the tag is an identifier sequence. As discussed above, said identifier sequence can be of variable length depending on the number of nucleic acid samples to be compared. A length of about 4 bases (4<sup>4</sup>= 256 possible different tag sequences) to distinguish between the origin of a limited number of samples (256 maximum), although it is preferred that the tag sequences differ by no more than one base between the samples to be distinguished. As needed, the length of the tag sequences can be adjusted.
In one embodiment, sequencing is carried out on a solid support, such as a bead (see, for example, WO patents No. 03/004690, No. 03/054142, No. 2004/069849, No. 2004 / 070005, No. 2004/070007 and No. 2005/003375 (all owned by 454 Corporation) This sequencing method is particularly suitable for the economical and efficient sequencing of many samples simultaneously.
In a preferred embodiment, sequencing comprises the steps of:
- bead binding adapter-linked fragments, each bead being attached to a single adapter-linked fragment,
- emulsifying the beads in water-in-oil microreactors, each water-in-oil microreactor comprising a single bead,
- loading the beads into wells, each well comprising a single bead, and
- generate a pyrophosphate signal.
In the first step, sequencing adapters are ligated to fragments within the combining library. Said sequencing adapter includes at least one key region for bead binding, a sequencing primer region, and a PCR primer region. In this way fragments linked to adapters are obtained.
In a further step, adapter-linked fragments are joined to beads, with each bead binding to a single adapter-linked fragment. Excess beads are added to the pool of adapter-linked fragments to ensure the binding of a single adapter-linked fragment per bead for most beads (Poisson distribution).
In the next step, the beads are emulsified in water-in-oil microreactors, each water-in-oil microreactor comprising a single bead. The PCR reagents present in the water-in-oil microreactors allow a PCR reaction to take place in the microreactors. The microreactors are then disrupted and enriched for beads comprising DNA (DNA positive beads).
In a later step, the beads are loaded into wells, each well comprising a single bead. The wells are preferably part of a PicoTiter ™ plate that allows simultaneous sequencing of a large number of fragments.
After the addition of enzyme-bearing beads, the sequence of the fragments is determined by pyrosequencing. In successive steps, the picotiter plate and the beads, as well as the beads with enzyme on it, are subjected to different deoxyribonucleotides in the presence of conventional sequencing reagents, and after the incorporation of a deoxyribonucleotide, a light signal is generated and recorded. Incorporation of the correct nucleotide generates a pyrosequencing signal that can be detected.
Pyrosequencing itself is known in the art and is described in others at www.biotagebio.com,
ES 2 393 318 T3 www.pyrosequencing.com/tab technology. The technology is further applied in, for example, patents WO No. 03/004690, No. 03/054142, No. 2004/069849, No. 2004/070005, No. 2004/070007 and No. 2005/003375 (all to name of 454 Corporation).
The high throughput screening of step k) is preferably carried out by immobilization of the probes designed in step h) on an array, followed by contacting the array comprising the probes with an assay library under conditions hybridization. Preferably, the contacting step is carried out under stringent hybridization conditions (see Kennedy et al., Nat. Biotech., Published on the internet on September 7, 2003, pages 1 to 5). The person skilled in the art is aware of the existence of suitable methods for the immobilization of probes on an array and of methods of contacting under hybridization conditions. Typical technology that is suitable for this purpose is reviewed in Kennedy et al., Nat. Biotech., Published online September 7, 2003, pages 1-5). 1-5.
A particular advantageous application is the cultivation of polyploid species. By sequencing polyploid cultures with high coverage, identifying SNPs and various alleles, and developing probes for allele-specific amplification, significant advances can be made in culturing polyploid species.
As part of the invention, the combination of generating randomly selected subsets by selective amplification for a plurality of samples and high throughput sequencing technology has been found to present certain complex problems that needed to be solved to further improve the method described herein. for efficient and high-throughput identification of polymorphisms. In more detail, it has been found that when combining multiple samples (that is, the first and second or later) in a group after performing a complexity reduction, the problem occurs that many fragments apparently come from two samples or, in other words, many fragments were identified that could not be uniquely assigned to a sample and thus could not be used in the polymorphism identification procedure. This led to a reduction in the reliability of the method and to polymorphisms (SNPs, indels, SSRs) that could not be adequately identified.
After careful and detailed analysis of the complete nucleotide sequence of the fragments that could not be located, it was found that those fragments contained two different adapters that comprised tags and that were likely formed between the generation of the reduced complexity samples and the ligation of sequencing adapters. The phenomenon is described as mixed labeling. The phenomenon described as mixed labeling, as used herein, thus refers to fragments that contain a label that relates the fragment to a sample on one side, while the opposite side of the fragment contains a label that relate the fragment to another sample. In this way, a fragment is apparently derived from two samples (quod non). This leads to misidentification of polymorphisms and is therefore undesirable.
It has been theorized that the formation of heteroduplex fragments between two samples is at the root of this anomaly.
The solution to this problem has been found in a redesign of the strategy for converting samples from which the complexity has been reduced into bead-bound fragments that can be amplified prior to high-throughput sequencing. In the present embodiment, each sample undergoes complexity reduction and optional purification. Then, blunt ends are generated on each sample (end polishing) followed by ligation of the sequencing adapter that is capable of binding to the bead. Sequencing adapter-ligated fragments from the samples are then pooled and ligated to the beads for emulsion polymerization and subsequent high-throughput sequencing.
As a further part of the present invention, it has been found that the formation of concatamers makes it difficult to correctly identify polymorphisms. Concatamers have been identified as fragments that are formed after blunt ends or polishing of complexity reduction products, for example with T4 DNA polymerase, and instead of linking them to the adapters that allow binding to the beads, they are ligated between yes, thus creating concatámers, that is to say, a concatámero is the result of the dimerización of fragments of blunt ends.
The solution to this problem was found in the use of certain specifically modified adapters. Amplified fragments obtained from complexity reduction typically contain a 3'-A overhang due to the characteristics of certain preferred polymerases, which do not exhibit 3'-5 'exonuclease error-correcting activity. The presence of such a 3'-A overhang is also the reason why blunt ends are formed in fragments prior to ligation of adapters. By providing an adapter that could be attached to a bead in which the adapter contains a 3'-T protruding end, it was found that both the problem of mixed labels and that of concatamers could be solved in one step. A further advantage of using such modified adapters is that the conventional blunt-ended step and subsequent phosphorylation step could be omitted.
ES 2 393 318 T3
Thus, in a further preferred embodiment, after the step of reducing the complexity of each sample, a step is carried out on the amplified restriction fragments linked with adapters that have been obtained from the step of reducing complexity, so that these fragments are linked by sequencing adapters, which contain a 3'-T overhang and are capable of binding to the beads.
It has also been found that, by phosphorylating the primers used in the complexity reduction step, the end polishing step (blunt end formation) and the phosphorylation of intermediates prior to ligation can be avoided.
Thus, in a highly preferred embodiment of the invention, the invention relates to a method for identifying one or more polymorphisms, said method comprising the steps of:
a) provide a plurality of nucleic acid samples of interest,
b) carry out a complexity reduction of each of the samples, providing a plurality of libraries of nucleic acid samples, in which the complexity reduction is carried out by:
- digestion of each nucleic acid sample with at least one restriction endonuclease to fragment it into restriction fragments,
ligation of the restriction fragments obtained with at least one synthetic double-stranded oligonucleotide adapter having an end compatible with one or both ends of the restriction fragments to produce adapter-ligated restriction fragments,
- contacting said adapter-ligated restriction fragments with one or more phosphorylated oligonucleotide primers under hybridization conditions, and
- amplification of said adapter-ligated restriction fragments by elongation of one or more oligonucleotide primers, wherein at least one of one or more oligonucleotide primers includes a nucleotide sequence presenting the same nucleotide sequence as the terminal parts of the strands at the ends of said adapter-ligated restriction fragments, including the nucleotides involved in the formation of the target sequence of said restriction endonuclease, and including at least part of the nucleotides present in the adapters, in which, optionally, at least one of said primers includes at its 3 'end a selected sequence comprising at least one nucleotide located immediately adjacent to the nucleotides involved in the formation of the target sequence of said restriction endonuclease and in which the adapter and / or the primer contain a label,
c) combining said libraries to form a combined library,
d) ligating the bead-binding sequencing adapters with the amplified adapter cap fragments in the combined library, using a sequencing adapter bearing a 3'-T overhang, and subjecting the bead-bound fragments to polymerization in emulsion, e) sequencing at least a part of the combined library,
f) aligning the sequences of each sample obtained in step e),
g) determining one or more polymorphisms among the plurality of nucleic acid samples in the alignment of step f),
h) use the polymorphism or polymorphisms determined in step g) to design detection probes,
i) providing a test sample nucleic acid of interest,
j) performing the complexity reduction of step b) on the nucleic acid of the test sample of interest to provide a test library of the nucleic acid of the test sample,
k) subjecting the test library to high throughput screening to identify the presence, absence or quantity of polymorphisms determined in step g) using the detection probes designed in step h).
Brief description of the drawings
Figure 1A shows a fragment according to the present invention attached to a bead (454 bead) and the sequence of the primer used for the preamplification of the two pepper plant lines. The term DNA fragment refers to the fragment obtained after digestion with a restriction endonuclease, Keygene adapter refers to an adapter that provides a binding site for the oligonucleotide (phosphorylated) primers used to generate a library, KRS refers to a identifier sequence (tag), adapter SEQ. 454 refers to a sequencing adapter, and 454 PCR adapter refers to an adapter that allows emulsion amplification of the DNA fragment. The PCR adapter allows bead binding and amplification, and may contain a 3'-T overhang.
Figure 1B shows a schematic primer used in the complexity reduction step. Said primer generally comprises a recognition site region indicated as (2), a constant region that may include a tag section indicated as (1) and one or more selective nucleotides in a selective region indicated as (3) at the 3 end ' thereof).
Figure 2 shows the estimation of DNA concentration using 2% agarose gel electrophoresis. S1 se
ES 2 393 318 T3 refers to PSP11; S2 refers to PI201234. 50, 100, 250 and 500 ng are referred to, respectively, 50 ng, 100 ng, 250 ng and 500 ng to estimate the amounts of S1 and S2 DNA. Figs. 2C and 2D show the determination of DNA concentration using NanoDrop spectrophotometry.
Figure 3 shows the results of the intermediary quality evaluations of Example 3.
Figure 4 shows flow charts of the sequence data processing, that is, the steps between the generation of the sequencing data and the identification of the putative SNPs, SSRs and indels, through steps of eliminating known sequence information in clipping. and tagging, resulting in tight sequence data that is grouped and assembled to provide contigs and singletons (fragments that cannot be assembled to form a contig), after which putative polymorphisms can be identified and evaluated. Figure 4B provides additional details of the polymorphism screening procedure.
Figure 5 refers to the problem of mixed labels and provides in panel 1 an example of a mixed label that includes labels associated with sample 1 (MS1) and sample 2 (MS2). Panel 2 provides a schematic explanation of the phenomenon.
AFLP restriction fragments derived from sample 1 (S1) and sample 2 (S2) are ligated using adapters (Keygene adapter) at both ends bearing specific labels from samples S1 and S2. After amplification and sequencing, the expected fragments have the S1-S2 tags and the S2-S2 tags. Furthermore, fragments bearing S1-S2 or S2-S1 tags were unexpectedly observed. Panel 3 explains the hypothetical cause of the generation of mixed tags, whereby heteroduplex products are formed from fragments of samples 1 and 2. Heteroduplexes are subsequently released, due to the 3-5 'exonuclease activity of DNA. T4 or Klenow polymerase, relative to the 3'-overhanging ends. During polymerization, the gaps are filled with nucleotides and the wrong label is entered. This works for heteroduplexes of roughly the same length (top panel), but also for heteroduplexes of more variable length. Panel 4 provides on the right hand side the conventional protocol leading to the formation of mixed labels and on the right hand side the modified protocol.
Figure 6 refers to the problem of concatemer formation, in which panel 1 provides a typical example of a concatemer, in which the various adapter and tag sections are underlined and with their origin (i.e. MS1, MS2, ES1 and ES2, corresponding respectively to an adapter-Msel restriction site from sample 1, adapter-Msel restriction site from sample 2, adapter-EcoRI restriction site from sample 1, adapter-EcoRI restriction site from sample 2). Panel 2 shows the expected fragments bearing the S1-S1 and S2-S2 tags and the observed but unexpected S1-S1-S2-S2, which is a concatamer of fragments from samples 1 and 2. Panel 3 provides the solution to avoid the generation of concatamers, as well as mixed labels, by introducing a protruding end in the AFLP adapters, modified sequencing adapters and bypassing the end polishing step when ligating the adapters sequencing. Concatamer formation was not observed because the ALP fragments cannot bind to each other and mixed fragments are not produced because the end polishing step is omitted. Panel 4 provides the modified protocol, which uses modified adapters to avoid the formation of concatamers, as well as mixed tags.
Figure 7. Multiple alignment 10037_CL989contig2 of AFLP fragment sequences from the pepper plant, containing a putative single nucleotide polymorphism (SNP). Note that the SNP (indicated by a black arrow) is defined by an A allele present in both sample 1 reads (PSP11), indicated by the presence of the MS1 tag in the name of the top two reads and a G allele. present in sample 2 (PI201234), indicated by the presence of the MS2 tag in the name of the two bottom reads. The reading names are displayed on the left. The consensus sequence of this multiple alignment is (5'3 '):
TAACACGACTTTGAACAAACCCAAACTCCCCCAATCGATTTCAAACCTAGAACA [A / G] TGTTGGTTTT GGTGCTAACTTCAACCCCACTACTGTTTTGCTCTATTTTTG Figure 8A. Schematic representation of the single sequence targeting repeat (SSR) enrichment strategy in combination with high-throughput sequencing for de novo identification of SSRs.
Figure 8B: Validation of a SNP G / A in the pepper plant using SNPWave detection. P1 = PSP11, P2 = PI201234. The eight RIL descendants are indicated with the numbers 1 to 8.
Examples
Example 1
An EcoRI / Msel (1) restriction ligation mixture was generated from genomic DNA of the pepper plant lines PSP-11 and PI20234. The restriction ligation mix was diluted 10-fold and 5
ES 2 393 318 T3 microliters of each sample (2) with primers EcoRI +1 (A) and Msel +1 (C) (set I). After amplification, the quality of the preamplification product of the two pepper samples was checked on a 1% agarose gel. Pre-amplification products were diluted 20-fold, followed by AFLP pre-amplification with KRSEcoRI +1 (A) and KRSMseI +2 (CA). The KRS segments (identifiers) are underlined and the selective nucleotides are in bold, at the 3 'end of the sEc ID 1 to 4 primer sequences, below. After amplification, the quality of the preamplification product of the two pepper samples was checked on a 1% agarose gel and by means of the genetic fingerprint technique using AFLP (4) with EcoRI +3 (A) and MseI + 3 (C) (3). The preamplification products of the two pepper lines were purified separately on a Qiagen PCR column (5). The concentration of the samples was measured on the NanoDrop. A total of 5,006.4 ng of PSP-11 and 5,006.4 ng of PI20234 were mixed and sequenced.
Primer set I used for PSP-11 preamplification:
E01LKRS1 5'-CGTCAGACTGCGTACCAATTCA-3 '[SEQ ID 1]
M15KKRS1 5'-TGGTGATGAGTCCTGAGTAACA-3 '[SEQ ID 2]
Primer set II used for preamplification of PI20234:
E01LKRS2 5'-CAAGAGACTGCGTACCAATTCA-3 '[SEQ ID 3]
M15KKRS2 5'-AGCCGATGAGTCCTGAGTAACA-3 '[SEQ ID 4] (1) EcoRI / Msel restriction ligation mixture
Restriction Mix (40ul / sample)
<td>DNA ECoRI (5U) MseI (2U) 5xRL Mq Total Incubation Addition of:</td><td>6 pl (± 300 ng) 0.1 μΙ 0.05 μΙ 8 μΙ 25.85 μΙ 40 μΙ for 1 hr at 37 ° C</td>
Ligation Mix (10 ul / sample):
<td>10 mM ATP T4 DNA ligase Ligation mix EcoRI adapter (5 pmol / pl) MseI adapter (50 pmol / pl) 5xRL Mq Total Incubation EcoRI 91M35 / 91M36 adapter:</td><td>1 pl 1 pl (10 pl / sample) 1 pl 1 pl 2 pl 4 pl 10 pl for 3 hours at 37 ° C * -CTCGTAGACTGCGTACC: 91M35 [SEQ ID 5]</td>
<td>± bio Adapter MseI 92A18 / 92A19:</td><td>CATCTGACGCATGGTTAA: 91M36 [SEQ ID 6] 5-GACGATGAGTCCTGAG-3: 92A18 [SEQ ID 7] 3-TACTCAGGACTCAT-5: 92A19 [SEQ ID 8]</td>
(2) Preamp
Preamp (A / C): Mix RL (10x)
EcoRI-pr E01L (50 ng / pl)
MseI-pr M02K (50ng / ul) dNTP (25mM) pol. Taq. (5U) 10X PCR
Mq
Total
Pl preamp thermal profile
0.6 pl
0.6 pl
0.16 pl
0.08 pl
2.0 pl
11.56 pl pl / reaction
Selective preamplification was performed in a reaction volume of 50 µl. PCR was carried out on a PE GeneAmp 9700 system and a 20 cycle profile that began with a denaturation step at 94 ° C for 30 seconds, followed by a hybridization step at 56 ° C for 60 seconds and a extension stage at 72 ° C for 60 seconds.
ES 2 393 318 T3
EcoRI + 1 (A)<sup>1</sup>
E01 L 92R11: 5-AGACTGCGTACCAATTCA-3 [SEQ ID 9]
Msel +1 (C)<sup>1</sup>
M02k 93E42: 5-GATGAGTCCTGAGTAAC-3 [SEQ ID 10]
<td colspan="2">A / AC Preamp:</td>
<td>Mix PA + 1 / + 1 (20x)</td><td>: 5 μί</td>
<td>EcoRI-pr.</td><td>: 1.5 μl</td>
<td>MseI-pr.</td><td>: 1.5 μl</td>
<td>dNTP (25 mM)</td><td>: 0.4 μl</td>
<td>Pol. Taq. (5 U)</td><td>: 0.2 μl</td>
<td>10X PCR</td><td>: 5 μl</td>
<td>Mq</td><td>: 36.3 μl</td>
<td>Total</td><td>: 50 μl</td>
Selective preamplification was performed in a reaction volume of 50 µΙ. PCR was carried out on a PE GeneAmp 9700 system and a 30 cycle profile that started with a denaturation step at 94 ° C for 30 seconds, followed by a hybridization step at 56 ° C for 60 seconds and a extension stage at 72 ° C for 60 seconds.
(3) KRSEcoRI +1 (A) and KRSMsel + 2 (CA)<sup>2</sup>
05F212 E01LKRS1
05F213 E01LKRS2
05F214 M15KKRS1
05F215 M15KKRS2
CGTCAGACTGCGTACCAATTCA
CAAGAGACTGCGTACCAATTCA
TGGTGATGAGTCCTGAGTAACA
AGCCGATGAGTCCTGAGTAACA
-3 '[SEQ ID 11]
-3 '[SEQ ID 12]
-3 '[SEQ ID 13]
-3 '[SEQ ID 14] selective nucleotides in bold and underlined labels (KRS)
PSP11 sample: E01LKRS1 / M15KKRS1
PI120234 sample: E01LKRS2 / M15KKRS2 (4) AFLP protocol
Selective amplification was performed in a reaction volume of 20 µ !. PCR was carried out on a PE GeneAmp 9700 PCR system. A 13-cycle profile was started with a denaturation step of 94 ° C for 30 seconds, followed by a hybridization step at 65 ° C for 30 seconds, with a reduction step in which the hybridization temperature was reduced by 0 , 7 ° C in each cycle, and an extension stage at 72 ° C for 60 seconds. This profile was followed by a 23 cycle profile with a denaturation step of 94 ° C for 30 seconds, followed by a hybridization step at 56 ° C for 30 seconds and an extension step at 72 ° C for 60 seconds.
EcoRI +3 (AAC) and MseI +3 (CAG)
E32 92S02: 5-GACTGCGTACCAATTCAAC-3 [SEQ ID 15]
M49 92G23: 5-GATGAGTCCTGAGTAACAG-3 [SEQ ID 16] (5) Qiagen column
A Qiagen purification was carried out following the manufacturer's instructions: QIAquick® Spin Manual (http://wwwl.qiagen.com/literature/handbooks/PDF/DNACleanupAndConcentration/QQ Spin / 1021422 HBQQSpin 072002WW.pdf)
Example 2: bell pepper
DNA from pepper lines PSP-11 and PI20234 was used to generate the AFLP product using Keygene's AFLP recognition site-specific primers. (These AFLP primers are essentially the same as conventional AFLP primers, for example those described in EP patent No. 0 534 858, and generally contain a recognition site region, a constant region and one or more selective nucleotides in a selective region). From the pepper lines PSP-11 or PI20234, 150 ng of DNA were digested with the restriction endonucleases EcoRI (5 U / reaction) and MseI (2 U / reaction) for 1 hour at 37 ° C, followed by inactivation for 10 minutes at 80 ° C. The restriction fragments obtained were ligated with a double-stranded synthetic oligonucleotide adapter, one end of which was compatible with one or both ends of the EcoRI and / or MseI restriction fragments. AFLP preamplification reactions (20 µl / reaction) were carried out with the AFLP + 1 / + 1 primers in 10-fold dilution-restriction mixture. PCR profile: 20 * (30 seconds at 94 ° C + 60 seconds at 56 ° C + 120 seconds at 72 ° C). AFLP reactions were carried out
Additional ES 2 393 318 T3 (50 μΙ / reaction) with different Keygene +1 EcoRI and +2 Msel recognition site primers (see Table, below; labels are shown in bold, selective nucleotides are underlined) in product AFLP EcoRI / MseI + 1 / + 1 preamplification diluted 20 times. PCR profile: PcR profile: 30 * (30 seconds at 94 ° C + 60 seconds at 56 ° C + 120 seconds at 72 ° C). The AFLP product was purified using the QIAquick PCR purification kit (QIAGEN) according to the QIAquick® Spin 07/2002 manual, page 18, and the concentration was measured with an ND-1000 NanoDrop® spectrophotometer. A total of 5 µg of AFLP PSP-11 + 1 / + 2 product and 5 µg of AFLP PI20234 + 1 / + 2 product were pooled and resolved in 23.3 µl TE. Finally, a mixture with a concentration of 430 ng ^ l of AFLP + 1 / + 2 product was obtained.
Table
<td>SEQ ID</td><td>PCR primer</td><td>Primer -3 '</td><td>Pepper</td><td>AFLP reaction</td>
<td>[SEQ ID 17]</td><td>05F21</td><td>CGTCAGACTGCGTACCAATTCA</td><td>PSP</td><td> 1</td>
<td>[SEQ ID 18]</td><td>05F21</td><td>TGGTGATGAGTCCTGAGTAACA</td><td>PSP</td><td> 1</td>
<td>[SEQ ID 19]</td><td>05F21</td><td>CAAGAGACTGCGTACCAATTCA</td><td>PI2023</td><td> 2</td>
<td>[SEQ ID 20]</td><td>05F21</td><td>AGCCGATGAGTCCTGAGTAACA</td><td>PI2023</td><td> 2</td>
Example 3: corn
DNA from maize lines B73 and M017 was used to generate the AFLP product using Keygene AFLP recognition site specific primers. (These AFLP primers are essentially the same as conventional AFLP primers, for example those described in EP 0 534 858, and generally contain a recognition site region, a constant region and one or more selective nucleotides in the 3 'end thereof). DNA from pepper lines B73 or M017 was digested with restriction endonucleases TaqI (5 U / reaction) for 1 hour at 65 ° C and MseI (2 U / reaction) for 1 hour at 37 ° C, followed by inactivation for 10 minutes at 80 ° C. The restriction fragments obtained were ligated with a double-stranded synthetic oligonucleotide adapter, one end of which is compatible with one or both ends of the TaqI and / or MseI restriction fragments. AFLP preamplification reactions (20 µl / reaction) were carried out with the AFLP + 1 / + 1 primers in 10-fold diluted restriction-ligation mixture. PCR profile: 20 * (30 seconds at 94 ° C + 60 seconds at 56 ° C + 120 seconds at 72 ° C). Additional AFLP reactions (50 μl / reaction) were carried out with different Keygene recognition site primers for FLP TaqI and MseI +2 (Table below; labels are shown in bold, selective nucleotides are underlined) in product of AFLP TaqI / MseI + 1 / + 1 diluted 20 times. PCR profile: PCR profile: 30 * (30 seconds at 94 ° C + 60 seconds at 56 ° C + 120 seconds at 72 ° C). The AFLP product was purified using the QIAquick PCR purification kit (QIAGEN) according to the QIAquick® Spin 07/2002 manual, page 18, and the concentration was measured with an ND-1000 NanoDrop® spectrophotometer. A total of 1.25 µg of each different AFLP + 2 / + 2 B73 product and 1.25 µg of each different AFLP + 2 / + 2 M017 product were pooled and resolved in 30 µl TE. Finally, a mixture with a concentration of 333 ng ^ l of AFLP + 2 / + 2 product was obtained.
Table
<td>SEQ ID</td><td>PCR primer</td><td>Primer sequence</td><td>Corn</td><td>Reaction of AFLP</td>
<td>[SEQ ID 21]</td><td>05G360</td><td>ACGTGTAGACTGCGTACCGAAA</td><td>B73</td><td> 1</td>
<td>[SEQ ID 22]</td><td>05G368</td><td>ACGTGATGAGTCCTGAGTAACA</td><td>B73</td><td> 1</td>
<td>[SEQ ID 23]</td><td>05G362</td><td>CGTAGTAGACTGCGTACCGAAC</td><td>B73</td><td> 2</td>
<td>[SEQ ID 24]</td><td>05G370</td><td>CGTAGATGAGTCCTGAGTAACA</td><td>B73</td><td> 2</td>
<td>[SEQ ID 25]</td><td>05G364</td><td>GTACGTAGACTGCGTACCGAAG</td><td>B73</td><td> 3</td>
<td>[SEQ ID 26]</td><td>05G372</td><td>GTACGATGAGTCCTGAGTAACA</td><td>B73</td><td> 3</td>
<td>[SEQ ID 27]</td><td>05G366</td><td>TACGGTAGACTGCGTACCGAAT</td><td>B73</td><td> 4</td>
<td>[SEQ ID 28]</td><td>05G374</td><td>TACGGATGAGTCCTGAGTAACA</td><td>B73</td><td> 4</td>
<td>[SEQ ID 29]</td><td>05G361</td><td>AGTCGTAGACTGCGTACCGAAA</td><td>M017</td><td> 5</td>
<td>[SEQ ID 30]</td><td>05G369</td><td>AGTCGATGAGTCCTGAGTAACA</td><td>M017</td><td> 5</td>
<td>[SEQ ID 31]</td><td>05G363</td><td>CATGGTAGACTGCGTACCGAAC</td><td>M017</td><td> 6</td>
<td>[SEQ ID 32]</td><td>05G371</td><td>CATGGATGAGTCCTGAGTAACA</td><td>M017</td><td> 6</td>
ES 2 393 318 T3
<td>[SEQ ID 33]</td><td>05G365</td><td>GAGCGTAGACTGCGTACCGAAG</td><td>M017</td><td> 7</td>
<td>[SEQ ID 34]</td><td>05G373</td><td>GAGCGATGAGTCCTGAGTAACA</td><td>M017</td><td> 7</td>
<td>[SEQ ID 35]</td><td>05G367</td><td>TGATGTAGACTGCGTACCGAAT</td><td>M017</td><td> 8</td>
<td>[SEQ ID 36]</td><td>05G375</td><td>TGATGATGAGTCCTGAGTAACA</td><td>M017</td><td> 8</td>
Finally, the 4 P1 samples and the 4 P2 samples were pooled and concentrated. A total amount of 25 µl of DNA product was obtained and a final concentration of 400 ng ^ l (total of 10 µg). Intermediary quality assessments are provided in Figure 3.
SEQUENCING BY 454
AFLP fragment samples from pepper and corn as described above were processed by 454 Life Sciences as described (Margulies et al., Genome sequencing in microfabricated high-density picoliter reactors, Nature 435 (7057): 376-80, published electronically July 31, 2005).
DATA PROCESSING
Processing line:
Input data
Raw sequence data was received for each analysis:
- 200,000 to 400,000 sequence results
- automatic nucleotide reading quality scores
Cropping and labeling
These sequence data were analyzed for the presence of Keygene recognition sites (KRS) at the beginning and end of the reading. These KRS sequences consist of AFLP adapter and sample tag sequences and are specific to a given combination of AFLP primers in a given sample. The KRS sequences were identified by BLAST and trimmed, and the restriction sites were restored. The readings were marked with a label for identification of the origin of kRs. Clipped sequences were selected from length (33 nt minimum) to participate in post-processing.
Grouping and assembling
A MegaBlast analysis of all trimmed and size-selected reads was performed to obtain clusters of homologous sequences. Consecutively all the groups were assembled with CAP3, resulting in assembled contigs. After both steps, single sequence reads had been identified that did not correspond to any other readings. These readings are designated as singletons. The processing line followed to carry out the steps described herein is shown in Figure 4A.
Polymorphism exploration and quality assessment
The contigs resulting from the assembly analysis form the basis for the detection of polymorphisms. Each mismatch in the alignment of each cluster is a potential polymorphism. Selection criteria were defined to obtain a quality score:
- number of reads in each contig
- allele frequency in each sample
- appearance of homopolymer sequence
- appearance of contiguous polymorphisms
SNPs and indels with a quality score above the threshold are identified as putative polymorphisms. The MISA (microsatellite identification) tool (http://pgrc.ipkgatersleben.de/misa) was used for the exploration of the SSRs. This tool identifies the dinucleotide, trinucleotide, tetranucleotide and SSR motifs of the compound using predefined criteria and summarizes the occurrences of these SSRs. The polymorphism screening and quality assignment procedure is shown in Figure 4B.
ES 2 393 318 T3
RESULTS
The Table below summarizes the results of the pooled sequence analysis obtained from 2 sequencing runs of 454 for the pooled pepper samples and 2 runs for the pooled corn samples.
<td></td><td>Pepper</td><td>Corn</td>
<td>Total number of sequence results</td><td> 457178</td><td> 492145</td>
<td>Number of trimmed sequences</td><td> 399623</td><td> 411008</td>
<td>Number of singletones</td><td> 105253</td><td> 313280</td>
<td>Number of contigs</td><td> 31863</td><td> 14588</td>
<td>Number of sequences in contigs</td><td> 294370</td><td> 97728</td>
<td>Total number of sequences containing SSR</td><td> 611</td><td> 202</td>
<td>Number of different sequences containing SSR</td><td> 104</td><td> 65</td>
<td>Number of different SSR motifs (di, tri, tetra and compound)</td><td> 49</td><td> 40</td>
<td>Number of SNPs with a Q score> 0.3 *</td><td> 1636</td><td> 782</td>
<td>Number of indels *</td><td> 4090</td><td> 943</td>
<td colspan="3">* both with selection against neighboring SNPs, flanking sequence of at least 12 bp and not present in homopolymer sequences greater than 3 nucleotides.</td>
Example 4. Identification of nucleotide polymorphisms (SNP) in pepper
DNA isolation
The genomic DNA of the two parental lines was isolated from a recombinant inbred population (RIL) of pepper and 10 descendants of RIL. The parental lines were PSP11 and PI201234. Genomic DNA was isolated from leaf material of individual seedlings using a modified CTAB procedure described by Stuart and Via (Stuart CN Jr. and Via LE, A rapid CTAB DNA isolation technique useful for RAPD fingerprinting and other PCR applications, Biotechniques, 14, 748-750, 1993). DNA samples were diluted to a concentration of 100 ng / µl in TE (10 mM Tris-HCl, pH 8.0, 1 mM EDTA) and stored at -20 ° C.
Template Preparation for AFLP Using Labeled AFLP Primers
AFLP templates were prepared from the parental pepper lines PSP11 and PI201234 using the combination of restriction endonucleases EcoRI / Msel as described in Zabeau and Vos, Selective restriction fragment amplification; a general method for DNA fingerprinting, EP patent No. 0534858-A1, B1, 1993; US Patent No. 6,045,994 and in Vos et al. (Vos P., Hogers R., Bleeker M., Reijans M., van de Lee T., Hornes M., Frijters A., Pot J., Peleman J., Kuiper M. et al., AFLP: a new technique for DNA fingerprinting, Nucleic Acids Research 23 (21): 44074414, 1995).
Specifically, restriction of genomic DNA with EcoRI and MseI was carried out as follows:
DNA restriction
DNA
EcoRI
MseI
5x MilliQ Water RL Buffer up to
100 at 500 ng units pl units pl
The incubation was carried out for 1 hour at 37 ° C. After enzyme restriction, the enzymes were inactivated by incubation for 10 minutes at 80 ° C.
Adapter ligation
ATP 10 mM 1pl
1pl T4 DNA ligase
EcoRI adapter (50 pmol / pl) 1pl
MseI adapter (5 pmol / pl) 1pl
5x RL buffer. 2pl
MilliQ water up to 40 pl
ES 2 393 318 T3
Incubation was carried out for 3 hours at 37 ° C.
Selective amplification of AFLP
After restriction-ligation, the restriction / ligation reaction was diluted 10-fold with ThioEo, and 5 µΙ of diluted mixture was used as a template in a selective amplification step. Note that because a + 1 / + 2 selective amplification was intended, a + 1 / + 1 selective preamplification step (with standard AFLP primers) was performed first. The reaction conditions of the + 1 / + 1 (+ A / + C) amplification were as follows.
Restriction-ligation mixture (diluted 10 times) 5 μl
Primer EcoRI +1 (50 ng / μΙ) 0.6 μΙ
Primer Msel +1 (50 ng / μΙ) 0.6 μΙ dNTP (20 mM) 0.2 μΙ
Taq polymerase (5 U / μΙ Amplitaq, PE) 0.08 μΙ
10χ 2.0 μΙ PCR buffer
MilliQ water up to 20 μΙ
The primer sequences were:
EcoRI + 1: 5'- AGACTGCGTACCAATTCA -3 '[SEQ ID 9] and
Msel + 1: 5'- GATGAGTCCTGAGTAAC -3 '[SEQ ID 10]
PCR amplifications were carried out using a PE9700 with a gold or silver block using the following conditions: 20 times (30 seconds at 94 ° C + 60 seconds at 56 ° C + 120 seconds at 72 ° C).
The quality of the + 1 / + 1 preamplification products generated was checked on a 1% agarose gel using a 100 base pair ladder and a 1 Kb ladder to check the distribution of fragment lengths. After + 1 / + 1 selective amplification, the reaction was diluted 20-fold with ThioEo, i and 5 µΙ of diluted mixture was used as template in the + 1 / + 2 selective amplification step using labeled AFLP primers.
Finally, selective amplifications were carried out by AFLP + 1 / + 2 (A / + CA): selective amplification product + 1 / + 1 (diluted 20 times): 5.0 μΙ.
KRS EcoRI-primer + A (50 ng / μΙ) 1.5 μΙ
KRS Msel- primer + CA (50 ng / μΙ) 1.5 μΙ dNTP (20 mM) 0.5 μΙ Taq polymerase (5 U / μΙ Amplitaq, Perkin Elmer) 0.2 μΙ
10X 5.0 μΙ PCR buffer
MQ up to 50 μΙ
The sequences of the tagged AFLP primers were:
PSP11:
05F212: EcoRI + 1: 5'-CGTCAGACTGCGTACCAATTCA-3 '[SEQ ID 1] and
05F214: Msel + 2: 5'-TGGTGATGAGTCCTGAGTAACA-3 '[SEQ ID 2]
PI201234:
05F213: EcoRI + 1: 5'-CAAGAGACTGCGTACCAATTCA-3 '[SEQ ID 3] and
05F215: Msel + 1: 5'-AGCCGATGAGTCCTGAGTAACA-3 '[SEQ ID 4]
Note that these primers contain 4 bp tags (underlined above) at their 5 prime ends to distinguish the amplification products originating from the respective pepper lines at the end of the sequencing procedure. Schematic representation of AFLP + 1 / + 2 pepper amplification products after amplification with AFLP primers containing 4 bp 5 prime tag sequences.
EcoRI Label Msel Label
PSP II; 5<sup>1</sup> -CGTC ------------------------------------— ACCA-3 '
3 '^ GCAG ----—---------------------------------- TGGT-5<sup>1</sup>
PI201234 5 <sup>T</sup>-CAñG ----------------------------------- GGCT-3 '* -GTTC ------ ----------------------------- CCGA-5 '
PCR amplifications (24 per sample) were carried out using a PE7900 with a gold or silver block under the following conditions: 30 times (30 seconds at 94 ° C + 60 seconds at 56 ° C + 120 seconds at 72 ° C).
ES 2 393 318 T3
The quality of the amplification products generated was checked on a 1% agarose gel using a 100 base pair ladder and a 1 Kb ladder to check the distribution of fragment lengths.
Purification and quantification of the AFLP reaction
After pooling two + 1 / + 2 selective AFLP reactions of 50 microliters for each pepper sample, the resulting 12 products from 100 μl AFLP reaction were purified using the QIAquick PCR purification kit (QIAGEN) following the QIAquick manual. ® Spin (page 18). A maximum of 100 µl of product was loaded on each column. Amplified products eluted at T10E0.1. The quality of the purified products was checked on a 1% agarose gel and the concentrations were measured on the NanoDrop (Figure 2).
NanoDrop concentration measurements were used to adjust the final concentration of each purified PCR product to 300 nanograms per microliter. Five micrograms of purified amplified product of PSP11 and 5 micrograms of PI201234 were mixed to generate 10 micrograms of template material for the preparation of the 454 sequencing library.
Sequence library preparation and high-throughput sequencing
Mixed amplification products from both pepper lines were subjected to high-throughput sequencing using 454 Life Sciences sequencing technology, as described in Margulies et al. (Margulies et al., Nature 437: 376-380 and Online Supplements). Specifically, the ends of the AFLP PCR products were first polished and then ligated with adapters to facilitate emulsion PCR amplification and subsequent sequencing of the fragments as described by Margulies et al.
The 454 adapter sequences, emulsion PCR primers, sequence primers, and sequencing operating conditions were all indicated by Margulies et al. The linear order of functional elements in an emulsion PCR fragment amplified on Sepharose beads in the 454 sequencing procedure was as follows, as exemplified in Figure 1A:
454 PCR adapter - 454 sequence adapter - AFLP primer 4 bp tag 1 - AFLP primer sequence 1 including the selective nucleotide (s) - AFLP fragment internal sequence - AFLP primer sequence 2 that includes one or more selective nucleotides, 4 bp tag 2 of AFLP primers - 454 sequence adapter - 454 PCR adapter - Sepharose bead.
Two high throughput 454 sequencing runs were performed by 454 Life Sciences (Branford, CT; United States).
454 Sequencing Operation Data Processing
Sequence data resulting from 2 sequencing runs of 454 were processed using a bioinformatics procedure (Keygene NV). Specifically, base-mapped raw sequence reads obtained from 454 were converted to FASTA format and inspected for the presence of tagged AFLP adapter sequences using a BLAST algorithm. Following high confidence matches to known tagged AFLP primer sequences, the sequences were trimmed, restriction endonuclease sites restored, and appropriate tags assigned (Sample 1 EcoRI (ES1), Sample 1 MseI (MS1), sample 2 EcoRI (ES2) or sample 2 MseI (MS2), respectively). Next, all trimmed sequences greater than 33 bases were pooled using a megaBLAST procedure based on overall sequence homologies. Clusters were then assembled into one or more contigs and / or singletons per cluster using a multiple alignment CAP3 algorithm. Contigs containing more than one sequence were inspected for sequence mismatches, representative of putative polymorphisms. Sequence mismatches were assigned quality scores based on the following criteria:
* number of reads in a contig * the observed distribution of alleles
The two criteria listed above form the basis for the so-called Q score assigned to each putative SNP / indel. Q scores are between 0 and 1; a Q score of 0.3 can only be achieved if both alleles are observed at least twice.
* localization in homopolymers of a certain length (adjustable; default value to avoid polymorphisms localized in homopolymers of 3 bases or longer).
* number of contigs in a pool.
* distance to closest neighboring sequence mismatches (adjustable; important for certain types of genotyping assays probing flanking sequences) * level of association of observed alleles with sample 1 or sample 2; in the case of a consistent perfect association between the alleles of a putative polymorphism and samples 1 and 2, the polymorphism (SNP) is indicated as putative
ES 2 393 318 T3 elite polymorphism (SNP). An elite polymorphism is considered to have a high probability of being located in a single or low copy number genomic sequence, if two homozygous lines have been used in the screening procedure. Conversely, a weak association of a polymorphism with the origin of the sample presents a high risk that false polymorphisms arising from the alignment of non-allelic sequences in a contig have been discovered.
Sequences containing SSR motifs were identified using the MISA search tool (microsatellite identification tool; available at http://pgrc.ipk-gatersleben.de/misa/).
The overall statistics of the operation are shown in the Table below.
Table. Global statistics from a 454 sequencing analysis for the identification of SNPs in pepper.
<td>Enzyme combination</td><td>Analysis</td>
<td>Cutout</td><td></td>
<td>All sequences</td><td> 254.308</td>
<td>Wrong</td><td> 5.293 (2 %)</td>
<td>Correct</td><td> 249.015 (98%)</td>
<td>Concatamers</td><td> 2.156 (8,5 %)</td>
<td>Mixed labels</td><td> 1.120 (0,4 %)</td>
<td>Correct sequences</td><td></td>
<td>One end cropped</td><td> 240.817 (97%)</td>
<td>Both ends trimmed</td><td> 8.198 (3 %)</td>
<td>Number of sample sequences 1</td><td> 136.990 (55%)</td>
<td>Number of sample sequences 2</td><td> 112.025 (45 %)</td>
<td>Grouped</td><td></td>
<td>Number of contigs</td><td> 21.918</td>
<td>Sequences in contigs</td><td> 190.861</td>
<td>Average number of sequences in each contig</td><td> 8,7</td>
<td>SNP scan</td><td></td>
<td>SNP with Q score> 0.3 *</td><td> 1.483</td>
<td>Indel with Q score> 0.3 *</td><td> 3.300</td>
<td>SNP scan</td><td></td>
<td>SSR scan</td><td></td>
<td>Total number of SSR motifs identified</td><td> 359</td>
<td>Number of sequences containing one or more SSR motifs</td><td> 353</td>
<td>Number of SSR motifs with unit size 1</td><td> 0</td>
<td>(homopolymer)</td><td></td>
<td>Number of SSR motifs with unit size 2</td><td> 102</td>
<td>Number of SSR motifs with unit size 3</td><td> 240</td>
<td>Number of SSR motifs with unit size 4</td><td> 17</td>
<td colspan="2">* SNP / indel screening criteria were as follows:</td>
No contiguous polymorphisms were found with a Q score greater than 0.1 at less than 12 bases on each side, not present in homopolymers of 3 or more bases. The screening criteria did not consider the consistent association with samples 1 and 2, that is, SNPs and indels are not necessarily elite putative SNP / indels.
An example of a multiple alignment containing a putative elite single nucleotide polymorphism is shown in Figure 7.
Example 5. Validation of SNPs by PCR amplification and Sanger sequencing
In order to validate the putative SNP A / G identified in Example 1, a tagged sequence site (STS) assay was designed for this SNP using flanking PCR primers. The sequences of the PCR primers were as follows:
Primer_1.2f: 5'- AAACCCAAACTCCCCCAATC-3 ', [SEQ ID 37] and Primer_1.2r: 5'- AGCGGATAACAATTTCACACAGGACATCAGTAGTCACACTGGTA CAAAAATAGAGCAAAACAGTAGTG -3' [SEQ ID 38]
Note that primer 1.2r contained an M13 sequence primer binding site and a filler fragment at its prime 5 end. PCR amplification was carried out using the AFLP + A / + CA amplification products of PSP11 and PI210234 prepared as described in Example 4 as a template. The conditions
ES 2 393 318 T3 PCR were as follows: for 1 PCR reaction the following components were mixed: 5 μl AFLP mix diluted 1/10 (approx. 10 ng / μl) μl 1 pmol / μΙ of primer 1,2f ( diluted directly from 500 μΜ stock solution) pl 1 pmol / μΙ 1.2r primer (diluted directly from 500 μΜ stock solution)
<td>5 pl PCR mix</td><td>-2 pl 10x PCR buffer -1 pl 5 mM dNTP -1.5 pl Mgcl<sub>2</sub>25 mM -0.5 pl of H<sub>2</sub>OR</td>
<td>5 pl of enzyme mix</td><td>-0.5 pl 10x PCR buffer (Applied Biosystems) -0.1 pl 5 U / μΙ AmpliTaq DNA polymerase (Applied Biosystems) -4.4 pl H<sub>2</sub>OR</td>
The following PCR profile was used:
<td>Cycle 1 Cycle</td><td>2 ': 94 ° C 2-34 20: 94 ° C 30: 56 ° C 2'30: 72 ° C</td>
<td>Cycle 35</td><td>7 ': 72 ° C ~: 4 ° C</td>
The PCR products were cloned into the vector pCR2.1 (TA cloning kit, Invitrogen) using the TA cloning method and transformed into INVaF 'competent E. coli cells. Transformants were subjected to blue / white screening. Three independent target transformants were selected from each of PSP11 and PI-201234 and cultured O / N in liquid selective medium for plasmid isolation.
Plasmids were isolated using the QIAprep Spin miniprep kit (QIAGEN). The insertions of these plasmids were then sequenced following the protocol indicated below and resolved on the MegaBACE 1000 (Amersham). The sequences obtained were inspected in the presence of the SNP allele. Two independent plasmids containing the PI-201234 insert and 1 plasmid containing the PSP11 insert contained the expected consensus sequence flanking the SNP. The sequence derived from the PSP11 fragment contained the expected A allele (underlined) and the sequence derived from the PI-201234 fragment contained the expected G allele (double underlined):
PSP11 (sequential): (5'-3 ')
AAACCCAAACTCCCC CAATCGATTTCAAACCTAGAACAAT GTTGGTT TTGGTGC TAACTTCAA
CCCCACTACT GTT TTGCTCTAT TTTTGT [SECID 39]
PI-201234 (sequence 1): (5'-3 ')
AAACCCAAACTCC CCCAATCGATTT CAAACC TAGAACAGTGT TGGTT TTGGTGCTAACTTCAA cc ceac tactgiτ ττgctctat ttttg [SEQ ID 40]
PI-201234 (sequence 2): (5'-3 ')
AAACCCAAACTCC CCC AATCG AT TTCAAACCT AGAACAgT GT TGGTTT TGGTGCTAACTT CAA
CCCCACTACTGTTTTGCTCTATTTTTG [SEQ ID 41]
This result indicates that the putative bell pepper SNP A / G represents a true genetic polymorphism detectable using the designed STS assay.
Example 6: Validating SNPs by Discovery with SNPWave
In order to validate the putative SNP A / G identified in Example 1, sets of SNPWave ligation probes were defined for both alleles of said SNP using the consensus sequence. The sequences of the ligation probes were as follows:
SNPWave probe sequences (5'-3 '):
06A162 GATGAGTCCTGAGTAACCCAATCGATTTCAAACCTAGAACAA (42 bases) [SEQ ID 42]
06A163 GATGAGTCCTGAGTAACCACCAATCGATTTCAAACCTAGAACAG (44 bases) [SEQ ID 43] 06A164
<img file="ES2393318T3_D0001.tif" />
ES 2 393 318 T3
Phosphate-TGTTGGTTTTGGTGCTAACTTCAACCAACATCTGGAATTGGTACGCAGTC (52 bases) [SEQ ID 44]
Note that the 06A162 and 06A163 allele-specific probes for alleles A and G, respectively, differ in size by 2 bases, such that, after ligation to the locus-specific common probe 06A164, ligation product sizes of 94 result (42 + 54) and 96 (44 + 52) bases.
The SNPWave and PCR ligation reactions were carried out as described by Van Eijk et al. (MJT van Eijk, JLN Broekhof, HJA van der Poel, RCJ Hogers, H. Schneiders, J. Kamerbeek, E. Verstege, JW van Aart , H. Geerlings, JB Buntjer, AJ van Oeveren and P. Vos, SNPWaveTM: SNPWave ™ - A flexible multiplex SNP genotyping technology. Nucleic Acids Research 32: e47) using 100 ng of genomic DNA from pepper lines PSP11 and PI201234 and 8 RIL descendants as starting material. The sequences of the PCR primers were:
93L01FAM (EOOk): 5-GACTGCGTACCAATTC-3 '[SEQ ID 45]
93E40 (MOOk): 5-GATGAGTCCTGAGTAA-3 '[SEQ ID 46]
After PCR amplification, the purification and detection of the PCR product on the MegaBACE1000 was as described by van Eijk et al. (See above). Figure 8B shows a gel pseudo-image of the amplification products obtained from PSP11, PI201234 and 8 RIL descendants.
The SNPWave results clearly demonstrate that SNP A / G is detected by the SNPWave assay, resulting in 92 bp products (= homozygous AA genotype) for P1 (PSP11) and RIL descendants 1, 2, 3, 4, 6 and 7) and in 94 bp products (GG homozygous genotype) for P2 (PI201233) and the RIL 5 and 8 descendants.
Example 7: Strategies for the enrichment of AFLP fragment libraries in low copy number sequences.
The present example describes several enrichment methods focused on single or low copy number genomic sequences in order to increase the yield of elite polymorphisms as described in Example 4. The methods can be classified into four categories:
1) Methods aimed at preparing high quality genomic DNA, excluding chloroplast sequences.
It is proposed to prepare nuclear DNA instead of total genomic DNA as described in Example 4, to exclude co-isolation of abundant chloroplast DNA, which could result in a reduced number of plant genomic DNA sequences, depending on the Restriction endonucleases and selective AFLP primers used during the fragment library preparation procedure. A protocol for the isolation of nuclear DNA from highly pure tomato has been described by Peterson DG, Boehm KS and Stack SM, Isolation of Milligram Quantities of Nuclear DNA From Tomato (Lycopersicon esculentum), A Plant Containing High Levels of Polyphenolic Compounds. Plant Molecular Biology Reporter 15 (2): 148-153, 1997.
2) Methods intended to use restriction endonucleases during the AFLP template preparation procedure that are expected to yield high levels of low copy number sequences.
It is proposed to use certain restriction endonucleases during the AFLP template preparation procedure, which are expected to target single or low copy number genomic sequences, resulting in polymorphism-enriched fragment libraries with an increased ability to be convertible into genotyping assays. An example of a restriction endonuclease targeting a low copy number sequence in plant genomes is PstI. Other methylation-sensitive restriction endonucleases may also preferentially target low copy number or unique genomic sequences.
3) Methods aimed at selectively removing highly duplicated sequences based on re-pairing kinetics of repeated sequences versus low copy number sequences.
It is proposed to selectively remove highly duplicated (repeat) sequences from the total genomic DNA sample or AFLP template material (cDNA) prior to selective amplification.
3a) High Cüt DNA preparation is a commonly used technique for enrichment for slow pairing low copy number sequences from a complex mixture of plant genomic DNA (Yuan et al., High-Cot sequence analysis of the maize genome. Plant J. 34: 249-255, 2003). It is suggested to use high C0t DNA instead of total genomic DNA for enrichment for polymorphisms located in low copy number sequences.
3b) As an alternative to the laborious high-C0t preparation, re-pairing denatured dcDNA can be incubated with a new Kamchatka crab nuclease, which cuts perfectly matching short DNA duplexes at a higher rate than not perfectly matching DNA duplexes ,
ES 2 393 318 T3 as described by Zhulidov et al (Simple cDNA normalization using Kamchatka crab duplex-specific nuclease. Nucleic Acids Research 32: e37, 2004) and Shagin et al. (A novel method for sNp detection using a new duplex-specific nuclease from crab hepatopancreas, Genome Research 12: 1935-1942, 2004). Specifically, it is proposed to incubate the AFLP restriction / ligation mixtures with said endonuclease to deplete the mixture in highly duplicated sequences, followed by selective AFLP amplification of the remaining low copy number or unique genomic sequences.
3c) Methyl filtration is a method to enrich for hypomethylated genomic DNA fragments using the restriction endonuclease McrBC, which cuts methylated DNA in the sequence [A / G] C, in which C is methylated (see Pablo D Rabinowicz, Robert Citek, Muhammad A. Budiman, Andrew Nunberg, Joseph A. Bedell, Nathan Lakey, Andrew L. O'Shaughnessy, Lidia U. Nascimento, W. Richard McCombie, and Robert A. Martienssen, Differential methylation of genes and repeats in land plants, Genome Research 15: 14311440, 2005). McrBC can be used to enrich for the fraction of low copy number sequences of a genome, which will be used as the starting material for screening for polymorphisms.
4) Use of cDNA and not genomic DNA for the recognition of gene sequences.
Finally, it is proposed to use oligo-dT primed cDNA and not genomic DNA as starting material for polymorphism screening, optionally in combination with the use of crab duplex specific nuclease indicated in 3b, above, for normalization. Note that the use of oligo-dT primed cDNA also excludes chloroplast sequences. Alternatively, AFLP-cDNA templates are used in place of oligo-dT primed cDNA to facilitate amplification of the remaining low copy number sequences analogously to AFLP (see also 3b, above).
Example 8: Strategy for the enrichment in repeats of simple sequences.
The present example describes the proposed strategy of single sequence repeat sequence discovery analogously to the SNP identification described in Example 4.
Specifically, restriction-ligation of genomic DNA from two or more samples is carried out, for example using PstI / MseI restriction endonucleases. Selective AFLP amplification is carried out as described in Example 4. It is then enriched for fragments containing the selected SSR motifs by one of the following two methods:
1) Southern blot hybridization on filters containing oligonucleotides corresponding to the desired SSR motifs (for example (CA) 15 in the case of enrichment for CA / GT repeats), followed by amplification of the joined fragments in a manner similar to that described by Armor et al. (Armor J., Sismani C., Patsalis P. and Cross G., Measurement of locus copy number by hybridization with amplifiable probes. Nucleic Acids Research 28 (2): 605-609) or by 2) enrichment using biotinylated capture oligonucleotide hybridization probes to capture fragments (from AFLP) in solution as described by Kijas et al. (Kijas JM, Fowler JC, Garbett CA and Thomas, MR, Enrichment of microsatellites from the citrus genome using biotinylated oligonucleotide sequences bound to streptavidin-coated magnetic particles. Biotechniques, 16: 656-662, 1994).
The SSR motif-enriched AFLP fragments are then amplified using the same AFLP primers used in the preamplification step, in order to generate a sequence library. An aliquot of the amplified fragments is cloned T / A and 96 clones are sequenced to estimate the fraction of positive clones (clones containing the desired SSR motif, eg cA / GT motifs of more than 5 repeat units. Another aliquot of the AFLP fragment-enriched mixture is detected by polyacrylamide gel electrophoresis (PAGE), optionally after further selective amplification to obtain a readable fingerprint, in order to visually inspect whether enrichment has been performed on fragments that contain SSR. Upon successful completion of these control steps, the sequence libraries are subjected to high throughput 454 sequencing.
The above strategy for de novo identification of SSRs is schematically illustrated in Figure 8A, and can be adapted for other sequence motifs by corresponding substitution of the capture oligonucleotide sequences.
Example 9. Strategy to avoid mixed labels.
Mixed tags refer to the observation that apart from the expected combination of AFLP-tagged primers in each sample, a reduced fraction of sequences is observed containing a sample 1 tag at one end and a sample 2 tag. at the other extreme (see also Table 1 in Example 4). Schematically the configuration of sequences containing mixed tags is illustrated below.
ES 2 393 318 T3
Schematic representation of the expected combinations of sample labels.
EcoRI label Mse label |
PSP 11; 5<sup>1</sup> -CGTC ------------------------------------------- ACCA-3 '* -GCAG --------------------------------------------- TGGT-5 '
PI-2O1234 5'-CAAG --------------------------------------- GGCT-3<sup>1</sup>
31-GTTC -----—---------------------------—----- CCGA-5<sup>1</sup>
Schematic representation of mixed labels.
EcoRl Label Msel Label
5<sup>r</sup>-CGTC -------------------------------—------— GGCT-3 '
3'-GCAG —------------------------------------- CCGA-5 '
5<sup>r</sup>-CAAG ------------------------------------------ ACCA-3 *
------ ----- m ------------— ------ „TGGT-5 '
The observation of mixed labels prevents the correct assignment of the sequences to PSP11 or PI-201234.
An example of a mixed tag sequence observed in the pepper sequence analysis described in Example 4 is shown in Figure 5A. An overview of the observed fragment configuration is shown in panel 2 of Figure 5A. they contain expected tags and mixed tags.
The proposed molecular explanation for mixed tags is that during the sequence library preparation step, T4 DNA polymerase or Klenow enzyme generate blunt ends in DNA fragments by removing the 3 prime protruding ends prior to ligation of adapters (Margulies et al., 2005). Although the above may work well in the case where a single DNA sample is processed, when processing a mixture of two or more differently labeled DNA samples, refilling by the polymerase results in the incorporation of an incorrect label sequence. in case a heteroduplex has formed between complementary strands derived from different samples (Figure 5B, panel 3, mixed labels). The solution found is to pool the samples after the purification step after ligation of adapters during the 454 fragment library construction step, as shown in Figure 5C, panel 4.
Example 10. Strategy to avoid mixed tags and concatamers using an improved 454 sequence library preparation design.
Apart from the observation of low frequencies of sequence results containing mixed tags as described in Example 9, a low frequency of concatenated AFLP fragment sequence results has been observed.
An example of a concatemer derived sequence result is shown in Figure 6A, panel 1. Schematically, the configuration of sequences containing expected tags and concatamers is shown in Figure 6A, panel 2.
The proposed molecular explanation for the presence of concatenated AFLP fragments is that, during the 454 sequence library preparation step, blunt ends are generated in the DNA fragments by removing the T4 DNA polymerase or the Klenow enzyme protuberant 3 prime prior to adapter ligation (Margulies et al., 2005). Consequently, the blunt-ended DNA fragments in the sample compete with the adapters during the ligation step and may bind to each other before binding to the adapters. This phenomenon is in fact independent of whether a single DNA sample or a mixture of multiple (labeled) samples is included in the library preparation step, and therefore could also occur during conventional sequencing as described by Margulies et al. . In the case where samples with multiple labels are used, as described in Example 4, concatamers complicate the correct assignment of sequence reads to samples based on label information and should therefore be avoided.
The proposed solution to the formation of concatamers (and mixed labels) is to replace the ligation of blunt-ended adapters with the ligation of adapters containing a 3-prime protruding end of T, analogously
ES 2 393 318 T3 to T / A cloning of PCR products, as shown in Figure 6B, panel 3. Conveniently, it is proposed that these modified adapters containing a T at the 3 'overhang end contain a C at the opposite 3 'overhang (which will not bind to the sample DNA fragment, to avoid the formation of concatemers between blunt ends of adapter sequences (see Figure 6B, panel 3)). The resulting adapted flow of operations for the sequence library construction procedure using the modified adapters approach is shown schematically in Figure 6C, panel 4.
Contents27
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
70 members in 11 offices
Priority claims15
| Document | Office | Kind | Date |
|---|---|---|---|
| 693053P | United States of America | – | |
| 69305305 | United States of America | P | |
| 69305305 | United States of America | P | |
| 06075104 | European Patent Office (EPO) | A | |
| 06075104 | European Patent Office (EPO) | A | |
| 06075104 | European Patent Office (EPO) | – | |
| 759034P | United States of America | – | |
| 75903406 | United States of America | P | |
| 75903406 | United States of America | P | |
| 06075104 | – | – | – |
| 693053P | – | – | – |
| 759034P | – | – | – |
| EP20060075104 | – | – | – |
| US20050693053P | – | – | – |
| US20060759034P | – | – | – |
Members70
| Document | Office | Kind | |
|---|---|---|---|
| AU2006259990A1 | Australia | A1 | |
| CA2613248A1 | Canada | A1 | |
| WO2006137733A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2006137734A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1910562A1 | European Patent Office (EPO) | A1 | |
| EP1910563A1 | European Patent Office (EPO) | A1 | |
| CN101278058A | China | A | |
| JP2008546404A | Japan | A | |
| JP2008546405A | Japan | A | |
| US2009036323A1 | United States of America | A1 | |
| US2009142758A1 | United States of America | A1 | |
| CN101641449A | China | A | |
| EP1910563B1 | European Patent Office (EPO) | B1 | |
| AT465274T | Austria | T | |
| ATE465274T1 | Austria | T1 | |
| DE602006013831D1 | Germany | D1 | |
| ES2344802T3 | Spain | T3 | |
| EP1910562B1 | European Patent Office (EPO) | B1 | |
| AT491045T | Austria | T | |
| ATE491045T1 | Austria | T1 | |
| DE602006018744D1 | Germany | D1 | |
| AU2006259990B2 | Australia | B2 | |
| EP2292788A1 | European Patent Office (EPO) | A1 | |
| DK1910562T3 | Denmark | T3 | |
| EP2302070A2 | European Patent Office (EPO) | A2 | |
| ES2357549T3 | Spain | T3 | |
| EP2302070A3 | European Patent Office (EPO) | A3 | |
| EP2292788B1 | European Patent Office (EPO) | B1 | |
| AT557105T | Austria | T | |
| ATE557105T1 | Austria | T1 | |
| DK2292788T3 | Denmark | T3 | |
| EP2302070B1 | European Patent Office (EPO) | B1 | |
| ES2387878T3 | Spain | T3 | |
| DK2302070T3 | Denmark | T3 | |
| US2012316073A1 | United States of America | A1 | |
| ES2393318T3This record | Spain | T3 | |
| CN102925561A | China | A | |
| JP2013063095A | Japan | A | |
| JP5220597B2 | Japan | B2 | |
| CN101641449B | China | B | |
| US8685889B2 | United States of America | B2 | |
| US8785353B2 | United States of America | B2 | |
| US2014213462A1 | United States of America | A1 | |
| US2014274744A1 | United States of America | A1 | |
| US9023768B2 | United States of America | B2 | |
| US2015232924A1 | United States of America | A1 | |
| CN102925561B | China | B | |
| CN105039313A | China | A | |
| JP5823994B2 | Japan | B2 | |
| US2016060686A1 | United States of America | A1 | |
| US2016145689A1 | United States of America | A1 | |
| US9376716B2 | United States of America | B2 | |
| US9447459B2 | United States of America | B2 | |
| US9453256B2 | United States of America | B2 | |
| US9493820B2 | United States of America | B2 | |
| US2017137872A1 | United States of America | A1 | |
| US2017206314A1 | United States of America | A1 | |
| US2018025111A1 | United States of America | A1 | |
| US9896721B2 | United States of America | B2 | |
| US9898576B2 | United States of America | B2 | |
| US9898577B2 | United States of America | B2 | |
| US2018137241A1 | United States of America | A1 | |
| US2018247017A1 | United States of America | A1 | |
| US10095832B2 | United States of America | B2 | |
| CN105039313B | China | B | |
| US10235494B2 | United States of America | B2 | |
| US2019147977A1 | United States of America | A1 | |
| US2020185056A1 | United States of America | A1 | |
| US10978175B2 | United States of America | B2 | |
| US2021202035A1 | United States of America | A1 |
Numbers
- Publication
- 2393318
- Publication, DOCDB
- 2393318
- Publication, EPODOC
- ES2393318T
- Application
- 10075564
- Application, DOCDB
- 10075564
- Application, EPODOC
- ES20100075564T
Titles2
- Spanish
- Estrategias para la identificación y detección de alto rendimiento de polimorfismos
- English
- Strategies for the identification and detection of high performance polymorphisms
Classification
- CPC, 13
- G16B20/20
- G16B30/00
- C12N15/1065
- C12Q1/6827
- G16B30/10
- C40B30/04
- C40B30/10
- C12Q1/6874
- C12Q1/6809
- C12Q1/6806
- C12Q1/6883
- C12Q2600/156
- C12Q1/6858
- IPC, 3
- C12Q1 68
- G16B20 20
- G16B30 10