Engineering of systems, methods and optimized guide compositions for sequence manipulation
Abstract
A vector system associated with CRISPR - Short Palindromic Repeats Grouped and Regularly Interspaced (CRISPR) (Cas) (CRISPR-Cas) that are not found in nature, engineering, comprising one or more vectors comprising: a) a first regulatory element operably linked to one or more nucleotide sequences encoding one or more polynucleotide sequences of the CRISPR-Cas system comprising a guide sequence, an RNAcrcr and a tracr pairing sequence, in which the sequence guide hybridizes with one or more target sequences at the sites of polynucleotides in a eukaryotic cell, b) a second regulatory element operably linked to a nucleotide sequence encoding a Type II Cas9 protein, in which components (a) and (b) are located in the same vectors or different system vectors, in which The CRISPR-Cas system comprises two or more nuclear localization signals (the NLS) expressed with the nucleotide sequence encoding the Cas9 protein, whereby one or more guide sequences target one or more polynucleotide sites in a eukaryotic cell and the Cas9 protein cleaves one or more polynucleotide sites, according to which the sequence of one or more polynucleotide sites is modified.

Term
7.2 yearsto projected expiry
Projected expiry 12 December 2033, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
16 claims: 4 independent, 12 dependent
- 1REIVINDICACIONES 1. Un sistema vector asociado a CRISPR -Repeticiones Palindrómicas Cortas Agrupadas y Regularmente Interespaciadas (CRISPR) (Cas) (CRISPR-Cas) que no se encuentran en la naturaleza, de ingeniería, que comprende uno o más vectores que comprenden:a) un primer elemento regulador unido de manera operable a una o más secuencias de nucleótidos que codifican una o más secuencias de polinucleótidos del sistema CRISPR-Cas que comprende una secuencia guía, un ARNtracr y una secuencia de emparejamiento tracr, en la que la secuencia gura se hibrida con una o mas secuencias objetivo en los sitios de los polinucleótidos en una célula eucariota, b) un segundo elemento regulador unido de manera operable a secuencia de nucleótidos que codifica una proteína Cas9 de Tipo 11 , en el que los componentes (a) y (b) están situados en los mismos vectores o diferentes vectores del sistema, en el que el sistema CRISPR-Cas comprende dos o más señales de localización nuclear (las NLS) expresadas con la secuencia de nucleótidos que codifica la proteína Cas9, según lo cual una o más secuencias guía fijan como objetivo uno o más sitios de polinucleótidos en una célula eucariota y la proteína Cas9 escinde uno o más sitios de polinucleótidos, según lo cual se modifica la secuencia de uno o más sitios de polinucleÓtidos
- 2Un sistema vector CRISPR-Cas de TIpo 11, que no se encuentra en la naturaleza, de ingeniería, según la reivindicación 1, en el que la proteína Cas9 muta con respecto a una proteína Cas9 de tipo natural correspondiente de manera que la proteína mutada es una nickasa que carece de capacidad para escindir una hebra de un polinucleótido objetivo, según lo cual una o más secuencias guía fijan como objetivo uno o más sitios de polinucleótidos en una célula eucariota y la proteína Cas9 sólo escinde una hebra de los sitios de polinucleótidos, según lo cual se modifica la secuencia de uno o más sitios de polinucleótidos.
- 3El sistema de la reivindicación 1 ó 2, en el que los vectores son vectores víricos
- 4El sistema de la reivindicación 3, en el que los vectores víricos son vectores víricos retrovirales, lentivíricos, adenovíricos, adenoasociados o de herpes simple
- 5El sistema según cualquiera de las reivindicaciones 2-4, en el que la proteína Cas9 comprende una o más mutaciones en los dominios catalíticos RuvC 1, RuvC 11 o RuvC 111
- 6El sistema según cualquiera de las reivindicaciones 2-4, en el que la proteína Cas9 comprende una mutación seleccionada del grupo que consiste en D10A, H840A, N854A y N863A con referencia a la numeración de la posición de una proteína Cas9 de Streptococcus pyogenes (SpCas9).
- 7El sistema según cualquier reivindicación precedente, en el que al menos una NLS está en o cerca del término amino de la proteína Cas9 y/o al menos una NLS está en o cerca del término carboxi de la proteína Cas9
- 8El sistema de la reivindicación 7, en el que al menos una NLS está en o cerca del término amino de la proteína Cas9 yJo al menos una NLS está en o cerca del término carboxi de la proteina Cas9
- 9El sistema según cualquier reivindicación precedente, en el que una o más secuencias de polinucleótidos del sistema CRISPR-Cas comprenden una secuencia guía fusionada a una secuencia cr de activación trans (traer)
- 10El sistema según cualquier reivindicación precedente, en el que la secuencia de polinucleótidos del sistema CRISPR-Cas es un ARN quimérico que comprende la secuencia guía, la secuencia traer y una secuencia de emparejamiento traer
- 11El sistema según cualquier reivindicación precedente, en el que la célula eucariota es una célula de mamífero o una célula humana
- 12El sistema según cualquier reivindicación precedente, en el que la proteína Cas9 es codón-optimizada para expresión en una célula eucariota.
- 13Uso del sistema según cualquiera de las reivindicaciones 1 a 12, para ingeniería de genomas, siempre que el uso no comprenda un procedimiento para modificar la identidad genética de la línea germinal de seres humanos y siempre que dicho uso no sea un método para tratamiento del cuerpo humano o de un animal por cirugía o terapia. 14 El uso de la reivindicación 13, en el que la ingeniería de genomas comprende modificar un polinucleótido objetivo en una célula eucariota, modificar la expresión de un polinucleótido en una célula eucariota, generar una célula eucariota modelo que comprenda un gen de enfermedad mutado o eliminar un gen
- 15El uso de la reivindicaciórJ 13, en el que el uso comprende además reparar dicho polinucleótido fijado como objetivo escindido insertando un polinucleótido de plantilla exágeno, en el que dicha reparación da como resultado 5 una mutación que comprende una inserción, supresión o sustitución de uno o más nucle6tidos de dicho polinucleótido objetivo.
- 16El uso de la reivindicaciórJ 13, en el que el uso comprende además editar dicho polinucle6tido objetivo escindido insertando un polinucleótido de plantilla exágeno, en el que dicha ediciórJ da como resultado una mutación que comprende una inserción, supresión o sustitución de uno o más nucleótidos de dicho polinucleótido objetivo 10 17. El uso según la reivindicación 15ó 16, en el que la inserción es por recombinación homóloga
- 18Uso del sistema según cualquiera de las reivindicaciones 1 a 12, en la producción de un animal transgénico no humano o planta transgénica .
Independent claims16
3,569 paragraphs in 50 sections, as filed
Systems engineering, methods and guide compositions optimized for sequence manipulation.
Field of the Invention
The present description generally refers to systems, methods and compositions used for the control of gene expression that involves targeting a sequence, such as genome disturbance or gene editing, that vector systems related to Short Palindromic Repeats can use Grouped and Regularly Interested (CRISPR) and their components
Report regarding federally sponsored research
This invention was carried out with government support granted by the NIH Pioneer Award, National Institutes of Health, OP1 MH100706. The government has certain rights in the invention.
Background of the invention
Recent advances in genome sequencing techniques and analysis methods have significantly accelerated the ability to catalog and map genetic factors associated with a diverse range of biological functions and diseases. Precise technologies are required to set the target genome to allow the systematic reverse engineering of causal genetic variations allowing selective disturbance of individual genetic elements, as well as advancing the applications of synthetic, biotechnological and medical biology. Although genome editing techniques such as designer zinc fingers, effectors of the transcription activator type (the TALE) or migration meganucleases to produce disturbances of the target genome are available, there is still a need for new genome engineering technologies that are affordable, easy to assemble, scalable and capable of setting multiple positions within the eukaryotic genome.
Summary of the invention
There is a pressing need for robust and alternative systems and techniques to target sequences with a wide range of applications. This invention studies this need and provides related advantages. The CRISPRlCas or the CRISPR-Cas system (both terms are used interchangeably throughout this application) does not require the generation of customized proteins for specific target sequences but rather a single Cas enzyme can be programmed using a short RNA molecule To recognize a specific AON target, in other words the Cas enzyme can be recruited for a specific AON target using said short RNA molecule. Adding the CRISPR-Cas system to the repertoire of genome sequencing techniques and analysis methods can significantly simplify the methodology and accelerate the ability to catalog and map genetic factors associated with a diverse range of biological functions and diseases. To use the CRISPR-Cas system effectively to edit genome without adverse effects, it is critical to understand engineering aspects and optimization of these genome engineering tools, which are aspects of the claimed invention.
In one aspect, the description provides a vector system comprising one or more vectors. In some embodiments, the system comprises: (a) a first regulatory element operably linked to a tracr pairing sequence and one or more insertion sites for inserting one or more guiding sequences upstream of the tracr pairing sequence, wherein when expressed, the guiding sequence it is directed to specific binding of sequences of a CRISPR complex to a sequence set as a target in a cell, for example, eukaryotic cell, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) the guiding sequence that hybridizes to the target-fixed sequence and (2) the tracr pairing sequence that hybridizes to the tracr sequence and (b) a second element regulator operably linked to an enzyme coding sequence encoding said CRISPR enzyme comprising a nuclear localization sequence; in which the components
<dl><dt>(to) </dt><dd>and (b) are located in the same or different vectors of the system. In some embodiments, the component</dd></dl>
<dl><dt>(to) </dt><dd>it further comprises the tracr sequence downstream of the tracr pairing sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, in which, when expressed, each of the two</dd></dl>
<dl><dt /><dd>or more guide sequences directs specific binding of the sequence of a CRISPR complex to a different target sequence in a euca riota cell. In some embodiments, the system comprises the sequence tracr under the control of a third regulatory element, such as a polymerase activator 111. In some embodiments, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% sequence complementarity along the length of the tracr pairing sequence when aligned optimally. In some embodiments, the CRISPR complex comprises one or more nucleic localization sequences of sufficient strength to drive the accumulation of said CRISPR complex in a detectable amount in the nucleus of a eukaryotic cell. Without wishing to be bound by theory, it is believed that a nuclear localization sequence is not necessarily for activity of the CRISPR complex in eukaryotes, but that including such sequences improves system activity, especially in terms of target nucleic acid molecules. in the core In some embodiments, the CRISPR enzyme is an enzyme of the type 11 CRISPR system. In some embodiments, the enzyme</dd></dl>
CRISPR is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Cas9 S. pneumoniae, S. pyogenes or S thermophilus and may include mutated Cas9 derived from these organisms. The enzyme can be a homolog or ortholog of Cas9. In some embodiments, the CRISPR enzyme is codon-optimized for expression in a eukaryotic cell. In some embodiments, the CRISPR enzyme directs the cleavage of one or two strands at the position of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA chain cleavage activity. In some embodiments, the first regulatory element is a polymerase 111 activator. In some embodiments, the second regulatory element is a polymerase 11 activator. In some embodiments, the guide sequence has at least 15, 16, 17, 18, 19, 20, 25 nucleotides or between 10-30 or between 15-25 or between 15-20 nucleotides in length. In general and throughout this specification, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, nucleic acid molecules that are single stranded, double stranded or partially double stranded; nucleic acid molecules comprising one or more free ends, no free end (for example, circular); Nucieic acid molecules comprising DNA, RNA or both and other varieties of polynucleotides known in the art. One type of vector is a ~ plasmid ~, which refers to a circular double stranded DNA loop into which additional segments of DNA can be inserted, such as by classical molecular cloning techniques. Another type of vector is a viral vector, in which virologically derived DNA or RNA sequences are present in the vector for packaging in a virus (eg, retrovirus, defective replication retrovirus, adenovirus, defective replication adenovirus and adeno-associated viruses) . Viric vectors also include polynucleotides supported by a virus for transfection into a host cell. Some vectors are capable of autonomous replication in a host cell into which they are introduced (for example, bacterial vectors with a bacterial origin of episomal mammals and vectors). Other vectors (for example, non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell and thereby replicated together with the host genome. On the other hand, some vectors are capable of directing the expression of the genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." Common expression vectors useful in recombinant DNA techniques are often in the form of plasmids.
Recombinant expression vectors may comprise a nucleic acid in a form suitable for expression of nucleic acid in a host cell, which means that recombinant expression vectors include one or more regulatory elements, which may be selected on the basis of host cells. which have to be used for expression, which are operatively linked to the nucleic acid sequence to be expressed. In a recombinant expression vector, "operably linked" means that the nucleotide sequence of interest is linked to the element or regulatory elements in a manner that allows the expression of the nucleotide sequence (eg, in a transcriptional transcription system in vitro or in a host cell when the vector is introduced into the host cell ).
The term "regulatory element" is intended to include activators, enhancers, internal ribosome entry sites (IRES) and other expression control elements (eg, transcription termination signals, such as polyadenylation signals and poly-U sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESS ION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct the constitutive expression of a nucleotide sequence in many types of host cell and those that direct the expression of the nucleotide sequence only in certain host cells (eg, tissue-specific regulatory sequences). A specific tissue activator can direct expression primarily in a desired tissue of interest, such as muscle, neuron, bone, skin, blood, specific organs (e.g., liver, pancreas) or particular cell types (e.g., lymphocytes ). Regulatory elements may also direct the expression in a time-dependent manner, such as in a cell cycle-dependent or developmentally dependent manner, which may or may not also be tissue specific or cell type. In some embodiments, a vector comprises one or more poi activators 111 (e.g., 1, 2, 3, 4, 5 or more poi activators 111), one or more poi II activators (e.g., 1, 2, 3, 4, 5 or more poi activators 11), one or more poly activators (eg, 1,2,3, 4, 5 or more poly activators) or combinations thereof. Examples of poi 111 activators include, but are not limited to, U6 and H1 activators Examples of poi 11 activators include, but are not limited to, the L TR activator of Rous sarcoma virus (RSV) ) retroviral (optionally with the RSV enhancer), the cytomegalovirus activator (CMV) (optionally with the CMV stimulator) [see, eg, Boshart et al, Cell, 41: 521-530 (1985)], the SV40 activator, the dihydrofolate reduclasse activator, the ~ -aclin activator, the phosphoglycerol kinase (PGK) activator and the EF1a activator are also included in the term "regulatory element ~ stimulatory elements, such as WPRE; CMV stimulators; the R-U5 'segment in LTR of HTLV-1 (Mol. Cell. BioL, Vol. 8 (1), p. 466-472, 1,988); SV40 stimulator and intron sequence between exons 2 and 3 of rabbit ~ -globin (Proc. NatL Acad. Sei. USA., Vol. 78 (3), p. 1,527-31, 1981). It will be appreciated by those skilled in the art that the design of the expression vector may depend on factors such as the choice of host cell to be transformed, the level of expression desired, etc. A vector can be introduced into host cells to thereby produce transcripts, proteins or peptides, including fusion proteins or peptides, encoded by nucleic acids as described herein (e.g., transcripts of clustered and regularly interspaced short palindromic repeats. (CRISPR), proteins, enzymes, mutant forms thereof, fusion proteins thereof, elc.).
Advantageous vectors include lentiviruses and adeno-associated viruses and types of such vectors can also be selected to target particular types of cells.
In one aspect, the description J provides a vector comprising a regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme comprising one or more nuclear localization sequences. In some embodiments, said regulatory element conducts transcription of the CRISPR enzyme in a eukaryotic cell so that said CRISPR enzyme accumulates it in a detectable amount in the nucleus of the eukaryotic cell. In some embodiments, the regulatory element is a polymerase 11 activator. In some embodiments, the CRISPR enzyme is an enzyme of the type 11 CRISPR system. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Cas9 of S. pneumoniae, S. pyogenes or S. thermophilus and may include mutated Cas9 from these organisms. In some embodiments, the CRISPR enzyme is codon-optimized for expression in a eukaryotic cell. In some embodiments, the CRISPR enzyme directs the cleavage of one or two strands at the position of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA chain cleavage activity.
In one aspect, the description provides a CRISPR enzyme comprising one or more nuclear localization sequences of sufficient strength to drive the accumulation of said CRISPR enzyme in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is an enzyme of the type 11 CRISPR system. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Cas9 of S. pneumoniae, S. pyogenes or S. thermophilus and may include mutated Cas9 from these organisms. The enzyme can be a homolog or ortholog of Cas9. In some embodiments, the CRISPR enzyme lacks the ability to cleave one or more strands of an objective sequence to which it binds.
In one aspect, the description provides a host eukaryotic cell comprising: (a) a first regulatory element operably linked to a tracr pairing sequence and one or more insertion sites J to insert one or more guiding sequences upstream of the tracr pairing sequence, in which, when expressed, the guiding sequence directs the specific binding of the sequence of a CRISPR complex to an objective sequence in a euca riota cell, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) the guiding sequence that hybridizes to the target sequence and (2) the tracr pairing sequence that hybridizes to the sequence bringing and / or (b) a second element regulator operably linked to an enzyme coding sequence encoding said CRISPR enzyme comprising a nuclear localization sequence. In some embodiments, the host cell comprises components (a) and (b). In some embodiments, component (a), component (b) or components (a) and (b) are stably integrated into a genome of the host eukaryotic cell. In some embodiments, component (a) further comprises the sequence brought downstream of the Iracr pairing sequence under the conlrol of the first regulatory element. In some embodiments, component (a) further comprises two or more guiding sequences operably linked to the first regulatory element, in which when expressed, each of the two or more guiding sequences directs the specific sequence junction of a sequence. CRISPR complex to a different target sequence in a eukaryotic cell. In some embodiments, the eukaryotic host cell further comprises a third regulatory element, such as a polymerase 111 activator, operably linked to said tracr sequence. In some embodiments, the bring sequence has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% sequence complementarity along the length of the tracr pairing sequence when aligned with optimally _ In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences of sufficient strength to drive the accumulation of said CRISPR enzyme in a detectable amount in the nucleus of a euca riota cell. In some embodiments, the CRISPR enzyme is an enzyme of the type 11 CRISPR system. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Cas9 of S. pneumoniae, s. pyogenes or
S. thermophilus and may include mutated Cas9 from these organisms. The enzyme can be a homolog or ortholog of Cas9. In some embodiments, the CRISPR enzyme is codon-optimized for expression in a eukaryotic cell. In some embodiments, the CRISPR enzyme directs the cleavage of one or two strands at the position of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is an activator of polymerase 111. In some embodiments, the second regulatory element is an activator of polymerase 11. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, 25 nucleotides or between 10-30 or between 15-25 or between 15-20 nucleotides in length. In one aspect, the description provides a non-human eukaryotic organism; preferably a multicellular eukaryotic organism, comprising a host eukaryotic cell according to any of the described embodiments. In other aspects, the description provides a eukaryotic organism; preferably a multicellular eukaryotic organism, comprising a host eukaryotic cell according to any of the described embodiments. The organism in some embodiments of these aspects may be an animal; for example a mammal. Also, the organism can be a arthropod lal like an inseclo. The organism can also be a plant. In addition, the organism can be a fungus
In one aspect, the description provides a case comprising one or more of the components described herein. In some embodiments, the case comprises a vector system and instructions for using the case. In some embodiments, the vector system comprises: (a) a first regulatory element operably linked to a bring pairing sequence and one or more insertion sites to insert one or more guiding sequences upstream of the tracr pairing sequence, in which when expressed, the guiding sequence directs the specific binding of the sequence of a CRISPR complex to a sequence set as a target in a eukaryotic cell, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) the guide sequence that hybridizes to the target sequence and (2) the tracr pairing sequence that hybridizes to the tracr sequence and / or (b) a second element regulator operably linked to an enzyme coding sequence encoding said CRISPR enzyme comprising a nuclear localization sequence. In some embodiments, the case comprises components (a) and (b) located in the same or different system vectors. In some embodiments, component (a) further comprises the tracr sequence downstream of the pairing sequence tracr under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, in which when expressed, each of the two or more guide sequences directs the specific binding of the sequence of a CRISPR complex to a different target sequence in a eukaryotic cell. In some embodiments, the system further comprises a third regulatory element, such as an activator of polymerase 111, an operably gone to said tracr sequence. In some embodiments, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% sequence complementarity along the length of the tracr pairing sequence when aligned with optimal way. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences of sufficient strength to drive the accumulation of said CRISPR enzyme in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is an enzyme of the type 11 CRISPR system. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Cas9 of s. pneumoniae, s. pyogenes or s. thermophilus and may include mutated Cas9 from these organisms. The enzyme can be a homolog or ortholog of Cas9. In some embodiments, the CRISPR enzyme is codon-optimized for expression in a eukaryotic cell. In some embodiments, the CRISPR enzyme directs the cleavage of one or two strands at the position of the target sequence. In some embodiments, the CRISPR enzyme lacks cleavage activity of the AON strand. In some embodiments, the first regulatory element is an activator of polymerase 111. In some embodiments, the second regulatory element is an activator of polymerase 11. In some embodiments, the guide sequence is at least 15, 16, 17, 18 , 19,20.25 nucle6tids or between 10-30 or between 15-25 or between 15-20 nudeotides in length.
In one aspect, the description provides a method for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method comprises allowing a CRISPR complex to be attached to the target polynucleotide to effect cleavage of said target polynucleotide thereby modifying the target polynucleotide, in which the CRISPR complex comprises a complexed CRISPR enzyme with a gluttonic sequence. to a target set sequence within said target polynucleotide, wherein said guide sequence is linked to a tracr pairing sequence which in turn is transcribed to a tracr sequence. In some embodiments, said cleavage comprises cleaving one or two strands at the position of the target sequence by said CRISPR enzyme. In some embodiments, said cleavage results in decreased transcription of a target gene. In some embodiments, the method further comprises repairing said target polynudeotide cleaved by homologous recombination with an exogenous plant polynucleotide, wherein said repair results in a mutation comprising insertion, deletion or replacement of one or more nucleotides of said polynucleotide. objective. In some embodiments, said mutation results in one or more amino acid changes in an expressed protein of a gene that comprises the target sequence. In some embodiments, the method further comprises supplying one or more vectors to said eukaryotic cell, in which one or more vectOfes conducts the expression of one or more of: the CRISPR enzyme, the guiding sequence linked to the tracr pairing sequence and the tracr sequence In some embodiments, said vectors are delivered to the eukaryotic cell in an individual. In some embodiments, said modification takes place in said eukaryotic cell in a cell culture. In some embodiments, the method further comprises isolating said eukaryotic cell from an individual before said modification. In some embodiments, the method further comprises returning said cell and eukaryotic cells from there to said individual
In one aspect, the description provides a method for modifying the expression of a polynucleotide in a eukaryotic cell. In some embodiments, the method comprises allowing a CRISPR complex to bind to the polynudeotide so that said binding results in the increased or decreased expression of said polynudeotide; wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence is transcribed to a target-targeted sequence within said polynucleotide, wherein said guide sequence is linked to a tracr pairing sequence that in turn is transcribed into a tracr sequence In some embodiments, the method further comprises providing one or more vectors to said eukaryotic cells, in which one or more vectors lead to the expression of one or more of: the CRISPR enzyme, the guiding sequence linked to the tracr pairing sequence and the tracr sequence
In one aspect, the description provides a method for generating a model eukaryotic cell that comprises a mutated gene of the disease. In some embodiments, a disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method comprises (a) introducing one or more vectors into an euca renal cell, in which one or more vectors lead to the expression of one or more of: a CRISPR enzyme, a guide sequence linked to a tracr pairing sequence and a tracr sequence and (b) allowing a CRISPR complex to bind to a target polynucleolide to effect cleavage of the target polynudeotide within said disease gene, wherein the CRISPR complex comprises the CRISPR enzyme complexed with (1) the guiding sequence that hybridizes to the target sequence within the target polynudeotide and (2) the tracr pairing sequence that hybridizes to the tracr sequence, thereby generating a model eukaryotic cell that comprises a mutated disease gene. In some embodiments, said cleavage comprises cleaving one or two strands at the position of the target sequence by said CRISPR enzyme. In some embodiments, said cleavage results in decreased transcription of a target gene. In some embodiments, the method further comprises repairing said target polynucleotide cleaved by homologous recombination with an exogenous template polynucleotide, wherein said repair results in a mutation comprising insertion, deletion or replacement of one or more nucleotides of said target polynucleotide. . In some embodiments, said mutation results in one or more amino acid changes in a protein expression of a gene comprising the target sequence.
In one aspect, the description provides a method for developing a biologically active agent that modulates a case of cellular signaling associated with a disease gene. In some embodiments, a disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method comprises (a) contacting a test compound with a model cell of any one of the described embodiments and (b) detecting a change in a reading that is indicative of a reduction
or an increase in the case of cellular signaling associated with said mutation in said disease gene, thereby developing said biologically active agent that modulates said case of cellular signaling associated with said disease gene.
In one aspect, the description provides a recombinant polynucleotide comprising a guide sequence upstream of a tracr pairing sequence, in which the guide sequence when expressed directs the specific binding of the sequence of a CRISPR complex to a corresponding target sequence present. in a eukaryotic cell. In some embodiments, the target sequence is a viral sequence present in a eukaryotic cell. In some embodiments, the target sequence is a proto-oncogene or an oncogene.
In one aspect, the description provides a method for selecting one or more prokaryotic cells by introducing one or more mutations in a gene into one or more prokaryotic cells, the method comprising · introducing one or more vectors into the cell or prokaryotic cells, in which one or more vectors drive the expression of one or more of: a CRISPR enzyme, a guide sequence linked to a bring matching sequence, a tracr sequence and an editing template; wherein the editing template comprises one or more mutations that nullifies the cleavage of the enzyme CRISPR; allow homologous recombination of the editing template with the target polynucleotide in the cell or cells to be selected; allowing a CRISPR complex to bind to an objective polynucleotide to effect cleavage of the target polynucleotide within said gene, wherein the CRISPR complex comprises the CRISPR enzyme complexed with (1) the guide sequence that hybridizes to the target sequence within the gene. target polynudeotide and (2) the tracr pairing sequence that hybridizes to the tracr sequence, in which the binding of the CRISPR complex to the target polynucleotide induces cell death, thereby allowing one or more prokaryotic cells to be selected in which one or more mutations have been introduced. In a preferred embodiment, the CRISPR enzyme is Cas9. In another aspect of the invention, the cell to be selected may be a eukaryotic cell. Aspects of the invention allow the selection of specific cells without requiring a selection marker or a two-stage procedure that may include a counter-selection system.
In some aspects, the description provides a composition that is not found in nature or engineering that comprises a chimeric RNA polynucleotide sequence (siRNA) of the CRISPR-Cas system, in which the polynucleotide sequence comprises: (a) a guide sequence capable of hybridizing a target set sequence in a eukaryotic cell, (b) a tracr pairing sequence and (c) a tracr sequence in which (a), (b) and (c) are arranged in a 5 'to 3' orientation, in which when transcribed, the tracr pairing sequence hybridizes to the tracr sequence and the guiding sequence directs the specific binding of the sequence of a CRISPR complex to the target sequence, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) the guide sequence that hybridizes to the target sequence and (2) the tracr pairing sequence that hybridizes to the tracr sequence or a CRISPR enzyme system, in the that the system is encoded by a vector system comprising one or more vectors comprising 1. a first regulatory element operably linked to a chimeric RNA polynucleotide sequence (siRNA i) of the CRISPR-Cas system, wherein the polynucleotide sequence comprises (a) one or more guide sequences capable of hybridizing one or more target sequences in a eukaryotic cell, (b) a tracr pairing sequence and (c) one or more tracr and 11 sequences. a second regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme comprising at least one or more nuclear localization sequences, in which (a),
(b) Y (c) are arranged in a 5 'to 3' orientation, in which components I and 11 are placed in the same or different system vectors, in which when transcribed, the matching sequence is brought hybridizes to the tracr sequence and the guide sequence directs the specific binding of the sequence of a CRISPR complex to the target sequence, wherein the CRISPR complex comprises the CRISPR enzyme complexed with (1) the guide sequence that hybridizes to the target sequence and (2) the tracr pairing sequence that hybridizes to the tracr sequence or a multiplexed CRISPR enzyme system, in which the system is encoded by a vector system comprising one or more vectors comprising 1. a first regulatory element operably linked to (a) one or more guide sequences capable of hybridizing a target sequence in a cell and (b) at least one or more tracr pairing sequences, 11. a second regulatory element operably linked to an enzyme coding sequence encoding a CRISPR and 111 enzyme. a third regulatory element an operably gone to a tracr sequence, in which components 1, ti and 111 are placed in the same or different system vectors, in which when transcribed, the tracr pairing sequence hybridizes to the tracr sequence and the guide sequence directs the specific binding of the sequence of a CRISPR complex to the target sequence, wherein the CRISPR complex comprises the CRISPR enzyme complexed with (1) the guide sequence that hybridizes to the target sequence and (2) the tracr pairing sequence that hybridizes to the tracr sequence and in which in the multiplexed system they use multiple guide sequences and a single tracr sequence and in which one or more of the guide, tracr and pairing sequences are modified to improve stability.
In aspects of the invention, the modification comprises a secondary engineering structure. For example, the modification may comprise a reduction in the hybridization region between the tracr pairing sequence and the tracr sequence. For example, the modification may also comprise fusing the tracr pairing sequence and the tracr sequence by an artificial loop. The modification may comprise the tracr sequence having a length between 40 and 120 bp. In embodiments, the tracr sequence is between 40 bp and the total length of the tracr sequence. In some embodiments, the length of RNAtrac includes at least 1-67 nucleotides and in some embodiments at least 1-85 nucleotides of wild-type RNAtrac. In some embodiments, at least the nucleotides corresponding to nucleotides 1.07 or 1-85 of wild-type Cas9 tRNA from S. pyogenes can be used. In the event that the CRISPR system uses enzymes other than Cas9 or other than SpCas9, then the corresponding nucleotides may be present in the relevant wild type ARNtrac. In some embodiments, the length of RNAtrac includes no more than nucleotides 1-67 or 1-85 of wild-type RNAtrac. The modification may comprise sequence optimization. In some aspects, sequence optimization may comprise reducing the incidence of polyT sequences in the tracr and / or tracr pairing sequence. Sequence optimization can be combined with the reduction in the hybridization region between the tracr pairing sequence and the tracr sequence; for example, a tracr sequence of reduced length
In one aspect, the description provides the CRISPR-Cas system or the CRISPR enzyme system in which the modification comprises the reduction of polyT sequences in the tracr and / or tracr pairing sequence. In some aspects, one or more T present in a poly-T sequence of the relevant natural type sequence (ie, an extension of more than 3, 4, 5, 6 or more contiguous T bases; in some embodiments, a stretch of not more than 10, 9, 8, 7, 6 contiguous T bases) may be substituted with a non-T nucle6id, eg. an A, so that the chain is broken into smaller sections of T having each section 4 or less than 4 (for example 3 or 2) T adjacent. Bases other than A can be used to replace J, for example C or G or nucleotides that are not found in nature or modified nucleotides. If the T chain is involved in the formation of a hairpin (or stalk), then it is advantageous that the complementary base for the non-T base changes to complement the non-T nucleotide. For example, if the non-T base is an A, then a complement can be changed to a T, for example to conserve or assist in the consideration of secondary structure. For example, 5'-TTTTT can be modified to convert it to 5'-TTIAT and the complementary 5'-AAAAA can be changed to 5'-ATAAA
In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding a polyT terminator sequence. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding a polyT terminator sequence in tracr and / or tracr pairing sequences. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding a polyT terminator sequence in the guide sequence. The polyT terminator sequence may comprise 5 contiguous T bases or more than 5.
In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises modifying loops and / or forks. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing a minimum of two forks in the guide sequence. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing a fork formed by complementation between the tracr and tracr pairing sequence (direct repeat). In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing one or more additional forks at or towards the 3 'end of the ARNtracr sequence. For example, a hairpin can be formed by providing autocomplementary sequences within the RNAtrac sequence linked by a loop so that a hairpin is formed in self-folding. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing additional forks added to the 3 'of the guide sequence. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises extending the 5 'end of the guide sequence. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing one or more forks at the 5 'end of the guide sequence. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding the sequence (S'-AGGACGAAGTCCTAA) to the 5 'end of the guide sequence. Other suitable sequences for forming forks will be known to the skilled person and can be used in certain aspects of the invention. In some aspects, at least 2, 3, 4, 5 or more additional forks are provided. In some aspects no more than 10, 9, 8, 7, 6 additional forks are provided. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises two forks. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises three forks. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises at most five forks
In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing cross-linking or providing one or more modified nucleotides in the polynucleotide sequence. Cross linked yfo modified nucleotides can be provided in any or all of the tracr, tracr coupling and / or guide yfo sequences in the enzyme coding sequence and / or in vector sequences. Modifications may include the inclusion of at least one nucleotide that is not found in nature or a modified nucleotide or analogs thereof. Modified nucleotides can be modified in the ribase, phosphate and / or base moiety. Modified nucleotides may include 2'-Q-methylographs, 2'-deoxy-analogs or 2'-fluoro-analogs. The main chain of nucleic acids can be modified, for example, a main chain of phosphorothioate can be used. The use of blocked nucleic acids (LNAs) or bridge nucleic acids (BNAs) may also be possible. More examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, 7-methylguallOsine
It will be understood that any or all of the foregoing modifications may be provided in isolation or in combination in a given CRISPR-Cas system or CRISPR enzyme system. Said system may include one, two, three, four, five or more of said modifications.
In one aspect, the invention provides the CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is an enzyme of the type 11 CRISPR system, e.g., a Cas9 enzyme. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is made up of less than one thousand amino acids, or less than four thousand amilOacids. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the Cas9 enzyme is StGas9 or StlCas9 or the Cas9 enzyme is a Cas9 enzyme of an organism selected from the group consisting of the genus Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Az.ospirillum, Sphaerochaeta, Lactobacillus, Eubacterium or Corynebacter. In one aspect, the invention provides the CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is a nuclease that directs the cleavage of both chains at the position of the target sequence.
In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the first regulatory element is a polymerase 111 activator. In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in that the second regulatory element is a polymerase activator 11.
In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the guide sequence comprises at least fifteen nucleotides.
In one aspect, The description provides the CRISPR-Cas system or CRISPR enzyme system in which the modification comprises optimized tracr sequence RNA and optimized guide sequence and / or complexed structure of tracr sequence and / or sequence or tracr coupling sequences and secondary stabilizing structures. of tracr sequence and / or tracr sequence with a reduced region of base pairing and condensed RNA elements of tracr sequence and / or in the multiplexed system there are two RNAs comprising a tracer and comprising a plurality of guides or RNA comprising a plurality of chimeras.
In aspects, the chimeric RNA architecture is further optimized according to the results of mutagenesis studies. In chimeric RNA with two or more hairpins, mutations in the proximal direct repeat to stabilize the hairpin can result in ablation of the activity of the CRISPR complex. Mutations in distal direct repetition to shorten or stabilize the hairpin may have no effect on the activity of the CRISPR complex. Randomization of sequences in the region of the bump between the proximal and distal repeats can significantly reduce the activity of the CRISPR complex. Changes in single base pairs or randomization of sequences in the linker region between hairpins can result in the complete loss of activity of the CRISPR complex. Fork stabilization of the distal forks that follow the first fork after the guide sequence can result in maintenance or improvement of the activity of the CRISPR complex. Accordingly, in preferred embodiments, the chimeric RNA architecture can be further optimized by generating a smaller chimeric RNA that may be beneficial for therapeutic delivery options and other uses and this can be achieved by modifying the direct distal repeat repeat in a manner the fork is shortened or stabilized. In further preferred embodiments, the chimeric RNA architecture can be further optimized by stabilization of one or more distal forks. Fork stabilization may include modifying suitable sequences to form forks. In some aspects, at least 2, 3, 4, 5 or more additional forks are provided. In some aspects, no more than 10, 9, 8, 7, 6 additional forks are provided. In some aspects, the stabilization may be cross-linking or other modifications. Modifications may include the inclusion of at least one nucleotide that is not in nature or a modified nucleotide or analogs thereof. The modified nucleotides can be modified in the ribose, phosphate and / or base moiety. Modified nucleotides may include 2'-O-methyl analogs, 2'-deoxy analogs or 2'-fluoro analogs. The main chain of nucleic acids can be modified, for example, a main chain of phosphorothioate can be used. The use of blocked nucleic acids (LNA) or bridge nucleic acids (BNA) may also be possible. More examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo · uridine, pseudouridine, inosine, 7-methylguanosine
In one aspect, the description provides the CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is codon-optimized for expression in a eukaryotic cell.
Accordingly, in some aspects, it is not necessarily required that the length of the RNAcrcr required in a construction of the invention be set, for example, a chimeric construction, and in some aspects it may have between 40 and 120 bp And in some aspects up to the total length of the tracr, for example, in some aspects up to the 3 'end of bringing as marked by the transcription tenancy signal in the bacterial genome. In some embodiments, the length of the RNAcrcr includes at least 1-67 nucleotides and in some embodiments at least 1-85 nucleotides of the wild type RNAcr. In some embodiments, at least nucleotides corresponding to nucleotides 1-67 or 1-85 of wild type Cas9 RNAcrcr of S. pyogenes can be used. In the event that the CRISPR system uses enzymes other than Cas9 or others other than SpCas9, then the corresponding nucleotides may be present in the relevant wild-type RNAcrcr. In some embodiments, the length of the RNAcrcr includes no more than 1-67 or 1-85 nucleotides of the wild type RNAcr. With respect to sequence optimization (for example, reduction in poly-T sequences), for example in terms of internal Ts chains to the tracr pairing sequence (direct repetition) or to the RNAcrcr, in some aspects, one or more Ts present in a poly-T sequence of the relevant wild-type sequence (that is, a segment of more than 3,4,5,6 or more contiguous T bases; In some embodiments, a segment of no more than 10, 9, 8, 7, 6 contiguous T bases) can be substituted with a non-T nucleotide, for example, an A, so that the chain is broken into smaller fragments of T having each fragment 4, or less than 4 (eg, 3 or 2) T contiguous. If the T chain is involved in the formation of a hairpin (or stem-loop), then it is advantageous that the complementary base for the non-T base is changed to complement the non-T nucleotide. For example, if the non-T base is an A, then its complement can be changed to a T, for example, to conserve or assist in the conservation of secondary structure. For example, 5'-TTTTT can be modified to become 5'TITAT and the complementary 5'-AAAAA can be changed to 5'-ATAAA. As for the presence of poly-T terminator sequences in transcript of bring + matching bring, for example, a poly.T terminator (TTTTT or more), in some aspects it is advantageous to add it to the end of the transcript, if it is in shape Two RNA (bring and bring matching) or single guide RNA. With respect to the loops and hairpins in the transcripts bring and tracr pairing, in some aspects it is advantageous that a minimum of two hairpins is present in the chimeric guide RNA. A first fork may be the fork formed by complementation between the tracr sequence and the matching pairing sequence (direct repetition). A second fork may be at the 3 'end of the ARNtracr sequence and this may provide secondary structure for interaction with Cas9. Additional forks can be added to the 3 'of the guide RNA, for example, in some aspects of the invention to increase the stability of the guide RNA. Additionally, the 5 'end of the guide RNA, in some aspects, can be extended. In some aspects, 20 bp can be considered at the 5 'end as a guide sequence. The 5 'portion can be extended One or more forks can be provided in the 5' portion, for example, in some aspects, this can also improve the stability of the guide RNA. In some aspects, the specific hairpin can be provided by attaching the sequence (5'-AGGACGAAGTCCTAA) to the 5 'end of the guide sequence and, in some aspects, this can help improve stability. Other suitable sequences for forming hairpins will be known to the expert and can be used in certain aspects. In some aspects, at least 2, 3, 4, 5 or more additional forks are provided. In some aspects no more than 10, 9, 8, 7, 6 additional forks are provided. The above also provides aspects that imply secondary structure in guide sequences. In some aspects there may be cross-linking and other modifications, for example, to improve stability. Modifications may include inclusions of at least one nucleotide that is not found in nature or a modified nucleotide or analogs thereof. The modified nucleotides can be modified in the ribose, phosphate and / or base moiety. Modified nucleotides may include 2'-O-methyl analogs, 2'-deoxy analogs or 2'-fluoro analogs. The nucleic acid main chain can be modified, for example, a phosphorothioate main chain can be used. The use of blocked nucleic acids (LNA) or bridge nucleic acids (BNA) may also be possible. More examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo · uridine, pseudouridine, inosine, 7-methylguanosine. Such modifications or cross-linking may be present in the guide sequence or other sequences adjacent to the guide sequence.
Brief description of the drawings
A better understanding of the features and advantages of the present invention will be gained by reference to the following detailed description explaining illustrative embodiments, in which the principles of the invention are used, and the accompanying drawings of which:
Figure 1 shows a schematic model of the CRISPR system. The Cas9 nuclease of Streptococcus pyogenes (yellow) is set as a target for genomic DNA by a synthetic guide RNA (sRNA) consisting of a 20-nt guide sequence (blue) and a scaffold (red) The base pairs of the sequence guide with the target DNA (blue), directly upstream of an adjacent proto-spacer unit (PAM); magenta) 5'-NGG requirement and Cas9 mediates a double chain break (OSB) -3 bp upstream of the PAM (red triangle)
Figure 2A-F illustrates an exemplary CRISPR system, a possible mechanism of action, an example adaptation for expression in eukaryotic cells and the results of tests evaluating nuclear localization and CRISPR activity.
Figure 3A-C illustrates an exemplary expression cassette for expression of elements of the CRISPR system in eukaryotic cells, predicted sample sequence structures and activity of the CRISPR system when measured in eukaryotic and prokaryotic cells_
Figure 4A-O illustrates the results of a SpCas9 specificity assessment for an example objective.
Figure 5A-G illustrates an exemplary vector system and the results for use in directing homologous recombination in eukaryotic cells.
Figure 6A-C illustrates a comparison of different RNAcrcr transcripts to target Cas9-mediated genes.
Figure 7 AO illustrates an exemplary CRISPR system, an example adaptation for expression in eukaryotic cells and the results of the tests evaluating the activity of CRISPR.
Figure 8A-C illustrates exemplary manipulations of a CRISPR system to target genomic sites in mammalian cells.
Figure 9A-B illustrates the results of an analysis by the Northem method of treating cRNA in mammalian cells.
Figure 10A-C illustrates a schematic representation of chimeric RNA and the results of SURVEYOR assays for CRISPR system activity in eukaryotic cells.
Figure lIA-B illustrates a graphical representation of the results of SURVEYOR trials for activity of the CRISPR system in eukaryotic cells
Figure 12 illustrates secondary structures envisioned for exemplary chimeric RNAs comprising a guide sequence, tracr coupling sequence and tracr sequence.
Figure 13 is a phylogenetic tree of genes Case
Figure 14A-F shows the phylogenetic analysis revealed by five Cas9 families, including three large Cas9 groups (-1,400 amino acids) and two small Cas9s (-1,100 amino acids).
Figure 15 shows a graph representing the function of different optimized guide RNAs
Figure 16 shows the sequence and structure of different guide chimeric RNAs.
Figure 17 shows the collapsed structure of the ARNtracr and directed repeat
Figure 18 A and B shows data of the chimeric guide RNA optimization of St1Cas9 in vitro.
Figure 19A-B shows the cleavage of unmethylated or methylated targets by SpCas9 cell lysate.
Figure 20A-G shows the optimization of the guide RNA architecture for SpCas9-mediated mammalian genome editing. (a) Bicistronic expression vector scheme (PX330) for single guide RNA driven by U6 activator (RNAg) and Cas9 of human codon-optimized Streptococcus pyogenes driven by CBh activator (hSpCas9) used for all subsequent experiments. The siRNA consists of a guide sequence of 20-nt (blue) and scaffold (red), truncated in various positions as indicated. (b) SURVEYOR assay for insertions or deletions mediated by SpCas9 in human EMX1 and PVALB sites. The arrows indicate the expected SURVEYOR fragments (n = 3). (e) Analysis by the Northem method for the four truncated architectures of ARNsg, with Ul as load control. (d) Both wild-type (wt) or nickase (D10A) spCas9-activated insertion mutants of a HindlU site in the human EMX1 gene. Single stranded oligonucleotides (the ssOONs) oriented in the sense or antisense direction in relation to the genomic sequence, were used as homologous recombination templates. (e) Human SERPINB5 site scheme. RNAs and PAMs are indicated by colored bars above the sequences; methylcytosine (Me) (pink) are highlighted and listed in relation to the transcription initiation site (TSS, +1). (f) SERPINB5 methylation status tested by bisulfite sequencing of 16 clones. Filled circles, methylated CpG; open circles, unmethylated CpG. (g) Efficacy of modification by three rRNAs targeting the methylated region of SERPINB5, tested by deep sequencing (n = 2). Error bars indicate Wilson intervals (Online Methods).
Figure 21A-B shows the additional optimization of the CRISPR-Cas RNAsg architecture. (a) Scheme of four additional architectures of ARNsg, I-IV. Each consists of a 20-nt guide sequence (blue) linked to direct repetition (RO, gray), which hybridizes to the RNAcrcr (red). The RD-ARNtracr hybrid is truncated at +12 or +22, as indicated with an artificial GAAA stem-loop. The truncated positions of ARNtracr are listed according to the transcription start site indicated previously for ARNtracr. ARNsg 11 and IV architectures support mutations within their poly-U tracts, which could serve as premature transcription terminators. (b) SURVEYOR assay for insertions or deletions mediated by SpCas9 in the human EMX1 site for target sites 1-3. The arrows indicate the SURVEYOR fragments expressed (n = 3).
Figure 22 illustrates the visualization of some target sites in the human genome.
Figure 23A-B shows (A) a schema of the rRNA and (B) the SURVEYOR analysis for five variants of rRNAs for SaCas9 for an optimal truncated architecture with the highest cleavage efficiency.
The figures herein are for illustrative purposes only and are not necessarily drawn to scale.
Detailed description of the invention
The terms "polynucleotide", "nucleotide", "nucleotide sequence", "nucleic acid" and "oligonucleotide" are used interchangeably. They refer to a polymeric form of nucleotides of any length, deoxyribonucleotides or ribonucleotides or analogs thereof. The polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. The following are not limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, sites (site) defined from binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interference RNA (siRNA), Short hairpin RNA (shRNA), micro-RNA (mRNA), ribozymes, and cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, DNA isolated from any sequence, RNA isolated from any sequence, nucleic acid probes and primers. A polynucleotide may comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If there are, modifications can be made to the nucleotide structure before or after the polymer assembly. The nucleotide sequence can be interrupted by non-nucleotide components. A polynucleotide can be further modified after polymerization, such as by conjugation with a labeled component.
In aspects of the invention, the terms "chimeric RNA", "chimeric guide RNA", "guide RNA", "single guide RNA" and "synthetic guide RNA" are used interchangeably and refer to the polynucleotide sequence comprising the guide sequence, the bring sequence and the matching sequence bring. The term "guide sequence" refers to the sequence of approximately 20 bp within the guide RNA that specifies the target site and can be used interchangeably with the terms "guide" or "spacer" The term "matching sequence bring" also It can be used interchangeably with the term "repetition or direct repetitions".
As used herein, the term "natural type" is a term of the art understood by experts and means the typical form of an organism, strain, gene or characteristic as found in nature as distinguished from mutant forms. or variants.
As used herein the term "variant" should be considered to mean the display of qualities that present a pattern that deviates from what occurs in nature.
The terms "not found in nature" or "engineering" are used interchangeably and indicate the involvement of the hand of man. The terms, when referring to nucleic acid molecules or polypeptides mean that the nucleic acid molecule or polypeptide is at least substantially free of at least one other component with which they are naturally associated in nature and as found in the nature
"Complementarity" refers to the ability of a nucleic acid to form the hydrogen bond or bonds with another nucleic acid sequence by traditional Watson-Crick base pairing or other non-traditional types. A percentage of complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 of a total of 10 being the complementarity of 50%, 60%, 70%, 80%, 90% and 100%). "Perfectly complementary" means that all contiguous residues of a nucleic acid sequence will form a hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. "Substantially complementary" as used herein refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98 %, 99% or 100% for a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21,22,23, 24, 25, 30, 35 , 40, 45, 50 or more nucleotides or refers to two nucleic acids that hybridize under stringent conditions
As used herein, "stringent conditions" for hybridization refers to the conditions in which a nucleic acid that has complementarity to an objective sequence predominantly hybridizes with the objective sequence and substantially does not hybridize to the non-objective sequences. Rigorous conditions are generally sequence dependent, and vary depending on a number of factors. In general, the longer the sequence, the higher the temperature at which the sequence is specifically hybridized to its target sequence Non-limiting examples of stringent conditions are described in detail in Tijssen (1993), Laboratory Techniques in Biochemistry and Molecular Biology- Hybridization With Nucleic Acid Probes Part 1, Second Chapter "Qverview of principies of hybridization and the strategy of nucleic acid probe assay", Elsevier, NY
"Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that stabilizes via hydrogen bonding between the bases of the nucleotide moieties. Hydrogen bonding can take place by base pairing of Watson Crick, Hoogstein binding or in any other sequence-specific way. The complex may comprise two strands that form a duplex structure, three or more strands that form a multicatenary complex, a single strand that is self-inhibited or any combination thereof. A hybridization reaction may constitute a stage in a more extensive procedure, such as PCR initiation.
or excision of a polynucleotide by an enzyme. A sequence capable of hybridization with a given sequence is referred to as the "complement" of the given sequence.
As used herein, "stabilization" or "increased stability" with respect to components of the CRISPR system refers to ensuring or stabilizing the structure of the molecule. This can be accomplished by introducing one or more mutations, including single or multiple base pair changes, increasing the number of forks, cross-linking, breaking particular stretches of nucleotides and other modifications. Modifications may include the inclusion of at least one nucleotide that is not found in nature or a modified nuotide or analogs thereof. Modified nucleotides can be modified in the ribose, phosphate and / or base moiety. The modified nucleotides may include 2'-Q-methyl-analogs, 2'-deoxy-analogs or 2'-fluoroangoles. The main chain of nucleic acids can be modified, for example, a main chain of phosphorothioate can be used. The use of blocked nucleic acids (LNAs) or bridge nucleic acids (BNAs) may also be possible. More examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, 7-methylguanosine. These modifications can be applied to any component of the CRSIPR system. In a preferred embodiment these modifications are made to the RNA components, for example, the guide RNA or chimeric polynucleotide sequence.
As used herein, "expression" refers to the procedure by which a polynucleotide is transcribed from a DNA template (such as in mRNA or other RNA transcript) and / or the method by which it is translated with subsequently an mRNA transcribed into peptides, polypeptides or proteins. Transcripts and encoded polypeptides can be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, the expression may include the mRNA splicing process in a eukaryotic cell.
The terms "polypeptide", "peptide" and "protein" are used interchangeably herein to refer to amino acid polymers of any length. The polymer can be linear or branched and can comprise modified amino acids and can be interrupted by non-amino acids. The terms also include an amino acid polymer that has been modified; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation or any other manipulation, such as conjugation with a labeled component. As used herein the term "amino acid · includes natural and / or unnatural or synthetic amino acids, including glycine and the two optical isomers O or L and amino acid analogs and peptidomimetics
The terms "individual," "individual" and "patient" are used interchangeably herein to refer to a vertebrate, preferably a mammal, more preferably a human being. Mammals include, but are not limited to, murine animals, apes, humans, farm animals, sports animals and pets. Tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro are also included. In some embodiments, an individual may be an invertebrate animal, for example, an insect or a nematode; while in others, an individual can be a plant or a fungus.
The terms "therapeutic agent", "capable therapeutic agent" or "treatment agent" are used interchangeably and refer to a molecule or compound that confers some beneficial effect on administration to an individual. The beneficial effect includes allowing diagnostic determinations; relief of a disease, symptom, disorder or pathological condition; reduce or prevent the onset of a disease, symptom, disorder or condition and in general counteract a disease, symptom, disorder or pathological condition
As used herein, "treatment" or "treat," or "alleviate" or "relieve" are used interchangeably. These terms refer to a proposal to obtain beneficial or desired results including but not limited to a therapeutic benefit and / or a prophylactic benefit. By therapeutic benefit is meant any therapeutically relevant improvement in, or effect on, one or more diseases, conditions or symptoms under treatment. For prophylactic benefit, the compositions may be administered to an individual at risk of developing a particular disease, condition or symptom or to an individual indicating one or more of the physiological symptoms of a disease, even though the disease, condition or symptom may not have been manifested yet
The term "effective amount" or "therapeutically effective amount" refers to the amount of an agent that is sufficient to effect beneficial or desired results. The therapeutically effective amount may vary depending on one or more of: the individual and the disease being treated, the weight and age of the individual, the severity of the disease, the manner of administration and the like, which can easily be determined by An expert in the field. The term also applies to a dose that will provide an image for detection by any one of the imaging methods described herein. The specific dose may vary depending on one or more of: the particular agent chosen, the dosage regimen to be followed, if administered in association with other compounds, the timing of administration, the tissue to be subjected to image formation and the physical supply system in which it is supported.
The practice of the present invention employs, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomic and recombinant DNA, which are within the skill of the art. See Sambrook, Fritsch and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd edition (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F
M. Ausubel, et al. eds., (1987); METHODS IN ENZYMOLOGY series (Academic Press, Inc.): PCR 2: A PRACTICAL APPROACH (MJ MacPherson, BD Hames and GR Taylor eds. (1995, Hartow and Lane, eds.
(1,988) ANT1BODIES, A LABORATORY MANUAL ANIMAL CELL CULTURE YEAR (R. 1. Freshney, ed. (1987 ».
Several aspects of the invention relate to vector systems comprising one or more vectors or vectors as such. Vectors can be designed for expression of CRISPR transcripts (eg, transcripts of nude acids, proteins or enzymes) in prokaryotic or eukaryotic cells. For example, CRISPR transcripts can be expressed in bacterial cells such as Escherichia eoli, insect cells (using baculovirus expression vectors), yeast cells or mammalian cells. Suitable host cells are further discussed in Goeddel, GENE EXPRESS ION TECHNOLOGY: METHODS IN ENZVMOLOGY 185, Academic Press, San Diego; Calif. (1,990). As an alternative, the recombinant expression vector can be transcribed and translated in vitro, for example using T7 activator and TI polymerase regulatory sequences.
Vectors can be introduced and propagated in a prokaryotic. In some embodiments, a prokaryotic was used to multiply copies of a vector that has to be introduced into a eukaryotic cell or as an intermediate vector in the production of a vector that has to be introduced into a eukaryotic cell (e.g., by multiplying a plasmid as part of a viral vector packaging system). In some embodiments, a prokaryotic is used to multiply copies of a vector and express one or more nucleic acids, such as to provide a source of a
or more proteins for delivery to a host cell or host organism. Protein expression in prokaryotes is most frequently carried out in Escherichia coli with vectors containing constitutive or inducible activators that direct the expression of fusion or non-fusion proteins. Fusion vectors add a series of amino acids to a protein encoded there, such as the amino acid term of the recombinant protein. Such fusion vectors can serve one or more purposes, such as: (i) to increase the expression of reco protein rolling ; (ii) to increase the solubility of the recombinant protein and (iii) to aid in the purification of the recombinant protein by acting as an affinity purification ligand. Frequently, in fusion expression vectors, a proteolytic cleavage site is introduced at the junction of the fusion moiety and the recombinant protein to allow separation of the recombinant protein from the fusion moiety after purification of the fusion protein. Such enzymes, and their recognition or similar origin sequences, include Factor Xa, thrombin and enterokinase. Example fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67: 31-40), pMAL (New England Biolabs, Beverly, Mass_) and pRIT5 (Pharmacia, Piscataway, NJ) that condense glutathione S-transferase (GST), maltose E binding protein or protein A, respectively, to the target recombinant protein.
Examples of suitable inducible, inducible, E. coli expression vectors include pTrc (Amrann et al., (1,988) Gene 69: 301-315) and pET 11d (Studier et al., GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calf. (1 .990) 60-89)
In some embodiments, a vector is a yeast expression vector. Examples of vectors for expression in yeast Saccharomyces cerivisae include pYepSec1 (Baldari, et al., 1987. EMBO J. 6: 229-234), pMFa (Kuijan and Herskowilz, 1982. Cell 30: 933-943), pJRY88 (Schultz et al. ., 1987. Gene 54: 113-123), pYES2 (Invitrogen Corporation, San Diego, Galif.) And picZ (lnVitrogen Corp, San Diego, Calif.).
In some embodiments, a vector conducts protein expression in insect cells using baculovirus expression vectors. Baculovirus vectors available for expression of proteins in cultured insect cells (eg, SF9 cells) include the pAe series (Smith, et al., 1983. Mol. CelL BioL 3: 2.156-2.165) and the pVL series (Lucklow and Summers, 1989. Virology 170: 31-39).
In some embodiments, a vector is capable of conducting the expression of one or more sequences in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987, Nature 329: 840) and pMT2PC (Kaufman, et al., 1987. EMBO J. 6: 187-195). When used in mammalian cells, the control functions of the expression vector are typically provided by one or more regulatory elements. For example, commonly used activators come from polyoma, adenovirus 2, cytomegalovirus, simian virus 40 and others described herein and cOllOcids in the art For other expression systems suitable for both prokaryotic and eukaryotic cells see, for example, Chapters 16 and 17 of Sambmok, et al., MOLECULAR CLONING: A LABORATORY MANUAL 28 ed., Cold Spring Harbar Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbar, N. Y, 1989
In some embodiments, the recombinant mammalian expression vector is capable of directing the expression of the nucleic acid preferably in a particular cell type (for example, tissue-specific regulatory elements are used to express the nucleic acid). Fabric specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific activators include the albumin activator (liver-specific; Pinkert, et al., 1987. Genes Dev. 1.268-277), lymphoid-specific activators (Calame and Eaton, 1988. Adv. ImmunoL 43: 235-275), in particular T cell receptor activators (Winoto and Baltimore, 1989. EMBO J. 8: 729-733) and immunoglobulins (Baneiji, et al., 1983. Cell 33: 729-740; Queen and Baltimore, 1983. Cell 33: 741-748), neuron-specific activators (e.g., neurofilament activator; Byrne and Ruddle, 1989. Proc. Natl. Acad Sci. USA 86: 5,473-5,477), pancreas-specific activators (Edlund , et al., 1985. Science 230: 912-916) and specific activators of the mammary glands (eg, whey activator; US Patent No. 4,873,192 and Publication of the Application European Patent No. 264,166). Also included are developmentally regulated activators, eg, murine hox activators (Kessel and Gruss, 1990. Science 249: 374-379) and the a-fetoprotein activator (Campes and Tilghman, 1989. Genes Dev. 3: 537-546)
In some embodiments, a regulatory element is operably linked to one or more elements of a CRISPR system in order to drive the expression of one or more elements of the CRISPR system. In general, the CRISPR (Short Palindromic Repeated Clusters and Regularly Interspaced), also known as the SPIDR (Separate Direct Interleaved Repeats, for its acronym in English), constitute a family of DNA sites that are normally specific to a particular bacterial species. The CRISPR site comprises a distinct class of intercalated short sequence repeats (the SSRs) that were recognized in E. col; (Ishino et al., J. Bacteriol., 169: 5.429-5.433 {1.98?] And Nakata et al., J. BacterioL, 171 '3.553-3.556 (1989)) and associated genes Similar interleaved SSRs have been identified in Haloferax mediterranei, Streptococcus pyogenes, Anabaena and Mycobacterium tuberculosis (See, Groensn et al., Mol. MicrobioL, 10: 1,057-1,065 (1993); Hoe et al., Emerg. Infect. Dis., 5: 254-263 (1999); Masepohl et al., Biochim Biophys Acta 1,307: 26--30 (1996) and Mojica et al., Mol. MicrobioL, 17: 85-93 {1,995]). The CRISPR site typically differs from the other SSRs by the structure of the repetitions, which have been called regularly short spaced repetitions (the SRSR) (Janssen et al. OMICS J. Integ. BioL, 6: 23-33 [2,002) Y Mojica et al., Mol. MicrobioL, 36: 244-246 (2,000]). In general, repetitions are short elements that occur in clusters that are regularly spaced by unique intervention sequences with a substantially constant length (Mojica et al., (2,000), supra), although repetition sequences are highly conserved between strains. , the number of interleaved repetitions and the sequences of the spacer regions typically differs from strain to strain (van Embden et al., J. BacterioL, 182: 2,393-2,401 (2,000)). The CRISPR site has been identified in more than 40 prokaryotes (See, eg, Jansen et al., Mol. Microbiol., 43: 1,565-1,575 [2,002] and Mojica et al., [2,005)) including, but not limited to Aeropyrum, Pyrobaculum, Sulfolobus, Archaeoglobus. Halocarcula, Methanobacterium, Methanococcus, Methanosarcina, Methanopyrus, Pyrococcus, Picrophilus, Thermoplasma, Corynebacterium, Mycobacterium, Streptomyces, Aquifex, Porphyromonas, Chlorobium, Thermus, Bacillus, Listeria, Staphylobacterium, Staphylobacteria, Staphylobacterium, Staphylobacterium Nitrosomonas, Desulfovibrio, Geobacter, Myxococcus, Campylobacter, Wolinella. Acinetobacter, Erwinia, Escherichia, Legionella, Methylococcus, Pastaurella, Photobacterium, Salmonella, Xanthomonas, Yersinia, Treponema and Thennotoga
In general, "CRISPR system" collectively refers to transcripts and other elements involved in the expression of or directed to the activity of genes associated with CRISPR ("Cas"), including sequences encoding a Cas gene, a bring sequence (CRISPR trans -activating) (e.g., ARNtracr or an active partial ARNtracr), a tracr pairing sequence (including a "directed repeat" and a partial direct repetition treated by RNAtracr in the context of an endogenous CRISPR system), a guide sequence (also referred to as a "spacer" in the context: or a CRISPR system endogenous) or other sequences and transcripts of a CRISPR site In some embodiments, one or more elements of a CRISPR system is derived from a CRISPR type 1, type 11 or type 111 system In some embodiments, one or more elements of a CRISPR system comes from a particular organism comprising an endogenous CRISPR system, such as Streptococcus pyogenes. In general, a CRISPR system is characterized by elements that activate the formation of a CRISPR complex at the site of an objective sequence (also referred to as a proto-spacer in the context of an endogenous CRISPR system). In the context of the formation of a CRISPR complex, "objective sequence" refers to a sequence for which a sequence is designed to guide complementarity, where hybridization between an objective sequence and a guide sequence activates the formation of a CRISPR complex. Complete complementarity is not necessarily required, provided there is sufficient complementarity to produce hybridization and activate the formation of a CRISPR complex. An objective sequence may comprise any polynucleotide, such as DNA or RNA polynucleotide. In some embodiments, an objective sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence may be within an organelle of a eukaryotic cell, for example, mitochondria or chloroplast. A sequence or template that can be used for recombination at the target site comprising the target sequences is referred to as "editing template" or ~ editing polynucleotide "or" editing sequence. "In aspects of the invention, a polynucleotide Exogenous template can be referred to as an editing template.In one aspect of the invention recombination is homologous recombination.
Typically, in the context of an endogenous CRISPR system, the formation of a CRISPR complex (comprising a guide sequence hybridized to an objective sequence and complexed with one or more Cas proteins) results in the cleavage of one or more strands at or near of (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 or more base pairs of) the target sequence. Without wishing to be bound by theory, the tracr sequence, which may comprise or consist of all or a portion of a natural-type sequence (e.g., about or more than about 20, 26, 32, 45, 48, 54 , 63, 67, 85 or more nucleotides of a wild-type bring sequence), may also be part of a CRISPR complex, such as by hybridization along at least a portion of the tracr sequence to all or a portion of a tracr pairing sequence that is operably linked to the guide sequence. In some embodiments, the tracr sequence presents sufficient complementarity to a matching sequence to bring to hybridize and participate in the formation of a CRISPR complex. As with the objective sequence, it is believed that complete complementarity is not necessary, provided there is enough to be functional. In some embodiments, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% sequence complementarity along the length of the matching sequence to bring when aligned with optimal way. In some embodiments, one or more vectors that drive the expression of one or more elements of a CRISPR system are introduced into a host cell so that the expression of the elements of the CRISPR system directs the formation of a CRISPR complex at one or more sites. objective. For example, a Cas enzyme, a guide sequence linked to a tracr pairing sequence and a bring sequence could each be operably linked to separate regulatory elements in separate vectors. Alternatively, two or more of the expressed elements thereof or Different regulatory elements can be combined into a single vector, providing one or more additional vectors any component of the CRISPR system not included in the first vector. The elements of the CRISPR system that are combined into a single vector can be arranged in any suitable orientation, such as an element located 5 'with respect to raises up "of) or 3' with respect to (" downstream "of) a Second element The coding sequence of an element may be located in the same or opposite strand of the coding sequence of a second element and oriented in the same or opposite direction. In some embodiments, a single activator drives the expression of a transcript encoding a CRISPR enzyme and one or more of the guide sequence, the tracr pairing sequence (optionally operably linked to the guide sequence) and a tracr sequence embedded within one or more intron sequences (eg, each in a different intron, two or more in at least one intron or all in a single intron). In some embodiments, the CRISPR enzyme, guide sequence, tracr pairing sequence and tracr sequence are operably linked to and expressed from the same activator.
In some embodiments, a vector comprises one or more insertion sites, such as a restriction endonuclease recognition sequence (also referred to as a "cloning site"). In some embodiments, one or more insertion sites (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. In some embodiments, a vector comprises an insertion site upstream of a pairing sequence bringing and optionally downstream of a regulatory element operably linked to the tracr pairing sequence, such that after insertion of a guide sequence in the insertion site and in the expression of the guide sequence direct the specific binding of the sequence of a CRISPR complex to an objective sequence in a eukaryotic cell. In some embodiments, a vector comprises two or more insertion sites, each insertion site being located between two tracr pairing sequences so as to allow insertion of a guide sequence at each site. In said arrangement, two or more guide sequences may comprise two or more copies of a single guide sequence, two or more different guide sequences or combinations thereof. When multiple different guide sequences are used, a single expression construct can be used to target CRISPR activity at multiple corresponding, different, target sequences within a cell. For example, a single vector may comprise approximately
or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more guide sequences. In some embodiments, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more of such vectors containing a guide sequence and optionally supplied to a cell can be provided.
In some embodiments, a vector comprises a regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme, such as a protein. Case Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, CsaS, Csn2, Csm2, Csm3, Csm4, Csm5, Csmr, Cmr1, Cmr1, Cmr1, Cmr1, Cmr1, Cmr1, Cmr1, Cmr1, Cmr1, Cmr3 Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, CSf2, CSf3, CSf4, counterparts thereof or modified versions thereof. These enzymes are known; for example, the amino acid sequence of S. pyogenes cas9 protein can be found in the database S "" '; ssProt with accession number 099ZW2. In some embodiments, the unmodified CRISPR enzyme exhibits DNA cleavage activity, such as Cas9. In some embodiments, the enzyme CRISPR is Cas9 and may be Cas9 of S. pyogenes or S. pneumoniae. In some embodiments, the CRISPR enzyme directs the cleavage of one or both strands at the position of an objective sequence, such as within the objective sequence and / or within the complement of the objective sequence. In some embodiments, the CRISPR enzyme directs the cleavage of one or both strands within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs of the first or last nucleotide of an objective sequence. In some embodiments, a vector encodes a CRISPR enzyme that mutates with respect to a corresponding wild-type enzyme so that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide that contains an objective sequence. For example, a substitution of aspartate to alanine (D10A) in the catalytic domain RuvC I of Cas9 de S. pyogenes co-converts Cas9 from a nuclease that cleaves both strands to a nickasa (cleaves a single strand). Other examples of mutations that make Cas9 a nickasa include, without limitation, H840A, N854A and N863A. In some embodiments, a Cas9 nickasa may be used in conjunction with sequence or guide sequences, for example, two guide sequences, which target transcribed and complementary strands respectively of the DNA target. This combination allows both strands to be nickened and used to induce NHEJ. Applicants have demonstrated (data not shown) the efficacy of two nickasa targets (ie, the RNAs set as targets in the same position but for different strands of DNA) in the induction of mutagenic NHEJ. A single nickasa (Cas9-D10A with a single RNAsg) is unable to induce NHEJ and create insertions or deletions but Applicants have demonstrated in double nickasa (Cas9-D10A and two RNAss targeted at different strands in the same position). it can do in human embryonic stem cells (the hESC). The efficacy is approximately 50% nuclease (ie, regular cas9 without mutation 010) in hESC
As a further example, two or more Cas9 catalytic domains (RuvC 1, RuvC 11 and RuvC 111) may mutate to produce a mule Cas9 that substantially lacks all DNA cleavage activity. In some embodiments, a D10A mutation is combined with one or more of the H840A, N854A or N863A mutations to produce a Cas9 enzyme that substantially lacks all DNA cleavage activity. In some embodiments, a CRISPR enzyme is considered to be substantially lacking in any DNA cleavage activity when the DNA cleavage activity of the mutated enzyme is less than about 25%, 10%, 5%, 1%, 0.1% , 0.01% or less with respect to its non-mutated form. Other mutations may be useful; in the case that Cas9 or another CRISPR enzyme is from a species other than S. pyogenes, mutations can be made in the corresponding amino acids to achieve similar effects
In some embodiments, an enzyme coding sequence encoding a CRISPR enzyme is codon optimized for expression in particular cells, such as eukaryotic cells. The euca riot cells can be those of, or from, a particular organism, such as a mammal, including but not limited to, human, mouse, rat, rabbit, dog or primate not to be human. In general, codon optimization refers to a method for modifying a nucleic acid sequence to improve expression in host cells of interest by replacing at least one codon (eg "about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, fifty or more cadooes) of the natural sequence with codons that are used more frequently or most frequently in the genes of that host cell while maintaining the natural amino acid sequence. Various species have particular bias for certain cadooes of a particular amino acid. Codon bias (differences in the use of codons between organisms) is often correlated with the translation efficiency of messenger RNA (mRNA), which in turn is believed to depend on, among other things, the properties of the cadons that They are being translated and the availability of particular transfer RNA molecules (tRNAs). The predominance of selected tRNAs in a cell is in general a reflection of the codons most frequently used in peptide synthesis. Accordingly, genes can be adapted for optimal expression of genes in a given organism based on codon optimization. The use of cadon tables is readily available, for example, in the "Database of Use of Cadones", and these tables can be adapted in a number of ways. See Nakamura Y., et al. "Codon usage tabulated from! He intemational DNA sequence databases: status far! He year 2000" Nud. Acids Res. 28: 292 (2,000). Computer algorithms for codons are also available optimizing a particular sequence for expression in a particular host cell, such as Gene Forge (Aptagen; Jacobus, PAlo. In some embodiments, one or more cores (e.g., 1, 2, 3, 4 , 5, 10, 15, 20, 25, 50 or more or all codons) in a sequence encoding a CRISPR enzyme correspond to the most frequently used cadon for a particular amino acid
In some embodiments, a vector encodes a CRISPR enzyme that comprises one or more nuclear localization sequences (the NLS), such as about or more than about 1, 2, 3, 4, 5, 6, 7 , 8, 9, 10 or more NLS. In some embodiments, the CRISPR enzyme comprises about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS at or near the amino terminal, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS in or near the carboxy terminal or a combination of these (for example, one or more NLS in the amino terminal and one or more NLS in the carboxy terminal). When more than one NLS is present, each can be selected independently of the others, so that a single NLS can be present in more than one copy and / or together with another or other NLS more present in one or more copies. In a preferred embodiment of the invention, the CRISPR enzyme comprises at most 6 NLS. In some embodiments, an NLS near the Non-C-terminal is considered when the closest amino acid in the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain of the terminal. Typically, an NLS consists of one or more short sequences of positively charged lysines or arginines exposed on the surface of the protein, but other types of NLS are known. Non-limiting examples of the NLS include a sequence of NLS from: the NLS of the large T antigen of the SV40 virus, which has the amino acid sequence PKKKRKV; the nucleoplasmin NLS (eg, the bipartite nucleoplasmin NLS with the sequence KRPAATKKAGQAKKKK); c-myc NLS having the amino acid sequence PAAKRVKLD or RQRRNELKRSP; the NLS M9 of hRNPA1 having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY; the RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV sequence of the lBS domain of importin-alpha; the VSRKRPRP and PPKKARED sequences of the myoma T protein; the POPKKKPL sequence of human p53; the SALlKKKKKMAP sequence of mouse c-abl IV; the DRLRR and PKQKKRK sequences of the NS 1 influenza virus, the RKLKKKIKKLL sequence of the Hepatitis virus delta antigen; the REKKKFLKRR sequence of the mouse Mx1 protein; the KRKGDEVDGVDEVAKKKSKK sequence of the human poly (ADP-ribose) polymerase and the RKCLQAGMNLEARKTKK sequence of the steroid hormone receptor glucocorticoid (human)
In general, one or more NLSs are of sufficient strength to drive the accumulation of the CRISPR enzyme in a detectable amount in the nucleus of a eukaryotic cell. In general, the strength of nuclear localization activity may come from the number of NLS in the CRISPR enzyme, the particular NLS or used or a combination of these factors. The detection of accumulation in the nucleus can be performed by any suitable technique. For example, a detectable marker can be condensed to the CRISPR enzyme, so that the location within a cell can be visualized, such as together with a means to detect the location of the nucleus (for example, a specific spot for the nucleus just like DAP!) Examples of detectable markers include fluofescent proteins (such as Green or GFP fluorescent proteins; RFP; CFP) and epitope tags (HA tag, flag tag, SNAP tag). Cell nuclei can also be isolated from cells, the contents of which can then be analyzed by any suitable method to detect protein, such as immunohistochemistry, Western assay or enzyme activity assay. The accumulation in the nucleus can also be determined indirectly, such as by an assay for the effect of CRISPR complex formation (eg, assay for DNA cleavage or mutation in the target sequence or assay for expression activity of modified genes affected by CRISPR complex formation and / or CRISPR enzyme activity), when compared to a control 00 exposed to the enzyme or CRISPR complex or exposed to a CRISPR enzyme that lacks one or more NLS.
In general, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence for hybridization with the target sequence and directs the specific binding of the sequence of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is approximately or more than about 50%, 60%, 75%, 80%, 85 %, 90%, 95%, 97.5%, 99% or more. The optimal alignment can be determined with the use of any suitable algorithm for aligning sequences, a non-limiting example of which includes the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, the algorithms based on the Burrows-Wheeler transformation (eg. , Surrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (1LIumina, San Diego, CA), SOAP (available at soap.genomics.org.cn) and Maq (available at maq.sourceforge.net) In some embodiments, a guide sequence is approximately or more than about 5, 10, 11, 12, 13, 14,15, 16, 17,18,19,20, 21,22, 23,24,25,26, 27, 28,29, 30, 35, 40, 45, 50, 75 or more nucleotides in length. In some embodiments, a guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12 or less nucleotides in length. The ability of a guide sequence to direct the specific binding of the sequence of a CRISPR complex to an objective sequence can be assessed by any suitable assay. For example, the components of a CRISPR system sufficient to form a CRISPR complex, including the guide sequence to be tested, can be provided to a host cell having the corresponding target sequence, such as by transfection with vectors encoding the components. of the CRISPR sequence, followed by an evaluation of the preferential cleavage within the target sequence, such as by Surveyor assay as described herein. Similarly, the cleavage of a target polynucleotide sequence can be evaluated in a test tube by providing the target sequence, the components of a CRISPR complex, including the guide sequence to be tested and a control guide sequence different from the test guide sequence and comparing the binding or cleavage rate in the target sequence between the reactions of the test and control guide sequence Other tests are possible, and will take place for experts in the field.
A guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within a genome of a cell. Exemplary target sequences include those that are unique in the target genome. For example, for Cas9 of S. pyogenes, a unique target sequence in a genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGG where NNNNNNNNNNNNXGG (N is A, G, T or CYX can be any) presents a single case in the genome. A single target sequence in a genome can include a S. pyogenes Cas9 target site of the MMMMMMMMMNNNNNNNNNNNXGG form where NNNNNNNNNNNXGG (N is A, G, T or C and X can be any) presents a single case in the genome. For Cas9 of CRISPR1 of S. thermophifus, a single target sequence in a geoome can include a Cas9 target site of the MMMMMMMMNNNNNNNNNNNNXXAGMW W where NNNNNNNNNNNNXXAGMW (N is A, G, T or C; X can be any and W is A or T) presents a single case in the genome A single target sequence in a genome can include a target S9 site of CRISPR1 of S thermophifus of the form MMMMMMMMMNNNNNNNNNNNXXAGMW where NNNNNNNNNNNXXAGMW (N is A, G, T or C; X can be any and W is A or T) has a single case in the genome. For the Cas9 de S. pyogenes, a single target sequence in a geoome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGGXG where NNNNNNNNNNNNXGGXG (N is A, G, T or CYX can be any) presents a single case in the genome.
A single target sequence in a genome can include a S. pyogenes Cas9 target site of the form MMMMMMMMNMNNNNNNNNNNXGGXG where NNNNNNNNNNNXGGXG (N is A, G, T or C and X may be any) presents a single case in the genome. In each of these sequences ~ M ~ it can be A, G, T or C and requires that it not be co-considered in the identification of a sequence as unique.
In some embodiments, a guide sequence was selected to reduce the degree of secondary structure within the guide sequence. The secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum free energy of Gibbs. An example of one such algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acid Res. 9 (1981), 133-148). Another example folding algorithm is the RNAfold webserver online, developed at the Institute for Theoretical Chemistry of the University of Vienna, using the prediction algorithm of centroid structure (see e.g., A. R Gruber et al., 2008, Cel ! 106 (1): 23-24 and PA Carr and GM Church, 2009, Nature 8iotechnology 27 (12): 1,151-62).
In general, a tracr pairing sequence includes any sequence that has sufficient complementarity with a tracr sequence to activate one or more of: (1) cleavage of a guide sequence flanked by tracr pairing sequences in a cell containing the corresponding tracr sequence and
(2) formation of a CRISPR complex in an objective sequence, in which the CRISPR complex comprises the tracr pairing sequence hybridized to the tracr sequence. In general, the degree of complementarity is with reference to the optimal alignment of the tracr pairing sequence and the tracr sequence, together with the shortening length of the two sequences. The optimal alignment can be determined by any suitable alignment algorithm and can also justify secondary structures, such as autocomplementarity within either the tracr sequence or the tracr pairing sequence. In some embodiments, the degree of complementarity between the tracr sequence and the tracr pairing sequence together with the length of the shortening of the two when aligned optimally is approximately or more than about 25%, 30%, 40%, 50% , 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or greater. Example illustrations of optimal alignment between a tracr sequence and a tracr pairing sequence are provided in Figures 128 and
138. In some embodiments, the tracr sequence is about or more than about S, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40 , 50 or more nucleotides in length. In some embodiments, the tracr sequence and tracr pairing sequence are contained within a single transcript, so that hybridization between the two produces a transcript with a secondary structure, such as a hairpin. Preferred loop-forming sequences for use in hairpin structures are four nucleotides in length and most preferably have the GAAA sequence. However, longer or shorter loop sequences can be used, as alternative sequences can be used. The sequences preferably include a nudeotide triplet (for example, AAA) and an additional nudeotide (for example C or G). Examples of loop-forming sequences include CAAA and AAAG. In one embodiment, the transcript or transcript polynucleotide sequence has at least two or more hairpins. In preferred embodiments, the transcript has two, three, four or five hairpins. In a further embodiment, the transcript presents at most five forks. In some embodiments, the single transcript also includes a transcription termination sequence; preferably this is a polyT sequence, for example six T nucleotides. An example illustration of said fork structure is provided in the lower portion of Figure 138, doode the portion of the sequence S 'of the final "N" and above the loop corresponds to the tracr pairing sequence and the portion of the sequence 3 'of the loop corresponds to the tracr sequence. Additional non-limiting examples of single polynucleotides comprising a guide sequence, a tracr pairing sequence and a tracr sequence are as follows (listed 5 'to 3'), where "N" represents a base of a guide sequence, the first block of Lowercase letters represent the tracr pairing sequence and the second block of lowercase letters sequence represents the tracr sequence and the final poly-T sequence represents the transcription terminator: (1) NNNNNNNNNNNNNNNNNNNNgttttlgtactctcaagatttaGAAAtaaatctlgcagaagctacaaagataaggctl catgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaa IIIII 1; (2) NNNNNNNNNNNNNNNNNNNNgttttlgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatca acaccctgtcattttatggcagggtgttttcgtlatttaa II! (3) NNNNNNNNNNNNNNNNNNNNgtltttgtactctcaGAAAtgcagaagctacaaagataaggctlcatgccgaaatca acaccctgtcattttatggcagggtgtl I ¡II¡; (4) NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAAtagcaagttaaaataaggctagtccgttatcaacttgaaaa agtggcaccgagtcggtgc IIIII 1; (5) NNNNNNNNNNNNNNNNNNNNgtlttagagctaGAAATAGcaagtlaaaataaggctagtccgttatcaactlgaa aaagtg IIIIII 1; And (6) NNNNNNNNNNNNNNNNNNNNgttttagagctagAAATAGcaagttaaaataaggctagtccgttatcaTTTTT TTI In some embodiments, sequences (1) to (3) are used together with Cas9 of CRISPR1 S. thermophilus. In some embodiments, sequences (4) to (6) are used together with Cas9 of S. pyogones. In some embodiments, the tracr sequence is a transcript separated from a transcript comprising the tracr pairing sequence (as illustrated in the upper portion of Figure 138).
In some embodiments, a recombination template is also provided. A recombination template can be a component of another vector as described herein, contained in a separate vector or provided as a separate polynucleotide. In some embodiments, a recombination template is designed to serve as a homologous recombination template, such as within or near a target sequence nickeated or cleaved by a CRISPR enzyme as part of a CRISPR complex. A template polynucleotide may be of any suitable length, such as about or more than about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1,000 or more nucleotides in length. In some embodiments, the template polynucleotide is complementary to a portion of a polynucleotide comprising the target sequence. When optimally aligned, a template polynucleotide can overlap with one or more nucleotides of a target sequence (e.g., about or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more nucleotides). In some embodiments, when a template sequence and a polynucleotide comprising an objective sequence are optimally aligned, the closest nucleotide of the template polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000 or more nucleotides of the target sequence.
In some embodiments, the CRISPR enzyme is part of a fusion protein comprising one or more heterologous protein domains (e.g., about or more than about 1,2,3,4,5,6,7,8,9 , 10 or more domains in addition to the enzyme CRISPR). A CRISPR enzyme fusion protein may comprise any additional protein sequence and optionally a linker sequence between any two domains. Examples of protein domains that can be condensed to a CRISPR enzyme include, without limitation, epitope tags, reporter gene sequences and protein domains with one or more of the following activities: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT) beta-galactosidase, betaglucuronidase, luciferase, green fluorescent protein (PFV), HcRed, OsRed, cyan ftuorescent protein (PFC) yellow fluorescent protein (PFAm) and autofluorescent proteins including blue fluorescent protein (PFA) A CRISPR enzyme can be condensed to a gene sequence that encodes a protein or a fragment of a protein that a DNA or other molecules cell molecules, including but not limited to (maltose binding protein (MBP), S tag, AON Lex A (OBO) binding domain fusions, AON GAL4 binding domain fusions and BP16 protein fusions of herpes simplex virus (HSV). Additional domains that may be part of a fusion protein comprising a CRISPR enzyme are described in US Pat. 20110059502. In some embodiments, a labeled CRISPR enzyme is used to identify the position of an objective sequence
In some aspects, the description provides methods comprising delivering one or more polynucleotides, such as or one or more vectors as described herein, one or more transcripts thereof, and one or proteins transcribed therein, to a host cell . In some aspects, the description further provides cells produced by such methods and organisms (such as animals, plants or fungi) that comprise or are produced from such cells. In some embodiments, a CRISPR enzyme is supplied together with (and optionally complexed with) a guide sequence to a cell. Conventional viral and non-viral based gene transfer methods can be used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to deliver components encoding nucleic acids of a CRISPR system to cells in culture or in a host organism. Non-viral vector delivery systems include AON plasmids, RNA (eg, a transcript of a vector described herein), naked nucleic acid and nucleic acid complexed with a delivery vehicle, such as a liposome. Viric vector delivery systems include AON and RNA viruses, which have episomal or integrated genomes after delivery to the cell. For a review of gene treatment procedures, see Anderson, Science
256: 808-813 (1992); Nabel & Felgner, TIBTECH 11: 211-217 (1993); Mitani & Caskey, TIBTECH 11166-166 (1,999); Dillon, TIBTECH 11 167-175 (1,993); Miller, Nature 357: 455-460 (1,992); Van Brunt, Biotechnology 6 (10) 1,149-1,154 (1,988); Vigne, Restorative Neurology and Neuroscience 8: 35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51 (1): 31-44 (1995); Haddada et al., In Current Topics in Microbiology and Immunology Ooerfler and B6hm (eds) (1995) and Yu et al., Gene Therapy 1. 13-26 (1994).
Non-viral nucleic acid delivery methods include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid conjugates: nucleic acid, naked DNA, artificial virions and increased AON absorption by agent. Lipofetion is described in e.g., U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355) And lipofeation reagents (e.g., Transfectam ™ and Lipofectin ™) are sold commercially. Cationic and neutral lipids that are suitable for effective recognition of polynucleotide receptor recognition include those of Felgner, International Patent WO 91117424; International Patent WO 91116024. The delivery can be to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vivo administration)
The preparation of lipid: nucleic acid complexes, including target fixed liposomes such as immunolipid complexes, is known to one skilled in the art (see, eg, Crystal, Science 270: 404410 (1999); Blaese et al. , Cancer Gene Ther. 2: 291-297 (1995); Behr et al., Bioconjugate Chem. 5: 382-389 (1,994); Remy et al., Bioconjugate Chem. 5: 647-654 (1,994); Gao et al., Gene Therapy 2: 710-722 (1995); Ahmad et al., Cancer Res. 52: 4,817-4,820 (1992); U.S. Pat. No. 4,186,183; 4,217,344; 4,235,871, 4,261,975;
4,485,054; 4,501,728; 4,774,085; 4,837,028 and 4,946,787).
The use of viral-based RNA or AON systems for the delivery of nucleic acids takes advantage of highly developed procedures to target a virus to specific cells in the body and to traffic the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or they can be used to treat cells in vitro and the modified cells can optionally be administered to patients (ex vivo). Conventional viral-based systems could include retroviral, lentivirus, adenoviral, adeno-associated and herpes simplex virus vectors for gene transfer. Integration into the host genome is possible with the methods of gene transfer of retroviruses, lentiviruses and adeno-associated viruses, often resulting in long-term expression of the inserted transgene. Additionally, high transduction efficiencies have been observed in many different cell types and target tissues.
The tropism of a retravirus can be modified by incorporating foreign envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that are capable of transducing or infecting non-dividing cells and typically produce high viral titers. The selection of a retroviral gene transfer system would therefore depend on the target tissue. Retroviral vectors consist of long terminal repeats that act in cis with packing capacity for up to 6 · 10 kb of foreign sequence. The minimal LTRs that act in cis are sufficient for replication and packaging of the vectors, which are then used to integrate the therapeutic gene into the target cell to provide permanent transgene expression. Widely used reliral vectors include those based on murine leukemia virus (MuLV), gibbon monkey leukemia virus (GaLV), simian immunodeficiency virus (VIS), human immunodeficiency virus (HIV) and combinations of themselves (see, e.g., Buchscher et al., J. Viral. 66: 2,731-2,739 (1992); Johann et al., J. Viral. 66: 1,635-1,640 (1992); Sommnerfelt et al., Viral 176: 58-59 (1990); Wilson et al., J. Virol. 63: 2,374-2,378 (1989); Miller et al., J. Virol 652,220-2,224 (1991); U.S. Patent PCTfUS94 (05700).
In applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors are capable of very high transduction efficiency in many cell types and do not require cell division. With such vectors, high titers and expression levels have been obtained. This vector can be produced in large quantities in a relatively simple system. AAV r adeno-associated virus vectors can also be used "to transduce cells with target nucleic acids, for example, in the in vitro production of nucleic acids and peptides and for in vivo and ex-gene treatment procedures. vivo (see, e.g., West et al., Virology 160: 38-47 (1987); U.S. Patent No. 4,797,368; International Patent WO 93f24641, Kotin, Human Gene Therapy 5: 793 -801 (1994); Muzyczka, J. Clin. Inves !. 94: 1,351 (1,994). The construction of recombinant AAV vectors is described in a number of publications, including US Pat. No. 5,173,414; Tralschin al., Mol. Cell Biol. 5: 3,251-3,260 (1985); Tratschin, et al., Mol. Cell Biol. 4: 2,072-2,081 (1,984); Hermonat & Muzyczka, PNAS 81 · 6,466-6,470 (1984) and Samulski et al., J. Virol. 633838-3828 (1989)
Typically packaging cells are used to form virus particles that are capable of infecting a host cell. Such cells include 293 cells, which package adenovirus, and 412 cells or PA317 cells, which package retroviruses. The viral vectors used in gene therapy are usually generated by producing a cell line that packages a nucleic acid vector into a viral particle. Vectors typically contain the minimum viral sequences required for packaging and subsequent integration into a host, other viral sequences being replaced by an expression cassette for the polynucleotide or polynucleotides to be expressed. Missing viral functions are typically supplied in trans through the packaging cell line. For example, AAV vectors used in gene therapy typically possess only ITR sequences of the AAV genome that are required for packaging and integration into the host genome. The viral DNA is packaged in a cell line, which contains an auxiliary plasmid that encodes the other AAV genes, ie rep and cap, but they lack ITR sequences. Cell line can also be infected with adenovirus as an auxiliary. The auxiliary virus activates AAV vector replication and AAV gene expression of the auxiliary plasmid. The auxiliary plasmid is not packaged in significant amounts due to the absence of ITR sequences. Contamination with adenovirus can be reduced by pO (for example, heat treatment to which adenovirus is more sensitive than AAV. Additional methods for the delivery of nucleic acids to cells are known to those skilled in the art. See, for example, US Pat. 20030087817
In some embodiments, a host cell is transinfected transiently or non-transiently with one or more vectors described herein. In some embodiments, a cell is transinfected as occurs naturally in an individual. In some embodiments, a cell that is transinfected is taken from an individual. In some embodiments, the cell is derived from cells taken from an individual, such as a cell line. A wide variety of cell lines of tissue cultures are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLa-S3, Huh1, HUh4, HUh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panc1, PC- 3, T F1, CTLL-2, C1R, Rat6, CV1, RPTE, A10, T24, J82, A375, ARH-77, Calu1, SW480, SW620, SKOV3, SK-UT, CaC02, P38801, SEM-K2, WEHI -231, HB56, TIB55, Jurkat, J45.01, LRMB, BcI-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, CaS-1, COS-6, CaS-M6A, monkey kidney epithelial BS-C-1, mouse embryo fibroblast BALBf 3T3, 3T3 Swiss, 3T3-L 1, human fetal fibroblasts 132-d5; Mouse fibroblasts 10.1, 293-T, 3T3, 721, 9L, A2780, A2780AoR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-1 cells, BEAS-2B, bEnd. 3, BHK-21, BR 293, BxPC3, C3H-10T1I2, C6f36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO ohfr -f-, COR -L23, COR-L23fCPR, COR-L23f5010, COR-L23fR23, COS-7, COV-434, CML T1, CMT, CT26, 017, oH82, OU145, DuCaP, EL4, EM2, EM3, EMT6fAR1, EMT6 / AR10. 0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepa1c1c7, HL-60, HMEC, HT-29, Jurkat, JY cells, K562, Ku812, KCL22, KG1, KY01, LNCap, MaMel 1-48, MC-38, MCF-7, MCF-10A, MoA-MB-231, MDA-MB-468, MDA-MB-435 cells , MDCK 11, MDCK 11, MORlO.2R, MONO-MAC 6, MTD-1A, MyEnd, NCI-H69fCPR, NCI-H69fLX10, NCI -H69fLX20, NCI -H69fLX4, NIH-3T3, NALM-1, NW-145 strains OPCN / OPCT, Peer, PNT-1A f PNT 2, RenCa, RIN-5F, RMAlRMAS, Saos-2, Sf9, SkBr3, T2, T-47o, T84, THP1, U373, U87, U937, VCaP cell lines , Vero cells, WM39, WT-49, X63 YAC-1, YAR and transgenic varieties thereof. Cell lines are available from a variety of sources known to those skilled in the art (see, e.g., The American Type Culture Collection (ATCC) (Manassus, Va.)). In some embodiments, a cell transinfected with one or more vectors described herein is used to establish a new cell line comprising one or more vector sequences. In some embodiments, an Iransinfection cell transiently infected with the components of a CRISPR system as described herein (such as by transient transfection of one or more vectOfes or RNA transfection) and modified by the activity of a CRISPR complex, It is used to establish a new cell line comprising cells that contain the modification but lack another exogenous sequence. In some embodiments, Iransinfected cells transiently or non-transiently with one or more vectors described herein or cell lines from said cells are used in the titration of one or more test compounds.
In some embodiments, one or more vectOfes described herein are used to produce a non-human transgenic animal or transgenic plant. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat or rabbit. In some embodiments, the organism or individual is a plant. In some embodiments, the organism or individual or plant is an algae. Methods for producing transgenic plants and animals are known in the art and generally begin with a method of cell transfection, as described herein. Transgenic animals are also provided such as transgenic plants, especially crops and algae. The transgenic animal or plant may be useful in applications outside providing a disease model. These may include food or feed production by expression of, for example, higher levels of protein, carbohydrate, nutrient or vitamins than would normally be observed in the natural type. With respect to this, transgenic plants are preferred, especially legumes and tubers and animals, especially mammals such as cattle (cows, sheep, goats and pigs), but also poultry and edible insects.
Transgenic algae or other plants such as rapeseed may be useful in particular in the production of vegetable oils or biofuels such as alcohols (especially methanol and ethanol), for example. These can be achieved by engineering to express or overexpress high levels of oil or alcohols for use in the oil or biofuel industries
In one aspect, the description provides methods for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method comprises allowing a CRISPR complex to bind to the target polynucleotide to effect the cleavage of said target polynucleotide thereby modifying the target polynucleolide, wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence hybridized to a target target sequence within said target polynucleotide, wherein said guide sequence is linked to a tracr pairing sequence that in turn hybridizes to a sequence. tracr
In one aspect, the description provides a method for modifying the expression of a polynucleotide in a eukaryotic cell. In some embodiments, the method comprises allowing a CRISPR complex to bind to the polynucleotide so that said binding results in increased or decreased expression of said polynucleotide; wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guiding sequence hybridized to a target-targeted sequence within said polynucleotide, wherein said guiding sequence is linked to a matching sequence bringing in turn hybridizing to a bringing sequence. .
With the recent advances in quantitative genetics, the abili ty to use CRISPR-Cas systems to carry out gene editing and efficient and cost effective manipulation will allow rapid selection and comparison of unique and multiplexed genetic manipulations to transform such genomes for improved production and improved baits. In this regard, reference is made to US Patents and Publications. US Pat. No. 6,603,061 -Agrobacterium-Mediated Plant Transformation Method; U.S. Patent No. 7,868,149 -Plant Genome Sequences and Uses Thereof and US Pat. 2009f0100536 -Transgenic Plants with Enhanced Agronomic Traits. Morrell el al "Crop genomics: advances and applications" Na! Rev Gene !. December 29, 2011, 13 (2): 85-96. In an advantageous embodiment of the invention, the CRISPRlCas9 system is used to achieve micro algae engineering (Example 14). Accordingly, the reference herein to animal cells can also be applied, mutatis mutandis, to plant cells unless it is evident otherwise.
In one aspect, the description provides methods for modifying a targ et polynucleotide in a eukaryotic cell, which can be in vivo, ex vivo or in vitro. In some embodiments, the method comprises sampling a cell or cell population of a human or non-human animal. or plant (including microalgae) and modify the cell or cells. Cultivation can take place in any ex vivo phase. The cell or cells can be reintroduced even into the non-human animal or plant (including microalgae)
In one aspect, the description provides cases containing one or more of any of the elements described in the above methods and compositions. In some embodiments, the case comprises a vector system and instructions for using the case. In some embodiments, the vector system comprises (a) a first regulatory element operably linked to a tracr pairing sequence and one or more insertion sites for inserting a guide sequence upstream of the bring pairing sequence, in which when is expressed, the guiding sequence directs the specific binding of the sequence of a CRISPR complex to a target sequence in a eukaryotic cell, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) the guiding sequence that hybridizes to the target sequence and (2) the tracr pairing sequence that hybridizes to the sequence bringing yfo (b) a second bound regulatory element operably to an enzyme coding sequence encoding said CRISPR enzyme comprising a nuclear localization sequence. The elements can be provided individually or in combinations and can be provided in any suitable container, such as a vial, a bottle or a tube. In some embodiments, the kit includes instructions in one or more languages, for example in more than one language.
In some embodiments, a kit comprises one or more reagents for use in a process using one or more of the elements described herein. The reagents may be provided in any suitable container. For example, a case may provide one or more reaction or storage buffers. The reagents can be provided in a form that is usable in a particular assay or in a form that requires the addition of another or other components before use (for example, in concentrated or lyophilized form). A buffer may be any buffer, including but not limited to sodium carbonate buffer, a sodium bicarbonate buffer, a buffer of bFate, a Tris buffer, a MOPS buffer, an HEPES buffer and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit comprises one
or more or ligonucleotides corresponding to a guide sequence for insertion into a vector so that it is operably linked to the guide sequence and a regulatory element In some embodiments, the kit comprises a homologous recombination template polynucleotide.
In one aspect, the description provides methods for using one or more elements of a CRISPR system. The CRISPR complex of the invention provides an effective means of modifying an objective polynucleotide. The CRISPR complex of the invention has a wide variety of utility including modification (eg, suppression, insertion, translocation, inactivation, activation) of a target polynucleotide in a multiplicity of cell types. As such, the CRISPR complex of the invention has a broad spectrum of applications in, for example, gene treatment, drug research, disease diagnosis and prognosis. An exemplary CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide. The guide sequence is linked to a tracr pairing sequence, which in turn hybridizes to a tracr sequence.
The target polynucleotide of a CRISPR complex can be any endogenous or exogenous polynucleotide to the eukaryotic cell. For example, the target polynucleotide may be a polynucleotide that resides in the nucleus of the eukaryotic cell. The target polynucleotide can be a sequence that encodes a germic product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). Without wishing to be bound by theory, it is believed that the objective sequence should be associated with a PAM (adjacent proto-spacer unit); that is, a short sequence recognized by the CRISPR complex The precise sequence and length requirements for the PAM differ depending on the CRISPR enzyme used, but the PAMs are typically 2-5 base pair sequences adjacent to the proto-spacer (that is, the target sequence) Examples of PAM sequences are provided in the examples section below and the expert can identify additional PAM sequences for use with a given CRISPR enzyme.
The target polynucleotide of a CRISPR complex can include a series of genes and polynucleotides associated with diseases as well as genes and polynucleotides associated with the biochemical signaling pathway.
Examples of target polynucleotides include a sequence associated with a biochemical signaling pathway, eg, a gene or polynucleotide associated with a biochemical signaling pathway. Examples of target polynucleotides include a gene or polynucleotide associated with a disease. A gene or polynucleotide "associated with a disease" refers to any gene or polynucleotide that is providing transcription or translation products at an abnormal level or in an abnormal form in cells from tissues affected by a disease compared to tissues or cells of a control not of empennedad. It may be a gene that expresses itself at an abnormally high level; It may be a gene that expresses itself at an abnormally low level, where the modified expression correlates with the onset and / or progression of the disease. A gene associated with the disease also refers to a gene that has mutation or mutations or genetic variation that is directly responsible or linked to imbalance with a gene, or genes, that is responsible for the etiology of a disease. Transcribed or translated products may be known or unknown and may be at a normal or abnormal level.
Examples of genes and polynucleotides associated with a disease are available at the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) And National Center for Biotechnology Information, National Library of Medicine (Betesda, Md.l, available in the World Computer Network
Examples of genes and polynucleotides associated with a disease are listed in Tables A and B. Information
5 Specific disease is available at the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and National Center for Biotechnology Information, National Library of Medicine (Betesda, Md.l, available on the Computer Network Global Examples of genes and polynucleotides associated with a biochemical signaling path are listed in Table C.
Mutations in these genes and pathways may result in the production of inappropriate proteins or
10 proteins in inappropriate amounts that affect function. Such genes, proteins and pathways can be the target polynucleotide of a CRISPR complex.
Table A Table B
<dl><dt>DISORDERS DISORDERS </dt><dd>GEN (ES) </dd></dl>
<dl><dt>Neoplasia </dt><dd>PTEN; ATM; ATR; EGFR; ERBB2; ERBB3; ERBB4;</dd></dl>
<dl><dt>Notch1; Notch2; Notch3; Notch4; AKT; AKT2; AKT3; HIF;</dt><dd /></dl>
<dl><dt>HIF1a; HIF3a; Met; HRG; Bcl2; PPAR alpha; PPAR</dt><dd /></dl>
<dl><dt>gamma; WT1 (Tumor Wilms l; Fam ilia FGF Receptor</dt><dd /></dl>
<dl><dt>members (5 members: 1, 2, 3, 4, 5); CDKN2a; APC; RB</dt><dd /></dl>
<dl><dt>(retinoblastoma); MEN1; VHL; BRCA1; BRCA2; AR</dt><dd /></dl>
<dl><dt>(Androgen receptor); TSG101; IGF; IGF receiver; Igf1 (4</dt><dd /></dl>
<dl><dt>variants); Igf2 (3 variants); Igf 1 receptor, Igf 2 receptor;</dt><dd /></dl>
<dl><dt>Bax; Bc12; Caspas family (9 members ·</dt><dd /></dl>
<dl><dt>1, 2, 3,4,6, 7, 8, 9,12); Kras; Apc</dt><dd /></dl>
<dl><dt>Degeneration </dt><dd>Abcr; Nc12; Cc2; cp (ceruloplasmin); Timp3; cathepsin D;</dd></dl>
<dl><dt>Age-related macular </dt><dd>Vine Go; Ccr2</dd></dl>
<dl><dt>Schizophrenia </dt><dd>Neuregulin1 (Nrg1); Erb4 (Neuregulin receptor);</dd></dl>
<dl><dt>Complexin1 (Cplx1); Tph1 Tryptophan hydroxylase; Tph2</dt><dd /></dl>
<dl><dt>Tryptophan hydroxylase 2; Neurexin 1, GSK3; GSK3a;</dt><dd /></dl>
<dl><dt>GSK3b </dt><dd /></dl>
<dl><dt>Disorders </dt><dd>5-HIT (Slc6a4); CDMT; DRD (Drd1a); SLC6A3; DADAIST;</dd></dl>
<dl><dt>DTNBP1; Dao (Oa01)</dt><dd /></dl>
<dl><dt>Trinucleotide Repeat </dt><dd>HIT (Huntington's Dx); SBMAlSMAX1 fAR (Kennedy's</dd></dl>
<dl><dt>Disorders </dt><dd>Dx); FXN / X25 (Friedrich's Ataxia); ATX3 (Machado</dd></dl>
<dl><dt>Joseph's Ox); ATXN1 and ATXN2 (ataxia</dt><dd /></dl>
<dl><dt>spinocerebellar); DMPK (Myotonic Dystrophy); Atrophin-1 and Atn1</dt><dd /></dl>
<dl><dt>(DRPLA Dx); CBP (Creb-BP - global instability); VLDLR</dt><dd /></dl>
<dl><dt>(Alzheimer's); Atxn7; Atxn10</dt><dd /></dl>
<dl><dt>Fragile syndrome X </dt><dd>FMR2; FXR1; FXR2; mGLUR5</dd></dl>
<dl><dt>Disorders </dt><dd>APH-1 (alpha and beta); Preseniline (Psen1); nicastrin</dd></dl>
<dl><dt>Secrelase related </dt><dd>(Ncstn); PEN-2</dd></dl>
<dl><dt>Others </dt><dd>Nos1, Parp1, Nat1, Nat2 </dd></dl>
<dl><dt>Prion related disorders </dt><dd>P ", </dd></dl>
<dl><dt>ALS </dt><dd>SOD1, ALS2; STEX; FUS; TAROBP; VEGF (VEGF-a;</dd></dl>
<dl><dt>VEGF-b; VEGF-c)</dt><dd /></dl>
<dl><dt>DISORDERS DISORDERS </dt><dd>GEN (ES) </dd></dl>
<dl><dt>Drug addiction </dt><dd>Prkce (to alcohol); Drd2; Ord4; ABAT (alcohol); GRIA2;</dd></dl>
<dl><dt>Grm5; Grin1; Htr1b; Grin2a; Drd3; Pdyn; Gria1 (to alcohol)</dt><dd /></dl>
<dl><dt>Autism </dt><dd>Mecp2; BZRAP1, MOGA2; SemaSA; Neurexin 1; Fragile X</dd></dl>
<dl><dt>(FMR2 (AFF2); FXR1, FXR2; Mglur5) </dt><dd /></dl>
<dl><dt>Alzheimer disease </dt><dd>E1; CHIP; UCH; UBB; Tau; LRP; PICALM _; Clusterin; PS1;</dd></dl>
<dl><dt>SORL 1; CR1, Vid Ir; Uba1; Uba3; CHIP28 (Aqp1,</dt><dd /></dl>
<dl><dt>Aquaporin 1); Uch11; Uch l3; APP</dt><dd /></dl>
<dl><dt>Swelling </dt><dd>IL-10; IL-1 (IL-1a; IL-1b); IL-13; IL-17 (IL-17a (CTLA8); IL</dd></dl>
<dl><dt>17b; IL-17c; IL-17d; IL-17f); 11-23; Cx3cr1; ptpn22; TNFa;</dt><dd /></dl>
<dl><dt>NOD2ICARD15 for IBD; IL-6; IL-12 (IL-12a; IL-12b);</dt><dd /></dl>
<dl><dt>CTLA4; Cx3cl1</dt><dd /></dl>
<dl><dt>Parkinson's disease </dt><dd>x-Synuclein; DJ -1; LRRK2; Parkin; PINK1</dd></dl>
Diseases and disorders
Anemia (CDAN1, CDA1, RPS19, DBA, PKLR, PK1, NT5C3, BUMPH1, PSN1, RHAG, blood and coagulation
RH50A, NRAMP2, SPTB, ALAS2, ANH1, ASB, ABCB7, ABC7, ASAT); naked lymphocyte syndrome (TAPBP, TPSN, TAP2, ABCB3, PSF2, RING11, MHC2TA, C2TA, RFX5, RFXAP, RFX5), Bleeding disorders (TBXA2R, P2RX1, P2X1); Factor H and factor H type -1 (HF1, CFH, HUS); Factor V and factor VIII (MCF02); Factor VII deficiency (F7); Factor X deficiency (F10); Factor XI deficiency (F11); Factor XII deficiency (F12, HAF); Factor XJIIA deficiency (F13A1, F13A); Factor XIIIB deficiency (F13B); Anemia Fanconi (FANCA, FACA, FA1, FA, FAA, FAAP9S, FAAP90, FLJ34064, FANCB, FANCC, FACC, BRCA2, FANC01, FANC02, FANCO, FACO, FAO, FANCE, FACE, FANCF, XRCC9, FANCG, BRIP1, BACH1 , FANCJ, PHF9, FANCL, FANCM, KIAA1596); Hemophagocytic lymphohistiocytosis disorders (PRF1, HPLH2, UNC13D, MUNC13-4, HPLH3, HLH3, FHL3); Hemophilia A (F8, F8C, HEMA); Hemophilia B (F9, HEMB), bleeding disorders (PI, An, F5); deficiencies and disorders of leuoocia (ITGB2, C018, LCAMB, LAO, EIF2B1, EIF2BA, EIF2B2, EIF2B3, EIF2BS, LVWM, CACH, CLE, EIF2B4); depranocltic anemia (HBB); Thalassemia (HBA2, HBB, HBD, LCRB, HBA1)
<dl><dt>Cellular detachment and </dt><dd>Non-Hodgkin lymphoma of cell B (BCL7A, BCL7); Leukemia (TAL 1, TCL5, SCL,</dd></dl>
<dl><dt>diseases and disorders </dt><dd>TAL2, FLT3, NBS1, NBS, ZNFN1A1, IK1, LYF1, HOXD4, HOX4B, BCR, CML, PHL, </dd></dl>
<dl><dt>oncological </dt><dd>ALL, ARNT, KRAS2, RASK2, GMPS, AF10, ARHGEF12, LARG, KlAA0382, CALM, </dd></dl>
<dl><dt>CLTH, CEBPA, CEBP, CHIC2, BTL, FLT3, KIT, PBT, LPP, NPM1, NUP214, D9S46E, </dt><dd /></dl>
<dl><dt>CAN, CAIN, RUNX1, CBFA2, AML1, WHSC1L 1, NS03, FLT3, AF1Q, NPM1, NUMA1, </dt><dd /></dl>
<dl><dt>ZNF 145, PLZF, PML, MYL, STAT5B, AF 10, CALM, CLTH, ARL11, ARLTS1, P2RX7, </dt><dd /></dl>
<dl><dt>P2X7, BCR, CML, PHL, ALL, GRAF, NF1, VRNF, WSS, NFNS, PTPN11, PTP2C, </dt><dd /></dl>
<dl><dt>SHP2, NS1, BCL2, CCN01, PRA01, BCL1, TCRA, GATA1, GF1, ERYF1, NFE1, </dt><dd /></dl>
<dl><dt>ABL 1, NQ01, IOM, NMOR1, NUP214, 09S46E, CAN, CAl N). </dt><dd /></dl>
Inflammation and diseases
AIDS (KIR3DL 1, NKAT3, NKB1, AMB11, KIR3DS1, IFNG, CXCL 12, SDF1); syndrome and disorders
auloimmune lymphoproliferative (TNFRSF6, APT1, FAS, CD9S, ALPS1A); immunorelated
combined immunodeficiency, (IL2RG, SCIDX1, SCIDX, IMD4); HIV-1 (CCL5, SCYA5, D17S136E, TCP228), susceptibility or infection by V1H (IL 10, CSIF, CMKBR2, CCR2, CMKBRS, CCCKR5 (CCRS)); Immunodeficiencies (C03E, C03G, AICOA, AIO, HIGM2, TNFRSF5, C040, UNG, OGU, HIGM4, TNFSFS, C040LG, HIGM1, IGM, FOXP3, IPEX, AIID, XPID, P1DX, TNFRSF14B, TACI); Inflation (IL10, IL-1 (IL-1a, IL-1b), IL-13, IL-17 (IL-17a (CTLA8), IL-17b, IL-17c, IL-17d, IL-17f), 1123 , Cx3cr1, ptpn22, TNFa, NOO2 / CAR01S for IBO, IL-6, IL-12 (IL-12a, IL-12b), CTLA4, Cx3c11); Severe combined immunodeficiencies (SCIOs) (JAK3, JAKL, OCLRE1C, ARTEMIS, SCIOA, RAG1, RAG2, ADA, PTPRC, C04S, LCA, IL7R, C030, T3D, IL2RG, SCIDX1, SCIDX, IMD4)
<dl><dt>Diseases and disorders </dt><dd>Neu amyloid clothing (TTR, PAlB); Amyloidosis (APOA1, APP, AAA, CVAP, AD1,</dd></dl>
<dl><dt>metabolic, hepatic, </dt><dd>GSN, FGA, LYZ, TIR, PAl B); Cirrhosis (KRT18, KRT8, C1RH1A, NAIC, TEX292,</dd></dl>
<dl><dt>renal and protein </dt><dd>K1M1988); cystic fibrosis (CFTR, ASCC7, CF, MRP7); glycogen storage diseases (SlC2A2, GlUT2, G6PC, G6PT, G6PT1, GAA, LAMP2, LAMPB, AGl, GOE, GBE1, GYS2, PYGl, PFKM); hepatic adenoma, 142330 (TCF1, HNF1A, MODY3), hepatic failure, early onset and neurological disorder (SCOD1, SC01), hepatic lipase deficiency (UPC), Hepatoblastoma, cancer and carcinomas (CTNNB1, PDGFRl, PDGRl, PRlTS, AXIN1, AXIN, CTNNB1, TP53, P53, lFS1, IGF2R, MPRI, MET, CASP, MCH5; medullary cystic kidney disease (UMOD, HNFJ, FJHN, MCKD2, AOMCKD2); Fen ilketonuria (PAH, PKU1, QDPR, DHPR, PTS); Polycystic kidney and liver disease (FCYT, PKH01, ARPKD, PKD 1, PKD2, PKD4, PKDTS, PRKCSH, G19P1, PClD, SEC63).</dd></dl>
<dl><dt>diseases and disorders </dt><dd>Becker muscular dystrophy (DMD, BMD, MYF6), Duchenne muscular dystrophy </dd></dl>
<dl><dt>skeletal muscle I </dt><dd>(DMD, BMD); Emery-Dreifuss muscular dystrophy (lMNA, lMN1, EMD2, FPlD, CMD1A, HGPS, lGMD1B, lMNA, LMN1, EMD2, FPlD, CMD1A); Facioscapulohumeral muscular dystrophy (FSHMD1A, FSH01A); Muscular dystrophy (FKRP, MDC1C, lGMD21, LAMA2, LAMM, LARGE, KlM0609, MDC1D, FCMD, ITID, MYOT, CAPN3, CANP3, OYSF, lGMD2B, SGCG, lGMD2C, DMDA1, SCG3, SGCA, DMG2, SGCA, DMG2, SGCA, DMG, ADG2, DMG, ADG2, DMG, ADG2 , SGCB, lGMD2E, SGCO, SGD, lGMD2F, CMD1 L, TCAP, lGMD2G, CMD1N, TRI M32, HT2A, lGMD2H, FKRP, MDC1C, lGMD21, TTN, CMD1G, TMD, lGMD2D1, P3, SEL3, P3, P3D1, P3, V3, P3 , RSMD1, PLEC1, PlTN, EBS1); Osteopetrosis (LRP5, BMN01, LRP7, lR3, OPPG, VBCH2, ClCN7, ClC7, OPTA2, OSTM1, Gl, TCIRG1, TIRC7, OC116, OPTB1); Muscular atrophy (VAPB, VAPC, AlSB, SMN1, SMA1, SMA2, SMA3, SMA4, BSCl2, SPG17, GARS, SMAD1, CMT2D, HEXB, IGHMBP2, SMUBP2, CATF1, SMARD1)</dd></dl>
<dl><dt>diseases and disorders </dt><dd>AlS (S001, AlS2, STEX, FUS, TARDBP, VEGF (VEGF-a, VEGF · b, VEGF-c); </dd></dl>
<dl><dt>neurological and neuronal </dt><dd>Alzheimer's disease (APP, AAA, CVAP, AD1, APOE, AD2, PSEN2, A04, STM2, APBB2, FE65l1, NOS3, PLAU, URK, ACE, DCP1, ACE1, MPO, PACIP1, PAXIP1 l, PTIP, A2M, BlMH , BMH, PSEN1, AD3); Autism (Mecp2, BZRAP1, MDGA2, Sema5A, Neu rexina 1, GL01, MECP2, RIT, PPMX, MRX16, MRX79, NlGN3, NlGN4, KIAA1260, AUTSX2); Fragile X syndrome (FMR2, FXR1, FXR2, mGlUR5); Huntington's disease and disease-like disorders (HO, 1T15, PRNP, PRIP, JPH3, JP3, HOL2, TBP, SCA17); Parkinson's disease (NR4A2, NURR1, NOT, TINUR, SNCAIP, TBP, SCA17, SNCA, NACP, PARK1, PARK4, DJ1, PARK7, lRRK2, PARK8, PINK1, PARK6, UCHl1, PARK5, SNCA, NACP, PARK1, PARK4, PRKN, PARK2, PDJ, DBH, NDUFV2); Rett syndrome (MECP2, RIT, PPMX, MRX16, MRX79, COKl5, STK9, MECP2, RIT, PPMX, MRX16, MRX79, x-Synuclein, DJ-1); Schizophrenia (Neuregulin1 (Nrg1), Erb4 (Neuregulin receptor), Complexin1 (Cplx1), Tph1 Tryptophan Hydroxylase, Tph2, Tryptophan Hydroxylase 2, Neurexin 1, GSK3, GSK3a, GSK3b, 5-HIT (Slc6a) (Slc6a (COM6) ), SlC6A3, DAOA, OTNBP1, Oao (Oao1)); Secretase-related disorders (APH-1 (a lfa and beta), Presenilin (Psen1), nicastrin, (Nestn), PEN -2, Nos1, Parp1, Nat1, Nat2); Disorders Trinucleotide repetition (HIT (Huntington's Ox), SBMAlSMAXlIAR (Kennedy's Ox), FXN / X25 (Friedrich's Ataxia), ATX3 (Machado-Joseph's Ox), ATXN1 and ATXN2 (Spinocerebellar ataxias), DMPK (Myotonic Dystrophy), Alotin) 1 and Atn1 (DRPLA Dx), CBP (Creb-BP-global instability), VLDlR (Alzheimer's), Atxn7, Atxn10). </dd></dl>
<dl><dt>eye diseases and disorders </dt><dd>age-related rnacular degeneration (Abcr, Ccl2, Cc2, cp (ceruloplasrnin), Timp3, cathepsinD, Vldlr, Ccr2); Cataract (CRYAA, CRYA1, CRYBB2, CRYB2, PITX3, BFSP2, CP49, CP47, CRYM, CRYA1, PAX6, AN2, MGDA, CRYBA1, CRYB1, CRYGC, CRYG3, CCl, UM2, MP19, CRYGO, CRYG, CRYG, FSP2 CP47, HSF4, CTM, HSF4, CTM, MIP, AQPO, CRYAB, CRYA2, CTPP2, CRYBB1, CRYGD, CRYG4, CRYBB2, CRYB2, CRYGC, CRYG3, CCl, CRYAA, CRYA1, GJAB, CX503, CAX CZP3, CAE3, CCM1, CAM, KRIT1); dystrophy and turbidity of the cornea (APOA1, TGFBI. CSD2, CDGG1, CSD, BIGH3, CDG2, TACSTD2, TROP2, M1S1, VSX1, RINX, PPCO, PPD, KTCN, COl8A2, FECO, PPCD2, PIP5K3, CFO); Congenital flat cornea (KERA, CNA2); Glaucoma (MYOC, TIGR, GlC1A, JOAG, GPOA, OPTN, GlC1E, FIP2, HYPl, NRP, CYP1B1, GlC3A, OPA1, NTG, NPG, CYP1 81, GlC3A); congenital leuros amaurosis (CRB1, RP12, CRX, CORD2, CRD, RPGRIP1, lCA6, CORD9, RPE65, RP20, AIPL 1, lCA4, GUCY2D, GUC2D, lCA1, COR06, ROH12, lCA3); macular dystrophy (ElOVl4, AOMD, STG02, STG03, ROS,</dd></dl>
IRP7, PRPH2, PRPH, AVMD, AOFMD, VMD2)
Table C
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>PI3KJAKT signaling </dt><dd>PRKCE; lTGAM; ITGAS; IRAK 1; PRKAA2; EIF2AK2;</dd></dl>
<dl><dt>PTEN; EI F4E; PRKCZ; GRK6; MAPK1; TSC1; PLK1;</dt><dd /></dl>
<dl><dt>AKT2; IKBKB; PIK3CA; CDKB; COKN1B; NFKB2; BCL2;</dt><dd /></dl>
<dl><dt>PIK3CB; PPP2R1A; MAPK8; BCL2L 1; MAPK3; TSC2;</dt><dd /></dl>
<dl><dt>lTGA1, KRAS; EIF4EBP1, REtA; PRKCO; NOS3;</dt><dd /></dl>
<dl><dt>PRKAA1; MAPK9; CDK2; PPP2CA; PIM1; lTGB7;</dt><dd /></dl>
<dl><dt>YWHAZ; ILK; TP53; RAF1; IKBKG; RELB; OYRK1A;</dt><dd /></dl>
<dl><dt>CDKN1A; ITGB1; MAP2K2; JAK1; AKT1; JAK2; PIK3R1;</dt><dd /></dl>
<dl><dt>CHUK; POPK1, PPP2R5C; CTNNB1, MAP2K1, NFKB1,</dt><dd /></dl>
<dl><dt>PAK3; ITGB3; CCN01; GSK3A; FRAP1, SFN; ITGA2;</dt><dd /></dl>
<dl><dt>TIK; CSNK1A1; BRAF; GSK3B; AKT3; FOX01; SGK;</dt><dd /></dl>
<dl><dt>HSP90AA1; RPS6KB1</dt><dd /></dl>
<dl><dt>ERKJMAPK signaling </dt><dd>PRKCE; lTGAM; ITGAS; HSPB1; IRAK1; PRKAA2;</dd></dl>
<dl><dt>EIF2AK2; RAC1, RAP1A; TLN1, EIF4E; ELK1, GRK6;</dt><dd /></dl>
<dl><dt>MAPK1; RAC2; PLK1; AKT2; PIK3CA; COKB; CREB1;</dt><dd /></dl>
<dl><dt>PRKCJ; PTK2; FOS; RPS6KA4; PIK3CB; PPP2R1A;</dt><dd /></dl>
<dl><dt>PIK3C3; MAPK8; MAPK3; ITGA1, ETS1; KRAS; MYCN;</dt><dd /></dl>
<dl><dt>EIF4EBP1, PPARG; PRKCO; PRKCA1; MAPK9; SRC;</dt><dd /></dl>
<dl><dt>CDK2; PPP2CA; PIM1; PIK3C2A; ITGB7; YWHAZ;</dt><dd /></dl>
<dl><dt>PPP1CC; KSR1, PXN; RAF1, FYN; OYRK1A; ITGB1,</dt><dd /></dl>
<dl><dt>MAP2K2; PAK4; PIK3R1; STAT3; PPP2R5C; MAP2K1;</dt><dd /></dl>
<dl><dt>PAK3; ITGB3; ESR1, ITGA2; MYC; TIK; CSNK1A1,</dt><dd /></dl>
<dl><dt>CRKL; BRAF; ATF4; PRKCA; SRF; STAT1; SGK</dt><dd /></dl>
<dl><dt>Signaling </dt><dd>RAC1; TAF4B; EP300; SMA02; TRAF6; PCAF; ELK1;</dd></dl>
<dl><dt>Glucocorticoid Receptor </dt><dd>MAPK1; SMAD3; AKT2; IKBKB; NCOR2; UBE21;</dd></dl>
<dl><dt>PIK3CA; CREB1; FOS; HSPA5; NFKB2; BCL2;</dt><dd /></dl>
<dl><dt>MAP3K14; STATSB; PIK3CB; PIK3C3; MAPKB; BCL2L 1,</dt><dd /></dl>
<dl><dt>MAPK3; TSC2203; MAPK10; NRIP1; KRAS; MAPK13;</dt><dd /></dl>
<dl><dt>RELAY; STAT5A; MAPK9; NOS2A; PBX1, NR3C1,</dt><dd /></dl>
<dl><dt>PIK3C2A; CDKN1C; TRAF2; SERPINE1; NCOA3;</dt><dd /></dl>
<dl><dt>MAPK14; TNF; RAF1; IKBKG; MAP3K7; CREBBP;</dt><dd /></dl>
<dl><dt>COKN1A; MAP2K2; JAK1; ILB; NCOA2; AKT1, JAK2;</dt><dd /></dl>
<dl><dt>PIK3R1; CHUK; STAT3; MAP2K1; NFKB1; TGFBR1;</dt><dd /></dl>
<dl><dt>ESR1, SMA04; CEBPB; JUN; AR; AKT3; CCL2; MMP1; STAT1, IL6; HSP90AA1</dt><dd /></dl>
<dl><dt>Axonal Guide Signaling </dt><dd>PRKCE; lTGAM; ROCK1, ITGAS; CXCR4; AOAM12;</dd></dl>
<dl><dt>IGF1; RAC1; RAP1A; EIF4E; PRKCZ; NRP1; NTRK2;</dt><dd /></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>ARHGEF7; SAINT; ROCK2; MAPK1; PGF; RAC2;</dt><dd /></dl>
<dl><dt>PTPN11; GNAS; AKT2; PIK3CA; ERBB2; PRKCI; PTK2;</dt><dd /></dl>
<dl><dt>CFL 1, GNAQ; PIK3CB; CXCL 12; PIK3C3; W NT11,</dt><dd /></dl>
<dl><dt>PRKD1, GNB2L 1, ABL1, MAPK3; ITGA1, KRAS; RHOA;</dt><dd /></dl>
<dl><dt>PRKCD; PIK3C2A; ITGB7; GLl2; PXN; VASP; RAF1;</dt><dd /></dl>
<dl><dt>FYN; ITGB1, MAP2K2; PAK4; ADAM17; AKT1; PIK3R1,</dt><dd /></dl>
<dl><dt>GU1, WNT5A; ADAM10; MAP2K1, PAK3; ITGB3;</dt><dd /></dl>
<dl><dt>CDC42; VEGFA; ITGA2; EPHA8; CRKL; RND1; GSK3B;</dt><dd /></dl>
<dl><dt>AKT3; PRKCA</dt><dd /></dl>
<dl><dt>Ephrin Receiver Signaling </dt><dd>PRKCE; ITGAM; ROCK1; ITGA5; CXCR4; IRAK 1;</dd></dl>
<dl><dt>PRKAA2; EIF2AK2; RAC 1; RAP1A; GRK6; ROCK2;</dt><dd /></dl>
<dl><dt>MAPK1; PGF; RAC2; PTPN11; GNAS; PLK1; AKT2;</dt><dd /></dl>
<dl><dt>DOK1, COK8; CREB1, PTK2; CFL 1; GNAQ; MAP3K14;</dt><dd /></dl>
<dl><dt>CXCL12; MAPK8; GNB2L 1; ABL1; MAPK3; ITGA1;</dt><dd /></dl>
<dl><dt>KRAS; RHOA; PRKCD; PRKAA1; MAPK9; SRC; CDK2;</dt><dd /></dl>
<dl><dt>PIM1, ITGB7; PXN; RAF1; FYN; OYRK1A; ITGB1;</dt><dd /></dl>
<dl><dt>MAP2K2; PAK4; AKT1; JAK2; STAT3; ADAM10;</dt><dd /></dl>
<dl><dt>MAP2Kl, PAK3; ITGB3; COC42; VEGFA; ITGA2;</dt><dd /></dl>
<dl><dt>EPHA8; TTK; CSNK1A1; CRKL; BRAF; PTPN1 3; ATF4;</dt><dd /></dl>
<dl><dt>AKT3; SGK</dt><dd /></dl>
<dl><dt>Signalization </dt><dd>ACTN4; PRKCE; ITGAM; ROCK1, ITGA5; IRAK1,</dd></dl>
<dl><dt>Acosta Cytoskeleton </dt><dd>PRKAA2; EIF2AK2; RAC1; INS; ARHGEF7; GRK6;</dd></dl>
<dl><dt>ROCK2; MAPK1, RAC2; PLK1, AKT2; PIK3CA; CDK8;</dt><dd /></dl>
<dl><dt>PTK2; CFL 1, PIK3CB; MYH9; OIAPH1, PIK3C3; MAPK8;</dt><dd /></dl>
<dl><dt>F2R; MAPK3; SLC9A1; ITGA1; KRAS; RHOA; PRKCD;</dt><dd /></dl>
<dl><dt>PRKAA1, MAPK9; COK2; PIM1, PIK3C2A; ITGB7;</dt><dd /></dl>
<dl><dt>PPP1CC; PXN; VIL2; RAF1; GSN; DYRK1A; ITGB1;</dt><dd /></dl>
<dl><dt>MAP2K2; PAK4; PIP5K1A; PIK3R1; MAP2K1; PAK3;</dt><dd /></dl>
<dl><dt>ITGB3; CDC42; APC; ITGA2; TTK; CSNK1A1; CRKL;</dt><dd /></dl>
<dl><dt>BRAF; VAV3; SGK</dt><dd /></dl>
<dl><dt>Signaling </dt><dd>PRKCE; IGF1; EP300; RCOR1; PRKCZ; HOAC4; TGM2;</dd></dl>
<dl><dt>huntington's disease </dt><dd>MAPK1; CAPNS1; AKT2; EGFR; NCOR2; SP1; CAPN2;</dd></dl>
<dl><dt>PIK3CA; HDAC5; CREB1, PRKCI; HSPA5; REST;</dt><dd /></dl>
<dl><dt>GNAQ; PIK3CB; PIK3C3; MAPKB; IGF1 R; PRK01,</dt><dd /></dl>
<dl><dt>GNB2L 1, BCL2L 1, CAPN1, MAPK3; CASP8; HDAC2;</dt><dd /></dl>
<dl><dt>HDAC7A; PRKCD; HDAC11; MAPK9; HDAC9; PIK3C2A;</dt><dd /></dl>
<dl><dt>HDAC3; TP53; CASP9; CREBBP; AKT1; PIK3R1;</dt><dd /></dl>
<dl><dt>PDPK1, CASP1, APAF1; FRAP1; CASP2; JUN; BAX;</dt><dd /></dl>
<dl><dt>ATF4; AKT3; PRKCA; CLTC; SGK; HOAC6; CASP3</dt><dd /></dl>
<dl><dt>Programmed cell death signaling </dt><dd>PRKCE; ROCK1; IDB; IRAK1; PRKAA2; EIF2AK2; BAK1;</dd></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>BIRC4; GRK6; MAPK1, CAPNS1, PLK1, AKT2; IKBKB;</dt><dd /></dl>
<dl><dt>CAPN2; CDK8; FAS; NFKB2; BCL2; MAP3K14; MAPK8;</dt><dd /></dl>
<dl><dt>BCL2L 1; CAPN1, MAPK3; CASP8; KRAS; RELAY;</dt><dd /></dl>
<dl><dt>PRKCO; PRKAA1, MAPK9; COK2; PIM1; TP53; TNF;</dt><dd /></dl>
<dl><dt>RAF1; IKBKG; RELB; CASP9; OYRK1A; MAP2K2;</dt><dd /></dl>
<dl><dt>CHUK; APAF1; MAP2K1; NFKB1; PAK3; LMNA; CASP2;</dt><dd /></dl>
<dl><dt>BIRC2; TIK; CSNK1A1, BRAF; BAX; PRKCA; SGK;</dt><dd /></dl>
<dl><dt>CASP3; BIRC3; PARPl</dt><dd /></dl>
<dl><dt>B cell receptor signaling </dt><dd>RAC1; PTEN; LYN; ELK1; MAPK1; RAC2; PTPNll;</dd></dl>
<dl><dt>AKT2; IKBKB; PIK3CA; CREB1; SYK; NFKB2; CAMK2A;</dt><dd /></dl>
<dl><dt>MAP3K14; PIK3CB; PIK3C3; MAPK8; BCL2Ll; ABL 1;</dt><dd /></dl>
<dl><dt>MAPK3; ETS1; KRAS; MAPK13; RELAY; PTPN6; MAPK9;</dt><dd /></dl>
<dl><dt>EGRI; PIK3C2A; BTK; MAPK14; RAF1, IKBKG; RELB;</dt><dd /></dl>
<dl><dt>MAP3K7; MAP2K2; AKT1; PIK3Rl; CHUK; MAP2Kl;</dt><dd /></dl>
<dl><dt>NFKB1; CDC42; GSK3A; FRAP1; BCL6; BCL10; JUN;</dt><dd /></dl>
<dl><dt>GSK3B; ATF4; AKT3; VAV3; RPS6KB1</dt><dd /></dl>
<dl><dt>Diademic signaling </dt><dd>ACTN4; C044; PRKCE; ITGAM; ROCK1; CXCR4; CYBA;</dd></dl>
<dl><dt>RAC1, RAP1A; PRKCZ; ROCK2; RAC2; PTPN11,</dt><dd /></dl>
<dl><dt>MMP14; PIK3CA; PRKCI; PTK2; PIK3CB; CXCL12;</dt><dd /></dl>
<dl><dt>PIK3C3; MAPK8; PRK01; ABL1, MAPK10; CYBB;</dt><dd /></dl>
<dl><dt>MAPK13; RHOA; PRKCO; MAPK9; SRC; PIK3C2A; BTK;</dt><dd /></dl>
<dl><dt>MAPK14; NOX1; PXN; VIL2; VASP; TTGB1; MAP2K2;</dt><dd /></dl>
<dl><dt>CTNND1, PIK3R1, CTNNB1, CLON1, CDC42; F11R; ITK;</dt><dd /></dl>
<dl><dt>CRKL; VAV3; CTTN; PRKCA; MMP1, MMP9</dt><dd /></dl>
<dl><dt>Integrina signaling </dt><dd>ACTN4; ITGAM; ROCK1; ITGA5; RAC1; PTEN; RAP1A;</dd></dl>
<dl><dt>TLN1; ARHGEF7; MAPK1, RAC2; CAPNS1, AKT2;</dt><dd /></dl>
<dl><dt>CAPN2; PIK3CA; PTK2; PIK3CB; PIK3C3; MAPK8;</dt><dd /></dl>
<dl><dt>CAV1; CAPN1; ABL1; MAPK3; ITGA1; KRAS; RHOA;</dt><dd /></dl>
<dl><dt>SRC; PIK3C2A; ITGB7; PPP1CC; ILK; PXN; VASP;</dt><dd /></dl>
<dl><dt>RAF1; FYN; ITGB1, MAP2K2; PAK4; AKT1, PIK3R1,</dt><dd /></dl>
<dl><dt>TNK2; MAP2K1; PAK3; ITGB3; C0C42; RN03; ITGA2;</dt><dd /></dl>
<dl><dt>CRKL; BRAF; GSK3B; AKT3</dt><dd /></dl>
<dl><dt>Signalization </dt><dd>IRAK1, S002; MY088; TRAF6; ELK1, MAPK1; PTPN11,</dd></dl>
<dl><dt>Acute phase response </dt><dd>AKT2; IKBKB; PIK3CA; FOS; NFKB2; MAP3K14;</dd></dl>
<dl><dt>PIK3CB; MAPK8; RIPK1, MAPK3; ILBST; KRAS;</dt><dd /></dl>
<dl><dt>MAPK13; IL6R; RELAY; SOCS1; MAPK9; FTL; NR3Cl;</dt><dd /></dl>
<dl><dt>TRAF2; SERPINE1; MAPK14; TNF; RAF1; POK1;</dt><dd /></dl>
<dl><dt>IKBKG; RELB; MAP3K7; MAP2K2; AKT1; JAK2; PIK3R1;</dt><dd /></dl>
<dl><dt>CHUK; STAT3; MAP2Kl; NFKB1; FRAP1; CEBPB; JUN;</dt><dd /></dl>
<dl><dt>AKT3; IL 1 Rl; ILB</dt><dd /></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>PTEN signaling </dt><dd>ITGAM; ITGA5; RAC1, PTEN; PRKCZ; BCL2L 11,</dd></dl>
<dl><dt>MAPK1; RAC2; AKT2; EGFR; IKBKB; CBL; PIK3CA;</dt><dd /></dl>
<dl><dt>CDKN1 B; PTK2; NFKB2; BCL2; PIK3CB; BCL2L 1;</dt><dd /></dl>
<dl><dt>MAPK3; ITGA1, KRAS; ITGB7; ILK; PDGFRB; INSR;</dt><dd /></dl>
<dl><dt>RAF1; IKBKG; CASP9; CDKN1A; ITGB1; MAP2K2;</dt><dd /></dl>
<dl><dt>AKT1, PIK3R1; CHUK; PDGFRA; PDPK1, MAP2K1,</dt><dd /></dl>
<dl><dt>NFKB1, ITGB3; CDC42; CCN01, GSK3A; ITGA2;</dt><dd /></dl>
<dl><dt>GSK3B; AKT3; FOX01; CASP3; RPS6KB1</dt><dd /></dl>
<dl><dt>P53 signaling </dt><dd>PTEN; EP300; BBC3; PCAF; FASN; BRCA1; GADD45A;</dd></dl>
<dl><dt>BIRC5; AKT2; PIK3CA; CHEK1; TP53INP1; BCL2;</dt><dd /></dl>
<dl><dt>PIK3CB; PIK3C3; MAPK8; THBS1; ATR; BCL2L1; E2F1;</dt><dd /></dl>
<dl><dt>PMAIP1; CHEK2; TNFRSF10B; TP73; RB1; HOAC9;</dt><dd /></dl>
<dl><dt>CDK2; PIK3C2A; MAPK14; TP53; LRDD; CDKN1A;</dt><dd /></dl>
<dl><dt>HIPK2; AKT1; PIK3R1; RRM28; APAF1; CTNNB1;</dt><dd /></dl>
<dl><dt>SIRT1; CCN01; PRKDC; ATM; SFN; CDKN2A; JUN;</dt><dd /></dl>
<dl><dt>SNAI2; GSK3B; BAX; AKT3</dt><dd /></dl>
<dl><dt>Signaling </dt><dd>HSPB1; EP300; FASN; TGM2; RXRA; MAPK1; NQ01;</dd></dl>
<dl><dt>of Arylhydrocarbon Receiver </dt><dd>NCOR2; SP1; TRNA; CDKN1 B; FOS; CHEK1,</dd></dl>
<dl><dt>SMARCA4; NFK82; MAPK8; ALDH1A1; ATR; E2F1;</dt><dd /></dl>
<dl><dt>MAPK3; NRIP1, CHEK2; RELAY; TP73; GSTP1; RB1,</dt><dd /></dl>
<dl><dt>SRC; CDK2; AHR; NFE2L2; NCOA3; TP53; TNF;</dt><dd /></dl>
<dl><dt>CDKN IA; NCOA2; APAF1; NFK81; CCND1; ATM; ESR1;</dt><dd /></dl>
<dl><dt>CDKN2A; MYC; JUN; ESR2; BAX; IL6; CYP181,</dt><dd /></dl>
<dl><dt>HSP90AA1 </dt><dd /></dl>
<dl><dt>Signaling of </dt><dd>PRKCE; EP300; PRKCZ; RXRA; MAPK1; NQ01;</dd></dl>
<dl><dt>Xenobiotic metabolism </dt><dd>NCOR2; PIK3CA; TRNA; PRKCI; NFKB2; CAMK2A;</dd></dl>
<dl><dt>PIK3CB; PPP2R1A; PIK3C3; MAPK8; PRKD1;</dt><dd /></dl>
<dl><dt>ALDH1A1; MAPK3; NRIP1; KRAS; MAPK13; PRKCD;</dt><dd /></dl>
<dl><dt>GSTP1; MAPK9; NOS2A; ABC81; AHR; PPP2CA; FTL;</dt><dd /></dl>
<dl><dt>NFE2L2; PIK3C2A; PPARGC1A; MAPK1 4; TNF; RAF1,</dt><dd /></dl>
<dl><dt>CRE8BP; MAP2K2; PIK3R1; PPP2R5C; MAP2K1;</dt><dd /></dl>
<dl><dt>NFK81; KEAP1; PRKCA; EIF2AK3; IL6; CYP1 81;</dt><dd /></dl>
<dl><dt>HSP90AA1 </dt><dd /></dl>
<dl><dt>SAPKlJNK signaling </dt><dd>PRKCE; IRAK1, PRKAA2; EIF2AK2; RAC I, ELK1;</dd></dl>
<dl><dt>GRK6; MAPK1, GADD45A; RAC2; PLK1, AKT2; PIK3CA;</dt><dd /></dl>
<dl><dt>FADD; CDK8; PIK3CB; PIK3C3; MAPK8; R1PK1;</dt><dd /></dl>
<dl><dt>GNB2L 1; IRS1; MAPK3; MAPK10; DAXX; KRAS;</dt><dd /></dl>
<dl><dt>PRKCD; PRKAA1, MAPK9; COK2; PIM1; PIK3C2A;</dt><dd /></dl>
<dl><dt>TRAF2; TP53; LCK; MAP3K7; DYRK1A; MAP2K2;</dt><dd /></dl>
<dl><dt>PIK3R1; MAP2K1; PAK3; CDC42; JUN; TIK; CSNK1A1;</dt><dd /></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>CRKL; BRAF; SGK</dt><dd /></dl>
<dl><dt>PPAr / RXR signaling </dt><dd>PRKAA2; EP300; IN S; SMAD2; TRAF6; PPARA; FASN;</dd></dl>
<dl><dt>RXRA; MAPK1, SMAD3; GNAS; IKBKB; NCOR2;</dt><dd /></dl>
<dl><dt>ABCA1, GNAQ; NFKB2; MAP3K14; STATSB; MAPK8;</dt><dd /></dl>
<dl><dt>IRS1; MAPK3; KRAS; RELAY; PRKAA1; PPARGC1A;</dt><dd /></dl>
<dl><dt>NCOA3; MAPK14; INSR; RAF1; IKBKG; RELB; MAP3K7;</dt><dd /></dl>
<dl><dt>CREBBP; MAP2K2; JAK2; CHUK; MAP2K1, NFKB 1,</dt><dd /></dl>
<dl><dt>TGFBR1; SMAD4; JUN; IL1R1; PRKCA; IL6; HSP90AA1; ADIPOQ</dt><dd /></dl>
<dl><dt>NF-KB signaling </dt><dd>IRAK1; EIF2AK2; EP300; INS; MYD88; PRKCZ; TRAF6;</dd></dl>
<dl><dt>TBK1; AKT2; EGFR; IKBKB; PIK3CA; BTRC; NFKB2;</dt><dd /></dl>
<dl><dt>MAP3K14; PIK3CB; PIK3C3; MAPK8; RIPK1; HDAC2;</dt><dd /></dl>
<dl><dt>KRAS; RELAY; PIK3C2A; TRAF2; TLR4; PDGFRB; TNF;</dt><dd /></dl>
<dl><dt>INSR; LCK; IKBKG; RELB; MAP3K7; CREBBP; AKT1,</dt><dd /></dl>
<dl><dt>PIK3R1; CHUK; PDGFRA; NFKB1; TLR2; BCL 10;</dt><dd /></dl>
<dl><dt>GSK3B; AKT3; TNFAIP3; IL 1 R1</dt><dd /></dl>
<dl><dt>Neuregulin signaling </dt><dd>ERBB4; PRKCE; ITGAM; ITGAS; PTEN; PRKCZ; ELK1;</dd></dl>
<dl><dt>MAPK1; PTPN11; AKT2; EGFR; ERBB2; PRKCI;</dt><dd /></dl>
<dl><dt>CDKN1B; STATSB; PRKD1, MAPK3; ITGA1, KRAS;</dt><dd /></dl>
<dl><dt>PRKCD; STAT5A; SRC; ITGB7; RAF1; ITGB1; MAP2K2;</dt><dd /></dl>
<dl><dt>ADAM17; AKT1, PIK3R1; PDPK1, MAP2K1, ITGB3;</dt><dd /></dl>
<dl><dt>EREG; FRAP1, PSEN1, ITGA2; MYC; NRG1, CRKL;</dt><dd /></dl>
<dl><dt>AKT3; PRKCA; HSP90AA 1; RPS6KB 1</dt><dd /></dl>
<dl><dt>Signaling of </dt><dd>CD44; EP300; LRP6; DVL3; CSNK1E; GJA1, SMO;</dd></dl>
<dl><dt>caten ina Wnt and Beta </dt><dd>AKT2; PIN1, CDH1, BTRC; GNAQ; MARK2; PPP2R1A;</dd></dl>
<dl><dt>WNT11; SRC; DKK1; PPP2CA; SOX6; SFRP2; ILK;</dt><dd /></dl>
<dl><dt>LEF1, SOX9; TP53; MAP3K7; CREBBP; TCF7L2; AKT1,</dt><dd /></dl>
<dl><dt>PPP2R5C; WNT5A; LRP5; CTNNB1; TGFBR1; CCND1;</dt><dd /></dl>
<dl><dt>GSK3A; DVL 1; APC; CDKN2A; MYC; CSNK1A 1; GSK3B;</dt><dd /></dl>
<dl><dt>AKT3; SOX2</dt><dd /></dl>
<dl><dt>Insulin Receptor Signaling </dt><dd>PTEN; INS; EIF4E; PTPN1, PRKCZ; MAPK1, TSC1,</dd></dl>
<dl><dt>PTPN11; AKT2; CBL; PIK3CA; PRKCI; PIK3CB; PIK3C3;</dt><dd /></dl>
<dl><dt>MAPK8; IRS1; MAPK3; TSC2; KRAS; EIF4EBP1;</dt><dd /></dl>
<dl><dt>SLC2A4; PIK3C2A; PPP1CC; INSR; RAF1, FYN;</dt><dd /></dl>
<dl><dt>MAP2K2; JAK1, AKT1, JAK2; PIK3R1, PDPK1; MAP2K1;</dt><dd /></dl>
<dl><dt>GSK3A; FRAP1, CRKL; GSK3B; AKT3; FOX01, SGK;</dt><dd /></dl>
<dl><dt>RPS6KB1 </dt><dd /></dl>
<dl><dt>IL-6 signaling </dt><dd>HSPB1; TRAF6; MAPKAPK2; ELK1; MAPK1; PTPN11;</dd></dl>
<dl><dt>IKBKB; FOS; NFKB2; MAP3K1 4; MAPK8; MAPK3;</dt><dd /></dl>
<dl><dt>MAPK10; IL6ST; KRAS; MAPK13; IL6R; RELAY; SOCS1;</dt><dd /></dl>
<dl><dt>MAPK9; ABCB1; TRAF2; MAPK14; TNF; RAF 1; IKBKG;</dt><dd /></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>RELB; MAP3K7; MAP2K2; ILB; JAK2; CHUK; STAT3;</dt><dd /></dl>
<dl><dt>MAP2K1; NFKB1; CEBPB; JUN; IL1R1; SRF; IL6</dt><dd /></dl>
<dl><dt>Hepatic Cholestasis </dt><dd>PRKCE; IRAK1, INS; MYOB8; PRKCZ; TRAF6; PPARA;</dd></dl>
<dl><dt>RXRA; IKBKB; PRKCI; NFKB2; MAP3K1 4; MAPK8;</dt><dd /></dl>
<dl><dt>PRK01; MAPK10; RELAY; PRKCO; MAPK9; ABCB1;</dt><dd /></dl>
<dl><dt>TRAF2; TLR4; TNF; INSR; IKBKG; RELB; MAP3K7; IL8;</dt><dd /></dl>
<dl><dt>CHUK; NR1H2; TJP2; NFKB1; ESR1, SREBF1, FGFR4;</dt><dd /></dl>
<dl><dt>JUN; IL1R1; PRKCA; IL6</dt><dd /></dl>
<dl><dt>IGF-1 signaling </dt><dd>IGF 1; PRKCZ; ELK1; MAPK1; PTPN11; NE004; AKT2;</dd></dl>
<dl><dt>PIK3CA; PRKCI; PTK2; FOS; PIK3CB; PIK3C3; MAPK8;</dt><dd /></dl>
<dl><dt>IGF1 R; IRS1; MAPK3; IGFBP7; KRAS; PIK3C2A;</dt><dd /></dl>
<dl><dt>YWHAZ; PXN; RAF1; CASP9; MAP2K2; AKT1; PIK3R1;</dt><dd /></dl>
<dl><dt>POPK1, MAP2K1, IGFBP2; SFN; JUN; CYR61; AKT3;</dt><dd /></dl>
<dl><dt>FOX01; SRF; CTGF; RPS6KB1</dt><dd /></dl>
<dl><dt>Stress Response </dt><dd>PRKCE; EP300; S002; PRKCZ; MAPK1; SQSTM1;</dd></dl>
<dl><dt>Oxidative mediated by NRF2 </dt><dd>NQ01, PIK3CA; PRKCI; FOS; PIK3CB; PIK3C3; MAPK8;</dd></dl>
<dl><dt>PRKD1; MAPK3; KRAS; PRKCD; GSTP1; MAPK9; FTL;</dt><dd /></dl>
<dl><dt>NFE2L2; PIK3C2A; MAPK14; RAF1, MAP3K7; CREBBP;</dt><dd /></dl>
<dl><dt>MAP2K2; AKT1; PIK3R1; MAP2K1; PPIB; JUN; KEAP1;</dt><dd /></dl>
<dl><dt>GSK3B; ATF4; PRKCA; EIF2AK3; HSP90AA1</dt><dd /></dl>
<dl><dt>Hepatic Fibrosis I </dt><dd>EON1, 1GF1, KOR; FLT1, SMA02; FGFR1, MET; PGF;</dd></dl>
<dl><dt>Stellar liver cell activation </dt><dd>SMAD3; EGFR; FAS; CSF1; NFKB2; BCL2; MYH9;</dd></dl>
<dl><dt>IGF1R; IL6R; RELAY; TLR4; PDGFRB; TNF; RELB; ILB;</dt><dd /></dl>
<dl><dt>POGFRA; NFKB1, TGFBR1, SMA04; VEGFA; BAX;</dt><dd /></dl>
<dl><dt>IL 1R1; CCL2; HGF; MMP1; STAT1; IL6; CTGF; MMP9</dt><dd /></dl>
<dl><dt>PPAR signaling </dt><dd>EP300; INS; TRAF6; PPARA; RXRA; MAPK1, IKBKB;</dd></dl>
<dl><dt>NCOR2; FOS; NFKB2; MAP3K14; STAT5B; MAPK3;</dt><dd /></dl>
<dl><dt>NRIP1; KRAS; PPARG; RELAY; STAT5A; TRAF2;</dt><dd /></dl>
<dl><dt>PPARGC1A; PDGFRB; TNF; INSR; RAF1; IKBKG;</dt><dd /></dl>
<dl><dt>RELB; MAP3K7; CREBBP; MAP2K2; CHUK; POGFRA;</dt><dd /></dl>
<dl><dt>MAP2K1; NFKB1; JUN; IL1R1; HSP90AA1</dt><dd /></dl>
<dl><dt>Fc Epsilon RI signaling </dt><dd>PRKCE; RAC1; PRKCZ; LYN; MAPK1; RAC2; PTPN11;</dd></dl>
<dl><dt>AKT2; PIK3CA; SYK; PRKCI; PIK3CB; PIK3C3; MAPK8;</dt><dd /></dl>
<dl><dt>PRKD1, MAPK3; MAPK10; KRAS; MAPK13; PRKCD;</dt><dd /></dl>
<dl><dt>MAPK9; PIK3C2A; BTK; MAPK14; TNF; RAF1; FYN;</dt><dd /></dl>
<dl><dt>MAP2K2; AKT1; PIK3R1; PDPK1; MAP2K1; AKT3;</dt><dd /></dl>
<dl><dt>VAV3; PRKCA</dt><dd /></dl>
<dl><dt>Receiver signaling </dt><dd>PRKCE; RAP1A; RGS16; MAPK1; GNAS; AKT2; IKBKB;</dd></dl>
<dl><dt>G protein coupled </dt><dd>PIK3CA; CREB1; GNAQ; NFKB2; CAMK2A; PIK3CB;</dd></dl>
<dl><dt>PIK3C3; MAPK3; KRAS; RELAY; SRC; PIK3C2A; RAF1;</dt><dd /></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>IKBKG; RELB; FYN; MAP2K2; AKT1; PIK3R1, CHUK;</dt><dd /></dl>
<dl><dt>PDPK1; STAT3; MAP2K1; NFKB1; BRAF; ATF4; AKT3;</dt><dd /></dl>
<dl><dt>PRKCA </dt><dd /></dl>
<dl><dt>Metabolism of </dt><dd>PRKCE; IRAK1, PRKAA2; EIF2AK2; PTEN; GRK6;</dd></dl>
<dl><dt>Inositol phosphate </dt><dd>MAPK1; PLK1; AKT2; PIK3CA; CDK8; PIK3CB; PIK3C3;</dd></dl>
<dl><dt>MAPK8; MAPK3; PRKCD; PRKAA1; MAPK9; CDK2;</dt><dd /></dl>
<dl><dt>PIM1, PIK3C2A; DYRK1A; MAP2K2; PIPSK1A; PIK3R1,</dt><dd /></dl>
<dl><dt>MAP2K1; PAK3; ATM; TIK; CSNK1A1; BRAF; SGK</dt><dd /></dl>
<dl><dt>PDGF signaling </dt><dd>EIF2AK2; ELK1; ABL2; MAPK1; PIK3CA; FOS; PIK3CB;</dd></dl>
<dl><dt>PIK3C3; MAPK8; CAV1; ABL 1; MAPK3; KRAS; SRC;</dt><dd /></dl>
<dl><dt>PIK3C2A; PDGFRB; RAF1; MAP2K2; JAK1; JAK2;</dt><dd /></dl>
<dl><dt>PIK3R1; PDGFRA; STAT3; SPHK1; MAP2K1; MYC;</dt><dd /></dl>
<dl><dt>JUN; CRKL; PRKCA; SRF; STAT1, SPHK2</dt><dd /></dl>
<dl><dt>VEGF signaling </dt><dd>ACTN4; ROCK1; KDR; FL T1; ROCK2; MAPK1; PGF;</dd></dl>
<dl><dt>AKT2; PIK3CA; TRNA; PTK2; BCL2; PIK3CB; PIK3C3;</dt><dd /></dl>
<dl><dt>BCL2L 1; MAPK3; KRAS; HIF1A; NOS3; PIK3C2A; PXN;</dt><dd /></dl>
<dl><dt>RAF1; MAP2K2; ELAVL1; AKT1; PIK3R1; MAP2K1; SFN;</dt><dd /></dl>
<dl><dt>VEGFA; AKT3; FOX01, PRKCA</dt><dd /></dl>
<dl><dt>Signalization of cells </dt><dd>PRKCE; RAC1; PRKCZ; MAPK1; RAC2; PTPN11;</dd></dl>
<dl><dt>Natural Iquilants </dt><dd>KIR2DL3; AKT2; PIK3CA; SYK; PRKCI; PIK3CB;</dd></dl>
<dl><dt>PIK3C3; PRKD1, MAPK3; KRAS; PRKCO; PTPN6;</dt><dd /></dl>
<dl><dt>PIK3C2A; LCK; RAF1; FYN; MAP2K2; PAK4; AKT1;</dt><dd /></dl>
<dl><dt>PIK3R1, MAP2K1, PAK3; AKT3; VAV3; PRKCA</dt><dd /></dl>
<dl><dt>R cell cycle: G1 / S </dt><dd>HOAC4; SMA03; SUV39H1, HOACS; COKN1B; BTRC;</dd></dl>
<dl><dt>Control regulation </dt><dd>ATR; ABL1; E2F1; HDAC2; HDAC7A; RB1; HDAC11;</dd></dl>
<dl><dt>HOAC9; COK2; E2F2; HOAC3; TP53; CDKN1A; CCN01,</dt><dd /></dl>
<dl><dt>E2F4; ATM; RBL2; SMAD4; CDKN2A; MYC; NRG1;</dt><dd /></dl>
<dl><dt>GSK3B; RBL1; HDAC6</dt><dd /></dl>
<dl><dt>T cell receptor signaling </dt><dd>RAC1; ELK1; MAPK1; IKBKB; CBL; PIK3CA; FOS;</dd></dl>
<dl><dt>NFKB2; PIK3CB; PIK3C3; MAPK8; MAPK3; KRAS;</dt><dd /></dl>
<dl><dt>RELAY; PIK3C2A; BTK; LCK; RAF1; IKBKG; RELB; FYN;</dt><dd /></dl>
<dl><dt>MAP2K2; PIK3R1; CHUK; MAP2K1; NFKB1; ITK; BCL10;</dt><dd /></dl>
<dl><dt>JUN; VAV3</dt><dd /></dl>
<dl><dt>Death Receiver Signage </dt><dd>CRADD; HSPB1, IDB; BIRC4; TBK1, IKBKB; FADO;</dd></dl>
<dl><dt>FAS; NFKB2; BCL2; MAP3K14; MAPK8; RIPK1, CASP8;</dt><dd /></dl>
<dl><dt>OAXX; TNFRSF10B; RELAY; TRAF2; TNF; IKBKG; RELB;</dt><dd /></dl>
<dl><dt>CASP9; CHUK; APAF1; NFKB1; CASP2; BIRC2; CASP3;</dt><dd /></dl>
<dl><dt>BIRC3 </dt><dd /></dl>
<dl><dt>FGF signaling </dt><dd>RAC1; FGFR1; MET; MAPKAPK2; MAPK1; PTPN11;</dd></dl>
<dl><dt>AKT2; PIK3CA; CREB1; PIK3CB; PIK3C3; MAPK8;</dt><dd /></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>MAPK3; MAPK13; PTPN6; PIK3C2A; MAPK14; RAF1;</dt><dd /></dl>
<dl><dt>AKT1; PIK3R1; STAT3; MAP2K1; FGFR4; CRKL; ATF4;</dt><dd /></dl>
<dl><dt>AKT3; PRKCA; HGF</dt><dd /></dl>
<dl><dt>GM-CSF signaling </dt><dd>LYN; ELK1; MAPK1, PTPN11; AKT2; PIK3CA; CAMK2A;</dd></dl>
<dl><dt>STAT5B; PIK3CB; PIK3C3; GNB2L 1; BCL2L 1; MAPK3;</dt><dd /></dl>
<dl><dt>ETS1, KRAS; RUNX1, PIM1, PIK3C2A; RAF1; MAP2K2;</dt><dd /></dl>
<dl><dt>AKT1, JAK2; PIK3R1, STAT3; MAP2K1, CCN01, AKT3;</dt><dd /></dl>
<dl><dt>STAT1 </dt><dd /></dl>
<dl><dt>Sclerosis signaling </dt><dd>IDB; IGF1; RAC1; BIRC4; PGF; CAPNS1; CAPN2;</dd></dl>
<dl><dt>amyotrophic lateral </dt><dd>PIK3CA; BCL2; PIK3CB; PIK3C3; BCL2L1; CAPN1;</dd></dl>
<dl><dt>PIK3C2A; TP53; CASP9; PIK3R1; RAB5A; CASP1;</dt><dd /></dl>
<dl><dt>APAF1; VEGFA; BIRC2; BAX; AKT3; CASP3; BIRC3</dt><dd /></dl>
<dl><dt>JAK / Stat signaling </dt><dd>PTPN1; MAPK1, PTPN11; AKT2; PIK3CA; STAT5B;</dd></dl>
<dl><dt>PIK3CB; PIK3C3; MAPK3; KRAS; SOCS1; STAT5A;</dt><dd /></dl>
<dl><dt>PTPN6; PIK3C2A; RAF1; CDKN1A; MAP2K2; JAK1;</dt><dd /></dl>
<dl><dt>AKT1, JAK2; PIK3R1, STAT3; MAP2K1; FRAP1; AKT3;</dt><dd /></dl>
<dl><dt>STAT1 </dt><dd /></dl>
<dl><dt>Metabolism of </dt><dd>PRKCE; IRAK1, PRKAA2; EIF2AK2; GRK6; MAPK1,</dd></dl>
<dl><dt>Nicotinate and </dt><dd /></dl>
<dl><dt>Nicotinamide </dt><dd>PLK1, AKT2; CDK8; MAPK8; MAPK3; PRKCD; PRKAA1,</dd></dl>
<dl><dt>PBEF1, MAPK9; CDK2; PIM1, DYRK1A; MAP2K2;</dt><dd /></dl>
<dl><dt>MAP2K1; PAK3; NT5E; TTK; CSNK1A1; BRAF; SGK</dt><dd /></dl>
<dl><dt>Chemokine signaling </dt><dd>CXCR4; ROCK2; MAPK1, PTK2; FOS; CFL 1, GNAQ;</dd></dl>
<dl><dt>CAMK2A; CXCL12; MAPK8; MAPK3; KRAS; MAPK13;</dt><dd /></dl>
<dl><dt>RHOA; CCR3; SRC; PPP1CC; MAPK14; NOX1; RAF1;</dt><dd /></dl>
<dl><dt>MAP2K2; MAP2K1, FLAN; CCL2; PRKCA</dt><dd /></dl>
<dl><dt>IL-2 signaling </dt><dd>ELK1; MAPK1; PTPN11; AKT2; PIK3CA; SYK; FOS;</dd></dl>
<dl><dt>STAT5B; PIK3CB; PIK3C3; MAPK8; MAPK3; KRAS;</dt><dd /></dl>
<dl><dt>SOCS1; STAT5A; PIK3C2A; LCK; RAF1; MAP2K2;</dt><dd /></dl>
<dl><dt>JAK1, AKT1, PIK3R1, MAP2K1, JUN; AKT3</dt><dd /></dl>
<dl><dt>Synaptic depression </dt><dd>PRKCE; IGF1; PRKCZ; PRDX6; LYN; MAPK1; GNAS;</dd></dl>
<dl><dt>long tie </dt><dd>PRKCI; GNAQ; PPP2R1A; IGF1 R; PRKD1; MAPK3;</dd></dl>
<dl><dt>KRAS; GRN; PRKCD; NOS3; NOS2A; PPP2CA;</dt><dd /></dl>
<dl><dt>YVVHAZ; RAF1, MAP2K2; PPP2R5C; MAP2K1, PRKCA</dt><dd /></dl>
<dl><dt>Signaling of </dt><dd>TAF4B; EP300; CARM1, PCAF; MAPK1, NCOR2;</dd></dl>
<dl><dt>EstrogellOs Receiver </dt><dd>SMARCA4; MAPK3; NRIP1; KRAS; SRC; NR3C1;</dd></dl>
<dl><dt>HDAC3; PPARGC1A; RBM9; NCOA3; RAF1; CREBBP;</dt><dd /></dl>
<dl><dt>MAP2K2; NCOA2; MAP2K1, PRKDC; ESR1, ESR2</dt><dd /></dl>
<dl><dt>Route of </dt><dd>TRAF6; SMURF1; BIRC4; BRCA1; UCHL 1; NEDD4;</dd></dl>
<dl><dt>Protein ubiquilination </dt><dd>CBL; UBE21; BTRC; HSPA5; USP7; USP10; FBXW7;</dd></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>USP9X; STUB1, USP22; B2M; BIRC2; PARK2; USP8;</dt><dd /></dl>
<dl><dt>USP1; VHL; HSP90AA1; BIRC3</dt><dd /></dl>
<dl><dt>IL-1 0 signaling </dt><dd>TRAF6; CCR1, ELK1, IKBKB; SP1, FOS; NFKB2;</dd></dl>
<dl><dt>MAP3K14; MAPK8; MAPK13; RELAY; MAPK1 4; TNF;</dt><dd /></dl>
<dl><dt>IKBKG; RELB; MAP3K7; JAK1; CHUK; STAT3; NFKB1;</dt><dd /></dl>
<dl><dt>JUN; IL1R1, IL6</dt><dd /></dl>
<dl><dt>VORlRXR activation </dt><dd>PRKCE; EP300; PRKCZ; RXRA; GA004SA; HES1,</dd></dl>
<dl><dt>NCOR2; SP1; PRKCI; COKN1B; PRK01; PRKCO;</dt><dd /></dl>
<dl><dt>RUNX2; KLF4; YY1; NCOA3; COKN1A; NCOA2; SPP1;</dt><dd /></dl>
<dl><dt>LRPS; CEBPB; FOX01; PRKCA</dt><dd /></dl>
<dl><dt>TGF-beta signaling </dt><dd>EP300; SMAD2; SMURF1; MAPK1; SMA03; SMA01;</dd></dl>
<dl><dt>FOS; MAPK8; MAPK3; KRAS; MAPK9; RUNX2;</dt><dd /></dl>
<dl><dt>SERPINE1; RAF1; MAP3K7; CREBBP; MAP2K2;</dt><dd /></dl>
<dl><dt>MAP2K1; TGFBR1; SMAD4; JUN; SMADS</dt><dd /></dl>
<dl><dt>TolI type receiver signaling </dt><dd>IRAK1; EIF2AK2; MYD88; TRAF6; PPARA; ELK1;</dd></dl>
<dl><dt>IKBKB; FOS; NFKB2; MAP3K1 4; MAPK8; MAPK13;</dt><dd /></dl>
<dl><dt>RELAY; TLR4; MAPK14; IKBKG; RELB; MAP3K7; CHUK;</dt><dd /></dl>
<dl><dt>NFKB1; TLR2; JUN</dt><dd /></dl>
<dl><dt>P38 MAPK signaling </dt><dd>HSPB1; IRAK1; TRAF6; MAPKAPK2; ELK1; FADO; FAS;</dd></dl>
<dl><dt>CREB1, 001T3; RPS6KA4; OAXX; MAPK13; TRAF2;</dt><dd /></dl>
<dl><dt>MAPK14; TNF; MAP3K7; TGFBR1, MYC; ATF4; IL1R1,</dt><dd /></dl>
<dl><dt>SRF; STAT1</dt><dd /></dl>
<dl><dt>NeurolrofinafTRK signaling </dt><dd>NTRK2; MAPK 1, PTPN11, PIK3CA; CREB1; FOS;</dd></dl>
<dl><dt>PIK3CB; PIK3C3; MAPK8; MAPK3; KRAS; PIK3C2A;</dt><dd /></dl>
<dl><dt>RAF1; MAP2K2; AKT1; PIK3R1; POPK1; MAP2K1;</dt><dd /></dl>
<dl><dt>C0C42; JUN; ATF4</dt><dd /></dl>
<dl><dt>FXRlRXR Activation </dt><dd>INS; PPARA; FASN; RXRA; AKT2; SDC1; MAPK8;</dd></dl>
<dl><dt>APOB; MAPK10; PPARG; MTTP; MAPK9; PPARGC1A;</dt><dd /></dl>
<dl><dt>TNF; CREBBP; AKT1; SREBF1; FGFR4; AKT3; FOX01</dt><dd /></dl>
<dl><dt>Synaptic Empowerment </dt><dd>PRKCE; RAP1A; EP300; PRKCZ; MAPK1; CREB1,</dd></dl>
<dl><dt>long-term </dt><dd>PRKCI; GNAQ; CAMK2A; PRK01; MAPK3; KRAS;</dd></dl>
<dl><dt>PRKCD; PPP1CC; RAF1; CREBBP; MAP2K2; MAP2K1;</dt><dd /></dl>
<dl><dt>ATF4; PRKCA</dt><dd /></dl>
<dl><dt>Calcium signaling </dt><dd>RAP1A; EP300; HOAC4; MAPK1, HDACS; CREB1,</dd></dl>
<dl><dt>CAMK2A; MYH9; MAPK3; HOAC2; HOAC7A; HOAC11,</dt><dd /></dl>
<dl><dt>HOAC9; HOAC3; CREBBP; CALR; CAMKK2; ATF4;</dt><dd /></dl>
<dl><dt>HOAC6 </dt><dd /></dl>
<dl><dt>EGF signaling </dt><dd>ELK1, MAPK1; EGFR; PIK3CA; FOS; PIK3CB; PIK3C3;</dd></dl>
<dl><dt>MAPK8; MAPK3; PIK3C2A; RAF1; JAK1; PIK3R1;</dt><dd /></dl>
<dl><dt>STAT3; MAP2K1; JUN; PRKCA; SRF; STAT1</dt><dd /></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>Hypoxia signaling in the </dt><dd>EDN1, PTEN; EP300; NQ01, UBE21; CREB1; TRNA;</dd></dl>
<dl><dt>Cardiovascular system </dt><dd>HIF1A; SLC2A4; NOS3; TP53; LDHA; AKT1; ATM;</dd></dl>
<dl><dt>VEGFA; JUN; ATF4; VHL; HSP90AA1</dt><dd /></dl>
<dl><dt>Inhibition mediated by LPS / IL-1 </dt><dd>IRAK1, MYD88; TRAF6; PPARA; RXRA; ABCA1;</dd></dl>
<dl><dt>RXR function </dt><dd>MAPK8; ALDH1A1; GSTP1; MAPK9; ABCB1; TRAF2;</dd></dl>
<dl><dt>TLR4; TNF; MAP3K7; NR1 H2; SREBF1, JUN; IL 1 R1</dt><dd /></dl>
<dl><dt>LXRlRXR Activation </dt><dd>FASN; RXRA; NCOR2; ABCA1, NFKB2; IRF3; RELAY;</dd></dl>
<dl><dt>NOS2A; TLR4; TNF; RELB; LDLR; NR1H2; NFKB1;</dt><dd /></dl>
<dl><dt>SREBF1; Il 1 R1; CCL2; IL6; MMP9</dt><dd /></dl>
<dl><dt>Amyloid process </dt><dd>PRKCE; CSNK1 E; MAPK1; CAPNS 1; AKT2; CAPN2;</dd></dl>
<dl><dt>CAPN1; MAPK3; MAPK13; MAPT; MAPK14; AKT1;</dt><dd /></dl>
<dl><dt>PSEN1; CSNK1A1; GSK3B; AKT3; APP</dt><dd /></dl>
<dl><dt>IL-4 signaling </dt><dd>AKT2; PIK3CA; PIK3CB; PIK3C3; IRS1, KRAS; SOCKS1,</dd></dl>
<dl><dt>PTPN6; NR3C1; PIK3C2A; JAK1; AKT1; JAK2; PIK3R1;</dt><dd /></dl>
<dl><dt>FRAP1; AKT3; RPS6KB1</dt><dd /></dl>
<dl><dt>Regulation </dt><dd>EP300; PCAF; BRCA1, GADD45A; PLK1; BTRC;</dd></dl>
<dl><dt>Damage control </dt><dd>CHEK1; ATR; CHEK2; YWHAZ; TP53; CDKN1A;</dd></dl>
<dl><dt>R cell cycle: G2IM DNA </dt><dd>PRKDC; ATM; SFN; CDKN2A</dd></dl>
<dl><dt>Signaling nitric oxide in the </dt><dd>KDR; Fl T1; PGF; AKT2; PIK3CA; PIK3CB; PIK3C3;</dd></dl>
<dl><dt>Cardiovascular system </dt><dd>CAV1, PRKCD; NOS3; PIK3C2A; AKT1, PIK3R1,</dd></dl>
<dl><dt>VEGFA; AKT3; HSP90AA1</dt><dd /></dl>
<dl><dt>Purina metabolism </dt><dd>NME2; SMARCA4; MYH9; RRM2; TO GIVE; EIF2AK4;</dd></dl>
<dl><dt>PKM2; ENTP01, RAD51, RRM2B; TJP2; RAD51C;</dt><dd /></dl>
<dl><dt>NT5E; POLD1; NME1</dt><dd /></dl>
<dl><dt>Signaling measured by cAMP </dt><dd>RAP1A; MAPK 1; GNAS; CREB1; CAMK2A; MAPK3;</dd></dl>
<dl><dt>SRC; RAF1, MAP2K2; STAT3; MAP2K1, BRAF; ATF4</dt><dd /></dl>
<dl><dt>Mitochondrial dysfunction </dt><dd>SOD2; MAPK8; CASP8; MAPK10; MAPK9; CASP9;</dd></dl>
<dl><dt>PARK7; PSEN1; PARK2; APP; CASP3</dt><dd /></dl>
<dl><dt>Notch signage </dt><dd>HES1; JAG1; NUMB; NOTCH4; ADAM 17; NOTCH2;</dd></dl>
<dl><dt>PSEN1, NOTCH3; NOTCH1, DLL4</dt><dd /></dl>
<dl><dt>Stress Route </dt><dd>HSPA5; MAPK8; XBP1; TRAF2; ATF6; CASP9; ATF4;</dd></dl>
<dl><dt>Endoplasmic reticulum </dt><dd>EIF2AK3; CASP3</dd></dl>
<dl><dt>Pyrimidine metabolism </dt><dd>NME2; AICDA; RRM2; EIF2AK4; ENTPD1, RRM2B;</dd></dl>
<dl><dt>NTSE; POlD1; NME1</dt><dd /></dl>
<dl><dt>Parkinson's signaling </dt><dd>UCHL1, MAPK8; MAPK13; MAPK14; CASP9; PARK7;</dd></dl>
<dl><dt>PARK2; CASP3</dt><dd /></dl>
<dl><dt>Signalization </dt><dd>GNAS; GNAQ; PPP2R1A; GNB2l1; PPP2CA; PPP1CC;</dd></dl>
<dl><dt>Ca rdíaca and beta adrenergic </dt><dd>PPP2RSC </dd></dl>
<dl><dt>Glycolysisl Gluconeogenesis </dt><dd>HK2; GCK; GPI; ALDH1A1; PKM2; lDHA; HK1</dd></dl>
<dl><dt>Interferon signaling </dt><dd>IRF1; SOCS1; JAK1; JAK2; IFITM1; STAT1; IFIT3</dd></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>Sonic Hedgehog signage </dt><dd>ARRB2; SAINT; GLl2; DYRK1A; GLl1, GSK3B; DYRK1B</dd></dl>
<dl><dt>Metabolism </dt><dd>PLD1; GRN; GPAM; YWHAZ; SPHK1; SPHK2</dd></dl>
<dl><dt>Gl icerophospholipids </dt><dd /></dl>
<dl><dt>Degradation of phospholipids </dt><dd>PRDX6; PLD1; GRN; YWHAZ; SPHK1; SPHK2</dd></dl>
<dl><dt>Iriplofan metabolism </dt><dd>SIAH2; PRMTS; NEDD4; ALDH1A1; CYP1B1; SIAH1</dd></dl>
<dl><dt>Lysine degradation </dt><dd>SUV39H1, EHMT2; NSD1, SET07; PPP2RSC</dd></dl>
<dl><dt>Rula </dt><dd>ERCCS; ERCC4; XPA; XPC; ERCC1</dd></dl>
<dl><dt>Nucleolide Excision Repair </dt><dd /></dl>
<dl><dt>Metabolism of </dt><dd>UCH L 1; HK2; GCK; GPI; HK1</dd></dl>
<dl><dt>Starch and sucrose </dt><dd /></dl>
<dl><dt>Amino sugar metabolism </dt><dd>NQ01; HK2; GCK; HK1</dd></dl>
<dl><dt>Metabolism of </dt><dd>PRDX6; GRN; YWHAZ; CYP1B1</dd></dl>
<dl><dt>Arachidonic acid </dt><dd /></dl>
<dl><dt>Signalization of circadian rhythm </dt><dd>CSNK1E; CREB1; ATF4; NR1D1</dd></dl>
<dl><dt>Coagulation system </dt><dd>BDKRB1; F2R; SERPINE1; F3</dd></dl>
<dl><dt>Dopamine Receiver </dt><dd>PPP2R1A; PPP2CA; PPP1CC; PPP2R5C</dd></dl>
<dl><dt>Signaling </dt><dd /></dl>
<dl><dt>Glutathione metabolism </dt><dd>IDH2; GSTP1, ANPEP; IDH1</dd></dl>
<dl><dt>Glycerolipid metabolism </dt><dd>ALDH1A1; GPAM; SPHK1; SPHK2</dd></dl>
<dl><dt>Linoleic acid metabolism </dt><dd>PRDX6; GRN; YWHAZ; CYP1B1</dd></dl>
<dl><dt>Methionine metabolism </dt><dd>DNMT1, ONMT3B; AHCY; ONMT3A</dd></dl>
<dl><dt>Pyruvate Metabolism </dt><dd>GL01; ALOH1A1; PKM2; IT HAS</dd></dl>
<dl><dt>Metabolism of </dt><dd>ALDH1A1, NOS3; NOS2A</dd></dl>
<dl><dt>Argin ina and Prolina </dt><dd /></dl>
<dl><dt>Eicosanoid signaling </dt><dd>PRDX6; GRN; YWHAZ</dd></dl>
<dl><dt>metabolism of </dt><dd>HK2; GCK; HK t</dd></dl>
<dl><dt>Fructose and Mannose </dt><dd /></dl>
<dl><dt>Galactose metabolism </dt><dd>HK2; GCK; HK1</dd></dl>
<dl><dt>Stilbene biosynthesis, </dt><dd>PRDX6; PRDX1; TYR</dd></dl>
<dl><dt>cuma rina and lignina </dt><dd /></dl>
<dl><dt>Route of </dt><dd>CALR; B2M</dd></dl>
<dl><dt>Antigen presentation </dt><dd /></dl>
<dl><dt>Steroid biosynthesis </dt><dd>NQ01, OHCR7 </dd></dl>
<dl><dt>Butanoate metabolism </dt><dd>ALDH1A1, NLGN1 </dd></dl>
<dl><dt>Citrate Cycle </dt><dd>IOH2; IOH1</dd></dl>
<dl><dt>Fatty acid metabolism </dt><dd>ALDH1A1; CYP1B1</dd></dl>
<dl><dt>Metabolism of </dt><dd>PRDX6; CHKA</dd></dl>
<dl><dt>Gl icerophospholipids </dt><dd /></dl>
<dl><dt>Histidine metabolism </dt><dd>PRMTS; ALDH1A1</dd></dl>
<dl><dt>Inositol metabolism </dt><dd>ER01L; APEX1</dd></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>Xenobiotic metabolism </dt><dd>GSTP1, CYP1 81 </dd></dl>
<dl><dt>by cytochrome p450 </dt><dd /></dl>
<dl><dt>Methane metabolism </dt><dd>PRDX6; PRDX1</dd></dl>
<dl><dt>Phenylalanine metabolism </dt><dd>PRDX6; PRDX1</dd></dl>
<dl><dt>Propanoate metabolism </dt><dd>ALDH1A1; LDHA</dd></dl>
<dl><dt>Metabolism of </dt><dd>PRMT5; AHCY</dd></dl>
<dl><dt>Selenoamino acid </dt><dd /></dl>
<dl><dt>Sphingolipid metabolism </dt><dd>SPHK1; SPHK2</dd></dl>
<dl><dt>Aminophosphonate </dt><dd>PRMT5 </dd></dl>
<dl><dt>Metabolism </dt><dd /></dl>
<dl><dt>Androgens and estrogens </dt><dd>PRMT5 </dd></dl>
<dl><dt>Metabolism of </dt><dd /></dl>
<dl><dt>Ascorbate and Aldarato </dt><dd>ALDH1A1 </dd></dl>
<dl><dt>Metabolism of </dt><dd /></dl>
<dl><dt>Biosyntis of bile acids </dt><dd>ALDH1A1 </dd></dl>
<dl><dt>Cysteine metabolism </dt><dd>LDHA </dd></dl>
<dl><dt>Biosinteis of fatty acids </dt><dd>FASN </dd></dl>
<dl><dt>Glutamate Receptor </dt><dd>GN82L1 </dd></dl>
<dl><dt>Signaling of </dt><dd /></dl>
<dl><dt>Stress response </dt><dd>PRDX1 </dd></dl>
<dl><dt>Oxidative mediated by NRF2 </dt><dd /></dl>
<dl><dt>Route of </dt><dd>GPI </dd></dl>
<dl><dt>Pentose phosphate </dt><dd /></dl>
<dl><dt>Interconversions of </dt><dd>UCHL1 </dd></dl>
<dl><dt>Pentose and Glucuronate </dt><dd /></dl>
<dl><dt>Retinol metabolism </dt><dd>ALDH1A1 </dd></dl>
<dl><dt>Riboflavin metabolism </dt><dd>TYR </dd></dl>
<dl><dt>Tyrosine Metabolism </dt><dd>PRMT5, TYR </dd></dl>
<dl><dt>Ubiquinone Biosynthesis </dt><dd>PRMT5 </dd></dl>
<dl><dt>Degradation of Valine, Leucine and </dt><dd>ALDH1A1 </dd></dl>
<dl><dt>Isoleucine </dt><dd /></dl>
<dl><dt>Metabolism of Wisteria, Serine and </dt><dd>CHKA </dd></dl>
<dl><dt>Threonine </dt><dd /></dl>
<dl><dt>Lysine degradation </dt><dd>ALDH1A1 </dd></dl>
<dl><dt>pain </dt><dd>TRPM5; TRPA1</dd></dl>
<dl><dt>pain </dt><dd>TRPM7; TRPC5; TRPC6; TRPC1; Cnr1; cnr2; Grk2;</dd></dl>
<dl><dt>Trpa1; Pomc; Cgrp; Crf; Pka; Was; Nr2b; TRPM5; Prkaca;</dt><dd /></dl>
<dl><dt>Prkacb; Prkar1a; Prkar2a</dt><dd /></dl>
<dl><dt>Mitochondrial function </dt><dd>IDA; CytC; SMAC (Devil); Aifm-1; Aifm-2</dd></dl>
<dl><dt>Developmental neurology </dt><dd>BMP-4; Chordin (Chrd); Noggin (Nog); WNT (Wnt2;</dd></dl>
<dl><dt>CELLULAR FUNCTION </dt><dd>GENES </dd></dl>
<dl><dt>Wnt2b; Wnt3a; Wnt4; Wnt5a; Wnt6; Wnt7b; Wnt8b;</dt><dd /></dl>
<dl><dt>Wnt9a; Wnt9b; Wnt10a; Wnt10b; Wnt16); beta-eatenin;</dt><dd /></dl>
<dl><dt>Dkk-1, Frizzled-related proteins; Otx-2; Gbx2; FGF-8;</dt><dd /></dl>
<dl><dt>Reelin; Dab1; unc-B6 (Pou4f1 or Brn3a); Numb; Reln</dt><dd /></dl>
The embodiments also refer to methods and compositions related to gene elimination, gene multiplication and repair of particular mutations associated with instability of DNA repeats and neurological disorders (Robert D. Wells, Tetsuo Ashizawa, Genetic Instabilities and Neurological Diseases, Second Edition, Academic Press October 13, 2011-Medical). Specific aspects of tandem repetition sequences have been found responsible for more than twenty human diseases (New knowledge in the instability of repetition: function of RNA-DNA hybrids. Mclvor El, Polak U, Napierala
M. RNA Biol. Sep-Oct 2010; 7 (5): 551-8). The CRISPR-Cas system can be used to correct these defects of gellomatic instability
A further aspect of the description refers to the use of the CRISPR-Cas system to correct defects in the EMP2A and EMP2B genes that have been identified to be associated with Lafora disease. Lafora disease is an autosomal recessive disorder that is characterized by progressive myoclonic epilepsy that can begin as epileptic seizures in adolescence. Some cases of the disease may be caused by mutations in genes that have to be identified however. The disease produces seizures, muscle spasms, difficulty walking, dementia and eventually death. At present there is no treatment that has been proven effective against the progress of the disease. Other genetic abnormalities associated with epilepsy can also be set as targets through the CRISPR-Cas system and the underlying genetics is also described in Genetics of Epilepsy and Genetic Epilepsies, edited by Giu liano Avanzini, Jeffrey L. Noebels, Mariani Foundation Paediatric Neurology: 20; 2,009)
In yet another aspect, the CRISPR-Cas system can be used to correct eye defects that arise from various genetic mutations described further in Genetic Diseases of the Eye, Second Edition, edited by Elias l. Traboulsi, Oxford University Press, 2.012
Various additional aspects refers to correcting defects associated with a wide range of genetic diseases that are also described on the website of the National Institute for Health under the topical subsection Genetic disorders Genetic brain diseases may include but are not limited to, Adrenoleukodystrophy, Causal Body Angenesis, Aicard I Syndrome, Alper's Disease, Alzheimer's Disease, Barth Syndrome, Balten's Disease, CADASIL, Cerebellar degeneration, Fabry disease, Gerstmann-Straussler-Scheinker disease, Huntington's disease and other Triple Repetition Trastomos, Leigh's disease, Lesch-Nyhan syndrome, Menkes disease, Mitochondrial myopathies and NINDS colpocephaly (by its acronym in English). These diseases are also described on the website of the National Institute of Health under the subsection of Genetic Brain Disorders.
In some embodiments, the condition may be neoplasia. In some embodiments, where the condition is neoplasia, the genes that have to be set as targets are any of those listed in Table A (in this case PTEN etc). In some embodiments, the condition may be Age-related Macular Degeneration. In some embodiments, the condition may be a Schizophrenic Disorder. In some embodiments, the condition may be a Trinucleotide Repeat disorder. In some embodiments, the condition may be Fragile X Syndrome. In some embodiments, the condition may be a Secretase-Related Disorder. In some embodiments, the condition may be a disorder related to the Prion. In some embodiments, the condition may be ALS (for its acronym in English). In some embodiments, the condition may be a drug addiction. In some embodiments, the condition may be Autism. In some embodiments, the condition may be Alzheimer's disease. In some embodiments, the condition may be inflammation. In some embodiments, the condition may be Par1 disease <inson
Examples of proteins associated with Pal1 <inson disease include but are not limited to a-synuclein, DJ-1, LRRK2, PINK1, Par1 <in, UCHL 1, Synphilin-1 and NURR1
Examples of addiction-related proteins may include ABAT for example.
Examples of inflammation-related proteins may include the monocyte chemoattractant protein -1 (MCP1) encoded by the Ccr2 gene, the type 5 chemokine receptor CC (CCR5) encoded by the Ccr5 gene, the IgG IIB receptor (FCGR2b, also called CD32) encoded by the Fcgr2b gene or the R1g Fc epsilon protein (FCER1g) encoded by the Fcer1g gene, for example
Examples of proteins associated with cardiovascular diseases may include IL1B (interleukin 1, beta), XDH (xanthine dehydrogenase), TP53 (tumor protein p53), PTGIS (prostaglandin 12 (prostacyclin) synthase), MB (myoglobin), IL4 (inteneucine 4) , ANGPT1 (angiopoietin 1), ABCG8 (ATP binding cassette, sub-family G (WHITE), member 8) or CTSK (cathepsin K), for example
Examples of proteins associated with Alzheimer's disease may include the very low density lipoprotein receptor protein (VLDLR) encoded by the VLDLR gene, the ubiquitin-type modifier activating enzyme (UBA1) encoded by the UBA1 gene or NEDD8 activating enzyme E1 enzyme lithic subunit protein (UBE1C) encoded by the UBA3 gene, for example
Examples of proteins associated with Autism Spectrum Disorder may include benzodiazapine receptor (peripheral) associated protein (BZRAP1) encoded by the BZRAP1 gene, member 2 protein of the AF4 / FMR2 (AFF2) family encoded by the AFF2 gene (also called MFR2), the autosomal fragile mental retardation X (FXR1) homologous protein encoded by the FXR1 gene or the autosomal homologous protein 2 of fragile mental retardation X (FXR2) encoded by the FXR2 gene, for example
Examples of proteins associated with Macular Degeneration may include the ATP binding cassette, member protein 4 (ABCA4) of sub-family A (ABC 1) encoded by the ABCR gene, apolipoprotein E protein (APOE) encoded by the APOE gene or chemokine Ligand 2 (CCL2) protein (CC unit) encoded by the CCL2 gene, for example.
Examples of Schizophrenia-associated proteins may include NRG1, ErbB4, CPLX1, TPH1, TPH2, NRXN1, GSK3A, BoNF, DISC1, GSK3B and combinations thereof.
Examples of proteins involved in tumor suppression may include ATM (mutated ataxia telangiectasia), ATR (telangiectasia ataxia and Rad3-related ataxia), EGFR (epidermal growth factor receptor), ERBB2 (viral oncogene homolog 2 of erythroblastic leukemia v- erb-b2), ERBB3 (homologous 3 of viral oncogene of erythroblastic leukemia v-erb-b2), ERBB4 (homologous 4 of viral oncogene of erythroblastic leukemia v-erb-b2), Notch 1, Notch2, Notch 3 or Notch 4, for example.
Examples of proteins associated with secretory disorder may include PSENEN (presenilin stimulator 2 homologue (C. elegans)), CTSB (cathepsin B), PSEN1 (presenilin 1), APP (beta amyloid precursor protein (A4 », APH1 B (homologous B defective 1 anterior pharynx (C. elegans)), P8EN2 (presenilin 2 (Alzheimer's disease 4)) or BACE1 (enzyme 1 of beta-site APP cleavage), for example.
Examples of proteins associated with Amyotrophic Lateral Sclerosis may include 8001 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARoBP (AoN TAR binding protein), VAGFA (endothelial growth factor A vascular), VAGFB (vascular endothelial growth factor B) and VAGFC (vascular endothelial growth factor C) and any combination thereof
Examples of proteins associated with prion diseases may include S001 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (AoN TAR binding protein), VAGFA (vascular Aalial growth factor A ), VAGFB (vascular endothelial growth factor B) and VAGFC (vascular endothelial growth factor C) and any combination thereof
Examples of proteins related to neurodegenerative diseases in prion disorders may include A2M (Alpha-2-Macroglobuline), AATF (Transcription factor for programmed cell death antagonization), ACPP (prostatic acid phosphatase), ACTA2 (Actin alfa 2 de smooth muscle aona), ADAM22 (AOAM metallopeptidase domain), ADORA3 (A3 adenosine receptor) or ADRA10 (Alpha-1 adrenergic receptor for Alpha-1D adrenoceptor), for example
Examples of proteins associated with immunodeficiency may include A2M [alpha-2-macroglobulin); AANAT [arylalkylamine N-acetyltransferase); ABCA1 [ATP binding cassette, sub-family A (ABC1), member 1]; ABCA2 [ATP binding cassette, sub-family A (ABC1), member 2) or ABCA3 [ATP binding cassette, sub-family A (ABC1), member 3); for example .
Examples of proteins associated with Trinucleotide Repetition Trastomos include AR (androgen receptor), FMR1 (1 fragile mental retardation X), HTT (huntingtin) or oMPK (myotonic dystrophy-protein kinase), FXN (frataxin), ATXN2 (ataxin 2 ), for example.
Examples of proteins associated with Neurotransmission Disorders include SST (somatostatin), NOS1 (nitric oxide synthase 1 (neuronal)), AORA2A (alpha-2A-adrenergic receptor), AORA2C (alpha-2C-adrenergic receptor), TACR1 (tachykinin receptor 1) or HTR2c (5C hydroxytryptamine (serotonin 2C) receptor, for example
Examples of neurodevelopmental-associated sequences include A2BP1 [ataxin 2 binding protein 1), AAOAT [amillOadipate aminotransferase), AANAT [arylalkylamine N-acetyltransferase), ABAT [4-aminobutyrate aminotransferase), ABCA1 [ATP binding cassette, sub- family A (ABC1). member 1) or ABCA13 [ATP binding cassette, sub-family A (ABe1, member 13), for example.
More examples of preferred disorders treatable with the present system can be selected from: Aicardi-Goutiéres syndrome; Alexander's disease; Allan-Hemdon + Dudley syndrome; POLG related disorders; Alpha-Mannosidosis (Type 11 and 111); Alstr6m syndrome; Angelman; Syndrome; Ataxia-Telangiectasia; Neuronal Ceroid Lipofuscinosis; Beta-Thalassemia; Bilateral Optic Atrophy and Optical Atrophy (Child) Type 1; Retinoblastoma (bilateral); Canavan disease; Syndrome 1 Cerebroculofacioskeletal [COFS1]; Cerebrotendinous xanthomatosis; Camelia de Lange syndrome; MAPT related disorders; Genetic Prion Diseases; Dravet syndrome; Early-onset Family Alzheimer's Disease; Friedreich's ataxia [FRDA]; Fryns syndrome; Fucosidosis; Fukuyama Coogenous Muscular Dystrophy; Galactosialidosis; Gaucher's disease; Organic Acidemias; Hemophagocytic lymphohistiocytosis; Hutchinson-Gilford Progeria syndrome; Mucolipidosis 11; Child Free Sialic Acid Storage Disease; Neurodegeneration Associated with PlA2G6; Jervell and Lange-Nielsen syndrome; Epidermolysis Bullosa on the Joints; Huntington's disease; Krabbe disease (childhood); Leigh Syndrome Associated with Mitochondrial DNA and NARP; Lesch-Nyhan syndrome; Lysencephaly Associated with LlS1; Lowe syndrome; Urinary Maple Syrup Disease; MECP2 Duplication Syndrome; Copper transport disorders related to ATP7 A; Muscular Dystrophy Related to LAMA2; Arilsulfatase A deficiency; Mucopolysaccharidoses Types 1, 11 or 111; Disorders of the Biogenesis of Peroxisome, Spectrum of Zellweger Syndrome; Neurodegeneration with Iron Accumulation Disorders in the Brain; Acid Sphingomyelinase Deficiency; Niemann-Pick Type C disease; Glycine encephalopathy; ARX Related Disorders; Urea Cycle Trastomos; Osteogenesis Imperfecta Related to COL 1A1 / 2; Mitochondrial DNA Suppression Syndrome; Disorders Related to PLP1; Perry syndrome; PhelanMcDermid syndrome; Type 11 Glycogen Storage Disease (Pompe Disease) (Childhood); MAPT related disorders; Disorders Related to MECP2; Type 1 Rhizomelic Chondrodysplasia Punctata, Roberts Syndrome; Sandhoff's disease; Schindler's disease -Type 1; Adenosine Deaminase Deficiency; Smith-Lemli-Opitz syndrome; Spinal Muscular Atrophy; Spinocerebellar Ataxia of Childhood Beginning; Hexosaminidase A deficiency; Type 1 Tanatophoreic Dysplasia; Disorders Related to Type VI Collagen; Usher Syndrome Type 1; Congenital Muscular Dystrophy; Wolf-Hirschhom syndrome; Lysosomal Acid Lipase Deficiency and Xeroderma Pigmentosa.
Chronic administration of protein therapy may cause unacceptable immune responses to the specific protein. The immunogenicity of protein drugs can be attributed to some epitopes of immunodominant helper T (HTL) lymphocytes. Reducing the MHC binding affinity of these HTL epitopes contained within these proteins may generate drugs with lower immunogenicity (Tangri S, al. CRationally engineered therapeutic proteins with reduced immunogenicity "J Immunol. March 15, 2005; 174 (6): 3.187-96.) The immunogenicity of the CRISPR enzyme in particular can be reduced by following the proposal explained first in Tangri et al with respect to erythropoietin and subsequently developed. Accordingly, direct evolution or rational design can be used to reduce the immunogenicity of the CRISPR enzyme (for example a Cas9) in the host species (human or other species).
In plants, pathogens are often host specific. For example, Fusarium oxysporum f. sp. Iycopersici produces tomato wilting but attacks only tomato and F. oxysporum r. Dianthii Puccinia graminis f sp. Tritici attacks only wheat. Plants have existing and induced defenses to resist most pathogens. Mutations and recombination cases throughout generations of plants lead to genetic variability that results in susceptibility, especially when pathogens reproduce more frequently than plants. In plants there may be resistance to the non-host, for example, the host and the pathogen are incompatible. It can also be Horizontal Resistance, for example, partial resistance against all races of a pathogen, typically controlled by many genes and Vertical Resistance, for example, complete resistance to some races of a pathogen but not to other races, typically controlled by some genes On a gene-by-gene level, plants and pathogens evolve together and genetic changes in one balance the changes in another. Accordingly, using Natural Variability, breeders combine most of the genes useful for Performance, quality, Uniformity, Hardness, Resistance. Sources of resistance genes include natural or foreign varieties, heirloom varieties, mutations related to wild and induced plants, eg, treating plant material with mutagenic agents. Using the present description, plant breeders are provided with a new tool for inducing mutations. Accordingly, a person skilled in the art can analyze the genome of sources of resistance genes and in Varieties with desired characteristics or traits employs the present invention to induce the increase of resistance genes, more accurately than previous mutagenic agents and therefore it accelerates and improves the cultivation programs
As will be apparent, it is envisioned that the present system can be used to target any polynucleotide sequence of interest. Some examples of conditions or diseases that could be treated positively using the present system are included in the above Tables and examples of genes currently associated with these diseases are also provided. However, the exemplified genes are not exhaustive.
Examples
The following examples are provided for illustrative purposes of various embodiments of the invention and are not intended to limit the present invention in any way. The present examples, together with the methods described herein are representative at present of preferred embodiments, are exemplary, and are not intended to limit the scope of the invention.
Example 1. Activity of the CRISPR complex in the nucleus of a eukaryotic cell.
An example CRISPR type 11 system is the CRISPR type II site of Streptococcus pyogenes SF370, which contains a cluster of four Cas9, Cas1, Cas2 and Csn1 genes, as well as two non-coding RNA elements, RNAcrcr and a characteristic series of repetitive sequences (direct repetitions) inter spaced by short stretches of non-repetitive sequences (spacers, approximately 30 bp each). In this system, AoN double chain break (RoC) is set as a target in four sequential stages (Figure 2A). First, two non-coding RNAs, the pre · ARNcr and ARNtracr series, are transcribed from the CRISPR site. Second, ARNtracr hybridizes to direct pre-cRNA repeats, which are then treated in mature cRNAs that contain individual spacer sequences. Third, the ARNcr: mature ARNtracr complex directs Cas9 to the target AON consisting of the proto-spacer and the corresponding PAM via heteroduplex phonation between the RNAcr spacer region and the proto-spacer AON. Finally, Cas9 mediates the cleavage of AoN target upstream of PAM to create an ROC within the proto-spacer (Figure 2A). This example describes an example procedure to adapt this programmable RNA nuclease system to direct the activity of the CRISPR complex in the nucleus of eukaryotic cells
To improve the expression of the CRISPR components in mammalian cells, two Streptococcus pyogenes (S. pyogenes) SF370 site 1, Cas9 (SpCas9) and RNase 111 (SpRNase 1/1) genes were codon-optimized. To facilitate nuclear localization, a nuclear localization signal was included in the (N) amino or (C) carboxyl tenin of both SpCas9 and SpRNase 111 (Figure 2B). To facilitate the visualization of protein expression, a nuorescent protein marker was also included in the N · or C-terminal of both proteins (Figure 28). A version of SpGas9 was also generated with an NLS attached to both Ny C-terminal (2xNLS-SpCas9). Constructions containing SpCas9 NLS-fused and SpRNase 111 were transinfected in human embryonic kidney (HEK) 293FT cells and the relative positioning of the NLS to SpCas9 and SpRNase 111 was found to affect its nuclear localization efficiency . While the C-terminal NLS was sufficient to set SpRNase 111 as a target for the nuclei, the union of a single copy of these particular NLSs in any of the SpCas9 Non-Cterminals could not achieve adequate nuclear localization in this system. In this example, the NLS C · terminal was that of nucleoplasmin (KRPAATKKAGOAKKKK) and the NLS C-terminal was that of the large T40 antigen SV40 (PKKKRKV). Of the SpCas9 versions tested, only 2xNLS-SpCas9 presented nuclear localization (Figure 2B)
The ARNtracr sequence of the S. pyogenes SF370 CRISPR site has two transcriptional start sites, resulting in two transcripts of 89 nucleotides (nt) and 171 nt that are subsequently treated in mature RNAcrcr of identical 75 nt. The shortest 89 nt RNAcrcr was selected for expression in mammalian cells (expression constructs illustrated in Figure 6, with functionality when determined by the results of the Surveryor test shown in Figure 68). Transcription start sites are indicated as +1 and the transcription terminator and the sequence tested by the northem assay are also indicated. The expression of treated siRNA was also confirmed by Northem assay. Figure 7C shows the results of a Northern assay of total RNA extracted from 293FT cells transinfected with U6 expression constructs that support long or short RNAcr, as well as SpCas9 and OR-EMX1 (1) -OR. The left and right panels are of 293FT cells transinfected without or with SpRNase 111, respectively. U6 indicates load control graphically represented with a probe that targets human U6 RNAs. The transfection of the short RNACRR expression construct leads to abundant levels of the treated RNAcr form (-75 bp). Very low amounts of long ARNtracr are detected in Northem's analysis.
To activate precise transcriptional initiation, the U6 activator based on RNA polymerase 111 was selected to drive the expression of RNAcr (Figure 2C). Similarly, a U6 activator-based construct was developed to express a pre-siRNA series consisting of a single spacer flanked by two direct repeats (the ROs, also induced by the term ~ bring pairing sequences "; Figure 2C) . The initial spacer was designed to target a target site of 33 base pairs (bp) (30 bp proto-spacer plus a 3 bp CRISPR unit (PAM) sequence that satisfies the Cas9 NGG recognition unit) at the site Human EMX1 (Figure 2C), a key gene in the development of the cerebral cortex
To test whether heter610ga expression of the CRISPR system (SpCas9, SpRNase 111, RNAcrcr and pre-siRNA) in mammalian cells can achieve the cleavage set as a target of mammalian chromosomes, HEK 293FT cells were transinfected with combinations of CRISPR components. Since RoCs in mammalian nuclei are partially repaired by the non-homologous end junction path (NHEJ), which leads to the formation of insertions or deletions, the Surveyor test was used to detect potential activity of cleavage at the target EMX1 site (see e.g., Guschin el al., 2010, Methods Mol Biol 649: 247). The joint transinfection of the four CRISPR components could induce a cleavage of up to 5.0% in the proto-spacer (see Figure 20). Joint transinfection of all CRISPR components except SpRNase tll also induced up to 4.7% insertion or deletion in the proto- spacer, which suggests that there may be endogenous mammalian Rnases that are capable of aiding cRNA maturation, such as for example related oicer and orosha enzymes. Removing any of the three remaining components cancels the genome excision activity of the CRISPR system (Figure 20). Sanger sequencing of amplicons containing the target site verified the excision activity: in 43 sequenced clones, 5 mutated alleles were found ( 11.6%). Similar experiments using a variety of guide sequences produced insertion or deletion rates as high as 29% (see Figures 4-8, 10 and 11). These results define a three component system for efficient CRISPR mediated genome modification in mammalian cells.
To optimize the efficiency of cleavage, applicants also tested whether different isoforms of RNAcrcr affected the effectiveness of cleavage and found that, in this example system, only the short transcribed form (89 bp) was able to mediate site cleavage. EMX1 human gellÓmico. Figure 9 provides an additional Northern analysis of rRNA treatment in mammalian cells. Figure 9A illustrates a scheme showing the expression vector for a single spacer nanqueado by two direct repetitions (RO-EMX1 (1} -RO). The 30 bp spacer that targets the proto-spacer 1 of the human EMX1 site and the Direct repeat sequences are shown in the sequence below Figure 9. A. The line indicates the region whose reverse complement sequence was used to generate Northern probes for detection of EMX1 cRNA (1). Figure 9B shows a Northem analysis of total RNA extracted from 293FT Iransinfected cells with U6 expression constructs that support RO-EMX1 (1) -RO. The left and right panels are of 293FT cells transinfected without or with SpRNase 111 respectively. RO-EMX1 (1) -RO was treated in mature cRNAs only in the presence of SpCas9 and short RNARcr and was not dependent on the presence of SpRNase 111. The mature RNARcr of total transfected 293FT RNA is -33 bp Y is shorter than the mature pRNA 39-42 bp of S. pyogenes. These results demonstrate that a CRISPR system can be transplanted into eukaryotic cells and reprogrammed to facilitate excision of target poly inudeotides from endogenous mammals.
Figure 2 illustrates the bacterial CRISPR system described in this example. Figure 2A illustrates a scheme showing the CRISPR site 1 of Streptococcus pyogenes SF370 and a proposed mechanism of AIS cleavage mediated by CRISPR by this system. The treated mature rRNA of the direct-spacer repeat series directs Cas9 to genomic targets consisting of complementary proto-spacers and a unit adjacent to the proto-spacer (PAM). In the pairing of the target-spacer base, Cas9 mediates a double chain break in the target AON. Figure 2B illustrates the achievement of Cas9 of s. pyogenes (SpCas9) and RNasa lit (SpRNasa 111) with nuclear localization signals (the NLS) to allow importation into the mammalian nucleus. Figure 2C illustrates the mammalian expression of SpCas9 and SpRNase 111 driven by the constitutive EF1a activator and RNAcr and the pre-RNAcr series (RO-Spacer-RO) driven by the U6 activator of Pol3 RNA to activate precise transcription initiation and termination . A human EMX1 site proto-spacer with a satisfactory PAM sequence is used as the spacer in the pre-siRNA series. Figure 20 illustrates sUNeyor nuclease assay for minor insertions and deletions mediated by SpCas9. SpCas9 was expressed with and without SpRNase 111, ARNtracr and a pre-RNAcr series that supports the EMX1 target spacer. Figure 2E illustrates a schematic representation of base pairing between the target site and cRNA setting as an EMX1 target, as well as an example chromatogram showing microsuppression adjacent to the SpCas9 cleavage site. Figure 2F illustrates identified mutated alleles of sequencing analysis of 43 clone amplicons showing a variety of microinsertions and deletions. Dashed lines indicate deleted bases and non-aligned or mismatched bases indicate insertions or mutations. Scale bar = 10 ~ m.
To further simplify the system of component Ires, a hybrid design of chimeric cRNA-cRNAcr was adapted, where a mature cRNA (comprising a guide sequence) was fused to a partial rRNA via a stem-loop to mimic the cRNA duplex: natural RNAcr (Figure 3A)
The guide sequences can be inserted into Bbsl sites using hybridized oligonucleotides. The proto-spacers in the transcribed and complementary strands are indicated above and below the AON sequences, respectively. A modification rate of 6.3% and 0.75% was achieved for human and Th mouse PVALB sites, respectively, which demonstrates the broad applicability of the CRISPR system in the modification of different sites across multiple organisms. Although cleavage was only detected with one of three spacers for each site using the chimeric constructs, all target sequences were cleaved with insertion or suppression production efficiency that reached 27% when the jointly expressed pre-RNAcr provision was used ( Figures 4 and 5).
Figure 5 provides a further illustration that SpCas9 can be reprogrammed to target multiple germ sites in mammalian cells. Figure 5A provides a schematic of the human EMX1 site showing the position of cyclic proto-spacers, indicated by the underlined sequences. Figure 58 provides a schematic of the pre-cRNA / tRNA complex showing hybridization between the direct repeating region of the pre-cRNA and cRNAcr (upper part) and a scheme of a chimeric RNA design comprising a 20 bp guide sequence And bring and bring pairing sequences that consist of partial direct repeat sequences and hybridized ARNtracf according to the hairpin structure (background). The results of a SUNeyor trial comparing the efficacy of Cas9-mediated cleavage in five proto-spacers in the human EMX1 site is illustrated in Figure 5C. Each pre-spacer is set as the target using treated preRNAcrfRNAtracr complex (cRNA) or chimeric RNA (siRNA).
Since the secondary RNA structure can be crucial for intermolecular interactions, a set of 801lzmann minimum free energy and heavy structure based structure prediction algorithm was used to compare the putative secondary scripting of all guiding sequences used in our experiment. genome target fixation (Figure 38) (see e.g., Gruber et al., 2008, Nucleic Acids Research, 36: W70). The analysis revealed that in most cases, effective guide sequences in the context of chimeric cRNA were substantially free of units of secondary structure, while ineffective guide sequences were more likely to form internal secondary structures that could avoid base pairing. with the target proto-spacer AON. Thus it is possible that variability in the secondary structure of the spacer may impact the effectiveness of CRISPR-mediated interference when a chimeric cRNA is used.
Figure 3 illustrates example expression vectors. Figure 3A provides a schematic of a bicistronic vector to drive the expression of a synthetic ARNcr-RNAcrcr chimera (chimeric RNA) as well as SpCas9. The chimeric guide RNA contains a 20 bp guide sequence corresponding to the proto-spacer at the genomic target site. Figure 38 provides a scheme showing guide sequences that target the EMX1, human PVALB and mouse Th sites, as well as their intended secondary structures. The effectiveness of the modification in each target site is indicated below the drawing of the secondary structure of the RNA (EMX1, n:; 216 amplicon sequencing readings; PVALB, n:; 224 readings; Th, n:; 265 readings). The folding algorithm produced an output with each colored base according to its probability of assuming the expected secondary structure, as indicated by a rainbow scale reproduced in Figure 38 in grayscale. More vector designs for SpCas9 are presented in Figure 3A, including unique expression vectors incorporating a U6 activator linked to an insertion site for a guideworm and a Cbh activator linked to SpCas9 coding sequence.
To test whether spacers that contain secondary structures are capable of functioning in prokaryotic cells in which CRISPRs operate naturally, interference transformation of plasmids supporting proto-spacers in an E. coli strain that heterologously expresses site 1 was tested. S. pyogenes CRISPR SF370 (Figure 3C). The CRISPR site was cloned into an expression vector E. low copy coli and the ARNcr series was replaced with a single spacer flanked by a pair of RO (pCRISPR). E. coli strains were transformed that harbored different pCRISPR plasmids with questioned plasmids containing the corresponding proto-spacer and PAM sequences (Figure 3C). In the bacterial assay, all spacers facilitated effective CRISPR interference (Figure 3C). These results suggest that there may be additional factors that affect the efficacy of CRISPR activity in mammalian cells.
To investigate the specificity of CRISPR-mediated cleavage, the effect of single nudeotide mutations in the guiding sequence on proto-spacer cleavage in the mammalian genome was analyzed using a series of chimeric cRNAs that target EMX1 with single point mutation. (Figure 4A). Figure 48 illustrates the results of a Surveyor nudease assay comparing the efficacy of Cas9 cleavage when paired with different mutant chimeric RNAs. The mismatch of single bases up to 12 bp 5 'of the PAM substantially annulled the gellomic cleavage by SpCas9, while spacers with mutations in upstream positions further retained activity against the target of the original proto-spacer (Figure 48). In addition to the PAM, SpCas9 has unique base specificity within the last 12 bp of the spacer. In addition, CRISPR is able to mediate genomic cleavage as efficiently as a pair of TALE (TALEN) nucleases targeting the same EMX1 proto-spacer. Figure 4C provides a schematic showing the design of the TALENs that target EMX1 and Figure 40 shows a Surveyor gel comparing the efficacy of TALEN and Cas9 (n = 3)
Having established a series of components to achieve CRISPR-mediated gene editing in mammalian cells through the error-predisposed NHEJ mechanism, CRISPR's ability to stimulate homologous recombination (RH), a gene repair pathway of High fidelity to prepare accurate measurements in the genome. Natural type SpCas9 is capable of mediating site-specific ROCs, which can be repaired through both NHEJ and RH. In addition, a substitution of aspartate to alanine (010A) in the RuvC t catalytic domain of SpCas9 was achieved to convert the nudease into a nickasa (SpCas9n; illustrated in Figure 5A) (see e.g. Sapranausaks et al., 2011, Cucleic Acis Research, 39: 9,275; Gasiunas et al., 2,012, Proc Natl. Acad. Sci. USA, 109: E2579), so that the nickeado genomic AON undergoes high-fidelity homology-directed repair (HOR) acronym in English). The Surveyor trial confirmed that SpCas9n did not generate insertions or deletions in the objective of the EMX1 proto-spacer. As illustrated in Figure 58, the joint expression of chimeric cRNA that sets EMX1 as a target with SpCas9 produced insertions or deletions at the target site, while joint expression with SpCas9n did not (n :: 3). On the other hand, the sequencing of 327 amplicooes did not detect any insertion or suppression induced by SpCas9n. The same site was selected to test CRIS-mediated RH by joint transfection of HEK 293FT cells with the chimeric RNA that targets EMX1, hSpCas9 or hSpCas9n, as well as an HR template to introduce a couple of restriction sites (Hindlll and Nhel ) near the proto-spacer. Figure SC provides a schematic illustration of the HR strategy, with relative positions of recombination points and primer hybridization sequences (arrows). SpCas9 and SpCas9n of course catalyzed the integration of the HR template into the EMX1 site. The PCR multiplication of the target region followed by restriction digestion with Hindlll revealed cleavage products corresponding to expected fragment sizes (net in polymorphism gel analysis of restriction fragment length shown in Figure 50), mediating SpCas9 and SpCas9n similar levels of HR efficiencies. The solitants also verified RH using Sanger sequencing of
3D
genomic amplicons (Figure 5E). These results demonstrate the usefulness of CRISPR to facilitate the insertion of genes set as targets in the mammalian genome. Given the objective specificity of 14 bp (12 bp of the spacer and 2 bp of the PAM) of wild-type SpCas9, the availability of a nickasa can significantly reduce the likelihood of out-of-projected modifications, since single-strand breaks do not they are substrates for the NHEJ route with predisposition to errors
Expression constructs that mimic the natural architecture of CRISPR sites with formed spacers (Figure 2A) were constructed to test the possibility of targeting multiplexed sequences. Using a single CRISPR series encoding a pair of spacers that target EMX1 and PVALB, effective cleavage was detected at both sites (Figure 4F, which shows both a schematic design of the ARNcr series and a SurveyOf analysis showing the effective mediation of cleavage). The target suppression of major genomic regions through ROe using spacers versus two targets within EMX1 spaced by 119 bp was also tested and a suppression efficiency of 1.6% was detected (3 of a total of 182 amplicons; Figure 5G). This demonstrates that the CRISPR system can mediate multiplexed editing within a single genome.
Example 2: Modifications and alternatives of the CRISPR system
The ability to use RNA to program sequence specific DNA cleavage defines a new class of genome engineering tools for a variety of research and industrial applications. Various aspects of the CRISPR system can also be improved to increase the efficiency and versatility of the fixation as a goal of CRISPR. The optimal activity of Cas9 may depend on the availability of free Mgz> at levels higher than those present in the mammalian nucleus (see e.g., Jinek et al., 2012, Science, 337: 816) and the preference for NGG unit immediately downstream of the proto-spacer restricts the ability to set an average target every 12 bp in the human genome. Some of these restrictions can be overcome by exploring the diversity of CRISPR sites through the microbial metagenome (see e.g., Makarova et al., 2011, Nat Rev Microbiol, 9: 467). Other CRISPR sites can be transplanted into the mammalian cell medium by a procedure similar to that described in Example 1 The effectiveness of the modification at each target site is indicated below the secondary RNA structures. The algorithm that generates the structures colors each base according to its probability of assuming the expected secondary structure. The 1 and 2 RNA guide spacers induced 14% and 6.4%, respectively. Statistical analysis of excision activity through biological replicates at these two proto-spacer sites is also provided in Figure 7.
Example 3: Sample Target Sequence Selection Algorithm
A computer program is designed to identify candidate CRISPR target sequences in both chains of an input AON sequence based on desired length of guide sequences and a CRISPR unit sequence (PAM) for a specified CRISPR enzyme. For example, the target sites for Cas9 of S. pyogenes, with PAM NGG sequences, can be identified by investigation for 5'-N.-NGG-3 'both in the input sequence and in the inverse complement of the input. Likewise, the target sites for Cas9 of CRISPR1 from S lhermophilus, with PAM NNAGAAW sequence, can be identified by research for 5'-N.-NNAGAAW-3 'both in the input sequence and in the inverse complement of the input. Likewise, the target sites for Cas9 of CRISPR3 of S. thennophilus, with PAM NGGNG sequence, can be identified by investigation for 5'-N.-NGGNG-3 'both in the input sequence and in the inverse complement of the input . The value ~ x "in N. It can be set by the program or specified by the user, such as 20.
Since multiple occurrences in the genome of the target DNA site can lead to unspecified genome editing, after identifying all potential sites, the program filters sequences based on the number of times they appear in the relevant reference genome. For those CRISPR enzymes for which sequence specificity is determined by a "seed" sequence, such as the 5 'of 11-12 bp of the PAM sequence, including the PAM sequence itself, the filtration step may be based on the seed sequence. Thus, to avoid editing in additional genomic sites, the results are filtered based on the number of occurrences of the seed sequence: PAM in the relevant genome. The user can be allowed to choose the length of the seed sequence. The user can also be allowed to specify the number of occurrences of the seed sequence: PAM in a genome for filter pass purposes. The defect is to investigate unique sequences. The level of filtration is modified by changing both the length of the seed sequence and the number of occurrences of the sequence in the genome. The program may additionally or alternatively provide the sequence of a complementary guide sequence to the indicated sequence or target sequences by providing the inverse complement of the identified sequence or target sequences.
Example 4: Evaluation of multiple chimeric ARNcr-ARNtracr hybrids
This example describes the results obtained for the chimeric RNAs (the siRNAs, which comprise a guide sequence, a matching sequence and a bring sequence in a single transcript) that have sequences that incorporate different lengths of wild-type RNAcr sequence. Figure 18a illustrates a schematic of a bicistronic expression vector for chimeric RNA and Cas9 Cas9 is driven by the activator
CBh and the chimeric RNA is co-coded by a U6 activator. The chimeric guide RNA consists of a 20 bp guide sequence (Ns) attached to the tracr sequence (sweeping from the first "U" of the minor strand to the end of the transcript). It is truncated in various positions as indicated. The guide and tracr sequences are separated by the GUUUUAGAGCUA tracr pairing sequence followed by the GAAA loop sequence. The results of the 5 SURVEYOR assays for insertions or deletions mediated by Cas9 at human EMX1 and PVALB sites are illustrated in Figure 18b and 18c, respectively. The arrows indicate the expected SURVEYOR fragments. The siRNAs are indicated by their designation "+ n" and cRNA refers to a hybrid RNA where the guide and bring sequences are expressed as separate transcripts. The quantification of these results. made in triplicate. They are illustrated by histogram in Figures 11a and 11b, which correspond to Figures 10b and 10c, respectively eN. D. "indicates no
10 insertions or deletions detected). The proto-spacer IDs and their corresponding genomic objective. proto-spacer sequence, PAM sequence and chain position are provided in Table o. The guide sequences are designed to be complementary to the entire proto-spacer sequence in the case of separate transcripts in the hybrid system or only to the underlined portion in the case of chimeric RNAs.
Table D:
<dl><dt>Proto-spacer ID </dt><dd>genomic target proto-spacer sequence (5 'to 3') PAM Strand </dd></dl>
<dl><dt>1 </dt><dd>EMX1 GGACATCGATGTCACCTCCAATGACTAG GG TGG + </dd></dl>
<dl><dt>2 </dt><dd>EMX1 CATIGGAGGTGACATCGATGTCCTCCCC AT TGG -</dd></dl>
<dl><dt>3 </dt><dd>EMX1 GGAAGGGCCTGAGTCCGAGCAGAAGAA GGG + </dd></dl>
<dl><dt>GAA </dt><dd /></dl>
<dl><dt>4 </dt><dd>PVALB GGTGGCGAGAGGGGCCGAGATTGGGTGT Te AGG + </dd></dl>
<dl><dt>5 </dt><dd>PVALB ATGCAGGAGGGTGGCGAGAGGGGCCGA TGG + </dd></dl>
<dl><dt>Cat </dt><dd /></dl>
Cell culture and transfection
The human embryonic kidney (HEK) cell line 293FT (Life Technologies) was maintained.
in Modified Eagle Medium from Oulbecco (OMEM) enriched with 10% fetal bovine serum (HyClone), GlutaMAX 2
mM (Life Technologies), 100 U / mi penicillin and 100 Jglml of streptomycin at 3rC with incubation in C02 at 15%
twenty 293FT cells were seeded in 24-well plates (Corning) 24 hours prior to transfection at a density of 150,000 cells per well. Cells were transinfected using Lipofectamine 2000 (Life Technologies) following the protocol recommended by the manufacturer. For each well of a 24-well plate, a total of 500 ng of plasmid was used
SURVEYOR assay for genome modification
25 293FT cells were transinfected with plasmid DNA as described above. Cells were incubated at 3rC for 72 hours post-transfection prior to genomic AON extraction. Genomic AON was extracted using QuickExtract DNA extraction solution (Epicenter) following the manufacturer's protocol. In summary, the button cells were resuspended in QuickExtract solution and incubated at 65 "c for 15 minutes and 98" c for 10 minutes. The genomic region flanking the CRISPR target site for each gene was multiplied.
30 by PCR (primers listed in Table E) and the products were purified using a QiaQuick Spin column (Qiagen) following the manufacturer's protocol. A total of 400 ng of the purified PCR products were mixed with 2 J110X of PCR buffer DNA Taq Polymerase (Enzymatics) and ultra pure water to a final volume of 20 JI Y were subjected to a rehybridization procedure to allow the formation of heteroduplex: 95 "C for 10 min, 95" C to 85 "C falling to -2" Cfs, 85 "C to 25" C to -0.25 "Cfs and 25" C held for 1 minute. After
35 rehybridization, the products were treated with SURVEYOR nuclease and S SURVEYOR stimulator (Transgenomics) following the protocol recommended by the manufacturer and analyzed on 420% Novex TBE polyacrylamide gels (Life Technologies). The gels were stained with SYBR Gold AoN staining (Life Technologies) for 30 minutes and images were formed with a Gel Doc (Bio-rad) gel imaging system. The quantification was based on relative band intensities
Table E:
<dl><dt>primer name </dt><dd>genomic target primer sequence (5 'to 3') </dd></dl>
<dl><dt>Sp-EMX1.F </dt><dd>EMX1 AAAACCACCCTTCTCTCTGGC </dd></dl>
<dl><dt>Sp-EMX1-R </dt><dd>EMX1 GGAGATTGGAGACACGGAGA G </dd></dl>
<dl><dt>Sp-PVALB-F </dt><dd>PVALB CTGGAAAGCCAATGCCTGAC </dd></dl>
<dl><dt>Sp-PVALB-R </dt><dd>PVALB GGCAGCAAACTCCTTGTCCT </dd></dl>
Computational identification of sites set as the objective of unique CRISPR
To identify unique target sites for the SF370 Cas9 (SpCas9) enzyme of S pyogenes in genome
5 From human, mouse, rat, zebrafish, fruit fly and C. e / egans, a computer package was developed to scan both strands of an AON sequence and identify all possible SpCas9 target sites. For this example, each SpCas9 target site was operationally defined as a 20 bp sequence followed by an adjacent NGG proto-spacer (PAM) unit sequence and all sequences that met this 5'-Nro-NGG- definition were identified. 3 'on all chromosomes. To avoid genome editing no
10 specific, after identifying all potential sites, all target sites were filtered based on the number of times they appeared in the relevant reference genome. To take advantage of the sequence specificity of the Cas9 activity conferred by a "seed" sequence, which may be, for example, 5 'sequence of approximately 11-12 bp of the PAM sequence, 5'-NNNNNNNNNN sequences were selected -NGG-3 'to be unique in the relevant genome. All genomic sequences were downloaded from UCSC Genome
fifteen Browser (human genome hg19, mouse genome mm9, genome of rat m5, zebrafish genome danRer7, genome dm4 of D. me / anogaster and genome ce10 of C. e / egans). The results of the full investigation are available to navigate using the UCSC Genome Browser information. An example visualization of some target sites in the human genome is provided in Figure 22.
Initially, three sites were set within the EMX1 site on human HEK 293FT cells. It was evaluated
twenty the effectiveness of genomic modification of each siRNA using the SURVEYOR nuclease assay, which detects mutations that result from breaks in the double strand of DNA (the RDCs) and its subsequent repair by the AON damage repair path of non-homologous end junction (NHEJ). The designated RNAi (+ n) constructs indicate that even the nucleotide + n of wild-type RNAcrcr is included in the chimeric RNA construct, with values of 48, 54, 67, and Y85 used for n. Chimeric RNAs containing fragments larger than
25 Natural-type RNAtracr (siRNA (+67) and siRNA (+85)) mediated AON excision at the three target sites of EMX1, demonstrating siRNA (+85) in particular significantly higher levels of AON excision than corresponding ARNcrfARNtracr hybrids which expressed guide and tracr sequences in separate transcripts (Figures 10b and 10a). Two sites were also targeted at the PVALB site that did not provide detectable cleavage using the hybrid system (guide sequence and tracr sequence expressed as separate transcripts)
30 using ARNquis. SiRNA (+67) and siRNA (+85) were able to mediate significant excision in the two PVALB proto-spacers (Figures 10c and 10b).
For the five objectives in the EMX1 and PVALB, a consistent increase in the efficiency of modifying the genome with increasing length of the tracr sequence was observed. Without wishing to be bound by any theory, the secondary structure formed by the 3 'end of the ARNtracr can play a role in the activation of the CRISPR complex formation rate. An illustration of the secondary structures envisioned for each of the chimeric RNAs used in this example is provided in Figure 21. The secondary structure was predicted using ARNfold (hllp: flrna.tbi.univie.ac.atlcgi-bin / RNAfold.cgi) using minimum free energy and partition function algorithm The pseudocolor for each base (reproduced in grayscale) indicates the pairing probability. Because the siRNAs with longer bring sequences were able to cleave targets that were not cleaved by natural CRISPR RNAcrfRNAtracr hybrids, it is possible that chimeric RNA can be loaded into Cas9 more effectively than its natural hybrid cootrapartide. To facilitate the application of Cas9 for site-specific genome editing in eukaryotic cells and organisms, all unique target sites planned for Cas9 of S. pyogenes were identified computationally in the genomes of the human being, mouse, rat, zebrafish, C. e / egans and D. me / anogaster. Chimeric RNA can be designed for Cas9 enzymes from other microbes
Four. Five to expand the target space of CRISPR RNA programmable nucleases.
Figures 11 and 21 illustrate exemplary bicistronic expression vectors for chimeric RNA expression including up to nucleotide +85 of wild type RNAcrcr sequence and SpCas9 with nuclear localization sequences. SpCas9 is expressed from a CBh activator and ends with the polyA bGH signal (bGH pA). The expanded sequence illustrated immediately below the scheme corresponds to the region surrounding the insertion site of the guide sequence and includes, from 5 'to 3', the 3 'portion of the activator U6 (first shaded region), Bbsl cleavage sites ( arrows), partial direct repetition (pairing sequence bring GTTTIAGAGCTA,
5 underlined), loop sequence GAAA and sequence tracr +85 (sequence underlined after loop sequence). An example guide sequence insert is illustrated below the insertion site of the guide sequence, with nucleotides of the guide sequence for a selected target represented by a ~ N "
The sequences described in the previous examples are as follows (the polynucleotide sequences are 5 'to 3'):
Short siRNA U6 (Streptococcus pyogenes SF370):
10 GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGA GAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTG ACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAAT GGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATC TTGTGGAAAGGACGAAACACCGGAACCATTCAAAACAGCATAGCAAGTTAAAAT AAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCIIIIIII
(bold = ARNtracr sequence; underline = terminator sequence)
Long ARNtracr U6 (Streptococcus pyogenes SF370) · GAGGGCCTATTTCCCATGATTCCTTCATATTrOCATATACGATACAAGGCTGTTAGA
GAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTG ACGTAOAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAAT GGACTATCATATGC1TACCGTAACTTGAAAOTATTTCGArnCTTGGCTTTATATATC TTGTGGAAAGGACGAAACACCGGTAGTATTAAGTATTGTTTTATGGCTGATAAATTT CTTTOAATTTCTCCTTGATTATTTGTTATAAAAGTTATAAAATAATCTTGTTGGAACC ATTCAAAACAGCATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAG
fifteen TGGCACCGAGTCGGTGC II 1 1111
U6-0R-Bbsl main chain-RO (Streptococcus pyogenes SF370)
GAGGOCCTATITCCCATOATTCcrrCAT A'rITOCAT ATACGATACAAOGCTGTTAOA OAOATAAITGOAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTG ACGTAGAAAOTAATAATTTCTTGGOTAOTTTGCAGTTTTAAAATTATGTTTTAAAAT GGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATC lTGTGGAAAGGACGAAACACCGGGTTTTAGAGCTATGCTGTTTTGAATGOTCCCAAA ACGGGTCrrCGAGAAGACGTTTTAGAGCTATGCTGTITroAATGGTCCCAAAAC
U6-chimeric RNA-Bbsl main chain (Streptococcus pyogenes SF370)
GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGA GAOATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTG ACOTAGAAAGTAATAATncrTGGGTAGTlTGCAGlTITAAAATTATGTIlTAAAAT GGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATC TTGTGGAAAGGACGAAACACCGGGTCITCGAGAAGACCTGTITrAGAOCTAGAAAT
AGCAAGTTAAAATAAGGCTAGTCCG
twenty NLS-SpCas9-EGFp
MDYKDHDGDYKDHDIDYKDDDDKMAPKKKRKVGlHGVPAADKKYSIGLDIGTNSVO
WAVITDEYKVPSK.KFKVLGNTDRHSIKK.NLIGALLFDSGETAEATRLK.RTARRRYTRRK
NRICYLQEIFSNEMAK VDDSFFHRLEESFL VEEDKKHERHPIFONIVD EYA YHEKYPTTYH LRKKL VDSTDK.ADLRLIYLALAHMIKFRGHFl.IEGDLNPDNSDVDKLFIQL VQTl'NQLFE ENPINASGVDAKA ILSARLSKSRRLENL TAQLPGEKKNGLFGNLlALSLGL TPNFKSNFDL AEDAKLQLSKDTYDDDLDNLLAQIGDQY ADLFLAAKNLSDAlLLSDILR VNTEITKAPLS A $ M IKRYDEHHQDL TLLKAL VRQQLPEK YKEIFFDQSKNGY AGYIDGGASQEEFYKFIK PILEKMDGTEELLVK.LNREDLLRKQRTFDNGSIPHQIHLGElHAlLRRQEDFYPFLKDNR EKlEKILTFRlPYYVGPlARGNSRfAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTN FDK NLPNEKVLPKHSLL YEYFTVYNEL TKVKYVTEGMRKP AFLSGEQKKAJVDLLFKTN
RKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHOLLKIIKDKDFLDNEENEDILE DIVLTLTLFEDREMIEERLKTY AHLFODK VMKQLURRYTGWGRLSRKUNG1RDKQSG KTILDFLKSDGFAt'lRNFMQLJHDDSL TFKEDlQKAQVSGQGDSLHEHIANLAGSPAIKKG ILQTVKVVDELVK VMGRHKPEN IVIEMARENQlTQKGQKNSRERM KRlEEGIKELGSQI LKEHPVENTQLQNEKL YL YYLQNGRDMYVDQELDINRLSDYDVDHlVPQSFLKDDSlD NKVL TRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLlTQRKFDNL TKAERGGLSE LDKAGFIKRQL VETRQITK.HV AQILDSRMNTKYDENDKUREVKVITLKSKL VSDFRKD F QFYK VREINNY HHAHDA YLNA VVGT ALIKKYPKLESEFVYGDYK VYDVRK.MlAKSEQE IGKATAKYFFYSNlMNFFKTEITLANGEIRKRPLlETNGETGEIVWOKGRDFATVRKVLS MPQVNrvKKTEVQTGGFSKESllPKRNSDKLlARKKDWDPKKYGGFDSPTV A YSVLVV AKVEKGKSKKLKSYKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEONEQKQLFVEQHKH YLDEIIEQISEFSKRVILADANLDKVLSA YNKHRDKP1REQAENII HLFTL TNLGAPAAFK And FDlTIDR.KRYTSTKEVLDATLIHQSITGlYETRIDLSQLGGOAAA YSKGEELFTGYVP IL V ELDGDVNGHKFSVSGEGEGOATI'GKLTLKFICTTGKLPVPWPTL VTTLTYGVQCFSRY P DHMKQHDFFKSAMPEGYVQERTIFFKODGNYKTRAEVKfEGDTL VNRlELKG IDFKED GNILGHKLEYNYNSHN VY lM ADKQKNGIKVNFKJRHNIEOGSVQLADHYQQNTPIGDGP VLLPDNHVLSTQSALSK.DPNEKRDHMVlLEFVTAAGITLGMDELYK
SpCas9-EGFP-NLS:
MDKKYSJGLDIGTNSVGWA VlTDEYKVPSKKFKVLGNTDRHSIK.KNLlGALLFDSGET A EATRLK.RT ARRR YTRRKNRJCYLQEIFSNEMAK VDDSFFHRLEESFL VEEDK.KHERHPIF GNIVDEVA YHEKYPTIY HLRKKL VDSTDKADLRLI YLALAHMIKFRGHFLlEGDLNPONS DVDKLFIQL VQTYNQLFEENPINASG VDAKAILSARLSKSRRLENLlAQlPGEKKNGLFG NLlALSLGL TPNFKSNFDLAEDAKLQlSKDTYDDDLDNLLAQIGOQY ADLFLAAKNLSD AlLLSDILR VNTEITKAPLSASMIKRYDEHHQDl TLLKAL VRQQLPEKYKE1FFDQSKNGY AGYIDGGASQEEFY KFIKPILEKMDGTEELL VKLNREDLLRKQRTFDNGSIPHQIH LGEL HAJ LRRQEDFYPFLKDNREKIEKIL TF RlPYYVGPLARONSRF A WMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNEL TKVKYVTEGMRKP A FLSGEQKKArvOLLFKTNRKVTVKQLKEDYFK.KIECFOSVE1SGVEDRFNASLGTYHOLL KIIKDKDFLDNEENEDlLEDIVL TLTLFEDREMIEERLKTYAHlFDOKVMKQLKRRRYTG WGRLSRKLlNGIRDKQSGKTfi.OFLKSOGfANRNFMQL IHDDSL TFKEDlQKAQVSGQG
DSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTIQKGQKN
KR SRERM IEEGIKELGSQI LKEHPVENTQLQNEKL YLYYLQNGRDMYVDQELDINRLSD YDVDHI VPQSFLKDDSIDNK VLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNA.KLIT QRKFDNLTK.AEROGLSELDKAOFJKRQL VETRQITKHV AQILDSRMNTKYOENOKLIRE VK V1TLKSKL VSDFRKDFQFYKVREJNNYHHAHDA YLNA VVOTALIKKYPKLESEFVYG DYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEffiKRPLlETNGETGEI VWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKK YGGFDSPTV A YSVlVV AK VEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE VKK.DLlI KLPKYSLFELENGRKRMLASAGELQKGNELALPSKVVNFLYLASHYEKLKGS PEDNEQKQLFVEQHKHYLDEJIEQISEFSKR V1LADANLDKVLSA YNKHRDKPIREQAENI IHLFTLTNLGAPAAFKYFDTIIDRKRYTSTKEVLDATLlHQSITGLYETRIDLSQLGGDAA AVSKGEELFTGVVPILVELDGnvNGHKFSVSGEGEGDATYOKLTLKF1CTfOKLPVPWP TLVTTLTYGVQCFSRYPDHMKQHDFFKSAMPEOYVQERTlFFKDDGNYKTRAEVKFEG DTLVNRIElKGIDFKEDGNILGHKLEYNYNSHNVYlMADKQKNGIKVNfKlRHNIEDGSV QLADHYQQNTPIGDGPVLLPDNHYLSTQSALSKDPNEKRDHMVLLEFVTAAOrrLGMD
ELYKKRPAATKKAGQAKKKK
NLS-SpCas9-E GFP-NLS ·
MDYKDHDGDYKDHDlDYKDDDDKMAPKKKRKVGIHGVPAADKKYSIGLD1GTNSVG
WA VITDEYKVPSKK.FKVLGNTDRHSlKKNLlGALLFDSGET AEA TR..LKRT ARRRYTRRK
NRJCYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKK.HERHPfFGNIVDEVA YHEK YPTrYH
LRKKlVDSTDKADlRLIYLAlAHMJKFRGHFLlEGDlNPDNSDVDKLFIQLVQTYNQLFE ENPINASGVDAKA ILSARLSKSRRLENLlAQLPGEKK.NGlfGNLlALSLGLTI'NFKSNFDL AEDA KlQLSKDTYDDDLDNLLAQJGDQYADLFLAAK.NLSDA ILLSOILR VNTETfKAPLS ASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGOASQEEFYKFIK PILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQJHLGELHAILRRQEDFYPFLKDNR EKIEKI L TFRIPYYVGPLARGNSRF A WMTRKSEETITPWNFEEVVDKGASAQSFIERMTN FDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAJVDlLFKTN RKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKllKDKDFLONEENEDILE DNLTLTLFEDREMIEERLKTY AHLFDDKVMKQLKRRRYTGWGRLSR.KLINGIROKQSG KTILDfLKSDGFANRNfMQlDiDDSLTFKEDlQKAQVSGQGDSLHEHIANLAGSPAJKKG ILQTVKVVDEL VK VMGRHKPENIVrEMARENQTTQKGQKNSRERMKR1EEGIKELGSQI
LKEHPVENTQLQNEKLYL VYLQNGRDMYVDQELDINRLSDYDVDHTVPQSFLKDDSTD NKVL TRSDKNROKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGlSE LDKAGFIKRQLVETRQlTKHVAQILDSRMNTKYDENDKLlREVKVITLKSKLVSDFRKDF QFY K VRElNNYHliAJiDAYLNAVVOT ALlKKYPKLESEFVYODYKVYDVRKMJAKSEQE IOKATAKYFFYSNlMNFfKTEITLANGEIRKRPLIETNGETGEIVWDKGROFATVRKVLS MPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTV A YSVL VV AKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELE
NGRKRMLASAGELQKGNELALPSKYVNFLYl..ASHYEKLKGSPEDNEQKQLFVEQHKH
and LDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPlREQAENIIHLFTLTNLGAPAAFKY FDlTIDRKRYTSTKEVLDA TLTHQSITGL YETRIDLSQLGGDAAA ELDGDVNGHKfSVSGEGEGDATYGKlTLKFICTIGKLPVrWPTLVITLTYGVQCFSRY VSKGEELFTGVVPIL V r DHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIOFKED ONILOHK.LEYNYNSHNVYIMAOKQKNOIKVNFKJRHNJEDGSVQLADHYQQNTrIGDGr VLLPDNHYLSTQSALSKDPNEKRDHMVLLEFVTAAGITLGMDELYKKRPAATKKAGQA
KKKK
NLS-SpCas9-NLS 'MDYKDHDGDYKDHDIDYKDDOOKMAPK.KKRKVGIHGVPAADK.KYSIGLDlGTNSVG \ VAVrrDEYKVPSKKFKVLGNTDRHSIKKNLlGALLFDSGETAEATRLKRTARRRYTRRK NRICYLQEIFSNEMAKVOOSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTTYH LRKK.LVDSTDKADLRUYLALAHMIKFRGHFLlEGDLNPDNSDVDK.LFIQL VQTYNQLFE ENPINASGVDAKAll..sARLSKSRRLENLlAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDL AEDA KLQLSKDTYDDDLDNLLAQIGDQY ADLFLAAKNLSDAILLSDlLR VNTEITKAPLS ASMlKR YDEHHQDL TLLKALVRQQLPEK YKEIFFDQSKNGY AGYIDGGASQEEFYKFIK PILEK MDGTEELLVKLNREDLLRKQRTFDNGSI PHQIHLGELHAlLRRQEOFYPFLK.DNR EK1EKIL TFRIPYVVGPLARGNSRFA WMTR KSEETITPWNFEEVVOKGASAQSFlERMTN FDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTN RKVTVKQLKEDYFKKJECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILE DIVLTLTLFEDREMIEERLKiY AHLFDDK VMKQlKRRR YiGWGRLSRKUNGIROKQSG KTILDFLKSDGFANRNFMQLlHDDSLTFKED1QKAQVSGQGDSLHEHIANlAGSPAJK.KG ILQTVKVVDELVKVMGRHKPENfVIEMARENQITQKGQKNSRERMKIDEEG! KELGSQI LKEHPVENTQLQNEKL YLYVLQNGRDMYVDQELDfNRLSDYDVDHrvPQSFLKDDS1D
NKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSE LDKAGFIKRQLVETRQITKHVAQILDSRMNTKYOENDKLIREVKVITLKSKLVSDFRKDF QFYKVREINNYHHAHDA YLNA VVGT ALlKKYPK LESEFVYGDYKVYDVRKM1AKSEQE JGKA T AKYFFYSNIMNFFKTEITLANGE1RKRPLIETNGETGEIVWDKG RDF A TVRKVLS MPQVNTVKKTEVQTGGFSKESILPKRNSDKUARKK.DWDPKKYGGFDSPTV A YSVL VV AKVEKGKSKKLKSVKElLGITIMERSSFEKNPIOFLEAKGYKEVKKDLIIKLPKYSLFELE NORKRMLASAOElQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKH
YLDEJIEQISEFSKRVILADANLDKVLSAYNKHR DKPfREQAENIIHLFTlTNLGAPAAFKY FDTIIDRKRYTSTKEVLDATLlHQSITGL YETRIDLSQLGGDKRPAA TKKA.GQA.K.KKK
NLS-mCherry-SpRNasa3:
MFlFLSLTSFLSSSR TL VSKGEEDNMAlIKEfMRFKVHMEGSVNGHEFEIEGEGEGRPYE
GTQTAKLKVTKGGPLPFAWDILSPQFMYGSKA YVKHPADlPDYLKLSFPEGFKWERVM
NFEDGGVVTVTQDSSlQDGEFIYKVKlROTNFPSDOPVMQKKTMGWEASSERMYPED
GALKGEIKQRLKLKDGGHYDAEVKTTVKA K KPVQLPOA YNVNIKLOITSHN EDYTIVE
QYERAEGRHSTGOMOELYKGSKQLEELLSTSFDlQFNDL TLLET AFTHTSY ANEHRLLN
VSHNERLEFLGDA VLQLlISEYLF AKy PKKTEGDMSKLRSMIVREESLAGFSRfCSFDAYI
KLGKGEEKSGGRRRDTll ... GDLFEAFLGALLLDKGIDA VRRFLKQVMlPQVEKGNFERVK
DYKTCLQEFLQTKGDVAIDYQVISEKGPAHAKQFEVSNVNGAVLSKGLGKSKKlAEQD
AAKNALAQlSEV
SpRNasa3-mCherry-NLS
MKQLEELLSTSFDlQFNDLTLLETAFTHTSYANEHRLLNVSHNERLEFl.GDAVLQLJISEY
LFAKYPKKTEGDMSKLRSMIVREESLAGFSRFCSFDA YIKLGKGEEKSGGRRRDTILGDL
FEAFLGALLLOKG IDA VRRFLKQVMIPQVEKGNFERVKDYKTCLQEFLQTKGDV AlDYQ
VISEKGPAH.AKQFEVSTVVNGA VLSKGLGKSKKLAEQDAAKNALAQLSEVGS VSKGEE
DNMAllKEFM RFK VHMEGSVNGHEFEIEGEGEGRPYEGTQT AK.LKVTKGGPLPF TO WDlL
SPQFMYGSKAYVKHPADlPDYLKLSFPEGFK WERVMNFEDGOVvrvfQDSSLQDGEFl
YKVK.LRGTNFPSDOPVMQK.KTMGWEASSERMYPEDGALKGEJKQRLKLK.DGGHYDAE
VK1TYKAKKPVQLPGA YNVNIKlDlTSHNEDYTIVEQYERAEGRHSTGGMDELYKKRP
AATKKAGQAKKK.K
NLS-SpCas9n-NLS (the nickase mutation D10A e,
lower case):
MDYKDHDGDYKDHDIDYKDDDDKMAPKKKRKVG IHGVPAAOKKYSTGLalGTNSVGW
TO VITDEYK VPSKKFK VLGNTDRI1SIK.KNLIGALLFDSGET AEATRLKR TARRR YTRRKN
RJCYLQEIFSNEMAK VDDSFFHRLEESFL VEEDKKHERHPlFONIVDEV A and HEK YPTIYHL
RKKL VDSTOKADLRLJYLALAHMlK.FRGHFLlEGDLNPDNSDVDKLFIQL VQTYNQLFEE
NPINASGVDAKAJLSARLSKSRRLENlIAQLPGEKKNG LFGNLlALSLGL TPNFKSNFDLA
EDAKLQLSKDTYDDDLDNLLAQTGDQY ADLFLAAKNLSDAlLLSDILR VNTEITKAPLSA
SMIKR YDEHHQDL TLLKALVRQQLPEKYKEIFFDQSKNGY AGYIDGGASQEEFYKFIKPI
LEKMDGTEELL VKlNREDLLRKQRTfDNGSIPHQIHLGELHAILRRQE DFYPFLKDNREK
IEKIL TFRIPYYVGPLARGNSRFA WMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFD
KNLPNEKVLPKHSLL YEYFTVYNElTKVKYVTEGMRKPAFLSGEQKKA TVDLlFKTNR
KVTVKQLKEDYFKKlECFDSVEISGVEDRfNASLGTYHDLLKlJKDKDFLDNEENEDILE
DTVLTLTLFEDREMIEERLKTY AHLFDDK VMKQLKRRRYTGWGRLSAAlfNGIRDKQSG
KTILDFLKSDGF ANRNFMQllHDDSlTFKEDlQK.AQVSGQGDSLHEHlANLAGSPA IKKG
ILQTVKVVDElVKVMGRHKPENIVIEMARENQTIQKGQKNSRERMKRlEEGIKELGSQI
LKEH PVENTQLQNEKL YLYYLQNGRDMYVDQELDINRLSDYDVDHNPQSFLKODSID
NKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLlTQRK.FONLTKAERGGLSE
LDKAGFIKRQLVETRQITKJNAQILDSRMNTKYDENDKLIREVKVITLKSKL VSDFRKDF
QFYK VREINNYHHAHDA YLNA VVGT ALlKKYPKLESEFVYGDYK VY DVRK}. {IAKSEQE
IGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLS
MPQVNIVKKTEVQTGGFSKESIlPKRNSDKLIARKKDWDPKK YGGFDSPTV TO YSVL VV
AKVEKGKSKKLKSVK.ELLOITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSlFELE
NGRK.RMLASAGELQKGNELALPSK and VNFL YLASHYEKLKGSPEDNEQKQLFVEQH KH
YLDE1IEQlSEFSKRV1LAOANLDK VLSA YNKHRDKPIREQAENIIHLFTL TNLGAPAAFK Y
FO "ITIDRKRYTSTK.EVLDATL1HQSITGL YETRJDLSQLGGDKRPAA TK KAGQAKKKK
hEMX1-Template
GAATGCTGCCCTCAGACCCGCTICCTCCCTGTCCITGTCTGTCCAAGGAGAATGAGG TCTCACTGGTGGATTTCGGACTACCCTGAGGAGCTGGCACCTGAGGGACAAGGCCC CCCACCTGCCCAGCTCCAGCCTCTGATGAGGGGTGGGAGAGAGCTACATGAGGTTG CTAAGAAAGCCTCCCCTGAAGGAGACCACACAGTGTGTGAGGTTGGAGTCTCTAGC AGCGGGTTCTGTGCCCCCAOGGATAGTCTOGCTGTCCAGGCACTGCTCTIGATATAA ACACCACCTCCTAGTTATGAAACCATGCCCATTCTGCCTCTCTGTATGGAAAAGAGC
RH-Hindll-Nhel "
ATGGGGCTGGCCCGTGGGGTGGTGTCCACTTTAGGCCCTGTGGGAGATCATGGGAA CCCACGCAGTGGGTCATAGGCTCTCTCA1ITACTACTCACATCCACTCTGTGAAGAA GCGA rrATGATCTCTTCTc r AGAAACTCGTAGAGTCCCATGTCTGCCGGCTICCAGA GCCTGCACTCCTCCACCTTGGCTTGGCTTTGCTGGGGCTAGAGGAGCTAGGATGCAC AGCAGCTCTGTGACCCTTTGTTTGAGAGGAACAGGAAAACCACCCTTCTCTCTGGCC CACTGTGTCCTCTTCCTGCCCTGCCATCCCCTTCTGTGAATGTTAGACCCATGGGAGC AGCTGGTCAGAGGGGACCCCGGCCTGOGGCCCCTAACCCTATGTAGCCTCAGTCTTC CCATCAOGCTCTCAGCTCAGCCTGAGTGTTGAGGCCCCAGTGOCTGCTCTGGGGGCC TCCTOAGTTTCTCATCTGTGCCCCTCCCTCCCTGGCCCAGGTGAAGGTGTGGTTCCAG AACCGGAGGACAAAGTACAAACGGCAGAAGCTGGAGGAGGAAGGGCCTGAGTCCG AGCAGAAGAAGAAGGGCTCCCATCACATCAACCGGTGGCGCATTGCCACGAAGCAG GCCAATGGGGAGGACATCGATGTCACCTCCAATGACaagcttgclagcGGTGGGCAACCAC AAACCCACGAGGGCAGAGTGCTGCTTGCTGCTGGCCAGGCCCCTGCGTGGGCCCAA GCTGGACTCTGGCCACrCCCTOOCCAGOCTTTGGGGAGGCcrGGAGTCATGGCCCCA CAGGGCTTGAAGCCCGGGGCCGCCATTGACAGAGGGACAAGCAATGGGCTGGCTGA GGCCTGGGACCACTTGGCCTTCTCCTCGGAGAGCCTGCCTGCCTGGGCGGGCCCGCC CGCCACCGCAGCCTCCCAGCTGCTCTCCGTGTCTCCAATCTCCcnTIGTITIGATGC ATTTCTGTTTTAA1TTATTTTCCAGGCACCACTGTAGTTTAGTGATCCCCAGTGTCCC CCTTCCCTATGGGAATAATAAAAGTCTCTCTCTTAATGACACGGGCATCCAGCTCCA GCCCCAGAGCCTGGGGTGGTAGATTCCGGCTCTGAOGGCCAGTGGGGGCTGGTAGA GCAAACGCGTTCAGGGCCTGGGAGCCTGGGGTOGGGTACTGGTGOAGGGGGTCAAG GGTAATTCATTAA CTCCTCTCTTTTGn "GGGGGACCCTGGTCTCTACCTCCAGCTCCA CAGCAGOAGAAACAGGCTAGACATAGGOAAGGGCCATCCTGTATCTTGAGGGAGGA CAGGCCCAGGTCTTTCTTAACGTATTGAGA GGTGGGAATCAGGCCCAGGTAGTTCAA TGGGAGAGGGAGAGTGCTTCCCTCTGCCTAGAGACTCTGGTGGCTTCTCCAGliGAG GAGAAACCAGAGGAAAGGGGAOGATTGGOOTCrGGGGGAGGGAACACCATICACA AAGGCTGACGGTTCCAGTCCGAAGTCGTGGGCCCACCAGGATGCTCACCTGTCCTTG GAGAACCGCTGGGCAGGTTGAGACTGCAGAGACAGGGCTTAAGGCTGAGCCTGCAA CCAGTCCCCAGTGAcrCAGGGccrCCTCAGCCCAAGAAAGAOCAACGTGCCAGGGC CCGCTGAGCTCTTGTGTTCACCTG
NLS-StCsn 1-NLS:
MKRPAATKKAGQAKKKKSDLVLGLDIGIGSVGVG1 LNKVTGEll HKNSRJFPAAQAENN L VRRTNRQGRRLARRKKHRR VRLNRLFEESGLn'DFrKISrNLNPYQLRVKGL TDELSNE ELFlALKNMVKHRGISYLDDASDDGNSSVGDYAQIVK.ENSKQLETKTPGQIQLERYQTY GQLRGDFTVEK.DGKKHRLINVFPTSA YRSEALRlLQTQQEFNPQITDEFINRYLEIL TGKR KYYHGPGNEKSRTDYG RYRTSGETLONlFGLLIGKCTFYPDEFRAAKASYTAQEFNLLND LNNL TVPTETK.KLSKEQKNQIJNYVKNEKAMGPAKLFKY1AKLLSCDYADIKGYRIDKS OKAE1HTFEA YRKMKTLETLDIEQMDRETLDKLA YVL TLNTEREG lQEALEHEF ADGSFS QKQVDEL VQFRKANSSIfGKGWHNFSVKLMMELlPEL YETSEEQMTIL TRLGKQKITSS SNKTKYIDEKLLTEEIYNPWAKSVRQAlKIVNAAlKEYGDFDNMEMARETNEDDEKK AlQK lQKA NK DEK DAAt '-. TLKAANQYNGKAELPHSVFHGHKQLATKlRL WHQQGERCL and. TGKTlSJHDLINNSNQfEVDHILPLSITF'DDSLANK VLVY TO TANQEKGQRTPYQALDSMO DA WSFRELUFYRESKTLSNKKKEYLL TEED1SKfDVRKKFIERNL VOTR and ASR VVLNA LQEHFRAHKJOTK VSVVRGQFrSQLR.R.HWGIEKTRDTYHHHA VDALI IAASSQLNL WKK QKNTLVSYSEDQLLDIETGELISDDEYKESVFKAPYQHFVDTLKSKEFEDSILFSYQVDSK FNRKISDATIY ATRQAKVGKDKADETYVLGK1KDIYTQDGYDAFMKIYKKDKSKfLMY RHDPQTFEKV1EPILENYPNKQINEKGKEVPCNPFLKYKEEHGYIRKYSKKGNGPEIKSLK YVDSKLGNHIDTTPKDSNNK YVLQSVSPWRADVYFNKlTGKYEILGLKy ADLQFEKGT GTYKJSQEK YNDIKKKEOVDSDSEFKFTL YKNDLLL VKDTETKEQQLFRFLSRTMPKQK I-IYVELKPYDKQKFEGGEALIKVLGNVANSGQCKKGLGKSNISIYKVRTDVLGNQHIIKN EGDKPKLDFKRPAATKKAGQAKK.KK
U6-St ARNtracr (7 -97):
GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGA GAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTG ACGTAGAAAGTAATAATTTCITGGGTAGTlTGCAGTITTAAAAlTATGITITAAAAT GGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATC 'ITGTGGAAAGGACGAAACACCGTTACTTAAATClTGCAGAAGCTACAAAGATAAOG CnCATGCCGAAATCAACAccc rGTCATlTIATGGCAGGGTGTTTTCGTTATITAA
U6-RD-Spacer-RD (S. Pyogenes SF370)
gagggcctamcccatgaltccncalatt1gcatatacgaf ~ caaggctgnagagagalaall ggaallaalllgactgtaaacacaaagalal
tagtacaaaalacgtgacglagaaagtaalaamcltgggtaglltgcaglltlaaaattalgttnaaaatggactatcatalgcttaccgtaaclI
5 gaaagtamcgamCtlggclttatatalcttgtggaaaggacgaaacaccgglrttttagagctatgctgtttlgaalggtcccaaaacNNN
NlINl'INNlINl'INNlJN1> INNlJN1> INNrNNNNlI.lNgttltagagcI31gcIgtltlgaatggtcccaaaacTTTTTTT (lowercase underline = direct repeat; N -sequence sequence; bold -terminator)
Chimeric RNA containing +48 RNAtracr (S pyogenes SF370)
1 O gagggcctatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatat tagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttaccgtaactt gaaaglatttcgafflettggctttalalatcttglggaaaggacgaaacaccNNNNNNNNNNNNNNNNNNNNgttttaga gctagaaatagcaagttaaaataaggctagtccg 11111 11 (N = leader sequence; underline = first matching sequence bring; second underlined -sequence bring; bold = terminator)
fifteen chimeric RNA containing ARNtracr +54 (S pyogenes SF370) gagggcctatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatat tagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagtttta aaattatgttttaaaatggactatcatatgctta ccgtaactt gaaaglatttcgatttcttggctttatatatcttglggaaaggacgaaacaccNNNNNNNNNNNNNNNNNNNNgttttaga ~ 1I 1I 1I aaatagcaagttaaaataaggclaglccgltatca I 1 (N = leader sequence; first underlined sequence =
twenty matching bring; second underline - sequence bring; bold = terminator) chimeric RNA containing ARNtracr +67 (S pyogenes SF370) gagggectatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatat tagtacaaaatacglgacgtagaaagtaataaltlcttgggtagltlgcagtltla aaattalgltltaaaalggactatcatalgctta ccgtaactt gaaaglatttcgattlcttggclttalalatcttglggaaaggacgaaacaccNNNNNNNNNNNNNNNNNNNNgttttaga
5 ~ aaalagcaagtlaaaataaggclagtccgttatcaacttgaaaaagtg 1111 11 1 (N = guide sequence; first underline = matching sequence bring; second underline -sequence to bring; bold = tenninator)
chimeric RNA containing ARNtracr +85 (S pyogenes SF370) gagggectatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatat tagtacaaaatacglgacgtagaaagtaataaltlcttggglagtttgcagtltlaaaattalgttttaaaaIggactatcatalgcttaccgtaactt
1 Or gaaaglatttcgattlcttggclttatatatcttglggaaaggacgaaacaccNNNNNNN NN NN NN NN NN N NNgttttaga ~ aaalagcaagtlaaataaggclagtccgttatcaacttgaaaaagtggcaccdsequence second sequence;
CBh-NLS-SpCas9-NLS
CGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCC ATTGACGTCAATAATGACGTATCljCCCATAGTAACCCCAATACCGACTTTCCATTG ACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTA
TCATATGCCAAGTACGCCCCCTATTGACGTCAATGACOOTAAATGGCCCOCCTGGCA TIATGCCCAGTACATGACCITATGGGACTITCCTAcn'GGCAGTACATeTACOTATTA OTCATeGCTATTAeCATGGTeGAGGTGAGCCCCACGTTCTGCTTCACTCTCCCCATCT CCCCCCCCTCCCCACccccAArrrroTATITATITA'1 111 IIAAITAITITQTQCAGCO ATGGGGGCGGGGGGGGOOOGGOGOCOCGCGCCAGOCGGGGCGOOGCGGGGCGAG GGGCGGGOCGOGGCGAGGCGGAGAGGTGCGGCO (; CAOCCAATCAGAGCGGCGCGC TCCGAAAGTTTCCTTTTATGGCGAGGCGGCGGCGOCGGCGGCCCTATAAAAAGCGA AGCGCGCGGCGGGCGGGAGTCGCTGCOACGCTGCCTTCGCCCCOTOCCCCGCTCCG CCGCCGCCTCGCGCCGCCCGCCCCGGCTCTGACTGACCGCGTTACTCCCACAGOTGA
GCGOGCGGGACGGCCCTTCTCCTCCGGOCTGTAATTAGCTGAGCAAOAOGTAAOGO TTTAAGGGATGOTTGOTTGGTGOOGTATTAATGTrrAATTACCTOGAGCACCTGCCT GAAATCAC III! Ill CAGOTTGGaccgglgccaccATGGACTATAAGGACCACQACGGAGA CTACAAOGATCATOATATTGATTACAAAGACGATGACGATAAGATGGCCQCAAAGA AGAAGCGGAAGGTCGGTATCeACOGAGTCCCAOCAOCCGACAAGAAGTACAOCATC GGCCTGGACAICGGCACCAACTCTGTGGGCTGGQCCGTGATCACCGACGAGTACAA GGTGCCCAOCAAGAAAlTCAAGGTGCTGGGCAACACCQACCGGCACAGCAICAAGA AGAACCTGATCGGAGCCCTGCTQTTCGACAGCGGCGAAACAGCCOAGGCCACCCGG CTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAOAACCGOATCTGCTATCT GCMGAGATCITCAGCAACOAGATGGCCAAGGTGGACGACAGcrrCITCCACAGAC TGQAAQAGICClTCCTGGTGGMGAGGATMQMGCACGAGCGGCACCCCATCITC GGCAACATCOTGGACGAGGTGGCCIACCACOAGAAGTACCCCACCATCTACCACCI GAGAAA GAAACTGGTQGACAGCACCGACAAGOCCGACCTGCGGCIGAICIATCTGG CCCTGGCCCACATGATCAAGTTCCGGGGCCACTICCTGATCOAOGOCGACCIGAACC CCGACAACAGCGACGTGGACAAGCfGTTCATCCAGCTQGTGCAGACCTACMCCAG CTGITCGAGGAAMCCCCATCAACGCCAGCOOCGTGGACGCCAAGGCCATCCTGTC TGCCAGACTGAGCAAGAGCAGACGGCTGGAAMTcroATCGCCCAGCfOCCCGGCG AGMGAAGAATGGCCTGTICGGCAACCTGATIGCCCTGAGCCTGGGCCTGACCCCC AACTICAAGAGCAACTTCGACCTGGCCGAGGATGC: CAAACTGCGACGGCAGCGACGGGGGGGGGGGGGGGG
ACCTOITTCTOOCCGCCAAGAACCTGTCCGAQGCC¡ ~ TCCTGCTGAGCGACAICCTGA
GAGTGAACACCGAGATCACCAAGGCCCCC (TGAGCGCCTCTATGATCAAGAGATAC
GACGAGCACCACCAGGACcrGACCCTGCTGAAAGCTcrcOTGCGGCAGCAGCIGCC TGAGAAGTACAAAGAGAITTTCTICGACCAGAGCAAGAACGGCIACGCCGGCIACA TIGACGGCGGAGCCAGCCAGGAAGAGITCfACAAGITCATCMGCCCATCCTGGAA AAGATOGACGOCACCOAGGMCTOCTCGTGMOCrGMCAOAOAGOACCIOCTOCG GAAGCAGCOGACCITCGACAACOGCAGCAICCCCCACCAGAICCACCfOOOAGAGC TGCACGCCATTCTGCGGCGGCAGGAAGATTTTTACCCAITCCTGAAGGACAAccGGO AAAAOATCOAGMOATCCIOACCTTCCGCAICCQCTACIACGTOGGCCCTCTGGCCA GGGGAAACAGCAGATICGCCTGGAIOACCAGMAGAGCGAGGAAACCATCACCCC CTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCrrCCGCCCAGAGCTTCATCGAGC GGATGACCAAClTCGATAAGAACCTGCCCMCGAOMGGTGCTGCCCAAGCACAGC CTGerGIAQGAGTACTICACCGTGTATAACGAGCIGACCAAAGTGAAATACGTGACC OAGGGMTGAGAMGCCCGCCTICCIGAGCGOCGAGCAGAAAAAGGCCAICGIOQ ACCTGCIGTTCAAGACCAACCGGMAGTGACCGTGMGCAGCTGMAGAGQACIAC lTCAAGAAAATCGAGTGC1TCGACTCCQTGQAAATCTCCGGCGTGGAAGATCGGITC AACGCCICCCDQGGCACAIACCACGATCTOCTGAAAATTATCAAGGACAAGGACTI CCIGGACAAIGAGOAAMCGAGGACATTCTGGMOAIATCOIGCTGACCCTGACAC IQITroAGGACAGAGAGATGAICGAGGAACOGCTGMAACCTATGCCCACcrGTTC OACOACAAAGIOAIOAAOCAGCIGAAOCGGCGOAGAIACACCGGCIGGGGCAOGC TGAoccOGAAGCTQAICAACOOCATCCGGGACAAGCAOTCCGOCAAGACAATCCTG GATTICCIGAAGICCGACGGCTrCGCCAACAGAA6CITCATGCAGCIGATCCACGAC GACAGCCTGACCTITAAAGAOGACAICCAGAAAGCCCAGGIGTCCOGCCAGGOCQA TAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCAITAAGAAGGGCA TCCTGCAGACAGTGAAGGTGGTQGACGAGCTCGTGAAAGTOAIGGOCCGGCACAAG CCCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAG6AGGGAC AGMGMCAGCCGCGAGAGAATGMGCGGATCGAAGAGGGCATCAAAGAGCIGGG CAGCCAGATCCIGAAAGAACACccCGTGGAAAAC¡ \ CCCAGCTGCAGAACGAGAAG CTGTACCTGTACTACCTGCAGMIGGGCGOGAIATGTACOTGGACCAGGAACTGGA CAICAACCGGCTQTCCGACIACOATGTGGACCAIATCGTGCCICAGAGCTTTCTGAA GGACOACTCCATCGACMCMGGTGCTGACCAGAJ \ GCGACMOAACCGGGGCAAG
AGCGACAACGIGCCCICCGAAGAGGTCGIGAAOAJ ~ GAIGMGAACTACTGGCGGCA
GCIOCTGMCGCCAAGCfGATIACCCAGAGAAAGTTCGACAATCTGACCAAGGCCG
AGAGAGOCOOCCIGAGCQAACTOOATAAOQCCOGCTTGATCAAGAQACAOCTOOTG oAAACCCGGCAGATCACAAAGCACGTGGCACAOATCCTQGACTCcCQGATQAAcAc TAAGIACGACGAGAATQACAAGCIGATCCGGGAAGTGAAAGTGATCACCCIQAAGT CCAAGCTGGTGICcOATTTCCGGAAGGATTICCAGTITTACAAAGTGCGCGAGATCA ACAACTACCACCACGCCCACGACGCCTACO <; AACOCCGTCQTQQGAACCGCCCTG ATCAAAAAGTACCCIAAGCIQGAAAGCGAGTJCQTGTAQQGCGACIACAAGGTGTA CGACGTGCGGAAGATGATCGQCAAGAGQQAGCAGGAAATCOGCAAGGCIACCGCC AAGTACITCTTCTACAGCAACATCATGAA (II II ICAAGACCOAOATTACCCTGGCC AACGGCGAGATCCGGAAGCGGCCTCIGATCGAGACAAAQGGCOAAACCGGGGAGA TCGTGTOGGAIAAOGGCCQGGATnTGCCACCOTGCQGAAAGTGCTQAGCATGQCC CAAGTGAATAICGTG6AAAAGACCQAQGTGCAGACAOOCOGCTICAOCAAAGAGTC TATcCTQccCAAGAGGAACAGCGATAAGCTQATCGCCAGAAAGAAGGACIGGGACC CTAAGAAGTACGGCGGCTICGACAGQCCCAQCGTGGCCIATTCIQTGCIGGTGGTGG CCAAAGTGGAAAAQQQCAAGTCCAAGAAACTGAAGAGTGTGAAAGAGCTGCIGGG GATCACCATCATQQAAAOAAOCAQCJTCGAGAAGAATCCCATCGACTTICTGGAAG CCAAGGGCIACAAAGAAGTQAAAAAGGAcCTQATCATCAAGCIGcCTAAGTACICC CIGTICGAGCTQGAAAACGGCCGGAAGAGAATGCIGGCCICIGCCGGCGAACIGCA GAAGGGAAACGAACTQGcQCTGccCICcAAATATGTGAAcTTCCIGTAQCTGGCCA GccACIATGAGAAGCTQAAGGGCICcccQQAGGAIAATGAGCAGAAACAGCTGTTT GTGGAACAGCACAAGCACTACCTQGACGAGAICAJCQAQCAGAICAGCGAGTTCTC CAAGAGAGIGATCCTGGCCQACGCIAATCIOGACAAAGIGCTGICCGCCTACAACA AGCAccooQAIAAGCCCAICAGAQAGCAGQQCGAGAAIATCATCCACCIGTTTACC CIGACCAATCIGGGAGCCCCIGCCQCCITCAAGTAC'rl'TGACACCACCATCGACCGQ AAOAGGTACACCAGCACCAAAGAGGIOCTGGACGQCAcccTGATCCACCAGAGCAT CAQCGGCCTQTACGAGACACQGATQGACCTGTCICAGCTGGGAGGCOACIIICJI JI TCITAGCITGACCAGCTITOTAGIAOCAGCAGGACGCTITM (underlined -NLS
hSpCas9-NLS)
RNA chimeric example for S. thermophilus CRISPR1 Cas9 LMD-9 (with PAM NNAGAAW) NNNNNNNNNNNNNNNNNNNNglttttglaclctcaagatttaGAAAtaaatcttgcagaaoctacaaagataaggctt catgccaaaatcaacaccclglcattllalggcagaalgtttloottatuaa I 1I 1I I (N -sequence greedily; underline = first matching sequence bring; second sequence underlined = bring; bold = terminator)
RNA chimeric example for S. thermophilus LMD-9 $ CRI PRl Cas9 (with PAM NNAGAAW) NNNNNNNNNNNNNNNNNNNNgltlttglactclcaGAAAlgcagaagclacaaagataaggcttcalgcooaaatca 1II acaccctgtcattttataacaoggtgttttoottaltlaa I I 1 (N -sequence guide; first underlined sequence pairing bring second underlined sequence = bring; bold = terminator)
Example chimeric RNA for S. thermophilus lMD-9 CRISPR1 Cas9 (with PAM of NNAGAAW) NNNNNNNNNNNNNNNNNNNNgltttlgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaalca acacccado sequence of the second sequence of the first sequence;
Example chimeric RNA for s. thermophilus CRISPR1 Cas9 LMD-9 (oon NNAGAAW PAM) NNNNNNNNNNNNNNNNNNNNgttatlgtactctcaagaltlaGAAAtaaalctlgcagaagctacaaagalaaggcttcatgccgaaalcaacaccclgtcattttatggcagogtgttttoottatttaa II 1I II (N greedily -sequence; underline = first matching sequence bring; second underlined -sequence bring; bold -terminador)
1 OR
Example chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (with PAM of NNAGAAW) NNNNNNNNNNNNNNNNNNNNgttattgtactctcaGAAAtgcagaaectacaaagataaoocttcatgccgaaatc aacaccattase sequencing or transduction of the second sequence;
Example chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (with PAM of NNAGAAW) NNNNNNNNNNNNNNNNNNNNgltaltgtactctcaGAAAtgcagaagctacaaagataagoctlcatgccgaaatc aacacctsese transsection of the second sequence;
Example chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (with PAM of NNAGAAW) NN NN NNNN NN NN NN NNN NN NgltaltglaclctcaagatttaGAAAlaaatcltgcagaagclacaatgataagqctt catgccaaacacatate sequence sequencing 11th sequence underlining (first sequence 11); bring; bold-terminator)
RNA chimeric example for S. thermophilus CRISPR1 Cas9 LMD-9 (with PAM NNAGAAW) NNNNNNNNNNNNNNNNNNNNgttattgtactctcaGAAAtgcagaagctacaatgataaggcttcatgccgaaatca acaccc1gtcattttatggcagggtgttttcgttatttaa 1IIIII (N -sequence guide; underline = first matching sequence bring; second underlined -sequence bring; bold = terminator)
Example chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (with NNAGAAW PAM) NNNNNNNNNNNNNNNNNNNNgltaltgtactctcaGAAAtgcagaagctacaatgataasscltcatgccgaaatca acaccctgtcatsequencing sequence;
Chimeric RNA sample for S thermophilus LMD-9 CRISPR3 Cas9 (PAM COfl NGGNG) NNNNNNNNNNNNNNNNNNNNgttttagagctgtgGAAAcacagcgagttaaaataaaoctlagtccgtactcaactt gaaaaggtggcaccgattcggtgtll III I (N -sequence guide; -sequence first underscore pairing bring; second underlined -sequence bring; bold ;;; terminator)
Codon-optimized version of Cas9 from CRISPR3 LMD-9 site of S. fhermophilus (with an NLS at both 5 'and 3' ends)
ATGAAAAGGCCGGCGGCCACGAAAAAGGCCGGCCAGGCAAAAAAOAAAAAOACCA
AGCCCTACAGCATCGGCCTGGACATCCGCACCAATAGCGTGGGCTGGGCCGTGACC
ACCGACAACTACAAGGTGCCCAGCAAGAAAATGAAOGTGCTGGGCAACACCTCCAA
OAAGTACATCAAOAAAAACCTGCTGOGCOTGCTGCTGTTCGACAGCOGCATTACAG
CCGAGGGCAOACGGCTOAAGAGAACCGCCAGACGGCGGTACACCCGGCGGAOAAA
CAGAATCCTGTATCTGCAAGAGATCTTCAOCACCGAGATGGCTACCCTGGACGACG
CCTTCTTCCAGCGGCTGGACGACAGCTTCCTGGTGCCCGACGACAAGCGGGACAGC
AAGTACCCCATCTTCGGCAACCTGGTGGAAGAGAAGGCCTACCACGACGAGTTCCC
CACCATCTACCACCTGAGAAAGTACCTGGCCGACAGCACCAAGAAGGCCGACCTGA
GACTGGTGTATCTGGCCCTGGCCCACATGATCAAGTACCGGOGCCACITCCTGATCG
AOCGCGAGTTCAACAGCAAGAACAACGACATCCAGAAGAACTTCCAGGACTTCCTG GACACCTACAACGCCATCTTCGAGAGCGACCTGTCCCTGGAAAACAGCAAGCAGCT GGAAGAGATCGTOAAOOACAAGATCAGCAAOCTGGAAAAGAAGGACCGCATCCTG AAGCTGTTCCCCOOCGAGAAGAACAOCOOAATCnrCAOCGAGTTTCTGAAOCTGAT CGTOOGCAACCAOGCCGACITCAGAAAGTGCTTCAACCTGGACOAGAAAOCCAGCC TOCAClTCAGCAAAGAGAOCTACGACOAOGACCTGOAAACCCTGCTGGGATATATC
GGCGACGAcrACAGCGACGTGTTCCTGAAGGCCAJ ~ GAAGCTGTACGACGCTATCCT
GCTOAOCOGClTccrGACCGTGACCGACAACGAGACAGAGGCCCCACTGAGCAGCG CCATGATTAAOCGGTACAACGAGCACAAAGAGGATCTGGCTCTGCTGAAAGAGTAC ATCCGGAACATCAGCCTGAAAACCTACAATGAGOTGlTCAAGGACGACACCAAGAA CGOCfACGCCOGCTACATCGACOOCAAGACCAACCAGGAAGATITCTATGTGTACC TOAAGAAGCTOCTGGCCGAGTICGAGGGGGCCGACTACTITCTGGAAAAAATCGAC CGCOAGGATTTCCTGCGGAAGCAOCGGACCTTCGACAACGGCAGCATCCCCTACCA OATCCATCTGCAGGAAATGCOOGCCATCCTGGACAAGCAOGCCAAGTTCTACCCAT TCCTGGCCAAGAACAAAGAGCGGATCGAGAAGATCCTGACCTTCCGCATCCCTTACT ACOTGGGCCCCCTOGCCAGAGGCAACAGCGATnnúCCTGGTCCATCCOGAAGCGC AATGAGAAOATCACCCCCTGGAACTICGAOGACGTGATCGACAAAGAGTCCAGCGC CGAGGCCTTCATCAACCGGATGACCAGCTTCGACCTGTACCTGCCCGAGGAAAAGG TGCTGCCCAAGCACAOCCTGCTOTACGAGACATTCAATOTGTATAACGAGCTGACCA AAGTGCOGTTTATCGCCOAGTCTATGCGGGACTACCAGTTCCTGGACTCCAAGCAGA AAAAGGACATCGTGCGGCTGTACnCAAGGACAAGCGGAAAGTGACCOATAAGGAC ATCATCGAGTACCTGCACGCCATCTACGGCTACGATGGCATCGAGCTGAAGGGCAT CGAGAAGCAGTICAACTCCAGCCTGAGCACATACCACGACCTGCTGAACATTATCA ACGACAAAGAATITCfOGACGACTCCAGCAACGAGGCCATCATCGAAGAGATCATC CACACCCTGACCATCrITGAGGACCGCGAGATGATCAAGCAGCGGCTGAGCAAGTI CGAGAACATCITCOACAAGAGCGTGCTGAAAAAGCTOAOCAGACUGCACTACACCG GCTGGGOCAAGCraAOCOCCAAGCTGATCAACGGCATCCGOGACGAGAAGTCCGGC AACACAATCCTOGACTACCTGATCGACGACGGCATCAGCAACCGGAACTTCATGCA GCTGATCCACGACGACGCCCTGAGCTTCAAGAAGAAGATCCAGAAGGCCCAGATCA TCGGGGACOAGGACAAGOGCAACATCAAAGAAGTCGTGAAGTCCCTGCCCGGCAGC
CCCGCCATCAAGAAGGGAATCCTGCAGAGCATCAJ ~ GATCGTGGACGAGCTCGTGAA
AGTGATGGGCGGCAGAAAGCCCGAGAGCATCGTGGTGGAAATGGCTAGAGAGAAC
60
CAGTACACCAATCAGGGCAAGAGCAACAOCCAGCAGAGACTGAAGAGACTGGAAA AGTCCCTGAAAGAGCTGGGCAGCAAGATrcrOAAAGAGAATATCCCTGCCAAOCTO TCCAAGATCGACAACAACGCCCTGCAGAACGACCGGCTGTACCTGTACTACCTOCA GAATGGCAAGGACATOTATACAGGCGACGACCTGGATATCGACCGCCTGAGCAACT ACGACATCGACCATATTATCCCCCAGGCCTTCCTGAAAGACACAGAC AAAGTOCTGGTGTCCTCCGCCAGCAACCGCGGCAAGTCCOATGATGTGCCCAGCCT GGAAGTCGTGAAAAAGAGAAAGACcn'CTGGTATCAGCTGCTGAAAAGCAAGCTGA TTAGCCAGAGGAAGTTCGACAACCTGACCAAGGCCGAGAGAGGCGGCCTGAGCCCT GAAGATAAGGCCGGCTTCATCCAOAGACAGCTGGTOGAAACCCOGCAGATCACCAA OCACGTGGCCAGACTGCTGOATGAGAAGTTTAACAACAAGAAOGACGAGAACAACC OGGCCOTOCOGACCOTGAAGATCATCACCCTGAAGTCCACCCTGGTGTCCCAGTTCC GGAAGGACTTCGAGCTGTATAAAGTGCGCGAOATCAATGACTTTCACCACGCCCAC GACGCCTACCTGAATGCCGTOOTGGCTTCCOCCCTGCTGAAOAAOTACCCTAAGCTG GAACCCGAGTICGTOTACGGCGACTACCCCAAGTACAACTCCTTCAGAGAGCOOAA GTCCOCCACCGAGAAGGTGTACTTCTACTCCAACATCATGAATATCTTTAAGAAGTC CATCTCCCTGGCCGATGGCAGAGTGATCGAGCGGCCCCTGATCGAAGTGAACGAAG AGACAGGCGAGAGCGTGTGGAACAAAGAAAGCGACCTGGCCACCGTGCGGCGGGT GCTGAGTTATCCTCAAGTGAATGTCOTGAAGAAGGTOGAAGAACAGAACCACGGCC TGGATCGGGGCAAGCCCAAGGGCCTGTTCAACGCCAACCTGTCCAGCAAGCCTAAG CCCAACTCCAACGAGAATCTCGTGGGGGCCAAAGAGTACCTGGACCCTAAGAAGTA CGGCGGATACGCCGGCATCTCCAATAGCTTCACCGTGCTCGTGAAGGGCACAATCO AGAAGGGCGCTAAGAAAAAOATCACAAACGTGCTGGAATTTCAGGGGATCTCTATC CTGGACCGGATCAACTACCGGAAGGATAAGCTGAACTTTCTGCTGGAAAAAGGCTA CAAGGACATTGAGCTGATTATCGAGCTGCCTAAGTACTCCCTGTTCOAACTGAGCGA CGGCTCCAGACGGATGCTGGCCTCCATCCTGTCCACCAACAACAAGCGGGGCGAGA TCCACAAGGGAAACCAGATCTTCCTGAGCCAGAAATTTGTGAAACTGCTGTACCACO CCAAOCGOATCTCCAACACCATCAATGAGAACCACCGGAAATACGTGGAAAACCAC AAGAAAGAUITrGAGOAACTGTrCTACTACATCCTGGAGTTCAACOAGAACTATOTO GGAOCCAAGAAGAACGGCAAACTOCTGAACTCCGCCTTCCAGAGCTGGCAGAACCA CAGCATCGACGAGCTGTOCAGCTCcrrCATCGOCCCl'ACCGGCAGCGAGCGGAAGG GACTGTlTGAGCTGACCTCCAGAGGCTCTGCCGCCOACTTTGAGTTCCTGGGAGTGA
AGATCCCCCGGTACAGAGACTACACCCCCTCTAGTCTGCTGAAGGACGCCACCCTGA
TCCA CCAGAGCGTGACCGGCCTGTACGAAACCCOGATCGACCTGGCTAAGCTGGGC
GAGGGAAAGCGTCCTGCTGCTACTAAGAAAGCTGGTCAAGCTAAGAAAAAGAAATA
TO
Example 5: Optimization of guide RNA for Cas9 of Streptococcus pyogenes (referred to as SpCas9)
5 Applicants mutated the RNAtracr and direct repeat sequences or mutated the chimeric guide RNA to activate the RNAs in the cells
The optimization is based on the observation that there are sections of thymine (Ts) in the RNAtracr and guide RNA, which could lead to termination of early transcription by the Poi 3 activator. Therefore the Applicants generated the following optimized sequences, ARNtracr optimized and the corresponding optimized direct repeat is
10 present in pairs,
Optimized ARNtracr 1 (underlined mutation):
GGAACCATTCAtAACAGCATAGCAAGTTAtAATAAGGCTAGTCCGTTATCAACTTGAA AAAGTGGCACCGAGTCGGTGCrnnT Optimized direct repeat 1 (underlined mutation): GTTaTAGAGCTATGCTGTTaTGAATGGTCCCAAAAC
5 Optimized ARNtracr 2 (underlined mutation) GGAACCA TTCAAtACAGCAT AGCAAGTTAAtA T AAGGCT AGTCCGTT ATCAACTTGAA AAAGTGGCACCGAGTCGGTGCrnnT Optimized direct repeat 2 (underlined mutation) · GTaTT AGAGCT ATGCTGTaTTGAA TGGTCCCAAAAC
10 Applicants also optimized the chimeric guide RNA for optimal activity in eukaryotic cells Original guide RNA
NNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGC TAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC I IIIIII
RNA sequence optimized chimeric guide 1
~~~~~~ NNNNNNNNNGTATTAGAGCTAGAAATAGCAAGTTAAIATAAGGC
TAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC II II III
fifteen RNA sequence optimized chimeric guide 2:
NNNNNNNNNNNNNNNNNNNNGTTTTAGAGCTATGCTGTTTTGGAAACAAAACAGCA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCG
GTGC II IIIII
Stream of optimized chimeric guide RNA 3:
NNNNNNNN ~ NNGTATTAGAGCTATGCTGTATTGGAAACAAIACAGC
ATAGCAAGTTAAIATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTC GGTGC II 1 111 [
The applicants showed that the optimized chimeric guide RNA works best as indicated in Figure 9. The
twenty Experiment was performed by co-transfection of 293FT cells with Cas9 and a U6 guide RNA DNA cassette to express one of the four forms of RNA shown above. The objective of the guide RNA is the same target site in the human Emx1 site: "GTCACCTCCAATGACTAGGG"
Example 6: Optimization of Cas9 CRISPR1 LMO-9 from Streptococcus thermophilus (referred to as StICas9).
Applicants designed the guide chimeric RNAs as shown in Figure 12
25 StlCas9 guide RNAs can undergo the same type of optimization as for SpCas9 guide RNAs, due to breakage of the poly thymine (Ts) sections.
Example 7: Improvement of the Cas9 system for in vivo application
The requestors performed a Metagenomic search for a Cas9 coo with a small molecular weight. Most of the Cas9 counterparts were quite large. For example, the SpCas9 is about 1368aa long, which is too long to be easily packaged into viral vectors for delivery. Some of the sequences
they may have been lifted and therefore the exact frequency for each length may not necessarily be accurate. However, it provides a glimpse of the distribution of Cas9 proteins and suggests that there are shorter Cas9 counterparts.
5 By computational analysis, the applicants found that in the bacterial strain Campylobacter, there are two Cas9 proteins with less than 1,000 amino acids. The sequence for a Cas9 from Campylobacter jejuni is presented below. At this length, CjCas9 can be easily packaged in AAV, lentivirus, Adenovirus, and other viral vectors for robust delivery in primary cells and in vivo in animal models
Cas9 by Campylobacter jejuni (CjCas9)
MARJLAFDIGISSIGWA FSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRLARR
KARLNHLKH LIANEFKLNY EDYQSFDESLAKA YKGSLlSPYELRFRALN ELLSKQDF AR
VILHIAKRRGYDDfKNSDDKEKGAlLKAIKQNEEKLANYQSVGEYLYKEYFQKfKENSK
EFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFS
HLVGNCSFFTDEKRAPKNSPLAFMFVAL TRJINLLNNLKNTEGIL YTKDDLNALLNEVLK
NGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFfKALGEHNLSQDDLNEIAKDIT
LIKDEIKLKKALAKYDLNQNQJDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNE
LNLKVAlNEDKKDFLPAFNETYYKDEVTNPVVLRAlKEYRKVLNALLKKYGKVHKJNIE
LAREVGKNHSQRAKlEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCA
YSGEKlKlSDLQDEKMLElDHIYPYSRSFDDS YMNK VL VFTKQNQEKLNQTPFEAFGNDS
AKWQKlEVLAKNLPTKKQKRlLDKNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYL
DFLPLSDDENTKLNDTQKGSK VHVEAKSGML TSALRHTWGFSAKDRNNHLHHAlDA VI
lA AND ANNSIVKAFSDFKKEQESNSAEL and AKKJSELDYKNKRlKFFEPFSGFRQKVLDKIDEIF
VSKPERKKPSGALHEETFRKEEEFYQSYGGKEGVLKALELGKIRKVNGKIVKNGDMFR
VDIFKHKKTNKfY TO VPIYTMDF ALKVLPNKA V ARSKKGEJKDWILMDENYEFCFSL YK
DSLILIQTKDMQEPEFVYYNAFTSSTVSLlVSKHDNKFETLSKNQKlLFKNANEKEVlAKS
IGIQNLKVFEKYIVSALGEVTKAEFRQREDFKK.
The putative ARNtracr element for this CjCas9 is:
TATAATCTCATAAGAAATTTAAAAAGGGACTAAAATAAAGAGTTTGCGGGACTCTG CGGGGTTACAATCCCCTAAAACCGCTTTTAAAATT
The direct repeat sequence is ·
~ IIIIACCATAAAGAAATTTAAAAAGGGACTAAAAC
fifteen The collapsed structure of the ARNtracr and direct repetition is provided in Figure 6
An example of a chimeric guide RNA for CjCas9 is:
NNNNNNNNNNNNNNNNNNNNGUUUUAGUCCCGAAAGGGACUAAAAUAAAGAGUU UGCGGGACUCUGCGGGGUUACAAUCCCCUAAAACCGCUUUU
Applicants also optimized Cas9 guide RNA using in vitro methods. Figure 18 shows optimization data for St1 Cas9 in vi / ro chimeric guide RNA.
twenty Although preferred embodiments of the present invention have been presented and described herein, it will be apparent to those skilled in the art that such embodiments are provided as an example only. Many variations, changes and substitutions will now take place for those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be employed in the practice of the invention. It is desired that the following
25 claims define the scope of the invention and that the methods and structures within the scope of the claims and their equivalents are thereby covered
Example 8: Optimization of ARNsg Sa
Applicants designed five variants of ARNsg for SaCas9 for optimal truncated architecture with the
Greater efficiency of excision. In addition, the Iracr duplex system was tested: natural direct repetition with the ARNsg
Guides with lengths indicated with SaCas9 were transinfected together and tested in HEK cells
293FT for the activity. A total of 100 ng of RNA6g U6-PCR amplicon was jointly transinfected (or 50
ng of direct repeat and 50 ng of ARNtracr) and 400 ng of SaCas9 plasmid in 200,000 mouse hepatocytes
Hepa1.-6 and DNA was collected in 72 hours post-transfection for SURVEYOR analysis. The results are
presented in Fig. 23
REFERENCES
1. Urnov, FD, Rebar, EJ., Holmes, MC, Zhang, HS & Gregory, PD Genome editing with engineered zinc finger nueleases. Na /. Re\'. Gene/. 11,636-646 (2010).
two. Bogdanove, AJ. & Voytas, DF TAL effectors: customizable proteins for ONA
targeting Science 333, 1843-1846 (2011).
3. Stoddard, BL Homing endonuclease S! Ructure and function. Q. Rev. Binphys. 38, 4995 (2005).
Four. Bae, T. & Schneewind, O. Allelic replacement in Slaphylococcus aureus with
inducible counter-selection. P / asmid SS, 58-63 (2006).
5. Sung, CK, Li, H., Claverys, JP & Monison, DA An rpsL cassene, janus, for gene
replacement through ncgative selection in Streptococcus pneumoniae. Appl. Enviro1l. Microbe/.
67,5190-5196 (2001).
6. Sharan SK, Thomason, LC., Kuznetsov, SO & Court, DL Recombineering: a homologous rccombination · based method of genetic cngineering. Nat. Protoc. 4, 206223
(2009).
7. Jinek M. et al. A programmable dual-RNA-guidcd DNA endonuclease in adaptive
bacterial immunity. Science 337.8 16-821 (2012).
<dl><dt>8. </dt><dd>Deveau, H., Garneau, JE & Moineau, S. CRlSPR-Cas system and its role in phagebacteria interacrions. Annu Rev. Microbe /. 64, 475-493 (2010) _</dd></dl>
<dl><dt>9. </dt><dd>Horvath, P. & Sarrangou, R. CRlSPR-Cas, the immune system of bacteria and archaea. Science 327, 167-170 (2010).</dd></dl>
<dl><dt>10. </dt><dd>Tems, MP & Tems, RM CR.ISPR · bascd adaptive irnrnune systems. Curro Opino Microbio. 14, 321327 (2011).</dd></dl>
<dl><dt>11. </dt><dd>van der 0051, J., Jore, MM, Westra, ER, Lundgren, M. & Brouns, S.1. CRISPR · based adaplive and heritable irnmunity in prok2lryotes. Trends Biochem Sci. 34,401407 (2009).</dd></dl>
<dl><dt>12. </dt><dd>Brouns, SJ et al. SmaJl CRlSPR RNAs!, 'Uide antiviroll defense in prokaryotes. Science3Z1, 960964 (2008).</dd></dl>
<dl><dt>13. </dt><dd>Cane, J., Wang, R., Li, H., Tems, IR.M. & Tems, MP Cas6 is an cndoribonuclease</dd></dl>
thal generoltcs guidc RNAs for invader dcfcnse in prokaryotes. Genes Del '. 22, 34893496 (2008).
<dl><dt>14. </dt><dd>Deltcheva, E. et al. CRlSPR RNA maturation by lrans · encoded small RNA and host factor RNasc JI !. Nature 471, 602-607 (20 11).</dd></dl>
<dl><dt>15. </dt><dd>Haloum · Aslan, A., Maniv, l. & Marraffini, LA Mature c1ustered, regularly inlerspaeed, short palindromic repeals RNA (crRNA) length is rneasurcd by a ruler mechanism anchored to the precursor processing sitc. Proc. NaIf. Acad. Sci. USA. 10H, 21218-2 1222 (20 11).</dd></dl>
<dl><dt>16. </dt><dd>Haurwitz, RE, Jinek, M., Wiedenheft, B., Zhou, K. & Doudna, JA Sequence-and structure-specific RNA processing by a CRlS PR endonuclease. Science 329, 1355-1358 (2010).</dd></dl>
<dl><dt>17. </dt><dd>Dcvcau, H. et al. Phagc response lO CRlSPR-encodcd rcsistance in Slreptococcus Ihermophilus. J. Bacteriol. 190, 1390-1400 (2008).</dd></dl>
<dl><dt>18. </dt><dd>Gasiunas, G., Barrangou, R., Horvath, P. & Siksnys, V. Cas9 · crRNA ribonucleoprolein complex mediales specific DNA cleavage for adaplive irnmunity in bacteria. </dd></dl>
Proc. Natl Acad. Sci. USA (2012).
<dl><dt>19. </dt><dd>Makarova, Ks, Aravind, L., Wolf, Y.1. & Koonin, EV Unification of Cas proteio furni lics and a simple scenario for the origin and cvolution of CRJSPR-Cas systems. Biol. Direcl. 6.38 (20 11).</dd></dl>
<dl><dt>20. </dt><dd>Barrangou, R. RNA-medialed programmable DNA cleavage. Nat. Biolechnol. 30, 836838 (20 12).</dd></dl>
<dl><dt>21. </dt><dd>Brouns SJ. Molecular biology A Swiss anny kni faith of immunity. Science 337. 808809 (2012).</dd></dl>
<dl><dt>22. </dt><dd>Carroll, D. A CRISPR Approach 10 Gene Targeling. Mol. Ther. 20, 1658-1660 (2012).</dd></dl>
<dl><dt>23. </dt><dd>Bikard, D., Hatoum-Aslan, A., Mucida, D. & Marraffini, LA CRJSPR interference can prevent natural transfonnation and virulence acquisition during in vivo bacterial infection. Cell Host Microbe 12, 177-186 (2012).</dd></dl>
<dl><dt>24. </dt><dd>Sapranauskas, R. el al. The StreplOCOCcus Ihermophilus CRJSPR-Cas syslem provides immunity in Escherichia eolio Nucleie Acids Res. (20 11).</dd></dl>
<dl><dt>25. </dt><dd>Semenova, E. et al. Interference by clustercd rcgularly interspaccd sh ort palindromic repcat (CRISPR) RNA is govemed by a sccd scquence. Proe. Nall Acod. Sci. U.SA. (20 11).</dd></dl>
<dl><dt>26. </dt><dd>Wiedenheft, B. et al. RNA-guided complex from a bacterial immune system enhances targct rccognition through secd sequcnce inleractions. Proc. Nal /. Acad. Sci. USA (2011).</dd></dl>
<dl><dt>27. </dt><dd>Zahncr, D. & Hakenbcck, R. Th (: Slreplococcus pneumoniae bCla-galactosidase is a surface protcin.J. Bacteriol. 182.5919-592 1 (2000). </dd></dl>
<dl><dt>28. </dt><dd>Marraffini, LA, Dedcnl, Ae &. Schneewind O. Sonases and the art of anchoring</dd></dl>
proteins 10 Ihe envelopes of gram-posilive bacteria. Microbe/. Mol. Biol. Re \ '. 70, 192-221 (2006).
<dl><dt>29. </dt><dd>Motamedi, MR, Szigety, SK & Rosenberg, SM Double-strand-break repair recombinalion in Escheriehia coli: physieal evidence for a DNA replicaicalion mechanism in vivo. Genes Dev. 13,2889-2903 (1999).</dd></dl>
<dl><dt>30. </dt><dd>Hosaka, T. et al. The novel mutalion K87E in ribosomal prolein S 12 enhanccs proteio synlhcsis activiry during Ihe late growth phasc in Eseherichia eolio Mo /. Genet Geflomies 271, 317–324 (2004).</dd></dl>
<dl><dt>31. </dt><dd>Coslantino, N. & Court, DL Enhaneed le veis of lambda Red-mediated recombinants in mismalch repair mulanlS. Proc. Nad. Acad. Set ". USA IOO, 15748-15753 (2003).</dd></dl>
<dl><dt>32. </dt><dd>Edgar, R. & Qimron, U. The Escherichio with CRJSPR syslem prulccls from lambda Iysogcnizalion, Iysogens, and prophage induclion. J. Bacterium /. 192, 6291-6294 (2010).</dd></dl>
<dl><dt>33. </dt><dd>Marruffini, LA & Sontheimer, EJ Self versus non-self discrimination during C RI SPR RNA-direclcd immunity. Nature 463, 568-571 (2010).</dd></dl>
<dl><dt>34. </dt><dd>Fischcr, S. al. An archaeal immune syslem can delecl multiple Protospacer Adjaccnt Motifs (PAMs) to target invader DNA. J. Biol. Chem. 281, 3335 1-33363 (20 12).</dd></dl>
<dl><dt>35. </dt><dd>Gudbergsdonir, S. el al. Dynamic properties of Ihe Sulf% bus CRISPR-Cas and CRlSPRlCmr syslems whcn challenged wilh vector-borne vir.il and plasmid genes and prolospacers. Mol. Microbe/. 79.35-49 (20 11).</dd></dl>
36. Wang, HH et al. Genome-scale promoter enginecring by coselection MAOE. Nar
Melhads 9, 59 1-593 (2012).
<dl><dt>37. </dt><dd>Cong, L. et al. Multiplex Genome Engineering Using CRJSPR-Cas Systems. Science In pr ... (2013).</dd></dl>
<dl><dt>38. </dt><dd>Mali, P. et al. RNA-Guided Human Genome Engineering via Cas9. Science ln press (2013).</dd></dl>
<dl><dt>39. </dt><dd>Hoskins, J. et al. Genome of the baeterium S / rep / ococcus pneumoniae strain R6. J. BaClerial. 183, 5709-57 17 (2001).</dd></dl>
<dl><dt>40. </dt><dd>Havarstein, LS, Coomaraswarny, G. & Morrison, OA An unmodified heptadecapeplide pheromone induces compctence for genelic trdnsfonnalion in S / replococCUs pneumoniae. Proc. Nal /. Acad. Sci. USA 92, 1 J140-1 J144 (1995).</dd></dl>
<dl><dt>41. </dt><dd>Horinouchi, S. & Weisblum, B. Nucleolide sequence and funclional map ofpCl94, a plasmid that specifics inducible ehlordmphenicol resislance. J. Bacterium /. 150,815-825 (1982).</dd></dl>
<dl><dt>42. </dt><dd>Horton, RM In Vitro Recombination and Mutagenesis of DNA: SOEing Together </dd></dl>
Tailor-Madc Genes. Melhods Mol. Biol. 15,251-261 (1993).
43 Podbielski, A., Spellerberg, 8., Woisehnik, M., Pohl, 8. & Lunicken, R. Novel series
of plasmid vectors for gene inactivation and expression analysis in group A streptococci (GAS). Gene 177,137-147 (1996).
<dl><dt>44. </dt><dd>Husmann, LK, Seon, JR, Lindahl, G. & Stcnberg, L. Expression oflhe Arp protein, a member of Ihe M protein family, is nol suffieicnt 10 inhibil phagocytosis of SlreplococCUs pyogenes. Infeclion and immunity 63, 345-348 (1995).</dd></dl>
<dl><dt>45. </dt><dd>Gibson, DG al. Enzyrnalic assembly of ONA molecules up 10 several hundred kilobases. Nm Melhods 6, 343-345 (2009).</dd></dl>
<dl><dt>46. </dt><dd>Tangri S, ct al. ("Ralionally cngineered Iherapeutic proteins with reduced immunogenicity" J lmmunol. 2005 Mar 15; 174 (6): 3187-96.</dd></dl>
Although preferred embodiments of the present invention have been presented and described herein, it will be apparent to those skilled in the art that such embodiments are provided as an example only. Many variations, changes and substitutions will now take place for those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be employed in the practice of the invention.
SEQUENCE LIST
<110> THE BROAD INSTITUTE, INe. MASSACHUSETIS TECHNOLOGY INSTITUTE
<120> ENGINEERING OF SYSTEMS, METHODS AND COMPOSITIONS OPTIMIZED GUIDELINES FOR HANDLING SEQUENCES
<130> PC927381 EPA
<140> U.S. Patent PCTfUS2013f074819
< 141> 2013-12-12
<150> 61f836 .127
< 151> 2013-06-17
<150> 61 f835.931
< 151> 2013-06-17
<150> 61 f828 .130
< 151> 2013-05-28
<150> 61 f819 .803
< 151> 2013-05-06
<150> 61 f81 4,263
< 151> 2013-04-20
<150> 61f806.375
< 151> 2013-03-28
<150> 61/802.174
< 151> 2013-03-15
<150> 61 f791 .409
< 151> 2013-03-15
<150> 61f769.046
< 151> 2013-02-25
<150> 61 f758 .468
< 151> 2013-01-30
<150> 61f748 .427
< 151> 2013-01-02
<150> 61f736.527
< 151> 2012-12-12
<160> 264
<170> Patentln version 3.5
<210> 1
<211> 15
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Synthetic Oligonucleotide Artificial Sequence ~ </dd></dl>
<400> 1 aggacgaagt cctaa 15
<210> 2 <211>7
<212> PRT
<213> Ape Virus 40
<400> 2
Pro Lys Lys Lys Arq Lys Val 1 5
<210> 3
<211> 16
<212> PRT
<213> Unknown
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Unknown · Nucleoplasmin bipartite NLS sequence ~ </dd></dl>
<400> 3
Lyt Arq Pro Wing Thr Wing Lys Lys Wing Gly Gln Wing Lys Lys Lys Lys 1 5 10 15
<210> 4 <211>9
<212> PRT
<213> Unknown
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Unknown · NLS sequence C-myc · </dd></dl>
<400> 4 Pro Ala Al.a Lys Arg Val Lys Leu Asp 1 5
<210> 5
<211> 11
<212> PRT
<213> Unknown
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Unknown · NLS C-myc sequence" </dd></dl>
<400> 5
Arq Gln Arg Arg Asn Glu Leu Lys Arg Ser Pro 1 5 10
<210> 6
<211> 38
<212> PRT
<213> Homo sapiens
<400> 6
Aan Gln Ser Ser Asn Phe Gly Pro Mat Lys Gly Gly Asn Phe Gly Gly 1 S 10 15
Arq Ser Ser Gly pro Tyr Gly G1y Gly Gly Gln Tyr Phe Ala Lys Pro 20 25 30
Arq Asn Gln Gly Gly Tyr 35
<210> 7
<211> 42 5 <212> PRT
<213> Unknown
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = -Unknown Description: lBS domain of importina-alpha sequence</dd></dl>
10 <400> 7
Arq Mat Arq Ila G1x Phe Lys Asn Lys Gly Lys Aap Thr Als Glu Leu 1 5 la 15
Arq ~ 9 ~ Arq Val Glu Val Ser Val Glu Lau Arq Lys Ala Lys Lys 20 25 30
Asp Glu Gln Ile Leu Lys ArO Arq Aan Val 3S 40
<210> 8
<211> 8
<212> PRT 15 <213> Unknown
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = -Unknown Description: Myoma T protein sequence</dd></dl>
<400> 8
Val Ser Arg Lys Arq Pro Arq Pro 20 1 5
<210> 9
<211> 8
<212> PRT
<213> Unknown
25 <220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = -Unknown Description: Myoma T protein sequence</dd></dl>
<400> 9
Pro Pro Ly8 Lys Ala Arq Glu ABp fifteen
30 <210> 10
<211> 8
<212> PRT
<213> Sapiens oven
<400> 10
Pro Gln Pro Lys Lys Lys Pro LBu 1 5
<210>11 <211>12
<dl><dt>< </dt><dd>212> PRT </dd></dl>
<dl><dt>< </dt><dd>213> Mus musculus </dd></dl>
<400> 11
Ser Ala LBu Ile Lys Lys Lys Lys Lys Met Ala Pro 1 5 10
<210> 12
<211> 5
<212> PRT
<213> Inftuenza virus
<400> 12 Asp Arg Leu Arg Arg 1 5
<210> 13 <211>7
<212> PRT
<213> Inftuenza virus
<400> 13
Pro Lys Gln Lys Lys Arq Lys 1 5
<210> 14
<211> 10
<212> PRT
<213> Hepatitis delta virus
<400> 14
Arg Lys Leu Lys Lys Lys Ile Lys Lys Leu 1 5 10
<210> 15
<211> 10
<212> PRT
<213> Mus musculus
<400> 15
Argo Glu Lys Lys Lys Phe Leu Lys Arg Arg 1 5 10
<210> 16 <211>20
<dl><dt>< </dt><dd>212> PRT </dd></dl>
<dl><dt>< </dt><dd>213> Homo sapiens </dd></dl>
<400> 16
Lys Arg Lya G1y Asp Glu Val Asp Gly val Asp Glu Val Ala Lys Lys 1 5 10 15
Lys Ser Ly8 Lys 20
<210> 17
<211> 17
<212> PRT
<213> Homo sapiens
5 <400> 17
ArO Lye Cye Leu Gln Wing Gly Met Asn Leu Glu Ale Arq Lye Thr Lye 1 S 10 lS
Lys
<210> 18 <211>27
<212> DNA 10 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Sinlic Oligonucleolide ~ </dd></dl>
<220> 15 <221> modified base
<222> (1) .. (20) <223> a, c, t og
<220>
<221> modified base 20 <222> (21) .. (22)
<223> a, c, t, g, Unknown or olro
<400> 18 nnnnnnnnnn nnnnnnnnnn nnagaaw 27
<210> 19 25 <211>19
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<221> source 30 <223> Inola = ~ Description of Artificial Sequence · Synthetic Oligonucleolide "
<220>
<221> modified base
<222> (1) ..(12)
<223> a, c, log
35 <220>
<dl><dt>< </dt><dd>221> modified base </dd></dl>
<dl><dt>< </dt><dd>222> (13) ..(1 4) </dd></dl>
<dl><dt>< </dt><dd>223> a, c, t, g, Unknown or other </dd></dl>
<400> 19 40 nnnnnnnnnn nnnnagaaw 19
<210> 20 <211>27
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = ~ Description of Artificial Sequence Ol synthetic igonucleotide ~ </dt><dd /></dl>
<dl><dt><220> <221> modified base <222> (1) .. (20) <223> a, c, tog </dt><dd /></dl>
<dl><dt><220> <221> modified base <222> (21) .. (22) <223> a, e, t, g, Unknown or other </dt><dd /></dl>
<dl><dt><400> 20 nnnnnnnnnn nnnnnnnnnn nnagaaw </dt><dd> 27 </dd></dl>
<dl><dt><210> 21 <211> 18 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota = ~ Description of Artificial Sequence · Synthetic oligonucleotide " </dt><dd /></dl>
<dl><dt><220> <221> modified base <222> (1) .. (11) <223> a, c, t og </dt><dd /></dl>
<dl><dt><220> <221> modified base <222> (12) .. (13) <223> a, e, t, g, Unknown or other </dt><dd /></dl>
<dl><dt><400> 21 nnnnnnnnnn nnnagaaw </dt><dd> 18 </dd></dl>
<dl><dt><210> 22 <211> 137 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota = "Description of Artificial Sequence · Synthetic polynucleotide" </dt><dd /></dl>
<dl><dt><220> <221> modified base <222> (1) .. (20) <223> a, e, t, g, Unknown or other </dt><dd /></dl>
<dl><dt><400> 22 nnnnnnnnnn </dt><dd>nnnnnnnnnn qtttttqtac tctcaaqatt taqaaataaa tcttqcaqaa 60 </dd></dl>
<dl><dt>qctacaaaqa taaqqcttca </dt><dd>tqccgaaatc aacaccctqt cattttatqq caqqqtqttt 120 </dd></dl>
<dl><dt>t cqttattta atttttt </dt><dd> 131 </dd></dl>
<210> 23
<21 1> 123
<212> DNA
<213> Artificial Sequence
5 <220>
<221> source
<223> Inota ::: Artificial Sequence MD: Synthetic polynucleotide "
<220>
<221> modified base 10 <222> (1) .. (20)
<223> a, c, t, g, Unknown or other
<400> 23
nnnnnnnnnn nnnnnnnnnn qtttttqtac tctcaqaaat qcaqaaqcta caaa; ataa; 60
octtcatgcc qaaatcaaca ccctqtcatt t tatggcagg gtgttttcqt tatttaattt 120
ttt 123
<210> 24 15 <211> 110
<212> DNA
<213> Artificial Sequence
<220>
<221> source 20 <223> Inota ::: Artificial Sequence MD: Synthetic polynucleotide "
<220>
<221> modified base
< 222> (1 ) .. (20)
<223> a, c, t, g, Unknown or other
25 <400> 24 nnnnnnnnnn nnnnnnnnnn qtttttgtac tctcaqaaat qcaqaaqcta caaagat aaq 60
octtcatocc oaaatcaaca ccctgtcatt ttatgqcaqg gtgttttttt 110
<210> 25
<211> 102
<212> DNA 30 <213> Artificial Sequence
<220>
<221> source
<223> Inota ::: Artificial Sequence MD: Synthetic polynucleotide "
<220> 35 <221> modified base
< 222> (1 )..(20)
<223> a, c, t, g, Unknown or other
<400> 25
nnnnnn ~ nn nnnn ~ nnnn qttttaqaqc taqaaataqc aaqttaaaat aaqqctaqtc 60
cgttatcaac ttqaaaaagt qqcaccqagt cqgtqctttt tt 102
40 <210>26
<211> 88 <212> DNA
<213> Artificial Sequence
<220>
<221> source 5 <223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide "
<220>
<221> modified base
< 222> (1 )..(20)
<223> a, c, t, g, Unknown or other
10 <400> 26
nnnnnnnnnn nnnnnnnnnn qttttagaqc tagaaataqc aaqttaaaat aaqqctaqtc 60
cqttatcaac ttqaaaaagt qttttttt 88
<210> 27
<211> 76
<212> DNA 15 <213> Artificial Sequence
<220>
<221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide "
<220> 20 <221> modified base
< 222> (1 )..(20)
<223> a, c, t, g, Unknown or other
<400> 27
nnnnnnnnnn nnnnnnnnnn gttttagaqc tagaaatagc aagttaaaat aaqgctagtc 60
cgttatcatt tttttt 76
25 <210>28
<211> 12
<212> RNA
<213> Artificial Sequence
<220> 30 <221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide "
<400> 28 guuuuagagc ua 12
<210> 29 35 <211> 33
<212> DNA
<213> Homo sapiens
<400> 29 ggacatcgat gtcacctcca atgactaggg Igg 33
40 <210>30
<211> 33
<212> DNA
<213> Sapiens oven <400> 30 cattggaggt gacatcgatg tcctccccat tgg 33
<210> 31 <211>33
<212> DNA
<213> Sapiens oven
<400> 31 ggaagggcct gagtccgagc agaagaagaa ggg 33
<210> 32 <211>33
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 32 ggtggcgaga ggggccgaga ttgggtgttc agg 33
<210> 33
<211> 33
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 33 atgcaggagg gtggcgagag gggccgagat tgg 33
<210> 34
<211> 21
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Artificial Sequence" Synthetic primer " </dd></dl>
<400> 34 aaaaccaccc ttctctctgg c 21
<210>35
<211> 21
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Artificial Sequence: Synthetic primer" </dd></dl>
<400> 35 ggagattgga gacacggaga 9 21
<210> 36
<211> 20
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Artificial Sequence: Synthetic primer" </dd></dl>
<dl><dt><400> 36 ctggaaagcc aatgcctgac </dt><dd> 20 </dd></dl>
<dl><dt>5 </dt><dd><210> 37 <211> 20 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = ~ Description of Artificial Sequence "Synthetic primer ~ </dt><dd /></dl>
<dl><dt>10 </dt><dd><400> 37 ggcagcaaac tccttgtcct twenty </dd></dl>
<dl><dt>15 </dt><dd><210> 38 <211> 12 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = ~ Artificial Sequence Description: Ol synthetic igonucieotide " </dt><dd /></dl>
<dl><dt>20 </dt><dd><400> 38 gttttagagc ta 12 </dd></dl>
<dl><dt><210> 39 <211> 335 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>25 </dt><dd><220> <221> source <223> Inota = ~ Artificial Sequence Description "Synthetic polynucleotide" </dd></dl>
<dl><dt><400> 39 qaqqqcctat </dt><dd>ttcccatgat tccttcatat ttqcatatilC qatilcaaggc tgttagagag 60 </dd></dl>
<dl><dt>ataattqgaa </dt><dd>ttaatttgac tgtaaacaca aaqatattaq tacaaaatac gtgacgtaqa 120 </dd></dl>
<dl><dt>aagtaataat </dt><dd>ttcttqqqt.a gtttqcaqtt ttaaaattat qttttaaaat qqactatc at 180 </dd></dl>
<dl><dt>atqcttaccq </dt><dd>taacttqaaa gtatttcqat ttcttqqctt tatatatctt qtqqaaaqqa 2.0 </dd></dl>
<dl><dt>cqaaacaccq </dt><dd>qaaccattca aaacaqcata qcaagttaaa ataaqqctag tcc qttatca 300 </dd></dl>
<dl><dt>acttqaaaaa </dt><dd>qtqqcaccqa qtcqqtqctt ttttt 335 </dd></dl>
<dl><dt>30 </dt><dd><210> 40 <211> 423 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>35 </dt><dd><220> <221> source <223> Inota = "Description of Artificial Sequence: Synthetic polynucleotide" </dd></dl>
<400> 40
gaqgqcctat ttcccatqat tccttcatat ttqcatatac qatacaaqqc tqttaqaqaq 60 ataattgqaa ttaatttgac tgtaaacaca aagatattaq tacaaaatac qtqacqtaga 120 aaqtaataat ttcttgqgta gtttgcaqtt ttaaaattat gttttaaólat ggactatcat 180 atgcttaccg taClcttqClaa gtatttcqat ttcttqqctt tatCltCltctt qt, qqaaaqgCl 240 cgaaacaccg qtClgtattaa qtattqtttt atqqctqata aatttctttq aatttctcct 300 tgattatttq ttataaaaqt tataaaataa tcttgttqqa accattcaaa acaqcataqc 360
aaqttaaaat aa9Qctaqtc cqttat caac ttgaaaaaqt qgcaccgagt cgqtqctttt 420 ttt 423
5 <210> 41 <211>339
<212> DNA
<213> Artificial Sequence
<220> 10 <221> source
<223> Inota = ~ Artificial Sequence Description l: Synthetic polynucleotide H
<400> 41
gagQ9cctat ttcccatqat tccttcatat ttqcatatac qatacaaqqc tqttaqagag 60 ataattqgaa ttaatttgac tqtaaacaca aaqatattaq tacaaaatac qtgacgtaga 120 aaqtaataat ttcttqqqta qtttqcaqtt ttaaaattat qttttaaaat qqactatcat 180 atgcttaccg taacttqaaa qtatttcgat ttcttqqctt tatatatctt qtqqaaa99a 240 cqaaacaccq qqttttaga.q ctatqctgtt ttga.a.tgqtc ccaaaacggg tcttcgagaa 300 gacqttttag agctatqc tg ttttgaatqg tcccaaaac 339
<210> 42 15 <211>309
<212> DNA
<213> Artificial Sequence
<220>
<221> source 20 <223> Inota = ~ Description of Artificial Sequence "Synthetic polynucleotide" <400> 42
qagqqcctat ttcccatgat tccttcatat ttgcatatac gatacaaqqc tqttaqagag 60 ataattqqaa ttaatttgac tqtaaacaca aagatattaq tacaaaatac qtgacqtaga 120 aaqta to, taat ttcttgqqta qtttgcagtt ttaaaattat gttt.taaaat ggaetatcat 180 atgcttaccg taacttgaaa qtatttcgat ttcttggctt tatatatctt qtgqaaag9a 240 cqaaacaccq qqtcttcqaq aaqacctqtt ttaqaqctac¡ aaatac¡caaq ttaaaataaq qctaqtccq 300 309
<210> 43
<211> 1648 5 <212> PRT
<213> Artificial Sequence
<220>
<221> source
<223> Inota = ~ Description of artificial sequence "synthetic polypeptide ~
<dl><dt><400> </dt><dd> 43 </dd></dl>
<dl><dt>Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lya Asp His ASp </dt><dd>Ila Asp </dd></dl>
<dl><dt>1 </dt><dd> 5 10 15 </dd></dl>
Tyr Lys Asp Asp Asp Asp Lys Met Wing Pro Lys Lys Lys Arq Lys Val 20 25 30
Gly Ile His Gly val Pro Wing Wing Asp LyS Lys Tyr Ser Ile Gly Leu 35 40 015
Asp Ile G ~ and Thr Asn Ser Va ~ G ~ and Trp A ~ a Val Ile Thr Asp Glu Tyr 50 55 60
Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg Bis ~ 70 75 80
Be Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Be Gly Glu 85 ~ ~
100 105 110
-~-~--~---~-----
Arg Arg Lys Aso Arq Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125
Met Wing Lys Val Asp Asp Ser Phe Phe Bis Arg Leu Glu Glu Ser Phe 130 135 140
Leu val Glu Glu Asp Lys Lys Bis Glu Arg Bis Pro Ile Phe Gly Asn 145 150 155 160
Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr Bis 165 170 175
Leu Arq Lys Lys Leu Val Asp Ser thr A9p Lys Wing Asp Leu Arg Leu 180 185 190
Ile Tyr Leu Wing Leu Wing Bis Met Ile Lys Phe Arq Gly His Phe Leu 195 200 205
Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val ASp Lys Leu Phe 210 21.5 220
Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240
<dl><dt>Aso Ala </dt><dd>Be Gly Val Asp Ala Lys Ala Ile Leu Being JUa Arg Leu Be </dd></dl>
<dl><dt>245 </dt><dd> 250 255 </dd></dl>
<dl><dt>81 </dt><dd /></dl>
Lys Ser Arq Arq Leu G1u Asn Leu I1e A1! L G1n Lau Pro Gly G1u Lys 260 265 270
Lys Asn Gly Leu Phe Gly Asn Lau Ile Al¡! L Lau Ser Leu G1y Leu Thr 275 280 285
"ro Asn! lhe Lys Ser Aan Phe Asp Leu Al¡ !!. Glu Asp Ala Lys Leu Gln 290 295 300
Leu Ser Lye Asp Thr Tyr Asp Asp Asp Le1.l Asp Asn! .Eu Leu Wing GIn 305 310 315 320
Ile Gly Asp Gln Tyr Wing Asp Leu Pha Leu Wing Wing Lys Asn Leu Ser 325 330 335
~ sp Wing I1a Leu Leu Ser Asp 11e Leu Arl; J Val Aan Thr G1u Il. orhr 340 345 350
Lys Ala Pro Leu Ser Ala Ser Met Ile Ly: s Arg Tyr Asp GIu His His 355 360 365
G1n Asp Leu Thr Leu Leu Lys Ala Leu Va: l ArQ G1n Gln Leu Pro GIu 370 375 380
Ly8 Tyr Lys Glu 119 Phe Phe Asp Gln Se: 1: 'Lys Asn Gly Tyr Ala Gly 385 390 395 400
Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe 11e Lys 405 410 415
Pro Ile Leu Glu Lys Met Asp Gly Thr Gl1l1 Gl.u Leu Leu Val. Lys Leu 420 425 430
P6sn Arg Glu Asp! .Eu Leu ~ g Lys Gl.n Ar9 Thr Phe Asp Asn Gly Ser 435 440 445
Ila Pro His Gl.n Ile His Leu Gl.y Gl.u Leu His Ala I1e Leu Arq Arch 450 455 460
Gln Glu Asp Phe Tyr Pro Pha Leu Lys AsIP Aso Arq Glu Lys Ile Glu 465 470 475 480
Lys Ile Leu Thr Phe Arq Ile Pro Tyr Ty: 1: 'Val Gly Pro: t.eu Ala Arg 485 490 495
Gly Asn Ser Arg Phe Wing Trp Met Thr Ar'iJ Lys Ser Gl.u Gl.u Thr Ile 82
500 you are
_ ~~ ~~~~ _ ~~~ _ll._
515 520 525
Ser Phe Ile Glu Arq Ket Thr Asn Phe ~ .8p Lys Asn Leu Pro Asn Glu 530 535 540
Lys Val Leu Pro Lys Bis Ser Leu Leu 'I'yr Glu Tyr phe thr Val Tyr 545 550 555 560
Asn Glu Leu Thr Lys val Lys 'l'yr val' I'hr Glu Gly Met Arq Lys Pro 565 570 575
Wing Phe Leu Ser Gly Glu Gln Lys Lys 1I.1a Ile Val Asp Leu Leu Pbe 580 585 590
Lys Thr Asn Arq Ly8 Val Thr Val Ly8 Giln Leu Ly8 Glu Asp Tyr Phe 595 600 605
Lys Lys Ile Glu Cys Phe Asp Ser Val G: lu IIa Ser Gly Val Glu Asp 610 615 620
Arq Phe Asn Ala Ser Leu Gly 'l'hr Tyr IIlia Asp Leu Leu Lys Ile Ile 625 630 635 .40
Lys Asp Lys Asp Phe Leu Asp Asn Gl.u Gilu Asn Glu Asp Il.e Leu Glu
645 E; SO 655
Asp Ile Val Leu 'l'hr Leu' l'hr Leu Phe Gau Asp Arg Glu Het Ile Glu 660 665 670
Glu Arg Leu Lys Thr Tyr Ala His Leu F'he Asp Asp Lys Val Met Lys 675 680 685
Gln Leu Lys Arg Arg Arq 'l'yr thr Gly' I'rp Gly Arg Leu Ser Arq Lys 650 695 700
Leu Ile Asn Gly Ile Arq Asp Lys Gln Sier Gly Lys' l'hr Ile Leu ASp 70S 710 715 720
Phe Leu Lys Ser Asp Gly Phe Ala Asn ~~ g Asn Phe Met Gln LBu Ile 725 1'30 735
His ASp Asp Ser Leu Thr Phe Lys Glu ~~ p Ile Gln Lys Wing Gln Val 740 145 750
Ser Gly Gln Gly Asp Ser Leu His Glu His Ile ~ a Asn Leu ~ a Gly 755 760 765
Be Pro Wing Ile Lys Lys Gly Ile Leu Gln Thr Val Lys val Val Asp 710 775 780
Glu Leu Val Lys Val Met Gly Arq His Lys Pro Glu Asn Ile Val Ile 785 790 795 800
Glu Met Wing Arch Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815
Arq Glu Arq Met Lys Arq Ile Glu Glu Gly Ila Lys Glu Leu Gly Ser 820 825 830
Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845
Lys Léu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arq Asp Het Tyr Val Asp 850 855 860
Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp Bis Ile 865 870 875 880
Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895
Thr Arq Ser Asp Lys Asn Arq Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910
Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arq Gln Leu Leu Asn Ala 915 920 925
Lys Leu Ile Thr Gln Arq Lys Phe Asp Asn Leu Thr Lys ~ a Glu Arq 930 935 940
Gly Gly Leu Ser Glu Leu Asp Lys ~ a Gly Phe Ile Lys Arq Gln Leu 945 950 955 960
Val Glu Thr Arq Gln Ile Thr Lys His Val ~ a Gln Ile Leu Asp Ser 965 970 975
Arq Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arq Glu Val 980 985 990
Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arq Lys Asp 995 1000 1005
Phe Gln Phe Tyr Lys Val Arq 1010 1015
Bis Asp Wing Tyr Leu ASn Al. 1025 1030
Lya Tyr Pro Lys Leu Glu Ser 1040 1045
Val Tyr Asp Val Arg Lys Met 1055 1060
Gly Lys Ala 'l'hr Ala Lys Tyr 1070 1,075
'he' he Lya 'l'hr Glu Ile Thr 1085 1,090
Arq Pro Leu lle Glu Thr Aan 1100 1105
AaP Lye Gly Arq Asp phé Ala 1115 1120
pro Gln Val Asn lle Val Lys 1130 1135
Phe Ser Lys Glu Ser Ile Leu 1145 1150
ne Al. Arq Lys Lys Asp Trp l160 1165
Asp Ser Pro Thr Val Ala Tyr 1175 1180
Glu Lys Gly Lys Ser Lys Lys 1190 1195
Gly ne Thr Ile Met Glu Ar9 1205 1210
Asp Phe Leu Glu Ala Lye Gly 1220 1225
n. ne Lys Leu Pro Lys Tyr 1235 1240
Glu 11e Asn Aso Tyr His His Ala 1020
Val Val Gly Thr Ala Leu 11e Lys 1035
Glu Phe Val '1'yr Gly Asp Tyr Lys 1050
lle Ala Lys Ser Glu Gln Glu rle 1065
Phe Phe Tyr Ser Asn Ile Met Asn wOLF
Leu Ala Asn Gly Glu He Arq I..ys 1095
Gly Glu Thr Gly Glu lle Val Trp
Thr Val Arq Lys Val Leu Sér Met 1125
Lys Thr Glu Val Gln rhr Gly Gly 1140
Pro Lys Arq Asn Ser Asp Lys I..eu 1155
Asp Pro Lys Lys Tyr Gly Gly Phe 1170
Ser Val Leu Val Val Ala Lys Val 1185
Leu Lys Ser Val Lys Glu Leu Leu 1200
Ser Ser Phe Glu Lys Asn Pro Ile 1215
Tyr Lys Glu Val Lys Lys Asp Leu 1230
Ser Leu Phe Glu Leu G1u Asn Gly 1245
Arq Lys Arq Met Leu Wing Being Wing Gly Glu Leu GI. Lys Gly Asn 12 50 1255 1260
Glu Leu AICl Lau Pro Ser Lya Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1215
Se r His Tyr Glu Lys Leu Lys Gly S-sr Pro Glu Asp Asn Glu Gln 1280 1285 1290
Lys GI. Leu Phe Val Glu GI. Bis Lys Bi s Tyr Leu Asp Glu Ile 1295 1300 1305
Ile Glu Gln Ile Ser Glu Phe Ser Lys Arq Val Ile Leu Ala Asp 1310 1315 1320
Wing As. Lau Asp Lys Val Lau Ser Wing Tyr Asn Lys Bis Arq Asp 1325 1330 1335
Lys Pro Ile Arq Glu Gln Ala Glu Asn Ile Ile Hi s Lau Phe Thr 1340 1345 1350
Leu Thr Asn Lau Gly Ala Pro Al.a Ala Phe LyS Tyr phe Asp Thr 1355 1360 1365
Thr Ile Asp Arq Lys Arq Tyr Thr Ser Thr Lys Glu Val Lau Asp 1370 1375 1380
Ala Thr Lau Ile His GIn Ser Ile Thr Gly Lau Tyr Glu Thr Arq 1385 1390 1395
Ile Asp Lau Ser Gln Leu Gly Gly Asp Wing Wing Wing Val Ser Lys 1400 140 5 1410
Gly Glu Glu Leu Pha Thr Gly Val Val Pro Ile Leu Val Glu Leu 1415 1420 1425
Asp Gly Asp Val Asn Gly Bis Lys Phe Ser Val Ser Gly Glu Gly 1430 1435 1440
Glu Gly Asp Wing Thr Tyr Gly Lys Leu Thr Leu Lys Phe Ile Cys 1445 1450 14 "
Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr 1460 1465 1470
Leu Thr Tyr Gly Val Gln Cys Phe Ser Arg Tyr Pro Asp His Met
Lys Gln 1490
Gln Glu 1505
Arg Al. 1520
Glu Leu 1535
His Ly. 1550
Al. Asp 1565
Bis Asn 1580
Gln Asn 1595
Bis Tyr 1610
Ly. Arch 1625
Ile Thr 1640
<210> 44
<211> 1625
<212> PRT
Bis Asp Phe Phe
Arg 'rhr Ile Phe
Glu Val Lys Phe
Lys Gly 11e Asp
Leu Glu Tyr Asn
Lys Gln Lys Asn
11e Glu Asp Gly
Thr Pro I 1e G1y
Leu Ser Thr Gln
Asp His Met Val
Leu Gly Met Asp
Ly. 1495
Phe 1510
Glu 1525
Phe 1540
Tyr 1555
Gly 1570
Be 1585
Asp 1 600
Be 1615
Leu 1630
Glu 1645
1485 Ser Ala Het Pro Glu 1500 Lys Asp Asp Gly Asn 1515 Gly Asp 'l'hr Leu Val 1530 Lys Glu Asp Gly Asn 1545 Asn Ser His Asn Val 1560 lle Lys Val Asn Phe 1575 Val Gln Leu Asp Wing 1590 Gly Pro val Leu Leu 1605 Ala Leu Ser Lys Asp 1620 Leu Glu Phe Val Thr 1635 Leu Tyr Lys
Gly 'ryr Val
Tyr Lys' l'hr
Asn Arq Ile
Ile LQu Gly
'l'yr 11e Met
Lys I1e Arq
Bis Tyr Gln
Pro Asp Asn
Pro Asn Glu
Wing Gly Wing
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Artificial Sequence Synthetic Polypeptide" </dd></dl>
<400> 44
H8t ABp Lya Lya Tyr Ser 1le Gly Leu Aap 1le Gly rhr ABn Ser Val 1 S 10 15
Gly Trp ~ a Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe
20 25 30
Lys Val Leu Gly Asn Thr Asp Arg Bis Ser Ile Lys Lys Asn Leu Ila
35 40 4S
Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu ~ a Thr Arg Leu 50 ~ 60
Ly8 Arq Thr Ala Arq Arg Arg Tyr Thr Arg Arq Lys Asn Arq Ile eye 65 70 75 80
Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95
Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110
His Glu Arq Bis Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125
His Glu Lys Tyr Pro Thr Ile Tyr Bis Leu Arg Lys Lys Leu Val Asp 130 135 140
Ser Thr Asp Lys Wing Asp Leu Arg Leu Ile Tyr Leu Wing Leu Wing His 145 150 155 160
Het Ile Lys Phe Arq Gly Bis Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175
Asp Asn Ser Aep Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190
Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205
Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arq Arq Leu Glu Asn 210 215 220
Leu Ile Wing Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn
225 230 235 240
Leu rle Ala Leu Ser Leu G ~ and Leu Thr Pro Asn Phe Lys Ser Asn Phe
245 250 255
Asp Leu Wing Glu Asp Wing Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp
2GO 265 270
ASp Asp Leu Asp Asn Leu Leu A ~ a G ~ n]: ~ e G ~ and Asp G ~ n Tyr A ~ a Asp 215 280 285
Leu Phe Leu A ~ a Ala Lys Asn Leu Ser} ~ sp Ile Leu Leu Ser Asp 290 295 300
Ile Leu Arq val Asn Thr Glu Ile Thr I, ys Ala Pro Leu Ser Ala Ser 305 310 315 320
Met Ile Lys Arq Tyr Asp Glu His Bis Gln Asp Leu Thr Leu Leu Lys 325 ~ 130 335
Wing Leu Val Arq Gln Gln Leu Pro Glu I, and Tyr Lys Glu Ile Phe Phe 340 345 350
_ ~ ~ _ ~~ _ ~ _ ~~ li._
355 360 365
Gln Glu Glu Phe Tyr Lys Phe l1e Lys E'ro Ile Leu Glu Lys Met Asp 310 315 380
Gly Thr Glu Glu Leu Leu Val Lys Leu l ~ n Ar9 Glu Asp Leu Leu Ar9 385 390 395 400
Lys Gln Arq Thr Phe Asp Asn G1y Ser 1:18 Pro Bis Gln I1e Bis Leu 405 ~ IIO 415
Gly Glu Leu Bis Ile Wing Leu Arq Arg Oln G1u Asp Phe Tyr Pro Phe 420 425 430
Leu Lys Asp Asn Arq Glu Lys Ila Glu I.ys Ile Leu Thr Phe Arq 11a .35 440 445
Pro Tyr Tyr Val Gly Pro Leu Wing Arch Cay Asn Ser Arq Phe Wing Trp 450 455 460
MAt. Thr Arq Lys Sar Glu Glu Thr Ile 1 ~ hr Pro Trp Aso Phe Glu Glu 465 470 475 480
Val Val Asp Lys Gly Ala Ser Gln Ser Pha Ila Glu Arq Met Thr 485 ~ 190 495
Aso Phe Asp Lys Aso Leu Pro Asn Glu] ~ and B Val Leu Pro Lys His Ser 500 505 510
Leu Leu Tyr Glu Tyr Phe Thr Val Tyr J ~ sn Glu Leu Thr Lys Val Lys 515 520 525
Tyr Val Thr Glu Gly Met Arq Lys Pro Wing Phe Leu Ser Gly Glu Gln 530 535 540
Lys Lys Wing Ile Val Asp Leu Leu Phe Lys Thr Asn Arch Lys Val Thr 545 550 555 560
Val Lys Gln Leu Lya Glu Asp Tyr Phe Lya Lys Ile Glu eye Phe ASp 565 570 575
Ser Val Glu Ile Ser Gly Val Glu Asp Arq pha Asn Ala Ser Leu Gly 580 585 590
Thr Tyr Ris Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605
Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620
Leu Phe Glu Asp Arq Glu Met Ile Glu Glu Arq Leu Lys Thr Tyr Ala 625 630 635 640
Ris Lau Phe Asp A8P Lys Val Met Lys Gln Leu Lys Arq Arq Arq Tyr 645 650 655
Thr Gly Trp Gly Arq Leu Ser Arq Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670
Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685
Wing Asn Arch Asn Phe Het Gln Leu Ile Bis Asp Asp Ser Leu Thr Phe 690 695 700
Lys Glu Asp Ile Gln Lys Wing Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720
Ris Glu His 11e Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735
l1e Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750
Arq Ris Lys Pro Glu Asn 11e Val Ile Glu Met Ala Arq Glu Aso G1n 755 760 765
Thr Thr Gln Lys Gly Gln Lys Asn Ser Arq Glu Arq Met Lys Arq Ile 770 775 780
Glu Glu Gly Ile Lys Glu Leu Gly Ser G, ln Ile Leu Lys Glu Uis Pro 785 790 795 800
val Glu Asn Thr Gln Leu Gln Asn Glu L: ys Leu Tyr Leu Tyr Tyr Léu 805 810 815
Gln Asn Gly Arq Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830
Leu Ser Asp Tyr Asp Val Asp 8is Ile v.al Pro Gln Ser Phe Leu Lys 835 8 ~ 0 845
Asp Asp Ser Ile Asp A9n Lys Val Leu Tjhr Arg Ser Asp LyS Asn Arg 850 8SS 860
Gly Lys Ser Asp Asn Val Pro Ser Glu G, lu Val Val Lys Lys Met Lys 865 870 875 880
Asn t'yr Trp Arg Gln Leu Leu Asn Ala L: ys Leu Ile t'hr Gln Arg Lys 885 890 89 ~
Phe Asp Asn Leu Thr Lys Ala Glu Arg G, ly Gly Leu Ser Glu Leu Asp .00 905 '10
Lys Wing Gly Phe Ile Lys Arg Gln Leu V. al Glu Thr Arq Gln Ile Thr '15 '20 .25
Lys His Val Wing Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp '30 935 9 ~ 0
Glu Asn Asp Lys Leu Ile Arq Glu Val L: ys val Ile Thr Leu Lys Ser '45 '50 .55 960
Lys Leu Val Ser Asp phe Arg Ly8 Asp Phe Gln Phe Tyr Lys val Arq 965 9'70 975
<dl><dt>Glu </dt><dd>Ile Asn Asn Tyr Bis 980 Bis Wing Bis Aisp Wing Tyr Leu Asn 985 990 Val wing </dd></dl>
<dl><dt>Val </dt><dd>Gly Thr Wing 995 Leu Ile Lys Lys 1000 Tyr Pro Lys I.eu Glu 1005 Be Glu Phe </dd></dl>
Val Tyr Gly A8p Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Wing 1010 1015 1020
Lys Ser Glu Gln Glu tle Gly Lys Ala Thr Ala Lys Tyr Phe Phe
1025 1030
Tyr Ser Asn r1e Met Asn Phe 1040 1045
Asn Gly Glu rle Arg Lys Arq 1055 1060
Thr Gly Glu Ile Val Trp Asp 1010 1015
Arg Lys val Leu Ser Met Pro 1085 1090
Glu Val Gln Thr Gly Gly Phe 1100 1105
Arq Asn Ser Asp Lys Leu ne 1115 1120
Lys Lys Tyr Gly Gly Phe Asp 1130 1135
LeU Val Val Wing Lys Val Glu 1145 1150
Ser Val Lys Glu Leu Leu Gly 1160 1165
Phe Glu Lys Asn Pro Ile Asp 1175 1180
Glu Val Lys Lys Asp Leu Ile 1190 1195
Phe Glu Leu Glu Asn Gly Arg 1205 1210
Glu LeU Gln Lys Gly Asn Glu 1220 1225
Asn Phe Leu Tyr Leu Ala Ser 1235 1240
Pro Glu Asp Aen Glu Gln Lys 1250 1255 1035
Phe Lys: 'l'hr G1u ne Thr Leu Ala 1050
Pro Leu Ile Glu Thr Asn Gly Glu 1065
Lys Gly Arq Asp Phe Ala Thr Val 1080
Gln Val. Aan Ile Val Lys Lys' lhr 1095
Ser Lysl Glu Ser ne Leu Pro Lys 1110
Wing Ar ~ r Lys Lys Asp trp Asp Pro 1125
Ser Prcl 'l'hr Val Ala' l'yr Ser Val 1140
Lys Gl ~ 'Lys Ser Lys Lys Leu Lys 1155
Ile Thl: Ile Met Glu Arq Ser Ser 1170
Phe Leu Glu Ala Ly. G1y 'l'yr Lys 1185
.ne Lyl! l Leu pro Lys Tyr Ser Leu 1200
Lys Art ;; r Met Leu Ala Ser Ala G1y 1215
Leu AleL Leu Pro Ser Lys Tyr Val 1230
His Tyl: Glu Lys Leu Lya Gly Ser 1245
Gln Let: l Phe Val Glu Gln Bis Lys 1260
Bis Tyr Lel, l Asp Glu Ile ne 1265 1270
Arq Val I1e Leu Wing Asp Wing 1280 1285
Tyr Asn Lys Bis Arq Asp Lys 1295 1300
ne ne Mia Leu Phe Thr Leu 1310 1315
Phe Lys Tyr Phe Asp Thr Thr 1325 1330
Thr Lys Glu Val Leu Asp Ala 1340 1345
Gly Leu Tyr Glu Thr Are¡ ne 1355 1360
Wing Wing Wing Val Ser Lys Gly 1310 1315
Pro no Leu Val Glu Leu Asp 1385 1390
Ser val Ser G1y Glu G1y Glu 1400 1405
Thr Leu Lys Phe Ile eys Thr 1415 1420
Pro Thr Leu Val Thr Thr Leu 1430 1435
Arq Tyr Pro Aep Mis Met Lys 1445 1450
"et Pro G1u G1y Tyr Val Gln 1460 1465
Asp Gly Aan Tyr Lys Thr Arq 1475 1480
Thr Leu Val Asn Are¡ Ile Glu 1490 1495
Glu Gl: n Ile Ser Glu Phe Ser Lys 1275
Asn Le'u Asp Lya Val Leu Ser Wing 1290
Pro Il, e Are¡ G1u Gln Ala G1u Asn 1305
Thr Asn Leu G1y Pro Wing Ala Wing 1320
Ile As: p Arq Lys Arq Tyr Thr Ser 1335
Thr Le'u I1e Mis Gln Ser Ile Thr 1350
Asp Leu Ser Gln Leu Gly Gly Asp 1365
Glu Glu Leu Phe Thr Gly Val Val 1380
Gly As, p Val Asn Gly His Lys Phe 1395
Gly As: p Wing Thr Tyr Gly Lys Leu 1410
Thr Gly Lys LQu Pro Val Pro Trp 1425
Thr Tyr Gly Val Gln eya Phe Ser 1440
G1n Bis Aap Phe .h. Ly. Ser Ala 1455
G1u Arg Thl; Ile Phe Phe Lys Asp 1410
Wing Glu Val Lys Phe G1u G1y Aap 1485
Leu Lys Gly Ile Asp Phe Lys Glu 1500
Asp G1y Asn Ile Leu Gly His Lys Leu Glu Tyr ASn Tyr Asn Ser 1505 1510 1515
Bis Asn Val Tyr I1e Met Wing Asp Lys Gln Lys Asn Gly I1e Lys 1520 1525 1530
Val Asn Phe Lys Ile Arg Bis Asn 11e Glu Asp Gly Ser Val Gln 1535 1540 1545
Leu Ala Asp His Tyr Gln G1n Aso Thr Pro 1le G1y ABp Gly Pro 155Q 1555 1560
Val Leu Leu Pro Asp Asn Hie Tyr Leu Ser Thr G1n Ser Ala. Leu 1565 1510 1575
Sr. Lys Asp Pro Asn Glu Lyo Arq Asp His Met Val Leu Leu Glu 1580 1585 1590
Phe val Thr Ala Wing Gly ne Thr Leu Gly Met ASp Glu Leu Tyr 1595 1600 1605
Lys Lys Arg Pro Wing Thr Wing Lys Lys Wing Gly G1n Wing Lys Lys 1510 1615 1620
Lys Lys
<210> 45
<211> 1664 5 <212> PRT
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Artificial Sequence" Synthetic polypeptide " </dd></dl>
10 <400> 45
Mat Asp Tyr Lys Asp Bis Asp Gly Asp '~ yr Lys Asp Bis Asp Ile Asp 1 S lO 15
Tyr Lys Asp Asp Asp Asp Lys Met Wing Pro Lys Lys Lys Are¡ Lys val 20 25 30
Gly Ile His Gly val Pro Ala Wing Asp l: 'ys Lys Tyr Ser Ile Gly Leu
35 .0 '5
Asp Ile Gly Thr Asn Ser Val Gly Trp JUa val Ile Thr Asp Glu Tyr 50 SS 60
Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arq His 65 70 75 80
Be Ile Lye Lys Asn Leu Ila Gly ~ a Leu Leu Phe Aep Be Gly Glu 85 90 95
Thr Wing Glu Wing Thr Arq Leu Lys Arg Thr Wing Arq Arq Arq Tyr Thr 100 105 110
Arg Arg Lys Asn Arq Ila eye Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125
Met Ala Lys Val Asp Asp Ser Phe Phe ais Arq Leu Glu Glu Ser Phe 130 135 140
Leu Val Glu Glu ASp Lys Lys ~ iB Glu Arq His Pro Xle Phe Gly Asn 145 150 155 160
Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr Bis 165 170 175
Leu Arq Lys LY8 Leu Val ASp Ser Thr Asp Ly8 Wing Asp Lau Arq Leu 180 185 190
Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly Mis Phe Leu 195 200 205
Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp val Asp Lys Leu Phe 210 215 220
Ila Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe GLu Glu Asn Pro Ile 225 230 235 240
Asn Wing Ser Gly Val Asp Wing Lys Wing Ile Leu Ser Wing Arg Leu Ser 245 250 255
Lys Ser Arg Arl} Leu Glu Asn Leu Ile Wing Gln Leu Pro Gly Glu Lys 26D 265 270
LyS Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 .285
Pro Asn Phe Ly8 Ser Asn Phe Asp Leu Wing Glu Asp Wing Lys Leu Gln
290 295 300
Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Wing Gln 305 310 315 320
Ile Gly ASp Gln Tyr Wing ASp Leu Phe Leu Wing Wing Lys Asn Leu Ser 325 330 335
Aap Wing Ile Leu Leu Ser ASp Ile Leu Arg Val Asn Thr Glu Ile rhr 340 345 350
Lys Ala Pro Leu Ser Ala Ser Met 11e Lys Arq Tyr Asp Glu His Bis 355 360 365
Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arq Gln Gln Leu Pro Glu 370 375 380
Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400
Tyr Ile Asp Gly Gly Ala Ser G1n Glu Glu Phe Tyr Lys Phe I1e Lys 405 410 415
Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430
Asn Arg Glu ASp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445
Ile Pro His Gl.n Ile His Leu Gly Glu Leu Bis Ile Leu Arq Arch 450 455 460
Gln Glu ASp Phe Tyr Pro Phe Leu Lys Asp ASn Arg Glu Lys Ile Glu 465 470 415 480
Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr val Gly Pro Leu Ala Ar9
485 490 495
Gly Aan Ser Arg Phe Wing Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510
Thr Pro Trp Asn phe Glu Glu val Val Aap Lys Gly Ala Ser Ala Gln 515 520 525
Ser Phe Ile Glu Arg Met Thr ABn Phe Asp Lys Asn Leu Pro Aan Glu 530 535 540
Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560
Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lya Pro
565 5'70
Wing Phe Leu Ser Gly Glu Gln Lys Lys Wing Ile Val Asp Leu Leu Phe 580 585 5.0
Lys Thr Asn Arq Lys Val Thr Val Lys Gln Leu Lys Glu Asp 'I'yr Phe 595 600 605
Lys Lys Ile Glu Cys Phe Asp Ser Val Gl \, 1 Ile Ser Gly Val Glu Asp 610 615 620
___ ~ _ ~ m. ~ _tis_ ~~~ yes ~ I ~
625 630 635 640
Lys Asp Lys Asp Phe Leu Asp Asn Glu G, lu Asn Glu Asp IIa Leu Glu 645 6.50 655
Asp He Val Leu Thr Leu Thr Leu Phe Glu ASp Arq Glu Met Ila Glu 650 665 670
Glu Arq Leu Ly. Thr Tyr Ala His Leu Phe Asp Asp Lya Val Met Lys 675 680 685
Gln Leu Lys Argo Arg Arq Tyr 'I'hr Gly T: tp Gly Arq Leu Ser Arq Lys 690 695 700
Leu rIe Asn Gly rle Argo Asp Lys Gln SI9r Gly Lys Thr Ile Leu Asp 705 110 715 720
Phe Leu Lys Ser Asp Gly Phe Ala Asn A ~ q Asn Phe Met Gln Leu Ile 725 7:30 735
His Asp A.sp Ser Leu 'I'hr Phe Lys Glu ~ 9p rle Gln Lys Ala Gln val
740 745 750
Be Gly Gln Gly Asp Be Leu Bis Glu His! Le Ala Asn Leu Ala Gly
7.5.5 760 76.5
Be Pro ~ a lle Lys Lys G '. Ile Leu G: Ln Thr Vd Lys Val Val Asp 770 775 780
Glu Lau Val Lys Val Met Gly Arq Mis L: í'8 Pro Glu Asn Ile Val ne 785 790 795 800
= ~~ _ = _ ~~~~~ s_ ~ __ _
Arg Glu Arg Met Lys Arq Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830
Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845
Lys Leu ryr Leu Tyr Tyr Leu Gln Asn Gly Arq Asp Met Tyr Val Asp 850 855 860
Gln Glu Leu Asp Ile Asn Arq Leu Ser Asp ryr Asp Val Asp Bis Ile 865 870 875 880
Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895
Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910
Glu val val Lys Lys Het Lys Asn ryr Trp Arg Gln Leu Leu Asn Ala 915 920 925
Lys Leu Ile rhr Gln Arq Lys Phe ASp Asn Leu rhr Lys Ala Glu Arq 930 935 940
Gly Gly Leu Ser Glu Leu Asp Lys Wing Gly Phe Ile Lys Arq Gln Leu 945 950 955 960
Val Glu Thr Arq Gln Ile Thr Lys Bis Val Ala Gln Ile Leu Asp Ser 965 970 975
Arq Met Asn rhr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arq Glu Val 980 985 990
Lys val Ile rhr Leu Lys Ser Lys Leu Val Ser Asp Phe Arq Lys A 995 1000 1005
Phe Gln Phe ryr Lys Val Arq Glu Ile Asn Asn Tyr Bis His Ala 1010 1015 1020
Bis Asp Wing Tyr Leu Asn Wing Val Val Gly Thr Wing Leu Ile Lys 1025 1030 1035
Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp t'yr Lys 1040 1045 1050
Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065
Gly Lys Ala Thr Ala Lys Tyr 1070 1015
Phe Phe Lys Thr Glu Ile Thr 1085 1090
Arq Pro Leu l1e Glu Thr A8n 1100 1105
A8p Lys Gly Arq Asp Phe Al. 1115 1120
Pro Gln Val Asn Ile Val Lys 1130 1135
Phe Ser Lys Glu Ser Ile Leu 1145 1150
ne Ala Arq Lya Lye Asp Trp
1160 1165
Asp Ser Pro Thr Val Ala Tyr 1175 1180
Glu Ly8 Gly Lys Ser Lys Lys 1190 1195
Gly He Thr I1e Met Glu Arg 1205 1210
Asp phe Leu Glu Ala Lys Gly 1220 1225
ne ne Lys Leu Pro Lys Tyr
1235 1240
Arq Lys Arq Met Leu Ala Ser
1250 1255
Glu Leu Ala Leu Pro Ser Lys 1265 1270
Ser His Tyr Glu Lys Leu Lys 1280 1285
Lys Gln Leu Phe Val Glu Gln 1295 1300
Phe Phe Tyr Ser Asn Ila Het Asn 1080
Leu Wing Asn Gly Glu Ile Arq Lya 1095
Gly Glu Thr Gly Glu l1e Val 'I'rp 1110
Thr Val Are¡ Lys Val Leu Ser Met 1125
Lys' rhr GIu Val Gln Thr Gly Gly 1140
Pro Lye Arg Asn Ser Asp Lys Leu 1155
ASp Pro Lys Lys Tyr Gly Gly Phe 1170
Ser Val Leu Val Val Wing Ly6 Val 1185
Leu Lys Ser Val Lys Glu Leu Leu l200
Ser Ser Phe Glu Lys Asn Pro Ile 1215
Tyr Lys Glu Val Lys Lys Asp Leu 1230
Ser Leu Phe Glu Leu Glu Asn Gly 1245
Wing Gly Glu Leu Gln Lys Gly Asn 1260
Tyr Val Asn Phe Leu Tyr Leu A1a 1215
Gly Ser Pro Glu Aap Aan Glu Gln 1290
Bis Lys His Tyr Leu Asp Glu Ile 1305
Ile Glu Gln Ile Ser Glu Phe Ser Ly! 1 Arg Val Ile Leu Ala Asp 1310 1315 1320
Al. Aan Leu Asp Lys Val Leu Ser Alu Tyr Asn Lys His Arg Asp 1325 1330 1335
Lys Pro Ila Are¡ GI u Gln Ala Glu Asn Ile IIa His Leu Pbe Thr 1340 1345 1350
Leu Thr Asn Leu Gly Wing Pro A1a Wing Phe Lys Tyr Phe Asp Thr 1355 1360 1365
Thr Ile Asp Arg Lys Arq Tyr Thr 581: 'Tbr Lya Glu Val Leu Asp 1370 1375 1380
Ala Thr Leu Ile His Gln Ser Ile 'lh) :, Gly Leu Tyr Glu Thr Arq 1385 1390 1395
Ile Asp Leu Ser Gln Leu Gly Gly Asp Wing Wing Wing Val Ser Lys 1400 1405 1410
Gly Glu Glu Leu Pbe Thr Gly Val val Pro Ile Leu Val Glu Leu 1415 1420 1425
Asp Gly Asp Val Asn Gly Ris Lys PhH Ser Val Ser Gly Glu Gly 1430 1435 1440
Glu Gly Asp Wing orbr Tyr Gly Lys Leu 'l'hr Leu Lys Phe Ile eys 1445 1450 1455
Thr Thr Gly Lye Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr 1460 1465 1470
Leu Thr Tyr Gly Val Gln Cys Phe Sel: 'Arq Tyr Pro Aap His Mat 1475 1480 1485
Lys Gln His Asp Phe Phe Lys Ser AJ.a Mat Pro Glu Gly Tyr Val 1490 1495 1500
Gln Glu Arq Thr Ile Phe Phe Lys Asp Asp Gly Aan Tyr Lys Thr 1'05 1510 1515
Arq Ala Glu Val Ly8 Phe Glu Gly Asp Thr Leu Val Aso ArO '1le 1520 1525 1530
G1u Leu Lys Gly l1e Asp Phe Lye Glu Asp G1y Asn 11e Leu Gly
<dl><dt>1535 </dt><dd> 1540 1545 </dd></dl>
<dl><dt>Bis </dt><dd>Ly. 1550 Leu Glu Tyr Asn Tyr 1555 Asn Sar Bis Asn val 1560 Tyr I1e Met </dd></dl>
<dl><dt>Al. Asp 1565 </dt><dd>Lys Gln Lys Asn Gly 1570 Ila Lys Val Asn Phe 1575 Lys I1e Arq </dd></dl>
<dl><dt>His Asn 1580 </dt><dd>11th Glu Asp Gly Be 1585 val Gln LeU Wing ASp 1590 Bis Tyr G1n </dd></dl>
<dl><dt>Gln Asn 1595 </dt><dd>'l'hr Pro Ile Gly Asp 1600 Gly Pro Val Leu Leu 1605 Pro Asp Asn </dd></dl>
<dl><dt>Bis </dt><dd>'yr 1610 Leu Be Thr Gln Be 1615 Wing Leu Ser Lys ASp 1620 Pro Asn Glu </dd></dl>
<dl><dt>Lys Arq 1625 </dt><dd>Asp Bis Het Val. Leu 1630 Leu Glu Phe Val Thr 1635 Wing wing Gly </dd></dl>
<dl><dt>Ile Thr 1640 </dt><dd>Leu Gly Met Asp Glu 1645 Leu 'l'yr Lys Lys Arq 1650 Pro wing To </dd></dl>
<dl><dt>Thr Ly. 1655</dt><dd>Lys Gly Gln wing Wing 1660 Lya Lys Lys Lys </dd></dl>
<dl><dt>5 </dt><dd><210> 46 <211> 1423 <212> PRT <213> Artificial Sequence </dd></dl>
<dl><dt>10 </dt><dd><220> <221> source <223> Inota = ~ Artificial Sequence Description l: Polypeptide if ntét i co ~ </dd></dl>
<dl><dt><400> 46 </dt><dd /></dl>
Met ASp Tyr Lya Asp 8is Asp Gly Asp 'l ~ yr Lys Asp His ASp 11e ASp 1 5 1.0 15
Tyr Lys Aep Asp Asp Asp Lys Met AlOl P'ro Lye Lye Ly "Arq Lys Val 20 25 30
Gly Ile Mis Gly val Pro Ala Wing Asp ¡, ys Lys Tyr Ser 11e Gly Leu 35 40 45
Asp Ile Gly Thr Asn Ser Val Gly Trp ~~ a Val Ile Thr Asp Glu Tyr 50 55 ~
Lys VOll Pro Ser Lys Lys Phe Lys Val l ~ u Gly Asn Thr Aep Ar9 Hie
70 75 80
Ser lle Lys Lys Asn Leu I1e Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95
Tbr Ala Glu Ala Thr Arq Leu Lys Ar. 'thx: Ala Arq Arq Arq Tyr 1'hr 100 105 110
Arq Arq Lys Asn Arq rle Cys Tyr Leu Gln Slu Ile Phe Ser Asn Glu 115 120 125
Met Ala Lys Val Asp Asp Ser Phe Phe Hisl Ar9 Leu Glu Glu Ser Phe 130 135 140
Leu Val Glu Glu Asp Lyo Lys His Glu Ar9r His Pro I1e Phe Gly Asn 145 150 155 160
I1e Val Asp Glu val Ala Tyr Hie Glu Lys: Tyr Pro Thr Ile t'yr Bis 165 1701 175
Leu Arq Lys Lys Leu Val Asp Ser Tbr Asp Lys Wing Asp Leu Arq Leu 180 185 190
Ila Tyr Leu Ala Leu Ala His Met Ile Lyl! L: Phe Arq Gly ais Phe Leu 195 200 205
Ile Glu Gly Asp Leu Asn Pro Asp Asn Se.I: Asp Val Asp Lys Lau Phe 210 215 220
rle Gln Leu Val Gln Thr Tyr Asn Gln Le \ ll Ph. Glu Glu Asn Pro n. 225 230 235 240
Asn Ala Ser Gly Val Asp Ala Lys Ala 11 .., Leu Ser Ala Arq Leu Ser 245 250 255
Lys Ser Arq Arq Leu Glu Asn Leu Ile Alal Gln Leu Pro Gly Glu Lys 260 265 270
Lye Asn Gly Leu Phe Gly Asn Leu Ila Ala, Leu Ser Leu Gly Leu Thr 215 280 285
Pro Asn Phe Lya Ser Aan phe Asp Leu Al ~ L Glu Asp A1a Lya Leu Gln 290 2.5 300
Len Ser Lys A8p Thr Tyr Asp A8p Asp Leu Asp Asn Leu Leu Gln Wing 305 310 315 320
Ile Gly Asp Gln Tyr A1a Asp Leu Phe Le \ :! Wing Wing Lys Asn Leu Ser 325 330 335
Asp Wing Ile Lau Lau Being Asp Ile Lau ~, Val Asn Thr G1u Ile Thr 340 345 350
Lys Ala Pro Leu Ser A1a Ser Met Ile Lyl! 1 Argo Tyr Asp Glu His His
355 360 365
Gln Asp Leu Thr Leu LeU Lys Ala Leu Val. Arg G1n Gln Leu Pro Glu
370 375 380
• and. 'l: yr Lys Glu Ile Phe Phe Asp Gln Sel: .y • Asn Gly Tyr A1a G1y 385 390 395 400
Tyr Ile Asp Gly Gly A1a Ser Gln Glu GI "; I Phe Tyr Lya phe ne .y. 405 41 (1 415
Pro Ile Leu G1u Lys Met Aep Gly Thr Gl "; l Glu Leu Leu Va1 Lys Leu 420 425 430
Asn Arg Glu Asp Leu Leu Arq Lys Gln Ar ~, Thr Phe Asp Asn Gly Ser 435 440 445
Ile Pro 8is Gln Ile Bis Leu Gly Glu Le "; 1 8is A1a Ile Leu Arq Arch
450 455 460
Gln Glu Asp Phe Tyr Pro Phe Leu Lys Aap Asn Arq Glu Lys Ile G1u 465 470 475 480
Lys Ile LeU Thr Phe Arq Ile pro Tyr Tyl: val Gly Pro Leu A1. Arq 485 490 495
Gly Asn Ser Arq Phe A1a Trp Met Thr ~, Lys Ser Glu Glu Thr Ile 500 505 510
Thr Pro Trp Asn Phe Glu Glu va1 val Asp Lys Gly "to Be Wing Gln 515 520 525
Ser Phe Ile Glu Argo Met '1'hr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540
Lys Val Leu Pro Lys His Ser Leu Leu Tyl: Glu Tyr Phe Thr Val Tyr 545 550 555 560
Asn Glu Leu Thr. And Val Lys Tyr Val Thr Glu Gly Met Arch. And Pro 565 57C1 575
Wing Phe Leu Ser Gly Glu Gln Lys Lys Wing lla Val Asp Leu Leu Phe S80 S8S 590
Lys Thr Aan Arq Lys Va1 Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595 600 605
Lys Lys lla Glu eye Phe Asp Ser Val Glu lla Ser Gly Val Glu Asp 610 615 620
Arq Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys lle n.
6.25 630 635 640
Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655
Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp ArO Glu Met lle Glu 660 665 670
Glu Arq Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685
Gln Leu Lys ArO ArO Arq Tyr 'l'hr Gly Trp Gly ArO Leu Ser Arq Lys 690 695 700
Leu Ile Asn Gly Ile Arq Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 70S 710 715 720
Phe Leu Lys Ser Aap Gly Phe Ala Asn ArO Asn Phe Met Gln Leu lle 725 730 735
Bis ASp Asp Ser Leu Thr Phe Lys Glu Asp tle Gln Lys Wing Gln Val 740 745 750
Ser Gly Gln Gly Asp Ser Leu His Glu Bis Lle Ala Asn Leu Ala Gly 755 '60 765
Be Pro Wing Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Va1 Asp 770 775 780
Glu Leu Val Lys Val Met Gly Arq His Lys Pro Glu Asn Ile Val ne 785 790 795 800
Glu Mot Wing Arg GIu Asn GIn Thr Thr G1n Lys G1y GIn Lys Asn Ser 805 810 815
Arg Glu Arq Met Lys Arq Ile Glu Glu Gly Ile Lye Glu Leu Gly Ser 820 '25 830
Gl.n Il.e Leu Lys Gl.u His Pro val Gl.u ASn Thr Gl.n Leu Gl.n Asn Glu 835 840 845
Lya Leu Tyr Leu Tyr Tyr Leu Gl.n Asn Gly Arq Asp Met Tyr Val Asp 850 855 860
Gln Glu Leu Aap Ile An Arq Leu Ser Asp Tyr Asp Val Asp Bis Ue 865 870 875 880
Val Pro Gln Ser Phe Leu Lys Asp Aap Ser Ile Asp Asn Lys Val Leu 885 890 .95
Thr Arq Ser Aap Lys Asn Arq Gly Ly8 Ser Asp Asn Val Pro Ser Glu 900 905 910
Glu val Val Lys Lys Met Lys Asn Tyr Trp Arq Gln Leu Leu Asn Ala 915 920 925
Lys Leu 11e Thr Gln Arq Lys Phe Asp Asn Leu Thr Lys Al.a Glu Arq 930 935 940
Gly Gly Leu Ser Glu Leu Asp Lya Ala Gly Ph. Ile Lys Arq Gln Leu 945 950 955 960
Val Glu Thr Arq Gln Ile Thr Lys Bis Val Ala Gln 11. Leu Asp Ser 965 970 975
Arq Met Aso Thr Ly5 Tyr Asp Glu Asn A4p Lye Leu Ile Arq Glu Val 980 9.5 990
Lys Val Ile Thr Leu Lys Ser LyS Leu Y-al Ser Asp Phe Arq Lys Aap 995 1000 1005
Phe Gln Phe Tyr Lys Val Arq Glu Ile Asn Asn Tyr Bis Bis Wing 1010 1015 1020
Bis Aap Wing Tyr Leu Asn Wing Val Val Gly Thr Wing Leu Ile Lys 1025 1030 103! I
Lys Tyr Pro Lys Leu Glu Ser Glu Phe val Tyr Gly Asp Tyr Lys 1040 1045 lOSO
Val Tyr Asp Yal Arq Lys Met Ile Ala Lys Ser Glu Gln Glu 11e 1055 1060 1065
Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser ASn Ile Met Asn
1070 1075 1080
Ph. Phe Lys Thr Glu Ile Thr Lau Alél Asn Gly Glu: Ue Arg Lys 1085 1090 1095
Arg Pro Leu Ile Glu Thr Asn Gly Glu Tbr Gly Glu tle Val Trp 1100 1105 1110
Asp Lys Gly Arq Asp Phe Ala Thr Val Arq Lys val Leu Ser Meto 1115 1120 1125
Pro Gln Val Asn Ile Val Lys Ly s ThJ: Glu Val Gln Thr Gly Gly 1130 1135 1140
Phe Ser Lys Glu Ser Ile Leu Pro Lyfl Arq Asn Ser Asp Lys Leu 1145 1150 1155
Ile Ala ArQ Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170
Asp Ser Pro Thr val Wing Tyr Ser Val Leu Val Val A1e Lys Val 1175 1180 1185
Glu Lys Gly Lys Ser Lys Lys Leu Ly! 1 Ser Val Lys Glu Leu Leu 1190 1195 1200
Gly Ile Thr Ile Meto Glu Arq Ser Sel: Phe Glu Lys Asn Pro Ila 1205 1210 1215
Asp Phe Leu Glu Ala Lys Gly Tyr Ly ~ 1 Glu Val Lys Lys Asp Leu 1220 1225 1230
or. Ile Lyo Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Aon Gly 1235 1240 1245
Arg Lys Arg Met Leu Wing Being Wing GI ~ r Glu Leu Gln Lys Gly Aso 1250 1255 1260
Glu Leu Ala Leu Pro Ser Lys 7yr Val. Asn Phe Leu Tyr Leu Ala 1265 1270 1275
Ser Bis Tyr Glu Lys Leu Lys Gly Se% ~ Pro Glu Asp Asn Clu Gln 1280 1285 1290
Lys Gln Leu pha Val Glu Gln His Lyfl Hia Tyr Leu Asp Glu Ile 1295 1300 1305
<dl><dt>ne Glu 1310 </dt><dd>Gln I1e Be Glu Ph. 1315 Be Lys Arg Val ne 1320 Leu Ala Asp </dd></dl>
<dl><dt>Al. Asn 1325 </dt><dd>Leu ASp Lys Val Leu 1330 Ser Ala Tyr Asn Lys 1335 Bis Arq Asp </dd></dl>
<dl><dt>LyS </dt><dd>Pro 1340 Ile Arg Glu Gln Wing 1345 Glu Asn Ile Ile Bis 1350 Leu Phe Thr </dd></dl>
<dl><dt>Leu </dt><dd>Thr 1355 Asn Leu Gly Ala Pro 1360 Wing Wing Phe Lys Tyr 1365 Phe Asp Thr </dd></dl>
<dl><dt>Thr </dt><dd>ne 1370 Asp Arq Lys Arg Tyr 1375 Thr Ser Thr Lys Glu 13 80 val Leu Asp </dd></dl>
<dl><dt>Al. Thr 1385 </dt><dd>Leu Ile Bis Gln Be 1390 Ile Thr Gly Leu Tyr 1395 Glu Thr Arq </dd></dl>
<dl><dt>n. Asp 1400</dt><dd>Leu Be Gln Lau Gly 1405 Gly Asp Lys Arq Pro 1410 Al.a Ala Thr </dd></dl>
<dl><dt>Lys </dt><dd>Lys 1415 Ala Gly Gln AJ.a Lys 1420 Lys Lys Lys </dd></dl>
<dl><dt>5 </dt><dd><210> 47 <211> 483 <212> PRT <213> Artificial Sequence </dd></dl>
<dl><dt>10 </dt><dd><220> <221> source <223> Inota = ~ Artificial Sequence Description "Synthetic polypeptide ~ </dd></dl>
<400> 47
Het Phe Leu Phe Leu Se ~ Leu Th ~ Se ~ Phe Leu Be Be Be Arq Thr 1 5 ID 15
Leu Val Ser Lys Gly Glu Glu Asp Asn Met Ile Wing Ile Lys Glu Phe 20 25 30
Met Arq Phe Lys Val Bis Met Glu Gly Ser Val Asn Gly His Glu Phe 35 40 45
Glu Ile Glu Gly Glu Gly Glu Gly Arg Pro Tyr Glu Gly Thr Gln Thr 50 55 60
~ a Lys Leu Lys Val Thr Lys Gly Gly Pro Leu Pro Phe ~ a Trp ASp 65 70 75 80
I ~ __ ~ mn ~ S ~~. ~ _ ~ _ ~
~ H e
Pro Wing Asp Ile Pro Asp Tyr Leu Lys lAU Ser Phe Pro Glu Gly Phe 100 105 110
Lys Trp Glu Arq Val Met Aso Phe Glu J. ~ p Gly Gly Val Val Thr Val 115 120 125
Thr Gln Asp Ser Ser Leu Gln Aap Gly 01u Phe na Tyr Lys val Lys 130 135 140
Leu Arq Gly Thr Asn Phe Pro Ser Asp c: ay Pro Val! <! Et Gln Lys Lys 145 150 155 UO
Thr Met Gly Trp Glu A1a Ser Ser Glu I ~ q Met Tyr Pro Glu Asp Gly 165 1.1'0 175
Wing Leu Lys Gly Glu Ile Lys Gln Arg l ~ u Lys Leu Lys Asp Gly Gly lBO lB5 190
Mis Tyr A8p Wing Glu val Lys Thr Thr 1 ~ r Lys Wing Lys Lys Pro val 195 200 205
Gln Leu Pro Gly Al.a Tyr Aso Val Aso 1: 1e Lys Leu Asp Ile Thr Ser 210 215 220
His Aso Glu Asp Tyr Thr Ile vOl1 Glu Gln Tyr Glu ArO Ala Glu Gly 225 230 235 240
Arg His Ser Thr Gly Gly Met Asp Glu] ~ u Tyr Lys Gly Ser Lys Gln 245 ~! 50 255
Leu Glu Glu Leu Leu Ser Thr Ser Phe} ~ SP Ile Gln Phe Aao Asp Leu 260 265 270
Thr Leu Leu Glu Thr A1a Phe Thr Nis, ~ h: r: Se: r: Ty: r: lia Asn Glu Hi. 275 280 285
Arg Leu Leu Asn Val Ser His Asn Glu} ~ q Leu Glu Phe Leu Gly Asp 290 295 300
Wing Val Leu Gln Leu Ile Ila Ser Glu '~ yr Leu Phe Al.a Lys Tyr Pro 305 310 315 320
Lys Lys Thr Glu Gly Asp Met Ser Lys l ~ u Arg Ser Met Ile Val Arg 325 330 335
Glu Glu Ser Leu Ala Gly Phe Ser Arq Phe Cys Ser Phe Asp Ala Tyr 340 345 350
11e Lys Leu Gly Lys Gly Glu Glu Lys Ser Gly Gly Arq Arq Arq ASp 355 360 365
Thr Ile Leu Gly Asp Leu Phe Glu Ala Phe Leu Gly ~ a Leu Leu Leu
370 375 380
ASp Lys Gly Ile Asp Ala val Arq Arq Phe Leu Lys Gln val Met Ile 385 390 395 400
Pro Gln Val Glu Lys Gly Asn Phe Glu Arq Val Lys Asp Tyr Lys Thr 405 410 415
Cys Leu Gln Glu Phe Leu G1n Thr Lys Gly Aep Val ~ a Ile Asp Tyr 420 425 430
Gln Val Ile Ser Glu Lys Gly Pro Wing My Wings Lye Gln Phe Glu Val
435 440 445
Ser Ile Val Val ABn Gly Ala Val LeU Ser Lys Gly Leu Gly Lys Ser 450 455 460
Lys Lys Leu Wing Glu Gln Aap Wing Wing Lys Asn Wing Leu Wing Gln Leu
465 470 475 480
Be Glu val
<210> 48 <211>483
<212> PRT 5 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description: Synthetic Polypeptide ~ </dd></dl>
<400> 48
Met Lys Gln Leu Glu Glu Leu Leu Ser Thr Ser Phe Asp Ile Gln Phe 1 5 10 15
Asn Asp Leu Thr Leu Leu Glu Thr Wing Pha Thr His Thr Ser Tyr Wing 20 25 3D
Asn Glu Bis Arq Leu Leu Asn Val Ser His Asn Glu Arq Leu Glu Phe 35 40 45
Leu Gly Asp Ala Val Leu Gln Leu Ile Ile Ser Glu 'I: yr Leu Phe Ala 50 55 60
Lys' 1'yr Pro Lys Lys Thr Glu Gly Asp Het 1 ~ er Lys Leu Arq Ser Net. 65 70 '15 80
Ile Val Are¡ Glu Glu Ser Lau Ala Gly Phfil Sfilr Arq Phe Cys Ser Phe 85 90 95
ASp Wing Tyr Ile Lys Leu Gly Lys Gly Glu I; lu Lys Ser Gly Gly Are¡ 100 105 110
Arg A.rq Asp Thr Ile Leu Gly Asp Leu Phe l:; lu Ala Phe Leu Gly Ala 115 120 125
Leu Leu Leu Asp Lys Gly Ile Aap Ala. Val Ju-g Arg Phe Leu Lys Gln 130 135 140
val Met Ile Pro Gln Val Glu Lys Gly Asn l? he Glu Are¡ Val Lys Asp 145 150 155 160
Tyr Lys Thr Cys Leu Gln Glu Phe Leu Gln ~ rhr Lys Gly Asp Val Ala
165 170 175
Ile Asp Tyr Gln Val Ila Ser Glu Lys Gly l? Ro Al.a. 8is Al.a Lys Gln 180 185 190
Phe Glu Val Ser Ile Val Val Asn Gly Ala Val Leu Ser Lys Gly Leu 195 200 205
Gly Lys Sar Lys Lys Leu Wing Glu Gln Asp JUa Ala Lys Asn Ala Leu 210 215 220
Wing Gln Leu Ser Glu Val Gly Ser Val Ser Lys Gly G1u Glu Asp Asn 225 230: ~ 35 240
Met Wing 11e Lys Glu Phe Mét Are¡ Phe Lys val His Met Glu Gly 245 250 255
Ser val ASR Gly Bis Glu Phe Glu Ile Glu I; ly Glu Gly Glu Gly Are¡ 260 265 270
Pro Tyr Glu Gly Thr Gln Thr Ala Lys Leu lt.ys Val Thr Lys Gly Gly 275 280 285
Pro Leu Pro Phe Ala 'I'rp Asp Ile Leu Ser] [) ro Gln Phe Met. 'l'yr Gly 2'0 2'5 300
Ser Lys ~ a Tyr Val Lys Bis Pro Wing Asp Ile Pro Asp Tyr Leu Lys 305 310 315 320
Leu Ser Phe Pro Glu Gly Phe Lys Trp Glu Arq Val Met Asn Phe Glu 325 330 335
ASp Gly Gly val Val Thr Val Thr Gln Asp Ser Ser Leu Gln Asp Gly 340 345 350
Glu Phe: J: le Tyr Lys Val Lys Leu Arg Gly Thr Asn Phe Pro Ser Asp 355 360 365
Gly Pro Val Het Gln Lys Lys Thr Het Gly Trp Glu Ala Ser Ser Glu 370 375 380
Arq Met Tyr Pro Glu Asp Gly Ala Leu Lys Gly Glu Ile Lys Gln Arq 385 390 395 400
Leu Lys Leu Lys Asp Gly Gly Bis Tyr Asp Wing Glu Val Lys Thr Thr 405 410 415
Tyr Lys Wing Lys Lys Pro Val Gln Leu Pro Gly Wing Tyr Asn Val Asn 420 425 430
Ile Lys Leu Asp Ile Thr Ser Ris Asn Glu Asp Tyr Thr Ile Val Glu 435 440 445
Gln Tyr Glu Arq A1a Glu Gly Arg Ris Ser Thr Gly Gly Met Asp Glu 450 455 460
Leu Tyr Lys Lys Arq Pro ~ a A1a Thr Lys Lys Ala Gly Gln A1a Lys 465 470 415 480
Lys Lys Lys
<210> 49
<21 1> 1423
<212> PRT 5 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description: Synthetic Polypeptide ~ </dd></dl>
<400> 49 Het Asp Tyr Lys Asp Bis Asp Gly Asp Tyr Lys Asp Bis Asp Ile Asp
1 5 10 15 10
Tyr Lys Asp Asp Asp Asp Lys Het Al.a Pro Lys Lys Lys Arg Lye Val
20 25 30
Gly Ile His Gly Val Pro Wing Wing Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45
Al.a U. Gly Tbr Asn Ser Val Gly Trp Wing Val Ue Tbr Asp Glu Tyr 50 55 60
Lys Val Pro Ser Lys Lys Pbe Lys Val Leu Gly Asn Tbr Asp Arg His 65 70 15 80
Ser Ile Lys Lys Asn Leu Ile Gly ~ a Leu Leu Pbe Asp Ser Gly Glu 85 90 95
Tbr Wing Glu Wing Thr Arg Leu Lys Arg Tbr Wing Arg Arg Are¡ Tyr Thr 100 105 110
Arg-Are¡ Lys Asn Arg Ile eye Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125
Met Ala Lys val Asp As p Ser Phe Pbe His Arg Leu Glu Glu Ser Phe 130 135 140
Leu Val Glu Glu Asp Lys Lys Bis Glu Are¡ Bis Pro Ile Phe Gly Asn 145 150 155 160
Ile Val Asp Glu Val Wing Tyr His Glu Lys Tyr Pro Tbr Ile Tyr Bis 165 170 175
Leu Arg Lys Lys Leu val Asp Ser Thr Asp Lys Wing Asp Leu Arg Leu
180 185 190
Ile Tyr Leu Ala Leu Aloa His Met Ile Lys Phe Are¡ Gly His Phe Leu 195 200 205
Ue .lu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220
Ila Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Uo 225 230 235 240
Asn Ala Ser Gly Val Asp Ala Lys Ala Uo Leu Ser ~ a Are¡ Leu Ser 245 250 255
Lys Sar Arq Arq Leu Glu Asn Leu Ile Wing aln Leu Pro Gly Glu Lys
260 265
Lya Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285
Pro Asn Phe Lys Ser Asn Phe As? Leu ~ a Glu Asp Ala Lys Leu Gln
2.0 2.5 300
Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Wing Gln 30S 310 315 320
Ile Gly Asp Gln Tyr Wing Asp Leu Phe Leu Wing Wing Lys Asn Leu Ser 325 330 335
ASp Wing Ile Leu Leu Ser Asp Ile Leu Ar9 Val Asn Thr Giu Ile Thr 340 345 350
Lys Ala Pro Leu Ser Ala Ser Met Ila Lys Arq Tyr Asp Glu 8is His 355 360 365
Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arq Gln Gln Leu Pro Glu 310 315 380
Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 3.0 395 400
Tyr Ile Asp Gly Gly Ala Ser Gln GIu Glu Phe Tyr Lys Phe Ile Lya 405 410 415
Pro Ile Leu Glu Lys Mat Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430
Asn Arq Glu Asp Leu Leu Arg Lys Gln Arq Thr Phe Asp Asn Gly Ser 435 440 445
Ile Pro His Gln Ile Bis Leu Gly GIu Leu His Wing Ile Leu Arq ArO '450 455 460
Gln GIu Asp Phe Tyr Pro Leu Lys Asp Asn Arg Glu Lys Ile Glu
., 0 Phe
465 415 480
Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr val Gly Pro Leu Ala Arg 485 495
4'.
Gly Asn Ser Arg Phe Wing Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510
Thr Pro Trp Asn Phe Glu Glu val Val Asp Lys Gly Ala Ser Ala Gln 515 5.20 525
Ser Phe Ile Glu Arq Met Thr Aan Phe Asp Lye Asn Leu Pro Asn Glu 530 535 540
Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560
Asn Glu Leu Thr Lys Val Lys Tyr Val Thr G1u Gly Met Arq Lys Pro 565 570 575
Wing Phe Leu Ser Gly Glu Gln Lys Lys Wing Ile Val Asp Leu Leu Phe 580 585 590
Lys Thr Asn Arq Lys val Thr Val Lye Gln Leu Lys Glu Asp Tyr Phe 595 600 605
Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620
ArO Phe Asn Ala Ser Leu Gly Thr Tyr His ABp Leu Leu Lye Ile Ile 625 630 635 640
Lys Asp Lys Asp Phe Leu ASp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655
Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arq Glu Met Ile Glu 660 665 670
Glu Arq Leu Lys Thr Tyr ~ a Hia Leu Phe Asp Asp Lys Val Met Ly & 675 680 685
G1n Leu Lys Arq Arq Arq Tyr Thr Gly Trp Gly Arq Leu Ser Arq Lys 690 695 700
Leu Ile Asn G1y Ile Arq Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720
Phe Leu Lys Ser Asp Gly Phe A1a Asn Arch Asn Phe Met Gln Leu! Le 725 730 735
His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys A1a Gln Val 740 745 750
Ser Gly Gln G1y Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 160 765
Be Pro Wing Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val A $ p 770 775 180
Glu Lau val Lya Val Met Gly Arq Bis Lys Pro Glu Aan Ile Val Ile 185 790 795 800
Glu Met Wing Arch Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 S10 815
Arq GIu Arq Met Lys Arq I ~ e Glu Glu Gly Ile Lys GIu Leu Gly Ser 820 825 830
Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Lau Gln Asn Glu 835 840 845
Lys Lau Tyr Leu Tyr Tyr Leu GIn Asn Gly Arg Aep Met Tyr Val Asp 850 855 860
Gln Glu Lau Asp Ile Asn Arq Lau Ser Asp Tyr Asp Val Asp His Ile 865 870 875 aao
Val Pro Gln Ser Phe Lau Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895
thr Arq Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910
Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arq Gln Lau Lau Asn Ala 915 920 925
Lys Lau Ile Thr Gln ArO Lys Phe Asp Asn Leu Thr Lys Ala Glu Arq 930 935 940
Gly Gly lAu Ser Glu Lau Asp Lys Wing Gly Phe 11e Lys Arg Gln Leu 945 950 955 960
Val Glu Thr Arq GIn Ile Thr Lys Bis Val Wing Gln Ile Lau Asp Ser 965 970 975
Arq Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Lau Ile Arg Gla Val 980 985 990
Lys Val Ile Thr Leu Lys Ser Lys Lau Val Ser Asp Phe Arq Lys Asp
995 1000 1005
Phe Gln Phe Tyr Lys Val Arq Glu Ile Asn Asn Tyr Bis His Ala 1010 1015 1020
His Asp Ala Tyr. Leu Asn Al. Val Val G1 and Thr Ala Leu I1e Lys 1025 1030 1035
Lys Tyr Pro Lya Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 lOSO
Val Tyr Asp Val Arq Lys Met Ile Ala Lys Ser Glu Gln Glu Ila 1055 1060 1065
Gly Lya Al.a Thr Al.a Lys Tyr Phe Phe Tyr Ser Asn Il.e Met. Asn 1070 1075 10aO
Phe Phe Lys Thr G1u Il.e Thr Leu Ala. Asn Gly Glu Ile Arq Lys
1085 1090 1095 Arq Pro Leu Ila Glu Thr Asn Gly Glu Thr Gly Glu Ila Val Trp 1100 1105 1110 Asp Lya Gl.y Arq Asp Phe Ala Thr Val Arq Lys Val Leu Ser Het 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys G1u Ser Ile Leu Pro Lys Arq Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arq Lys Lys Asp Trp Asp Pro Lys Lys Tyr G1y G1y Phe 1160 1165 1170 Asp Be Pro Thr Val Ala Tyr Ser Val Leu Val. Val Al.a Lys Val. 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val. Lys Glu Leu Leu
1190 1195 1200 Gly Ile Tbr r1e Met Glu Arq Ser Ser Phe Glu Ly. AI! In Pro rla
120.5 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Gl.u Val Lys Lys Asp Leu 1220 1225 1230 11e Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Gl.u Asn Gly
1235 1240 1245 Arg Lys Arq Met Leu Wing Being Wing Gly Glu Leu Gln Lys Gly Asn
G1u Leu 1265
Be His 1280
Lys G1n 1295
ne G1u 1310
A1a Asn 1325
Lys Pro 1340
Leu Thr 1355
Thr n. 1310
Ah Thr 1385
n. Asp 1400
Lys LyS 1415
<210> 50
<211 > 2012
<212> DNA
Ala Leu Pro Ser
Tyr Glu Lys Leu
Leu Phe Val Glu
Gln Ile Ser Glu
Leu Asp Lys Val
Ile Arq Glu Gln
Asn Leu Gly Ala
Asp ArQ Lys Arq
Leu 11a His Gln
Leu Ser Gln Leu
Wing G1 and Gln Wing.
Lys 1210
Lys 1285
G1n 1300
Phe 1315
Leu 1330
A1. 1345
Pro 1360
Tyr 1375
Be 1390
G1 and 1405
Lys 1420
1260 Tyr Val Asn Phe Leu 1215 Gly Ser Pro Glu Asp 12 90 His Lys His Tyr Leu 1305 Ser Lys ArQ Val n. 1320 Se r Wing Tyr Asn Lys 1335 Glu Asn Ile Ile His 1350 Wing Wing Phe Lys Tyr 1365 Thr Ser Thr Lys G1u 1380 Ile Thr Gly Leu Tyr 1395 Gly Asp Lys Arq Pro 1410 Lys Lys Lys
Tyr Leu Ala
A $ n Glu Gln
Aep Glu Ile
Leu Ala Asp
His Arq. Asp
Leu Phe Thr
Phe Asp Thr
Val Leu Asp
Glu Thr Arq
Wing Thr wing
<213> Artificial Sequence 10 <220>
<dl><dt>< </dt><dd>221 > source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description l: Synthetic polynucleotide " </dd></dl>
<400> 50
qaatqctqcc ctcaqacccg cttcctccct gtccttqtct qtccaaoqaq aatqaqqtct 60 cactqqtqqa tttcqqacta ccetqaqqaq ctqqcacetq aqqqaeaaqq ecccccaoct 120aeqaqaqaqaqaqq
tcccctgaaq gaqaccacac aqtqtqtgaq qttg'gagtct ctagcagcgq qttctqtqcc 240 cceagqgata gtetggetgt eeagqeactg etct, tgatat aaaeaeeaee teetagttat 300 qaaaccatqc ccattctqcc tctctgtatg gaaa, agagca tgqqqctgqc ccgtgqqgtg 360 gtgtecactt taqqcectgt GGQ & qatcat qooa, acccac qcagtOQgtc ataq9ctctc 420 etcacatcca teatttacta ctctqtgaa ( '480 jJ aagc'gattat gatctctcct etagaaactc gtagilgtcec atgtctgccq qettccaqaq cctg'cactcc tccaccttqq cttqqetttq 540 etgggQ'ctaq aqqagctagg atgcacagca qctc: tqtgac cctttgtttg aqaqqaacaq 600 qaaaaccaec cttctctctq qcccactqtq tcctcttcct qccetgccat ccccttetgt 660 eccatqgqaq qaatgttaga cagctqqtca gagg'qqaecc cqqcctgqqq cccetaaccc 720 tatgtaqeet eagtcttcee ateaqgctct EAGC: teagcc tqagtgttga ggccecaqtg 7.0 qctgctctgg qqgcctcctq agtttctcat ctqt, qccect ccctccctqq cccaqqtgaa .40 caqaaccqqa qqtqt_tc qqacaaaqta CAAA, cgqcaq aaqctqgagq aggaaqqqcc 900 tgagtccgag agggctccca cagaagaaga tcac.atcaac cqgtggcqca ttqccacgaa 960 qeaggccaat qgqqaqqaca tcqatgtcac ctcc: aatgac aaqcttqcta qcqqtqqgca 1020 accaeaaacc cacgagggca gagtqctqct tgctgctqqc cagqcccctq cqtqqqccca 1080 aqctqgactc tqqccactcc ctqqccaggc TTTG: qqqaqq tqgccccaca cctqqagtca 1140 gqqcttqaaq cccgqqqccq ccattgacaq aqqgracaaqc aatqqgctqg ctqaqqcctg 1200 qgaccacttq qccttctcct cgO'aqagcct qcct, gcctqg gcqqqcccqc ccgccaceqc 1260 aqectcccaq ctgctctccq tqtctccaat CTCC: cttttq ttttgatqca tttctqtttt 1320 ccaqqcacca aatttatttt ctqtagttta qtga, tcccca qtgtccccct tccctatgqg 1380 aataataaaa gtctetctct taatqacacq qgca, tccaqc tccagcccca qaqcctqggg 1440 t.gqtaqattc cqqctct.qaq qqccaqt.qqq O'octqqtaqa qcaaacqcqt tcaO'qqcctq 1500 ggaqcctqgg qtgqqqtact gqtqgagqgg qtca.aq9Qta attcattaac tcctctcttt 1560 tqtt.qqggqa ccctqc¡tctc tacetccaqc tcca.caqcaq qaqaaacaqq ctagacatag 1620 tcctqt qqaagggcca: ATCT tgaqqgaqqa tctttcttaa caqq'cccaqq eqtattqaqa 1680 qqtqggaatc agqcccaqgt aqttcaatgg gaga, gqgaga qtqcttccct ctgcctagaq 1740 actctgqtqq cttctccaqt tqaqqagaa. ccaq'aqqaaa qqggaqqatt qqggtctggq 1800 Qqag9gaaca ccattcacaa agqctqacqq TTCC, aqteeg aaqtcqtqqq cccaccagqa 1860 tgctcacctq tccttgqaqa accqetqgqc aqqttgagac tgcagaqaca qqgcttaaqq 1920 ctgagcctgc aaccaqtccc caqt.qactca 99qc, ctcctc aqcccaaqaa aga9caaeqt 1980 qccaqqqccc qctqaqctct tqtqttcacc TQ 2012
<210> S1
<211> 1153
<212> PRT
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = MD artificial sequence description: synthetic polypeptide M </dd></dl>
<400> S1
Met Lys Arq Pro Wing Thr Wing Lys Lys Wing Gly Gln Wing Lys Lys Lys 1 5 10 15
Lys Ser Aap Leu Val Leu Gly Leu Asp 11e Gly Ile Gly Ser Val Gly 20 25 30
V ~ l Gly Ile Leu Asn Lys Val Thr Gly Glu Ile Ile Bis Lys Asn Ser 35 40 45
Arg Ile Phe Pro Wing Gln Wing Glu Asn Asn Leu Val Arq Arq orhr 50 55 ~
Asn Arg Gln Gly Arq Arg Leu Wing Arg Arg Lys Lys Bis Arg Arq Val 65 70 75 80
Arg Leu Asn Arg Leu Phe Glu Glu Ser Gly Leu 11. Thr Asp Phe orhr S5 90 n
Lys Ile Ser Ile Asn Leu Asn Pro Tyr Gln Leu Arq Val Lys Gly Leu 100 105 110
Thr Asp Glu Leu Ser Asn Glu Glu Leu Phe Ile Ala Leu Lys Asn Met 115 120 125
Val Lys Uis Arq Gly Ile Ser Tyr Leu Asp Asp Ala Ser Asp Asp Gly 130 135 140
Aan Ser Ser Val Gly Asp Tyr Wing Gln Ile Val Lys Glu A3n Ser Lys 145 150 155 160
Gln Leu Glu Thr Lys Thr Pro Gly Gln Ile Gln Leu Glu Arq Tyr Gln 165 1 "70 175
Thr Tyr Gly Gln Leu Arq Gly Asp Phe Thr Val Glu Lys Asp Gly Lys 180 185 190
Ly8 His Arq Leu Ile Asn Val phe Pro Tbr Ser Ala Tyr Arq Ser Glu 195 200 205
Ile Leu Thr Arq Leu Gly Lys Gln Lye Thr Thr Be Be Be Asn Ly. 465 470 475 4S0
Thr Lys Tyr Ile Asp Glu Ly. Leu Leu Thr Glu Glu Ile Tyr Asn Pro 485 490 495
Val Val Wing Lys Ser Val Arq Gln Wing 11. Lys Ile Val Asn Wing Wing 500 505 510
Ile Lys Glu Tyr Gly Asp Phe Asp Asn Ile Val 118 Glu Met Wing Arch 515 520 525
Glu Thr Asn Glu ASp Asp Glu Lys Lys Ile Gln Lys Ile Gln Lye 530 535 540
Wing Asn Lys Asp Glu Lys Asp Wing Wing Met Leu Lys Wing Wing Asn Gln 545 550 555 560
Tyr Asn Gly Lys Wing Glu Leu Pro Bis Ser Val Phe His Gly His Lys 565 570 575
Gln Leu Ala Thr Lye Ile Arq Leu Trp Kis Gln GIn Gly GIu Arq Cys
580 585 590
Leu Tyr Thr Gly Lys Thr Ile Ser Ile His Asp Leu Ile Asn Asn Ser 595 600 605
Asn Gln Phe Glu Val Asp His lle Leu Pro Le: u Ser Ile Tbr Phe Asp 610 615 620
Aap Ser Leu Wing Asn Lys Val Leu Val Tyr Wing Thr Wing Asn GIn GIu 625 630 635 640
Lys Gly Gln Arq Thr Pro Tyr Gln Ala Leu Asp Ser Mat Asp Asp Ala 645 650 655
Trp Ser Phe Arq Glu Lau Ly8 Ala Phe Val Arq Glu Ser Lys Thr Leu 660 665 670
Ser Asn Lys Lys Lys Glu Tyr Leu Leu Thr Glu Glu Asp Ile Ser Lys 675 680 685
Phe Asp Val Arq Lys Lye Phe Ile Glu Arq Asn Leu Val Asp Thr Arq 690 695 700
Tyr Ala Ser Arq Val Val Leu Asn Ala Leu Gln Glu His Phe Arq Ala
70S 710 715
His Lys Ile Asp rhr Lys val Ser val Val Arq G1y Gln Phe Thr Ser
725 '730 135
Gln Leu Arq Arq His Trp Gly Ile Glu: [. Ys Thr Arq ASp Thr Tyr His 740 745 750
His His Ala Val ASp Ala Leu Ile l1e JUa Ala Ser Ser Gln Leu Asn 755 760 765
Leu Trp Lys Lys Gln Lys Asn Thr Leu 1 ~ Ser Tyr Ser Glu Asp GIn 710 715 780
Leu Leu Asp I1e Glu Thr G1y Glu Leu: Ue Ser Asp Asp Glu Tyr Lys 785 790 795 900
Glu Ser Val Phe Lys Ala Pro Tyr GIn l ~ i8 Phe Val Asp Thr Leu Lys 805 ~ lO 815
Ser Lys Glu Phe Glu Asp Ser Ile Leu l ~ he Ser Tyr Gln Val Asp Ser 820 825 830
Lys Phe Asn Arq Lys l1e Ser Asp Ala 'rhr I1e Tyr Ala Thr Arg Gln 835 840 845
Wing Lys Val Gly Lys Asp Lys ~ a Asp t ~ lu Thr Tyr Val Leu Gly Lys 850 855 860
Ile Lys Asp I1e Tyr Thr Gln Asp G1y 'ryr Asp Ala Phe Met Lys Ila 865 870 875 880
Tyr LyS Lys ASp Lys Ser Lys Phe Leu l:> Iet Tyr Arq His Asp Pro GIn
885 :890 895
Thr Phe Glu Lys Val Ile Glu Pro Ile: Leu Glu Asn Tyr Pro Asn Lys 900 905 910
Gln Ile Asn Glu Lys Gly Lys Glu Val IEtro Cys Asn Pro Phe Leu Lys 915 920 925
Tyr Lys Glu Glu Bis Gly Tyr Ile Arch: Lys ryr Ser Lys Lys Gly Asn 930 935 940
Gly Pro Glu Ile Lys Ser Leu Lys Tyr 'ryr ASp Ser Lys Leu Gly Asn 945 950 955 960
Hi8 Ile Asp Ile Thr Pro Lye Aep Ser Asn Asn Lys Val Val Leu Gln 965 970 975
Ser Val Ser Pro Trp Arq ~ a Asp Val Tyr Phe As n Lys Thr Thr Gly 980 985 990
Lys Tyr Glu Ile Leu Gly Leu Lys Tyr Ala Asp Leu Gln Phe Glu Lya 995 1000 1005
Gly Thr Gly Th r Tyr Lys Ile Ser Gln Glu Lys Tyr Asn Asp Ile 1010 1015 1020
Lys Lys Lys Glu Gly Val Asp Se r Asp Ser Glu Phe Lys Phe Thr 1025 1030 1035
I.eu Tyr Lys Asn Asp Leu Leu Leu Val Lys Asp Thr Glu Thr Lys 1040 1045 1050
G1u G1n G1n Leu Phe Arq Phe Leu Ser Arq Thr Met Pro Lys Gln 1055 1060 1065
Lys His Tyr Val Glu Leu Lys Pro Tyr Asp LyS Gln Ly & Phe Glu 1070 1075 1080
Gly Gly Glu Al.a Leu Ile Lys Val Leu Gly Asn Val Ala Asn Ser 1085 1090 1095
Gly Gln Cys Lys Lys Gly Leu Gly Lys Ser As n Ile Ser Ile Tyr 1100 1105 1110
Lys Val Argo Thr Asp va l Leu Gly Asn Gln His Ile Ile Lys Asn 1115 1120 1125
Glu Gly Asp Lys Pro Lys Leu Asp Phe Lys Arq Pro Ala Wing Thr 1130 1135 1140
Lys Lys Wing Gly Gln Wing Lys Lys Lys Lys 1145 1150
<210> 52
<211> 340
<212> DNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Synthetic Polynucleotide Artificial Sequence Description ~ </dd></dl>
<dl><dt><400> 52 qaqqqcctat </dt><dd>ttcccatqat t ccttcatat ttqcatatac qatacaaqqc t.qttaqa9aq 60 </dd></dl>
<dl><dt>ataat.tggaa ttaatttgac tgtaaaca.ca aaqatattaq tacaaa.atac qtqacqtaqa </dt><dd> 120 </dd></dl>
<dl><dt>aaqtaataat ttcttqqqta qtttgcaqtt ttaaaattat qttttaaaat qqactatcat </dt><dd> 180 </dd></dl>
<dl><dt>atqcttaccq taacttqaaa qtatttcqat ttcttqgctt tat atatct t </dt><dd>qtgqaaaqga 2.0 </dd></dl>
<dl><dt>cqaaacaccq ttacttaaat cttqcaqaaq ctacaaaqat aaqqcttcat qccqaaatca </dt><dd> 300 </dd></dl>
<dl><dt>acaccctgtc attttatqqc aqgqtgtttt cgttatttaa </dt><dd> 340 </dd></dl>
<dl><dt>5 </dt><dd><210> 53 <211> 360 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = "Artificial Sequence Description: Synthetic polynucleotide" </dt><dd /></dl>
<dl><dt>10 </dt><dd><220> <221> modified base <222> (288) .. (317) <223> a, c, t, g, Unknown or other </dd></dl>
<dl><dt><400> 53 gafJ0'9cctat </dt><dd>ttcccatqat tccttcatat ttqcatatac qatacaaqge tqttaqaqaq 60 </dd></dl>
<dl><dt>a.ta.attqqaa ttaatttqac tqtaaacaca aaqatattaq tacaaaatae qtqacqt.aga </dt><dd> 120 </dd></dl>
<dl><dt>aaqtaataat ttctt.gqqta qtttqcaqtt. t.taaaattat qttttaaaat qgactatcat</dt><dd> 180 </dd></dl>
<dl><dt>atqcttaccg taacttgaaa qtatttcgat ttcttggctt tatatatctt qtgqaaagqa </dt><dd> 2.0 </dd></dl>
<dl><dt>cqaaacaccq qqttttaqaq ctatqctgtt ttgaa.tqgtc </dt><dd>ccaaaacnnn nnnnnnnnnn 300 </dd></dl>
<dl><dt>15 </dt><dd>nnnnnnnnnn nnnnnnnqtt ttaqaqctat getqttttqa atqqt.cccaa aacttttttt 360 </dd></dl>
<dl><dt><210> 54 <211> 318 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>20 </dt><dd><220> <221> source <223> Inota = "Artificial Sequence Description: Synthetic polynucleotide · </dd></dl>
<dl><dt>25 </dt><dd><220> <221> modified base <222> (250) .. (269) <223> a, c, t, g, Unknown or other </dd></dl>
<dl><dt><400> 54 gaqqqcctat ttcceatqat tccttcatat ttqcatatac oatacaaq; c tgttagagaq </dt><dd> 60 </dd></dl>
<dl><dt>ataattqqaa ttaatttqac tqtaaacaca aaqatattaq t acaaaatac </dt><dd>qtqacqtaoa 120 </dd></dl>
<dl><dt>129 </dt><dd /></dl>
<dl><dt>aaqtaataat ttcttgqqta qtttqcaqtt </dt><dd>ttaaaattat qttttaaaat qqactatcat 180 </dd></dl>
<dl><dt>atqcttaccq taacttqaaa qtatttcqat ttcttqqctt tatatatctt otqqaaaqqa </dt><dd> 240 </dd></dl>
<dl><dt>c qaaacaccn </dt><dd>nnnnnnnnnn nnnnnnnnoq ttttaqaqc t agaaatagea agtt aaaata 300 </dd></dl>
<dl><dt>aqqctaqtce qttttttt </dt><dd> 318 </dd></dl>
<dl><dt>5 </dt><dd><2 10> 55 <211> 325 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> were <223> Inola = -Artificial Sequence Description: Synthetic polynucleotide ~ </dt><dd /></dl>
<dl><dt>10 </dt><dd><220> <221> modified base <222> (250) .. (269) <223> a, c, 1, g, Unknown or olro </dd></dl>
<dl><dt><400> 55 qaqqgeetat </dt><dd>ttcceatqat tccttcatat ttqeatatac qatacaaqqe tqttaqaqaq 60 </dd></dl>
<dl><dt>ataattqqaa </dt><dd>ttaatttqae tgtaaacaca aaqatattaq tacaaa ata e qtqacqtag & 120 </dd></dl>
<dl><dt>aaqtaataat </dt><dd>ttcttqqqta qtttqcaqtt ttaaaattat qttttaaaat qqactatcat 180 </dd></dl>
<dl><dt>at.qctt.ac cq </dt><dd>taacttqaaa qtatttcgat ttcttqqctt tatatatctt qt9qaaaqqa 2.0 </dd></dl>
<dl><dt>cqaaacaecn </dt><dd>nnno.nnnnnn nnnnnnnnnq t.ttta qagct agaaataqca aqttaaaat.a 300 </dd></dl>
<dl><dt>aqqctaqtcc </dt><dd>gttatcattt ttttt 325 </dd></dl>
<dl><dt>15 </dt><dd><210> 56 <211> 337 <212> AON <213> Artificial Sequence </dd></dl>
<dl><dt>20 </dt><dd><220> <221> fuenle <223> Inota = -Artificial Sequence Description: Polynucleic solid ~ </dd></dl>
<dl><dt>25 </dt><dd><220> <221> modified base <222> (250) .. (269) <223> a, c, l. g, Unknown or pot</dd></dl>
<dl><dt><400> 56 </dt><dd /></dl>
qaqqqcctat ttc: c: catqat tccttcatat ttqcatatae oatacaagqc tgttaqagaq 60 at && ttogaa tta & tqta tttgac & ACACA aagatattag tacaaaatac qtqacqtaqa 120 aaqtaataat ttcttqgqta qtttgcagtt ttaaaattat qttttaaaat ggactatcat. 1.0 atqcttaccq taacttqaaa qtatttcgat ttctt9gctt tatatatc: tt gtggaaaqqa 240 nnnnnnnnnn cqaaacaccn nnnnnnnnnq ttttagaqct aqaaataqca agttaaaata 300 aggctaqtcc qttatcaact tqaaaaaqt: q ttttttt 337
<210> 57 <21 1> 352 5 <212> DNA
<213> Artificial Sequence
<220>
<221> source
<223> Inota = "Artificial Sequence Description: Synthetic polynucleotide ·
10 <220>
<221> modified base
< 222> (250) .. (269)
<223> a, c, t, g, Unknown or other
<400> 57
gagqgcctat ttcccatqat tccttc8tat TTGE & t & tao gatac & aq9c tgttaqaqaq 60 ataattqqa8 tt8attt98e tqtaaacaca aaqatattaq tacaa.aatae qt.qaeqtaga 120 aaotaataat ttcttqqqta qtttgcagtt ttaaaattat qttttaaaat qgactatcat '.0 & t90ttaecg taaettgaaa gtatttcqat ttcttgqctt t8tat & tctt qtqqaaaqqa 240 cqaaacaccn nnnnnnnnnn nnnnnnnnnq ttttaqagct aqaaa.taqca aqttaaaata 300 aggctaqtce qttatcaact tgaaaaaqtq qcac: cqagtc : qqtgcttttt tt 352 15
<210> 58
< 21 1> 5101
<212> AON
<213> Artificial Sequence
20 <220>
<221> source
<223> Inota = "Description of Artificial Sequence: Synthetic polynucleotide · <400> 58
c: qttac: ataa cttacgqtaa atggcccgcc tggctgaccg cccaacqacc cccgcccatt 60 gacqtc: aata atgacgtatg ttcccatagt aacgccaata qqgactttcc attqacqtca 120 atqqqtqgag tatttacqgt 180 aaqtac: GCCC cctattqacq tcaatgacgq taaatqqccc qcctgqcatt atqc: c: caqta 2.0 catqacctta tqqqactttc ctacttqqca qtacatctac qtattaqtca tcqctattac 300 catqgtcgag qtgaqcccca cgttctgctt cac: tc: tcccc atctcccccc cctccccacc 360 cccaatttt.g tat.t.tattt.a t.tttt.taatt attttgtgca gcgatqgggq c9Q9g99999. 20 ggqq.qggcgc qcqccaqqcq q.qgcqqqqcq qgqcoaqQ09 'cgqQgcqqQq cQaqqcqqaq .80
aqgtqcggcg gcagccaatc & gaqcggcgc qctccqaa.aq tttcctttta tqgcqaqqcq 5.0 qCCiJqcqgcc¡q cc¡c¡ccctata aaaaqc: qaaq cgcc¡cqqcqq CiJc ctq ccq aka: gc: c: qcc: q ccteqcqccq cc :: c: gc: ctc cccgq :: tqactqa 660 ccqcqttact cccacagqtg agcgqqcqqq aeqqcccttc tcctccqqqc tgtaattagc 720 tgagcaagag gtaaqgqttt aaqgqatqqt tgqttqqt.gg qqtattaatq tttaattacc 7'0 tgqaqcacct qcctqaaatc actttttttc agqttt; Jqacc qqtgccacca tqgactataa 840 ggaccacqac qqaqactaca agqat catqa tattg; "aaagacgatq acqataaqat TTAC 900 ggccccaaaq aaqaaqeqqa aqqtcgqtat ccacql; Jaqtc ccagcaqccq acaaqaaqta 960 caqcatcqqc ctgqacatcq gcaccaactc tqtqqqctqq gccqtqatca ccqacqaqta 1020 caaqgtqc: c: c agcaaqaaat tcaaqqtqct qCiJqc: a; CCAEC qcatcaaqaa gaccqgcaea qaacc 1080 :: tqatc qqaqccctqc tgttcgacaq cgqcqaaaca CiJccqaCiJqcca cccqqctqaa 1140 qaqaaccqcc aqaaqaaqat acaccagacq qaaqaaccqq atctqctatc tqcaaqagat 1200 qagatqqcca cttcaqcaac agqtqqacqa c: AQC: ttcttc cacaqactqq aagaqtcctt 1260 eetqqtqqaa ageacgagcq gaqqataaga qcaceceate ttcqqcaaca tcqtqqacga 1320 qqtqqcetac cacqaqaaqt accccaccat ctaccacctq agaaaqaaac tqqtqqaeaq 1380 caccqacaaq gccgacctgc qqctgatcta tctqgccctq qcceaeatqa tc: aaqttccg 1440 . Qqqccacttc ctgatcgagg qcqacctqaa ccccq "caac aqcqacqtgg acaaqctgtt 1500 catccagctq qtgcaqacct acaaccaqct qttcgaggaa aaccccatca acO'ccaqcgq 1560 cqtc¡gacgcc aaqqccatcc tqtctqccag actga '; Jcaaq agcagacqqc tggaaaatct 1620 qatcgcccag ctgcccgqcq aqaaqaaqaa tqqcc't: QTTC qqcaacctqa ttqccctqaq 1680 cctqqqcctq acccccaact teaagagcaa cttcqacctg gccqaqgatq ccaaactgca 1740 c Ctgagcaaq qacacctacq acqacgacct qgacaacctq ctq9cccaga tcggcgacca 1800 qtacqccqac ctgtttct99 ccqccaaqaa cctqtccqac 9ccatcctqc tqagcqacat 1960 ec.tgaqaqtg aacaccqaqa tca.ccaagqc ccccctqaqc qc.ctctatqa tcaagagata 1920 cgacqaqcac caccaqgacc tgaccctqct qaaagctctc Qt.qcqqcaqc agctqcctqa 1980 qaaqtacaaa qagattttct tcgaccagag caagaacqgc tacqccc¡qct acattqacqq 2040 cqCiJac¡ccagc caqqaaQaqt tctacaaqtt catca..aqcec atcctgqaaa aqatqgacqq 2100 ctqctcqtga caccqaqqaa & qctgaacaq aqaqgacctq ctqcqgaagc aqcggacctt 2160 cgacaacqqc aqcatccccc accaqatcca cctqqqagaq ctgcacgcca ttctqcqgcq 2220 qcaqgaaqat ttttacCCat tcctQaaqga caaccqqqaa aagatcqaqa aqatcctqac 2280 cttccqC: Atc: ccctActacq tqq9ccctct qqcca.qqqqa Aaeaqcaqat tcqcctqqat 2340 aqcqagqaaa gaccaqaaaq ccatcaeccc ctgqaacttc qaIJIJaaQtIJO tlJlJ & caalJlJg 2400 cgcttccqce calJaqcttca tcqagcgqat qaceaaette qataaqaacc tqceeaacga 2460 cccaaqcaca c¡aac¡qtc¡ctq qcctqctqta. cqa.9tacttc accCJt.qtata acc¡aqetqac 2520 caaaqtqaaa tacqtqaceq agqqaatgaq aaac¡cceqcc ttcctqac¡cc¡ gcqaliJcaliJaa 2580 aaaqgccat c qtqqacct qe tgttcaaqac gtqaccqtqa caaecqqaaa aqcagctqaa 2640 agagqactae ttcaagaaaa tcgagtqctt cqact: ccqtg qaaatetccq qeqtggaaga 2700 teggtteaac gccteectqq qeaeataeca cgatc: tqctg aaaattatca aqgacaaqqa 2760 cttcctqqac aat9a99aaa acgaqqacat tctq9aaqat atcgtqct9a ccctqacact 2820 qtttgaqgilc ilqaqaqatqa tcqaggaacg qctgllaaacc tatqcccac: c tqttcqaega 2880 eaaagtgatg aagcaqctg & aqcqqcggag ataCi: lccggc tqqqgcagge tqaqccggaa 2940 gctqatc: aac qqcatecggq acaagcaqtc cggCi: laqaca atcctggatt tcctgaagtc 3000 cqaeq9cttc ctcca cagaaagccc aggtgtccqq ccaqgqcgat agcctgcaC: 1iJ aqcaeattqe 3120 caatctqqcc ggcagccccg ccattaagaa qqge.ateetg cagacaqtga aggtqqtqqa 3180 eqac¡ctcqtg aaaqtqatqq qccqgcaca & qcccuagaac atcqtqatcg aaatggccaq 3240 agaqaaccaq EACEA: cc: AQA 3300 agqqacagaa gagagaatqa gaacuqccqc aqcqqatcqa aqaqqgcatc aaagagetgg gcagccaqat cc: tguaagaa cacccc: aaaaeaeeca gtqq 3360 gctgcagaac gagugctqt acctqt.aeta cetq <: agaat qqqc9'ggata tgtaegtgga 3420 ccaggaaetg gacatcaacc qqctqtcc9a ctaC ~ Jatqtq qaccatatcq tgcctcagag 3480 ctttctgaaq gacgactcca tegacaacaa 9qt.q <; CTGA aqallqcqac: a aqaaecq999 3540 c & aqagcqac aacqtgccct ccqaaqagqt cqtqi: laqaag atqaagaact actgqcliJqC & qccaaqctga gctqctgaae 3600 ttacccagaq aaaqt: tcqac & a.tctga.cca a.qqcegagaq 3660 & gqcqg: cctq agcqaactgg ataaqqccqq cttcntcaaq aqac: agctqg tqqaaacccq 3720 qcaqatc .ea aagcacqtqg c: acaqatcet qgacteccqq atqaaeaeta aqtaeqacq & 3780 gaatgaeaaq ctgatccqgg aaqtgaaaqt qatcnccctq aaqtccaage tqqt.qtc: cga 3840 tttce9gaa. accacgeeca 3900 cqacqcctac ctqaacqccq tcqtqqgaac cgecctqate aaaaagtacc ct.aagctgga 3960 aagcqaqttc qtqtacqqcg actacaagqt. qtacqacqt.g eggaagatga tegecaagag 4020 C949c499aa atc9Qce.aqq ctaccqccaa qte.c1; tcttc tacaqcaaca tcatqaactt 4080 t.ttcaaqacc qaqattaccc tqqccaacqq cgaqlltccqq aagcggcctc tgatcg & QAC 4140 taag aaaeqgcqaa accq9qQaqa tcqtgtqqqa "l; lqccq9 gattttqcea ccqtqcqqaa 4200 agtgctgagc atqccccaag tqaatatcqt gaaallagacc qaqgtqcaqa eaggcqgctt 4260
caqeaaagag tctatcctqc ccaagaqgaa ctqatcqcca caqcqataaq qaaaqaaqga 4320 ctgqqaccct aaqaaqtacq qcqgcttcga cagccccacc qtqqcctatt ctqtgctgqt 4380 qqtoqccaaa 9tqqaaaaqq qcaaqtccaa qaaactqaaq aqtqtqaaaq aQctqctq9Q 4440 qatcaccatc atc; rqaaagaa gce-qcttcga qaagaatccc atcqactttc tggaagccaa 4500 qe-aqtgaaaa gqgctacaaa aggacctgat catcaaqctg CCTEs-aqtact ccctgttcga 4560 gctqgae-ae-c 9qcC9qaaga qaatgctqqc ctctqccqqc gaactqcaqa aqgqaaacga 4620 actqgccctg ccctccaaat atqtgaactt cctqtacctq gccaqccact atqaqaaqct 4680 qaagqqctcc atqagcaqaa cccqagqata acagctqttt qt: ggaacagc acaagcacta 41 40 cetqqaeqaq atcatcqaqc agatcaqcqa qttctccaaq aqaqtqatcc tgqccqacqc 4800 taatctqqac ••• qtgctqt ccqcctacaa caaqcac cgg tcaqaqagca gataaqccca 4860 qqccqaqaat atcatccacc tgtttaccct qaccaatctq qoaqcccctg ccgccttcaa 4920 qtactttgac accaccatcq accgqaagag qtacaccagc accaaaqaqq tqctggacqc 4980 caccaqagca caccctqatc tcaccgqcct qtacqaqaca cqqatcqacc tqtctcaqct 5040 qqgaqgcgac tttctttttc ttaqcttgac caqctttctt aqtaqcaqca qqacqcttta 5100 • 5101
<210> 59
<21 1> 137 5 <212> DNA
<213> Artificial Sequence
<220>
<221> source
<223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide "
10 <220>
<221> modified base
< 222> (1 )..(20)
<223> a, c, 1, g, Unknown or other
<400> 59
nnnnnnnnnn nnnnnnnnnn qtttttqtac tctcaaqatt taqaaataaa tcttgcagaa 60 9cta.caaaqa taaggcttca tgccqaaatc aacaccctqt cattttatqq caq9qtqt.tt 120 att
<210> 60
<21 1> 123 <; 212> DNA
<213> Artificial Sequence
20 <220>
<221> source
<223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide "
<dl><dt><220> <221> modified base <222> (1) .. (20) <223> a, c, 1, g, Unknown or other </dt><dd /></dl>
<dl><dt>5 </dt><dd><400> 60 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcaqaaat qcaqaaqcta caaaqataaq 60 </dd></dl>
<dl><dt>gcttcatqcc qaaatcaaca ccctgtcatt ttatqqcagq qtqttttcqt tatttaattt </dt><dd> 120 </dd></dl>
<dl><dt>ttt </dt><dd> 123 </dd></dl>
<dl><dt>10 </dt><dd><210> 61 <211> 110 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = "Artificial Sequence Description: Synthetic polynucleotide · </dt><dd /></dl>
<dl><dt>15 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, t, g, Unknown or other </dd></dl>
<dl><dt><400> 61 nnnnnnnnnn </dt><dd>nnnnnnnnnn gtttttgtac tctcaqaaat qcaqaaqcta eaaaqataaq 60 </dd></dl>
<dl><dt>qcttcatqcc gaaatcaaca ccctqtcatt ttatqqcaqq qtqttttttt </dt><dd>Il0 </dd></dl>
<dl><dt>20 </dt><dd><210> 62 <21 1> 137 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>25 </dt><dd><220> <221> source <223> Inota = "Artificial Sequence Description: Synthetic polynucleotide · </dd></dl>
<dl><dt>30 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, t, g, Unknown or other </dd></dl>
<dl><dt><400> 62 nnnnnnnnnn </dt><dd>nnnnnnnnnn qttattqtac tctcaaqatt taqaaataaa tcttgcaqaa 60 </dd></dl>
<dl><dt>gctacaaaqa taaqqcttca tgecgaaatc aacaecctqt cattttatg9 caggqtgttt </dt><dd> 120 </dd></dl>
<dl><dt>tcqttattta atttttt </dt><dd> 137 </dd></dl>
<dl><dt>35 </dt><dd><210> 63 <21 1> 123 <212> AON <213> Artificial Sequence </dd></dl>
<dl><dt>136 </dt><dd /></dl>
<dl><dt><220> <221> in te <223> Inola = ~ Artificial Sequence Description: Synthetic polynucleotide ~ </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, 1, g, Unknown or other </dd></dl>
<dl><dt><400> 63 nnnnnnnnnn </dt><dd>nnnnnnnnnn qttattqtac tctcaqaaat gcagaa9cta caaaq.taaq 60 </dd></dl>
<dl><dt>qottcatqcc qaaatcaaca ccctqtcatt ttatqqoagg qtqttttcqt tatttaattt </dt><dd> 120 </dd></dl>
<dl><dt>ttt </dt><dd> 123 </dd></dl>
<dl><dt>10 </dt><dd><210> 64 <21 1> 1 10 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>15 </dt><dd><220> <221> in te <223> Inola = "Description of Artificial Sequence: Synthetic polynucleotide ~ </dd></dl>
<dl><dt>20 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, 1, g, Unknown or other </dd></dl>
<dl><dt><400> 64 nnnnnnnnnn </dt><dd>nnnnnnnnnn qttattqtac tctcaqaaat qcagaagcta caaaqataaq 60 </dd></dl>
<dl><dt>gcttcatgcc gaaatcaaca ccctgtcatt ttat9qcaqq gtgttttttt </dt><dd> 110 </dd></dl>
<dl><dt>25 </dt><dd><210> 65 <21 1> 137 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = "Description of Artificial Sequence: Polynucleic acid syllable ~ </dt><dd /></dl>
<dl><dt>30 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, 1, g, Unknown or other </dd></dl>
<dl><dt><400> 65 nnnnnnnnnn </dt><dd>nnnnnnnnnn gttattgtac tctcaagatt tagaaataaa tcttqcagaa 60 </dd></dl>
<dl><dt>qctaeaatga taaggcttca tqccgaaatc aacaccctgt cattttatgg cagggtqttt </dt><dd> 120 </dd></dl>
<dl><dt>tcqttattta atttttt </dt><dd> 137 </dd></dl>
<dl><dt>35 </dt><dd /></dl>
<dl><dt><210> 66 <21 1>123 </dt><dd /></dl>
<dl><dt><212> AON <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide. </dd></dl>
<dl><dt><220> <221> modified base <222> (1) .. (20) <223> a, c, t, g, Unknown or other </dt><dd /></dl>
<dl><dt>10 </dt><dd><400> 66 nnnnnnnnnn nnnnnnnnnn qttattgtac tctcagaaat qcaoaagcta eaatgataag 60 </dd></dl>
<dl><dt>qcttcatqcc q_aatcaaea ecctgtcatt ttatQ9caqq qtgttttcqt tatttaattt </dt><dd> 120 </dd></dl>
ttt
<210> 67
<21 1> 110
<212> DNA 15 <213> Artificial Sequence
<220>
<221> source
<223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide.
<220> 20 <221> modified base
<222> (1) ..(20)
<223> a, c, 1, g, Unknown or other
<400> 67 nnnnnnnnnn nnnnnnnnnn gttattgtae tetcaqaaat qcaqaaqct a
qcttcatqcc qaaateaaea ecctgtcatt ttatqqcaqq qtqttttttt
25 <210>68
<dl><dt>< </dt><dd>211 > 107 </dd></dl>
<dl><dt>< </dt><dd>212> AON </dd></dl>
<213> Artificial Sequence
<220> 30 <221> source
<223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide.
<220>
<221> modified base
<222> (1) .. (20) 35 <223> a, e, 1, g, Unknown or other
<400> 68
caatqataaq 60 110
<dl><dt>nnnnnnnnnn </dt><dd /><dt>nnnnnnnnnn </dt><dd>gttttagaqc tqtqqaaaca c & qcqaqtta aaataaqqct 60 </dd></dl>
<dl><dt>taqtecgtac tcaacttqaa aaggtqqeac cqattegqtg ttttttt </dt><dd> 107 </dd></dl>
<dl><dt>138 </dt><dd /></dl>
<210> 69
< 211> 4263
<212> DNA 5 <213> Artificial Sequence
<220>
<221> source
<223> Inota = Artificial Sequence MD: Synthetic polynucleotide "
<400> 69
atqaaaaqqc cqqeqqccac qaaaaaqgcc qqceaggeaa aaaaQaaaaa Q'aeeaaqccc 60 tacaqcatcq qcctqqacat cqqcaccaat aqcgtqqqct qqqccqtgac caccqacaac 120 tacaaqgtgc ccaqcaaqaa aatq-aaggtq ctqgqcaaca cctccaagaa gtacatcaaq 180 aaaaacctqc tqgqcgtqct gctgttcgac aqcqqcatta cagccqagqq cagacqgctg 240 aagagaaccq ccaqacggcq qtacacccqq gaatcctqta cqqagaaaca tctqcaaqaq 300 atcttcaqca ccqaqatqgc taccctgqac gacgccttct tccaqcqgct qgacgacaqc 360 ccgacgacaa ttcctgqtgc gcqggacagc aaqtacccca tcttcqgcaa cctggtqqaa 420 gagaaqgcct accacqacga qttccccacc atctaccacc tgaqaaagta cctgqccgac 480 aqcaccaaqa aqqccqacct qaqactqqtq tatctqqccc tqqcccacat qatcaaqtac 540 cgqqqccact tcctgatcqa qqqcgaqttc aacaqcaaga acaacqacat ccagaagaac 600 ttccaqqact tcctggacac ctacaacqcc atcttcqaqa qcqac :: ctqtc cctqqaaaac 660 agcaagcaqc tqgaaqaqat cqtqaaggac aqctgqaaaa aagatcagca gaaqqaccqc 720 atcctgaaqc tgttccccqq cgagaagaac agcggaatct tcaqcqaqtt tctqaaqctg 780 atcqtgqgca accaggccqa cttc8qaaag tgcttcaacc tgqacqagaa aqccaqcctq 840 cacttcaqea aaqagageta cqac.qaqqac ctqgaaaccc tqctqgqata tatcqqeqac 900 qactacaqcq acqtqttcct gaagqccaaq aagctgtacq acgctatcct qctgaqcqqc 960 ttcctgaccq tgaccg8caa cqaqacagaq qccccactga gcaqcqccat gattaagcqq 1020 taca.acgaqc acaaaqaqqa tctqqctctq ctqaaaqaQt acatccggaa catcaqcctq Wolf aaaacctaca atgagqtgtt caaqqacqac accaagaacg qctacqcC99 ctacatcqac 1140 qqcaagacca Bccaqqaaga tttctatgtg tacctgaaga aqctgctgqc cgaqttcgag 1200
qqgqccgact actttctqqa aaaaatcqac cqcqaqqatt tcctgcqqaa gcagcqqacc 1260 ttcgacaacq qcaqcatccc ctaccaqatc catctqcaqg aaatqcqgqc catcctqqac 1320 aaqcagqcca aqttctaccc attcctgqcc aaqaacaaag agcgqatcqa qaagatcctq 1380 tcccttacta accttccqca cqtqqqcccc ctqqccaqag qcaacagcga ttttqcctgg 1440 agcgcaatqa tccatccqqa qaaqatcacc ccctqqaact tcgaqgacqt gatcqacaaa 1500 qaqtccaqcq ccqaggcctt catcaaccqg atgaccaqct tcgacctgta cctgcccgaq 1560 tqcccaaqca gaaaagqtgc caqcctgctq tacqaqacat tcaatgtgta taacqaqctg 1620 accaaagtgc qgtttatcqc cqaqtctatq cqqqactacc aqttcctqqa ctccaaqcaq 1680 aaaaillqqaca tcqtgcqqct qtacttcaaq gacaagcgga aaqtqaccqa taaqqacatc 1740 atcqagtacc tqcacqccat ctacqqctac qatqqcatcg agctgaaqqg catcqagaag 1800 ccaqcctgag cagttcaact cacataccac gacctqctga acattatcaa cgacaaagaa 1860 tttctqqacq actccaqcaa cgagqccatc atcgaagaqa tcatccacac cctqaccatc 1920 tttgaqqacc gcgagatgat caaqcaqcgq ctgaqcaaqt tcgagaacat cttcgacaaq 1980 agcgtqctga aaaaqctgaq caqacgqcac tacaccqqct ggqqcaaqct gagcqccaaq 2040 ctgatcaacq gcatccqqqa cqaqaaqtcc qqcaacacaa tcctgqacta cctqatcqac 2100 qacqqcatca qcaaccggaa cttcatqcag ctgatccacg acqa c <¡ccct. gagcttcaaq 2160 aagaagatcc aqaaqqccca qatcatcqgq qacqaqqaca aqqqcaacat caaaqaaqtc 2220 qtqaaqteec tqeecqqcaq eeccqecatc aaqaaqgqaa tcctgaqq gatqqqcqgc agaaaqcccq agagcatcqt. gqtqqaaatg 2340 qctaqaqaga aceaqtacac aaqagcaaca caatcaqqgc gccagcaqaq actqaagaga 2400 ccctgaaaqa ct.ggaaaagt gctgqqca <¡C aaqattctqa aagagaatat ccctqccaag 2460 tcqacaacaa ctqtccaaga cqccctqcaq aacqaccqqc tgtacctgta ctacctqcaq 2520 aatgqcaag <acatqtatac aggcgacgac ctgqatatcq accqcctgag caactacgac 2580 atcgaccata ttatccccca aaaqacaaca ggccttcctg gcattgacaa caaagtgctg 2640 qtgtcctccq ccaqcaaccg cgqcaaqtcc qatqatqtgc aqtcgtqaaa ccaqcctqga 2700 aaqaqaaaqa ccttctqqta tcaqctqctg aaaaqcaaQ'c tgattaqcca gaqqaaqttc 2760 <acaacctqa ccaa9Q'ccqa qaqaqgcqqc ctqagccctq aaqataaggc cqgcttcatc 2820 caqagacaqc tqqt.qqaaa.c cCQ'qcaQ'atc accaaqcacq tggccagact qctqqatgag 2880 acaaqaagqa aaqtttaaca cgaqaacaac cgqQ'ccqt. <¡C qqaccqtqaa qatcatcacc 2940 ct <aaqtcca ccctqgtgtc ccaqttccgq aaggacttcg aqctgtataa aqt. <¡cqcqaq 3000 atcaatgact ttcaccacqe ccacgacqcc tacctgaatq ccgtqqtCJO'c ttccqccctq 3060 ctqaaqaaqt accctaagct gqaacccqaq ttcqtgtacg gcgactaccc 31 caaqq
tccttcaqaq aqcgqaagtc cgccaccgag aaqqtgtact tctact CAAC catcatgaat 3180 atctttaaga agtccatctc cctqqccqat qqcaqagtga tcgagcqgcc cctgatcqaa 3240 qtgaacqaaq agacaggcqa gagcgtgtqg aacaaagaaa gcgacctgqc caccgtqcgq 3300 qttatcctca cgggtqctga aqtgaatgtc qtqaagaagq tqqaagaaca qaaccacggc 3360 ctggatcqgg gcaaqcccaa qgqcctqttc aac "lccaacc tgtccaqcaa qcctaagccc 3420 aactccaacg aqaatctcqt qqqggccaaa gaqtacctgg accctaagaa gtacggcgga 3480 tacgecqqea tctceaataq ctteaecqt.g etcgtqaagq gcacaatcqa gaa "lgqcqct 3540 aagaaaaaga tcacaaaeqt qctgqaattt eaqqqgatct ctat.cctqqa CCQ "latcaac 3600 taccqqaaqq ataaqctgaa ctttctqctg qaaaaaqgct acaaqqacat tgagctgatt 3660 atcgaqctqc ctaagtactc cctgttcgaa ctqagcgacg gctccagacg gatgctqqcc 3720 ccaccaacaa tccatcctgt qaqatecaea caaqeqgqqe aqqoaaaeca gatcttcctg 3780 aqccaqaaat ttqtqaaact qctgtaccac qccaaqcgga tctccaacac catcaatgag 3840 aatacqtgga aaccacegqa aaaccacaaQ aaaqaqtttq aqqaactgtt ctactacatc 3900 acqagaacta ctqoaqttca tqtqqgagcc aagaagaacg gcaaactqct gaactccgcc 3960 ttccaqaqct qqcaqaacca caqcatcqac qaqctqtqca qctccttcat cqgccctacc 4020 ggcagcgagc qqaaqqqact qtttqagctq acctccaQag qctctgccqc cqactttgaq 4080 ttcctqqqag tqaaqatccc ccqqt.acaqa qactacaccc cctctaqtct gctqaaqqac 4140 qccaccctqa tccaccaqaq cqtgaccc¡qc ctqtacqaaa ccc9gatcqa cctggctaaq 4200 ctqqqcqaqq qaaaqcqtcc tgctqctact aaqaaaqctc¡ qtcaac¡ctaa qaaaaaqaaa 4260 taa 4263
<210> 70 5 <211> 84
<212> DNA
<213> Artificial Sequence
<220>
<221> source 10 <223> Inota = ~ Artificial Sequence Description "Synthetic oligonucleotide"
<400> 70
gqaaccattc ataacaqcat agcaagttat aataa9gcta qtccqttatc aacttgaaaa 60 aqtgqcaccq aqtcc¡otc¡ct tttt 94
<210> 71 <211>36
<dl><dt><212> AON <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic oligonucleotide ~ </dd></dl>
<dl><dt><400> 71 gtlatagagc tatgctgtla tgaatggtcc caaaac </dt><dd> 36 </dd></dl>
<dl><dt>10 </dt><dd><210> 72 <211> 84 <212> AON <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic oligonucleotide " </dt><dd /></dl>
<dl><dt>'5 </dt><dd><400> 72 99aaec attc aat aeageat agcaaqttaa tataa99cta qtccqttatc to acttgaaaa 60 </dd></dl>
<dl><dt>aqtqgcaccq aqtcqqtgct tttt </dt><dd> 84 </dd></dl>
<dl><dt>20 </dt><dd><2 10> 73 <211> 36 <212> AON <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = "Artificial Sequence Description: Synthetic oligonucleotide" </dt><dd /></dl>
<dl><dt>25 </dt><dd><400> 73 gtatlagagc tatgclgtat tgaatggtcc caaaac 36 </dd></dl>
<dl><dt><210> 74 <21 1> 103 <212> AON <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>30 </dt><dd><220> <221> source <223> Inota = "Artificial Sequence Description: Synthetic polynucleotide" </dd></dl>
<dl><dt>35 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, 1, g, Unknown or other </dd></dl>
<dl><dt><400> 74 nnnnnnnnnn </dt><dd>nnnnnnnnnn gttttagaqc tagaaatagc aaqttaaaat aaqqctagtc 60 </dd></dl>
<dl><dt>cqttatcaac ttgaaaaagt ggcaecgagt cgqtgctttt </dt><dd>ttt 103 </dd></dl>
<dl><dt>40 </dt><dd><210> 75 <211> 103 <212> AON <213> Artificial Sequence </dd></dl>
<dl><dt>'42 </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inola = ~ Artificial Sequence Description: Synthetic polynucleotide ~ </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> modified base <222> (1) "(20) <223> a, c, 1, g, Unknown or other </dd></dl>
<dl><dt><400> 75 nnnnnnnnnn </dt><dd>nnnnnnnnnn qtattaqaqc taga aataqc aaqt ~ aat.t aaggctaqtc 60 </dd></dl>
<dl><dt>cqttatcaac ttgaaaaagt 9qcaccgagt cqgtqctttt ttt </dt><dd> 103 </dd></dl>
<dl><dt>10 </dt><dd><210> 76 <211> 123 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>15 </dt><dd><220> <221> fuenle <223> Inota = "Artificial Sequence Description: Synthetic polynucleotide ~ </dd></dl>
<dl><dt>20 </dt><dd><220> <221> modified base <222> (1) "(20) <223> a, c, 1, g, Unknown or other </dd></dl>
<dl><dt><400> 76 nnnnnnnnnn </dt><dd>nnnnnnnnnn qttttaqaqc tatqctqttt tooaaacaaa acaqcataqc 60 </dd></dl>
<dl><dt>: laqttaaaat aaqqctaqtc cqttatcaac ttqaaaaaqt 9qcaccqaqt cqqtqctttt </dt><dd> 120 </dd></dl>
<dl><dt>ttt </dt><dd> 123 </dd></dl>
<dl><dt>25 </dt><dd><210> 77 <211> 123 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = "Artificial Sequence Description: Synthetic polynucleotide ~ </dt><dd /></dl>
<dl><dt>30 </dt><dd><220> <221> modified base <222> (1) "(20) <223> a, c, 1, g, Unknown or other </dd></dl>
<dl><dt>35 </dt><dd> <400> 77 </dd></dl>
<dl><dt>aaqttaatat aaqqctagtc cgttatcaac ttqaaaaaqt oqcaccqaqt cgqtqctttt </dt><dd> 120 </dd></dl>
<dl><dt>ttt </dt><dd> 123 </dd></dl>
<dl><dt>nnnnnnnnnn </dt><dd /><dt>nnnnnnnnnn </dt><dd>qtattaqagc tatqctqtat tgqaaaca at acagcatagc 60 </dd></dl>
<dl><dt>40 </dt><dd> <210> 78 <211>20 </dd></dl>
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 78 gtcacctcca atgactaggg 20
5 <210> 79
<211> 984
<212> PRT
<213> Campylobacter jejuni
<400> 79
Met Wing Arch Ile Leu Wing Phe Asp Ile Gly Ile Ser Ser Ile Gly Trp 1 5 10 15
Wing Phe Ser Glu Asn Asp Glu Leu Lys Asp eys Gly Val Arq Ile Phe 20 25 30
Thr Lys Val Glu Asn Pro Lys Thr Gly Glu Ser Leu Ala Leu Pro Arq 35 40 45
Arq Lau Ala Arg Ser Ala Arq Lys Arq Leu Ala Arq Arq Lys Ala Arq 50 55 60
Lau Aan His Lau Lys His Lau Ile Ala Asn Glu Phe Lys Leu Asn Tyr 65 70 75 60
Glu Asp Tyr Gln Ser Phe Asp Glu Ser Leu Ala Lys Ala Tyr Lys Gly 85 90 95
Ser Lau Ile Ser Pro Tyr Glu Lau Arq Phe Arq Ala Lau Asn Glu Leu 100 105 110
Lau Ser Lys Gln Asp Phe Ala Arq Val Ile Lau His Ile Ala Lys Arq 115 120 125
Arq Gly Tyr Asp Asp 11e LY8 Asn Ser Asp Asp Lys Glu Lys Gly Ala 130 135 140
Ile Leu Lys Wing Ile Lys Gln Asn Glu Glu Lys Lau Wing Asn Tyr Gln 145 150 155 160
Ser Val Gly Glu Tyr Lau Tyr Lys Glu Tyr phQ G1n Lys Phe Lys Glu 165 170 175
Asn Ser Lys Glu Phe Thr Asn Val Arg Asn Lys Lys Glu Ser Tyr Glu 180 185 190
Arg eye Ile Gln Ser Phe Leu Lys Asp Glu Leu Lys Leu Ile Phe 195 200 205
Lys Lys Gln Arq Glu Phe Gly Phe Ser Phe Ser Lys Lys Phe Glu Glu 210 215 220
Glu Val Leu Ser Val ~ a Phe Tyr Lys Arq Ala Leu Lys Asp Phe Ser 225 230 235 240
His Leu Val Gly Asn eye Ser Phe Phe Thr Asp G1u Lys Arg Ala Pro 245 250 255
Lys Asn Ser Pro Leu Wing Phe Met Phe Val Wing Leu Thr Arg Ile 260 265 270
Asn Leu Leu Asn Asn Leu Lys Aso Thr Glu Gly Ile Leu Tyr Thr Lys 275 280 285
Asp Asp Leu Asn Ala Leu Leu Asn Glu Val Leu Lys Asn Gly Thr Leu 290 295 300
Thr Tyr Ly. Gln Thr Lya Lys Leu Leu Gly Leu Ser Asp Aap Tyr Glu 305 310 315 320
Phe Lya Gly Glu Lys Gly Thr Tyr Pha Ila Glu Pha Lys Lys Tyr Lys 325 330 335
Glu Phe Ile Lys Ala Leu Gly Glu Bis Asn Leu Ser Glo Asp Asp Leu 340 345 350
Asn Glu Ile Wing Lys Asp Ile Thr Leu Ile Lys Asp Glu Ile Lys Leu 355 360 365
Lys Lys Wing Léu Wing Lys Tyr Asp Leu Asn Glo Aso Gln Ile Asp Ser 370 375 380
Leu Ser Lys Leu Glu Phe Lys Asp Bis Leu Aso Ile Ser Pha Lys Wing 385 390 395 400
Leu Lys Leu Val Thr Pro Leu Met Leu Glu Gly Lys Lys Tyr Asp Glu 405 410 415
Wing eye Asn Glu Leu Asn Leu Lys Val Wing Ila Asn Glu Asp Lys Lys 420 425 430
Asp Phe Leu Pro Ala Phe Asn Glu Thr Tyr Tyr Lys Asp Glu Val Thr
435 440 445
Asn Pro Val Val Leu Arg Ile Lys Glu Tyr Arg Lys Val Leu Asn 450 455 460
Ala Leu Leu Lys Lys Tyr Gly Lys Val His Lys 118 Asn Ile Glu Leu
465 470 475 480
Wing Arg Glu val Gly Lys Asn His Ser Gln Arg Wing Lys Ile Glu Lys 485 490 495
Glu Gln ~ n Glu Asn Tyr Lys Ala Lys Lys Asp Ala Glu Leu Glu eys 500 505 510
Glu Lys Leu Gly Leu Lys Ile Asn Ser Lys ABn Ile Leu Lys Leu Arch 515 520 525
Leu Phe Lys Glu Gln Lys Glu Phe eys Ala Tyr Ser Gly Glu Lys Ile 530 535 540
Lys Ile Ser Asp Leu Gln Asp Glu Lys Mat Leu Glu Ile Asp Hia Ile 545 550 555 560
Tyr Pro Tyr Ser Arq Ser Phe ASp Asp Ser Tyr Met Asn Lys Val Leu 565 570 515
Val Phe Thr Lys Gln Asn Gln Glu Lys Leu Asn Gln Thr Pro Phe Glu 580 585 590
Wing Phe Gly Asn Asp Ser Wing Lys Trp Gln Lys Ile Glu Val Leu Wing 595 600 605
Lys Asn Leu Pro Thr Lys Lys Gln LyS ArO Ile Leu Asp Lys Aso Tyr 610 615 620
Lys ASp Lys Glu Gln Lys Asn Phe Lys Asp Arch Asn Leu Asn Asp Thr 625 630 635 640
Arg Tyr Ile Ala Arg Leu Val Leu Aso Tyr Thr Lys Asp Tyr Leu Asp 645 650 655
Phe Leu Pro Leu Ser Asp Asp Glu Asn Thr Lys Leu Asn Asp Thr Gln 660 665 670
Lys Gly Ser Lys val His Val Glu Ala Lys Ser Gly Met Leu Thr Ser
675 680 685
<dl><dt>Ala Leu Ar9 Mis </dt><dd>Thr Trp Gly Phe Be ~ a Lys Asp Arq Asn Asn His </dd></dl>
<dl><dt>690 </dt><dd> 695 700 </dd></dl>
<dl><dt>Leu </dt><dd>Bis His Ala Ile Asp Ala val Ile l1a To Tyr Ala Asn Asn Be </dd></dl>
<dl><dt>705 </dt><dd> 710 715 720 </dd></dl>
<dl><dt>Ile val Lys Ala Phe </dt><dd>Be Asp Phe Lys Lys Glu Gln Glu Ser Asn Be </dd></dl>
<dl><dt>725 </dt><dd> 730 135 </dd></dl>
Wing Glu Leu Tyr Wing Lys Lys Ile Ser Glu Leu Asp Tyr Lys Asn Lys 7 ~ O 745 150
Arq Lys phe Phe Glu Pro Phe Ser Gly Phe Arq Gln Lys Val Leu ASp 755 760 165
Lys Ile Asp Glu Ile Phe Val Ser Lys Pro Glu Arq Lys Lys Pro Ser 770 175 780
Gly Ala Leu Mis Glu Glu Thr Phe Arq Lys Glu Glu Glu Phe Tyr Gln
785 790 795 800
Be Tyr Gly Gly Lys Glu Gly Val Leu Lys ~ a Leu Glu Leu Gly Lys 805 810 815
eleven. Arq Lys val Asn Gly Ly. Ile Val Lys Asn Gly Aap Met Phe Arq 820 825 830
Val Asp Ile Phe Lys His Lys Lys Thr Asn Lys Phe Tyr Al.a Val Pro 835 840 B45
<dl><dt>Ile Tyr Thr Met Asp Pho Ala </dt><dd>Leu Lys Val What u Pro Asn Lya Wing Val </dd></dl>
<dl><dt>850 </dt><dd> 855 860 </dd></dl>
<dl><dt>Ala Arq Ser Lys Lys Gly Glu </dt><dd>Ile Lye Asp Trp Ile Leu Met Asp Glu </dd></dl>
<dl><dt>865 </dt><dd> 8"10 875 880 </dd></dl>
<dl><dt>Aan Tyr Glu Phe Cys Phe Ser Leu Tyr </dt><dd>Lys Asp Ser Leu Ile Leu Ile </dd></dl>
<dl><dt>BBS </dt><dd>B90 B95 </dd></dl>
<dl><dt>Gln Thr Lys Asp Met Gln Glu Pro Glu </dt><dd>Phe Val Tyr Tyr Asn A1a Phe </dd></dl>
<dl><dt>900 </dt><dd> 905 910 </dd></dl>
Thr Ser Ser Thr Val Ser Leu Ile Val Ser Lys My Asp Asn Lys Phe 915 920 925
<dl><dt>Glu Thr Leu Ser Ly8 930 </dt><dd>Asn Gln Ly8 935 Ile Leu Phe LyB Asn Ala Asn Glu 940 </dd></dl>
<dl><dt>Lys 945 </dt><dd>Glu Val Ile Ala Lys 950 Be Ile Gly Ile Gln Asn LeU Lys Val Phe 955 960 </dd></dl>
<dl><dt>Glu Lys Tyr </dt><dd>Ile Val 965 Ser A1a Leu Gly Glu Val 910 Thr Ly8 A1a Glu Phe 915 </dd></dl>
<dl><dt>Arq Gln Arq Glu Asp Phe Lys 980 </dt><dd>Lys </dd></dl>
<dl><dt>5 </dt><dd><210> 80 <211> 91 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>10 </dt><dd><220> <221> source <223> Inota :: - Description of Artificial Sequence: Ol synthetic oligonucleotide " </dd></dl>
<dl><dt><400> 80 tataatetea taaqaaattt aaaaaqqqac taaaataaaq agtttqcqqq actctgcgqq </dt><dd> 60 </dd></dl>
<dl><dt>qttacaatcc cctaaaaccq cttttaaaat t </dt><dd> 91 </dd></dl>
<dl><dt>15 </dt><dd><210> 81 <211> 36 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>20 </dt><dd><220> <221> source <223> Inola = -Artificial Sequence Description · Synthetic oligonucleotide " </dd></dl>
<dl><dt><400> 81 attttaccat aaagaaattt aaaaagggac taaaac </dt><dd> 36 </dd></dl>
<dl><dt>25 </dt><dd><210> 82 <211> 95 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota :: - Description of Artificial Sequence Ol synthetic igonucleotide " </dt><dd /></dl>
<dl><dt>30 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, u, g, Unknown or other </dd></dl>
<400> 82
nnnnnnnnnn nnnnnnnnnn quuuuagucc cgaaagggac uaaaauaaaq aquuugcgqg
ac ucuqcqqo quuacaaucc ccuaaaaccg cuuuu
<210> 83
<211> 69 5 <212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
10 <400>83
qucaccucca augacuaqgg quuuuagaqc uaqaaauaqc aaquuaaaau aagqcuaguc
cguuuuuuu
<210> 84 <211>69
<212> RNA 15 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 84
qgqqccgaga uuggguquuc guuuuaqagc uagaaauagc aaguuaaaau aaqgcuaguc
cquuuuuuu
<210> 85
<211> 69
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
25 <220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 85
guqgcgagag gggccgagau guuuuagagc uagaaauagc aaquuaaaau aaggcuaguc
cquuuuuuu
30 <210>86 <211>69
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220> 35 <221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleolide "
60 95
60 69
bU 69
60 69
<400> 86
qgqqccqaqa uuqgguquuc quuuuaqaqc uaqaaauaqc aaquuaaaau aaqgcuaquc
cguuuuuuu
<210> 87
<211> 69 5 <212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide ~ </dd></dl>
10 <400> 87
qugqcqaqaq qqqccqaqau guuuuaqaqc uaqaaauaqc aaquuaaaau aaqqcuaguc
cguuuuuuu
<210> 88
<211> 76
<212> RNA 15 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 88
qucaccucca augacuaqqq quuuuaqaqc uaqaaauaqc aaquuaaaau aaqgcuaquc
cguuaucauu uuuuuu
<210> 89 <211>76
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
25 <220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 89
qacaucqauq uccuccccau quuuuaqaqc uaqaaauaqc aaquuaaaau aaggcuaquc
cguuaucauu uuuuuu
30 <210>90
<211> 76
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220> 35 <221> source
<223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide"
bU
60 69
60 76
60 76
<400> 90
gaquccgaqc aqaagaagaa guuuuagagc uaqaaauagc aaquuaaaau aaqqcuaquc
cquuaucauu uuuuuu
<210> 91
<21 1> 76 5 <212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic Oligonucleotide ~ </dd></dl>
10 <400> 91
gaquccgaqc aqaagaagaa quuuuagagc uaqaaauagc aaquuaaaau aaqqcuaquc
cquuaucauu uuuuuu
<210> 92 <211>76
<212> RNA 15 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota :: ~ Description of Artificial Sequence: Ol synthetic igonucleotide " </dd></dl>
<400> 92
ggqqccgaqa uugqguquuc quuuuagagc uagaaauagc aaquuaaaau aagqcuaquc
cquuaucauu uuuuuu 20
<210> 93
<211> 88
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
25 <220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description "Ol synthetic igonucleotide ~ </dd></dl>
<400> 93
qucaccucca augacuaggg quuuuagaqc uaqaaauagc aaquuaaaau aaqqcuaguc
cguuaucaac uuqaaaaagu quuuuuuu
30 <210>94
<211> 88
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220> 35 <221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide "
60 16
60 16
60 76
60 88
<400> 94
qacaucqauq uccuccccau quuuu & qagc
cguuaucaac uugaaaaaqu quuuuuuu
<210> 95
<211> 88 5 <212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence </dd></dl>
10 <400> 95
uaqaaauaqc aaquuaaaau aaqqcuaquc
aa
"Synthetic oligonucleotide ~
qaquccqaqc aqaaqaagaa quuuuaqaqc uaqaaauaqc aaquuaaaau aa9qcuaquc 60
cquuaucaac uuqaaaaaqu quuuuuuu aa
<210> 96 <211>88
<212> RNA 15 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota :: ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 96
qoqqccqaqa uuqgququuc guUUU & q & qc u & oaaauaqc aaguuaaaau aagqcuaquc 60
cquuaucaac uuqaaaaaqu quuuuuuu 88
<210> 97
<211> 88
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
25 <220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide ~ </dd></dl>
<400> 97
guqqcqaqaq 90qccqaqau guuuuaqaqc uaqaaauaqc aaquuaaaau Baqgcuaguc 60
cguuaucaac uuqaaaaaqu guuuuuuu 88
30 <210>98
<211> 103
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220> 35 <221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide "<400> 98
qucaccucca augacuaggq quuuuaqaqc uaqaaauagc aaquuaaaau aaggcuaquc
cquuaucaac uugaaaaaqu qqcaccqaqu oqquqcuuuu uuu
<210> 99
<211> 103 5 <212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description: Synthetic oligonucleotide ~ </dd></dl>
10 <400> 99
qucaccucca augacuaggg quuuuaqaqc uaqaaauagc aaquuaaaau aagqcuaquc
cquuaucaac uuqaaaaaqu qqcaccqaqu oqqugcuuuu uuu
<210> 100 <211>103
<212> RNA 15 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence l: Synthetic oligonucleotide " </dd></dl>
<400> 100
qaquccgagc aqaaqaaqaa quuuuagaqc uaqaaauaqc aaquuaaaau aaqqcuaquc
cquuaucaac uuqaaaaaqu qgcaccgaqu cgqugcuuuu uuu 20
<210> 101 <211>103
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
25 <220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Artificial Sequence" Synthetic oligonucleotide " </dd></dl>
<400> 101
gggqccgaga uugqququuc quuuuagaqc uaqaaauaqc aaquuaaaau aaq9Cu & qUC
cquuaucaac uuqaaaaaqu qqcaccqaqu cqquqcuuuu uuu
30 <210> 102
<211> 103
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220> 35 <221> source
<223> Inota = "Description of Artificial Sequence" Synthetic oligonucleotide "
60 103
60 103
60 103
60 103
<dl><dt><400> 102 9 \ 1Qg-Cg- · 9 · 9 Q99cc9agau quuuuaqaqc uaqaaauaqc aaquu •••• u </dt><dd>aag-9Cu.quC 60 </dd></dl>
<dl><dt>cquuauc •• c </dt><dd>uuqa •• aaqu qqcaccqaqu cqQ'Uqcuuuu uuu 103 </dd></dl>
<dl><dt>5 </dt><dd><2 10> 103 <211> 102 <212> AON <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide " </dt><dd /></dl>
<dl><dt>10 </dt><dd><400> 103 qttttaqaqc tatqctqtt.ttqaat.qqtcc caaaacqqaa qqqcctqaqt ccqaqcagaa 60 </dd></dl>
<dl><dt>qaaqaaqttt taqaqct.atq ctgttttqaa tqqt.ccca.a </dt><dd>& C 102 </dd></dl>
<dl><dt>15 </dt><dd><210> 104 <21 1> 100 <212> DNA <213> Furnace sapiens </dd></dl>
<dl><dt><400> 104 cqgaqqacaa aqtacaaacq qcaqaa9ct.q qa9qaqqaaq qqcct.qaqtc cqaqcaqaaq </dt><dd> 60 </dd></dl>
<dl><dt>aaqaaqqqet </dt><dd>cccat.cacat caacC {lgtqq cqcat.tqcca 100 </dd></dl>
<dl><dt>20 </dt><dd><210> 105 <211> 50 <212> AON <213> Furnace sapiens </dd></dl>
<dl><dt><400> 105 agctggagga ggaagggcct gagtccgagc agaagaagaa gggctcccac </dt><dd> 50 </dd></dl>
<dl><dt>25 </dt><dd><210> 106 <21 1> 30 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt>30 </dt><dd><220> <221> source <223> Inota = -Artificial Sequence Description: Synthetic oligonucleotide · </dd></dl>
<dl><dt><400> 106 gaguccgagc agaagaagaa guuuuagagc </dt><dd> 30 </dd></dl>
<dl><dt>35 </dt><dd><2 10> 107 <211> 49 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> fu <<23>> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dt><dd /></dl>
<400> 107 agctggagga ggaagggcct gagtccgagc agaagagaag ggctcccat 49
<210> 108 <211>53
<212> DNA
<213> Sapiens oven
<400> 108 ctggaggagg aagggcctga gtccgagcag aagaagaagg gctcccatca cat 53
<210> 109 <211>52
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 109 ctggaggagg aagggcctga gtccgagcag aagagaaggg ctcccatcac at 52
<210> 110
<211> 54
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 110 ctggaggagg aagggcctga gtccgagcag aagaaagaag ggctcccatc acat 54
<210>111
<211> 50
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 111 ctggaggagg aagggcctga gtccgagcag aagaagggct cccatcacat 50
<210> 112 <211>47
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 112 ctggaggagg aagggcctga gcccgagcag aagggctccc atcacat 47
<210> 113
<211> 66
<212> DNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Description of Artificial Sequence l: Synthetic oligonucleotide" </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = "Combined DNA / RNA molecule description: Synthetic oligonucleotide" </dd></dl>
<220>
<dl><dt>< </dt><dd>221> modified base </dd></dl>
<dl><dt /><dd>< 222> (1) (20) </dd></dl>
<dl><dt>< </dt><dd>223> a, c, t, g, Unknown or other </dd></dl>
<dl><dt><400> 113 nnnnnnnnnn </dt><dd>nnnnnnnnnn guuuuagaqc uagaaauaqc aaguuaaaau aaggctagtc 60 </dd></dl>
<dl><dt>cquuuu </dt><dd> 66 </dd></dl>
<dl><dt><210> 114 <21 1> 20 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota = ~ Description of Artificial Sequence "Ol synthetic igonucleotide ~ </dt><dd /></dl>
<dl><dt><400> 114 gaguccgagc agaagaagaa </dt><dd> 20 </dd></dl>
<dl><dt><210> 115 <211> 20 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota "'~ Description of Artificial Sequence Synthetic oligonucleotide" </dt><dd /></dl>
<dl><dt><400> 115 gacaucgauguccuccccau </dt><dd> 20 </dd></dl>
<dl><dt><210> 116 <211> 20 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic oligonucleotide " </dt><dd /></dl>
<dl><dt><400> 116 gucaccucca augacuaggg </dt><dd> 20 </dd></dl>
<dl><dt><210> 117 <211> 20 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota ::: ~ Artificial Sequence Description: Synthetic oligonucleotide " </dt><dd /></dl>
<dl><dt><400> 117 auuggguguu cagggcagag </dt><dd> 20 </dd></dl>
<dl><dt><210> 118 <211> 20 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Synthetic Oligonucleotide Artificial Sequence ~ </dd></dl>
<400> 118 guggcgagag gggccgagau 20
<210> 119 <211>20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 119 ggggccgaga uuggguguuc 20
<210> 120
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence l: Synthetic oligonucleotide ~ </dd></dl>
<400> 120 gugccauuag cuaaaugcau 20
<210> 121 <211>20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description: Synthetic oligonucleotide ~ </dd></dl>
<400> 121 guaccaccca caggugccag 20
<210> 122
<211> 20
<212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence l: Synthetic oligonucleotide ~ </dd></dl>
<400> 122 gaaagccucu gggccaggaa 20
<210> 123 <211>48
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 123 ctggaggagg aagggcctga gtccgagcag aagaagaagg gctcccat 48
<210> 124 <211>20
<212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 124 gaguccgagc agaagaagau 20
<210> 125
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 125 gaguccgagc agaagaagua 20
<210> 126
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 126 gaguccgagc agaagaacaa 20
<210> 127
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 127 gaguccgagc agaagaugaa 20
<210> 128
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 128 gaguccgagc agaaguagaa 20
<210> 129 <211>20
<212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 129 gaguccgagc agaugaagaa 20
<210> 130
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 130 gaguccgagc acaagaagaa 20
<210> 131
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 131 gaguccgagg agaagaagaa 20
<210> 132
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 132 gaguccgugc agaagaagaa 20
<210> 133
<211> 20
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 133 gagucggagc agaagaagaa 20
<210> 134 <211>20
<212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 134 gagaccgagc agaagaagaa 20
<210> 135
<211> 24
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 135 aatgacaagc ttgctagcgg tggg 24
<210> 136
<211> 39
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide" </dd></dl>
<400> 136 aaaacggaag ggcctgagtc cgagcagaag aagaagttt 39
<210> 137
<211> 39
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<400> 137 aaacaggggc cgagattggg tgttcagggc agaggtttt 39
<210> 138
<211> 38
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide " </dd></dl>
<dl><dt>< </dt><dd>222> (1) "" (20) </dd></dl>
<dl><dt>< </dt><dd>223> a, c, u, g, Unknown or other </dd></dl>
<dl><dt><400> 138 aaaacggaag ggcclgaglc cgagcagaag aagaagtt </dt><dd> 38 </dd></dl>
<dl><dt>5 </dt><dd><210> 139 <211> 40 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> fuenle <223> Inota = ~ Artificial Sequence Description "Ol synthetic igonucleolide" </dt><dd /></dl>
<dl><dt>10 </dt><dd><400> 139 aacggaggga ggggcacaga tgagaaactc agggttttag 40 </dd></dl>
<dl><dt>15 </dt><dd><210> 140 <211> 38 <212> DNA <213> Furnace sapiens </dd></dl>
<dl><dt><400> 140 agccctLctt cttctgctcg gaclcaggcc cttcctcc </dt><dd> 38 </dd></dl>
<dl><dt>20 </dt><dd><210> 141 <211> 40 <212> DNA <213> Furnace sapiens </dd></dl>
<dl><dt><400> 141 cagggaggga ggggcacaga tgagaaactc aggaggcccc </dt><dd> 40 </dd></dl>
<dl><dt>25 </dt><dd><210> 142 <211> 80 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>30 </dt><dd><220> <221> Fuenle <223> Inola = ~ Description of Artificial Sequence: Ol Synthetic Igonucleotide " </dd></dl>
<dl><dt><400> 142 qqcaatqcqc caccqqttqa tqtqatqgga gcccttctag gaggccccca gagcagccac </dt><dd> 60 </dd></dl>
<dl><dt>tgqqgcctca acactcagqc </dt><dd> 80 </dd></dl>
<dl><dt>35 </dt><dd><210> 143 <211> 98 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> fuenle <223> Inota = ~ Description of Artificial Sequence "Synthetic Oligonucleolide" </dt><dd /></dl>
<dl><dt>40 </dt><dd><400> 143 qqacgaaaca cC9gsaccat tcaaaacagc atagcaaqtt aaaataag9c tagtccgtta 60 </dd></dl>
<dl><dt>tcaacttgaa & aagtqqcac cqagtcgqtg cttttttt </dt><dd> 98 </dd></dl>
<dl><dt>161 </dt><dd /></dl>
<dl><dt><210> 144 <211> 186 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide " </dd></dl>
<dl><dt><400> 144 gqacgaaBca ccggtagtat taaqtattgt tttatqqctq ataaatttct ttgaatttct </dt><dd> 60 </dd></dl>
<dl><dt>ccttqattat ttgttataaa aqttataaaa taatcttqtt gqaaccattc aaaacaqcat </dt><dd> 120 </dd></dl>
<dl><dt>aqcaaqttaa aataaqqcta qtccqttatc aacttqaaaa aqtqqcaccq aqtcqgtgct </dt><dd> 180 </dd></dl>
<dl><dt>tttttt </dt><dd> 186 </dd></dl>
<dl><dt>10 </dt><dd><210> 145 <211> 46 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt>15 </dt><dd><220> <221> source <223> Inota = "Description of Artificial Sequence: Synthetic oligonucleotide" </dd></dl>
<dl><dt>20 </dt><dd><220> <221> modified base <222> (1) .. (19) <223> a, c, u, g, Unknown or other </dd></dl>
<dl><dt><400> 145 nnnnnnnnnn nnnnnnnnng uuauuguacu cucaagauuu auuuuu </dt><dd> 46 </dd></dl>
<dl><dt>25 </dt><dd><210> 146 <21 1> 91 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = "Description of Artificial Sequence: Synthetic oligonucleotide" </dt><dd /></dl>
<dl><dt>30 </dt><dd><400> 146 quuacuuaaa ucuuqcaqaa qcuacaaaqa uaaqqcuuca uqccqaaauc aacacccuqu 60 </dd></dl>
<dl><dt>eauuuuauqq </dt><dd>caqqququuu ucguuauuua a 91 </dd></dl>
<dl><dt>35 </dt><dd><2 10> 147 <211> 70 <212> DNA <213> Furnace sapiens </dd></dl>
<dl><dt><400> 147 ttttctagtq ctqaqtttct qtqactcctc tacattctac ttctctgtgt ttctqtatac </dt><dd> 60 </dd></dl>
<dl><dt>tacct cct cc </dt><dd> 70 </dd></dl>
<dl><dt>162 </dt><dd /></dl>
<dl><dt><210> 148 <21 1> 122 <212> DNA <213> Homo sapiens </dt><dd /></dl>
<dl><dt>5 </dt><dd> <400> 148 </dd></dl>
<dl><dt>ggaq9a8999 cctqaqtccq aqcaqaagaa qaaq9qetce eateaeatea aeegqtg9cq </dt><dd> 60 </dd></dl>
<dl><dt>cattqecacq aaqcagqeca atqqqqaqqa eatcqatqtc aceteeaatq actagqgtg9 </dt><dd> 120 </dd></dl>
<dl><dt>that </dt><dd> 122 </dd></dl>
<dl><dt>10 </dt><dd><210> 149 <211> 48 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt>15 </dt><dd><220> <221> source <223> Inota = "Artificial sequence descriptor: Synthetic oligonucleotide" </dd></dl>
<dl><dt><220> <221> modified base <222> (3) .. (32) <223> a, c, u, 9, Unknown or other </dt><dd /></dl>
<dl><dt>20 </dt><dd><400> 149 acnnnnnnnn nnnnnnnnnn nnnnnnnnnn nnguuuuaga gcuaugcu 48 </dd></dl>
<dl><dt>25 </dt><dd><2 10> 150 <21 1> 67 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = "Description of Artificial Sequence: Synthetic oligonucleotide" </dt><dd /></dl>
<dl><dt>30 </dt><dd><220> <221> source <223> Inota = "Combined DNA / RNA molecule description: Synthetic oligonucleotide" </dd></dl>
<dl><dt><4qQ ": 150 ageauageaa quuaaaauaa qqetaguccg </dt><dd>uuaucaaeuu qaaaaaqu99 cacegaqucq 60 </dd></dl>
<dl><dt>qugcuuu </dt><dd> 67 </dd></dl>
<dl><dt>35 </dt><dd><2 10> 151 <211> 62 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt>40 </dt><dd><220> <221> source <223> Inota = "Description of Artificial Sequence: Ol synthetic igonucleotide" </dd></dl>
<dl><dt><220> <221> modified base </dt><dd /></dl>
<dl><dt>163 </dt><dd /></dl>
<400> 151
nnnnnnnnnn nnnnnnnnnn quuuuaqaqc uagaaauaqc aaquuaaaau aaqqcuaguc
cq
5 <210> 152 <211>73
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220> 10 <221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic oligonucleotide "
<400> 152
tqaatgqtcc caa88c9gaa ggqcctqaqt ccgaqcaqaa qaagaagttt taqagctatg
ctgttttgaa t90
<210> 153 15 <211>99
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<221> source 20 <223> Inota = ~ Artificial Sequence Description "Synthetic oligonucleotide"
<220>
<dl><dt>< </dt><dd>221> modified base </dd></dl>
<dl><dt>< </dt><dd>222> (1) (20) </dd></dl>
<dl><dt>< </dt><dd>223> a, c, u, g, Unknown or other </dd></dl>
25 <400> 153
nnnnnnnnnn nnnnnnnnnn quuuuaqagc uagaaauaqc aaquuaaaau aaqqcuaquc
cquuaucaac uugaaa •• qu ggcaccqagu cgquqcuuu
<210> 154
<211> 127
<212> RNA 30 <213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic polynucleotide" </dd></dl>
<400> 154
quuuuuquac ucucaaqauu uaaquaacuq uacaaeguu. cuuaaaueuu qeagaaqcua
caaagauaaq gcuucaugcc qaaaucaaca cceuqucauu uuauqgcaqq ququuuucqu
uauuuaa 35
6.
6.
6.
60 120 127
<dl><dt><210> 155 <211> 56 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic oligonucleotide " </dd></dl>
<dl><dt>10 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, t, g, Unknown or other </dd></dl>
<dl><dt><400> 155 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcaagatt taagtaactg tacaac </dt><dd> 56 </dd></dl>
<dl><dt>15 </dt><dd><210> 156 <211> 91 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>20 </dt><dd><220> <221> source <223> Inota ::: Artificial Sequence MD · Synthetic oligonucleotide " </dd></dl>
<dl><dt><400> 156 gttacttaaa tcttqcaqaa qctac aaaga taaggcttca tgccqaaatc aacaccctgt </dt><dd> 60 </dd></dl>
<dl><dt>cattttatqq caqqqtgttt tcqttattta a </dt><dd> 91 </dd></dl>
<dl><dt>25 </dt><dd><210> 157 <211> 134 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic polynucleotide " </dt><dd /></dl>
<dl><dt>30 </dt><dd><220> <221> base modified to <222> (1) .. (20) <223> a, c, t, g, Unknown or other</dd></dl>
<dl><dt><40º ~ J57 nnnnnnnnnn </dt><dd>nnnnnnnnnn gtttttqtae teteaagatt taaqgaaaet aaatettqea 60 </dd></dl>
<dl><dt>qaagctaeaa aqataaqget teatgcegaa ateaacacee tgtcatttta tggcaggqtg </dt><dd> 120 </dd></dl>
<dl><dt>35 </dt><dd>ttttcqttat ttaa 134 </dd></dl>
<dl><dt><210> 158 <211> 131 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<220>
<221> source
<223> Inota = ~ Description of Artificial Sequence Synthetic Polynucleotide "
<220> 5 <221> modified base
< 222> (1) .. (20)
<223> a, c, t, g, Unknown or olro
<400> 158
nnnnnnnnnn nnnnnnnnnn qtttttgtac tctcaaqatt tagaaataaa tcttqcaqaa 60
qctacaaaqa taaqqcttca tqccqaaatc aacaccctgt cattttatqq caqqqtqttt 120
tcqttattta to 131
10 <210> 159
<211> 125
<212> DNA
<213> Artificial Sequence
<220> 15 <221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic polynucleotide "
<220>
<221> modified base
<222> (1) .. (20) 20 <223> a, c, t, g, Unknown or airo
<400> 159
nnnnnnnnnn nnnnnnnnnn qtttttqtac tctcaaqatq aaaatcttgc aqaaqctaca 60
aaqataaqqc ttcatqccqa aatcaacacc ctgtcatttt atqqcaq9gt gttttcgtta 120
tttaa 125
<210> 160
<211> 112 25 <212> DNA
<213> Artificial Sequence
<220>
<221> source
<223> Inota = ~ Description of Artificial Sequence: Synthetic polynucleotide "
30 <220>
<221> modified base
< 222> (1) .. (20)
<223> a, c, t, g, Unknown or airo
<400> 160
nnnnnnnnnn nnnnnnnnnn gtttttgtac tctgaaaaqa aqctacaaaq ataaqqcttc 60
atqccqaaat caacaccctq tcattttatq qcaqqgtgtt ttcgttattt aa ll2
<dl><dt><210> 161 <21 1> 107 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic polynucleotide " </dd></dl>
<dl><dt>10 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, t, g, Unknown or olro </dd></dl>
<dl><dt><400> 161 nnnnnnnnnn </dt><dd>nnnnnnnnnn qtttt tqtec tqaaaaqcta caaaqataaq qcttcatqcc 60 </dd></dl>
<dl><dt>qaaat caaca </dt><dd>ccctqtcatt ttatqqcagq qtqttttcqt tatttaa 101 </dd></dl>
<dl><dt>15 </dt><dd><210> 162 <211> 108 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt>20 </dt><dd><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic polynucleotide " </dd></dl>
<dl><dt><220> <221> modified base <222> (1) .. (20) <223> a, c, t, g, Unknown or airo </dt><dd /></dl>
<dl><dt>25 </dt><dd><400> 162 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcaagatt tagaaataaa tcttgcagaa 60 </dd></dl>
<dl><dt>qctacaaaqa taa; qcttca </dt><dd>tqccqaaatc aacaccctqt cattttat loa </dd></dl>
<dl><dt>30 </dt><dd><210> 163 <211> 86 <212> DNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic oligonucleotide M </dt><dd /></dl>
<dl><dt>35 </dt><dd><220> <221> modified base <222> (1) .. (20) <223> a, c, t, g, Unknown or other </dd></dl>
<dl><dt><400> 163 nnnnnnnnnn </dt><dd>nnnnnnnnnn gtttttgtac tctcaagatt tagaaataaa tcttgcaqaa 60 </dd></dl>
<dl><dt>qctacaaaqa taaqqcttca tqccqa </dt><dd> 86 </dd></dl>
<dl><dt>40 </dt><dd><210> 164 <211> 79 <212> DNA <213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Synthetic Oligonucleotide Artificial Sequence ~ </dd></dl>
<220>
<dl><dt>< </dt><dd>221> modified base </dd></dl>
<dl><dt>< </dt><dd>222> (1) "" (20) </dd></dl>
<dl><dt>< </dt><dd>223> a, c, t, g, Unknown or olro </dd></dl>
<400> 164
nnnnnnnnnn nnnnnnnnnn
qctacaaaqa taaqqctte
10 <210> 165
<211> 73
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220> 15 <221> source
qtttttqtac t ctcaaqatt taqaaataaa tcttqcaqaa 60 19
<223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide ~
<220>
<221> modified base
<222> (1) .. (20) 20 <223> a, c, t, g, Unknown or other
<400> 165
nnnnnnnnnn nnnnnnnnnn qtttttgtae teteaaqatt taqaaataaa tettqeaqaa 60
qcteeaaaga taa 13
<210> 166
<211> 125 25 <212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide " </dd></dl>
30 <400> 166
quuuuaqucc cuuuuuaaau uucuuuaugg uaaaauuaua aucucauaaq aaauuuaaaa 60
aqgqacuaaa auaaaqaquu ugcgqqacuc uqcqqqquua caaucceeua aaaeeqcuuu 120
uaaaa 125
<210> 167 35 <211> 91
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<221> source 40 <223> Inota = ~ Description of Artificial Sequence "Synthetic oligonucleotide ~
<400> 167
quuacuuaaa ucuuqcaqaa qcuacaaaqa uaaqqcuuca uqccgaaauc aacaeecugu 60
cauuuuauqq caqqququuu ucquuauuua a 91
<dl><dt><210> 168 <211> 56 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic oligonucleotide " </dd></dl>
<dl><dt><400> 168 gggacucaac caagucauuc guuuuuguac ucucaagauu uaaguaacug uacaac 56 </dt><dd /></dl>
<dl><dt>10 </dt><dd><210> 169 <211> 147 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt>15 </dt><dd><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic polynucleotide " </dd></dl>
<dl><dt><400> 169 qgqacucaac </dt><dd>caaqucauuc quuuuuqu ~ c ucucaaqa uu uaaqua acuq uacaa eguua 60 </dd></dl>
<dl><dt>cuuaaaue uu </dt><dd>geaqaaqcua caaaqauaag qc uucaugcc gaaaucaaea cccuqucauu 120 </dd></dl>
<dl><dt>uuauqqcagg ququuuucqu </dt><dd>u still 147 </dd></dl>
<dl><dt>20 </dt><dd><210> 170 <211> 70 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt>25 </dt><dd><220> <221> source <223> Inota ::: Artificial Sequence MD "Ol Synthetic Igonucleotide" </dd></dl>
<dl><dt><400> 170 cuuqeaqaag cuacaaagau </dt><dd>aqgcuuea u gccgaaauca to cacccuquc uggc auuuua 60 </dd></dl>
<dl><dt>aqqququuuu </dt><dd> 70 </dd></dl>
<dl><dt>30 </dt><dd><210> 171 <211> 42 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic oligonucleotide " </dt><dd /></dl>
<dl><dt>35 </dt><dd><400> 171 gggacucaac caagucauuc guuuuuguac ucucaagauu ua 42 </dd></dl>
<dl><dt>40 </dt><dd><210> 172 <211> 112 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota ::: Artificial Sequence MD "Synthetic polynucleotide" 169 </dt><dd /></dl>
<400> 172
gggacucaac caagucauuc guuuuuguac ucucaa9auu uacuugcaga & gcuacaaag
auaaqqcuuc augccgaaau caacacccug ucauuuuauq gcaqgguquu uu
<210> 173
<21 1> 116
<212> RNA
<213> Artificial Sequence
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota :: ~ Description of Artificial Sequence "Synthetic polynucleotide" </dd></dl>
<400> 173
gggacucaac caaqucauuc quuuuuquac ucucaaqauu ua9aaacuuq ca9aaqcuac
aaaqauaago cuucauqccg aaaucaacae eeugueauull uauqqeaqqq uguuuu
<210> 174
<211> 116
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic polynucleotide" </dd></dl>
<400> 174
ggqacucaac caaqucauuc quuuuuquac ucueaagauu uagaaacuuq cagaaqcuac
aaagauaagg cuucaugccg aaaucaacac ccugucauuu UBuqgcaggg uguuuu
<210> 175 <211>102
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Description of Artificial Sequence "Synthetic polynucleotide" </dd></dl>
<400> 175
99gacucaac caaqucauue guuuuuquaq aaauaeaaa; auaaggcuuc augcegaaau
caacacecug ucauuuuaug gcagQ9Uguu uucquuauuu aa
<210> 176
<211> 102
<dl><dt>< </dt><dd>212> RNA </dd></dl>
<dl><dt>< </dt><dd>213> Artificial Sequence </dd></dl>
<220>
<dl><dt>< </dt><dd>221> source </dd></dl>
<dl><dt>< </dt><dd>223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide " </dd></dl>
<400> 176
gggacucaac caaqucauuc quuuuuguag aaauacaaag auaaggcuuc augccqaaau
caacaceeug ucauuuuauq qcaqqguguu uucquuauuu aa
60 112
60 116
60 102
60 102
<dl><dt><210> 177 <211> 57 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic oligonucleotide " </dt><dd /></dl>
<dl><dt><400> 177 gggacucaac caagucauuc guuuuuguag aaauacaaag auaaggcuuc augccga </dt><dd> 57 </dd></dl>
<dl><dt><210> 178 <211> 57 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota ::: Artificial Sequence MD: Synthetic oligonucleotide " </dt><dd /></dl>
<dl><dt><400> 178 gggacucaac caagucauuc guuuuuguag aaauacaaag auaaggcuuc augccga </dt><dd> 57 </dd></dl>
<dl><dt><210> 179 <211> 23 <212> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota :: Artificial Sequence MD "Synthetic oligonucleotide" </dt><dd /></dl>
<dl><dt><400> 179 gtggtgtcac gctcgtcgtt t99 </dt><dd> 23 </dd></dl>
<dl><dt><210> 180 <211> 23 <21 2> DNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inota ::: Synthetic Oligonucleotide Artificial Sequence MD " </dt><dd /></dl>
<dl><dt><400> 180 tccagtctat taattgllgc cgg </dt><dd> 23 </dd></dl>
<dl><dt><210> 181 <211> 64 <212> DNA <213> Furnace sapiens </dt><dd /></dl>
<dl><dt><400> 181 caagaggett gagtaggaga gqagtgecgc cgagqcgggg eggggcgggg cqtggagctg </dt><dd> 60 </dd></dl>
<dl><dt>gqct </dt><dd> 64 </dd></dl>
<dl><dt><210> 182 <211> 99 </dt><dd /></dl>
<dl><dt>171 </dt><dd /></dl>
<dl><dt><212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt>5 </dt><dd><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic oligonucleotide " </dd></dl>
<dl><dt><220> <221> modified base <222> (1) .. (20) <223> a, c, u, g, Unknown or other </dt><dd /></dl>
<dl><dt>10 </dt><dd><400> 182 nnnnnnnnnn nnnnnnnnnn quauuaqaqc uaqaaauagc aaquuaauau aaggcuaquc 60 </dd></dl>
<dl><dt>cguuaucaac </dt><dd>uugaaaaaqu ggcaccqaqu cgguqcuuu 99 </dd></dl>
<dl><dt>15 </dt><dd><210> 183 <211> 119 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt><220> <221> source <223> Inota = ~ Artificial Sequence Description: Synthetic polynucleotide " </dt><dd /></dl>
<dl><dt>20 </dt><dd><220> <221> modified base <222> (1) .. {20) <223> a, c, u, g, Unknown or other </dd></dl>
<dl><dt><400> 183 nnnnnnnnnn </dt><dd>nnnnnnnnnn quuuuagagc uaugcuquuu ugqaaacaaa acagcauaqc 60 </dd></dl>
<dl><dt>aaquuaaaau </dt><dd>aaqqcuaquc cquuaucaac uugaaaaaqu qqcaccqagu cqquqcuuu 119 </dd></dl>
<dl><dt>25 </dt><dd><210> 184 <211> 119 <212> RNA <213> Artificial Sequence </dd></dl>
<dl><dt>30 </dt><dd><220> <221> source <223> Inota = ~ Description of Artificial Sequence · Synthetic polynucleotide " </dd></dl>
<dl><dt>35 </dt><dd><220> <221> modified base <222> (1) (20) <223> a, c, u, g, Unknown or other </dd></dl>
<dl><dt><400> 184 nnnnnnnnnn </dt><dd>nnnnnnnnnn gu & uuaqaqc uauqcuquau uqqaaacaau acagcauagc 60 </dd></dl>
<dl><dt>aaquuaauau </dt><dd>aagqcuaquc cquuaucaac uuqaaaaaqu qqcaccgaqu cqquqcuuu 119 </dd></dl>
<dl><dt>40 </dt><dd><210> 185 <211> 12 <212> DNA <213> Homo sapiens </dd></dl>
<dl><dt><400> 185 tagcgggtaa gc </dt><dd> 12 172 </dd></dl>
<210> 186
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 186 tcggtgacat 9t 12
<210> 187 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 187 actccccgta 99 12
<210> 188 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 188 aclgcglgtl aa 12
<210> 189
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 189 acgtcgcclg at 12
<210> 190
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 190 tagglcgacc ag 12
<210> 191 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 191 ggcgttaalg at 12
<210> 192 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 192 Igtcgcatgl ta 12
<210> 193 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 193 atggaaaege at 12
<210> 194 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 194 gecgaattee te 12
<210> 195 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 195 geatggtaeg ga 12
<210> 196
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 196 cggtaetett ae 12
<210> 197
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 197 geetgtgeeg ta 12
<210> 198
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 198 taeggtaagt eg 12
<210> 199
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 199 cacgaaatta ce 12
<210> 200
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 200 aaccaagata cg 12
<210> 201
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 201 gaglcgalac gc 12
<210> 202 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 202 glclcacgal cg 12
<210> 203 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 203 tcgtcgggtg ca 12
<210> 204
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 204 aclccgtagl ga 12
<210> 205
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 205 caggacgtcc gt 12
<210> 206 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 206 tcglatccct ac 12
<210> 207 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 207 tltcaaggcc gg 12
<210> 208 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 208 cgccgglgga at 12
<210> 209 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 209 gaacccglcc la 12
<210> 210 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 210 gattcalcag eg 12
<210>211
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 211 aeaeeggtet le 12
<210> 212
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 212 ateglgccet aa 12
<210> 213
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 213 gegtcaatgt te 12
<210> 214
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 214 ctecgtatel cg 12
<210> 215
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 215 eegattcctt cg 12
<210> 216
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 216 Igcgcctcca gl 12
<210> 217 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 217 laacglcgga gc 12
<210> 218 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 218 aagglcgccc at 12
<210> 219
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 219 glcggggact at 12
<210> 220
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 220 ttcgagcgattt 12
<210> 221 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 221 tgagtcgtcg ag 12
<210> 222 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 222 tltacgcaga gg 12
<210> 223 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 223 aggaagtatc gc 12
<210> 224 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 224 actcgatacc to 12
<210> 225 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 225 cgclacalag ca 12
<210> 226
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 226 ttcataaccg gc 12
<210> 227
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 227 ccaaacggtt aa 12
<210> 228
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 228 cgattccttc gt 12
<210> 229
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 229 cglcatgaal aa 12
<210> 230
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 230 agtggcgatg ac 12
<210> 231
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 231 eecctacgge ae 12
<210> 232 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 232 gecaaccege ae 12
<210> 233 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 233 Igggaeaecg gl 12
<210> 234
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 234 ttgaetgcgg eg 12
<210> 235
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 235 aetalgcgta 99 12
<210> 236 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 236 lcaeccaaag cg 12
<210> 237 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 237 geaggacgle eg 12
<210> 238 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 238 acaccgaaaa cg 12
<210> 239 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 239 cgglgtattg ag 12
<210> 240 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 240 cacgaggtat gc 12
<210> 241
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 241 taaagcgacc cg 12
<210> 242
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 242 cttagtcggc ca 12
<210> 243
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 243 cgaaaacglg gc 12
<210> 244
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 244 cglgccclga ac 12
<210> 245
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 245 tttaccatcg aa 12
<210> 246
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 246 cgtagccalg ti 12
<210> 247 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 247 cccaaacggl ta 12
<210> 248 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 248 gcgltalcag aa 12
<210> 249
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 249 Icgatgglaa ac 12
<210> 250
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 250 cgactlttlg ca 12
<210> 251 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 251 Icgacgaclc ac 12
<210> 252 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 252 acgcgtcaga la 12
<210> 253 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 253 cgtacggcac ag 12
<210> 254 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 254 ctatgccgtg ca 12
<210> 255 <211>12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 255 cgcgtcagat at 12
<210> 256
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 256 12 aagatcggta gc 12
<210> 257
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 257 cttcgcaagg ag 12
<210> 258
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 258 gtcgtggact ac 12
<210> 259
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 259 gglcgtcatc aa 12
<210> 260
<211> 12
<dl><dt>< </dt><dd>212> DNA </dd></dl>
<dl><dt>< </dt><dd>213> Sapiens oven </dd></dl>
<400> 260 gttaacagcg tg 12
<dl><dt><210> 261 <211> 12 <212> DNA <213> Furnace sapiens </dt><dd /></dl>
<dl><dt><400> 261 tagctaaccg ti </dt><dd> 12 </dd></dl>
<dl><dt><210> 262 <211> 12 <212> DNA <213> Furnace sapiens </dt><dd /></dl>
<dl><dt><400> 262 aglaaaggcg the </dt><dd> 12 </dd></dl>
<dl><dt><210> 263 <211> 12 <212> DNA <213> Furnace sapiens </dt><dd /></dl>
<dl><dt><400> 263 gglaaltlcg 19 </dt><dd> 12 </dd></dl>
<dl><dt><210> 264 <211> 147 <212> RNA <213> Artificial Sequence </dt><dd /></dl>
<dl><dt><220> <221> source <223> Inola = ~ Description of Artificial Sequence · Synthetic polynucleotide " </dt><dd /></dl>
<dl><dt><220> <221> modified base <222> (1) .. (20) <223> a, e, u, g, Unknown or other </dt><dd /></dl>
<dl><dt><400> 264 nnnnnnnnnn </dt><dd>nnnnnnnnnn quuuuaquac ucuquaauuu uagquaugag quagacgaaa 60 </dd></dl>
<dl><dt>auuquacuua </dt><dd>uaccuaaaau uacagaaucu acuaaaacaa 9gcaaaau9c cququuuauc 120 </dd></dl>
<dl><dt>uc quC & acuu </dt><dd>guugqcgaqa uuuuuUu '47</dd></dl>
Contents50
44 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44
601 members in 19 offices
Priority claims24
| Document | Office | Kind | Date |
|---|---|---|---|
| 201261736527P | United States of America | – | |
| 201261736527 | United States of America | P | |
| 201361748427P | United States of America | – | |
| 201361748427 | United States of America | P | |
| 201361758468P | United States of America | – | |
| 201361758468 | United States of America | P | |
| 201361769046P | United States of America | – | |
| 201361769046 | United States of America | P | |
| 201361791409P | United States of America | – | |
| 201361802174P | United States of America | – | |
| 201361791409 | United States of America | P | |
| 201361802174 | United States of America | P | |
| 201361806375P | United States of America | – | |
| 201361806375 | United States of America | P | |
| 201361814263P | United States of America | – | |
| 201361814263 | United States of America | P | |
| 201361819803P | United States of America | – | |
| 201361819803 | United States of America | P | |
| 201361828130P | United States of America | – | |
| 201361828130 | United States of America | P | |
| 201361835931P | United States of America | – | |
| 201361836127P | United States of America | – | |
| 201361835931 | United States of America | P | |
| 201361836127 | United States of America | P |
Members601
| Document | Office | Kind | |
|---|---|---|---|
| US8697359B1 | United States of America | B1 | |
| CA2894668A1 | Canada | A1 | |
| CA2894681A1 | Canada | A1 | |
| CA2894684A1 | Canada | A1 | |
| CA2894688A1 | Canada | A1 | |
| CA2894701A1 | Canada | A1 | |
| US2014170753A1 | United States of America | A1 | |
| WO2014093595A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014093622A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2014093635A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014093655A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2014093661A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2014093694A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014093701A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014093709A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014093712A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014093718A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2014179006A1 | United States of America | A1 | |
| US2014179770A1 | United States of America | A1 | |
| US2014186843A1 | United States of America | A1 | |
| US2014186919A1 | United States of America | A1 | |
| US2014186958A1 | United States of America | A1 | |
| US2014189896A1 | United States of America | A1 | |
| US8771945B1 | United States of America | B1 | |
| WO2014093635A9 | World Intellectual Property Organization (WIPO) | A9 | |
| WO2014093655A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US8795965B2 | United States of America | B2 | |
| WO2014093622A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2014093661A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP2764103A2 | European Patent Office (EPO) | A2 | |
| US2014227787A1 | United States of America | A1 | |
| US2014234972A1 | United States of America | A1 | |
| WO2014093712A9 | World Intellectual Property Organization (WIPO) | A9 | |
| US2014242664A1 | United States of America | A1 | |
| US2014242699A1 | United States of America | A1 | |
| US2014242700A1 | United States of America | A1 | |
| EP2771468A1 | European Patent Office (EPO) | A1 | |
| US2014248702A1 | United States of America | A1 | |
| US2014256046A1 | United States of America | A1 | |
| US2014273231A1 | United States of America | A1 | |
| US2014273232A1 | United States of America | A1 | |
| US2014273234A1 | United States of America | A1 | |
| EP2784162A1 | European Patent Office (EPO) | A1 | |
| US2014310830A1 | United States of America | A1 | |
| WO2014093661A9 | World Intellectual Property Organization (WIPO) | A9 | |
| WO2014093694A9 | World Intellectual Property Organization (WIPO) | A9 | |
| US8865406B2 | United States of America | B2 | |
| US8871445B2 | United States of America | B2 | |
| WO2014093622A9 | World Intellectual Property Organization (WIPO) | A9 | |
| WO2014093701A9 | World Intellectual Property Organization (WIPO) | A9 | |
| US2014335620A1 | United States of America | A1 | |
| US8889356B2 | United States of America | B2 | |
| US8889418B2 | United States of America | B2 | |
| US8895308B1 | United States of America | B1 | |
| US2014357530A1 | United States of America | A1 | |
| WO2014093622A8 | World Intellectual Property Organization (WIPO) | A8 | |
| WO2014093655A9 | World Intellectual Property Organization (WIPO) | A9 | |
| US8906616B2 | United States of America | B2 | |
| CA2915795A1 | Canada | A1 | |
| CA2915834A1 | Canada | A1 | |
| CA2915837A1 | Canada | A1 | |
| CA2915842A1 | Canada | A1 | |
| CA2915845A1 | Canada | A1 | |
| WO2014204723A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014204724A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014204725A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014204726A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014204727A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014204728A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014204729A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US8932814B2 | United States of America | B2 | |
| US2015020223A1 | United States of America | A1 | |
| EP2825654A1 | European Patent Office (EPO) | A1 | |
| US2015031134A1 | United States of America | A1 | |
| US8945839B2 | United States of America | B2 | |
| EP2771468B1 | European Patent Office (EPO) | B1 | |
| EP2840140A1 | European Patent Office (EPO) | A1 | |
| EP2840140A1 | European Patent Office (EPO) | A1 | |
| WO2014204727A9 | World Intellectual Property Organization (WIPO) | A9 | |
| EP2848690A1 | European Patent Office (EPO) | A1 | |
| US2015079681A1 | United States of America | A1 | |
| US8993233B2 | United States of America | B2 | |
| US8999641B2 | United States of America | B2 | |
| EP2784162B1 | European Patent Office (EPO) | B1 | |
| ES2536353T3 | Spain | T3 | |
| DK2771468T3 | Denmark | T3 | |
| WO2014204725A8 | World Intellectual Property Organization (WIPO) | A8 | |
| PT2771468E | Portugal | E | |
| US2015184139A1 | United States of America | A1 | |
| WO2014204728A8 | World Intellectual Property Organization (WIPO) | A8 | |
| DK2784162T3 | Denmark | T3 | |
| EP2896697A1 | European Patent Office (EPO) | A1 | |
| US2015203872A1 | United States of America | A1 | |
| EP2898075A1 | European Patent Office (EPO) | A1 | |
| ES2542015T3This record | Spain | T3 | |
| AU2013359123A1 | Australia | A1 | |
| AU2013359199A1 | Australia | A1 | |
| AU2013359212A1 | Australia | A1 | |
| AU2013359238A1 | Australia | A1 | |
| AU2013359262A1 | Australia | A1 |
Numbers
- Publication
- 2542015
- Application
- 14170383
Titles2
- Spanish
- Ingeniería de sistemas, métodos y composiciones de guía optimizadas para manipulación de secuencias
- English
- Systems engineering, methods and guide compositions optimized for sequence manipulation
Classification
- CPC, 30
- C12N9/16
- C12N9/22
- A61K48/00
- C12N15/1082
- C12N15/63
- C12N15/79
- C12N15/52
- C12Y301/00
- C12N15/01
- C12N15/102
- C12N15/113
- C12N15/907
- C12N2310/10
- C12N2310/20
- C12N2320/11
- C12N2320/30
- C12N2750/14143
- G16B20/00
- G16B20/20
- G16B20/30
- G16B20/50
- G16B30/00
- G16B30/10
- A61K48/005
- C12N15/902
- C12N2800/10
- C12N15/86
- C12N15/85
- C12N2810/50
- C12Q1/6806
- IPC, 1
- C12N15 63