Campylobacter glycosyltransferases for biosynthesis of gangliosides and ganglioside mimics
Abstract
An isolated or recombinant nucleic acid molecule comprising a polynucleotide sequence that encodes a polypeptide selected from the group consisting of: a) a polypeptide having acyltransferase activity for the biosynthesis of lipid A, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 350-1234 (ORF 2a) of the LOS biosynthesis locus of strain OH4384 of C. jejouni as shown in SEQ ID NO: 1; b) a polypeptide having the activity of glycosyltransferase, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 1234-2487 (ORF 3a) of the locus for biosynthesis of the LOS of strain OH4384 of C. jejouni as shown in SEQ ID NO: 1; c) a polypeptide having the activity of glycosyltransferase, wherein the polypeptide contains an amino acid sequence that is approximately at least 50% identical to an amino acid sequence encoded by nucleotides 2786-3952 (ORF 4a) of the locus for LOS biosynthesis of strain OH4384 of C. jejouni as shown in SEQ ID NO: 1 over a region of at least 100 amino acids in length; d) a polypeptide having the activity of a1,4-GalNAc transferase, wherein the GalNAc transferase polypeptide has an amino acid sequence that is approximately at least 77% identical to an amino acid sequence as set forth in SEQ ID NO: 13 over a region of at least 50 amino acids in length.

Term
Term ended
Projected expiry passed 1 February 2020, 6.6 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
39 claims: 3 independent, 36 dependent
- 1ES 2 269 098 T3 REIVINDICACIONES 1. Una molécula aislada o recombinante de ácido nucleico que comprende una secuencia de polinucleótidos que codifica a un polipéptido seleccionado del grupo que consiste de:a) un polipéptido que tiene la actividad de aciltransferasa para la biosíntesis del lípido A, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 350-1234 (ORF 2a) del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;b) un polipéptido que tiene la actividad de glicosiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 1234-2487 (ORF 3a) del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;c) un polipéptido que tiene la actividad de glicosiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 50% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 2786-3952 (ORF 4a) del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1 sobre una región de al menos 100 aminoácidos de longitud;d) un polipéptido que tiene la actividad de la β 1,4-GalNAc transferasa, en donde el polipéptido de la GalNAc transferasa tiene una secuencia de aminoácidos que es aproximadamente al menos 77% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 13 sobre una región de al menos 50 aminoácidos de longitud;e) un polipéptido que tiene la actividad de la β 1,3-Galactosiltransferasa, en donde el polipéptido de la Galactosiltransferasa tiene una secuencia de aminoácidos que es aproximadamente al menos 75% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 15 o SEQ ID NO: 17 sobre una región de al menos 50 aminoácidos de longitud;f) un polipéptido que tiene ya sea la actividad de la a2,3-sialiltransferasa o tanto a la actividad de la α2,3sialiltransferasa como de la a2,8-sialiltransferasa, en donde el polipéptido de la Galactosiltransferasa tiene una secuencia de aminoácidos que es aproximadamente al menos 66% idéntica sobre una región de al menos 60 aminoácidos de longitud hasta una secuencia de aminoácidos como la expuesta en una o más de las SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 7 o SEQ ID NO: 10;g) un polipéptido que tiene la actividad de la ácido siálico sintasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 6924-7961 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;h) un polipéptido que tiene la actividad de biosíntesis del ácido siálico, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 8021-9076 del locus para la biosíntesis del LOS de OH4384 de la cepa de C. jejouni como se muestra en la SEQ ID NO: 1;i) un polipéptido que tiene la actividad del ácido siálico-CMP sintetasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 65% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 9076-9738 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;j) un polipéptido que tiene la actividad de la acetiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 65% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 9729-10559 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;y k) un polipéptido que tiene la actividad de la glicosiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 65% idéntica a una secuencia de aminoácidos codificada por medio de un complemento inverso de los nucleótidos 10557-11366 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1.
- 2La molécula aislada o recombinante de ácido nucleico de la reivindicación 1, en donde el ácido nucleico comprende una secuencia de polinucleótidos que codifica a uno o más polipéptidos seleccionados del grupo que consiste de:a) un polipéptido de sialiltransferasa que tiene tanto actividad de una a2,3-sialiltransferasa como actividad de una a2,8-sialiltransferasa, en donde el polipéptido de la sialiltransferasa comprende una secuencia de ES 2 269 098 T3 aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos como la expuestas en la SEQ ID NO: 3 sobre una región de al menos 50 aminoácidos se longitud;b) un polipéptido de GalNAc transferasa que tiene actividad de /11,4-GalNAc transferasa, en donde el polipéptido de la GalNAc transferasa comprende una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos como la expuestas en la SEQ ID NO: 13 sobre una región de al menos 50 aminoácidos se longitud;y c) un polipéptido de galactosiltransferasa que tiene actividad de,61,3-galactosiltransferasa, en donde el polipéptido de la galactosiltransferasa comprende una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos como la expuestas en la SEQ ID NO: 15 sobre una región de al menos 50 aminoácidos se longitud.
- 3La molécula de ácido nucleico de la reivindicación 1, en donde las comparaciones de las secuencias se llevan a cabo utilizando un algoritmo BLASTP versión 2.0 con una longitud de palabra (W) de 3, G=11, E=1, y una matriz de sustitución BLOSUM62.
- 4La molécula de ácido nucleico de la reivindicación 1, en donde la región extiende la longitud total de la secuencia de aminoácidos del polipéptido.
- 5La molécula de ácido nucleico de la reivindicación 1, en donde:a) el polipéptido de la sialiltransferasa comprende una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 3, la SEQ ID NO: 5, la SEQ ID NO: 7 o la SEQ ID NO: 10;b) el polipéptido de la GalNAc transferasa comprende una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 13;y c) el polipéptido de la galactosiltransferasa comprende una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 15 o la SEQ ID NO: 17.
- 6La molécula de ácido nucleico como la de la reivindicación 5, en donde:a) la secuencia del polinucleótido que codifica al polipéptido de la sialiltransferasa es al menos 75% idéntica a una secuencia de ácido nucleico como la expuesta en la SEQ ID NO: 2, la SEQ ID NO: 4, o la SEQ ID NO: 6, sobre una región de al menos 50 nucleótidos de longitud;b) la secuencia del polinucleótido que codifica al polipéptido de la β 1,4-GalNAc transferasa es aproximadamente al menos 75% idéntica a una secuencia de ácido nucleico como la expuesta en la SEQ ID NO: 12, sobre una región aproximadamente de al menos 50 nucleótidos de longitud;y c) la secuencia del polinucleótido que codifica al polipéptido de la β 1,3-galactosiltransferasa es aproximadamente al menos 75% idéntica a una secuencia de ácido nucleico como la expuesta en la SEQ ID NO: 14 o en la SEQ ID NO: 16, sobre una región aproximadamente de al menos 50 nucleótidos de longitud.
- 7La molécula de ácido nucleico de la reivindicación 6, en donde las comparaciones de las secuencias se llevan a cabo utilizando un algoritmo BLASTN versión 2.0 con una longitud de palabra (W) de 1, G=5, E=2, q=-2, y r=1.
- 8La molécula de ácido nucleico de la reivindicación 6, en donde:a) la secuencia del polinucleótido que codifica al polipéptido de la sialiltransferasa tiene una secuencia de ácido nucleico como la expuesta en la SEQ ID NO: 2, la SEQ ID NO: 4, o la SEQ ID NO: 6;b) la secuencia del polinucleótido que codifica al polipéptido de la GalNAc transferasa tiene una secuencia de ácido nucleico como la expuesta en la SEQ ID NO: 12;y c) la secuencia del polinucleótido que codifica al polipéptido de la galactosiltransferasa tiene una secuencia de ácido nucleico como la expuesta en la SEQ ID NO: 14 o en la SEQ ID NO: 16.
- 9La molécula de ácido nucleico de la reivindicación 5, en donde la sialiltransferasa es una sialiltransferasa bifuncional que tiene tanto actividad de a2,3-sialiltransferasa como actividad de a2,8-sialiltransferasa y la secuencia del polinucleótido que codifica al polipéptido de la sialiltransferasa es al menos 75% idéntica a una secuencia de ácido nucleico como la expuesta en la SEQ ID NO:2 o la SEQ ID NO: 4.
- 10Un casete de expresión que comprende una molécula de ácido nucleico de la reivindicación 1.
- 11Un vector de expresión que comprende al casete de expresión de la reivindicación 10. ES 2 269 098 T3
- 12Una célula huésped que comprende al vector de expresión de la reivindicación 11.
- 13Un polipéptido aislado o producido en forma recombinante seleccionado del grupo que consiste de:a) un polipéptido que tiene la actividad de aciltransferasa para la biosíntesis del lípido A, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 350-1234 (ORF 2a) del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;b) un polipéptido que tiene la actividad de glicosiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 1234-2487 (ORF 3a) del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;c) un polipéptido que tiene la actividad de glicosiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 50% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 2786-3952 (ORF 4a) del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1 sobre una región de al menos 100 aminoácidos de longitud;d) un polipéptido que tiene la actividad de la /11,4-GalNAc transferasa, en donde el polipéptido de la GalNAc transferasa tiene una secuencia de aminoácidos que es aproximadamente al menos 77% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 13 sobre una región de al menos 50 aminoácidos de longitud;e) un polipéptido que tiene la actividad de la β 1,3-Galactosiltransferasa, en donde el polipéptido de la Galactosiltransferasa tiene una secuencia de aminoácidos que es aproximadamente al menos 75% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 15 o SEQ ID NO: 17 sobre una región de al menos 50 aminoácidos de longitud;f) un polipéptido que tiene actividad de a2,3-sialiltransferasa, en donde el polipéptido de la sialiltransferasa tiene una secuencia de aminoácidos que es aproximadamente al menos 66% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 3, la SEQ ID NO: 5, la SEQ ID NO: 7 o la SEQ ID NO: 10, sobre una región de al menos 60 aminoácidos de longitud;g) un polipéptido que tiene la actividad de la ácido siálico sintasa, en donde el polipéptido contiene una secuencia de aminoácidos que es al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 6924-7961 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;h) un polipéptido que tiene actividad para la biosíntesis del ácido siálico, en donde el polipéptido contiene una secuencia de aminoácidos que es al menos 70% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 8021-9076 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;i) un polipéptido que tiene actividad para el ácido siálico-CMP sintetasa, en donde el polipéptido contiene una secuencia de aminoácidos que es aproximadamente al menos 65% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 9076-9738 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;j) un polipéptido que tiene actividad de la acetiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es al menos 65% idéntica a una secuencia de aminoácidos codificada por los nucleótidos 9729-10559 del locus para la biosíntesis del LOS de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1;y k) un polipéptido que tiene la actividad de la glicosiltransferasa, en donde el polipéptido contiene una secuencia de aminoácidos que es al menos 65% idéntica a una secuencia de aminoácidos codificada por medio de un complemento inverso de los nucleótidos 10557-11366 del locus para la biosíntesis del lOs de la cepa OH4384 de C. jejouni como se muestra en la SEQ ID NO: 1.
- 14El polipéptido aislado o producido en forma recombinante de la reivindicación 13, en donde el polipéptido es producido en forma recombinante y al menos parcialmente purificado.
- 15El polipéptido aislado o producido en forma recombinante de la reivindicación 13, en donde el polipéptido se expresa por medio de una célula huésped heteróloga.
- 16El polipéptido aislado o producido en forma recombinante de la reivindicación 15, en donde la célula huésped es E. coli. ES 2 269 098 T3
- 17El polipéptido aislado o producido en forma recombinante de la reivindicación 13, en donde el polipéptido es un polipéptido del serotipo O:2 de C. jejuni.
- 18El polipéptido aislado o producido en forma recombinante de la reivindicación 13, en donde el polipéptido es un polipéptido de la sialiltransferasa de acuerdo a g) y el polipéptido se selecciona del grupo que consiste de:un polipéptido que tiene tanto actividad de a2,3 sialiltransferasa como actividad de a2,8 sialiltransferasa y comprende una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos de una sialiltransferasa cstII codificada por ORF 7a del locus para la biosíntesis de LOS de la cepa OH4384 de C. Jejuni como la expuesta en la SEQ ID NO: 3;un polipéptido que tiene una actividad de a2,3 sialiltransferasa y comprende una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos de una sialiltransferasa cstII del serotipo O:10 de C. jejuni como la expuesta en la SEQ ID NO: 5;un polipéptido que tiene actividad de a2,3 sialiltransferasa y comprende una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos de una sialiltransferasa cstII del serotipo O:41 de C. Jejuni como la expuesta en la SEQ ID NO: 7;y un polipéptido que tiene actividad de a2,3 sialiltransferasa y comprende una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos de una sialiltransferasa cstII del serotipo O:2 de C. Jejuni como la expuesta en la SEQ ID NO: 10.
- 19El polipéptido aislado o producido en forma recombinante de la sialiltransferasa de la reivindicación 18, en donde el polipéptido de la sialiltransferasa tiene una secuencia de aminoácidos seleccionada del grupo que consiste de la SEQ ID NO:3, la SEQ ID NO: 5, la SEQ ID NO: 7, y la SEQ ID NO: 10.
- 20El polipéptido de la reivindicación 13, en donde:a) el polipéptido de la sialiltransferasa de f) tiene una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 3, la SEQ ID NO: 5, la SEQ ID NO: 7, o la SEQ ID NO: 10;b) el polipéptido de la //1,4-GalNAc transferasa de d) tiene una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 13;y c) el polipéptido de la//1,3-galaclosillransferasa de e) tiene una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 15 o la SEQ ID NO: 17.
- 21Una mezcla de reacción para la síntesis de un oligosacárido sialilado, la mezcla de reacción comprendiendo un péptido de la sialiltransferasa que tiene una secuencia de aminoácidos que es aproximadamente al menos 75% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO:3 sobre una región de aproximadamente al menos 50 aminoácidos de longitud, el polipéptido de la sialiltransferasa teniendo tanto actividad de la a2,3 sialiltransferasa como actividad de la a2,8 sialiltransferasa;una fracción aceptora galactosilada;y un azúcar sialil-nucleótido en donde la sialiltransferasa transfiere un primer residuo de ácido siálico desde el azúcar sialil-nucleótido hasta la fracción aceptora galactosilada en un enlace a2,3, y además transfiere un segundo residuo de ácido siálico hasta el primer residuo de ácido siálico en un enlace a2,8.
- 22La mezcla de reacción de la reivindicación 21, en donde el azúcar sialil-nucleótido es ácido siálico-CMP.
- 23La mezcla de reacción de la reivindicación 21, en donde el polipéptido de la sialiltransferasa tiene una secuencia de aminoácidos como la expuesta en la SEQ ID NO:3.
- 24La mezcla de reacción de la reivindicación 21, en donde el aceptor galactosilado comprende a un compuesto que tiene la fórmula Gal/1,4-R o Gal//1,3-R, en donde R se selecciona del grupo que consiste de H, un sacárido, un oligosacárido, o un grupo aglicón que tiene al menos un átomo de carbohidrato.
- 25La mezcla de reacción de la reivindicación 21, en donde el aceptor galactosilado se une a una proteína, lípido, o proteoglicano.
- 26La mezcla de reacción de la reivindicación 21, en donde el oligosacárido sialilado es un gangliósido, una imitación de gangliósido, o una porción carbohidrato de un gangliósido.
- 27La mezcla de reacción de la reivindicación 21, en donde el oligosacárido sialilado es un lisogangliósido, una imitación de lisogangliósido, o una porción carbohidrato de un lisogangliósido.
- 28La mezcla de reacción de la reivindicación 26, en donde la fracción aceptora galactosilada comprende un compuesto que tiene una fórmula seleccionada del grupo que consiste de Gal4Glc-R 1 y Gal3GalNAc-R 2 , en donde R 1 ES 2 269 098 T3 se selecciona del grupo que consiste de ceramida o de otro glicolípido, y R 2 se selecciona del grupo que consiste de Gal4GlcCer, (Neu5Ac3)Gal4GlcCer, y (Neu5Ac8Neu5Ac3)Gal4GlcCer.
- 29La mezcla de reacción de la reivindicación 28, en donde el aceptor galactosilado se selecciona del grupo que consiste de Gal4GlcCer, Gal3GalNAc4(Neu5Ac3)Gal4GlcCer, y Gal3GalNAc4(Neu5Ac8Neu5Ac3)Gal4GlcCer.
- 30La mezcla de reacción de la reivindicación 21, en donde el aceptor galactosilado se forma poniendo en contacto un aceptor sacárido con UDP-Gal y un polipéptido de la galactosil transferasa, en donde el polipéptido de la galactosiltransferasa transfiere el residuo Gal desde el UDP-Gal hasta el aceptor.
- 31La mezcla de reacción de la reivindicación 30, en donde el polipéptido de la galactosiltransferasa tiene actividad de j61,3-galactosiltransferasa y tiene una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO:15 o en la SeQ ID NO: 17 sobre una región de al menos 50 aminoácidos de longitud.
- 32La mezcla de reacción de la reivindicación 31, en donde la galactosiltransferasa tiene una secuencia de aminoácidos como la expuesta en la SEQ ID NO:15 o en la SEQ ID NO: 17.
- 33La mezcla de reacción de la reivindicación 30, en donde el sacárido aceptor comprende un residuo GalNAc terminal.
- 34La mezcla de reacción de la reivindicación 33, en donde el sacárido aceptor para la galacitosiltransferasa se forma poniendo en contacto un aceptor para una GalNAc transferasa con UDP-GalNAc y un polipéptido de GalNAc transferasa, en donde el polipéptido de la GalNAc transferasa transfiere el residuo GalNAc desde el UDP-GalNAc hasta el aceptor para la GalNAc transferasa.
- 35La mezcla de reacción de la reivindicación 34, en donde el polipéptido de la GalNAc transferasa tiene actividad de /'1,4-GalNAc transferasa y tiene una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO:13 sobre una región de al menos 50 aminoácidos de longitud.
- 36La mezcla de reacción de la reivindicación 28, en donde el polipéptido de la GalNAc transferasa tiene una secuencia de aminoácidos como la expuesta en la SEQ ID NO:13. Un método para sintetizar un oligosacárido sialilado, el método comprendiendo la incubación bajo condiciones estables de una mezcla de reacción que comprende un polipéptido de la sialiltransferasa que tiene una secuencia de aminoácidos que es al menos 75% idéntica a una secuencia de aminoácidos como la expuesta en la SEQ ID NO: 3 sobre una región de al menos 50 aminoácidos de longitud, teniendo el polipéptido de la sialiltransferasa tanto actividad de la α2,3 sialiltransferasa como actividad de la α2,8 sialiltransferasa, una fracción aceptora galactosilada;y un azúcar sialil-nucleótido, en donde el polipéptido de la sialiltransferasa transfiere un primer residuo de ácido siálico desde el azúcar sialil-nucleótido hasta la fracción aceptora galactosilada en un enlace α2,3, y además transfiere un segundo residuo de ácido siálico hasta el primer residuo de ácido siálico en un enlace α2,8.
- 37El método de la reivindicación 37, en donde el oligosacárido sialilado es un gangliósido.
- 38El método de la reivindicación 38, en donde el polipéptido de la sialiltransferasa tiene una secuencia de aminoácidos como la expuesta en la SEQ ID NO:3.
- 39El método de la reivindicación 37, en donde el oligosacárido sialilado es un gangliósido, un lisogangliósido, una imitación de gangliósido, o una imitación de lisogangliósido.
Independent claims39
365 paragraphs in 36 sections, as filed
ES 2 269 098 T3
DESCRIPTION
Campylobacter glycosyl transferases for ganglioside biosynthesis and ganglioside mimics.
Background of the Invention Field of the Invention
This invention belongs to the field of enzymatic oligosaccharide synthesis, which includes gangliosides and ganglioside mimics.
Background
Gangliosides are a class of glycolipids often found in cell membranes, consisting of three elements. One or more salicylic acid residues that are attached to an oligosaccharide or a fraction of the nucleus of a carbohydrate, which in turn is attached to a hydrophobic lipid structure (ceramide) that is generally embedded in the cell membrane. The ceramide fraction includes the base portion of a long chain (LCB) and a fatty acid portion (FA). Gangliosides, as well as other glycolipids and their structures in general, are discussed in, for example, Lehninger, Biochemistry (Worth Publishers, 1981) pages 287-295 and Devlin, Textbook of Biochemistry (Wiley-Liss, 1992). Gangliosides are classified according to the number of monosaccharides in the carbohydrate fraction, as well as the number and location of sialic acid groups present in the carbohydrate fraction. Monosialogangliosides are designated "GM", disialogangliosides are designated "GD", trisialogangliosides are designated "GT", and tetrasialogangliosides are designated "GQ". Gangliosides can be further classified depending on the position (s) of the sialic acid residue or linked residues. Another classification is based on the number of saccharides present in the oligosaccharide nucleus, with the subscript "1" that designates a ganglioside that has four saccharide residues (Gal-Gal-NAc-Gal-Glc-Ceramide), and the subscripts "2" , "3" and "4" representing gangliosides trisaccharides (Gal-NAc-Gal-Glc-Ceramide), disaccharides (Gal-Glc-Ceramide) and monosaccharides (Gal-Ceramide), respectively.
Gangliosides are more abundant in the brain, particularly in the nerve endings. They are believed to be present at receptor sites for neurotransmitters, including acetylcholine, and can also act as specific receptors for other biological macromolecules, including interferon, hormones, viruses, bacterial toxins, and the like. Gangliosides have been used to treat disorders of the nervous system, including attacks of cerebral ischemia. See, for example, Mahadnik et al. (1988) Drug Development Res. 15: 337-360; US Patent Nos. 4,710,490 and 4,347,244; Horowitz (1988) Adv Exp. Med. And Biol. 174: 593-600; Karpiatz et al. (1984) Adv Exp. Med. And Biol. 174: 489-497. Certain gangliosides are found on the surface of human hematopoietic cells (Hildebrand et al. (1972), Biochim. Biophys. Acta 260: 272-278; Macher et al. (1981) J. Biol. Chem. 256: 1968-1974; Dacremont and co-workers, Biochim. Biophys. Acta 424: 315-322; Klock et al. (1981), Blood Cells 7: 247) that may play a role in the terminal granulocytic differentiation of these cells. Nojiri et al. (1988) J. Biol. Chem. 263: 7443-7446. These gangliosides, referred to as the "neolacto" series, have neutral nucleus oligosaccharide structures that have the formula [Gale- (1,4) GlcNAce (1,3)]<sub>n</sub>Gale (1,4) Glc, where n = 1-4. Included within these gangliosides of the neolacto series are 3'-nLMi (NeuAea (2,3) Gal / l (1,4) GleNAe / l (1,3) Gal / l (1,4) -Gle / l ( 1,1) -Ceramide) and 6'-nLMi (Neu-Aca (2,6) Gale (1,4) GlcNAc /? (1,3) Gate (1,4) -Glc /? (1,1) -Ceramide).
Ganglioside "mimics" are associated with some pathogens. For example, oligosaccharides from low molecular weight LPS nuclei of Campylobacter jejuni strains O: 19 were shown to exhibit ganglioside molecular mimicry. Since the late 1970s, Campylobacter jejuni has been recognized as an important cause of acute gastroenteritis in humans (Skirrow (1977) Brit. Med. J. 2: 9-11). Epidemiological studies have shown that Campylobacter infections are more common in developed countries than Salmonella infections and that they are also an important cause of diarrheal diseases in developing countries (Nachamkin et al. (1992), Campylobacter jejuni: Current Status and Future Trends. American Society for Microbiology, Washington, DC.). In addition to causing acute gastroenteritis, C. jejuni have been implicated as a frequent antecedent for the development of Guillain-Barré syndrome, a form of neuropathy that is the most common cause of generalized paralysis (Ropper (1992) N. Engl. J. Med. 326: 1130-1136) . The most common serotype of C. jejuni associated with Guillain-Barré syndrome is O: 19 (Kuroki (1993), Ann. Neurol. 33: 243-247) and this prompted the detailed study of the lipopolysaccharide (LPS) structure of the strains belonging to this serotype (Aspinall et al. (1994a), Infect. Immun. 62: 2122-2125; Aspinall et al. ( 1994b), Biochemistry 33: 241-249; and Aspinall et al. (1994c) Biochemistry 33: 250-255).
The terminal oligosaccharide fractions identical to those of the glycosides GD1a, GD3, GM1 and GT1a have been found in different O: 19 strains of C. jejuni. OH4384 of C. jejuni belongs to serotype O: 19 and was isolated from a patient who developed Guillain-Barré syndrome followed by an attack of diarrhea (Aspinall et al. (1994a), supra.). It was shown to have an outer core LPS that mimics the trisialylated ganglioside GT1a. The molecular mimicry of host structures by the saccharide portions of LPS is considered to be a virulence factor of different mucosal pathogens that would use this strategy to evade the immune response (Moran et al. (1996a) FEMS Immunol. Med. Microbiol 16: 105-115; Moran et al. (1996b) J. Endotoxin Res. 3: 521-531).
ES 2 269 098 T3
Consequently, the identification of the genes involved in the synthesis of LPS and the study of its regulation are of considerable interest for a better understanding of the mechanisms of pathogenesis used by these bacteria. Furthermore, the use of gangliosides as therapeutic reagents, as well as the study of ganglioside function, would be facilitated by convenient and efficient methods of synthesizing the desired gangliosides and ganglioside mimics. A combined chemical and enzymatic approach for the synthesis of 3'-nLMi and 6'-nLM has been described (Gaudini and Paulson (1994), J. Am. Chem. Soc. 116: 1149-1159). However, previously available enzymatic methods for ganglioside synthesis suffer from difficulties in efficiently producing enzymes in sufficient quantities, at low enough cost, for practical ganglioside synthesis on a large scale. Therefore, there is a need for new enzymes involved in ganglioside synthesis that is flexible for large-scale production. There is also a need for more efficient methods for the synthesis of gangliosides. The present invention meets these and other needs.
Brief description of the drawings
Figures 1A-1C show the structures of the outer core of the lipooligosaccharide (LOS) of the O: 19 strains of C. jejouni. These structures were described by Aspinall et al. (1994), Biochemistry 33, 241-249, and the portions showing similarity to the oligosaccharide portion of gangliosides are bounded by boxes. Figure 1A: The LOS of C. jejuni sero-strain O: 19 (ATCC # 43446) has structural similarity to the oligosaccharide portion of ganglioside GD1a. Figure 1B: The LOS of C. jejuni strain O: 19 OH4384 has structural similarity to the oligosaccharide portion of ganglioside GT1a. Figure 1C: The LOS of C. jejuni OH4382 has structural similarity to the oligosaccharide portion of ganglioside GD3.
Figures 2A-2B show the genetic organization of the OH4384 cst-I locus and the sharing of the loci for the biosynthesis of OH4384 and NCTC 11168 LOS. The distance between the marks on the scale is 1 kb. Figure 2A shows a schematic representation of the cst-I locus of OH4384, based on the nucleotide sequence that is available from GenBank (# AF130466). The prfB partial gene is somewhat similar to a Helicobacter pylori peptide chain releasing factor (GenBank # AE000537), while the cysD gene and cysN partial gene are similar to E. coli genes that encode subunits of adenylyl transferase sulfate (GenBank # AE000358). Figure 2B shows a schematic representation of the locus for OH4384 LOS biosynthesis, which is based on the nucleotide sequence of GenBank (# AF130984). The nucleotide sequence of the locus for LOS biosynthesis of OH4384 is identical to that of OH4384 except for the cgtA gene, which is missing an "A" (see text and GenBank # AF167345). The sequence of the locus for the biosynthesis of LOS NCTC 11168 is available from the Sanger Center (URL: http // www.sanger.ac.uk / Projects / C jejuni /). The corresponding homologous genes have the same number with an "a" migration for the OH4384 genes and a "b" migration for the NCTC 11168 genes. A single gene for strain OH4384 is shown in black and genes unique to NCTC 11168 are shown in gray. The # 5a and # 10a of the OH4384 ORF are found as a fusion ORF in the structure (# 5b / 10b) in NCTC 11168 and are denoted by a (*). The proposed functions for each ORF are found in Table 4.
Figure 3 shows an alignment of the deduced amino acid sequences for sialyltransferases. The cst-I gene from H4384 (first 300 residues), the cst-II gene from OH4384 (identical to cst-II from OH4382), the cstII gene from O: 19 (sero-strain) (GenBank # AF167344), the cst- II from NCTC 11168 and a putative ORF from H. influenzae (GenBank # U32720) were aligned using the Clusta1X alignment program (Thompson et al. (1997), Nucleic Acids Res. 25, 4876-82). The shading was produced through the GenDoc program (Nicholas, KB, and Nicholas, HB (1997) URL: http://www.cris.com/~ketchup/gendoc.shtml).
Figure 4 shows a scheme for the enzymatic synthesis of ganglioside mimics using the C. jejuni glycosyltransferases of OH4384. Starting from a synthetic acceptor molecule, a series of ganglioside mimics were synthesized with an α-2,3-sialyltransferase (Cst-I), aj6-1,4-N-acetylgalactosaminyltransferase (CgtA), ajd-1,3-galactosyltransferase (CgtB), and a recombinant bifunctional α-2,3 / α-2,8-sialyltransferase (Cst-II) using the sequences shown. All products were analyzed by mass spectrometry and the observed monoisotopic masses (shown in parentheses) were all within 0.02% of the theoretical masses. The GM3, GD3, GM2 and GM1a mimics were also analyzed by means of NMR spectroscopy (see Table 4). Summary of the invention
The present invention provides prokaryotic glycosyltransferase enzymes and nucleic acids encoding the enzymes. In one embodiment, the invention provides isolated and / or recombinant nucleic acid molecules that include a polynucleotide sequence that encodes a polypeptide selected from the group consisting of:
(a) a polypeptide having acetyltransferase activity for lipid A biosynthesis, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 350-1234 (ORF 2a ) from the locus for LOS biosynthesis of C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
(b) a polypeptide having glycosyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 1234-2487 (ORF 3a) of the locus for biosynthesis of the LOS of C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
ES 2 269 098 T3 (c) a polypeptide having glycosyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 50% identical to an amino acid sequence encoded by nucleotides 2786-3952 (ORF 4a) of the locus for LOS biosynthesis of C. jejouni strain OH4384 as shown in SEQ ID NO: 1 over a region of at least 100 amino acids in length;
(d) a polypeptide having the / 11,4-GalNAc transferase activity, wherein the GalNAc transferase polypeptide has an amino acid sequence that is approximately at least 77% identical to an amino acid sequence as set forth in SEQ ID NO: 13 over a region of at least 50 amino acids in length;
(e) a polypeptide having jB1,3-Galactosyltransferase activity, wherein the Galactosyltransferase polypeptide has an amino acid sequence that is approximately at least 75% identical to an amino acid sequence as set forth in SEQ ID NO: 15 or SEQ ID NO: 17 over a region of at least 50 amino acids in length;
(f) a polypeptide having either α2,3-sialyltransferase activity or both α2,3sialyltransferase and α2,8-sialyltransferase activity, wherein the Galactosyltransferase polypeptide has an amino acid sequence that is approximately 66% identical over a region of at least 60 amino acids in length to an amino acid sequence as set forth in one or more of SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 7, or SEQ ID NO: 10;
(g) a polypeptide having sialic acid synthase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 6924-7961 of the locus for biosynthesis of the LOS from C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
(h) a polypeptide having sialic acid biosynthesis activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 8021-9076 of the locus for biosynthesis of the LOS of C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
(i) a polypeptide having sialic acid-CMP synthetase activity, wherein the polypeptide contains an amino acid sequence that is approximately 65% identical to an amino acid sequence encoded by nucleotides 9076-9738 of the locus for biosynthesis of the LOS of C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
(j) a polypeptide having acetyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 65% identical to an amino acid sequence encoded by nucleotides 9729-10559 of the locus for LOS biosynthesis of strain OH4384 of C. jejouni as shown in SEQ ID NO: 1; and (k) a polypeptide having glycosyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 65% identical to an amino acid sequence encoded by a reverse complement of nucleotides 10557-11366 of the locus for LOS biosynthesis of C. jejouni strain OH4384 as shown in SEQ ID NO: 1.
In presently preferred embodiments, the invention provides an isolated nucleic acid molecule that includes a polynucleotide sequence that encodes one or more of the polypeptides selected from the group consisting of: a) a sialyltransferase polypeptide having both α2,3 sialyltransferase activity and α2,8 sialyltransferase activity, wherein the sialyltransferase polypeptide has an amino acid sequence approximately at least 76% identical to a sequence of amino acids as set forth in SEQ ID NO: 3 over a region approximately at least 60 amino acids in length; b) a GalNAc transferase polypeptide having a / 11,4-GalNAc transferase activity, wherein the GalNAc transferase polypeptide has an amino acid sequence that is approximately at least 75% identical to an amino acid sequence such as set forth in SEQ ID NO: 13 over a region approximately at least 50 amino acids long; and c) a galactosyltransferase polypeptide having the / 11,3 galactosyltransferase activity, wherein the galactosyltransferase polypeptide has an amino acid sequence that is approximately at least 75% identical to an amino acid sequence as set forth in SEQ. ID NO: 15 over a region approximately at least 50 amino acids long.
The invention also provides expression cassettes and expression vectors in which a nucleic acid of the galactosyltransferase of the invention is operably linked to a promoter and other control sequences that facilitate the expression of glycosyltransferases in a desired host cell. Recombinant host cells expressing the glycosyltransferases of the invention are also provided.
The invention also provides polypeptides produced in isolation and / or recombinantly, selected from the group consisting of:
ES 2 269 098 T3
a) a polypeptide having acetyltransferase activity for lipid A biosynthesis, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 350-1234 (ORF 2a ) from the locus for LOS biosynthesis of C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
b) a polypeptide having glycosyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 1234-2487 (ORF 3a) of the locus for biosynthesis of the LOS of C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
c) a polypeptide having glycosyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 50% identical to an amino acid sequence encoded by nucleotides 2786-3952 (ORF 4a) of the locus for biosynthesis of the LOS of C. jejouni strain OH4384 as shown in SEQ ID NO: 1 over a region approximately at least 100 amino acids long;
d) a polypeptide having the / 11,4-GalNAc transferase activity, wherein the GalNAc transferase polypeptide has an amino acid sequence that is approximately at least 77% identical to an amino acid sequence as set forth in SEQ ID NO: 13 over a region approximately at least 50 amino acids long;
e) a polypeptide having the / 11,3-galaclosillransferase activity, wherein the galactosyltransferase polypeptide has an amino acid sequence that is approximately at least 75% identical to an amino acid sequence as set forth in SEQ ID NO: 15 or SEQ ID NO: 17 over a region approximately at least 50 amino acids in length;
f) a polypeptide having either α2,3-sialyltransferase activity or, both α2,3sialyltransferase and α2,8-sialyltransferase activity, wherein the polypeptide has an amino acid sequence that is approximately at least 66% identical to an amino acid sequence as set forth in SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 7 or SEQ ID NO: 10 over a region of at least 60 amino acids in length;
g) a polypeptide having sialic acid synthase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 6924-7961 of the locus for LOS biosynthesis from C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
h) a polypeptide having the biosynthetic activity of sialic acid, wherein the polypeptide contains an amino acid sequence that is approximately at least 70% identical to an amino acid sequence encoded by nucleotides 8021-9076 of the locus for LOS biosynthesis from C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
i) a polypeptide having sialic acid-CPM synthetase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 65% identical to an amino acid sequence encoded by nucleotides 9076-9738 of the locus for biosynthesis of the LOS of C. jejouni strain OH4384 as shown in SEQ ID NO: 1;
j) a polypeptide having acetyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 65% identical to an amino acid sequence encoded by nucleotides 9729-10559 of the locus for LOS biosynthesis of the C. jejouni strain OH4384 as shown in SEQ ID NO: 1; Y
k) a polypeptide having glycosyltransferase activity, wherein the polypeptide contains an amino acid sequence that is approximately at least 65% identical to an amino acid sequence encoded by a reverse complement of nucleotides 10557-11366 of the para locus the LOS biosynthesis of C. jejouni strain OH4384 as shown in SEQ ID NO: 1.
In presently preferred embodiments, the invention provides glycosyltransferase polypeptides including: a) a sialyltransferase polypeptide having both α2,3-sialyltransferase activity and α2,8-sialyltransferase activity, wherein the polypeptide of sialyltransferase has an amino acid sequence that is approximately at least 76% identical to an amino acid sequence as set forth in SEQ ID NO: 3 over a region of at least 60 amino acids in length; b) a GalNAc transferase polypeptide having the activity of the / 11,4-GalNAc transferase, wherein the GalNAc transferase polypeptide has an amino acid sequence that is approximately at least 75% identical to an amino acid sequence such as set forth in SEQ ID NO: 13 over a region approximately at least 50 amino acids in length; and c) a galactosyltransferase polypeptide having the / 11,3-galactosyltransferase activity, wherein the galactosyltransferase polypeptide has an amino acid sequence that is approximately at least 75% identical to an amino acid sequence.
ES 2 269 098 T3 as set forth in SEQ ID NO: 13 or SEQ ID NO: 17 over a region approximately at least 50 amino acids long.
The invention also provides reaction mixtures for the synthesis of a sialylated oligosaccharide. Reaction mixtures include a sialyltransferase polypeptide having both α2,3-sialyltransferase activity and α2,8-sialyltransferase activity. A galactosylated acceptor moiety and a sialyl nucleotide sugar are also present in the reaction mixtures. The sialyltransferase transfers a first sialic acid residue from the sialyl sugar nucleotide (e.g., CMP-sialic acid) to the acceptor moiety galactosylated at an α2,3 bond, and a second sialic acid residue is further added to the first residue of sialic acid at an α2,8 bond.
In another embodiment, the invention provides methods for synthesizing a sialylated oligosaccharide. These methods involve the incubation of a reaction mixture that includes a sialyltransferase polypeptide that has both α2,3 sialyltransferase activity and α2,8 sialyltransferase activity, a galactosylated acceptor moiety, and a sialyl nucleotide sugar. , under suitable conditions wherein the sialyltransferase polypeptide transfers a first sialic acid residue from the sialyl sugar nucleotide to the acceptor moiety galactosylated at an α2,3 bond, and further transfers a second sialic acid residue to the first sialic acid residue at an α2,8 bond. Detailed description
Glycosyltransferases, reaction mixtures, and methods of the invention are useful for transferring a monosaccharide from a donor substrate to an acceptor molecule. The addition generally takes place at the non-reducing end of an oligosaccharide or carbohydrate moiety on a biomolecule. Biomolecules as defined herein include, but are not limited to, biologically significant molecules such carbohydrates, proteins (eg, glycoproteins), and lipids (eg, glycolipids, phospholipids, sphingolipids, and gangliosides).
The following abbreviations are used here:
Ara = arabinosil,
Fru = fructosyl;
Fuc = fucosyl;
Gal = galactosyl;
GalNAc = N-acetylgalactosaminyl;
Glc = glucosyl;
GlcNAc = N-acetylglucosaminyl;
Man = manosil; Y
NeuAc = sialyl (N-acetylneuraminyl).
The term "sialic acid" refers to any member of a family of carboxylated nine carbon sugars. The most common member of the sialic acid family is N-acetylneuraminic acid (2-keto-5-acetamindo-3,5-dideoxy-D-glycero-D-galactononulopyrans-1-ionic acid (often abbreviated as Neu5Ac, NeuAc, or NANA) A second member of the family is N-glycolylneuraminic acid (Neu5Gc or NeuGc), in which the N-acetyl group of NeuAc is hydroxylated. A third member of the sialic acid family is 2-keto-3-deoxy-nonulosonic acid (KDN) (Nadano et al. (1986) J. Biol. Chem. 261: 11550-11557; Kanamori et al. (1990), J. Biol. Chem. 265: 21811-21819 Also included are 9-substituted sialic acids such as a 9-O-Ci-C<sub>6</sub> acyl-Neu5Ac such as 9-O-lactyl-Neu5Ac or 9-O-acetyl-Neu5Ac, 9-deoxy-9-fluoro-Neu5Ac and 9-azido-9-deoxy-Neu5Ac. For a review of the sialic acid family, see, for example, Varki (1992) Glycobiology 2: 25-40; Sialic Acids: Chemistry, Metabolism and Function, R. Schauer, Ed. (Springer-Verlag, New York (1992); Schauer, Methods in Enzymology, 50: 64-89 (1987), and Schaur, Advances in Carbohydrate Chemistry and Biochemistry , 40: 131-234. The synthesis and use of sialic acid compounds in a sialylation process are described in international application WO 92/16640, published October 1, 1992.
Donor substrates for glycosyltransferases are activated nucleotide sugars. Such activated sugars generally consist of uridine and guanine diphosphates, and cystidine monophosphate derivatives of sugars in which the nucleoside diphosphate or monophosphate serve as a leaving group. Bacterial, plant, and fungal systems can sometimes utilize other activated nucleotide sugars.
Oligosaccharides are considered to have a reducing end and a non-reducing end, whether or not the saccharide at the reducing end is actually a reducing sugar or not. In accordance with accepted nomenclature, oligosaccharides are described here with the non-reducing end on the left and the reducing end on the right.
ES 2 269 098 T3
All the oligosaccharides described here are described with the name or abbreviation for the non-reducing saccharide (for example, Gal), followed by the configuration of the glycosidic bond (a or β), of the ring bond, the position of the reducing saccharide ring involved in the link, and then the name or abbreviation of the reducing saccharide (eg, GlcNAc). The bond between two sugars can be expressed, for example, as 2,3, 2 ^ 3, or (2,3). Each saccharide is a pyranose or furanose.
The term "nucleic acid" refers to a deoxyribonucleotide or ribonucleotide polymer in either single or double stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in form. similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence includes the sequence complementary to it.
The term "operably linked" refers to the functional linkage between a control sequence for nucleic acid expression (such as a promoter, signal sequence, or arrangement of transcription factor binding sites) and a second acid sequence. nucleic acid, wherein the expression control sequence affects the transcription and / or translation of the nucleic acid corresponding to the second sequence.
A "heterologous polynucleotide" or a "heterologous nucleic acid", as used herein, is one that originates from a source external to the particular host cell, or, if from the same source, is modified from its original form. Thus, a heterologous glycosyltransferase gene in a host cell includes a glycosyltransferase gene that is endogenous to the particular host cell but has been modified. Modification of the heterologous sequence can occur, for example, by treating DNA with a restriction enzyme to generate a DNA fragment that is capable of being operably linked to a promoter. Techniques such as site-directed mutagenesis are also useful for modifying a heterologous sequence.
The term "recombinant" when used in reference to a cell indicates that the cell replicates a heterologous nucleic acid, or expresses a heterologous acid-encoded peptide or protein. Recombinant cells can contain genes that are not found within the native (non-recombinant) form of the cell. Recombinant cells also include those that contain genes that are found in the cell's native form, but that are modified and introduced into the cell through artificial means. The term also encompasses cells that contain a nucleic acid endogenous to the cell that has been modified without removing the nucleic acid from the cell; Such modifications include those obtained through gene replacement, site-specific mutation, and related techniques known to those skilled in the art.
A "recombinant nucleic acid" is a nucleic acid that is in a form that is altered from its natural state. For example, the term "recombinant nucleic acid" includes a coding region that is operably linked to a promoter and / or other expression control region, processing signal, other coding region, and the like, to which the acid nucleic is not bound in its naturally occurring form. A "recombinant nucleic acid" also includes, for example, a coding region or other nucleic acid in which one or more nucleotides have been substituted, deleted, inserted, compared to the corresponding naturally occurring nucleic acid. Modifications include those introduced by way of in vitro manipulation, in vivo modification, synthetic methods, and the like.
A "recombinantly produced polypeptide" is a polypeptide that is encoded by a heterologous and / or recombinant nucleic acid. For example, a polypeptide that is expressed from a nucleic acid encoding C. jejouni glycosyltransferase that is introduced into E. coli is a "recombinantly produced polypeptide." A protein expressed from a nucleic acid that is operably linked to a non-native promoter is an example of a "recombinantly produced polypeptide." The recombinantly produced polypeptides of the invention can be used to synthesize gangliosides and other polysaccharides in their non-purified form (eg, as a cell lysate or an intact cell), or after being partially or completely purified.
A "recombinant expression cassette" or simply an "expression cassette" is a nucleic acid construct generated recombinantly or synthetically, with elements of the nucleic acid that are capable of affecting the expression of a structural gene in hosts compatible with such sequences. Expression cassettes include at least promoters and optionally, transcription termination signals. Typically, the recombinant expression cassette includes the nucleic acid to be transferred (eg, a nucleic acid encoding a desired polypeptide), and a promoter. Additional factors necessary or useful to effect expression may also be used as described herein. For example, an expression cassette can also include nucleotide sequences that encode a signal sequence that directs the secretion of an expressed protein from the host cell. Transcription termination signals, enhancers, and other nucleic acid sequences that influence gene expression can also be included in an expression cassette.
A "subsequence" refers to a sequence of nucleic acids or amino acids that contain a portion of a longer sequence of nucleic acids or amino acids (eg, a polypeptide), respectively.
The term "isolated" is to refer to material that is substantially or essentially free of the components that normally accompany the material as it is in its native state. Typically, the isolated proteins or nucleic acids of the invention are about at least 80% pure, usually about at least 90%, and preferably about at least 95% pure. Purity or homogeneity can be indicated
ES 2 269 098 T3 by means of a number of ways well known in the state of the art, such as agarose gel or polyacrylamide gel electrophoresis of a protein or nucleic acid sample, followed by visualization by means of coloring. For certain purposes high resolution and HPLC or a similar medium will be needed for the purification used. An "isolated" enzyme, for example, is one that is substantially or essentially free of components that interfere with the activity of the enzyme. An "isolated nucleic acid" includes, for example, one that is not present on the chromosome of the cell in which the nucleic acid occurs naturally.
The terms "identical" or percent "identity", in the context of two or more acidic nucleic acid or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues. or of nucleotides that are the same, when compared and aligned for maximum correspondence, when measured using one of the following sequence comparison algorithms, or by visual inspection.
The phrase "substantially identical", in the context of two nucleic acids or polypeptides, refers to two or more sequences or subsequences that have at least 60%, preferably 80%, more preferably 90-95% identity of the amino acid residues or nucleotide numbers, when compared and aligned for maximum correspondence, when measured using one of the following sequence comparison algorithms, or by visual inspection. Preferably, substantial identity exists over a region of the sequences that is about at least 50 residues in length, more preferably over a region about at least 100 residues, and most preferably, the sequences are substantially identical over a region of about 100 residues. at least 150 residues. In the most preferred embodiment, the sequences are substantially identical throughout the length of the coding regions.
For sequence comparison, a sequence typically acts as a reference sequence, to which the test sequences are compared. When using a sequence comparison algorithm, the test and reference sequences are entered into a computer, the subsequence coordinates are designated, if necessary, and the sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity for the test sequence (s), relative to the reference sequence, based on the designated program parameters.
Optimal alignment of the sequences for comparison can be performed, for example, by means of the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2: 482 (1981), by means of the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48: 443 (1970), by means of the search for similarity method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. United States of America 85: 2444 (1988), through computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by visual inspection (see generally, Current Protocols in Molecular Biology, FM Ausubel et al., eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1995 Supplement) (Ausubel)).
Examples of algorithms that are suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (1990) J. Mol. Biol. 215: 403-410 and in Altschuel et al. (1977) Nucleic Acids Res. 25: 3389-3402, respectively. The software to perform the analyzes with the BLAST algorithms is publicly available through the National Center for Biotechnology Information (http://www.ncbi.nlm.nih.gov/). For example, the comparison can be performed using a BLASTN version 2.0 algorithm with a word length (W) of 11, G = 5, E = 2, q = -2, and r = 1, and a comparison of both strings. For amino acid sequences, the BLASTP version 2.0 algorithm can be used, with default word length (W) values of 3, G = 11, E = 1, and a BLOSUM62 substitution matrix. (see, Henikoff & Henikoff, Proc. Natl. Acad. Sci. United States of America 89: 10915 (1989)).
In addition to calculating percent sequence identity, the BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, for example, Karlin & Altschul, Proc. Nat'l. Acad. Sci. United States of America 90: 5873-5787 (1993)). A measure of similarity supplied by the BLAST algorithm is the probability of the smallest sum (P (N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the probability of the smallest sum in a comparison of the test nucleic acid with the reference nucleic acid is approximately less than 0.1, more preferably approximately less than 0. .01, and more preferably about less than 0.001.
The phrase "specifically hybridize with" refers to the binding, duplication, or hybridization of a molecule only with a particular sequence of nucleotides under stringent conditions when that sequence is present in a complex (eg, total cellular) mixture of DNA or RNA. . The term "stringent conditions" refers to the conditions under which a probe will hybridize to its target subsequence, but not to other sequences. Stringent conditions are sequence dependent and will be different in different circumstances. Longer sequences hybridize specifically at higher temperatures. Generally, stringent conditions are selected to be about 5 ° C lower than the thermal melting point (Tm) for a specific sequence with defined pH and ionic strength. The Tm is the temperature (under a defined ionic strength, pH, and nucleic acid concentration) at which 50% of the probes complementary to the target sequence hybridize to the target sequence at equilibrium. (since the target sequences are generally present in excess, in the
ES 2 269 098 T3
Tm, 50% of the probes are occupied in equilibrium). Typically, stringent conditions will be those in which the salt concentration is about less than 1.0 M Na ions, typically about 0.01 to 1.0 M concentration (or other salts) at pH between 7 , 0 and 8.3 and the temperature is at least about 30 ° C for short probes (eg, 10 to 50 nucleotides) and about at least 60 ° C for long probes (eg, greater than 50 nucleotides). Stringent conditions can also be achieved with the addition of destabilizing agents such as formamide.
A further indication that two nucleic acid sequences or polypeptides are substantially identical is that the polypeptide encoded by the first nucleic acid is immunologically cross-reactive with the polypeptide encoded by the second nucleic acid, as described below. Thus, one polypeptide is typically substantially identical to a second polypeptide, for example, where the two peptides differ only in conservative substitutions. Another indication that two nucleic acid sequences are substantially identical is that the two molecules hybridize to each other under stringent conditions, as described below.
The phrases "specifically binds to a protein" or "is specifically immunoreactive with" when referring to an antibody refers to a binding reaction that is determinative of the presence of the protein in the presence of a heterogeneous population of proteins and other biologicals. Therefore, under the designated conditions of the immunoassay, the specified antibodies preferentially bind to a particular protein and do not bind to a significant amount of other proteins present in the sample. Specific binding to a protein under such conditions requires an antibody that is selected for its specificity for a particular protein. A variety of immunoassay formats can be used to screen for antibodies specifically immunoreactive to a particular protein. For example, solid phase ELISA immunoassays are routinely used to screen for monoclonal antibodies specifically immunoreactive with a protein. See, Harlow and Lane (1988), Antibodies, A Laboratory Manual, Cold Spring Harbor Publications, New York, for a description of immunoassay formats and conditions that can be used to determine specific immunoreactivity.
"Conservatively modified variations" of a particular polynucleotide sequence refer to those polynucleotides that encode identical, or essentially identical amino acid sequences, or where the polynucleotide does not encode an amino acid sequence, for essentially identical sequences. Due to the degeneracy of the genetic code, a large number of essentially identical nucleic acids code for any given polypeptide. For example, the codons CGU, CGC, CGA, CGG, AGA, and AGG all code for the amino acid arginine. Therefore, at each position where an arginine is specified by a codon, the codon can be altered by any of the corresponding codons described, without altering the encoded polypeptide. Such nucleic acid variations are "silent variations", which are a kind of "conservatively modified variations". Each polynucleotide sequence described herein that encodes a polypeptide also describes every possible silent variation, except where noted otherwise. One of skill will recognize that every codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine) can be modified to produce a functionally identical molecule by standard techniques. Therefore, each "silent variation" of a nucleic acid that encodes a polypeptide is implicit in each described sequence.
Furthermore, one of skill will recognize that individual substitutions, deletions, or additions that alter, add, or delete a single amino acid or a small percentage of amino acids (typically less than 5%, more typically less than 1%) in an encoded sequence are "conservative variations. modified ”where the alterations result in the substitution of an amino acid with a chemically similar amino acid. Conservative substitution tables that provide amino acids of similar functionality are well known in the art. One of skill will appreciate that many conservative variations of the fusion proteins and of the nucleic acid encoding the fusion proteins give essentially identical products. For example, due to the degeneracy of the genetic code, "silent substitutions" (that is, substitutions of a nucleic acid sequence that do not result in an alteration in an encoded polypeptide) are an implicit feature of each nucleic acid sequence that encodes to an amino acid. As described herein, the sequences are preferably optimized for expression in a particular host cell used to produce the enzymes (eg, yeast, human, and the like). Similarly, "conservative amino acid substitutions", at one or a few amino acids in an amino acid sequence, which are substituted with different amino acids with very similar properties (see, definition section, supra), are also easily identified. because they are very similar to a particular amino acid sequence, or to a particular nucleic acid sequence that encodes an amino acid. Such conservatively substituted variations of any particular sequence are a feature of the present invention. See also, Creighton (1984) Proteins, WH Freeman and Company. Furthermore, individual substitutions, deletions or additions that alter, add or delete a single amino acid or a small percentage of amino acids in an encoded sequence are also "conservatively modified variations".
Description of the preferred modalities
The present invention provides new glycosyltransferase enzymes, as well as other enzymes that are involved in enzyme-catalyzed oligosaccharide synthesis. The glycosyltransferases of the invention include sialyltransferases, including a bifunctional sialyltransferase having both α2,3 sialyltransferse activity and α2,8 sialyltransferse activity. Also provided are jB1,3-galactosyltransferases, / 11,4-GalNAc transfera9
ES 2 269 098 T3 sas, sialic acid synthases, sialic acid-CPM synthesizes, acetyltransferases, and other glycosyltransferases. The enzymes of the invention are prokaryotic enzymes, which include those involved in the biosynthesis of lipooligosaccharides (LOS) in different strains of Campylobacter jejuni. The invention also provides nucleic acids encoding these enzymes, as well as expression cassettes and expression vectors for use in the expression of glycosyltransferases. In additional embodiments, the invention provides reaction mixtures and methods in which one or more of the enzymes are used to synthesize an oligosaccharide.
The glycosyltransferases of the invention are useful for different purposes. For example, glycosyltransferases are useful as tools for the chemoenzymatic synthesis of oligosaccharides, including gangliosides and other oligosaccharides that have biological activity. The glycosyltransferases of the invention, and nucleic acids encoding glycosyltransferases, are also useful for studies of the mechanisms of pathogenesis of organisms that synthesize ganglioside mimics, such as C. jejouni. Nucleic acids can be used as probes, for example, to study the expression of genes involved in the synthesis of ganglioside mimics. Antibodies raised against glycosyltransferases are also useful for the analysis of the expression patterns of these genes that are involved in pathogenesis. Nucleic acids are also useful in designing antisense oligonucleotides to inhibit the expression of Campylobacter enzymes that are involved in the biosynthesis of ganglioside mimics that can mask pathogens from the host immune system.
The glycosyltransferases of the invention provide several advantages over previously available glycosyltransferases. Bacterial glycosyltransferases such as those of the invention can catalyze the formation of oligosaccharides that are identical to corresponding structures in mammals. Furthermore, bacterial enzymes are easier and less expensive to produce in quantity compared to mammalian glycosyltransferases. Therefore, bacterial glycosyltransferases such as those of the present invention are attractive replacements for mammalian glycosyltransferases, which can be difficult to obtain in large quantities. That the glycosyltransferases of the invention are of bacterial origin facilitates the expression of large amounts of the enzymes using relatively inexpensive prokaryotic expression systems. Typically, prokaryotic systems for the expression of polypeptide products involve much lower costs than the expression of the polypeptides in mammalian cell culture systems.
Furthermore, the novel bifunctional sialyltransferases of the invention simplify the enzymatic synthesis of biologically important molecules, such as gangliosides, which have a sialic acid linked via an α2,8 bond to a second sialic acid, which is itself α2,3 -bound to a galactosylated acceptor. While previous methods for the synthesis of these structures require two separate sialyltransferases, only one sialyltransferase is required when using the bifunctional sialyltransferase of the present invention. This avoids the costs associated with obtaining a second enzyme, and can also reduce the number of steps involved in the synthesis of these compounds.
A. Glycosyltransferases and associated enzymes
The present invention provides prokaryotic glycosyltransferase polypeptides, as well as other enzymes that are involved in glycosyltransferase catalyzed synthesis of oligosaccharides, including gangliosides and ganglioside mimics. In presently preferred embodiments, polypeptides include those that are encoded by open reading frames within the lipopolysaccharide (LOS) locus of Campylobacter species (Figure 1). Included within the enzymes of the invention are glycosyltransferases, such as sialyltransferases (including a bifunctional sialyltransferase), / 11,4-GalNAc transferases, and jd1,3-galactosyltransferases, among other enzymes such as those described herein. Also provided are accessory enzymes such as, for example, sialic acid CMP synthetase, sialic acid synthase, acetyltransferase, an acetyltransferase that is involved in lipid A biosynthesis, and an enzyme involved in sialic acid biosynthesis.
The glycosyltransferases and accessory polypeptides of the invention can be purified from natural sources, for example prokaryotes such as Campylobacter species. In presently preferred embodiments, glycosyltransferases are derived from C. jejuni, in particular from C. jejuni serotype O: 19, including strains OH4384 and OH4382. Glycosyltransferases and accessory enzymes derived from C. jejuni serotypes 0:10, 0:41, and 0: 2 are also provided. Methods by which glycosyltransferase polypeptides can be purified include standard protein purification methods including, for example, ammonium sulfate precipitation, affinity columns, column chromatography, gel electrophoresis, and the like ( see generally R. Scopes, Protein Purification, Springer-Verlag, NY (1982) Deutscher, Methods in Enzymology Vol. 182: Guide to Protein Purification., Academic Press, Inc. NY (1990)).
In presently preferred embodiments, the glycosyltransferase and accessory polypeptides of the enzyme of the invention are obtained by means of recombinant expression using the nucleic acids encoding the glycosyltransferase and accessory enzyme described herein. Expression vectors and methods for producing glycosyltransferase are described in detail below.
In some embodiments, the glycosyltransferase polypeptides are isolated from their natural environment, whether they have been recombinantly produced or purified from their natural cells. Substantially pure compositions of about at least 90 to 95% homogeneity are preferred for some applications, and
ES 2 269 098 T3 more preferred, 98 to 99% or more homogeneity. Once purified, or to homogeneity as desired, the polypeptides can then be used (e.g., as immunogens for the production of antibodies, or for the synthesis of oligosaccharides, or other uses as described herein, or apparent to those skilled in the art. the art). Glycosyltransferases do not, however, need to be partially purified yet to be used for the synthesis of a desired saccharide structure. For example, the invention provides recombinantly produced enzymes that are expressed in a heterologous host cell and / or from recombinant nucleic acid. Such enzymes of the invention can be used when they are present in a cell lysate or an intact cell, as well as in purified form.
1. Sialyltransferases
In some embodiments, the invention provides sialyltransferase polypeptides. Sialyltransferases have α2,3-sialyltransferase activity, and in some cases they also have α2,8-sialyltransferase activity. These bifunctional sialyltransferases, when placed in a reaction mixture with a suitable saccharide acceptor (for example, a saccharide having a terminal galactose) and a sialic acid donor (for example, CMP-sialic acid) can catalyze the transfer of a first sialic acid from donor to acceptor at an α2,3 bond. The sialyltransferase then catalyzes the transfer of a second sialic acid from the sialic acid donor to the first sialic acid residue at an α2,8 bond. This type of Siaa2,8-Siaa2,3-Gal structure is often found in gangliosides, including GD3 and GT1 as shown in Figure 4.
Examples of bifunctional sialyltransferases of the invention are those found in Campylobacter species, such as C. jejuni. A currently preferred bifunctional sialyltransferase of the invention is that of C. jejuni serotype O: 19. An example of a bifunctional sialyltransferase is that of C. jejuni strain OH4384; this sialyltransferase has an amino acid sequence as shown in SEQ ID NO: 3. Other bifunctional sialyltransferases of the invention generally have an amino acid sequence that is about at least 76% identical to the amino acid sequence of C. jejuni OH4384 bifunctional sialyltransferase over a region about at least 60 amino acids in length. More preferably, the sialyltransferases of the invention are approximately at least 85% identical to the amino acid sequence of the OH4384 sialyltransferase, and even more preferably approximately at least 95% identity to the amino acid sequence of SEQ ID NO: 3, on a region approximately 60 amino acids long. In presently preferred embodiments, the region of percent identity spans a region longer than a 60 amino acid region. For example, in more preferred embodiments, the region of similarity spans a region approximately at least 100 amino acids in length, more preferably a region approximately at least 150 amino acids in length, and most preferably over the entire length of the sialyltransferase. . Therefore the bifunctional sialyltransferases of the invention include polypeptides that have either or both of the α2,3- and α2,8-sialyltransferase activities, and are about at least 65% identical, more preferably about 70% identical, more preferably about at least 80% identical, and most preferably about at least 90% identical to the amino acid sequence of the Cstll OH4384 sialyltransferase from C. jejuni (SEQ ID NO: 3) on a region of the polypeptide that is required to retain the respective sialyltransferase activities. In some embodiments, the bifunctional sialyltransferases of the invention are identical to the C. jejuni OH4384 sialyltransferase Cstll over the full length of the sialyltransferase.
The invention also provides sialyltransferases having α2,3 sialyltransferase activity, but little or no α2,8 sialyltransferase activity. For example, Cstll sialyltransferase from C. jejuni sero-strain O: 19 (SEQ ID NO: 9) differs from that of strain OH4384 by eight amino acids, but nevertheless substantially lacks α2,8 sialyltransferase activity (Figure 3). The corresponding sialyltransferase of strain NCTC 11168 of serotype 0: 2 (SEQ ID NO: 10) is 52% identical to that of OH4384, and also has little or no α2,8 sialyltransferase activity. Also provided are sialyltransferases that are substantially identical to the Cstll sialyltransferase from C. jejuni strain 0:10 (SEQ ID NO: 5) and from strain a0: 41 (SEQ ID NO: 7). The sialyltransferases of the invention include those that are about at least 65% identical, more preferably about at least 70% identical, more preferably about at least 80% identical, and most preferably about at least 90% identical to the amino acid sequences. of C. jejuni 0:10 (SEQ ID NO: 5), 0:41 (SEQ ID NO: 7), sero-strain O: 19 (SEQ ID NO: 9), or strain NCTC 11168 of serotype 0: 2 (SEQ ID NO: 10). The sialyltransferases of the invention, in some embodiments, have an amino acid sequence that is identical to that of the 0:10, 0:41, O: 19 sero-strains or the C. jejuni NCTC 11168 strains.
The percentage of identities can be determined by inspection, for example, or it can be determined using an alignment algorithm such as the BLASTP version 2.0 algorithm using default parameters, such as word length (W) of 3, G. = 11, E = 1, and a BLOSUM62 substitution matrix.
The sialyltransferases of the invention can be identified, not only by sequence comparison, but also by preparing antibodies against C. jejuni bifunctional sialyltransferase OH4384, or other sialyltransferases provided herein, and which determine whether the antibodies they are specifically immunoreactive with a sialyltransferase of interest. To obtain a particular bifunctional sialyltransferase, an organism that is likely to produce a bifunctional sialyltransferase can be identified by determining whether the organism displays both α2,3 and 2,8 linkages of sialic acid on its cell surfaces. Alternatively, or additionally, one can simply assay the enzyme from an isolated sialyltransferase to determine whether both sialyltransferase activities are present.
ES 2 269 098 T3
2. f1,4-GalNAc transferase
The invention also provides the deβ 1,4-GalNAc transferase polypeptides (eg, CgtA). The β 1,4-GalNAc transferases of the invention, when placed in a reaction mixture, catalyze the transfer of a GalNAc residue from a donor (e.g., UDP-GalNAc) to a suitable acceptor saccharide (typically a saccharide having a terminal galactose residue). The resulting GalNAc ^ 1,4-Gal structure is often found in gangliosides and other sphingoids, among many other saccharide compounds. For example, CgtA transferase can catalyze the conversion of ganglioside GM3 to GM2, (Figure 4).
Examples of the 1,4-GalNAc transferases of the invention are those that are produced by means of the Campylobacter species, such as C. jejuni. An example of a /> 1,4-GalNAc transferase polypeptide is that of C. jejuni strain OH4384, which has an amino acid sequence as shown in SEQ ID NO: 13. The /> 1,4GalNAc transferases of the invention generally include an amino acid sequence that is approximately at least 75% identical to an amino acid sequence as set forth in SEQ ID NO: 13 over a region of approximately at least 50 amino acids in length. . More preferably, the β 1,4-GalNAc transferases of the invention are approximately at least 85% identical to this amino acid sequence, and even more preferably approximately at least 95% identical to the amino acid sequence of SEQ ID nO: 13, over a region approximately at least 50 amino acids long. In presently preferred embodiments, the region of percent identity extends over a region greater than 50 amino acids, more preferably over a region of about at least 100 amino acids, and most preferably, over the full length of the GalNAc transferase. Therefore, the ^ 1,4-GalNAc transferases of the invention include polypeptides that have ^ 1,4GalNAc transferase activity and are approximately at least 65% identical, more preferably approximately at least 70% identical, more preferably approximately equal to less 80% identical, and most preferably approximately at least 90% identical to the amino acid sequence of the ^ 1,4-GalNAc transferases of OH4384 from C. jejuni (SEQ ID NO: 13) on a region of the polypeptide that is required to retain the activity of ^ 1,4-GalNAc transferase. In some embodiments, the ^ 1,4-GalNAc transferases of the invention are identical to the C. jejuni OH4384 ^ 1,4-GalNAc transferase over the full length of the ^ 1,4-GalNAc transferase.
Again, the percentage of identities can be determined by inspection, for example, or it can be determined using an alignment algorithm such as the BLASTP version 2.0 algorithm with a word length (W) of 3, G = 1, E = 1, and a BLOSUM62 substitution matrix.
The β 1,4-GalNAc transferases of the invention can also be identified by immunoreactivity. For example, one can prepare antibodies against C. jejuni OH4384 ^ 1,4-GalNAc transferase of SEQ ID NO: 13 and determine whether the antibodies are specifically immunoreactive with a β 1,4-GalNAc transferase of interest.
3. β1,3-Galactosyltransferases
The invention also provides / 10-galaclosyl transferase (CgtB). When placed in a suitable reaction medium, the β 1,3-galactosyltransferases of the invention catalyze the transfer of a galactose residue from a donor (eg, UDP-Gal) to a suitable saccharide acceptor (eg, saccharides having a terminal GalNAc residue). Among the reactions catalyzed by β 1,3-galactosyltransferases is the transfer of a galactose residue to the GM2 oligosaccharide fraction to form the GM1a oligosaccharide fraction.
Examples of the β 1,3-galactosyltransferases of the invention are those produced by the Campylobacter species, such as C. jejuni. For example, a β 1,3-galactosyltransferase of the invention is that of C. jejuni strain OH4384, which has the amino acid sequence shown in SEQ ID NO: 15.
Another example of a /> 1,3-galaclosillransferase of the invention is that of the strain NCTC 11168 of the serotype 0: 2 of C. jejuni. The amino acid sequence of this galactosyltransferase is set forth in SEQ ID NO: 17. This galactosyltransferase is well expressed in E. coli, for example, and exhibits a high amount of soluble activity. Also, unlike CgtB from OH4384, which can add more than one galactose if a reaction mix contains excess donor and is incubated for a long enough period of time, the β 1,3-galactose from NCTC 11168 does not have a significant amount of polygalactosyltransferase activity. For some applications, the polygalactosyltransferase activity of the OH4384 enzyme is desirable, but in other applications such as the synthesis of GM1 mock-ups, the addition of only one terminal galactose is desirable.
The /> 1,3-galaclosillransferases of the invention generally have an amino acid sequence that is approximately at least 75% identical to an amino acid sequence of the CgtB of OH4384 or NCTC 11168 as set forth in SEQ ID NO: 15 and in SEQ ID NO: 17, respectively, over a region approximately at least 50 amino acids in length. More preferably, the β 1,3-galactosyltransferases of the invention are approximately at least 85% identical to any of these amino acid sequences, and even more preferably approximately at least 95% identical to the amino acid sequences of SEQ ID NO: 15 or SEQ ID NO: 17, over a region approximately at least 50 amino acids in length. In presently preferred embodiments, the region of percent identity extends over a region longer than 50 amino acids, more preferably over a region of about at least 100 amino acids, and most preferably, over the full length of the galactosyltransferase. Thus, the β 1,3-galactosyltransferase of the invention includes polypeptides that have /> 1,3-galaclosyl transferase activity and are approximately at least 65% identical, more preferably
ES 2 269 098 T3 about at least 70% identical, more preferably about at least 80% identical, and most preferably about at least 90% identical to the amino acid sequence of C. jejuni OH4384 laj61,3-galactosyltransferase (SEQ ID NO: 15) or the galactosyltransferase of NCTC 11168 (SEQ ID NO: 17) on a region of the polypeptide that is required to retain the activity of laj61,3-galactosyltransferase. In some embodiments, the β 1,3-galactosyltransferase of the invention is identical to the β 1,3-galactosyltransferase from OH 4384 or NCTC 11168 from C. jejuni over the full length of the J61,3-galactosyltransferase.
The percentage of identities can be determined by inspection, for example, or it can be determined using an alignment algorithm such as the BLASTP version 2.0 algorithm with a word length (W) of 3, G = 11, E = 1 , and a BLOSUM62 substitution matrix.
Laj61,3-galactosyltransferase of the invention can be obtained from the respective Campylobacter species, or it can be produced recombinantly. Galactosyltransferases can be identified by enzyme activity assays, for example, or by detecting specific immunoreactivity with antibodies raised against C. OH4384 β 1,3-galactosyltransferase. jejuni having an amino acid sequence as set forth in SEQ ID NO: 15 or C. jejuni NCTC 11168 β 1,3-galactosyltransferase as set forth in SEQ ID NO: 17.
Four. Additional enzymes involved in the LOS biosynthetic pathway
The present invention also provides additional enzymes that are involved in the biosynthesis of oligosaccharides such as those found on bacterial lipooligosaccharides. For example, the enzymes involved in the synthesis of CMP-sialic acid, the donor for sialyltransferases, are provided. A sialic acid synthase is encoded by an 8a open reading frame (ORF) of C. jejuni (SEQ ID NO: 21) and by means of an 8b open reading frame of strain NCTC 11168 (see, Table 3). Another enzyme involved in the synthesis of sialic acid is encoded by ORF 9a of OH4384 (SEQ ID NO: 22) and 9b of NCTC 11168. A sialic acid-CMP synthetase is encoded by ORF 10a (SEQ ID NO: 23) and 10b of OH4384 and NCTC 11168, respectively.
The invention also provides an acyltransferase that is involved in lipid A biosynthesis. This enzyme is encoded by means of an open reading frame 2a of C. jejuni strain OH4384 (SEQ ID NO: 18) and by means of a frame 2B open reader from strain NCTC 11168. An acetyltransferase is also provided; this enzyme is encoded by an ORF 11a from strain OH4384 (SEQ ID NO: 24); no homologue is found at the LOS biosynthesis locus of strain NCTC 11168.
Three additional glycosyltransferases are also provided. These enzymes are encoded by ORF 3a (SEQ ID NO: 19), 4a (SEQ ID NO: 20), and 12a (SEQ ID NO: 25) of strain OH4384 and ORF 3b, 4b and 12b of strain NCTC 11168.
The invention includes, for each of these enzymes, polypeptides that include an amino acid sequence that is approximately at least 75% identical to an amino acid sequence as set forth herein over a region approximately at least 50 amino acids in length. More preferably, the enzymes of the invention are approximately at least 85% identical to the respective amino acid sequence, and even more preferably approximately at least 95% identical to the amino acid sequence, over a region of approximately at least 50 amino acids in length. In presently preferred embodiments the region of percent identity spans a region longer than 50 amino acids, more preferably over a region of about at least 100 amino acids, and most preferably over the entire length of the enzyme. Therefore, the enzymes of the invention include polypeptides that have the respective activity and are approximately at least 65% identical, more preferably approximately at least 70% identical, more preferably approximately at least 80% identical, and most preferably about at least 90% identical to the amino acid sequence of the corresponding enzyme as set forth herein over a region of the polypeptide that is required to retain the respective enzyme activity. In some embodiments, the enzymes of the invention are identical to the corresponding C. jejuni OH4384 enzymes over the full length of the enzyme.
B. Nucleic acids encoding glycosyltransferases and related enzymes
The present invention also provides isolated and / or recombinant nucleic acids encoding the glycosyltransferases and other enzymes of the invention. The nucleic acids of the invention encoding glycosyltransferase are useful for a variety of purposes, including recombinant expression of corresponding glycosyltransferase polypeptides, and as probes to identify nucleic acids encoding other glycosyltransferases, and to study regulation and the expression of enzymes.
Nucleic acids of the invention include those that encode a full-length glycosyltransferase enzyme such as those described above, as well as those that encode a subsequence of a glycosyltransferase polypeptide. For example, the invention includes nucleic acids that encode a polypeptide that is not the full length of the enzyme, but nevertheless has glycosyltransferase activity. The nucleotide sequences of the LOS locus of C. jejuni strain OH4384 are provided herein as SEQ ID NO: 1, and are
ES 2 269 098 T3 identify the respective reading frames. Additional nucleotide sequences are also provided, as discussed below. The invention includes not only nucleic acids that include nucleotide sequences as set forth herein, but also nucleic acids that are substantially identical to, or substantially complementary to, the exemplified embodiments. For example, the invention includes nucleic acids that include a nucleotide sequence that is approximately at least 70% identical to one that is disclosed herein, more preferably at least 75%, even more preferably at least 80%, more preferably at least 85%. %, still more preferably at least 90%, and still more preferably about at least 95% identical to an exemplified nucleotide sequence. The region of identity spans about at least 50 nucleotides, more preferably about at least 100 nucleotides, even more preferably about at least 500 nucleotides. The region of a specified percent identity, in some embodiments, encompasses the coding region of a sufficient portion of the encoded enzyme to retain the respective enzyme activity. The specified percent identity, in preferred embodiments, spans the entire length of the coding region of the enzyme.
The nucleic acids of the invention encoding glycosyltransferases can be obtained using methods that are known to those skilled in the art. Suitable nucleic acids (eg, cDNA, genomic, or subsequences (probes)) can be cloned, or amplified by means of in vitro methods such as the polymerase chain reaction (PCR), the ligase chain reaction (LCR), the transcription-based amplification systems (TAS), the self-sustained sequence replication system (SSR). A wide variety of in vitro cloning and amplification methodologies are known to trained persons. Examples of these techniques and sufficient instructions to guide trainees through many cloning exercises are found in Berger and Kimmel, Guide to Molecular Cloning Techniques, Methods in Enzymology 152 Academic Press, Inc., San Diego, CA (Berger) ; Sambrook et al. (1989) Molecular Cloning - A Laboratory Manual (Second Edition) Vol. 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor Press, NY, (Sambrook et al.); Current Protocols in Molecular Biology, FM Ausubel et al., Eds., Current Protocols, a joint venture between Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., (1994 Supplement) (Ausubel); Cashion et al., US Patent No. 5,017,478; and Carr, European Patent No. 0,246,864. Examples of techniques sufficient to direct trainees through in vitro amplification methods are found in Berger, Sambrook, and Ausubel, as well as in Mullis et al., (1987) US Patent No. 4,683,202; PCR Protocols A Guide to Methods and Applications (Innis et al., Eds) Academic Press Inc. San Diego, CA (1990) (Innis); Arnheim & Levinson (October 1, 1990) C&EN 36-47; The Journal OfNIH Research (1991) 3: 81-94; (Kwoh et al. (1989) Proc. Natl. Acad. Sci. United States of America 86: 1173; Guatelli et al. (1990) Proc. Natl. Acad. Sci. United States of America 87, 1874; Lomell et al. (1989 ) J. Clin. Chem., 35: 1826; Landegren et al., (1988) Science 241: 1077-1080; Van Brunt (1990) Biotechnology 8: 291-294; Wu and Wallace (1989) Gene 4: 560; and Barringer et al. (1990) Gene 89: 117. Improved methods of cloning amplified nucleic acids in vitro are described in Wallace et al., US Patent No. 5,426,039.
The nucleic acids of the invention encoding the glycosyltransferase polypeptides, or the subsequences of these nucleic acids, can be prepared by any suitable method as described above, including, for example, cloning and restricting the appropriate sequences. . As an example, a nucleic acid encoding a glycosyltransferase of the invention can be obtained by means of routine cloning methods. A known nucleotide sequence of a gene encoding the glycosyltransferase of interest, as described herein, can be used to provide probes that specifically hybridize to a gene encoding a suitable enzyme in a genomic DNA sample, or to an mRNA. in a total RNA sample (eg, on a Southern or Northern blot). Preferably, the samples are obtained from prokaryotic organisms, such as the Campylobacter species. Examples of the Campylobacter species of particular interest include C. jejuni. Many C. jejuni O: 19 strains synthesize ganglioside mimics and are useful as a source of the glycosyltransferase of the invention.
Once the target glycosyltransferase nucleic acid is identified, it can be isolated according to standard methods known to those skilled in the art (see, for example, Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual, 2nd Ed. , Vol. 1-3, Cold Spring Harbor Laboratory; Berger and Kimmel (1987) Methods in Enzymology, Vol. 152: Guide to Molecular Cloning Techniques, San Diego: Academic Press, Inc .; or Ausubel et al. (1987) Current Protocols in Molecular Biology, Greene Publishing and Wiley-Interscience, New York).
A nucleic acid encoding a glycosyltransferase of the invention can also be cloned to detect its expressed product by means of assays based on physical, chemical or immunological properties. For example, a nucleic acid encoding the cloned bifunctional sialyltransferase can be identified by the ability of a polypeptide encoded by the nucleic acid to catalyze the coupling of a sialic acid at an α2,3 bond to a galactosylated acceptor, followed by coupling a second sialic acid residue to the first sialic acid at an α2,8 bond. Similarly, a cloned nucleic acid encoding a / 11,4-GalNAc transferase or a j61,3-galactosyltransferase can be identified by the ability of the encoded polypeptide to catalyze the transfer of a GalNAc residue from UDP-GalNAc, or a galactose residue of UDP-Gal, respectively, to a suitable acceptor. Suitable test conditions are known in the art, and include those described in the Examples. Other physical properties of a polypeptide expressed from a particular nucleic acid can be compared to the properties of known glycosyltransferase polypeptides.
ES 2 269 098 T3 of the invention, such as those described herein, to provide another method of identifying nucleic acids encoding the glycosyltransferases of the invention. Alternatively, a putative gene for a glycosyltransferase can be mutated, and its role as a glycosyltransferase established by detecting a variation in the ability to produce the respective glycoconjugate.
In other embodiments, nucleic acids encoding glycosyltransferase can be cloned using DNA amplification methods such as polymerase chain reaction (PCR). Therefore, for example, the nucleic acid sequence or subsequence is amplified by means of PCR, preferably using a sense primer that contains one restriction site (e.g. Xbal) and an antisense primer that contains other restriction sites ( eg HindIII). This will produce a nucleic acid encoding the amino acid sequence or subsequence of the desired glycosyltransferase and having terminal restriction sites. This nucleic acid can then be easily ligated into a vector containing a nucleic acid encoding the second molecule and having the appropriate corresponding restriction sites. Appropriate primers for PCR can be determined by one of ordinary skill in the art using the sequence information provided here. Appropriate restriction sites can also be added to the nucleic acid encoding the glycosyltransferase of the invention, or amino acid subsequence, by site-directed mutagenesis. The plasmid containing the nucleotide sequence or subsequence encoding glycosyltransferase is cleaved with the appropriate restriction endonuclease and then ligated into an appropriate vector for amplification and / or expression according to standard methods.
Examples of suitable primers suitable for the amplification of the nucleic acids of the invention encoding glycosyltransferase are shown in Table 2; some of the primer pairs are designed to provide a 5 'Ndel restriction site and a 3' Sall site on the amplified fragment. The plasmid containing the sequence or subsequence encoding the enzyme is cleaved with the appropriate restriction endonuclease and then ligated into an appropriate vector for amplification and / or expression according to standard methods.
As an alternative to cloning a nucleic acid encoding glycosyltransferase, an appropriate nucleic acid can be chemically synthesized from a known sequence encoding a glycosyltransferase of the invention. Direct methods of chemical synthesis include, for example, the phosphodiester method of Narang et al. (1979) Meth. Enzymol. 68: 90-99; the phosphodiester method of Brown et al. (1979) Meth. Enzymol. 68: 109-151; the diethylphosphoramidite method of Beaucage et al. (1981) Tetra. Lett., 22: 18591862; and the solid support method of US Patent No. 4,458,066. Chemical synthesis produces a single-stranded oligonucleotide. This can be converted into double-stranded DNA by hybridization with a complementary sequence, or by polymerization with a DNA polymerase using the single-stranded template. One trained would recognize that while chemical DNA synthesis is often limited to sequences of approximately 100 bases, longer sequences can be obtained by ligation of shorter sequences. Alternatively, the subsequences can be cloned and the appropriate subsequences excised using the appropriate restriction enzymes. The fragments can then be ligated to produce the desired DNA sequence.
In some embodiments, it may be desirable to modify the nucleic acids encoding the enzyme. Someone trained will recognize many ways to generate alterations in a given nucleic acid construct. Such well-known methods include site-directed mutagenesis, PCR amplification using degenerate oligonucleotides, exposure of nucleic acid-containing cells to radiation or mutagenic agents, chemical synthesis of a desired oligonucleotide (e.g., in conjunction with ligation and / or cloning to generate large nucleic acids) and other well known techniques. See, for example, Gilman and Smith (1979) Gene 8: 81-97, Roberts et al. (1987) Nature 328: 731-734.
In a current preferred embodiment, the recombinant nucleic acids present in the cells of the invention are modified to provide preferred codons that enhance nucleic acid translation in a selected organism (e.g., E. coli preferred codons are substituted within a coding nucleic acid for expression in E. coli).
The present invention includes nucleic acids that are isolated (that is, not in their native chromosomal location) and / or are recombinant (that is, modified from their original form, present in a non-native organism, etc.).
1. Sialyltransferases
The invention provides nucleic acids encoding sialyltransferases such as those described above. In some embodiments, the nucleic acids of the invention encode bifunctional sialyltransferase polypeptides that have both α2,3 sialyltransferase activity and α2,8 sialyltransferase activity. These sialyltransferase nucleic acids encode a sialyltransferase polypeptide having an amino acid sequence that is approximately at least 76% identical to an amino acid sequence as set forth in SEQ ID NO: 3 over a region of approximately at least 60 amino acids in length. More preferably, the sialyltransferases encoded by the nucleic acids of the invention are approximately at least 85% identical to the amino acid sequence of SEQ ID NO: 3, and even more preferably approximately at least 95% identical to the amino acid sequence of the SEQ ID NO: 3, over a region of at least 60 amino acids in length. In presently preferred embodiments, the region of percent identity spans a region greater than
ES 2,269,098 T3 amino acids, more preferably over a region of about at least 100 amino acids, and most preferably over the full length of the sialyltransferase. In the presently preferred embodiment, the nucleic acids encoding the sialyltransferase of the invention encode a polypeptide having the amino acid sequence shown in SEQ ID NO: 3.
An example of a nucleic acid of the invention is an isolated and / or recombinant form of a nucleic acid encoding C. jejuni OH4384 bifunctional sialyltransferase. The nucleotide sequence of this nucleic acid is shown in SEQ ID NO: 2. The sequences of the polynucleotide of the invention encoding the sialyltransferase are typically about at least 75% identical to the nucleic acid sequence of SEQ ID NO: 2 over a region of about at least 50 nucleotides in length. More preferably, the nucleic acids of the invention encoding the sialyltransferase are approximately at least 85% identical to this nucleotide sequence, and even more preferably they are approximately at least 95% identical to the nucleotide sequence of SEQ ID NO: 2 , over a region of at least 50 amino acids in length. In presently preferred embodiments, the specified percent identity threshold region extends over a region greater than 50 nucleotides, more preferably over a region of about at least 100 nucleotides, and more preferably over the full length of the region for coding. of sialyltransferase. Therefore, the invention provides nucleic acids encoding bifunctional sialyltransferase that are substantially identical to those cstll of OH4384 from C. jejuni strain as set forth in SEQ ID NO: 2 or strain 0:10 (SEQ ID NO: 4).
Other nucleic acids of the invention that encode sialyltransferase encode sialyltransferases that have α2,3 sialyltransferase activity but lack substantial α2,8 sialyltransferase activity. For example, nucleic acids encoding a Cstll α2,3 sialyltransferase from sero-strain O: 19 of C. jejuni (SEQ ID NO: 8) and NCTC 11168, are provided by the invention; these enzymes have little or no α2,8-sialyltransferase activity (Table 6).
To identify the nucleic acids of the invention, visual inspection can be used, or a suitable alignment algorithm can be used. An alternative method by which a nucleic acid of the invention encoding bifunctional sialyltransferase can be identified is by hybridization, under stringent conditions, of the nucleic acid of interest to a nucleic acid that includes a polynucleotide sequence of a sialyltransferase like the one discussed here.
2.11,4-GalNAc transferases
The invention also provides nucleic acids that include polynucleotide sequences that encode a GalNAc transferase polypeptide having a / 11,4-GalNAc transferase activity. Polynucleotide sequences encoding a GalNAc transferase polypeptide having an amino acid sequence that is approximately at least 70% identical to the C. OH4384 / 11,4-GalNAc transferase. jejuni, which has an amino acid sequence as set forth in SEQ ID NO: 13, over a region approximately at least 50 amino acids in length. More preferably, the GalNAc transferase polypeptide encoded by the nucleic acids of the invention is approximately at least 80% identical to this amino acid sequence, and even more preferably approximately at least 90% identical to the amino acid sequence of SEQ ID NO: 13, over a region at least 50 amino acids long. In presently preferred embodiments, the region of percent identity spans a region greater than 50 amino acids, more preferably over a region of about at least 100 amino acids, and most preferably, over the full length of the GalNAc transferase polypeptide. In a presently preferred embodiment, the nucleic acids of the invention that encode the GalNAc transferase polypeptide, encode a polypeptide having the amino acid sequence as shown in SEQ ID NO: 13. To identify nucleic acids of the invention, visual inspection can be used, or an appropriate alignment algorithm can be used.
An example of a nucleic acid of the invention encoding GalNAc transferase is an isolated and / or recombinant form of the nucleic acid encoding C. jejuni OH4384 GalNAc transferase. This nucleic acid has a nucleotide sequence as shown in SEQ ID NO: 12. The sequences of the polynucleotides of the invention that encode GalNAc transferase are typically about at least 75% identical to the nucleic acid sequence of SEQ ID NO: 12 over a region of about at least 50 nucleotides in length. More preferably, the nucleic acids of the invention encoding GalNAc transferase are approximately at least 85% identical to this nucleotide sequence, and even more preferably approximately at least 95% identical to the nucleotide sequence of SEQ ID NO: 12 , over a region of at least 50 amino acids in length. In presently preferred embodiments, the region of percent identity spans a region greater than 50 nucleotides, more preferably over a region of about at least 100 nucleotides, and most preferably, over the full length of the region encoding the GalNAc transferase.
To identify the nucleic acids of the invention, visual inspection can be used, or a suitable alignment algorithm can be used. An alternative method by which a nucleic acid of the invention encoding a GalNAc transferase can be identified is through hybridization, under stringent conditions, of the nucleic acid of interest to a nucleic acid that includes a polynucleotide sequence of SEQ ID NO: 12.
ES 2 269 098 T3
3. fl1,3-Galactosyltransferases
The invention also provides nucleic acids including polynucleotide sequences that encode a polypeptide having j61,3-galactosyltransferase (CgtB) activity. The β1,3-galactosyltransferase polypeptides encoded by these nucleic acids of the invention preferably include an amino acid sequence that is approximately at least 75% identical to a P 1,3-galactosyltransferase from C. strain OH4384. jejuni as disclosed in SEQ ID NO: 15, or as that of a β 1,3-galactosyltransferase from strain NCTC 11168 as disclosed in SEQ ID NO: 17, over a region approximately at least 50 amino acids long . More preferably, the galactosyltransferase polypeptides encoded by these nucleic acids of the invention are approximately at least 85% identical to this amino acid sequence, and even more preferably they are approximately at least 95% identical to the amino acid sequence of SEQ ID NO. : 15 or SEQ ID NO: 17, over a region of at least 50 amino acids in length. In presently preferred embodiments, the region of percent identity spans a region greater than 50 amino acids, more preferably over a region of about at least 100 amino acids, and most preferably, over the full length of the polypeptide-encoding region. galactosyltransferase.
An example of a nucleic acid of the invention encoding a j61,3-galactosyltransferase is an isolated and / or recombinant form of the nucleic acid encoding C. jejuni OH4384 β 1,3-galactosyltransferase. This nucleic acid includes a nucleotide sequence as shown in SEQ ID NO: 14. Another nucleic acid that encodes suitable j61,3-galactosyltransferase includes a nucleotide sequence from a C. jejuni, for which the nucleotide sequence is shown in SEQ ID NO: 16. The polynucleotide sequences of the invention encoding laj61,3-galactosyltransferase are typically at least 75% identical to the nucleic acid sequence of the SEQ ID NO: 14 or that of SEQ ID NO: 16 over a region approximately at least 50 nucleotides in length. More preferably, the nucleic acids of the invention encoding j61,3-galactosyltransferase are about at least 85% identical to at least one of these nucleotide sequences, and even more preferably, about at least 95% identical to the sequences of nucleotides of SEQ ID NO: 14 and / or SEQ ID NO: 16, over a region of at least 50 amino acids in length. In presently preferred embodiments, the region of percent identity spans a region greater than 50 nucleotides, more preferably over a region of at least 100 nucleotides, and most preferably over the full length of the j61 coding region. , 3-galactosyltransferase.
To identify the nucleic acids of the invention, visual inspection can be used, or a suitable alignment algorithm can be used. An alternative method by which a nucleic acid of the invention encoding the galactosyltransferase polypeptide can be identified is by hybridization, under stringent conditions, of the nucleic acid of interest to a nucleic acid that includes a polynucleotide sequence. of SEQ ID NO: 14 or SEQ ID NO: 16.
Four. Additional enzymes involved in the LOS biosynthetic pathway
Also provided are nucleic acids that encode other enzymes that are involved in the prokaryotic LOS biosynthetic pathway such as Campylobacter. These nucleic acids encode enzymes such as, for example, sialic acid synthase, which is encoded by means of an 8a open reading frame (ORF) of the C. OH4384 strain. jejuni and by means of an open reading frame 8b of the strain NCTC 11168 (see, Table 3), another enzyme involved in the synthesis of sialic acid that is encoded by ORF 9a of OH4384 and 9b of NCTC 11168, and a sialic acidCMP synthetase that is encoded by ORF 10a and 10b of OH4384 and NCTC 11168, respectively.
The invention also provides nucleic acids that encode an acyltransferase that is involved in lipid A biosynthesis. This enzyme is encoded by means of an open reading frame 2a of C. jejuni strain OH4384 and by means of a reading frame open 2B of strain NCTC 11168. Nucleic acids encoding an acetyltransferase are also provided; this enzyme is encoded by ORF 11a from strain OH4384; no homolog is found at the LOS biosynthesis locus of strain NCTC 11168.
Nucleic acids encoding three additional glycosyltransferases are also provided. These enzymes are encoded by ORFs 3a, 4a, and 12a of strain OH4384 and ORFs 3b, 4b and 12b of strain NH 11168 (Figure 1).
C. Glycosyltransferase Expression and Expression Cassettes
The present invention also provides expression cassettes, expression vectors, and recombinant host cells that can be used to produce the glycosyltransferases and other enzymes of the invention. A typical expression cassette contains a promoter operably linked to a nucleic acid encoding glycosyltransferase or other enzymes of interest. Expression cassettes are typically included on expression vectors that are introduced into suitable host cells, preferably prokaryotic host cells. More than one glycosyltransferase polypeptide can be expressed in a single host cell by placing multiple transcription cassettes in a single expression vector, by constructing a gene encoding a fusion protein consisting of more than one glycosyltransferase, or through the use of different expression vectors for each glycosyltransferase.
ES 2 269 098 T3
In a preferred embodiment, the expression cassettes are useful for the expression of glycosyltransferases in prokaryotic host cells. Commonly used prokaryotic control sequences, which are defined herein to include promoters for transcription initiation, optionally with an operator, along with ribosome binding site sequences, include commonly used promoters such as beta promoter systems. -lactamase (penicillinase) and lactose (lac) (Change et al., Nature (1977) 198: 1056), to the tryptophan (trp) promoter system (Goeddel et al., Nucleic Acids Res. (1980) 8: 4057), to the tac promoter (DeBoer, et al., Proc. Natl. Acad. Sci. United States of America (1983 ) 80: 21-25); and promoter P<sub>L</sub> derived from lambda and to the ribosome binding site in the N gene (Shimatake et al., Nature (1981) 292: 128). The particular promoter system is not critical to the invention, and any available promoter that works in prokaryotes can be used.
Either of the two promoters, constitutive or regulated, can be used in the present invention. Regulated promoters can be convenient because host cells can be grown to high densities before expression of glycosyltransferase polypeptides is induced. The high level of expression of heterologous proteins shows cell growth in some situations. Regulated promoters especially suitable for use in E. coli include the bacteriophage lambda P promoter<sub>L</sub>, to the hybrid trp-lac promoter (Amann et al., Gene (1983) 25: 167; de Boer et al., Proc. Natl. Acad. Sci. United States of America (1983) 80: 21, and to the bacteriophage T7 promoter ( Studier et al., J. Mol. Biol. (1986); Tabor et al., (1985) These promoters and their use are discussed in Sambrook et al., Supra. A currently preferred regulatable promoter is the dual tac-gal promoter, which described in PCT / US97 / 20528 (International Publication No. WO 9820111).
For the expression of glycosyltransferase polypeptides in prokaryotic cells other than those of E. coli, a promoter is required that functions in the particular prokaryotic species. Such promoters can be obtained from genes that have been cloned from the species, or heterologous promoters can be used. For example, a hybrid trp-lac promoter works in Bacillus in conjunction with E. coli. Promoters suitable for use in eukaryotic host cells are well known to those of skill in the art.
A ribosome binding site (RBS) is conveniently included in the expression cassettes of the invention that are intended for use in prokaryotic host cells. An RBS in E. coli, for example, consists of a nucleotide sequence 3-9 nucleotides in length located at nucleotides 3-11 upstream of the initiation codon (Shine and Dalgarno, Nature (1975) 254: 34; Steitz , In Biological regulation and development: Gene expression (ed. RF Goldberger), vol. 1, p. 349, 1979, Plenum Publishing, NY).
Coupling for translation can be used to enhance expression. The strategy uses an upstream short open reading frame derived from a native highly expressed gene for the translation system, which is placed downstream of the promoter, and a ribosome binding site followed after a few amino acid codons by a codon. termination. Just before the stop codon is a second ribosome binding site, and after the stop codon is a start codon for translation initiation. The system dissolves the secondary structure in RNA, allowing efficient initiation of translation. See, Squires et al. (1988) J. Biol. Chem. 263: 16297-16302.
The glycosyltransferase polypeptides of the invention can be expressed intracellularly, or can be secreted from the cell. Intracellular expression often results in high yields. If necessary, the amount of soluble active glycosyltransferase polypeptides can be increased by means of replication procedures (see, for example, Sambrook et al., Supra .; Marston et al., Biol Technology (1984) 2: 800; Schoner et al. , Biol Technology (1985) 3: 151). In embodiments in which glycosyltransferase polypeptides are secreted from the cell, either within the periplasm or within the extracellular medium, the polynucleotide sequence encoding the glycosyltransferase is linked to a polynucleotide sequence encoding a cleavable sequence. of the signal peptide. The signal sequence directs the transposition of the glycosyltransferase polypeptide across the cell membrane (see, for example, Sambrook et al., Supra .; Oka et al., Proc. Natl. Acad. Sci. United States of America (1985) 82 : 7212; Talmadge et al., Proc. Natl. Acad. Sci. United States of America (1980) 77: 3988; Takahara et al., J. Biol. Chem. (1985) 260: 2670).
The glycosyltransferase polypeptides of the invention can also be produced as fusion proteins. This approach often results in high yields, because normal prokaryotic control sequences direct transcription and translation. In E. coli, lacZ fusions are often used to express heterologous proteins. Suitable vectors are readily available, such as the pUR, pEX, and pMR100 series (see, eg, Sambrook et al., Supra). For certain applications, it may be desirable to cleave the non-glycosyltransferase amino acids from the fusion protein after purification. This can be achieved by any of the different methods known in the art, including cleavage by means of cyanogen bromide, a protease, or by means of Factor X<sub>to</sub> (see, for example, Sambrook et al., supra .; Itakura et al., Science (1977) 198: 1056; Goeddel et al., Proc. Natl. Acad. Sci. United States of America (1979) 76: 106; Nagai and co-workers, Nature (1984) 309: 810; Sung et al., Proc. Natl. Acad. Sci. United States of America (1986) 83: 561). Cleavage sites can be engineered within the gene by the fusion protein at the desired point of cleavage.
ES 2 269 098 T3
A suitable system for obtaining recombinant proteins from E. coli that maintains the integrity of its N-terminals has been described by Miller et al., Biotechnology 7: 698-704 (1989). In this system, the gene of interest is produced as a C-terminal fusion to the first 76 residues of the yeast ubiquitin gene that contains a peptidase cleavage site. Cleavage at the junction of the two fractions results in the production of a protein having an intact authentic N-terminal residue.
The glycosyltransferases of the invention can be expressed in a variety of host cells, including E. coli, other bacterial hosts, yeast, and different higher eukaryotic cells such as COS, CHO and HeLa cell lines and myeloma cell lines. Examples of useful bacteria include, but are not limited to, Escherichia, Enterobacter, Azotobacter, Erwinia, Bacillus, Pseudomonas, Klebsielia, Proteus, Salmonella, Serratia, Shigella, Rhizobia, Vitreoscilla, and Paracoccus. The nucleic acid encoding the recombinant glycosyltransferase is operably linked to the appropriate expression control sequences for each host. For E. coli, this includes a promoter such as T7, trp or lambda promoters, a ribosome binding site and, preferably, a transcription termination signal. For eukaryotic cells, the control sequences will include a promoter and preferably an enhancer derived from the genes for immunoglobin, SV40, cytomegalovirus, etc., and a polyadenylation sequence, and may include a splice donor and acceptor sequences.
The expression vectors of the invention can be transferred into selected host cells by well known methods such as transformation by calcium chloride for E. coli and calcium phosphate treatment or electroporation for mammalian cells. . Cells transformed by the plasmids can be selected by means of the resistance to antibiotics conferred by the genes contained in the plasmids, such as the amp, gpt, neo and hyp genes.
Once expressed, recombinant glycosyltransferase polypeptides can be purified according to standard state-of-the-art procedures, including ammonium sulfate precipitation, affinity columns, column chromatography, gel electrophoresis, and the like (generally see R Scopes, Protein Purification, Springer-Verlag, NY (1982), Deutscher, Methods in Enzymology Vol. 182: Guide to Protein Purification, Academic Press, Inc. NY (1990)). Substantially preferred are pure compositions of about at least 90 to 95% homogeneity, and 98 to 99% or more homogeneity. Once purified, partially or to homogeneity as desired, the polypeptides can then be used (eg, as immunogens for the production of antibodies). Glycosyltransferases can also be used in a non-purified or semi-purified state. For example, a host cell expressing glycosyltransferase can be used directly in a glycosyltransferase reaction, either with or without processing such as permeabilization or other cell disruption.
One skilled in the art would recognize that modifications can be made to the glycosyltransferase proteins without diminishing their biological activity. Some modifications can be made to facilitate cloning, expression, or incorporation of the target molecule into a fusion protein. Such modifications are well known to those skilled in the art, and include, for example, the addition of methionine to the amino terminal to provide an initiation site, or additional amino acids (eg, poly His) placed on either terminal to create restriction sites. conveniently located either stop codons or purification sequences.
D. Methods and reaction mixtures for the synthesis of oligosaccharides
The invention provides reaction mixtures and methods in which the glycosyltransferases of the invention are used to prepare the desired oligosaccharides (which are composed of two or more saccharides). The glycosyltransferase reactions of the invention take place in a reaction medium containing at least one glycosyltransferase, a donor substrate, an acceptor sugar, and typically, a soluble divalent metal cation. The methods rely on the use of glycosyltransferase to catalyze the addition of a saccharide to a saccharide substrate (also called an "acceptor"). A number of methods are known to use glycosyltransferases to synthesize the desired oligosaccharide structures. Examples of methods are described, for example, in WO 96/32491, Ito et al. (1993) Pure Appl. Chem. 65: 753, and in US Patent Nos. 5,352,670, 5,374,541, and 5,545,553.
For example, the invention provides methods for adding sialic acid at an α2,3 bond to a galactose residue, contacting a reaction mixture containing an activated sialic acid (eg, CMP-NeuAc, CMP-NeuGc, and similar) to an acceptor moiety that includes a terminal galactose residue in the presence of a bifunctional sialyltransferase of the invention. In presently preferred embodiments, the methods also result in the addition of a second sialic acid residue that is linked to the first sialic acid via an α2,8 bond. The product of this method is Siaa2,8-Siaa2,3-Gal-. Examples of suitable acceptors include a terminal Gal that is linked to GlcNAc or Glc via a / 11.4 linkage, and a terminal Gal that is / 11.3 linked to either GlcNAc or GalNAc. The terminal residue to which sialic acid binds can itself bind to, for example, H, a saccharide, oligosaccharide, or an aglycone group having at least one carbohydrate atom. In some embodiments, the acceptor residue is a portion of an oligosaccharide that is linked to a protein, a lipid, or a proteoglycan, for example.
In some embodiments, the invention provides reaction mixtures and methods for the synthesis of gangliosides, lysogangliosides, ganglioside mimics, lysoganglioside mimics, or the carbohydrate portions of these molecules. These reaction methods and mixtures typically include as the galactosylated acceptor moiety a compound having a formula selected from the group consisting of Gal4Glc-R<sup>1</sup> and Gal3GalNAc-R<sup>2</sup>; where
ES 2 269 098 T3
R<sup>1</sup> is selected from the group consisting of ceramide or other glycolipids, R<sup>2</sup> is selected from the group consisting of Gal4GlcCer, (Neu5Ac3) Gal4GlcCer, and (Neu5Ac8Neu5c3) Gal4GlcCer. For example, for ganglioside synthesis the galactosylated acceptor can be selected from the group consisting of Gal4GlcCer, Gal3GalNAc4 (Neu5Ac3) Gal4GlcCer, and Gal3GalNAc4 (Neu5Ac8Neu5c3) Gal4GlcCer.
The methods and reaction mixtures of the invention are useful for producing any of a large number of gangliosides, lysogangliosides, and related structures. Many gangliosides of interest are described in Oettgen, HF, ed., Gangliosides and Cancer, VCH, Germany, 1989, pages 10-15, and the references cited there. Gangliosides of particular interest include, for example, those found in the brain as well as other sources that are listed in Table 1.
TABLE 1
Ganglioside Formulas and Abbreviations
Structure
Abbreviation
Neu5Ac3Gal4GlcCer GM3
GalNAc4 (Neu5Ac3) Gal4GlcCer GM2
Gal3GalNAc4 (Neu5Ac3) Gal4GlcCer GMla
Neu5Ac3Gal3GalNAc4Gal4GlcCer GMlb
Neu5Ac8Neu5Ac3Gal4GlcCer GD3
GalNAc4 (Neu5Ac8Neu5Ac3) Gal4GlcCer GD2
Neu5Ac3Gal3GalNAc4 (Neu5Ac3) Gal4GlcCer GDla
Neu5Ac3Gal3 (Neu5Ac6) GalNAc4Gal4GlcCer GDla
Gal3GalNAc4 (Neu5Ac8Neu5Ac3) Gal4GlcCer GDlb
Neu5Ac8Neu5Ac3Gal3GalNAc4 (Neu5Ac3) Gal4GlcCer GTla
Neu5Ac3Gal3GalNAc4 (Neu5Ac8Neu5Ac3) Gal4GlcCer GTlb
Gal3GalNAc4 (Neu5 Ac8Neu5Ac8Neu5 Ac3) Gal4GlcCer GT1 c
Neu5Ac8Neu5Ac3Gal3GalNAc4 (Neu5Ac8Neu5c3) Gal4GlcCer_GQlb
Glycolipid Nomenclature, IUPAC-IUB Joint Commission on Biochemical Nomenclature (1997 Recommendations); PureAppl. Chem. (1997) 69: 2475-2487; Eur, J. Biochem (1998) 257: 293-298) (www.chem.qmw.ac.uk/iupac/misc/glvlp.html).
The bifunctional sialyltransferases of the invention are particularly useful for the synthesis of the gangliosides GD1a, GD1b, GT1a, GT1b, GT1c and GQ1b, or the carbohydrate portions of these gangliosides, for example. The structures of these gangliosides, which are shown in Table 1, require both α2,3-sialyltransferase and α2,8-sialyltransferase activity. An advantage provided by the methods and reaction mixtures of the invention is that both activities are present in a single polypeptide.
The glycosyltransferases of the invention can be used in combination with additional glycosyltransferases and with other enzymes. For example, a combination of sialyltransferase and galactosyltransferases can be used. In some embodiments of the invention, the galactosylated acceptor that is utilized by the bifunctional sialyltransferase is formed by contacting a suitable acceptor with UDP-Gal and a galactosyltransferase. The galactosyltransferase polypeptide, which can be one as described herein, transfers the Gal residue from the UDP-Gal to the acceptor.
Similarly, the / 11,4-GalNAc transferases of the invention can be used to synthesize an acceptor for galactosyltransferase. For example, the saccharide acceptor for galactosyltransferase can be formed by contacting an acceptor for a GalNAc transferase with UDP-GalNAc and a GalNAc transferase polypeptide, wherein the GalNAc transferase polypeptide transfers the GalNAc residue from UDP-GalNAc up to the acceptor for GalNAc transferase.
In this group of embodiments, the enzymes and substrates can be combined in an initial reaction mixture, or the enzymes and reagents for a second cycle of glycosyltransferase can be added to the reaction medium after the first cycle of glycosyltransferase it has almost been completed. By conducting two cycles of glycosyltransferase in sequence in a single vessel, overall yields are improved over postprocedures in which an intermediate species is isolated. Additionally, cleaning and disposal of extra solvents and by-products is reduced.
ES 2 269 098 T3
The products produced by the above processes can be used without purification. However, it is usually preferred to recover the product. Known standard techniques can be used for the recovery of glycosylated saccharides such as thin layer or thick layer chromatography, or ion exchange chromatography. The use of membrane filtration is preferred, more preferably using a reverse osmotic membrane, or one or more column chromatographic techniques for recovery.
E. Uses of Glycoconjugates Produced Using Glycosyltransferases and Methods of the Invention
The oligosaccharide compounds that are made using the glycosyltransferases and methods of the invention can be used in a variety of applications, for example, as antigens, diagnostic reagents, or as therapeutics. Therefore, the present invention also provides pharmaceutical compositions that can be used in the treatment of a variety of conditions. The pharmaceutical compositions contain oligosaccharides made according to the methods described above.
The pharmaceutical compositions of the invention are suitable for use in a variety of drug delivery systems. Formulations suitable for use in the present invention are found in Remington's Pharmaceutical Sciences, Mace Publishing Company, Philadelphia, PA, 17th edition, (1985). For a brief review of drug delivery methods, see, Langer, Science 249: 1527-1533 (1990).
The pharmaceutical compositions are intended for parenteral, intranasal, topical, oral or local administration, such as by means of an aerosol or in transdermal form, for prophylactic and / or therapeutic treatment. Commonly, pharmaceutical compositions are administered parenterally, eg, intravenously. Thus the invention provides compositions for parenteral administration comprising the compound dissolved or suspended in an acceptable vehicle, preferably an aqueous vehicle, eg, water, buffered water, physiological saline, PBS, and the like. The compositions may contain pharmaceutically acceptable adjuncts as required for approximate physiological conditions, such as pH adjusting agents and buffers, tonicity adjusting agents, wetting agents, detergents, and the like.
These compositions can be sterilized by conventional sterilization techniques, or they can be sterilized by filtration. The resulting aqueous solutions can be packaged for use as is, or lyophilized, the lyophilizate preparation being combined with a sterile aqueous vehicle prior to administration. The pH of the preparations will typically be between 3 and 11, more preferably between 5 and 9, and more preferably between 7 and 8.
In some embodiments, the oligosaccharides of the invention can be incorporated into liposomes formed from standard lipids that are formed in the vesicle. A variety of methods are available for preparing liposomes, as described in, for example, Szoka et al., Ann. Rev. Biophys. Bioeng. 9: 467 (1980), US Patent Nos. 4,235,871,4,501,728 and 4,837,028. The targeting of liposomes using a variety of targeting agents (for example, the sialyl galactosides of the invention) is well known in the state of the art (see, for example, US Patent Nos. 4,957,773 and 4,603,044) .
The compositions containing the oligosaccharides can be administered for prophylactic and / or therapeutic treatments. In therapeutic applications, the compositions are administered to a patient already suffering from a disease, as described above, in an amount sufficient to cure or at least partially interrupt the symptoms of the disease and its complications. An amount adequate to accomplish this is defined as a "therapeutically effective dose." Amounts effective for this use will depend on the severity of the disease and the weight and general condition of the patient, but generally range from about 0.5 mg to about 40 g of oligosaccharide per day per 70 kg of the patient, with doses from about 5 mg to about 20 g of the compounds per day are the most commonly used.
Single or multiple administrations of the compositions can be performed with dose levels and a pattern selected by the treating physician. In any case, the pharmaceutical formulations must provide an amount of the oligosaccharides of this invention, sufficient to effectively treat the patient.
Oligosaccharides can also find use as diagnostic reagents. For example, the labeled compounds can be used to locate areas of inflammation or metastasis of a tumor in a patient suspected of having inflammation. For this use, the compounds can be labeled with suitable radioisotopes, for example,<sup>121</sup>1,<sup>14</sup>C, or tritium.
The oligosaccharides of the invention can be used as an immunogen for the production of monoclonal or polyclonal antibodies specifically reactive with the compounds of the invention. A multitude of techniques available to those skilled in the art can be used in the present invention for the production and manipulation of different immunoglobulin molecules. Antibodies can be produced in a variety of ways well known to those of skill in the art.
The production of non-human monoclonal antibodies, for example murine, lagomorphic, equine, etc., is well known and can be achieved, for example, by means of immunization of the animal with a preparation containing the oligosaccharide of the invention. Cells that produce antibodies, obtained from immunized animals
ES 2 269 098 T3 are immortalized and evaluated, or first evaluated for the production of the desired antibody and then immortalized. For a discussion of general procedures for monoclonal antibody production, see, Harlow and Lane, Antibodies: A Laboratory Manual, Cold Spring Harbor Publications, NY (1988).
Example
The following example is offered by way of illustration, but not to limit the present invention.
This Example describes the use of two strategies for the cloning of four genes responsible for the biosynthesis of GT1, a ganglioside mimic in the LOS of a bacterial pathogen, OH4384 from Campylobacter jejuni, which has been associated with Guillain-Barré syndrome (Aspinall et al. (1994) Infect. Immun. 62: 2122-2125). Aspinal et al. ((1994) Biochemistry 33: 241-249) showed that this strain has an outer core LPS that mimics the trisialylated ganglioside GT1a. We first cloned a gene encoding α-2,3-sialyltranferase (cst-I) using a strategy for evolution of activity. We then use raw nucleotide sequence information from the recently completed sequence of C. jejuni NCTC 11168 to amplify a region involved in LOS biosynthesis from C. jejuni OH4384. Using primers that are localized to heptosyltransferases I and II, the locus for biosynthesis of the 11.47 kb LOS of C. jejuni OH4384 was amplified. Sequencing revealed that the locus encodes 13 partial or complete open reading frames (ORFs), while the corresponding locus in C. jejuni NCTC 11168 spans 13.49 kb and contains 15 ORFs, indicating a different organization between these two strains.
Potential glycosyltransferase genes were individually cloned, expressed in Escherichia coli, and examined using synthetic fluorescent oligosaccharides as acceptors. We identified the genes that encode a jB-1,4-N-acetylgalactosaminyltransferase (cgtA), a jB-1,3-galactosyltransferase (cgtB) and a bifunctional sialyltransferase (cst-II) that transfers sialic acid to O-3 of galactose and to O-8 of a sialic acid that is α-2,3 linked to a galactose. The binding specificity of each identified glycosyltransferase was confirmed by 600 MHz NMR analysis on nanomole amounts of the in vitro synthesized model compounds. Using a reverse gradient broadband NMR nanoprobe, sequence information could be obtained by detecting correlations<sup>3</sup>J (C, H) through the glycosidic bond. The role of cgtA and cst-II in the synthesis of the GT1a mimic in OH4384 of C. jejuni was confirmed by comparing their sequences and activity with the corresponding homologues in two related strains of C. jejuni that express shorter mimics of gangliosides in their LOS. Therefore, these three enzymes can be used to synthesize a GT1a mimic starting from lactose.
The abbreviations used are: CE, capillary electrophoresis; CMP-Neu5Ac, cytidine monophosphate-N-acetylneuraminic acid; COZY, correlation spectroscopy; FCHASe, 6- (5-fluorescein-carboxamido) -hexanoic acid succimidyl ester; GBS, Guillain-Barré syndrome; HMBC, heteronuclear multiple bond coherence; HSQC, heteronuclear coherence of unique quantum; LIF, induced laser fluorescence; LOS, lipooligosaccharide; NOE, nuclear Overhauser effect; NOESY, NOE spectroscopy; TOCSY, total correlation spectroscopy.
Experimental procedures
Bacterial strains
The following C. jejuni strains were used in this study: serostain O: 19 (ATCC # 43446); serotype O: 19 (strains OH4382 and OH4384 were obtained with the Laboratory Center for Disease Control (Health Canada, Winnipeg, Manitoba)); and serotype O: 2 (NCTC # 11168). Escherichia coli strain DH5a was used for the HindIII library while E. coli AD202 (CGSG # 7297) was used to express the different cloned glycosyltransferases.
Basic recombinant DNA methods
Isolation of genomic DNA from C. jejuni strains was performed using Qiagen Genomic-tip 500 / G (Qiagen Inc., Valencia, CA) as previously described (Gilbert et al. (1996) J. Biol. Chem. 271: 2827128276 ). The isolation of the DNA from the plasmid, the digestions with the restriction enzyme, the purification of the DNA fragments for cloning, the ligations and the transformations were carried out as recommended by the supplier of the enzyme, or the manufacturer of the kit used for the procedure. particular. Long PCR reactions (> 3 kb) were carried out using the Expand ™ long template PCR system as described by the manufacturer (Boehringer Mannheim, Montreal). PCR reactions to amplify specific ORFs were carried out using Pwo DNA polymerase as described by the manufacturer (Boehringer Mannheim, Montreal). DNA restriction and modification enzymes were purchased from New England Biolabs Ltd. (Mississauga, ON). DNA sequencing was carried out using an Applied Biosystems (Montreal) Model 370A Automatic DNA Sequencer and the manufacturer's cyclic sequencing kit.
ES 2 269 098 T3
Evaluation of activity for C. jejuni sialyltransferase
The genomic library was prepared using a partial HindIII digest of C. jejuni OH4384 chromosomal DNA. The partial digest was purified on a QIAquick column (QIAGEN Inc.) and ligated with HindIII-digested pBluescript SK-. The E. coli strain DH5a was electroporated with the ligation mixture, and the cells were seeded on plates in LB medium with 150 jUg / mL of ampicillin, 0.05 mM IPTG and 100 jUg / mL of X-Gal (5 -Bromo-4-chloro-indolyl-β-D-galactopyranoside). The white colonies were collected in 100 wells and resuspended in 1 mL of medium with 15% glycerol. Twenty pL from each well were used to inoculate 1.5 mL of LB medium supplemented with 150 jUg / mL of ampicillin. After 2 h of growth at 37 ° C, IPTG was added to 1 mM and the cultures grew for another 4.5 h. The cells were recovered by centrifugation, resuspended in 0.5 mL of 50 mM Mops (pH 7, MgCl<sub>2</sub> 10 mM) and sonicated for 1 min. The extracts were evaluated for sialyltransferase activity as described below, except that the incubation period and temperature were 18 h and 32 ° C, respectively. Positive wells were seeded on single colony plates, and 200 colonies were picked and 10 wells analyzed for activity. Finally, the colonies from the positive wells were analyzed individually which led to the isolation of two positive clones, pCJH9 (5.3 kb insert) and pCJH101 (3.9 kb insert). Using various subcloned fragments and custom primers, the inserts of the two clones on both strains were fully sequenced. Clones with individual HindIII fragments were also analyzed for sialyltransferase activity and the insert of the only positive (a 1.1 kb HindIII fragment cloned into pBluescript SK-) was transferred to pUC118 using KpnI and PstI sites for the purpose of obtain the insert in the opposite orientation with respect to the plac promoter.
Cloning and sequencing of the locus for LPS biosynthesis
The primers used to amplify the locus for C. jejuni OH4384 LPS biosynthesis were based on preliminary sequences available on the website (URL: http://www.sanger.ac.uk/Projects/Cjejuni/) of the C. jejuni sequencing group (Sanger Center, UK) that sequenced the entire genome of strain NCTC11168. Primers CJ-42 and CJ-43 (all primer sequences are described in Table 2) were used to amplify an 11.47 kb locus using the Expand ™ long template PCR system. The PCR product was purified on an S-300 spin column (Pharmacia Biotech) and completely sequenced on both strands using a combination of primer walking and HindIII fragment subcloning. Specific ORFs were amplified using the primers described in Table 2 and Pwo DNA polymerase. The PCR products were digested using the appropriate restriction enzymes (see Table 2) and were cloned into pCWori +.
(Table goes to next page)
ES 2 269 098 T3
TABLE 2
Initiators used for Amplification of Open Reading Frames
Primers used to amplify the locus for core biosynthesis of LPS
CJ42: HeptosylTase-Π primer (Error! Reference source not found).
5 'GC CAT TAC CGT ATC GCC TAA CCA GG 3' 25 mer
CJ43: HeptosylTase-I primer (Error! Reference source not found).
5 'AAA GAA TAC GAA TTT GCT AAA GAG G 3' 25 mer
Primers used to amplify and clone ORF 5a:
CJ-106 (3 'initiator, 41mer) (Error! Reference source not found).
I left
5 'CCT AGG TCG ACT TAA AAC AAT GTT AAG AAT ATT TTT TTT AG 3'
CJ-157 (5 'initiator, 37 mer) (Error! Reference source not found).
Ndel
5 'CTT AGG AGG TCA TAT GCT ATT TCA ATC ATA CTT TGT G 3'
Primers used to amplify and clone ORF 6a:
CJ-1O5 (3 'initiator, 37 mer) (Error! Reference source not found).
I left
5 'CCT AGG TCG ACC TCT AAA AAA AAT ATT CTT AAC ATT G 3'
CJ-133 (5 'initiator, 39mer) (Error! Reference source not found).
Ndel
5 'CTTAGGAGGTCATATGTTTAAAATTTCAATCATCTTACC 3'
Primers used to amplify clone ORF 7a:
CJ-131 (5 'initiator, 41mer) (Error! Reference source not found).
Ndel
5 'CTTAGGAGGTCATATGAAAAAAGTTATTATTGCTGGAAATG 3'
CJ-132 (3 'initiator, 41mer) (Error! Reference source not found).
I left
5 'CCTAGGTCGACTTATTTTCCTTTGAAATAATGCTTTATATC 3'
Expression in E. coli and in glycosyltransferase assays
The different constructs were transferred to E. coli strain AD202 and analyzed for the expression of glycosyltransferase activity after 4 h of induction with 1 mM IPTG. The extracts were made by sonication and the enzymatic reactions were carried out overnight at 32 ° C. FCHASE-labeled oligosaccharides were prepared as previously described (Wakarchuk et al. (1996) J. Biol. Chem. 271: 19166-19173). Protein concentration was determined using the Bicinchoninic Acid Protein Assay Kit (Pierce, Rockford, IL). For all enzymatic analyzes, a unit of activity was defined as the amount of enzyme that generated one // mole of product per minute.
The assay for α-2,3-sialyltransferse activity in clone pools contained 1 mM LacFCHASE, 0.2 mM CMP-Neu5Ac, 50 mM Mops pH 7, MnCl<sub>2</sub> 10 mM and MgCl<sub>2</sub> 10 mM in a final volume of 10 juL. The different subcloned ORFs were analyzed by the expression of glycosyltransfe24 activity.
ES 2 269 098 T3 rasa after 4 h of induction of the cultures with 1 mM IPTG. The extracts were made by sonication and the enzymatic reactions were carried out overnight at 32 ° C.
JB-1,3-galactosyltransferase was assayed using 0.2 mM GM2-FCHASE, 1 mM UDP-Gal, 50 mM Month pH 6, MnCl<sub>2</sub> 10 mM and 1 mM DTT. The / l-í ', 4-GalNAc transferase was assayed using GM3-FCHASE 0.5 mM, UDP-GalNAc 1 mM, Hepes 50 mM pH 7 and MnCl<sub>2</sub> 10 mM. Α-2,3-sialyltransferase was assayed using 0.5 mM LacFCHASE, 0.2 mM CMP-Neu5Ac, 50 mM Hepes pH 7 and MgCl<sub>2</sub> 10 mM. Α-2,8-sialyltransferase was assayed using 0.5 mM GM3-FCHASE, 0.2 mM CMP-Neu5Ac, 50 mM Hepes pH 7 and MnCl<sub>2</sub> 10 mM.
The reaction mixtures were appropriately diluted with 10 mM NaOH and analyzed by capillary electrophoresis carried out using the separation and detection conditions previously described (Gilbert et al. (1996) J. Biol. Chem. 271: 2871 -28276). The peaks of the electropherograms were analyzed using manual integration of the peaks with the P / ACE Station software. For rapid detection of enzyme activity, samples of the transferase reaction mixtures were examined by thin layer chromatography on silica-60 TLC plates (E. Merck) as previously described (Id.).
NMR spectroscopy
NMR experiments were performed on a Varian INOVA 600 NMR spectrometer. Most of the experiments were performed using a 5 mm triple resonance Z gradient probe. NMR samples were prepared from 0.3-0.5 mg (200-500 nanomoles) of FCHASE-glycoside. The compounds were dissolved in H2O and the pH was adjusted to 7.0 with dilute NaOH. After freeze-drying, the samples were dissolved in 600 pL of D<sub>2</sub>O. All NMR experiments were carried out as previously described (Pavliak et al. (1993) J. Biol. Chem. 268: 14146-14152; Brisson et al. (1997) Biochemistry 36: 3278-3292) using standard techniques such as COZY, TOCSY, NOESY, 1D-NOESY, 1D-TOCSY and HSQC. As a reference to the chemical proton shift, the internal acetone methyl resonance was set at 2,225 ppm (1H). As a reference for the chemical shift of the<sup>13</sup>C, the methyl resonance of internal acetone relative to external dioxane was set at 31.07 ppm at 67.40 ppm. Homonuclear experiments were on the order of 5-8 hours each. The 1D NOESY experiments for GD3-FCHASE, [0.3 mM], with 8000 scans and a mixing time of 800 ms were performed for a duration of 8.5 h each and processed with a line spreading factor of 2-5 Hz. For the 1D NOESY of the resonances at 4.16 ppm, 3000 scans were used. The following parameters were used to acquire the HSQC spectrum: a relaxation delay of 1.0 s, spectral widths in F<sub>2</sub> and F<sub>1</sub> 6000 and 24147 Hz, respectively, acquisition times in t<sub>2</sub> 171 ms. For dimension t<sub>1</sub>, 128 complex points were acquired using 256 scans per increment. Discrimination of the signal in F1 was achieved by means of the States method. The total acquisition time was 20 hours. For GM2-FCHASE, due to the wide lines, the number of scans per increment was increased so that the HSQC was performed for 64 hours. The phase sensitive spectrum was obtained after zeroing up to 2048 x 2048 spots. Undisplaced Gaussian window functions were applied in both dimensions. The HSQC spectra were plotted with a resolution of 23 Hz / point in the dimension of the<sup>13</sup>C and 8 Hz / point in the proton dimension. For the observation of the multiple unfolds, the dimension of the<sup>1</sup>H with 2 Hz / point resolution using advanced linear prediction and a π / 4 shifted square Sinebell function. All NMR data were acquired using standard Varian sequences supplied with VRMN 5.1 or VRMN 6.1 software. The same program was used for processing.
A broadband reverse gradient NMR nanoprobe (Varian) was used to perform the HMBC gradient (Bax and Summers (1986) J. Am. Chem. Soc. 108,2093-2094; Parella et al. (1995) J. Mag Reson. A 112.241245) were experienced for the GD3-FCHASE sample. The NMR nanoprobe, which is a high resolution magic angle spin probe, produces a high resolution spectrum of liquid samples dissolved in only 40 pL (Manzi et al. (1995) J. Biol. Chem. 270, 9154-9163). The sample for GD3-FCHASE (mass = 1486.33 Da) was prepared by lyophilizing the original 0.6 mL (200 nanomoles) sample and dissolving it in 40 pL of D2O for a final concentration of 5 mM. The final pH of the sample could not be measured.
The HMBC gradient experiment was performed at a spin speed of 2990 Hz, 400 increments of 1024 complex points, 128 scans per increment, acquisition time of 0.21 s, <sup>1</sup>J (C, H) = 140 Hz and <sup>n</sup>J (C, H) = 8 Hz, for a duration of 18.5 h.
Mass spectrometry
All mass measurements were obtained using an Elite-STR MALDITOF instrument from Perkin-Elmer Biosystems (Fragmingham, MA). Approximately two pg of each oligosaccharide was mixed with a matrix containing a saturated solution of dihydroxybenzoic acid. Positive and negative mass spectra were acquired, using the reflector mode.
Results
Detection of glycosyltransferase activity in C. jejuni strains
Before cloning of the glycosyltransferase genes, cells of C. jejuni strains OH4384 and NCTC 11168 were examined for different enzymatic activity. When the activity of an enzyme was detected, it was opti25
ES 2 269 098 T3 mized the assay conditions (described in the Experimental Procedures) to ensure maximum activity. The capillary electrophoresis assay used was extremely sensitive and allowed the detection of enzymatic activity in the pU / ml range (Gilbert et al. (1996) J. Biol. Chem. 271: 28271-28276). Both the sequenced strain NCTC 11168 and the GBS-associated strain OH4384 were screened for the enzymes required for the synthesis of the GT1a ganglioside mimic. As predicted, strain OH4384 possessed the enzymatic activity required for the synthesis of this structure: / -1,4-N-acetylgalactosaminyltransferase, / '- 1,3-galaclosillransferase, a2,3-sialyltransferase and a-2,8- sialyltransferase. The genome of strain NCTC 11168 lacked the / -1,3galactosyltransferase activity and the α-2,8-sialyltransferase activity.
Cloning of an α-2,3-sialyltransferase (cst-I) using a strategy for the evaluation of activity
A plasmid library made from unfractionated partial HindIII digestion of chromosomal DNA from C. jejuni strain OH4384 produced 2,600 white colonies that were harvested to form 100 wells. A "divide and conquer" evaluation protocol was used from which two positive clones were obtained and designated as pCJH9 (a 5.3 kb insert, 3 HindIII sites) and pCJH101 (3.9 kb insert, 4 sites HindIII). Open reading frame (ORF) analysis and chromosomal DNA PCR reactions of C. jejuni strain OH4384 indicated that pCJH9 contained inserts that were not contiguous in the chromosomal DNA. The sequence downstream of nucleotide # 1440 in pCJH9 was not studied further while the first 1439 nucleotides were found to be completely contained within the sequence of pCJH101. ORF analysis and chromosomal DNA PCR reactions indicated that all fragments of pCJH101 with HindIII were contiguous on the chromosomal DNA of OH4384 from C. jejuni.
Four ORFs, two partial and two complete, were found in the sequence of pCJH101 (Figure 2). The first 812 nucleotides encode a polypeptide that is 69% identical to at least 265 aa residues of Helicobacter pylori peptide chain release factor RF-2 (prfB gene, GenBank # AE000537). The last base of the chain release factor TAA stop codon is also the first base of the ATG start codon of an open reading frame spanning nucleotides # 812 to # 2104 in pCJH101. This ORF was designated cst-I (Campylobacter sialyltransferase I) and encodes a 430 amino acid polypeptide that is homologous to a putative Haemophilus influenzae ORF (GenBank # U32720). The putative ORF of H. influenzae encodes a 231 amino acid polypeptide that is 39% identical to the mid-region of the Cst I polypeptide (amino acid residues # 80 to # 330). The downstream sequence of cst-I includes an ORF and a partial ORF encoding polypeptides that are homologous (> 60% identity) to the two subunits, CysD and CysN, of the E. coli adenylyltransferase sulfate (GenBank # AE000358 ).
In order to confirm that the cst-I ORF encodes sialyltransferase activity, it was subcloned and overexpressed in E. coli. The expressed enzyme was used to add sialic acid to Gal - / - 1,4-Glc - / - FCHASE (Lac-FCHASE). This product (GM3-FCHASE) was analyzed by NMR to confirm binding specificity to Neu5Ac-a-2,3-Gal of Cst-I.
Locus sequencing for C. jejuni OH4384 LOS biosynthesis
Analysis of preliminary sequence data available on the website of the C. jejuni NCTC 11168 sequencing group (Sanger Center, UK (http://www.sanger.ac.uk/ Projects / C_jejuni /)) revealed that the two heptosyltransferases involved in the synthesis of the inner core of LPS were easily identifiable through sequence homology with other bacterial heptosyltransferases. The region between the two heptosyltransferases spans 13.49 kb in NCTC 11168 and includes at least seven potential glycosyltransferases based on BLAST searches of GenBank. Since there is no available structure for the outer core of LOS from NCTC 11168, it was impossible to suggest functions for the putative glycosyltransferase genes in that strain.
Based on the conserved regions in the heptosyltransferase sequences, primers (CJ-42 and CJ-43) were designed to amplify the region between them. A 13.49 kb PCR product was obtained using chromosomal DNA from C. jejuni NCTC 11168 and an 11.47 PCR product using chromosomal DNA from C. jejuni OH4384. The size of the PCR product from strain NCTC 11168 was consistent with the Sanger Center data. The smaller size of the PCR product from strain OH4384 indicated heterogeneity between strains in the region between the two heptosyltransferase genes and suggested that genes for some of the strain OH4384-specific glycosyltransferases might be present at that location. The 11.47 kb PCR product was sequenced using a combination of a primer walk and subcloning of HindIII fragments (GenBank # AF130984). The G / C content of the DNA was 27%, typical of Campylobacter DNA. Sequence analysis revealed eleven complete ORFs in addition to the two partial ORFs encoding the two heptosyltransferases (Figure 2, Table 3). When the deduced amino acid sequences are compared, the two strains are found to share six genes that are over 80% identical and four genes that are 52% to 68% identical (Table 3). Four genes are unique to C. jejuni while one gene is unique to C. jejuni OH4384 (Figure 2). Two genes that are present as separate ORFs (ORF # 5a and # 10a) in C. jejuni OH4384 are found in a fusion ORF in the framework (# 5b / 10b) in C. jejuni NCTC 11168.
ES 2 269 098 T3
TABLE 3
Location and description of the locus ORFs for the biosynthesis of LOS of OH4384 from C. jejuni
<td>ORF #</td><td>Location</td><td>Homolog in Strain NCTC11168<sup>to </sup>(% identity in aa sequence)</td><td>Homologues found in GenBank (% identity in aa sequence)</td><td>Function<sup>b</sup></td>
<td>the</td><td> 1-357</td><td>ORF # Ib (98%)</td><td>rfaC (GB # AE000546) from Helicobacter pylori (35%)</td><td>Heptosyltransferase I</td>
<td>2nd</td><td> 350-1.234</td><td>ORF # 2b (96%)</td><td>waaM (GB # AE001463) from Helicobacter pylori (25%)</td><td>Acyltransferase for lipid A biosynthesis</td>
<td>3rd</td><td> 1.234-2.487</td><td>ORF # 3b (90%)</td><td>lgtf (GB # U58765) from Neisseria meningitidis (31%)</td><td>Glycosyltransferase</td>
<td>4th</td><td> 2.786-3.952</td><td>ORF # 4b (80%)</td><td>cpsl4J (GB # X85787) from Streptococcus pneumoniae (45% on the first 100 aa)</td><td>Glycosyltransferase</td>
<td>5th</td><td> 4.025-5.065</td><td>N-terminal ORF # 5b / 10b (52%)</td><td>ORF # HP0217 (GB # AE000541) from Helicobacter pylori (50%)</td><td>Bl, 4-N- acetylgalactosaminyltransferase (cgt4)</td>
<td>6th</td><td>5,057-5,959 (complement)</td><td>ORF # 6b (60%)</td><td>Cps23FU (GB # AF030373) from Streptococcus pneumoniae (23%)</td><td>Bl, 3-Galactosyltransferase {cgtBj</td>
<td>7a</td><td> 6.048-6.920</td><td>ORF # 7b (52%)</td><td>ORF # HI0352 (GB # U32720) from Haemophilus influenzae (40%)</td><td>Bi-functional α-2,3 / α-2,8 sialyltransferase (cts-ΙΓ)</td>
<td>8a</td><td> 6.924-7.961</td><td>ORF # 8b (80%)</td><td>siaC (GB # U40740) from Neisseria meningitidis (56%)</td><td>Sialic acid synthase</td>
<td>9a</td><td> 8.021-9.076</td><td>ORF # 9b (80%)</td><td>siaA (GB # M95053) from Neisseria meningitidis (40%)</td><td>Sialic acid biosynthesis</td>
<td>10a</td><td> 9.076-9.738</td><td>C-terminal ORF # 5b / 10b (68%)</td><td>neuA (GB # U54496) from Haemophilus ducreyi (39%)</td><td>Sialic acid-CMP synthetase</td>
<td>lia</td><td> 9.729-10.559</td><td>Not homologous</td><td>Putative ORF (GB # AF010496) of Rhodobacter capsulatus (22%)</td><td>Acetyltransferase</td>
<td>12a</td><td>10,557-11,366 (complement)</td><td>ORF # 12b (90%)</td><td>ORF # HI0868 (GB # U32768) from Haemophilus influenzae (23%)</td><td>Glycosyltransferase</td>
<td>13a</td><td> 11.347-11.474</td><td>ORF # 13b (100%)</td><td>Helicobacter pylori rfaF (GB # AE000625) (60%)</td><td>Heptosyltransferase II</td>
<td colspan="5"><sup>to</sup> The sequence of the C. jejuni NCTC 11168 ORFs can be obtained from the Sanger Center (URL: htto / www.saneer.ac.uk / Proiects / C jejuni /).<sup>b</sup> Functions that were determined experimentally are in bold. Other functions are based on GenBank higher scoring homologs.</td>
ES 2 269 098 T3
Identification of the outer core glycosyltransferases
Different constructs were made to express each of the potential genes for glycosyltransferase located among the C. jejuni OH4384 heptosyltransferases. Plasmid pCJL-09 containing ORF # 5a and a culture of this construct showed GalNAc transferase activity when tested using GM3-FCHASE as acceptor. GalNAc transferase was specific for a sialylated acceptor since Lac-FCHASE was a poor substrate (less than 2% of the activity seen with GM3-FCHASE). The reaction product obtained from GM3-FCHASE had the correct mass determined with a MALDITOF mass spectrometer, and the elution time identical in the CE assay as that of the standard GM2-FCHASE. Considering the LPS structure of the outer core of OH4384 from C. jejuni, this GalNAc transferase (cgtA for Campylobacter glycosyltransferase A), has a specificity of 6-1.4 for the terminal Gal residue of GM3-FCHASE. The specificity of CgtA binding was confirmed by GM2-FCHASE NMR analysis (see text below, Table 4). The in vivo role of cgtA in the synthesis of a GM2 mimic is confirmed by the deleted natural mutant provided by C. jejuni OH4382 (Figure 1). By sequencing the C. jejuni OH4382 cgtA homolog an in-frame shift mutation was found (a stretch of seven A's instead of 8 A's after base # 71) that would result in the expression of a truncated version of cgtA (29 aa instead of 347 aa). The structure of the outer core of the LOS of OH4382 of C. jejuni is consistent with the absence of /'-1,4-GlaNAc transferase while the inner galactose residue is replaced only with sialic acid (Aspinall et al. (1994) Biochemistry 33, 241249).
Plasmid pCJL-04 containing ORF # 6a and an IPTG-induced culture of this construct showed galactosyltransferase activity using GM2-FCHASE as an acceptor thereby producing GM1a-FCHASE. This product was sensitive to / '- 1,3-galaclosidase and was found to have the correct mass by MALDI-TOF mass spectrometry. Considering the structure of the outer core of the LOS of OH4384 of C. jejuni, it is suggested that this galactosyltransferase (cgtB by Campylobacter glycosyltransferase B) has / 1-1.3 specificity for the GalNAc terminal residue of GM2-FCHASE. The specificity of the CgtA binding was confirmed by NMR analysis of GM1a-FCHASE (see text below, Table 4) which was synthesized through the sequential use of Cst-I, CgtA and CgtB.
Plasmid pCJL-03 that included ORF # 7a and an IPTG-induced culture showed sialyltransferase activity using both Lac-FCHASE and GM3-FCHASE as acceptors. This second OH4384 sialyltransferase was designated cst-II. Cst-II was shown to be bi-functional since it could transfer α-2,3 sialic acid to the terminal Gal of Lac-FCHASE and also α-2,8- to the terminal sialic acid of GM3-FCHASE. NMR analysis of a reaction product formed with Lac-FCHASE confirmed the α-2,3 bond of the first sialic acid on Gal, the α-2,8 bond of the second sialic acid (see text below, Table 4).
(Table goes to next page)
ES 2 269 098 T3
TABLE 4
NMR Proton Chemical Shifts for Fluorescent Derivatives of Ganglioside Mimics Synthesized Using Cloned Glycosyltransferases
<td>Residue</td><td colspan="4">Chemical Shift (ppm) H Lac- GM3- GM2- GMla-</td><td>GD3-</td>
<td>PGlc</td><td colspan="2"> 1 4,57 4,70</td><td> 4,73</td><td> 4,76</td><td> 4,76</td>
<td>to</td><td colspan="2"> 2 3,23 3,32</td><td> 3,27</td><td> 3,30</td><td> 3,38</td>
<td></td><td colspan="2"> 3 3,47 3,54</td><td> 3,56</td><td> 3,58</td><td> 3,57</td>
<td></td><td colspan="2"> 4 3,37 3,48</td><td> 3,39</td><td> 3,43</td><td> 3,56</td>
<td></td><td colspan="2"> 5 3,30 3,44</td><td> 3,44</td><td> 3,46</td><td> 3,50</td>
<td></td><td colspan="2"> 6 3,73 3,81</td><td> 3,80</td><td> 3,81</td><td> 3,85</td>
<td></td><td colspan="2"> 6’ 3,22 3,38</td><td> 3,26</td><td> 3,35</td><td> 3,50</td>
<td>pGal (l-4)</td><td colspan="2"> 1 4,32 4,43</td><td> 4,42</td><td> 4,44</td><td> 4,46</td>
<td>b</td><td colspan="2"> 2 3,59 3,60</td><td> 3,39</td><td> 3,39</td><td> 3,60</td>
<td></td><td colspan="2"> 3 3,69 4,13</td><td> 4,18</td><td> 4,18</td><td> 4,10</td>
<td></td><td colspan="2"> 4 3,97 3,99</td><td> 4,17</td><td> 4,17</td><td> 4,00</td>
<td></td><td> 5 3,1</td><td> 31 3,77</td><td> 3,84</td><td> 3,83</td><td> 3,78</td>
<td></td><td> 6 3,1</td><td> 36 3,81</td><td> 3,79</td><td> 3,78</td><td> 3,78</td>
<td></td><td> 6’ 3,1</td><td> 31 3,78</td><td> 3,79</td><td> 3,78</td><td> 3,78</td>
<td>áNeu5Ac (2-3)</td><td>3ax</td><td> 1,81</td><td> 1,97</td><td> 1,96</td><td> 1,78</td>
<td>c</td><td> 3<sub>eq</sub></td><td> 2,76</td><td> 2,67</td><td> 2,68</td><td> 2,67</td>
<td></td><td> 4</td><td> 3,69</td><td> 3,78</td><td> 3,79</td><td> 3,60</td>
<td></td><td> 5</td><td> 3,86</td><td> 3,84</td><td> 3,83</td><td> 3,82</td>
<td></td><td> 6</td><td> 3,65</td><td> 3,49</td><td> 3,51</td><td> 3,68</td>
<td></td><td> 7</td><td> 3,59</td><td> 3,61</td><td> 3,60</td><td> 3,87</td>
<td></td><td> 8</td><td> 3,91</td><td> 3,77</td><td> 3,77</td><td> 4,15</td>
<td></td><td> 9</td><td> 3,88</td><td> 3,90</td><td> 3,89</td><td> 4,18</td>
<td></td><td> 9’</td><td> 3,65</td><td> 3,63</td><td> 3,64</td><td> 3,74</td>
<td></td><td>Nac</td><td> 2,03</td><td> 2,04</td><td> 2,03</td><td> 2,07</td>
<td>PGalNAc (l-4)</td><td> 1</td><td></td><td> 4,77</td><td> 4,81</td><td></td>
<td>d</td><td> 2</td><td></td><td> 3,94</td><td> 4,07</td><td></td>
<td></td><td> 3</td><td></td><td> 3,70</td><td> 3,82</td><td></td>
<td></td><td> 4</td><td></td><td> 3,93</td><td> 4,18</td><td></td>
<td></td><td> 5</td><td></td><td> 3,74</td><td> 3,75</td><td></td>
<td></td><td> 6</td><td></td><td> 3,86</td><td> 3,84</td><td></td>
<td></td><td> 6’</td><td></td><td> 3,86</td><td> 3,84</td><td></td>
<td></td><td>NAc</td><td></td><td> 2,04</td><td> 2,04</td><td></td>
<td>Residue</td><td colspan="4">Chemical Shift (ppm) H Lac- GM3- GM2- GMla-</td><td>GD3-</td>
<td>pGal (l-3) and aNeu5Ac (2-8)</td><td>1 2 3 4 5 6 6 ' 3ax</td><td></td><td></td><td> 4,55 3,53 3,64 3,92 3,69 3,78 3,74</td><td> 1,75</td>
ES 2 269 098 T3
<td>f 3<sub>eq</sub></td><td> 2,76</td>
<td> 4</td><td> 3,66</td>
<td> 5</td><td> 3,82</td>
<td> 6</td><td> 3,61</td>
<td> 7</td><td> 3,58</td>
<td> 8</td><td> 3,91</td>
<td> 9</td><td> 3,88</td>
<td> 9’</td><td> 3,64</td>
<td>NAc</td><td> 2,02</td>
<td colspan="2"><sup>to</sup> in ppm from the HSQC spectrum obtained at 600 MHz, D<sub>2</sub>Or, pH 7.28 ° C for Lac-, 25 ° C for GM3-, 16 ° C for GM2-, 24 ° C for GMla-, and 24 ° C for GD3-FCHASE. The methyl resonance of</td>
<td>the internal acetone is at 2225 ppm ('H). The error is ± 0.02 ppm for the chemical shifts of 'H and ± 5 ° C for the sample temperature. The error is ± 0.1 ppm for the H-6 resonances of residues a, b, d, and e due to overlap.</td><td></td>
Comparison of sialyltransferases
The in vivo role of C. jejuni OH4384 cst-II in the synthesis of an imitation of the trisialylated ganglioside GT1a is supported by comparison with the C. jejuni (sero-strain) O: 19 cst-II homologue that expresses the mimicry of the disialylated ganglioside GD1a. There are 24 nucleotide differences that translate into 8 amino acid differences between these two cst-II homologues (Figure 3). When expressed in E. coli, the cst-II homolog of O: 19 from C. jejuni (sero-strain) has α-2,3-sialyltransferase activity but very low α-2,8-sialyltransferase activity (Table 5) that is consistent with the absence of the terminal α-2,8-linked sialic acid in the nucleus exterior of LOS (Aspinall et al. (1994) Biochemistry 33,241-249) of O: 19 from C. jejuni (sero-strain). The cst-II homolog of NCTC 11168 from C. jejuni expressed much lower α-2,3-sialyltransferase activity than the O: 19 (sero-strain) or OH4384 homolog and no detectable α-2,8-sialyltransferase activity. An inducible IPTG band could be detected on an SDS-PAGE gel when NCTC 11168 cst-II was expressed in E. coli (data not shown). The Cst-II protein of NCTC 11168 shares only 52% identity with the homologues of O: 19 (sero-strain) or OH4384. It could not be determined whether the differences in the sequence could be responsible for the lower activity expressed in E. coli.
Although cst-I externally mapped the locus for LOS biosynthesis, it is obviously homologous to cst-II since its first 300 residues share 44% identity with Cst-II, either from OH4384 from C. jejuni or from NCTC 11168 of C. jejuni (Figure 3). The two homologous Cst-II share 52% identical residues with each other and the C-terminal 130 amino acids of Cst-I are absent. A truncated version of Cst-I that lost 102 amino acids at the C-terminal was found to be active (data not shown) indicating that the C-terminal domain of Cst-I is not required for sialyltransferase activity. Although the 102 residues at the C-terminus are dispensable for enzymatic activity in vitro, they can interact with other cellular components in vivo either for regulatory purposes or for proper cell localization. The low level of conservation between the C. jejuni sialyltransferases is very different from what was previously observed for the α-2,3-sialyltransferases of N. meningitidis and N. gonorrhoeae, where the former transferases are more than 90 identical. % with the protein level between the two species and between different isolates of the same species (Gilbert et al., supra.).
ES 2 269 098 T3
TABLE 5
Comparison of the activity of C. jejuni sialyltransferases. The different sialyltransferases were expressed in E. coli as fusion proteins with the maltose binding protein in the pCWori + vector (Wakarchuk et al. (1994) Protein. Sci. 3, 467-475). The sonicated extracts were analyzed using 500 pM of either Lac-FCHASE or GM3-FCHASE.
<td rowspan="2">Sialyltransferase gene</td><td colspan="2">Activity (pU / mg)<sup>to</sup></td><td rowspan="2">Relationship (%)<sup>6</sup></td>
<td>Lac-FCHASE</td><td>GM3-FCHASE</td>
<td>cst-I (OH4384) cs yes - // (OH4384) cst-II (sero-strain 0:19) cs yes - // (NCTC 11168)</td><td> 3.744 209 2.084 8</td><td> 2,2 350,0 1,5 0</td><td> 0,1 167,0 0,1 0,0</td>
<td colspan="4">"The activity is expressed in pU (pmol of product per minute) per mg of total protein in the extract.<sup>b</sup> Ratio (in percentage) of the activity on GM3-FCHASE.</td>
NMR analysis on nanomole quantities of the synthesized model compounds
In order to properly evaluate the binding specificity of an identified glycosyltransferase, its product was analyzed by means of NMR spectroscopy. In order to reduce the time required for the purification of the enzymatic products, NMR analyzes were carried out on amounts in nanomole. All compounds are soluble and produce sharp resonances with line widths of a few Hz since the anomeric doublets H-1 (J<sub>or</sub> = 8 Hz) are well resolved. The only exception is for GM2-FCHASE which has wide lines (~ 10 Hz), probably due to aggregation. For the proton spectrum of the 5 mM GD3-FCHASE solution in the NMR nanoprobe, the line widths of the anomeric signals were of the order of 4 Hz, due to the higher concentration. Also, additional peaks were observed, probably due to degradation of the sample over time. There were also some slight changes in the chemical shifts, probably due to a change in pH by the concentration of the sample from 0.3 mM to 5 mM. The proton spectra were acquired at different temperatures in order to avoid the overlap of the HDO resonance with the anomeric resonances. As can be evaluated from the proton spectra, all compounds were pure and impurities or degradation products that were present did not interfere with NMR analyzes that were performed as previously described (Pavliak et al. (1993) J Biol. Chem. 268, 14146-14152; Brisson et al. (1997) Biochemistry 36, 3278-3292).
For all FCHASE glycosides, the assignments by <sup>13</sup>C from similar glycosides (Sabesan and Paulson (1986) J. Am. Chem. Soc. 108, 2068-2080; Michon et al. (1987) Biochemistry 26, 8399-8405; Sabesan et al. (1984) Can. J. Chem. 62, 1034-1045) were available. For the FCHASE glycosides, the assignments with<sup>13</sup>C were verified by first assigning the proton spectrum of the homonuclear 2D standard experiments, COZY, TOCSY, and NOESY, and then verifying the assignments with <sup>13</sup>C from an HSQC experiment, which detects CH correlations. The HSQC experiment does not detect the C-1 and C2-like quaternary carbons of sialic acid, but the HMBC experiment does. Mainly due to Glc resonances, the chemical shifts of the proton obtained from the HSQC spectrum differed from those obtained from homonuclear experiments due to heating of the sample during uncoupling of the<sup>13</sup>C. From a series of proton spectra acquired at different temperatures, the chemical shifts of the Glc residue were found to be the most sensitive to temperature. In all compounds, the resonances of Glc H-1 and H2 changed by 0.004 ppm / ° C, Gal (1-4) H-1 by 0.002 ppm / ° C, and less than 0.001 ppm / ° C for the H-3 of Neu5Ac and other anomeric resonances. For LAC-FCHASE, the Glc H-6 resonance changed by 0.008 ppm / ° C.
The large temperature coefficient for the Glc resonances is attributed to the current ring shifts induced by binding to the aminophenyl group of FCHASE. The temperature of the sample during the HSQC experiment was measured from the chemical shift of the Glc H-1 and H-2 resonances. For GM1aFCHASE, the temperature changed from 12 ° C to 24 ° C due to the presence of the Na + counterion in the solution and the NaOH used to adjust the pH. Other samples had less severe heating (<5 ° C). In all cases, the changes of the proton shifts with temperature did not cause any problems in the assignments of the resonances in the HSQC spectrum. In Table 4 and Table 6, all chemical shifts are taken from the HSQC spectrum.
ES 2 269 098 T3
The binding site on the aglycone was determined primarily from a comparison of the chemical shifts of the <sup>13</sup>C of the enzyme product with those of the precursor to determine glycosylation shifts as previously done for ten sialyloligosaccharides (Salloway et al. (1996) Infect. Immun., 64, 2945-2949). Here, instead of comparing the spectra of the<sup>13</sup>C, the spectra of HSQC are compared, as it would take a hundred times more material to obtain a spectrum of <sup>13</sup>C. When chemical shifts of the <sup>13</sup>C from the HSQC spectra of the parent compound are compared to those of the enzyme product, the main downfield shift always occurs at the binding site, while other chemical shifts of the precursor do not change substantially. Differences in the chemical shift of the proton are much more susceptible to long-range conformational effects, sample preparation, and temperature. The identity of the newly added sugar can be quickly identified from a comparison of its chemical shifts from <sup>13</sup>C than that of monosaccharides or any terminal residue, since only the anomeric chemical shift of the glycon changes substantially due to glycosidation (Sabesan and Paulson, supra).
The spin-spin coupling with the neighboring proton (J<sub>H H</sub>) obtained from the 1D TOCSY or 1D NOESY experiments is also useful for determining the identity of the sugar. NOE experiments are performed to sequence sugars by observing NOEs between anomeric glycon protons (H-3s for sialic acid) and aglycon proton resonances. The largest NOE is usually on the binding proton, but other NOEs can also occur on the aglycon proton resonances that are close to the binding site. Although at 600 MHz, the NOEs of many tetra and pentasaccharides are positive or very small, all of these compounds gave very negative NOEs with a mixing time of 800 ms, probably due to the presence of the large FCHASE fraction.
For synthetic Lac-FCHASE, the assignments of <sup>13</sup> C for the lactose fraction of Lac-FCHASE were confirmed by means of the 2D methods outlined above. All Glc unit proton resonances were mapped from a 1D-TOCSY experiment on Glc H-1 resonance with a mixing time of 180 ms. A 1D-TOCSY experiment for H-1 of Gal was used to map the resonances of H-1 to H-4 of the Gal unit. The remaining H-5 and H-6 of the Gal unit were then assigned from the HSQC experiment. The neighboring spin-spin coupling values (J<sub>H H</sub>) for the sugar units were in agreement with the previous data (Michon et al., supra). Chemical shifts for the FCHASE fraction have been given previously (Gilbert et al. (1996) J. Biol. Chem. 271, 28271-28276).
Accurate determination of the mass of the Lac-FCHASE Cst-I enzyme product was consistent with the addition of sialic acid to the Lac-FCHASE acceptor (Figure 4). The product was identified as GM3-FCHASE since the spectrum of the proton and the chemical shifts of the<sup>13</sup>C of the sugar fraction of the product (Table 6) were very similar to those of the GM3 oligosaccharide or sialillactose, (aNeuAc (2-3) j6Gal (1-4) 6Glc; Sabesan and Paulson, supra). The GM3-FCHASE proton resonances were assigned from the COZY spectrum, the HSQC spectrum, and the comparison of the proton and chemical shifts of the<sup>13</sup>C with those of aNeu5Ac (2-3) 6Gal (14) 6G1cNAc-FCHASe (Gilbert et al., Supra.). For these two compounds, the proton and the chemical shifts of the<sup>13</sup>C for the Neu5Ac and Gal residues were within the errors linked to each other (Id.). From a comparison of the HSQC spectra of Lac-FCHASE and GM3-FCHASE, it is obvious that the binding site is a C-3 of Gal due to the large downfield shift of H-3 from Gal H-3 and of Gal C-3 due to sialylation typical for sialyloligosaccharides (2-3) (Sabesan and Paulson, supra.). Also, as observed above for aNeu5Ac (2-3) j6Gal (1-4) 6GlcNAc-FCHASE (Gilbert et al., Supra.), The NOE of sialic acid H-3ax for Gal H-3 was typically observed from link aNeu5Ac (2-3) Gal.
(Table goes to next page)
ES 2 269 098 T3
<img file="ES2269098T3_D0001.tif" />
ES 2 269 098 T3
<img file="ES2269098T3_D0002.tif" />
ES 2 269 098 T3
<img file="ES2269098T3_D0003.tif" />
ES 2 269 098 T3
Accurate determination of the mass of the Cst-II enzyme product from Lac-FCHASE indicated that two sialic acids had been added to the Lac-FCHASE acceptor (Figure 4). The resonances of the proton were assigned from COZY, ID TOCSY and ID NOESY and the comparison of the chemical shifts, with known structures. Glc H-1 through H-6 and Gal H-1 through H-4 resonances were mapped from ID TOCSY on H-1 resonances. Neu5Ac resonances were assigned from COZY and confirmed by ID NOESY. The NOESY ID of the H-8 Neu5Ac resonances, H-9 at 4.16 ppm, was used to locate the H-9 and H-7 resonances (Michon et al., Supra). The appearance of the singlet of the Neu5Ac (2-3) H-7 resonance arising from the small neighboring coupling constants is typical of the 2-8 bond (Id.). The other resonances were assigned from the HSQC spectrum and the assignments of the <sup>13</sup>C for terminal sialic acid (Id.). The proton and the chemical shifts of carbon<sup>13</sup>C of the Gal unit were similar to those in GM3FCHASE, indicating the presence of the aNeu5Ac (2-3) Gal bond. The J values<sub>H</sub>h, the chemical shifts of <sup>13</sup>C and the proton of the two sialic acids, were similar to those of aNeu5Ac (2-8) Neu5Ac in the Neu5Ac a (2-8) -linked trisaccharide (Salloway et al. (1996) Infect. Immun. 64, 2945-2949 ) indicating the presence of that link. Therefore, the product was identified as GD3-FCHASE. Neu5Ac C-8 sialylation caused a 6.5 ppm downfield shift in its C-8 resonance from 72.6 ppm to 79.1 ppm.
The NOEs between residues for GD3-FCHASE were also typical of the sequence aNeu5Ac (2-8) aNeu5Ac (2-3) / Gal. The largest NOEs between residues of the two H-3 resonances<sub>ax</sub> at 1.7-1.8 ppm Neu5Ac (2-3) and Neu5Ac (2-8) are for the H-3 Gal and H-8 resonances of -8) Neu5Ac. The smallest NOEs are also observed between residues for H-4 of Gal and H-7 of -8) Neu5Ac. NOEs on FCHASE resonances are also observed due to the overlap of a FCHASE resonance with H-3 resonances, (Gilbert et al., Supra.). NOEs are also observed between H-3 residues<sub>eq</sub> from Neu5Ac (2-3) to H-3 of Gal. Also, the internal residues confirmed the assignments of the protons. The NOEs for the 2-8 bond are the same as those observed for the -8Neu5Aca2- polysaccharide (Michon et al., Supra).
The glycosidic linkages of sialic acid could also be confirmed by using the HMBC experiment that detects the correlations. <sup>3</sup>J (C, H) through the glycosidic bond. The results for both links, a-2,3 and a-2,8, indicate that the correlations<sup>3</sup>J (C, H) between the two anomeric C-2 resonances of Neu5Ac and the H-3 resonances of Gal, and H-8 of -8) Neu5Ac. Correlations within the residuals were also observed for the H-3 resonances.<sub>m</sub> and H-3<sub>eq</sub> of the two Neu5Ac residues. The correlation (C-1, H-2) of Glc is also observed since there was a partial overlap of the cross-linked peaks at 101 ppm with the cross-linked peaks at 100.6 ppm in the HMBC spectrum.
Accurate determination of the mass of the CgtA enzyme product from GM3-FCHASE indicated that an N-acetylated hexose unit had been added to the acceptor GM3-FCHASE (Figure 4). The product was identified as GM2-FCHASE since the glycoside proton and the chemical shifts of the<sup>13</sup>C were similar to those for the Gm2 oligosaccharide (GM2OS) (Sabesan et al. (1984) Can. J. Chem. 62, 1034-1045). From the HSQC spectrum for GM2-FCHASE, and the integration of its proton spectrum, there are now two resonances at 4.17 ppm and 4.18 ppm along with a new anomeric "d1" and two NAc groups at 2, 04 ppm. From the TOCSY and NOESY experiments, the resonance at 4.18 ppm was unambiguously assigned to Gal H-3 due to the strong NOE between H-1 and H-3. For / -galactopyranose, strong NOEs are observed within the residues between H1 and H-3 and H-1 and H-5 due to the axial position of the protons, and their short distances between the protons (Pavliak et al. (1993 ) J. Biol. Chem. 268, 14146-14152; Brisson et al. (1997) Biochemistry 36, 32783292; Sabesan et al. (1984) Can. J. Chem. 62, 1034-1045). From the TOCSY spectrum and comparison of GM2-FCHASE and GM2OS H1 chemical shifts (Sabesan et al., Supra) the resonance at 4.17 ppm is assigned as Gal H-4. Similarly, from the TOCSY and NOESY spectra, H-1 through H-5 of GalNAc and Glc, and H-3 through H-6 of Neu5Ac were assigned. Due to the width of the lines, the multiplet pattern of the resonances could not be observed. The other resonances were assigned from comparison with the HSQC spectrum of the precursor and the assignments by <sup>13</sup>C for GM2OS (Sabesan et al., Supra). By comparing the HSQC spectrum for the GM3- and GM2-FCHASE glycosides, a -9.9 ppm downfield shift occurred between the precursor and the product on the Gal C-4 resonance. Along with the NOEs within the residue for H-3 and H-5 of / -GalNAc, the NOE was also observed between residues from H1 of GalNAc to H-4 of Gal at 4.17 ppm, confirming the sequence / GalNAc ( 1-4) Gal. The NOEs observed were those expected from the conformational properties of ganglioside GM2 (Sabesan et al., Supra).
Accurate determination of the mass of the CgtB enzyme product from GM2-FCHASE indicated that one unit of hexose had been added to the GM2-FCHASE acceptor (Figure 4). The product was identified as GM1aFCHASE since the chemical shifts of the<sup>13</sup>C of the glycoside, were similar to those for oligosaccharide GM1a (Id.). The proton resonances were assigned from COZY, 1D TOCSY, and 1D NOESY. From 1D TOCSY on the additional "e1" resonance of the product, four resonances were observed with a multiplet pattern of / -galactopyranose. Starting with 1D TOCSY and 1D NOESY on the H-1 resonances of / GalNAc, the resonances of H-1 through H-5 were assigned. The multiplet pattern from H-1 to H-4 of pGalNAc was typical of the / -galactopyranosyl configuration, confirming the identity of this sugar to GM2-FCHASE. It was clear that because of glycosidation, the main perturbations occurred for the / GalNAc resonances, and there was a -9.1 ppm downfield shift between the acceptor and the product on the GalNAc C-3 resonance. Also, together with the NOEs within the residue for H-3, H-5 of Gal, a NOE between residues from H-1 of Gal to H3 of GalNAc, and a smaller one for H-4 of GalNAc were observed, confirming the / Gal (1-3) GalNAc sequence. The NOEs
ES 2 269 098 T3 observed were those expected from the conformational properties of ganglioside GM1a (Sabesan et al., Supra).
There was some discrepancy with the assignment of C-3 and C-4 resonances for j6Gal (1-4) in GM2OS and GM1OS, which are contrary to those of the published data (Sabesan et al., Supra). Previously, assignments were based on the comparison of chemical shifts of the<sup>13</sup>C with known compounds. For GM1aFCHASE, the H-3 assignment of Gal (1-4) was confirmed by observing its large neighborhood coupling, J<sub>2</sub>,<sub>3</sub> = 10Hz, directly in the HSQC spectrum processed with 2Hz / point in the proton dimension. The H-4 multiplet is much narrower (<5 Hz) due to the equatorial position of H-4 in galactose (Sabesan et al., Supra.). In Table 6, the C-4 and C-6 assignments of one of the sialic acids in (-8Neu5Ac2-)<sub>3</sub>, they also had to be reversed (Michon et al., supra) as confirmed by the H-4 and H-6 assignments.
Chemical shifts of <sup>13</sup>C of the FCHASE glycosides obtained from the HSQC spectrum were perfectly in agreement with those of the reference oligosaccharides shown in Table 6. Differences above 1 ppm were observed for some resonances and these are due to different aglycons in the reducing end. Excluding these resonances, the averages of the differences in chemical shifts between the FCHASE glycosides and their reference compounds were less than ± 0.2 ppm. Therefore, the comparison of the chemical shifts of the proton, of the values J<sub>H</sub>h and the chemical shifts of <sup>13</sup>C with known structures, and the use of NOEs or HMBC, were all used to determine binding specificity for different glycosyltransferases. The advantage of using the HSQC spectrum is that the assignment of the proton can be independently verified to confirm the assignment of the resonances of the<sup>13</sup>C of the atoms at the bond site. In terms of sensitivity, proton NOEs are the most sensitive, followed by HSQC and HMBC. Using an NMR nanoprobe instead of a 5mm NMR probe on the same amount of material significantly reduced the total acquisition time, making it possible to acquire an overnight HMBC experiment.
Discussion
For the purpose of cloning the C. jejuni LOS glycosyltransferases, an evaluation strategy similar to that previously used to clone Neisseria meningitidis α-2,3-sialyltransferase was employed (Gilbert et al., Supra). The activity evaluation strategy produced two clones that encoded two versions of the same α-2,3-sialyltransferase (cst-I) gene. ORF analysis suggested that a polypeptide at residue 430 is responsible for α-2,3-sialyltransferase activity. To identify other genes involved in LOS biosynthesis, a locus for LOS biosynthesis in the entire genome sequence of C. jejuni NCTC 11168 was compared to the corresponding locus of C. jejuni OH4384. Full open reading frames were identified and analyzed. Several of the open reading frames were individually expressed in E. coli, including a β1,4-N-acetylgalactosaminyl-transferase (cgtA), aj6-1,3-galactosyltransferase (cgtB), and a bifunctional sialyltransferase (cst-II).
In vitro synthesis of fluorescent derivatives of nanomole amounts of ganglioside mimics and their NMR analyzes unequivocally confirmed the binding specificity of the four cloned glycosyltransferases. Based on these data, it is suggested that the pathway described in Figure 4 is the one used by OH4384 from C. jejuni to synthesize a GT1a mimic. This role for cgtA is further supported by the fact that OH4342 from C. jejuni, which carries an inactive version of this gene, does not have / -M, 4-GalNAc in its central LOS core (Figure 1). The cst-II gene of OH4384 from C. jejuni exhibited both α-2,3- and α-2,8-sialyltransferase in an in vitro assay while cst-II from C. jejuni O: 19 ( sero-strain) showed only α-2,3-sialyltransferase activity (Table 5). This is consistent with a role for cst-II in the addition of a terminal α-2,8-linked sialic acid at OH4382 and OH4384 from C. jejuni, which have identical cst-II genes, but not in C. jejuni O: 19 (sero-strain, see Figure 1). There are 8 amino acid differences between the Cst-II homologues of C. jejuni (sero-strain) O: 19 and those of OH4382 / 84.
The bifunctionality of cst-II may have an impact on the success of C. jejuni infection since it has been suggested that the expression of the disialylated epitope may be involved in the development of neuropathic complications such as Guillain syndrome. Barré (Salloway et al. (1996) Infect. Immun. 64, 2945-2949). It is worth noting that its bifunctional activity is new among the sialyltransferases described so far. However, a bifunctional glycosyltransferase activity has been described for E. coli 3-deoxy-D-manno-octulosonic acid transferase (Belunis, CJ, and Raetz, CR (1992) J. Biol. Chem. 267: 9988 -9997).
The mono / bi-functional activity of cst-II and the activation / inactivation of cgtA appear to be two forms of the phase variation mechanisms that allow C. jejuni to make different surface carbohydrates that are presented to the host. In addition to those small gene alterations found among the three O: 19 strains (sero-strain, OH4382 and OH4384), there are major genetic rearrangements when loci between OH4384 and NCTC 11168 of C. jejuni (an O: 2 strain) are compared. Except for the prfB gene, the cst-I locus (including cysN and cysD) is found only in OH4384 of C. jejuni. There are significant differences in the organization of the locus for LOS biosynthesis between the strains OH4384 and NCTC 11168. Some of the genes are well conserved, some of them are poorly conserved while the others are unique for one or another strain. Two genes that are present as separate ORFs (# 5a: cgtA and # 10a: NeuA) in OH4384, are found as in-frame fusion ORFs in NCTC
ES 2 269 098 T3
11168 (ORF # 5b / # 10b). JB-N-acetylgalactosaminyltransferase activity was detected in this strain, suggesting that at least the cgtA part of the fusion may be active.
In summary, this Example describes the identification of different open reading frames that encode enzymes involved in the synthesis of lipooligosaccharides in Campylobacter.
Contents36
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
95 members in 12 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 11821399 | United States of America | P | |
| 11821399 | United States of America | P | |
| 19990118213P | United States of America | – | |
| 20000495406 | United States of America | – | |
| 49540600 | United States of America | A | |
| 49540600 | United States of America | A | |
| 00901455118213P | – | – | – |
| 495406 | – | – | – |
| US19990118213P | – | – | – |
| US20000495406 | – | – | – |
Members95
| Document | Office | Kind | |
|---|---|---|---|
| CA2360205A1 | Canada | A1 | |
| WO0046379A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2274300A | Australia | A | |
| WO0046379A8 | World Intellectual Property Organization (WIPO) | A8 | |
| EP1147200A1 | European Patent Office (EPO) | A1 | |
| US2002042369A1 | United States of America | A1 | |
| CA2441570A1 | Canada | A1 | |
| WO02074942A2 | World Intellectual Property Organization (WIPO) | A2 | |
| JP2002535992A | Japan | A | |
| US6503744B1 | United States of America | B1 | |
| WO02074942A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO02074942B1 | World Intellectual Property Organization (WIPO) | B1 | |
| US2003148459A1 | United States of America | A1 | |
| US2003157655A1 | United States of America | A1 | |
| US2003157656A1 | United States of America | A1 | |
| US2003157657A1 | United States of America | A1 | |
| US2003157658A1 | United States of America | A1 | |
| MXPA01007853A | Mexico | A | |
| EP1385941A2 | European Patent Office (EPO) | A2 | |
| US6699705B2 | United States of America | B2 | |
| US6723545B2 | United States of America | B2 | |
| AU772569B2 | Australia | B2 | |
| MXPA03008565A | Mexico | A | |
| JP2004524033A | Japan | A | |
| AU2004203474A1 | Australia | A1 | |
| US2004180406A1 | United States of America | A1 | |
| US2004203103A1 | United States of America | A1 | |
| US2004203112A1 | United States of America | A1 | |
| US2004203113A1 | United States of America | A1 | |
| US2004219638A1 | United States of America | A1 | |
| US2004229263A1 | United States of America | A1 | |
| US2004229272A1 | United States of America | A1 | |
| US2004229313A1 | United States of America | A1 | |
| US6825019B2 | United States of America | B2 | |
| US2004259140A1 | United States of America | A1 | |
| US2004259203A1 | United States of America | A1 | |
| US2004265875A1 | United States of America | A1 | |
| US2005048630A1 | United States of America | A1 | |
| US2005064550A1 | United States of America | A1 | |
| US2005084891A1 | United States of America | A1 | |
| US6905867B2 | United States of America | B2 | |
| US6911337B2 | United States of America | B2 | |
| US2005227248A1 | United States of America | A1 | |
| US7026147B2 | United States of America | B2 | |
| EP1652927A2 | European Patent Office (EPO) | A2 | |
| EP1147200B1 | European Patent Office (EPO) | B1 | |
| AT329036T | Austria | T | |
| ATE329036T1 | Austria | T1 | |
| US7078207B2 | United States of America | B2 | |
| EP1652927A3 | European Patent Office (EPO) | A3 | |
| DE60028541D1 | Germany | D1 | |
| US2006166317A1 | United States of America | A1 | |
| DK1147200T3 | Denmark | T3 | |
| PT1147200E | Portugal | E | |
| US7138258B2 | United States of America | B2 | |
| US7166717B2 | United States of America | B2 | |
| US7169593B2 | United States of America | B2 | |
| US7169914B2 | United States of America | B2 | |
| US2007048854A1 | United States of America | A1 | |
| US7189836B2 | United States of America | B2 | |
| US7192756B2 | United States of America | B2 | |
| AU2002237122B2 | Australia | B2 | |
| ES2269098T3This record | Spain | T3 | |
| US7202353B2 | United States of America | B2 | |
| US7208304B2 | United States of America | B2 | |
| US7211657B2 | United States of America | B2 | |
| US7217549B2 | United States of America | B2 | |
| US7220848B2 | United States of America | B2 | |
| DE60028541T2 | Germany | T2 | |
| US7238509B2 | United States of America | B2 | |
| AU2007202898A1 | Australia | A1 | |
| AU2004203474B2 | Australia | B2 | |
| AU2004203474B9 | Australia | B9 | |
| EP1914302A2 | European Patent Office (EPO) | A2 | |
| US7371838B2 | United States of America | B2 | |
| EP1652927B1 | European Patent Office (EPO) | B1 | |
| AT395415T | Austria | T | |
| ATE395415T1 | Austria | T1 | |
| US7384771B2 | United States of America | B2 | |
| DE60038917D1 | Germany | D1 | |
| EP1914302A3 | European Patent Office (EPO) | A3 | |
| JP2008259515A | Japan | A | |
| JP2008259516A | Japan | A | |
| ES2308364T3 | Spain | T3 | |
| US7462474B2 | United States of America | B2 | |
| JP2009000122A | Japan | A | |
| US7608442B2 | United States of America | B2 | |
| JP4397955B2 | Japan | B2 | |
| JP4398502B2 | Japan | B2 | |
| JP4431283B2 | Japan | B2 | |
| JP4431312B2 | Japan | B2 | |
| JP4460615B2 | Japan | B2 | |
| CA2441570C | Canada | C | |
| CA2360205C | Canada | C | |
| EP1914302B1 | European Patent Office (EPO) | B1 |
Numbers
- Publication
- 2269098
- Publication, DOCDB
- 2269098
- Publication, EPODOC
- ES2269098T
- Application
- 901455
- Application, DOCDB
- 00901455
- Application, EPODOC
- ES20000901455T
Titles2
- Spanish
- GLICOSIL TRANSFERASAS DE CAMPILOBACTER PARA LA BIOSINTESIS DE GANGLIOSIDOS E IMITACIONES DE GANGLIOSIDOS.
- English
- GLICOSIL CAMPILOBACTER TRANSFERS FOR BIOSYNTHESIS OF GANGLIOSIDES AND IMITATIONS OF GANGLIOSIDS.
Classification
- CPC, 5
- C12P19/18
- C12N9/1029
- C12N9/1051
- C12N9/1081
- C12N9/1241
- IPC, 8
- C12N15 09
- C12N15 54
- C12N1 21
- C12N9 10
- C12N9 12
- C12N9 88
- C12P19 18
- C12Q1 68