Fusion proteins displayable on the surface of filamentous phage and a recombinant, filamentous phage bearing the same
Abstract
This record has no abstract on file.
Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
8 claims: 2 independent, 6 dependent
- 1120,940/5 245 WHAT IS CLAIMED IS:1. A fusion protein, comprising: (a) a carrier protein moiety essentially identical in amino acid sequencewith at least a functional portion of a mature coat protein of afilamentous phage and having the same activity said carrier proteinmoiety acting, when the fusion protein is produced in a suitable hostcell infected by the phage, to cause the display of the fusion proteinor a processed form thereof on the surface of the phage, and (b) a peptide or proteinaceous binding domain foreign to said coatprotein, with the proviso that when the coat protein is the gene illprotein of said filamentous phage, (b) is coupled essentially to theamino terminal of said carrier protein moiety, and the carrier proteinmoiety comprises essentially the entire gene III protein.
- 6The fusion protein where (b) is foreign to said filamentous phage. 120,940/4 246
Independent claims2
1,216 paragraphs in 231 sections, as filed
<img img-format="tif" img-content="drawing" file="IL120940AD00021.tif" id="idf0001" />
120,940/2
FUSION PROTEINS DISPLAYABLE ON THE SURFACE OF FILAMENTOUSPHAGE AND A RECOMBINANT, FILAMENTOUS PHAGEBEARING THE SAME οχω oxo υο cnxswn >ιγρχ υη>ηnx χ^υη 1
Field of the Invention
The present invention relates to a fusion protein.
The present application is divided from Israel Specification No. 91501,filed September 1, 1989. In order that the invention may be better understoodand appreciated, description from Israel Specification No. 91501 is includedherein, it being understood that this is for background purposes only, the subjectmatter of Israel Specification 91501 being specifically disclaimed and notforming a part of the present invention. ΙΑ
Information Disclosure Statement
The amino acid sequence of a protein determines its three-dimensional (3D) structure, which in turn 15 determines protein functioning (EPST63, ANFI73). The system of classification of protein structure of Schulz and Schirmer (SCHU79. ch 5) is adopted herein.
The 3D structure of a protein is essentially 20 unaffected by the identity of the amino acids at someloci; at other loci only one or a few types of aminoacid is allowed (SHOR8S, EISE8S, REID88). Generally,loci where wide variety is allowed have the amino acidside group directed toward the solvent. While limited 25 variety is allowed where the side group is directedtoward other parts of the protein. (See also SCHU79.P169-171 and CREI84, p239-245, 314-315).
The secondary structure (helices, sheets, turns, 3 0 loops) of a protein is determined mostly by localsequence. Certain amino acids tend to be correlatedwith certain secondary structures and the commonly usedChou-Fasman (CHOU74, CHOU78a, CHOU78b) rules depend onthese correlations. However, every amino acid type has
<img img-format="tif" img-content="drawing" file="IL120940AD00022.tif" id="idf0002" />
.../Α been observed in helices and in both parallel and antiparallel sheets. Pentapeptides of identical sequence are found in different proteins; in some cases the conformations of the pentapeptides are very 5 different (KABS84, ARGO87).
Turns and loops tolerate insertions and deletionsmore readily than do other secondary structures(RICH81, TH0R88, SUTC87a); related proteins differ most 10 in loops and turns.
Changing three residues in subtilisin fromBacillus amvloliquefaciens to be the same as thecorresponding residues in subtilisin from B. 15 licheniformis produced a protease that had nearly thesame activity as the subtilisin from the latterorganism; 82 differences remained in the sequences.The three residues changed were chosen because theywere the only differences within 7 Angstroms (A) of the 20 active site (WELL87a).
Schulz and Schirmer summarize many observations onthe binding of proteins to other molecules (SCHU79.p98-105). For example, haemoglobin alpha chains bind 25 very tightly to haemoglobin beta chains (delta G morenegative than -11.0 Kcal/mole); antibodies bind tightlyto antigens (K^s range from 10~6 to 10-14 M, is thedissociation constant equal to [A][B]/[A:B]); basicbovine pancreatic trypsin inhibitor (BPTI) binds 30 tightly to trypsin (K^ = 6.0 x 10“14 M (TSCH87), deltaG = -18.0 Kcal/mole); and avidin binds to biotin (Kj =1.3 x 10-15 M (CREI84. p362)). In each case the binding results from complementarity of the surfacesthat come into contact: bumps fit into holes, unlike 35 charges come together, dipoles align, and hydrophobic atoms contact other hydrophobic atoms. Although bulkwater is excluded, individual water molecules arefrequently found filling space in intermolecularinterfaces; these waters usually form hydrogen bonds toone or more atoms of the protein or to other boundwater. 10 15 20
The factors affecting protein binding are known,(CHOT75, CHOT76, SCHU79. p98-107, and CREI84. Ch8), butdesigning new complementary surfaces has proveddifficult. Although some rules have been developed forsubstituting side groups (SUTC87b), the side groups ofproteins are floppy and it is difficult to predict whatconformation a new side group will take. Further, theforces that bind proteins to other molecules are allrelatively weak and it is difficult to predict theeffects of these forces. Hence, it is difficult todesign superior binding proteins based on theory alone(QUIO87).
Enzyme-substrate affinity, however, hasfortuitously been increased by protein engineering (WILK84).Bacillus A point mutant of tyrosyl tRNA synthetase of 25 exhibits a 100-foldSubstitution of onea surface locus may 30 stearothermophilusincrease in affinity for ATPamino acid for another at profoundly alter binding properties of the proteinother than substrate binding, without affecting thetertiary structure of the protein. For example, insickle-cell haemoglobin the change of the surfaceresidue E6 to V in the beta chains causesdeoxyhaemoglobin-S to form fibers through self binding(DICK83. pl25-145) ; the tertiary and quaternary structure of the haemoglobin are not changed (PADL85,WISH75, WISH76). 35
Changing a single amino acid in 3PTI greatlyreduces its binding to trypsin, but some of the newmolecules retain the parental characteristics ofbinding to and inhibiting chymotrypsin, while othersexhibit new binding to elastase (TANK77; TSCH87).Changes of single amino acids on the surface of thelambda Cro repressor greatly reduce its affinity forthe natural operator 0r3, but greatly increase thebinding of the mutant protein to a mutant operator(EISE85). Thus changing the surface of a bindingprotein may alter its specificity without abolishingbinding activity.
The recently developed techniques of "reversegenetics" have been used to produce single specificmutations at precise base pair loci (OLIP86, OLIP87,and AU5U87). Mutations are generally detected bysequencing and in some cases by loss of wild-typefunction. These procedures allow researchers toanalyze the function of each residue in a protein(MILL38) or of each base pair in a regulatory DNAsequence (CHEN88). In these analyses, the norm hasbeen to strive for the classical goal of obtainingmutants carrying a single alteration (AUSU87).
Reverse genetics is often applied to codingregions to determine which residues are most importantto protein structure and function; isolation of asingle mutant at each residue of the protein gives aninitial estimate of which residues play crucial roles.
Prior to the method of Israel Specification 91501, twogeneral approaches have been developed to create novelmutant proteins through reverse genetics. In one approach, dubbed "protein surgery" (DILL87), a specificsubstitution is introduced at a single protein residueto determine the effects on structure and function ofspecific substitutions (CRAI85) (RAOS87) (BASH87) .However, many desirable protein alterations requiremultiple amino acid substitutions and thus are notaccessible through single base changes or even throughall possible amino acid substitutions at any oneresidue.
The other approach has been randomly to generate avariety of mutants at many loci within a cloned geneusing mutagenic chemicals or radiation. The specificlocation and nature of the change are determined by DNAsequencing. (PAKU86) This approach is limited by thenumber of colonies that can be examined. Also, it doesnot take advantage of any knowledge of the proteinstructure and its relationship to binding activity.
Progress toward rules governing substitutions ofamino acids (ULME83) has been greatly hampered by theextensive efforts involved in using either method andthe practical limitations on the number of coloniesthat can be inspected (ROBE86).
The term "saturation mutagenesis" with referenceto synthetic DNA is generally taken to mean generationof a population in which: a) every possible single-basechange within a fragment of a gene of DNA regulatoryregion is represented, and b) most mutant genes containonly one mutation. Thus a set of all possible singlemutations for a 6 base pair length of DNA comprises apopulation of 18 mutants. Oliphant et al. (OLIP86) andOliphant and Struhl (OLIP87) have demonstrated ligationand cloning of highly degenerate oligonucleotides and have applied saturation mutagenesis to the study ofpromoter sequence and function. They suggest thatsimilar methods could be used to study geneticexpression of proteins, but they do not say how to: a)choose protein residues to vary, or b) select or screenmutants with desirable properties.
Reidhaar-Olson and Sauer (REID88) have usedsynthetic degenerate oligo-nts to vary simultaneouslytwo or three residues through all twenty amino acids inthe dimer interface of cl repressor from bacteriophagelambda. They give no discussion of the limits on howmany residues could be varied at once nor do theymention the problem of unequal abundance of DNAencoding different amino acids. They looked forproteins that either had wild-type dimerization or thatdid not dimerize. They did not seek proteins havingnovel binding properties and did not report any.
Several researchers have designed and synthesizedproteins de novo. These designed proteins are smalland most have been synthesized in vitro as polypeptidesrather than genetically. Gutte and colleagues havemade a polypeptide that binds DDT in 55% ethanol(MOSE83). Recently Moser et al. (MOSE87) reportedgenetic expression in E. coli both of the designed 24residue DDT-binding protein and of fusions of the DDT-binding sequence to LacZ. They state that design ofbiologically active proteins is currently impossible.
Erickson et al. (ERIC86) have designed andsynthesized a series of proteins that they have namedbetabellins, that are meant to have beta sheets. Theysuggest use of polypeptide synthesis with mixedreagents to produce several hundred analogous <<7\ betabellins, and use of a column to recover analogueswith high affinity for a chosen target compound boundto the column. They envision successive rounds ofmixed synthesis of variant proteins and purification byspecific binding. They do not discuss how residuesshould be chosen for variation. Because proteinscannot be amplified, the researchers must sequence therecovered protein to learn which substitutions improvebinding. The researchers must limit the level ofdiversity so that each variety of protein will bepresent in sufficient quantity for the isolatedfraction to be sequenced.
Methods have been developed to separate cellsthrough their affinity to various substances. Methodsapplied to animal cells reveal common problems: a) non-specific interactions between cells and affinitysupports, and b) irreversible binding of cells toaffinity matrices (BONN85).
Ferenci and collaborators have published a seriesof papers on the chromatographic isolation of mutantsof the maltose-transport protein LamB of E. coli(WAND79, FERE80a, FERE80b, FERE80C, FERE82a, FERE82b,FERE83, CLUN84, FERE86a, FERE86b, FERE86C, FERE87a,FERE87b, HEIN87, and HEIN88). The papers report thatspontaneous and induced mutants at the lamB geneticlocus can be isolated by chromatography over a columnsupporting immobilized maltose, maltodextrins, orstarch. The reports speculate that other applicationsare possible, but specifically mention only theelucidation of the residues responsible for theselectivity of the maltodextrin pore or similar poreproteins. The mutant proteins were non-chimeric, andno attempt was made to obtain binding to a new target. 8
Both FERE86a and CLUN84 point up thedifficulties of working with live bacteria that canmetabolize chemicals and change their physiologicalbehavior during the chromatographic experiment. A fragment of a heterologous gene can beintroduced into bacteriophage Fl gene III (SMIT85). Ifthe inserted gene preserves the original reading frame,expression of the altered gene III causes an inserteddomain to appear in the gene III protein. The resulting strain of fl virions are adsorbed by anantibody against the protein encoded by theheterologous DNA. The phage were eluted at pH 2.2 andretained some infectivity. However, the single copy offl gene III was used for insertion of the heterologousgene so that all copies of gene III protein wereaffected; infectivity of the resultant phage wasreduced 25-fold.
Smith presented his method as a way to isolatecloned genes using antibodies to the gene products. Hemade no mention of mutagenizing the inserted geneticmaterial or of inducing novel binding properties in theinserted protein domain. A fragment of the repeat region of thecircumsporozoite protein from Plasmodium falciparum hasbeen expressed on the surface of M13 as an insert inthe gene III protein (CRUZ88). The recombinant phagewere both antigenic and immunogenic in rabbits. Theauthors do not suggest mutagenesis of the insertedmaterial.
Gene fragments coding - for hepatitis B virusantigens have been fused to fragments of lamB. and ifthe fusion is in a region coding for exposed domains ofLamB, the HBV antigens appear on the cell surface and 5 are immunogenic (CHAR87). Charbit et al. (CHAR87)suggest use of these engineered strains for developmentof a live bacterial vaccine; they did not suggestmutagenesis of the fused heterologous gene fragments,nor development of binding capabilities. 10
Ladner, US Patent No. 4,704,692, "Computer BasedSystem and Method for Determining and DisplayingPossible Chemical Structures for Converting Double- orMultiple-Chain Polypeptides to Single-Chain 15 Polypeptides" describes a design method for convertingproteins composed of two or more chains into proteinsof fewer polypeptide chains, but with essentially thesame 3D structure. There is no mention of variegatedDNA and no genetic selection. Ladner and Bird, 20 W088/01649 (Publ. March 10, 1988) disclose the specific application of computerized design of linker peptidesto the preparation of s,ingle chain antibodies.
Ladner, Glick and Bird, W088/06630 (publ. 7 Sept. 25 1988) (LGB) speculate that diverse single chain antibody domains may be screened for binding to aparticular antigen by varying the DNA encoding thecombining determining regions of a single chainantibody, subcloning the SCAD gene into the gpV gene of 30 phage lambda so that a SCAD/gpV chimera is displayed onthe outer surface of the phage, and selecting phagewhich bind to the antigen through affinitychromatography. The only antigen mentioned is bovinegrowth hormone. No other binding molecules, targets, 35 carrier organisms, or outer surface proteins are 120,940/2 10 discussed, nor is there any mention of the method or degree of mutagenesis.
Ladner and Bird, W088/06601 (published September 7, 1988) suggestthat single chain ‘pseudodimeric’ repressors (DNA-binding proteins) may beprepared by mutating a putative linker peptide, followed by in vivo selection andthat mutation and selection may be used to create a dictionary of recognitionelements for use in the design of asymmetric repressors. The repressors are notdisplayed on the outer surface of an organism.
No admission is made that any cited reference is prior art or pertinent priorart, and the dates given are those' appearing on the reference and may not beidentical to the actual publication date.
Israel Specification No. 91501, the present specification, and furtherdivisional specifications 120,939 and 120,941 relate to the construction,expression and selection of mutated genes that specify novel proteins withdesirable binding properties, as well as these proteins themselves. Thesubstances bound by these proteins, hereinafter referred to as ‘targets,’ may be,but need not be, proteins. Targets may include other biological or syntheticmacromolecules as well as organic and inorganic molecules.
The novel binding proteins may be obtained: (1) by mutating a geneencoding a known binding protein within the sub-sequence encoding a knownbinding domain; or (2) by taking such a sub-sequence of the gene for a firstprotein and combining it with all or part of a gene for a second protein (whichmay, or may not, itself be a known binding protein); or (3) by mutating a gene 11 encoding a protein which, while not possessing a known binding activity,possesses a secondary or higher structure that lends itself to binding activity(clefts, grooves, etc.); or (4) by mutating a gene encoding a known bindingprotein but not in the sub-sequence known to cause the binding. The proteinfrom which the novel binding protein is derived need not have any specificaffinity for the target material.
More specifically, according to the invention described and claimed inIsrael Specification 91501, there is provided a method of obtaining a nucleic acidencoding a proteinaceous binding domain that binds a predetermined targetmaterial, other than the antigen combining site of an antibody which specificallybinds said domain, comprising: a) preparing a variegated population of amplifiable genetic packages,said genetic packages being selected from the group consisting of cells, sporesand viruses, each said genetic package being genetically alterable and having anouter surface including a genetically determined outer surface protein, eachpackage including a first nucleic acid construct coding for a chimeric potentialbinding protein, each said chimeric protein comprising, and each said constructcomprising, DNA encoding (i) a potential binding domain which is a mutant of astable predetermined domain of a predetermined parental protein, other than asingle chain antibody, comprising one or more identifiable surface residues, andfor which both an affinity molecule and an amino acid sequence are eitheravailable or obtainable; and (ii) an outer surface transport signal for obtaining thedisplay of the potential binding domain on the outer surface of the geneticpackage, the expression of which construct results in the display of said chimericpotential binding protein and its potential binding domain on the outer surface ofsaid genetic package; and wherein said variegated population of genetic packagescollectively displays a plurality of different potential binding domains, the 120*940/2 12 differentiation among said plurality of different potential binding domainsoccurring through the at least partially random variation of one or morepredetermined amino acid positions of said parental binding domain to randomlyobtain at each said position an amino acid belonging to a predetermined set oftwo or more amino acids, the amino acids of said set occurring at said position instatistically predetermined expected proportions, the genetic messageencapsulated by said genetic packages being amplifiable in vitro or by cell cultureof said genetic packages and separable on the basis of the potential bindingdomain displayed thereon; b) causing the expression of said chimeric potential binding proteinsand the display of said potential binding domains on the outer surface of saidpackages; c) contacting said packages with the predetermined target materialsuch that said potential binding domains and the target material may interact; d) separating packages displaying a potential binding domain thatbinds the target material from packages that do not so bind, on the basis of theirability to bind with the target material in step (c), and e) recovering at least one package displaying on its outer surface achimeric binding protein comprising a stable successful binding domain (SBD)which bound said target, said package comprising nucleic acid encoding saidsuccessful binding domain, and amplifying said SBD-encoding nucleic acid invivo or in vitro.
In Israel Specification No. 120*939 , also divided from Israel
Specification No. 91501, there is described and claimed a method of obtaining anucleic acid encoding a proteinaceous binding domain that binds a predeterminedtarget, which comprises: 13 a) providing a population of amplifiable genetic packages, saidgenetic packages being selected from the group consisting of cells, spores andviruses, each said genetic package being genetically alterable and having an outersurface including a genetically determined outer surface protein, each packageincluding a first nucleic acid construct coding for a chimeric potential bindingprotein, each said chimeric protein comprising, and each said constructcomprising, DNA encoding (i) a potential binding domain which is a mutant of astable predetermined domain of a predetermined parental protein, comprising oneor more identifiable surface residues, and for which both an affinity molecule andan amino acid sequence are either available or obtainable, and (ii) an outersurface transport signal for obtaining the display of the potential binding domainon the outer surface of the genetic package, the expression of which constructresults in the display of said chimeric potential binding protein and its potentialbinding domain on the outer surface of said genetic package; and wherein saidpopulation of genetic packages collectively display the genetic messageencapsulated by said genetic packages being amplifiable in vitro or by cell cultureof said genetic packages or host cells transformed therewith and separable on thebasis of the potential binding domain displayed thereon; b) causing the expression of said chimeric potential binding proteinsand the display of said potential binding domains on the outer surface of saidpackages; c) contacting said packages with the predetermined target materialsuch that said potential binding domains and the target material may interact; d) separating packages displaying a potential binding domain thatbinds the target material from packages that do not so bind, on the basis of theirability to bind with the target material in step (c), and e) recovering at least one package displaying on its outer surface achimeric binding protein comprising a stable successful binding domain (SBD) 120,940/5
13A which bound said target, said package comprising nucleic acid encoding saidsuccessful binding domain, and amplifying said SBD-encoding nucleic acid in vivoor in vitro; with the proviso that when the parental protein is a single chain antibody, (i)the genetic package is a filamentous phage, and/or (ii) the outer surface of thegenetic package presents, not only said chimeric protein, but also the cognate wildtype outer surface protein.
Israel Specification 120,941, also divided from Israel Specification 91501,provides a chimeric binding protein, comprising (i) a proteinaceous binding domainwhich binds to a target sufficiently strongly so that the disassociation constant ofthe binding domain:target complex is less than 10'6 moles/liter, and (ii) at least afunctional portion of a coat protein of a virus, said portion acting, when the chimericprotein is produced in a suitable host cell, to cause the display of the chimericbinding protein or a processed form thereof on the outer surface of the virus, saidbinding domain being capable of binding to a target material which said coatprotein does not preferentially bind, said bindjng domain being foreign to the nativecoat proteins of said virus, with the proviso that when the proteinaceous bindingdomain is a single chain antibody, the virus is a filamentous phage.
SUMMARY OF THE INVENTION
The present provides a fusion protein, comprising (a) a carrier proteinmoiety essentially identical in amino acid sequence with at least a functionalportion of a mature coat protein of a filamentous phage and having the sameactivity said carrier protein moiety acting, when the fusion protein is produced in asuitable host cell infected by the phage, to cause the display of the fusion proteinor a processed form thereof on the surface of the phage, and (b) a peptide orproteinaceous binding domain foreign to said coat protein, with the proviso thatwhen the coat protein is the gene III protein of said filamentous phage, (b) iscoupled essentially to the amino terminal of said carrier protein moiety, and thecarrier protein moiety comprises essentially the entire gene III protein. 120,940/4
13B
In a first preferred embodiment, there is provided a fusion protein whereinthe coat protein is the gene III protein.
In a second preferred embodiment, there is provided a fusion proteinwherein the coat protein is the gene VIII protein.
For the purposes of this invention, the term “potential binding protein” refersto a protein encoded by one species of DNA molecule in a population of variegatedDNA wherein the region of variation appears in one or more subsequencesencoding one or more segments of the polypeptide having the potential of servingas a binding domain for the target substance.
From time to time, it may be helpful to speak of the “parent sequence” of thevariegated DNA. When the novel-binding domain sought is an analogue of aknown binding domain, the parent sequence is the sequence that encodes theknown binding domain. The variegated DNA will be identical with this parentsequence at most 14 loci, but will diverge from it at chosen loci. When apotential binding domain is designed from firstprinciples, the parent sequence is a sequence whichencodes the amino acid sequence that has been predictedto form the desired binding domain, and the variegatedDNA is a population of "daughter DNAs" that are relatedto that parent by a high degree of sequence similarity.
The fundamental principle of the invention is oneof forced evolution. The efficiency of the forcedevolution is greatly enhanced by careful choice ofwhich residues are to be varied. The 3D structure ofthe potential binding domain is a key determinant inthis choice. First a set of residues that cansimultaneously contact one molecule of the target isidentified. Then all or some of the codons encodingthese residues are varied simultaneously to produce avariegated population of DNA. The variegatedpopulation of DNA is used to transform cells so that avariegated population of genetic packages is produced.
The mixed population of genetic packagescontaining genes encoding possible binding proteins isenriched for packages containing genes that expressproteins that in fact bind to the target ("successfulbinding domains"). After one or more rounds of suchenrichment, one or more of the chosen genes areexamined and sequenced. If desired, new loci ofvariation are chosen. The selected daughter genes ofone generation then become the parent sequences for thenext generation of variegated DNA, beginning the next"variegation cycle." Such cycles are continued until aprotein with the desired target affinity is obtained.
14A 91501/1
The following publications are mentioned herein for background purposes, but as explained further below, do not teach or suggest the present invention as herein defined. CA 108:183391u (Heine, et al., 1988) describesgeneration of LamB mutants by spontaneous mutation, or byhydroxylamine or random linker mutagenesis. None of thesetechniques result in "variegation," as herein defined andclaimed. The researcher has no control over where mutationsoccur, or what mutations are made. At best, the researcherknows the preferred sites of hydroxylamine attack, and thetype of mutations it induces at these sites. Claim 1emphasizes that in the method of the present invention,predetermined residues are mutated so that the substitutedresidues belong to a predetermined set and occur inpredetermined expected proportions. With Heine's method,mutation can occur at any residue; there is no control overthe nature of the mutation. Moreover, Heine, et al.prepared LamB mutants, not chimeras in which a new bindingdomain was grafted onto LamB.
Moreover, CA 108:18339lu describes mutation of a nativebacterial membrane protein. The mutations do not convertthis protein into a chimera of a native membrane protein (orof just its outer surface transport signal) and a bindingdomain of a different protein. CA 101:126595b (Clune, et al., 1984) describes affinitychromatography selection of LamB mutants (prepared asdescribed above) with higher affinity for the normal LamBligands such as starch and/or maltose. The same objectionsapply to this reference as to Heine, et al. It is alsoworth noting that Clune, et al. did
not screen the LamB
14B 91501/1
<img img-format="tif" img-content="drawing" file="IL120940AD00023.tif" id="idf0003" />
mutants for binding to any ligand which is not specificallybound by wild-type LamB. CA 107:54416m (Heine, et al., 1987) characterizes theeffects of specific mutations on starch and maltose bindingand transport. It adds nothing of significance to theabove-mentioned references. CA 107:128472a indeed teaches expression of aheterologous protein in a host cell. However, it does notprovide motiation to variegate a parental binding domain andto express the resulting mutant domains as fusions to nativeviral coat or cell membrane protein for display on the coatof the virus or the membrane of the cell. CA 109:209280e discusses the sequence similaritybetween a sequence in the reovirus type 3 cell attachmentprotein and a monoclonal antibody which mimics reovirus byattaching to the same cell surface receptor. The attachmentprotein is a binding protein, and it lies on the surface ofa virus, however, its binding domain was not mutated, andthis protein was wholly native to the virus on which it wasdisplayed. CA 106:169916w describes a gene which encodes aprecursor of IgA protease. The Examiner apparently believesthat this precursor protein includes the transport signal ofclaim 1. That is not correct. The protein is clearlyidentified as being a secreted protein; it is well-known inthe art that with such proteins, the leader is cleaved off,leaving the mature protein, as the protein passes throughthe inner membrane. Since the protein leaves the cell (notethe reference to "extracellular"), there is no analogy withthe protein of the present invention. Even, however, if the 9 120,940/2
14C protein were a membrane protein such as LamB, it is merely a wild type proteinin which the cell localization signals and the binding region are nativelyassociated.
The appended claims are hereby incorporated by reference into thisspecification as an enumeration of the preferred embodiments.
As indicated above, while described, exemplified and illustrated herein, thesubject matter of Israel Specification No. 91501 and divided Israel Specifications120,939 anj 120,941 no longer constitutes a part of the present invention, and is specifically disclaimed.
With specific reference now to the examples in detail, it is stressed that theparticulars described are by way of example and for purposes of illustrativediscussion of the preferred embodiments of the present invention only, and arepresented in the cause of providing what is believed to be the most useful and readilyunderstood description of the principles and conceptual aspects of the invention. Inthis context, it is to be noted that only subject matter embraced in the scope of theclaims appended hereto, whether in the manner defined in the claims or in a mannersimilar thereto and involving the main features as defined in the claims, is intendedto be included in the scope of the present invention, while subject matter of IsraelSpecification 91501 and divided Specifications 120,939 ancj 120,941 , although described and exemplified to provide background and better understandingof the invention, is not intended for inclusion as part of the present invention. 15
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is a schematic showing the relationshipsbetween various types of Binding Domains (3D) . 10 Figure 2 is a flow chart showing the major steps usedto create a novel protein with affinity for a pre-determined target.
Figure 2 is a schematic of a P3D contacting a molecule 15 of taroet material.
Figure 4 is a schematic of the construction of pLG3from K13mplS and p3R322. 20 Figure 5 is a schematic of the construction of pLG7from pLG3 and synthetic DNA.
DETAILED DESCRIPTION OF TEE INVENTION 25 Sec. 0.1: Overview:
The present invention separates mutated genes thatspecify novel proteins with desirable bindingproperties from closely related genes that specify 30 proteins with no or undesirable binding properties, by:1) arranging that the product of each mutated gene bedisplayed on the outer surface of a replicable geneticpackage that contains the gene, and 2) using affinityseparation incorporating a desirable target material to 35 enrich the population of packages for those packages 16 containing genes specifying proteins with improvedbinding to that target material.
Let Κθ (x,y) be a dissociation constant, [x] [y] KD(x,y) ----------.
[x:y]
For the purposes of the appended claims, a proteinP is a binding protein if (1) for one molecular, ionic or atomic species A,the dissociation constant KD (P,A) < 10“6 moles/liter, and (2) for a different molecular, ionic or atomic species B, KD (Ρ,Β) > 10”1 moles/liter.
As a result of these two conditions, the protein Pexhibits specificity for A over B, and a minimum degreeof affinity (or avidity) for A.
When a domain of a protein is primarilyresponsible for the protein's ability to specificallybind a chosen target, it is referred to herein as a"binding domain" (BD). We engineer the appearance of astable protein domain, denoted as an "initial potentialbinding domain" (IPBD), on the surface of a geneticpackage. The present invention is concerned with theexpression of numerous, diverse, variant "potentialbinding domains" (PBD), all related to a "parentalpotential binding domain" (PPBD) such as the bindingdomain of a known binding protein, and with selection
<img img-format="tif" img-content="drawing" file="IL120940AD00024.tif" id="idf0004" />
17 and amplification of the genes encoding the mostsuccessful mutant PBDs. An IPBD is chosen as PPBD tothe first round of variegation. Selection-through-binding isolates one or more "successful binding 5 domains" (SBD). An SBD from one round of variegationand selection-through-binding is chosen to be the PPBDfor the next round. The invention is not, however,limited to proteins with a single BD since the methodmay be applied to any or all of the BDs of the protein, 10 sequentially or simultaneously. The relationships ofthe various BDs are illustrated in Figure 1.
The term "variegated DNA" refers to a populationof molecules that have the same base sequence through 15 most of their length, but that vary at a limited numberof defined loci, preferably 5-10 codons. A molecule ofvariegated DNA can be introduced into a plasmid so thatit constitutes part of a gene (OLIP86, OLIP87, AUSU87,REID88). When plasmids containing variegated DNA are 20 used to transform bacteria, each cell makes a versionof the original protein. Each colony of bacteria mayproduce a different version from any other colony. Ifthe variegations of the DNA are concentrated at lociknown to be on the surface of the protein or in a loop, 25 a population of proteins will be generated, manymembers of which will fold into roughly the same 3Dstructure as the parent protein. The specific bindingproperties of each member, however, may be differentfrom each other member. It remains to sort out the 30 colonies containing genes for proteins with desirablebinding properties from those that do not exhibit thedesired affinities. A "single-chain antibody" is a single chain 35 polypeptide comprising at least 200 amino acids, said 18 amino acids forming two antigen-binding regionsconnected by a peptide linker that allows the tworegions to fold together to bind the antigen. Eitherthe two antigen-binding regions must be variabledomains of known antibodies, or they must (1) each foldinto a beta barrel of nine strands that are spatiallyrelated in the same way as are the nine strands ofknown antibody variable light or heavy domains, and (2)fit together in the same way as do the variable domainsof said known antibody. Generally speaking, this willrequire that, with the exception of the amino acidscorresponding to the hypervariable region, there is atleast 88% homology with the amino acids of the variabledomain of a known antibody.
The term "affinity separation means" includes, butis not limited to: a) affinity column chromatography,b) batch elution from an affinity matrix material, c)batch elution from an affinity material attached to aplate, d) fluorescence activated cell sorting, and e)electrophoresis in the presence of target material."Affinity material" is used to mean a material withaffinity for the material to be purified, called the"analyte". In most cases, the association of theaffinity material and the analyte is reversible so thatthe analyte can be freed from the affinity materialonce the impurities are washed away.
Affinity column chromatography, batch elution froman affinity matrix material held in some container, andbatch elution from a plate are very similar andhereinafter will be treated under "affinitychromatography." 19
Fluorescent-activated cell sorting involves use ofan affinity material that is fluorescent per se or islabeled with a fluorescent molecule. Currentcommercially available cell sorters require 800 to 1000molecules of fluorescent dye, such as Texas red, boundto each cell. FACS can sort 103 cells or viruses/sec.
Electrophoretic affinity separation involveselectrophoresis of viruses or cells in the presence oftarget material, wherein the binding of said targetmaterial changes the net charge of the virus particlesor cells. It has been used to separate bacteriophageson the basis of charge. (SERW87).
The present invention makes use of affinityseparation of bacterial cells, or bacterial viruses (orother genetic packages) to enrich a population forthose cells or viruses carrying genes that code forproteins with desirable binding properties.
In the present invention, the words "select" and"selection" are used exclusively in the genetic sense;i.e. a biological process whereby a phenotypiccharacteristic is used to enrich a population for thoseorganisms displaying the desired phenotype.
The process of the present invention comprisesthree major parts: I. design and production of a replicablegenetic package (GP) that displays an IPBD onthe surface of the GP, denoted GP(IPBD), II. design and implementation of an affinityseparation process that separates GP(IPBD)s 20 that bind to a known affinity molecule fromwild-type GPs or GP(IPBD“)s, neither of whichbinds the known affinity molecule, and III. design and implementation of a geneticvariegation method, denoted structure-directed mutagenesis, wherein a population of106 or more different GP(PBD)s, denotedGP(vgPBD), is produced. 10 15 20 25 30
One affinity separation is called a "separation cycle";one pass of variegation followed by as many separationcycles as are needed to isolate an SBD, is called a"variegation cycle". The amino acid sequence of onebecomes the PPBD to the nextWe perform variegation cycles iteratively until the desired affinity and specificityof binding between an SBD and chosen target areachieved. SBD from one roundvariegation cycle.
Part I is a strain construction in which we dealwith a single IPBD sequence. Variability may beintroduced into DNA subsequences adjacent to the ipbdsubsequence and within the osp-ipbd gene so that theIPBD will appear on the GP surface. A molecule, suchas an antibody, having high affinity for correctlyfolded IPBD is used to: a) detect IPBD on the GPsurface, b) screen colonies for display of IPBD on theGP surface, or c) select GPs that display IPBD from apopulation, some members of which might display IPBD onthe GP surface. In one preferred embodiment, Part I ofthe process involves: 1) choosing a GP such as a bacterial cell (Sec. 1.1.1), bacterial spore (1.2.1), or phage (1.3.1), 35 21 having a suitable outer surface protein (Secs. 1.1.3, 1.2.3, and 1.3.3), 2) choosing a stable IPBD (Sec. 2), 3) designing an amino acid sequence that: a)includes the IPBD as a subsequence and b) willcause the IPBD to appear on the GP surface (Secs.1.1.2, 1.2.2, 1.3.2, and 4), 4) engineering a gene, denoted osp-ipbd. that: a)codes for the designed animo acid sequence, b)provides the necessary genetic regulation, and c)introduces convenient sites for geneticmanipulation (Secs. 4.1, 4.2, 4.3, 5.1, and 5.2), 5) cloning the osp-ipbd gene into the GP (Sec. 6.1), and 6) harvesting the transformed GPs (Sec. 7) andtesting them for presence of IPBD on the GPsurface (Sec. 8) ; this test is performed with anaffinity molecule having high affinity for IPBD,denoted AfM(IPBD).
In another preferred embodiment, Part I of the processinvolves: 1) and 2) as above 3) designing a DNA sequence that: a) encodes theIPBD as a subsequence and b) contains suitablerestriction sites so that random DNA may beoperably linked to the ipbd gene fragment; and c)provides the necessary genetic regulations; this 22 DNA sequence is called a "display probe", (Secs. 1.1.4, 1.2.4, 1.3.4 and 4), 4) constructing that display probe, 5) cloning the display probe into and amplifyingit in a suitable host into the OCV, 6) cloning random or pseudorandom DNA into one ofthe restriction sites provided in the displayprobe, (Sec. 6.2), whereby the random orpseudorandom DNA functions as a potential osp. and 7) harvesting GPs (Sec. 7) screening colonies ofthe transformed GPs for presence of IPBD on the GPsurface; this screening is performed with anaffinity molecule having high affinity for IPBD,denoted AfM(IPBD), (Sec. 8); or, alternatively; 8) selecting GPs that display IPBD by use of anaffinity separation using AfM(IPBD), (Sec. 8).
Once a GP(IPBD) is produced, it can be used manytimes as the starting point for developing differentnovel proteins that bind to a variety of differenttargets. The knowledge of how we engineer theappearance of one IPBD on the surface of a GP can beused to design and produce other GP(IPBD)s that displaydifferent IPBDs.
Although Part I deals with only a single IPBD,many preparations are made for Part III where weintroduce numerous mutations into the potential bindingdomain. References to PBD or pbd in Part I are toindicate a preparatory intent. 23
In Part II we optimize separation of GP(IPBD) fromwild-type GP, denoted wtGP, based on the affinity ofIPBD for AfM(IPBD) and establish the sensitivity of theaffinity separation process. In a preferredembodiment, Part II of the process of the presentinvention involves: 1) preparing affinity columns bearing AfM(IPBD) atvarious densities of AfM(IPBD)/(volume of matrix),(Sec. 10.1), 2) preparing GP(IPBD)s with various amounts of IPBD per GP, 3) picking a gradient regime for eluting thecolumns (Sec. 10.1), 4) determining which combination of: a) IPBD/GP, b) density of AfM(IPBD)/(volume of support), c)initial ionic strength, d) elution rate, and e)(amount of GP)/(volume of support) loaded, givesthe best separation of GP(IPBD) from wtGP (Sec. 10.1), 5) determining the smallest amount of GP(IPBD)that can be isolated from a much larger amount ofwtGP using the optimal condition, (Sec. 10.2), and 6) determining the efficiency of the affinityseparation procedure (Sec. 10.3).
Part II optimizes separation of a single type ofGP(IPBD) from a large excess of a single different GP.The optimum conditions will be used in Part III to 24 separate GP(PBD)s that bind the target from GP(PBD)sthat do not bind the target. The optimization will beat one or more specific temperatures and at one or morespecific pHs. In Part III, the user must specify the 5 conditions under which the selected SBD should bind thetarget. If the conditions of intended use differmarkedly from the conditions for which affinityseparation was optimized, the user must return to PartII and optimize the affinity separation for conditions 10 similar to the conditions of intended use of theselected SBD.
In Part III, we choose a target material and aGP(IPBD) that was developed by the method of Part I and 15 that is suitable to the target material. Using IPBD asthe PPBD to the first cycle of variegation, we preparea wide variety of osp-pbd genes that encode a widevariety of PBDs. We use an affinity separation,developed by the method of Part II, to enrich the 20 population of GP(vgPBD)s for GPs that display PBDs withbinding properties relative to the target that aresuperior to the binding properties of the PPBD. An SBDselected from one variegation cycle becomes the PPBD tothe next variegation cycle. In a preferred embodiment, 25 Part III of the process of the present inventioninvolves: 1) picking a target molecule (Sec. 11), 30 2) picking a GP(IPBD) (Sec. 12), 3) picking a set of several residues in the PPBDto vary based on a) the 3D structure of the IPBD,b) sequences of homologous proteins, and c)computer or theoretical modeling that indicates 35 25 which residues can tolerate different amino acidswithout disrupting the. underlying structure (Sec. 13.1), 4) picking a subset of the residues to be variedsimultaneously based on the number of differentvariants and which variants are within thedetection capabilities of the affinity separation;(Sec. 13.2); 5) implementing the variegation by: a) synthesizing the part of the osp-pbd genethat encodes the residues to be varied using aspecific mixture of nucleotide substrates forsome or all of the bases encoding residuesslated for variation, thereby creating apopulation of DNA molecules, denoted vgDNA(Sec. 13.3), b) ligating this vgDNA, by standard methods,into the operative cloning vector (OCV) (e.g.a plasmid or bacteriophage) (Sec. 14.1), c) using the ligated DNA to transform cells,thereby producing a population of transformedcells (Sec. 14.2), d) culturing (i.e. increasing in number) thepopulation of transformed cells and harvestingthe population of GP(PBD)s, said populationbeing denoted as GP(vgPBD), (Sec. 14.3), e) enriching the population for GPs that bindthe target by using the affinity separation 26 process developed in Part II, with the chosentarget molecule as affinity molecule (Sec. 15), f) repeating steps III.5.d and III.5.e until aGP(SBD) having improved binding to the targetis isolated (Sec. 15), and g) testing the isolated SBD or SBDs foraffinity and specificity for the chosen target(Sec. 15.8), 6) repeating steps III.3, III.4, and III.5 untilthe desired degree of binding is obtained.
Part III is repeated for each new target material.Part I need be repeated only if no GP(IPBD) suitable toa chosen target is available. Part II need be repeatedfor each newly-developed GP(IPBD) and for previously-developed GP(IPBD)s if the intended conditions of useof a novel binding protein differ significantly fromthe conditions of previous optimizations.
Sec. 0.2: Abbreviations:
The following abbreviations will be usedthroughout the present invention:
Abbreviation Meaning
GP
Genetic Package, e.g. abacteriophage
Any protein
The gene for protein X 27 IPBD Initial Potential Binding Domain, e.q. BPTI PBD Potential Bindinq Domain, e.q. a derivative of BPTI SBD Successful Binding Domain, e.q. a derivative of BPTI selected for binding to a target PPBD Parental Potential Binding Domain, i.e. an IPBD or an SBD from a previous selection OSP Outer Surface Protein, e.q. coat protein of a phage or LamB from E. coli OSP-PBD Fusion of an OSP and a PBD, order of fusion not specified OSTS Outer Surface Transport Signal GP(x) A genetic package containing the x gene GP(X) A genetic package thatdisplays X on its outer surface <Q) An affinity matrix supporting"O", e.q. {T4 lysozyme} is T4 28
lysozyme attached to an affinity matrix AfM(W) A molecule having affinity for "W" . e.cr. trvpsin is an AfM(BPTI) XINDUCE A chemical that can induce expression of a crene, e.cr. IPTG for the lacUV5 promoter OCV Operative Cloning Vector KT Kt = [T][SBD]/[T:SBD] (T is atarget) KN KN = [N][SBD]/[N:SBD] (N is anon-target) DoAMoM Density of AfM(W) on affinity matrix Abun(x) Abundance of DNA molecules encoding amino acid x OMP Outer membrane protein nt nucleotide Kd A bimolecular dissociation constant, Kj = [A][B]/[A:B] serr Error level in synthesizing vgDNA 29
Sec. 0.3: Standard sequencing method:
The present invention is not limited to a singlemethod of determining the sequence of nucleotides (nts)in DNA subsequences. Sequencing reactions, agarose gelelectrophoresis, and polyacrylamide gel electrophoresis(PAGE) are performed by standard procedures (AUSU87).
The present invention is not limited to a singlemethod of determining protein sequences, and referencein the appended claims to determining the amino acidsequence of a domain is intended to include anypractical method or combination of methods, whetherdirect or indirect. The preferred method, in mostcases, is to determine the sequence of the DNA thatencodes the protein and then to infer the amino acidsequence. In some cases, standard methods of protein-sequence determination may be needed to detect post-translational processing. ——— *** ---
The major steps in the process of making andisolating a novel binding protein with affinity for achosen target material are illustrated in Figure 2.
Sec. 1: Specification of Genetic Package and Means for
Displaying a Heterologous Binding Domain On Its Outer
Surface:
Sec. 1.0: General Reguirements for Genetic Packages
It is emphasized that the GP on which selection-through-binding will be practiced must be capable, 30 after the selection, either of growth in some suitableenvironment or of in vitro amplification and recoveryof the encapsulated genetic message. During at leastpart of the growth, the increase in number must be . 5 approximately exponential with respect to time. Thecomponent of a population that exhibits the desiredbinding properties may be quite small, for example, onein 104 * 6 or less. Once this component of the populationis separated from the non-binding components, it must 10 be possible to amplify it. Culturing viable cells isthe most powerful amplification of genetic materialknown and is preferred. Genetic messages can also beamplified in vitro, but this is not preferred. 15 A GP may typically be a vegetative bacterial cell, a bacterial spore or a bacterial DNA virus. A strainof any living cell or virus is potentially useful ifthe strain can be: 20 1) maintained in culture, 2) affinity separated and retain its viability, 3) genetically altered with reasonable facility, 25 and 4) manipulated to display the potential bindingprotein domain where it can interact with thetarget material during affinity separation. 30 DNA encoding the IPBD sequence may be operablylinked to DNA encoding at least the outer surfacetransport signal of an outer surface protein (OSP)native to the GP so that the IPBD is displayed on theouter surface of the GP. It should be possible to 35 31 cause a genetic package to display the IPBD or PBD onits outer surface without adversely affecting theviability of the GP or the binding characteristics ofthe IPBD or PBD, if the fusion is near domain 5 boundaries (BECK83, CRAW87, TOTH86, SMIT85, MANO86; and cf. ROSS81, • HOLL83). Those characteristics of a protein that are 10 recognized by a cell and that cause it to be transported out of the cytoplasm and displayed on the cell surface will be termed "outer-surface transportsignals". 15 The replicable genetic entity (phage or plasmid) that carries the osp-pbd genes (derived from the osp-ipbd gene) through the selection-through-bindingprocess, see Sec. 14, is referred to hereinafter as theoperative cloning vector (OCV). When the OCV is a 2 0 phage, it may also serve as the genetic package. Thechoice of a GP is dependent in part on the availabilityof a suitable OCV and suitable OSP.
Preferably, the GP is readily stored, for example, 25 by freezing. If the GP is a cell, it should have ashort doubling time, such as 20-40 minutes. If the GPis a virus, it should be prolific, e.g., a burst sizeof at least 100/infected cell. GPs which are finickyor expensive to culture are disfavored. The GP should 30 be easy to harvest, preferably by centrifugation. TheGP is preferably stable for a temperature range of -70to 42°C (stable at 4°C for several days or weeks) ;resistant to shear forces found in HPLC; insensitive toUV; tolerant of desiccation; and resistant to a pH of 35 2.0 to 10.0, surface active agents such as SDS or .7>, 32
Triton, chaotropes such as 4M urea or 2M guanidiniumHCI, common ions such as K+, Na+, and SO4 , commonorganic solvents such as ether and acetone, anddegradative enzymes. Finally, there must be a suitableOCV (see Sec. 3) .
Preferably, the 3 D structure of the OSP, and thesequence of the OSP gene p. 47 are known. If the 3Dstructure is not known, there is preferably knowledgeof which residues are exposed on the cell surface, thelocation of the domain boundaries within the OSP,and/or of successful fusions of the OSP and a foreigninsert. The OSP preferably appears in numerous copieson the outer surface of the GP, and preferably serves anon-essential function. It is desirable that the OSPnot be post translationally processed, or at least thatthis processing be understood.
The preferred GP, OCV and OSP are those for whichthe fewest serious obstacles can be seen, rather thanthe one that scores highest on any one criterion.
Next, we consider general answers to the questionsposed in this step for the cases of: a) vegetativelygrowing bacterial cells (Sec. 1.1), b) bacterial spores(Sec. 1.2), and c) (Sec. 1.3). Preferred OSPs forseveral GPs are given in Table 2.
Sec. 1.1: Bacterial Cells as Genetic Packages:
One may choose any well-characterized bacterialstrain which may be grown in culture. The importantquestions in this case are: a) do we know enough aboutmechanisms that localize proteins on the outside of thecell, b) will the IPBD fold in the environment of the 33 outer membrane, and c) will cells change expression ofosp-pbd. derived from osp-ipbd. during affinityseparation? Some IPBDs may need large or insolubleprosthetic groups, such as an Fe4S4 cluster, that areavailable within the cell, but not in the medium. Theformation of Fe4S4 clusters found in some ferrodoxinsis catalyzed by enzymes found in the cell (BONO85).IPBDs that require such prosthetic groups may fail tofold or function if displayed on bacterial cells.
Sec. 1.1.1: Preferred Bacterial Cells as GP :
In view of the extensive knowledge of E. coli. astrain of E. coli. defective in recombination, is thestrongest candidate as a bacterial GP. Other preferredcandidates are Salmonella typhimurium. Bacillussubtilis. and Pseudomonas aeruginosa.
Sec. 1.1.2: Preferred Outer Surface Proteins for
Displaying IPBDs on Bacterial Cells:
Gram-negative bacteria have outer-membraneproteins (OMP), that form a subset of OSPs. Many OMPsspan the membrane one or more times. The signals thatcause OMPs to localize in the outer membrane areencoded in the amino acid sequence of the matureprotein. Fusions of fragments of omp genes withfragments of an x gene have led to X appearing on theouter membrane (BENS84, CLEM81). If no fusion data areavailable, then we fuse an ipbd fragment to variousfragments of the osp gene and obtain GPs that displaythe osp-ipbd fusion on the cell outer surface byscreening or selection for the display-of-IPBDphenotype. 34
Oliver has reviewed mechanisms of proteinsecretion in bacteria (OLIV85 and OLIV87). Nikaido andVaara (NIKA87) have reviewed mechanisms by whichproteins become localized to the outer membrane ofGram-negative bacteria. For example, the LamB proteinof E. coli is synthesized with a typical signal-sequence which is subsequently removed. Benson et al.(BENS84) showed that LamB-LacZ fusion proteins would bedeposited in the outer membrane of E. coli whenresidues 1-49 of the mature LamB protein are includedin the fusion, but that residues 1-43 are insufficient.
LamB of E. coli is a porin for maltose andmaltodextrin transport, and serves as the receptor foradsorption of bacteriophages lambda and K10. Thisprotein has been purified to homogeneity (ENDE78) andshown to function as a trimer (PALV79). Mutations tophage resistance have been used to define the parts ofthe LamB protein that adsorb each phage (ROAM80,CLEM81, CLEM83, GEHR87).
Topological models have been developed thatdescribe the function of phage receptor andmaltodextrin transport. The models describe thesedomains and their locations with respect to thesurfaces of the outer membrane (CLEM81, CLEM83, CHAR84,HEIN88).
LamB is transported to the outer membrane if afunctional N-terminal sequence is present; further, thefirst 49 amino acids of the mature sequence arerequired for successful transport (BENS84). Homologybetween parts of LamB protein and other outer membraneproteins OmpC, OmpF and PhoE has been detected(NIKA84), including homology between LamB amino acids 35 39-49 and sequences of the other proteins. Thesesubsequences may label the proteins for transport tothe outer membrane. Further, monoclonal antibodiesderived from mice immunized with purified LamB, havebeen used to characterize four distinct topological andfunctional regions, two of which are concerned withmaltose transport (GABA82).
Sec. 1.1.3 Choice of Insertion site for IPBD in
Bacterial Cell OSP:
For fusions of the phoA into the coding sequencefor an integral membrane protein, the PhoA domain islocalized according to where in the integral membraneprotein the phoA gene was inserted (BECK83 and MANO86)That is, if phoA is inserted after an amino acid whichnormally is found in the cytoplasm, then PhoA appearsin the cytoplasm. If phoA is inserted after an aminoacid normally found in the periplasm, however, then thePhoA domain is localized on the periplasmic side of themembrane, and anchored in it. Beckwith and colleagues(BECK88) have extended these observations to the lacZ . . A. gene that can be inserted into genes for integralmembrane proteins such that the LacZ domain appears ineither the cytoplasm or the periplasm according towhere the lacZ gene was inserted. OSP-IPBD fusion proteins need not fill astructural role in the outer membranes of Gram-negativebacteria because parts of the outer membranes are nothighly ordered. For large OSPs there is likely to beone or more sites at which osp can be truncated andfused to ipbd such that cells expressing the fusionwill display IPBDs on the cell surface. If fusionsbetween fragments of osp and x have been shown to 36 display X on the cell surface, we can design an osp-ipbd gene by substituting ipbd for x in the DNAsequence. Otherwise, successful OMP-IPBD fusion ispreferably sought by fusing fragments of the best omp 5 to an ipbd. expressing the fused gene, and testing theresultant GPs for display-of-IPBD phenotype. We usethe available data about OMP to pick , the point orpoints of fusion between omp and ipbd to maximize thelikelihood that IPBD will be displayed. Alternatively, 10 we truncate osp at several sites or in a manner thatproduces osp fragments of variable length and fuse theosp fragments to ipbd; cells expressing the fusion arescreened or selected which display IPBDs on the cellsurface. An additional alternative is to include short 15 segments of random DNA in the fusion of omp fragmentsto ipbd and then screen or select the resultingvariegated population for members exhibiting thedisplay-of-IPBD phenotype. 20 The promoter for the osp-ipbd gene, preferably, is subject to regulation by a small chemical inducer, suchas isopropyl thiogalactoside (IPTG) (lac UV5 promoter).It need not come from a natural osp gene; anyregulatable bacterial promoter can be used (MANI82). 25
Once a genetic packaging system employingvegetative bacterial cells has been designed, it istime to choose an IPBD (Sec. 2). 30 Sec. 1.1.4: In Vivo Selection for Pseudo-osp Gene From
Random DNA Inserts in Bacterial Cells:
As an alternative to choosing a natural OSP and aninsertion site in the OSP, we can construct a gene 35 comprising: a) a regulatable promoter (e.g. lacUV5), b) 37 a Shine-Dalgarno sequence, c) a periplasmic transportsignal sequence, d) a fusion of the ipbd gene with asegment of random DNA (as in Kaiser et al. (KAIS87)),e) a stop codon, and f) a transcriptional terminator.The random DNA, which preferably comprises 90-300bases, encode numerous potential OSTS. (EF. KAIS87)The fusion of ipbd and the random DNA could be ineither order, but ipbd upstream is slightly preferred.Isolates from the population generated in this way canbe screened for display of the IPBD. Preferably, aversion of selection-through-binding is used to selectGPs that display IPBD on the GP surface, and thuscontain a DNA insert encoding a functional OSTS.Alternatively, clonal isolates of GPs may be screenedfor the display-of-IPBD phenotype.
The preference for ipbd upstream of the random DNAarises from consideration of the manner in which thesuccessful GP(IPBD) will be used. In Part III, we willintroduce numerous mutations into the pbd region of theosp-pbd gene, some of which might include gratuitousstop codons. If pbd precedes the random DNA, thengratuitous stop codons in pbd lead to no OSP-PBDprotein appearing on the cell surface. If pbd followsthe random DNA, then gratuitous stop codons in pbdmight lead to incomplete OSP-PBD proteins appearing onthe cell surface. Incomplete proteins often are non-specifically sticky so that GPs displaying incompletePBDs are easily removed from the population.
Sec. 1.2: Displaying IPBD on bacterial spores:
Bacterial spores have desirable properties as GPcandidates. Bacillus spores neither activelymetabolize nor alter the proteins on their surface. 38
However, spores are much more resistant than vegetativebacterial cells or phage to chemical and physicalagents. Spores have the disadvantage that themolecular mechanisms that trigger sporulation are lesswell worked out than is the formation of M13 or theexport of protein to the outer membrane of E. coli.
Sec. 1.2.1.: Preferred Bacterial Spores for Use as GPs:
Bacteria of the genus Bacillus form endosporesthat are extremely resistant to damage by heat,radiation, desiccation, and toxic chemicals (reviewedby Losick et al. (LOSI86)). These spores have complexstructure and morphogenesis that is species-specificand only partially elucidated. The followingobservations are relevant to the use of Bacillus sporesas genetic packages.
Plasmid DNA is commonly included in spores.Plasmid encoded proteins have been observed on thesurface of Bacillus spores (DEBR86). Sporulationinvolves complex temporal regulation that is moderatelywell understood (LOSI86). The sequences of severalsporulation promoters are known; coding sequencesoperatively linked to such promoters are expressed onlyduring sporulation (RAYC87).
Donovan et al. have identified several polypeptidecomponents of B. subtilis spore coat (DONO87); thesequences of two complete coat proteins and amino-terminal fragments of two others have been determined.Some components of the spore are synthesized in theforespore, e.g. small acid-soluble spore proteins(ERRI88), while other components are synthesized in themother cell and appear in the spore (e.g. the coat 39 proteins). This spatial organization of synthesis iscontrolled at the transcriptional level.
Spores self-assemble, but the signals that causevarious proteins to localize in different parts of thespore are not well understood; presumably, the signalscontrolling deposition of the coat proteins from thecytoplasm of the mother cell onto the spore coat areembedded in the polypeptide sequence. Some, but notall, of the coat proteins are synthesized as precursorsand are then processed by specific proteases beforedeposition in the spore coat (DONO87). Viable sporesthat differ only slightly from wild-type are producedin B. subtilis even if any one of four coat proteins ismissing (DONO87). Disulfide bonds form within thespore (thiol reducing agents are needed to solubilizeseveral of the proteins of the coat) . The 12kd coatprotein, CotD, contains 5 cysteines. CotD alsocontains an unusually high number of histidines (16)and prolines (7). The llkd coat protein, CotC,contains only one cysteine and one methionine. CotChas a very unusual amino-acid sequence with 19 lysines(K) appearing as 9 K-K dipeptides and one isolated K.There are also 20 tyrosines (Y) of which 10 appear as 5Y-Y dipeptides. Peptides rich in Y and K are known tobecome crosslinked in oxidizing environments (DEVO78,WAIT83, WAIT86). CotC contains 16 D and E amino acidsthat nearly equals the 19 Ks. There are no A, F, R, I,L, N, P, Q, S, or W amino acids in CotC. Neither CotCnor CotD is post-translationally cleaved. The proteinsCotA and CotB are post-translationally cleaved.
Endospores from the genus Bacillus are more stablethan are exospores from Streptomvces. Bacillus subtilis forms spores in 4 to 6 hours, but Streptomvces 40 species may require days or weeks to sporulate. Inaddition, genetic knowledge and manipulation is muchmore developed for B. subtilis than for other spore-forming bacteria. Thus Bacillus spores are preferredover Streptomvces spores. Bacteria of the genusClostridium also form very durable endospores, butClostridia, being strict anaerobes, are not convenientto culture. The choice of a species of Bacillus isgoverned by knowledge and availability of cloningsystems and by how easily sporulation can becontrolled. A particular strain is chosen by thecriteria listed in Sec. 1.0. Many vegetativebiochemical pathways are shut down when sporulationbegins so that prosthetic groups might not beavailable.
Sec. 1.2.2_Preferred outer-surface proteins for
Displaying IPBD on Bacterial Spores:
If a spore is chosen as GP, the promoter is themost important part of the osp gene, because thepromoter of a spore coat protein is most active: a)when spore coat protein is being synthesized anddeposited onto the spore and b) in the specific placethat spore coat proteins are being made. In B.subtilis. some of the spore coat proteins are post-translationally processed by specific proteases. It isvaluable to know the sequences of precursors and maturecoat proteins so that we can avoid incorporating therecognition sequence of the specific protease into ourconstruction of an OSP-IPBD fusion. The sequence of amature spore coat protein contains information thatcauses the protein to be deposited in the spore coat;thus gene fusions that include some or all of a mature 41 coat protein sequence are preferred for screening orselection for the display-of-IPBD phenotype.
Fusions of ipbd fragments to cote or cotDfragments are likely to cause IPBD to appear on thespore surface. The genes cote and cotD are preferredosp genes because CotC and CotD are not post-translationally cleaved. Subsequences from cotA orcotB could also be used to cause an IPBD to appear onthe surface of B. subtilis spores, but we must take thepost-translational cleavage of these proteins intoaccount. DNA encoding IPBD could be fused to afragment of cotA or cotB at either end of the codingregion or at sites interior to the coding region.Spores could then be screened or selected for thedisplay-of-IPBD phenotype.
To date, no Bacillus sporulation promoter has beenshown to be inducible by an exogenous chemical induceras the lac promoter of E. coli. Nevertheless, thequantity of protein produced from a sporulationpromoter can be controlled by other factors, such asthe DNA sequence around the Shine-Dalgarno sequence orcodon usage.
Sec. 1.2.3: Choice of Insertion site for IPBD in OSP of Bacterial Spore:
The considerations governing insertion site in thespore OSP are the same as those given in Section 1.1.3.
Sec. 1.2.4: In Vivo Selection for Pseudo-osp Genes
From Random DNA Inserts in Bacterial Spores: 42
Although the considerations for spores are nearlyidentical to the considerations for vegetativebacterial cells (Sec. 1.1), the available informationon the mechanisms that cause proteins to appear onspores is meager so that use of the random-DNA approachbecomes a more attractive option.
We can use the approach described above at 1.1.4for attaching an IPBD to an E. coli cell, except that:a) a sporulation promoter is used, and b) noperiplasmic signal sequence should be present.
Sec. 1.3: Displaying IPBD on Outer Surface of Phages:
Sec. 1.3.1: Preferred Phages for Use as GPs:
Unlike bacterial cells and spores, choice of aphage depends strongly on knowledge of the 3D structureof an OSP and how it interacts with other proteins inthe capsid. The size of the phage genome and thepackaging mechanism are also important because thephage genome itself is the cloning vector. The osp-ipbd gene must be inserted into the phage genome;therefore: 1) the virion must be capable of accepting theinsertion or substitution of genetic material, and 2) the genome of the phage must be small enough toallow convenient manipulation.
Additional considerations in choosing phage are: 1)the morphogenetic pathway of the phage determines theenvironment in which the IPBD will have opportunity to 43 fold, 2) IPBDs containing essential disulfides may notfold within a cell, 3) IPBDs needing large or insolubleprosthetic groups may not fold if secreted because theprosthetic group is lacking,- and 4) when variegation is 5 introduced in Part III, multiple infections could generate hybrid GPs that carry the gene for one PBD buthave at least some copies of a different PBD on theirsurfaces; it is preferable to minimize this possibility. 10
Bacteriophages are excellent candidates for GPsbecause there is little or no enzymatic activityassociated with intact mature phage, and because thegenes are inactive outside a bacterial host, rendering 15 the mature phage particles metabolically inert. Thefilamentous phage M13 and bacteriophage PhiX174 are ofparticular interest.
Filamentous phage : 20
The entire life cycle of the filamentous phageM13, a common cloning and sequencing vector, is wellunderstood. M13 and fl are so closely related that weconsider the properties of each relevant to both 25 (RASC86); any differentiation is for historicalaccuracy. The genetic structure (the complete sequence(SCHA78), the identity and function of the ten genes,and the order of transcription and location of thepromoters) of M13 is well known as is the physical 30 structure of the virion (BANN81, BOEK80, CHAN79, ITOK79, KAPL78, KUHN85b, KUHN87, MAK080, MARV78,MESS78, OHKA81, RASC86, RUSS81, SCHA78, SMIT85, WEBS78,and ZIMM82) ; see RASC86 for a recent review of thestructure and function of the coat proteins. 35 44
Relevant facts about M13 are disclosed in Example I.
Bacteriophage PhiX174 : 5
The bacteriophage PhiX17 4 is a very smallicosahedral virus which has been thoroughly studied bygenetics, biochemistry, and electron microscopy (SeeThe Single-Stranded DNA Phages (DENH78)). To date, no 10 proteins from PhiX174 have been studied by X-raydiffraction. PhiX174 is not used as a cloning vectorbecause PhiX174 can accept almost no additional DNA;the virus is so tightly constrained that several of itsgenes overlap. Chambers et al. (CHAM82) showed that 15 mutants in gene G are rescued by the wild-type G genecarried on a plasmid so that the host supplies thisprotein.
Three gene products of PhiX174 are present on the 20 outside of the mature virion: F (capsid), G (majorspike protein, 60 copies per virion), and H (minorspike protein, 12 copies per virion). The G proteincomprises 175 amino acids, while H comprises 328 aminoacids. The F protein interacts with the single- 25 stranded DNA of the virus. The proteins F, G, and Hare translated from a single mRNA in the viral infectedcells.
Large DNA Phages 30
Phage such as lambda or T4 have much largergenomes than do M13 or PhiX174. Large genomes are lessconveniently manipulated than small genomes. A phagewith a large genome, however, could be used if genetic 35 manipulation is sufficiently convenient. Phage such as 45 lambda and T4 have more complicated 3D capsidstructures than M13 or PhiX174, with more OSPs tochoose from. Phage lambda virions and phage T4 virionsform intracellularly, so that IPBDs requiring large orinsoluble prosthetic groups might fold on the surfacesof these phage. Phage lambda and phage T4 are notpreferred, however, derivatives of these phages couldbe constructed to overcome these disadvantages. RNA Phages RNA phage, such as Qbeta, are not preferredbecause manipulation of RNA is much less convenientthan is the manipulation of DNA. Although competentRNA bacteriophage are not preferred, useful geneticallyaltered RNA-containing particles '’co.uld be derived fromRNA phage, such as MS2.
To use MS2 as a GP, we would need to eliminatemost of the natural viral genome so that an osp-ipbdgene could fit into the protein capsid. It is knownthat the A protein binds sequence-specifically to asite at the 5' end of the + RNA strand triggeringformation of RNA-containing particles if coat proteinis present. If a message containing the A proteinbinding site and the gene for a chimera of coat proteinand a PBD were produced in a cell that also contained Aprotein and wild-type coat protein (both produced fromregulated genes on a plasmid), then the RNA coding forthe chimeric protein would get packaged. A packagecomprising RNA encapsulated by proteins encoded by thatRNA satisfies the major criterion that the geneticmessage inside the package specifies something on theoutside. The particles by themselves are not viable. 46
After isolating the packages that carry an SBD, wewould need to: 1) separate the RNA from the protein capsid, 2) reverse transcribe the RNA into DNA, using AMVor MMTV reverse transcriptase, and 3) amplify the DNA by several cycles of polymerasechain reaction (PCR) until there is enough tosubclone the recovered genetic message into aplasmid for sequencing and further work.
Alternatively, helper phage could be used to rescue theisolated phage.
Sec. 1.3.2:_Preferred Outer-Surface Proteins for
Displaying IPBDs on Phages:
For a given bacteriophage, the preferred OSP isusually one that is present on the phage surface in thelargest number of copies, as this allows the greatestflexibility in varying the ratio of OSP-IPBD to wildtype OSP and also gives the highest likelihood ofobtaining satisfactory affinity separation. Moreover,a protein present in only one or a few copies usuallyperforms an essential function in morphogenesis orinfection; mutating such a protein by addition orinsertion is likely to result in reduction in viabilityof the GP.
It is preferred that the wild-type osp gene bepreserved. The ipbd gene fragment may be insertedeither into a second copy of the recipient osp gene orinto a novel engineered osp gene. The preferred OSP 47 for use when the GP is M13 is the gene III protein (seeExample 1).
Sec. 1.3.3: Choice of Insertion site for IPBD in OSP:
The user must choose a site in the candidate OSPgene for inserting a ipbd gene fragment. The coats ofmost bacteriophage are highly ordered. Thus in bacteriophage, unlike the cases of bacteria and spores,it is important to retain most or all of the residuesof the parental OSP in engineered OSP-IPBD fusionproteins. A preferred site for insertion of the ipbdgene into the phage osp gene is one in which: a) theIPBD folds into its original shape, b) the OSP domainsfold into their original shapes, and c) there is nointerference between the two domains.
If there is a 3D model of the phage that indicatesthat either the amino or carboxy terminus of an OSP isexposed to solvent, then the exposed terminus of thatmature OSP becomes the prime candidate for insertion ofthe ipbd gene. A low resolution 3D model suffices.
In the absence of a 3D structure, the amino andcarboxy termini of the mature OSP are the . bestcandidates for insertion of the ipbd gene. Afunctional fusion may require additional residuesbetween the IPBD and OSP domains to avoid unwantedinteractions between the domains. Random-sequence DNAor DNA coding for a specific sequence of a proteinhomologous to the IPBD or OSP, can be inserted betweenthe osp fragment and the ipbd fragment if needed.
Fusion at a domain boundary within the OSP is alsoa good approach for obtaining a functional fusion. « 'x 48
Smith exploited such a boundary when subcloningheterologous DNA into gene III of fl (SMIT85).
There are several methods of identifying domains.5 Methods that rely on atomic coordinates have beenreviewed by Janin and Chothia (JANI85) see also ROSE85, RASH84, VITA84, PABO79, POTE83, and SCOT87.
If the only structural information available is10 the amino acid sequence of the candidate OSP, we usethe sequence to predict turns and loops. There is ahigh probability that some of the loops and turns willbe correctly predicted (cf. Chou and Fasman, (CHOU72));these locations are also candidates for insertion of 15 the ipbd gene fragment.
Sec. 1.3.4: In Vivo Selection for Pseudo-OSP Gene from
Random DNA Inserts in Bacterial Spores: 20 Alternatively, a functional insertion site may be determined by generating a number of recombinantconstructions and selecting the functional strain byphenotypic characteristics. Because the OSP-IPBD mustfulfill a structural role in the phage coat, it is 25 unlikely that any particular random DNA sequencecoupled to the ipbd gene will produce a fusion proteinthat fits into the coat in a functional way.Nevertheless, random DNA inserted between largefragments of a coat protein gene and the ipbd gene will 3 0 produce a population that is likely to contain one ormore members that display the IPBD on the outside of aviable phage. A display probe, similar to that definedin 1.1.4, is constructed and random DNA sequencescloned into appropriate sites. 35 49
Sec, 2: Choice of IPBD :
An IPBD may be chosen from naturally occurringproteins or domains of naturally occurring proteins, or 5 may be designed from first principles. A designedprotein may have advantages over natural proteins if:a) the designed protein is more stable, b) the designedprotein is smaller, and c) the charge distribution ofthe designed protein can be specified more freely. ' 10 A candidate IPBD must meet the following criteria:1) stablility under the conditions of its intended use(the domain may comprise the entire protein that willbe inserted, e.g. BPTI), 2) knowledge of the amino acid 15 sequence is obtainable, 3) identification of theresidues on the outer surface, and their spatialrelationships, and 4) availability of a molecule,AfM(IPBD) having high specific affinity for the IPBD. 20 Preferably, the IPBD is no larger than necessary because it is easier to arrange restriction sites insmaller amino-acid sequences. The usefulness ofcandidate IPBDs that meet all of these requirementsdepends on the availability of the information 25 discussed below.
Information used to judge IPBD suitabilityincludes: 1) a 3D structure (knowledge strongly preferred) , 2) one or more sequences homologous to the 30 IPBD (the more homologous sequences known, the better), 3) the pi of the IPBD (knowledge necessary in somecases), 4) the stability and solubility as a functionof temperature, pH and ionic strength (preferably knownto be stable over a wide range and soluble in 35 conditions of intended use), 5) ability to bind metal 50 ions such as Ca++ or Mg++ (knowledge preferred; bindingper se. no preference), 6) enzymatic activities, if any(knowledge preferred, activity per se has uses but maycause problems), 7) binding properties, if any(knowledge preferred, specific binding also preferred), 8) availability of a molecule having specific andstrong affinity ( Kj < 10-11 M) for the IPBD(preferred), 9) availability of a molecule havingspecific and medium affinity ( 10“8 M < Kj < 10-6 M)for the IPBD (preferred) , 10) the sequence of a mutantof IPBD that does not bind to the affinity molecule(s)(preferred), and 11) absorption spectrum in visible,UV, NMR, etc. (characteristic absorption preferred).
If only one species of molecule having affinityfor IPBD (AfM(IPBD)) is available, it will be used to;a) detect the IPBD on the GP surface, b) optimizeexpression level and density of the affinity moleculeon the matrix (Sec. 10.1), and c) determine theefficiency and sensitivity of the affinity separation(Secs. 10.2 and 10.3). As noted above, however, onewould prefer to have available two species ofAfM(IPBD), one with high and one with moderate affinityfor the IPBD. The species with high affinity would beused in initial detection and in determining efficiencyand sensitivity (10.2 and 10.3), and the species withmoderate affinity would be used in optimization (10.1).
For at least 20 candidate IPBDs the aboveinformation is available or is practical to obtain, forexample, bovine pancreatic trypsin inhibitor (BPTI, 58residues), crambin (46 residues), third domain ofovomucoid (56 residues), T4 lysozyme (164 residues),and azurin (128 residues). 51
Most of the PBDs derived from a PPBD according tothe process of the present invention affect residueshaving side groups directed toward the solvent.Exposed residues can accept a wide range of aminoacids, while buried residues are more limited in thisregard (REID88). Surface mutations typically have onlysmall effects on melting temperature of the PBD, butmay reduce the stability of the PBD. Hence the chosenIPBD should have a high melting temperature (60°Cacceptable, the higher the better) and be stable over awide pH range (8.0 to 3.0 acceptable; 11.0 to 2.0preferred) , so that the SBDs derived from the chosenIPBD by mutation and selection-through-binding willretain sufficient stability. Preferably, thesubstitutions in the IPBD yielding the various PBDs donot reduce the melting point of the domain below 50°C.
Two general characteristics of the targetmolecule, size and charge, make certain classes ofIPBDs more likely than other classes to yieldderivatives that will bind specifically to the target.Because these are very general characteristics, one candivide all targets into six classes: a) large positive,b) large neutral, c) large negative, d) small positive,e) small neutral, and f) small negative. A smallcollection of IPBDs, one or a few corresponding to eachclass of target, will contain a preferred candidateIPBD for any chosen target.
Alternatively, the user may elect to engineer aGP(IPBD) for a particular target; Sec 2.1 givescriteria that relate target size and charge to thechoice of IPBD.
Sec. 2.1: Influence of target size on choice of IPBD: ·5?>, 52 . If the target is a protein or other macromoleculea preferred embodiment of the IPBD is a small proteinsuch as BPTI from Bos taurus (58 residues) , crambin 5 from rape seed (46 residues) , or the third domain ofovomucoid from Coturnix coturnix Japonica (Japanesequail) (56 residues) (PAPA82), because targets fromthis class have clefts and grooves that can accommodatesmall proteins in highly specific ways. If the target 10 is a macromolecule lacking a compact structure, such asstarch, it should be treated as if it were a smallmolecule. Extended macromolecules with defined 3Dstructure, such as collagen, should be treated as largemolecules. 15
If the target is a small molecule, such as asteroid, a preferred embodiment of the IPBD is aprotein the size of ribonuclease from Bos taurus (124residues), ribonuclease from Aspergillus orvzae (104 20 residues), hen egg white lysozyme from Gallus callus(129 residues), azurin from Pseudomonas aeruginosa (128residues), or T4 lysozyme (164 residues), because suchproteins have clefts and grooves into which the smalltarget molecules can fit. The Brookhaven Protein Data 25 Bank contains 3D structures for these proteins. Genesencoding proteins as large as T4 lysozyme can bemanipulated by standard techniques for the purposes ofthis invention. 30 If the target is a mineral, insoluble in water, one must consider the nature of the mineral's molecularsurface. Smooth surfaces, (such as crystallinesilicon) require medium to large proteins (such asribonuclease) as IPBD in order to have sufficient 35 contact area and specificity. Rough, grooved surfaces 53 (zeolites), could be bound either by small proteins(BPTI) or larger proteins (T4 lysozyme).
Sec. 2.2: Influence of target charge on choice of 5 IPBD:
Electrostatic repulsion between molecules of likecharge can prevent molecules with highly complementarysurfaces from binding. Therefore, it is preferred 10 that, under the conditions of intended use, the IPBDand the target molecule either have opposite charge orthat one of them is neutral. Inclusion of counter ionscan reduce or eliminate electrostatic repulsion. 15 Sec. 2.3: Other aspects of choice of IPBD:
If the chosen IPBD is an enzyme, it may benecessary to change one or more residues in the activesite to inactivate enzyme function. For example, if 20 the IPBD were T4 lysozyme and the GP were E. coli cellsor M13, we would inactivate the lysozyme lest it lysethe cells. If, on the other hand, the GP were PhiX174,then inactivation of lysozyme may not be needed becauseT4 lysozyme can be overproduced inside E. coli cells 25 without detrimental effects and PhiX174 formsintracellularly. It is preferred to inactivate enzymeIPBDs that might be harmful to the GP or its host bysubstituting mutant amino acids at one or more residuesof the active site. It is permitted to vary one or 30 more of the residues that were. changed to abolish theoriginal enzymatic activity of the IPBD. Those GPsthat receive osp-pbd genes encoding an active enzymemay die, but the majority of sequences will not bedeleterious. 35 54
Sec. 3: Choice of OCV:
The OCV is preferably small, e.g. , less than 10KB. It is desirable that cassette mutagenesis bepractical in the OCV; preferably, at least 25restriction enzymes are available that do not cut theOCV· It is likewise desirable that single-strandedmutagenesis be practical. Finally, the OCV preferablycarries a selectable marker. A suitable OCV isobtained or is engineered by manipulation of availablevectors. Plasmids are preferred over the bacterialchromosome because genes on plasmids are much moreeasily constructed and mutated than are chromosomalgenes. When bacteriophage are to be used, the osp-ipbdgene must be inserted into the phage genome.
For phage such as M13, an antibiotic resistancegene is engineered into the genome (HINE80) . Morevirulent phage, such as PhiX174, make discernableplaques that can be picked, in which case a resistancegene is not essential; furthermore, there is no room inthe PhiX174 virion to add any new genetic material.Inability to include an antibiotic resistance gene is adisadvantage because it limits the number of GPs thatcan be screened.
It is preferred that GP(IPBD) carry a selectablemarker not carried by wtGP. It is also preferred thatwtGP carry a selectable marker not carried by GP(IPBD).
Sec. 4: Designing the osp-ipbd gene insert:
We design an amino acid sequence that will causethe IPBD to appear on the GP surface when it is
<img img-format="tif" img-content="drawing" file="IL120940AD00025.tif" id="idf0005" />
55 expressed. This amino acid sequence may determine theentire coding region of the osp-ipbd gene, or it maycontain only the ipbd sequence adjoining restrictionsites into which random DNA will be cloned (Sec. 6.2). 5
The actual gene may be produced by any means. Thepbd segment, derived from the ipbd segment, must beeasily genetically manipulated in the ways described inPart III. Synthetic ipbd segments are preferred 10 because they allow greatest control over placement ofrestriction sites.
Sec. 4.1 Genetic regulation of the osp-ipbd gene: 15 Regarding regulation of the osp-ipbd gene, the two important questions are: a) how much OSP-IPBD do weneed on each GP, and b) how accurately must we regulatethe amount? z 20 The essential function of the affinity separation is to separate GPs that bear PBDs (derived from IPBD)having high affinity for the target from GPs bearingPBDs having low affinity for the target. If a gradientof some solute, such as increasing salt, changes the 25 conditions, then all weakly-binding PBDs will cease tobind before any strongly-binding PBDs cease to bind.Regulation of the osp-pbd gene must be such that allpackages display - sufficient PBD to effect a goodseparation in Sec 15. If the amount of PBD/GP had an 30 effect on the elution volume of the GP from theaffinity matrix, then we would need to regulate theamount of PBD/GP accurately. The following analysisshows that there is no strong linear effect of IPBD/GPon elution volume and assumes only: a) that all GPs are 35 the same size, b) that interactions between the PBDs 56 and the affinity matrix dominate differential elutionof GPs, c) that the system is at equilibrium, and d)that all PBDs on any one GP are identical.
If Np identical PBDs on a GP each have access totarget molecules, and each PBD has a free-energy ofbinding to the target of delta Gj-,, then the total freeenergy of binding is delta Gj^ot = Np * ^elta Gp .
Delta Gj-, is a function of parameters of the solvent,such as: 1) concentration of ions, 2) pH, 3)temperature, 4) concentration of neutral solutes suchas sucrose, glucose, ethanol, etc., 5) specific ions,such as, calcium, acetate, benzoate, nicotinate, etc.If conditions are altered during affinity separation sothat delta Gp approaches zero, delta Gptot approacheszero Np times faster. As delta goes to or abovezero, the packages will dissociate from the immobilizedtarget molecules and be eluted. GPs bearing more PBDs have a sharper transitionbetween bound and unbound than packages with fewer ofthe same PBDs. For equilibrium conditions, the mid-point of the transition is determined only by thesolution conditions that bring the individualinteractions to zero free-energy. The number ofPBDs/GP determines the sharpness of the transition.
It should also be noted that the number of PBDs/GPis usually influenced by physiological conditions sothat a sample of genetically identical GP(PBD)s maycontain GPs having different numbers of PBDs on the GPsurface. In a population of GP(vgPBD)s each PBD 57 sequence will appear on more that one GP, and theactual number of PBDs/GP will vary from GP to GP withinsome range. Within a variegated population of PBDs,let PBDX be the PBD with maximum affinity for thetarget. If there is a linear effect on elution volumeof number of PBDs/GP, then the GPs having the greatestnumber of PBDX will be most retarded on the column.When we culture the enriched population the GP(PBDX)will be amplified and give rise to new GP(PBDx)s havingvarying numbers of PBDX/GP. Thus the affinityseparation process of the present invention couldtolerate a linear effect of number of PBDs/GP on theelution volume of the GP(PBD) unless strong binding totarget fortuitously causes the PBD to be displayed, onthe GP only in low number.
Since there is no linear effect on elution volumefrom the number of IPBDs/GP, need for highly accurateregulation of IPBD/GP is not anticipated. Reproduciblegene expression is more easily controlled usingregulated rather than constitutive genetic elements.The analysis above assumes that GP(IPBD)s are inequilibrium between solution in buffer and bound to theaffinity matrix. Rate of elution may be an importantparameter in column affinity chromatography. In batchelution from an affinity matrix or elution from anaffinity plate, the time that each buffer is in contactwith the affinity material may be an importantThe density of affinity molecules on thean important variable in optimizing the affinity separation. Because the analysis above isqualitative, in Sec. 10 of the preferred embodiment weexperimentally optimize: 1) the density of IPBD on theGP surface, 2) the density of affinity molecules on theaffinity matrix, 3) the initial ionic strength, 4) the variable,matrix is 58 elution rate, and 5) the quantity of GP/(volume ofmatrix) to be loaded on the column.
Transcriptional regulation of gene expression isbest understood and most effective, so we focus ourattention on the promoter. A number of promoters areknown that can be controlled by specific chemicalsadded to the culture medium. For example, the lacUV5promoter is induced if isopropylthiogalactoside isadded to the culture medium, for example, at between1.0 uM and 10.0 mM. Hereinafter, we use "XINDUCE" as ageneric term for a chemical that induces expression ofa gene. If transcription of the osp-ipbd gene iscontrolled by XINDUCE, then the number of OSP-IPBDs perGP increases for increasing concentrations of XINDUCEuntil a fall-off in the number of viable packages isobserved or until sufficient IPBD is observed on thesurface of harvested GP(IPBD)s.
The attributes that affect the maximum number ofOSP-IPBDs per GP are primarily structural in nature.There may be steric hindrance or other unwantedinteractions between IPBDs if OSP-IPBD is substitutedfor every wild-type OSP. Excessive levels of OSP-IPBDmay also adversely affect the solubility ormorphogenesis of the GP. For cellular and viral GPs,as few as five copies of a protein having affinity foranother immobilized molecule have resulted insuccessful affinity separations (FERE82a, FERE82b, andSMIT85).
Another consideration of promoter regulation isthat it is useful later to know the range of regulationof the osp-ipbd. (Sec. 8) In particular, one shoulddetermine how nearly the absence of XINDUCE leads to 59 the absence of IPBD on the GP surface; a non-leakypromoter is preferred. Non-leakiness is useful: a) toshow that affinity of GP(osp-ipbd)s for AfM(IPBD) isdue to the osp-ipbd gene, and b) to allow growth ofGP(osp-pbd) in the absence of XINDUCE if the expressionof osp-pbd is disadvantageous. The lacUV5 promoter inconjunction with the Lacl1? repressor is a preferredexample.
Sec. 4.2: DNA sequence design:
The present invention is not limited to a singlemethod of gene design. The following procedure is anexample of one method of gene design that fills theneeds of the present invention.
If the amino-acid sequence of OSP-IPBD is adefinite sequence, then the entire gene will beconstructed (Sec. 6.1). If random DNA is to be fusedto ipbd, then a "display probe" is constructed first;the random DNA is then inserted to complete thepopulation of putative osp-ipbd genes (Sec. 6.2) fromwhich a functional osp-ipbd gene is identified by invivo selection or kindred techniques.
One may use any genetic engineering method toproduce the correct gene fusion, so long as one caneasily and accurately direct mutations to specificsites in the pbd DNA subsequence (Sec. 14.1). For themethods of mutagenesis considered here, however, theDNA sequence for the osp-ipbd gene must be differentfrom any other DNA in the OCV. The degree and natureof difference needed is determined by the method ofmutagenesis. One replaces subsequences coding for thePBD with vgDNA, then subsequences to be mutagenized 60 must be bounded by restriction sites that are uniquewithin the OCV. If single-stranded-oligonucleotide-directed mutagenesis is to be used, then the DNAsequence of the subsequence coding for the IPBD must beunique within the OCV.
Regulatory elements include: a) promoters, b)Shine-Dalgarno sequences, and c) transcriptionalterminators, and may be isolated from nature ordesigned from knowledge of consensus sequences ofnatural regulatory regions.
The coding portions of genes to be synthesized aredesigned at the protein level and then encoded in DNA.The amino acid sequences are chosen to achieve variousgoals, including: a) display of a IPBD on the surfaceof a GP, b) change of charge on a IPBD, and c)generation of a population of PBDs from which to selectan SBD. The ambiguity in the genetic code is exploitedto allow optimal placement of restriction sites and tocreate various distributions of amino acids atvariegated codons.
Sec. 4.3: Specific DNA sequence assignment: A computer program may be used to identify allpossible ambiguous DNA sequences coding for an amino-acid sequence given by the user and to identify placeswhere recognition sites for site-specific restrictionenzymes could be provided without altering the amino-acid sequence.
Restriction sites are positioned within the osp-ipbd gene so that the longest segment between sites isas short as possible. Enzymes the produce cohesive 61 ends are preferred. The codon preferences of theintended host and the secondary structure of themessenger RNA are also considered. 5 Sec. 5.1: Organization of gene synthesis;
An established strategy for gene synthesis is tosynthesize both strands of the entire gene inoverlapping segments of 20 to 50 nucleotides (nts) 10 (THER88). We prefer an alternative method that is moresuitable for synthesis of vgDNA. Our method differsfrom previous methods (OLIP86, OLIP87, AUSU87) in thatwe: a) use two synthetic strands, and b) do not cut theextended DNA in the middle. Our goals are: a) to 15 produce longer pieces of dsDNA than can be synthesizedas ssDNA on commercial DNA synthesizers, and b) toproduce strands complementary to single-stranded vgDNA.By using two synthetic strands, we remove therequirement for a palindromic sequence at the 3' end. 20 DNA synthesizers can produce oligo-nts of up to100 nts in reasonable yield, Mqna = 100 · The parameters Nw (the length of overlap needed to obtainefficient annealing) and Ns (the number of spacer bases 25 needed so that a restriction enzyme can cut near theend of blunt-ended dsDNA) are determined by DNA andenzyme chemistry. Nw = 10 and Ns = 5 are reasonablevalues. 30 We divide the DNA sequence to be synthesized into two nearly equal parts, each 5-8 bases longer than halfthe total length, so that there is an overlap betweenthe two parts of 10 to 16 bp (Nw) containing novariegated bases. The overlap preferably, is not 35 palindromic and has high GC content. We synthesize the Οι 62 overlap portion and the 5' extension of each strand.When these strands are annealed and completed withKlenow enzyme and all four NTPs, we obtain the desiredsequence as blunt-ended dsDNA. If the DNA is to beligated to other DNA having cohesive ends, five to ten(Ns) bases are added to that end. The synthetic dsDNAcan then be cut efficiently with an appropriaterestriction enzyme (OLIP87).
Because MDNA is not rigidly fixed at 100, thecurrent limits of 190 (= 2 MDNA - Nw) nts overall and100 in each fragment are not rigid, but can be exceededby 5 or 10 nts. Going beyond the limits of 190 and 100will lead to lower yields, but these may be acceptablein certain cases.
Sec. 5.2: DNA synthesis and purification methods :
The present invention is not limited to anyparticular method of DNA synthesis or construction.
In the preferred embodiment, DNA is synthesized bystandard means on a Milligen 7500 DNA synthesizer. TheMilligen 7500 has seven vials from whichphosphoramidites may be taken. Normally, the firstfour contain A, C, T, and G. The other three vialsmay contain unusual bases such as inosine or mixturesof bases, the so-called "dirty bottle". The standardsoftware allows programmed mixing of two, three, orfour bases in equimolar quantities.
The present invention is not limited to anyparticular method of purifying DNA for geneticengineering. Agarose gel electrophoresis andelectroelution on an IBI device (International 63
Biotechnologies, Inc., New Haven, CT) is, preferably,used to purify large dsDNA fragments. For oligo-nts,PAGE and electroelution with an Epigene device (EpigeneCorp., Baltimore, MD) are an alternative to HPLC.
Sec. 6.1: Cloning of Known OSP-ipbd gene into OCV:
In the preferred method, the synthetic gene isconstructed using plasmids that are transformed intobacterial cells by standard methods (MANI82, p250) orslightly modified standard methods. Alternatively, DNAfragments derived from nature are operably linked toother fragments of DNA derived from nature or tosynthetic DNA fragments. In most cases of thepreferred method, gene synthesis involves constructionof a series of plasmids containing larger and largersegments of the complete gene.
Sec. 6.2 Cloning of Random DNA (Potential osp) Into
Display Probe:
If random DNA and phenotypic selection orscreening are used to obtain a GP(IPBD), then we clonerandom DNA into one of the restriction sites that wasdesigned into the display probe.
The random DNA may be obtained in a variety ofways. Degenerate synthetic DNA is one possibility.Alternatively, pseudorandom DNA may be taken fromnature. If, for example, an Sph I site (GCATG/C) hasbeen designed into the display probe at one end of theipbd fragment, then we would use Nla III (CATG/) topartially digest DNA that contains a wide variety ofsequences, generating a wide variety of fragments withCATG 3' overhangs. Preferably, the display probe has 64 different restriction sites at each end of the ipbdgene so that random DNA can be cloned at either end. A plasmid carrying the display probe is digestedwith the appropriate restriction enzyme and thefragmented, random DNA is annealed and ligated bystandard methods. The ligated plasmids are used totransform cells that are grown and selected forexpression of the antibiotic-resistance gene. Plasmid-bearing GPs are then selected for the display-of-IPBDphenotype by the procedure given in Sec. 15 of thepresent invention using AfM(IPBD) as if it were thetarget. Sec. 15 is designed to isolate GP(PBD)s thatbind to a target from a large population that do notbind.
Sec. 7: Harvest of GPs :
Cells are transformed with ligated OCVs andselected for uptake of OCV after an appropriateincubation with an agent appropriate to the selectablemarkers on the OCV. GPs are harvested by methodsappropriate to the GP at hand, generally,centrifugation to pelletize GPs and resuspension of thepellets in sterile medium (cells) or buffer (spores orphage).
Sec. 8: Verification of Display Strategy:
The harvested packages are now tested for displayof IPBD on the surface; any ions or cofactors known tobe essential for the stability of IPBD or AfM(IPBD)must be included at appropriate levels. The tests canbe done: a) by affinity labeling, b) enzymatically, c)spectrophotometrically, d) by affinity separation, or 65 e) by affinity precipitation. The AfM(IPBD) in thisstep is one picked to have strong affinity(preferably, < 10”11 M) for the IPBD molecule andlittle or no affinity for the wtGP. For example, ifBPTI were the IPBD, trypsin, anhydrotrypsin, orantibodies to BPTI could be used as the AfM(BPTI) totest for the presence of BPTI. Anhydrotrypsin, atrypsin derivative with serine 195 converted todehydroalanine, has no proteolytic activity but retainsits affinity for BPTI (AKOH72 and HUBE77).
Preferably, the presence of the IPBD on thesurface of the GP is demonstrated through the use of asoluble, labeled derivative of a AfM(IPBD) with highaffinity for IPBD. The labeled derivative of AfM(IPBD)is denoted as AfM(IPBD)*.
If random DNA has been used, then the proceduresof Sec. 15 are used to obtain a clonal isolate that hasthe display-of-IPBD phenotype. Alternatively, clonalisolates may be screened for the display-of-IPBDphenotype. The tests of this step are applied to oneor more of these clonal isolates.
If no isolates that bind to the affinity moleculeare obtained we take corrective action as disclosed inSec. 9.
If one or more of the tests indicates that theIPBD is displayed on the GP surface, we verify that thebinding of molecules having known affinity for IPBD isdue to the chimeric osp-ipbd gene through the use ofstandard genetic and biochemical techniques, such as: 66 1) transferring the osp-ipbd gene into the parent GP to verify that osp-ipbd confers binding, 2) deleting the osp-ipbd gene from the isolated GPto verify that loss of osp-ipbd causes loss ofbinding, 3) showing that binding of GPs to AfM(IPBD)correlates with [XINDUCE] (in those cases thatexpression of osp-ipbd is controlled by[XINDUCE]), and 4) showing that binding of GPs to AfM(IPBD) isspecific to the immobilized AfM(IPBD) and not tothe support matrix.
Presence of IPBD on the GP surface is indicated bya strong correlation between [XINDUCE] and thereactions that are linear in the amount of IPBD (suchas; a) binding of GPs by soluble AfM(IPBD)*, b)absorption caused by IPBD, and c) biochemical reactionsof IPBD) . The demonstration (4) that binding is toAfM(IPBD) and the genetic tests (1) and (2) areimportant; the test with XINDUCE (3) is less so.
We sequence the relevant ipbd gene fragment fromeach of several clonal isolates to determine theconstruction.
We establish the maximum salt concentration and pHrange for which the GP(IPBD) binds the chosenAfM(IPBD). 67
If the IPBD is displayed on the outside of the GP,and if that display is clearly caused by the introducedosp-ipbd gene, we proceed to Part II, otherwise we mustanalyze the result and adopt appropriate corrective 5 measures.
Sec. 9: Perfecting the Display System:
If we have attempted to fuse an ipbd fragment to a10 natural osp fragment, our options are : 1) pick a different fusion to the same osp by a) using opposite end of osp, b) keeping more or fewer residues from osp in 15 the fusion; for example, in increments of 3 or 4 residues, c) trying a known or predicted domainboundary, d) trying a predicted loop or turn position, 20 2) pick a different osp. or 3) switch to random DNA method. 25 If we have just tried the random DNA method unsuccessfully, our options are :
1) choose a different relationship between ipbdfragment and random DNA (ipbd first, random DNA 30 second or vice versa), 2) try a different degree of partial digestion, adifferent enzyme for partial digestion, adifferent degree of shearing or a different source 35 of natural DNA, or 68 3) switch to the natural OSP method.
If all reasonable OSPs of the current GP have been5 tried and the random DNA method has been tried, both without success, we pick a new GP.
Part II 10 Sec. 10.0: Affinity Separation Means:
In Part II we optimize an affinity separationsystem that will be used in Part III to enrich apopulation of GP(vgPBD)s for those GP(PBD)s that 15 display PBDs with increased affinity for the target.
Affinity chromatography is the preferred means,but FACS, electrophoresis, or other means may also beused. 20
Sec. 10.1: Optimization of Affinity Chromatography
Separation:
Changes in eluant concentration cause GPs to elute25 from the column. Elution volume, however, is moreeasily measured and specified. It is to be understoodthat the eluant concentration is the agent causing GPrelease and that an eluant concentration can becalculated from an elution volume and the specified 30 gradient.
Using a specified elution regime, we compare theelution volumes of GP(IPBD)s with the elution volumesof wtGP on affinity columns supporting AfM(IPBD). 35 Comparisons are made at various: a) amounts of IPBD/GP, 69 b) densities of AfM(IPBD)/(volume of matrix) (DoAMoM), c) initial ionic strengths, d) elution rates, e)amounts of GP/(volume of support), f) pHs, and g)temperatures, because these are the parameters mostlikely to affect the sensitivity and efficiency of theseparation. We then pick those conditions giving thebest separation.
We do not optimize pH or temperature; rather werecord optimal values for the other parameters for oneor more values of pH and temperature. The conditionsof intended use, specified by the user (Sec. 11) , mayinclude a specification of pH or temperature. If pH isspecified, then pH will not be varied in eluting thecolumn (Sec. 15.3). Decreasing pH may be used toliberate bound GPs from the matrix. If the intendeduse specifies a temperature, we will hold the affinitycolumn at the specified temperature during elution, butwe might vary the temperature during recovery.
The AFM (IPBD) is preferably one known to havemoderate affinity for the IPBD (¾ in the range 10“6 Mto 10”8 M). When populations of GP(vgPBD)s arefractionated, there will be roughly threesubpopulations; a) those with no binding, b) those thathave some binding but can be washed off with high saltor low pH, and c) those that bind very tightly and mustbe rescued in situ. We optimize the parameters toseparate (a) from (b) rather than (b) from (c) . LetPBDW be a PBD having weak binding to the target andPBDS be a PBD having strong binding. Higher DoAMoMmight, for example, favor retention of GP(PBDW) butalso make it very difficult to elute viable GP(PBDS) .We will optimize the affinity separation to retainGP(PBDW) rather than to allow release of GP(PBDS) 70 because a tightly bound GP(PBDS) can be rescued by insitu growth. If we find that DoAMoM strongly affectsthe elution volume, then in part III we may reduce theamount of target on the affinity column when an SBD hasbeen found with moderately strong affinity (1¾ on theorder of 10”7 M) for the target.
In this step, we measure elution volumes ofgenetically pure GPs that elute from the affinitymatrix as sharp bands that can be detected by UVabsorption. Samples from effluent fractions are platedon suitable medium (cells or spores) or on sensitivecells (phage) and colonies or plaques counted.
Several values of IPBD/GP, DoAMoM, elution rates,initial ionic strengths, and loadings should beexamined. We anticipate that optimal values of IPBD/GPand DoAMoM will be correlated and therefore should beoptimized together. The effects of initial ionicstrength, elution rate, and amount of GP/(matrixvolume) are unlikely to be strongly correlated, and sothey can be optimized independently.
For each set of parameters to be tested, thecolumn is eluted in a specified manner. For example,we may use a regime called Elution Regime 1: a KC1gradient runs from lOmM to maximum allowed for theGP(IPBD) viability in 100 fractions of 0.05 Vy (voidvolume), followed by 20 fractions of 0.05 Vy at maximumallowed KCl; pH of the buffer is maintained at thespecified value with a convenient buffer such as Tris.It is important that the conditions of thisoptimization be similar to the conditions that are usedin Part III for selection for binding to target (Sec. ·:·τ>···'.· 1 71 15.3) and recovery of GPs from the chromatographicsystem (Sec. 15.4).
When the osp-ipbd gene is regulated by [XINDUCE],IPBD/GP can be controlled by varying [XINDUCE].Appropriate values of [XINDUCE] depend on the identityof [XINDUCE] and the promoter; if, for example, XINDUCEis isopropylthiogalactoside (IPTG) and the promoter islacUV5. then [IPTG] = 0, 0.1 uM, 1.0 uM, 10.0 UM, 100.0uM, and 1.0 mM are appropriate levels to test. Therange of variation of [XINDUCE] is extended until anoptimum is found or an acceptable level of expressionis obtained.
DoAMoM is varied from the maximum that the matrixmaterial can bind to 1% or 0.1% of this level inappropriate steps. We anticipate that the efficiencyof separation will be a smooth function of DoAMoM sothat it is appropriate to cover a wide range of valuesfor DoAMoM with a coarse grid and then explore theneighborhood of the approximate optimum with a finergrid.
Several values of initial ionic strength aretested, such as 1.0 mM, 5.0 mM, 10.0 mM and 20.0 mM.
The elution rate is varied, by successive factorsof 1/2, from the maximum attainable rate to 1/16 ofthis value. The fastest elution rate giving the goodseparation is optimal.
The goal of the optimization is to obtain a sharptransition between bound and unbound GPs, triggered byincreasing salt or decreasing pH or a combination ofboth. This optimization need be performed only: a) for 72 each temperature to be used, b) for each pH to be used,and c) when a new GP(IPBD) is created.
Regulatable promoters are available for allgenetic packages except, possibly, bacterial spores. Apromoter functional in bacterial spores might beprepared by constructing a hybrid of a sporulationpromoter and a regulatable bacterial promoter (e.g.,lac) , or by saturation mutagenesis of a sporulationpromoter followed by screening for regulatable promoteractivity (cf. OLIP86, OLIP87). When the promoter ofthe osp-ipbd gene is not regulatable, we optimizeDoAMoM, the elution rate, and the amount of GP/volumeof matrix. If the optimized affinity separation is notacceptable, we must develop a means to alter the amountof IPBD per GP.
Sec. 10.2:_Measuring the sensitivity of affinity separation:
We determine the sensitivity of the affinityseparation (Csens·^ by measuring the minimum quantityof GP(IPBD) that can be detected in the presence of alarge excess of wtGP. The user chooses a number ofseparation cycles, denoted Ncbroin, that will beperformed before an enrichment is abandoned;preferably, Ncbrom is in the range 6 to 10 and Ncbrommust be greater than 4. Enrichment can be terminatedby isolation of a desired GP(SBD) before Ncbrom passes.
The measurement of sensitivity is significantlyexpedited if GP(IPBD) and wtGP carry differentselectable markers. 73
Mixtures of GP(IPBD) and wtGP are prepared in theratios of where ranges by an appropriate factor (e.g. 1/10) over an appropriate range, typically1011 through 104. Large values of Viim are testedfirst; once a positive result is obtained for one valuevlim, no smaller values of Vj£m need be tested.
Each mixture is applied to a column supporting, at theoptimal DoAMoM, an AfM(IPBD) having high affinity forIPBD and the column is eluted by the specified elutionregime. The last fraction that contains viable GPs andan inoculum of the column matrix material are cultured.If GP(IPBD) and wtGP have different selectable markers,then transfer onto selection plates identifies eachcolony. Otherwise, a number (e.g. 32) of GP clonal isolates are tested for presence of IPBD by thetechniques discussed in Sec. 8.
If IPBD is not detected on the surface of any ofthe isolated GPs, then GPs are pooled from: a) the lastfew (e.g. 3 to 5) fractions that contain viable GPs,and b) an inoculum taken from the column matrix. Thepooled GPs are cultured and passed over the same columnand enriched for GP(IPBD) in the manner described.This process is repeated until Ncbroin passes have beenperformed, or until the IPBD has been detected on theGPs. If GP(IPBD) is not detected after Ncbroitl passes,Vlim is decreased and the process is repeated.
Cgensi equals the highest value of for which the user can recover GP(IPBD) within Ncbrom passes.The number of chromatographic cycles (KCyC) that wereneeded to isolate GP(IPBD) gives a rough estimate ofceff; ceff Is approximately the KCyCth root of Vlim: ceff = (approx.) exp( loge(Vlim)/Kcyc ) 74
For example, if Vjim were 4.0 x 108 and threeseparation cycles were needed to isolate GP(IPBD), thenCeff = (approx.) 736.
Sec. 10.3: Measuring the efficiency of separation :
To determine Ceff more accurately, we determinethe ratio of GP(IPBD)/wtGP loaded onto an AfM(IPBD)column that yields approximately equal amounts ofGP(IPBD) and wtGP after elution.
Sec. 10.4: Other Separation Means
Other separation means are optimized in a mannerparallel to the used for affinity chromatography. FACS (e.g. FACStar from Beckton-Dickinson,Mountain View, CA) is most appropriate for bacterialcells and spores because the sensitivity of themachines requires approximately 1000 molecules offluorescent label bound to each GP to accomplish aseparation. To optimize FACS separation of GPs, we usea derivative of Afm(IPBD) that is labeled with afluorescent molecule, denoted Afm (IPBD)*. Thevariables that must be optimized include: a) amount ofIPBD/GP, b) concentration of Afm(IPBD)*, c) ionicstrength, d) concentration of GPs, and e) parameterspertaining to operation of the FACS machine. BecauseAfm(IPBD)* and GPs interact in solution, the bindingwill be linear in both [Afm(IPBD)*] and [displayedIPBD]. Preferably, these two parameters are variedtogether. The other parameters can be optimizedindependently. 75
Electrophoresis is most appropriate tobacteriophage because of their small size (SERW87).Electrophoresis Is a preferred separation means if thetarget is so small that chemically attaching it to acolumn or to a fluorescent label would essentiallychange the entire target. For example, chloroacetateions contain only seven atoms and would be essentiallyaltered by any linkage. GPs that bind chloroacetatewould become more negatively charged than GPs that donot bind the ion and so these classes of GPs could beseparated.
The parameters to optimize for electrophoresisinclude: a) IPBD/GP, b) concentration of gel material,e.g. agarose, c) concentration of Afm (IPBD), d) ionicstrength, e) size, shape, and cooling capacity of theelectrophoresis apparatus, f) voltages and currents,and f) concentration of GPs. Preferably, IPBD/GP and[Afm(IPBD) ] are varied at the same time and otherparameters are optimized independently.
Part III
Sec. 11.0: Choice of target material :
Any material may be chosen as target material,subject only to the following restrictions:
If affinity chromatography is to be used, then: 1) the molecules of the target material must be ofsufficient size and chemical reactivity to beapplied to a solid support suitable for affinityseparation, ····.: 76 2) after application to a matrix, the targetmaterial must not react with water, 3) after application to a matrix, the targetmaterial must not bind or degrade proteins in anon-specific way, and 4) the molecules of the target material must besufficiently large that attaching the material toa matrix allows enough unaltered surface area(generally at least 500 fi2, excluding the atomthat is connected to the linker) for proteinbinding.
If FACS is to be used as the affinity separationmeans, then: 1) the molecules of the target material must be ofsufficient size and chemical reactivity to beconjugated to a suitable fluorescent dye or thetarget must itself be fluorescent, 2) after any necessary fluorescent labeling, thetarget must not react with water, 3) after any necessary fluorescent labeling, thetarget material must not bind or degrade proteinsin a non-specific way, and 4) the molecules of the target material must besufficiently large that attaching the material toa suitable dye allows enough unaltered surfacearea (generally at least 500 A2, excluding theatom that is connected to the linker) for proteinbinding. 77
If affinity electrophoresis is to be used, then: 1) the target must either be charged or of such anature that its binding to a protein will changethe charge of the protein, 2) the target material must not react with water, 3) the target material must not bind or degradeproteins in a non-specific way, and 4) the target must be compatible with a suitablegel material.
Possible target materials include, but are notlimited to: a) soluble proteins (such as horse heart myoglobin, human neutrophil elastase, activated (bloodclotting) factor X, alpha-fetoprotein, alphainterferon, melittin, Bordetella pertussis adenylatecyclase toxin, any retroviral pol protease or anyretroviral gag protease), b) lipoproteins (such ashuman low density lipoprotein), c) glycoproteins (suchas a monoclonal antibody), d) lipopolysaccharides (suchas O-antigen of Salmonella enteritidis), e) nucleicacids (such as tRNAs, ribosomal RNAs, messenger RNAsdsDNA or ssDNA, possibly with sequence specificity); f)soluble organic molecules (such as cholesterol ,aspartame, bilirubin, morphine, codeine,dichlorodiphenyltrichlorethane (DDT), benzo(a)pyrene,prostaglandin PGE2, protoporphyrin IX, or actinomycinD) , g) organometallic complexes (such as iron haem orcobolt haem), h) organic polymers (such as cellulose orchitin), i) insoluble minerals (such as asbestos,zeolites, or hydroxylapatite), j) viral and phage coat 78 proteins (such as influenza haemaggutinin or phagelambda capsid), and k) bacterial membrane or outermembrane proteins (such as LamB from E. coli orflagella proteins). A supply of several milligrams of pure targetmaterial is desired. Impure target material could beused, but one might obtain a protein that binds to acontaminant instead of to the target.
The following information about the targetmaterial is highly desirable: 1) stability as a function of temperature, pH, andionic strength, 2) stability with respect to chaotropes such asurea or guanidinium Cl, 3) pi, 4) molecular weight, 5) requirements for prosthetic groups or ions,such as haem or Ca+2, and 6) proteolytic activity, if any.
In addition to this most desirable information, itis useful to know: 1) the target's sequence, if thetarget is a macromolecule, 2) the 3D structure of thetarget, 3) enzymatic activity, if any, and 4) toxicity,if any. 79
The user of the present invention specifiescertain parameters of the intended use of the bindingprotein: 1) the acceptable temperature range, 2) the acceptable pH range, 3) the acceptable concentrations of ions andneutral solutes, 4) the maximum acceptable dissociation constantfor the target and the SBD: KT = [Target][SBD]/[Target:SBD]
In some cases, the user may require discriminationbetween T, the target, and N, some non-target. Let
Kt = [T][SBD]/[T:SBD] , and KN = [N][SBD]/[N:SBD] , then KT/KN = ([T][N:SBD])/([N][T:SBD]).
The user then specifies a maximum acceptable value forthe ratio Κφ/ΚΝ.
If the target material is a general protease, onemust consider the following points: 1) a highly specific protease can be treated likeany other target, 2) a general protease, such as subtilisin, maydegrade the OSPs of the GP including OSP-PBDs; ·??>> •·ΐ>' 80 there are several alternative ways of dealing withgeneral proteases, including: a) a chemicalinhibitor may be used to prevent proteolysis (e.g.phenylmethylfluorosulfate (PMFS) that inhibitsserine proteases), b) one or more active-siteresidues may be mutated to create an inactiveprotein (e.g. a serine protease in which theactive serine is mutated to alanine), or c) one ormore active-site amino-acids of the protein may bechemically modified to destroy the catalyticactivity (e.g. a serine protease in which theactive serine is converted to anhydroserine), 3) SBDs selected for binding to a protease neednot be inhibitors; SBDs that happen to inhibitthe protease target are a fairly small subset ofSBDs that bind to the protease target, 4) the more we modify the target protease, theless like we are to obtain an SBD that inhibitsthe target protease, and 5) if the user requires that the SBD inhibit thetarget protease, then the active site of thetarget protease must not be modified any more thannecessary; inactivation by mutation or chemicalmodification are preferred methods of inactivationand a protein protease inhibitor becomes a primecandidate for IPBD. For example, BPTI could bemutated, by the methods of the present invention,to bind to proteases other than trypsin (TANK77and TSCH87). 81
Sec. 12.0: Choice of GP(IPBD) :
The user must pick a GP(IPBD) that is suitable tothe chosen target according to the criteria of Sec. 2.It is anticipated that a small collection of aGP(IPBD)s can be assembled such that, for any chosentarget, at least one member of the collection will be asuitable starting point for engineering a protein thatbinds to the chosen target by the methods of thepresent invention. The user should optimize theaffinity separation for conditions appropriate to theintended use by the methods described in Part II.
Sec. 13.0: Identification of Family of PBDs, Related to PPBD, to Be Generated
Sec. 13.1: Choosing residues on IPBD (or other PPBD) to vary:
We choose residues in the IPBD to vary throughconsideration of several factors, including: a) the 3Dstructure of the IPBD, b) sequences homologous to IPBD,and c) modeling of the IPBD and mutants of the IPBD.Because the number of residues that could stronglyinfluence binding is always greater than the numberthat can be varied simultaneously, the user must pick asubset of those residues to vary at one time. The usermust also pick trial levels of variegation andcalculate the abundances of various sequences. Thelist of varied residues and the level of variegation ateach varied residue are adjusted until the compositevariegation is commensurate with Csens£ and Μη^ν· A key concept is that only structured proteinsexhibit specific binding, i.e. can bind to a particular 82 chemical entity to the exclusion of most others. Thusthe residues to be varied are chosen with an eye topreserving the underlying IPBD structure.Substitutions that prevent the PBD from folding will 5 cause GPs carrying those genes to bind indiscriminatelyso that they can easily be removed from the population.
Burial of hydrophobic surfaces so that bulk wateris excluded is one of the strongest forces driving the 10 binding of proteins to other molecules. Bulk water canbe excluded from the region between two molecules onlyif the surfaces are complementary. We must test asmany surfaces as possible to find one that iscomplementary to the target. The selection-through- 15 binding isolates those proteins that are more nearlycomplementary to some surface on the target. Theeffective diversity of a variegated population ismeasured by the number of different surfaces, ratherthan the number of protein sequences. Thus we should 20 maximize the number of surfaces generated in our population, rather than the number of protein sequences.
In hypothetical example 1, we consider a 25 hypothetical PBD, shown in Figure 3 binding to a hypothetical target. Figure 3 is a 2D schematic of 3Dobjects; by hypothesis, residues 1, 2, 4, 6, 7, 13, 14,15, 20, 21, 22, 27, 29, 31, 33, 34, 36, 37, 38, and 39of the IPBD are on the 3D surface cf the IPBD, even 30 though shown well inside the circle. Proteins do nothave distinct, countable faces. Therefore we define an"interaction set" to be a set of residues such that allmembers of the set can simultaneously touch onemolecule of the target material without any atom of the 35 target coming closer than van der Waals distance to any S3 main-chain atom of the IPBD. The concept of a residue"touching" a molecule of the target is discussed below.One hypothetical interaction set, Set A, in Figure 3comprises residues 6, 7, 20, 21, 22, 33, and 34,represented by squares. Another hypotheticalinteraction set, Set B, comprises residues 1, 2, 4, 6,31, 37, and 3S, represented by circles.
If we vary one residue, number 21 for example,through all twenty amino acids, we obtain 20 proteinsequences and 20 different surfaces for interaction setA. Note that residue 6 is in two interaction sets andvariation of residue 6 through all 20 amino acidsyields 20 versions of interaction set A and 20 versionsof interaction set B.
Now consider varying two residues, each throughall twenty amino acids, generating 400 proteinsequences. If the two residues varied were, forexample, number 1 and number 21, then there would beonly 40 different surfaces because interaction set Adoes not depend on residue 1 and interaction set B doesnot depend on residue 21. If the two residues varied,however, were number 7 and number 21, then 400 surfaceswould be generated.
If N spatially separated residues are varied atone time, 20 x N surfaces are generated. Variation ofN residues in the same interaction set yields 20Ksurfaces. For example, if N = 7, variation ofseparated residues yields 140 surfaces while variationof interacting residues yields 207 = 1.28 x 109surfaces. Thus, to maximize the number of surfacesgenerated when N residues are varied, all residuesshould be in the same interaction set. 84
The amount of surface area buried in strongprotein-protein interactions ranges from 1000 82 to2000 82 (SCHU79, pl03ff). Individual amino acids havetotal surface areas that depend mostly on type of aminoacid and weakly on conformation. These areas rangefrom about 180 82 for glycine to about 360 82 fortryptophan. From amino-acid solvent exposures ofpublished protein structures, we calculate that 100082on a protein surface comprises between 4 and 30 amino-acid residues. Varied amino acid sequences, as foundin actual proteins, involve between 10 and 25 residuesin forming 1000 82 of protein surface. Schulz andSchirmer estimate that 100 82 of protein surface canexhibit as many as 1000 different specific patterns(SCHU79, pl05). The number of surface patterns risesexponentially with the area that can be variedindependently. One of the BPTI structures recorded inthe Brookhaven Protein Data Bank (6PTI), for example,has a total exposed surface area of 3997 82 (using themethod of Lee and Richards (LEEB71) and a solventradius of 1.4 8 and atomic radii as shown in Table 7).If we could vary this surface freely and if 100 82 canproduce 1000 patterns, we could construct 10120different patterns by varying the surface of BPTI!This calculation is intended only to suggest the hugenumber of possible surface patterns based on a commonprotein backbone.
One protein framework cannot, however, display allpossible patterns over any one particular 100 82 ofsurface merely by replacement of the side groups ofsurface residues. The protein backbone holds thevaried side groups in approximately constant locationsso that the variations are not independent. We can, ····.?' aiR'n · 85 nevertheless, generate a vast collection of differentprotein surfaces by varying those protein residues thatface the outside of the protein.
Examination of a model of BPTI in contact withmyoglobin shows that residues 3, 7, 8, 10, 13, 39, 41,and 42 can all simultaneously contact a molecule thesize and shape of myoglobin. Residue 49 cannot touch asingle myoglobin molecule simultaneously with any ofthe first set even though all are on the surface ofBPTI. It is not the intent of the present invention,however, to use models to determine wThich part of thetarget molecule will actually be the site of binding bya PBD.
For cassette mutagenesis, the protein residues tobe varied are, preferably, close enough in sequencethat the variegated DNA (vgDNA) encoding all of themcan be made in one piece. The present invention is not·limited to a particular length of vgDNA that can besynthesized. With current technology, a stretch of 60amino acids (180 DNA bases) can be spanned.
One can use other mutational means, such assingle-stranded-oligonucleotide-directed mutagenesis(BOTS85) using two or more mutating primers to mutatewidely separated residues.
Alternatively, to vary residues separated by morethan sixty residues, two cassettes may be mutated. Afirst cassette is mutagenized to produce a populationhaving, for example, up to 30,000 members. Usingvariegated OCV, we mutagenize a second cassette toproduce a second variegated population having thedesired diversity. 86
The composite level of variation must not exceedthe prevailing capabilities to a) produce very largenumbers of independently transformed cells or b) detectsmall components in a highly varied population. Thelimits on the level of variegation are discussed inSec. 13.2.
We assemble the data about the IPBD and the targetthat are useful in deciding which residues to vary 1)3D structure, or at least a list of residues on thesurface of the IPBD, 2) list of sequences homologous toIPBD, and 3) model of the target molecule or a stand-infor the target.
These data and an understanding of the behavior ofdifferent amino acids in proteins will be used toanswer two questions: 1) which residues of the IPBD are on the outsideand close enough together in space to touch thetarget simultaneously? 2) which residues of the IPBD can be varied withhigh probability of retaining the underlying IPBDstructure?
Although an atomic model of the target materialfrom X-ray crystallography, NMR, etc. is preferred insuch examination, it is not necessary. For example, ifthe target were a protein of unknown 3D structure, itwould be sufficient to know the molecular weight of theprotein and whether it were a soluble globular protein,a fibrous protein, or a membrane protein. One can thenchoose a protein of known structure of the same class 87 and similar size and shape to use as a molecular stand-in and yardstick. At low resolution, all proteins of agiven size and class look much the same. The specificvolumes are the same, all are more or less sphericaland therefore all proteins of the same size and classhave about the same radius of curvature. The radii ofcurvature of the two molecules determine how much ofthe two molecules can come into contact.
The most appropriate method of picking theresidues of the protein chain at which the amino acidsshould be varied is by viewing, with interactivecomputer graphics, a model of the IPBD. A stick-figurerepresentation of molecules is preferred. A suitableset of hardware is an Evans &amp; Sutherland PS390 graphicsterminal (Evans &amp; Sutherland Corporation, Salt LakeCity, UT) and a MicroVAX II supermicro computer(Digital Equipment Corp., Maynard, MA). Suitableprograms for viewing and manipulating protein modelsinclude: a) PS-FRODO, written by T. A. Jones (JONE85)and distributed by the Biochemistry Department of RiceUniversity, Houston, TX; and b) PROTEUS, developed byDayringer, Tramantano, and Fletterick (DAYR86).
Theoretical calculations, such as dynamicsimulations of proteins, are used to estimate theeffect of substitution at a particular residue of aparticular amino-acid type on the 3D structure of theparent protein. Such calculations might also indicatewhether a particular substitution will greatly affectthe flexibility of the protein.
Sec. 13.1.1: The principal set: 88
Using the knowledge of which residues are on thesurface of the IPBD, we pick residues that are closeenough together on the surface of the IPBD to touch amolecule of the target simultaneously without havingany IPBD main-chain atom come closer than van der Waalsdistance (viz. 4.0 to 5.0 A) from any target atom. Aresidue of the IPBD "touches" the target if: a) a main-chain atom is within van der Waals distance, viz. 4.0to 5.0 8 of any atom of the target molecule, or b) thecbeta is within DcutOff °f any atom of the targetmolecule so that a side-group atom could make contactwith that atom. Because side groups differ in size(cf. Table 35) , some judgment is required in pickingDCutoff· In tlle preferred embodiment, we will useDCutoff = 8,0 but other values in the range 6.0 8 to10.0 8 could be used. If IPBD has G at a residue, weconstruct a pseudo C^g^-a with the correct bond distanceand angles and judge the ability of the residue totouch the target from this pseudo C^g^g.
Alternatively, we choose a set of residues on thesurface of the IPBD such that the curvature of thesurface defined by the residues in the set is not sogreat that it would prevent contact between allresidues in the set and a molecule of the target. Thismethod is appropriate if the target is a macromolecule,such as a protein, because the PBDs derived from theIPBD will contact only a part of the macromolecularsurface.
We prefer that there be some indication that theunderlying IPBD structure will tolerate substitutionsat each residue in the principal set of residues.Indications could come from various sources, including: 89 a) homologous sequences, b) static computer modeling,or c) dynamic computer simulations.
The residues in the principal set need not becontiguous in the protein sequence. We require onlythat the amino acids in the residues to be varied allbe capable of touching a molecule of the targetmaterial simultaneously without having atoms overlap.If the target were, for example, horse heart myoglobin,and if the IPBD were BPTI, any set of residues in oneinteraction set of BPTI defined in Table 34 could bepicked.
Preferably, the principal set contains eight tosixteen residues. This number of residues allowssufficient variability that a. surface that iscomplementary to the target can be found, but is smallenough that a significant fraction of the surface canbe varied at one time.
Sec. 13.1.2: The secondary set:
The secondary set comprises residues that touchresidues in the primary set, and are excluded from theprimary set because the residue: a) is internal, b) ishighly conserved, or c) is on the surface, but thecurvature of the IPBD surface prevents the residue frombeing in contact with the target at the same time asone or more residues in the primary set.
Internal residues, although frequently conservedand may tolerate some conservative changes such as I toL or F to Y. These changes affect the detail placementand dynamics of adjacent protein residues and suchvariation may be useful once an SBD is found. 90
Surface residues in the secondary set are mostoften located on the periphery of the principal set,which do not make direct contact with the targetsimultaneously with all other residues of the principalset. The charge on the amino acid in one of theseresidues could, however, have a strong effect onbinding. It is appropriate to vary the charge of someor all of these residues to improve an SBD. Forexample, the variegated codon containing equimolar Aand G at base 1, equimolar C and A at base 2, and A atbase 3 yields amino acids T, A, K, and E with equalprobability.
Sec. 13.1.3: Choice of residues to vary initially:
The allowed level of variegation that assuresprogressively determines how many residues can bevaried at once; geometry determines which ones.
The user picks residues to vary in many ways; thefollowing is a preferred manner. Pairs of residues arepicked that are diametrically opposed across the faceof the principal set. Two such pairs are used todelimit the surface, up/down and right/left.Alternatively, three residues that form an inscribedtriangle, having as large an area as possible, on thesurface are picked. One to three other residues arepicked in a checkerboard fashion across the interactionsurface. Choice of widely spaced residues to varycreates the possibility for high specificity becauseall the intervening residues must have acceptablecomplementarity before favorable interactions can occurat widely-separated residues.
<img img-format="tif" img-content="drawing" file="IL120940AD00026.tif" id="idf0006" />
91
The number of residues picked is coupled to therange through which each can be varied by therestrictions discussed in Sec. 13.2. In the firstround, we do not assume any binding between IPBD and 5 the target and so progressivity is not an issue. Atthe first round, the user may elect to produce a levelof variegation such that each molecule of vgDNA ispotentially different through, for example, unlimitedvariegation of 10 codons (2010 approx. = 1013). One 10 run of the DNA synthesizer produces approximately 1013molecules of length 100 nts. Inefficiencies in ligation and transformation will reduce the number ofproteins actually tested to between 107 and 5 x 108.Multiple iterations of the process with such very high 15 levels of variegation will not yield repeatableresults; the user must decide whether this isimportant.
Sec. 13.2: Range of variation at Each Site of 20 Mutation:
The total level of variegation is the product ofthe number of variants at each varied residue. Eachvaried residue can have a different scheme of 25 variegation, producing 2 to 20 different possibilities.We require that the process be progressive, i.e. eachvariegation cycle produces a better starting point forthe next variegation cycle than the previous cycleproduced. 30 N.B. : Setting the level of variegation suchthat the ppbd and many sequences related tothe ppbd sequence are present in detectableamounts insures that the process isprogressive. If the level of variegation is 35 92 so high that the ppbd sequence is present atsuch low levels that there is an appreciablechance that no transformant will display thePPBD, then the best SBD of the next round 5 could be worse than the PPBD. At excessively high level of variegation, each round ofmutagenesis is independent of previous roundsand there is no assurance of progressivity.
This approach can lead to valuable binding 10 proteins, but repetition of experiments with this level of variegation will not yieldprogressive results. Excessive variation isnot preferred. 15 If the level of variegation is such that the parental sequence and each single amino-acid change ispresent for selection, then we know that a selectedsequence is closer to optimal or the same as theparent. If, on the other hand, very high levels of 2 0 variegation are used, a sequence may be selected, notbecause it is superior to the parental sequence, butbecause the parental and improved sequences are, bychance, absent. 25 Progressivity is not an all-or-nothing property.
So long as most of the information obtained fromprevious variegation cycles is retained and manydifferent surfaces that are related to the PPBD surfaceare produced, the process is progressive. If the level 30 of variegation is so high that the ppbd gene may not bedetected, the assurance of progressivity diminishes.If the probability of recovering PPBD is negligible,then the probability of progressive behavior is alsonegligible. 35 93
An opposing force in our design considerations isthat PBDs are useful in the population only up to theamount that can be detected; any excess above thedetectable amount is wasted. Thus we produce as manysurfaces, related to PPBD as possible within theconstraint that the PPBD be detectable.
We defer specification of exactly how muchvariegation is allowed until we have: a) specified realnt distributions for a variegated codon, and b)examined the effects of discrepancies between specifiednt distributions and actual nt distributions.
Sec. 13.3: Design of vgDNA Encoding PBD Family:
We must now decide how to distribute thevariegation within the codons for the residues to bevaried. These decisions are influenced by the natureof the genetic code. When vgDNA is synthesized,variation at the first base of a codon creates apopulation containing amino acids from the same columnof the genetic code table (as shown in the Table 3-6 onp87 of WATS87) ; variation at the second base of thecodon creates a population containing amino acids fromthe same row of the genetic code table; variation atthe third base of the codon creates a populationcontaining amino acids from the same box. If two orthree bases in the same codon are varied, the patternis more complicated. Work with 3D protein structuralmodels may suggest definite sets of amino acids tosubstitute at a given residue, but the method ofvariation may require either more or fewer kinds ofamino acids be included. For example, examination of amodel might suggest substitution of N or Q at a givenresidue. Combinatorial variation of codons requires
<img img-format="tif" img-content="drawing" file="IL120940AD00027.tif" id="idf0007" />
94 that mixing N and Q at one location also include K andH as possibilities at the same residue. One mustchoose to put: 1) N only, 2) Q only, or 3) a mixture ofN, K, H, and Q. The present invention does not rely onaccurate predictions of which amino acids should beplaced at each residue, rather attention is focused onwhich residues should be varied.
There are many ways to generate diversity in aprotein. (See RICH86, CARU85, and OLIP86.) One extremecase is that one or a few residues of the protein arevaried as much as possible (inter alia see CARU85,CARU87, RICH86, and WHAR86). We will call this limit"Focused Mutagenesis". Focused Mutagenesis isappropriate when the IPBD or other PPBD shows little orno binding to the target, as at the beginning of thesearch for a protein to bind to a new target material.When there is no binding between the PPBD and thetarget, we preferably pick a set of five to sevenresidues and vary each through all 20 possibilities.
An alternative plan of mutagenesis ("DiffuseMutagenesis") is to vary many more residues through amore limited set of choices (See Vershon et al. . Chl5of INOU8 6 and PAKU86) . This can be accomplished byspiking each of the pure nts activated for DNAsynthesis (e.q. nt-phosphoramidites) with a smallamount of one or more of the other activated nts.Contrary to general practice, the present inventionsets the level of spiking so that only a smallpercentage ( 1% to .00001%, for example ) of the finalproduct contains the initial DNA sequence. Manysingle, double, triple, and higher mutations occur, butrecovery of the basic sequence is a possible outcome.Let Nj-, be the number of bases to be varied, and let Q 95 be the fraction of all sequences that should have theparental sequence, then M, the fraction of the mixturethat is the majority component, is M = exp{ loge(Q)/Nb } = 10 (lo9l0(Q)/Nb).
If, for example, thirty base pairs on the DNAchain were to be varied and 1% of the product is tohave the parental sequence, then each mixed ntsubstrate should contain 86% of the parental nt and 14%of other nts. Table 8 shows the fraction (fn) of DNAmolecules having n non-parental bases when 30 bases aresynthesized with reagents that contain fraction M ofthe majority component. When M=.63096, f24 and higherare less than 10“8. The entry "most" in Table 8 is thenumber of changes that has the highest probability.Note that substantial probability for multiplesubstitutions only occurs if the fraction of parentalsequence (fO) is allowed to drop to around 10~6.Mutagenesis of this sort can be applied to any part ofthe protein at any time, but is most appropriate whensome binding to the target has been established. TheNb base pairs of the DNA chain that are synthesizedwith mixed reagents need not be contiguous. They arepicked so that between Nb/3 and Nb codons are affectedto various degrees. The residues picked for mutationare picked with reference to the 3D structure of theIPBD, if known. For example, one might pick all ormost of the residues in the principal and secondaryset. We may impose restrictions on the extent ofvariation at each of these residues based on homologoussequences or other data. The mixture of non-parentalnts need not be random, rather mixtures can be biasedto give particular amino acid types specificprobabilities of appearance at each codon. For 96 example, one residue may contain a hydrophobic aminoacid in all known homologous sequences; in such a case,the first and third base of that codon would be varied,but the second would be set to T. This diffusestructure-directed mutagenesis will reveal the subtlechanges possible in protein backbone associated withconservative interior changes, such as V to I, as wellas some not so subtle changes that require concomitantchanges at two or more residues of the protein.
For Focused Mutagenesis, we now consider thedistribution of nts that will be inserted at eachvariegated codon. Each codon could be programmeddifferently. If we have no information indicating thata particular amino acid or class of amino acid isappropriate, we strive to substitute all amino acidswith equal probability because representation of onepbd above the detectable level is wasteful. Equalamounts of all four nts at each position in a codonyields the amino acid distribution in which each aminoacid is present in proportion to the number of codonsthat code for it. This distribution has thedisadvantage of giving two basic residues for everyacidic residue. In addition, six times as much R, S,and L as W or M occur. If five codons are synthesizedwith this distribution, sequences encoding five Rs are7776-times more abundant than sequences encoding fiveWs. To have W-W-W-W-W present at detectable levels, wemust have R-R-R-R-R present in 7776-fold excess.
Let Abun(x) be the abundance of DNA sequencescoding for amino acid x, defined by the distribution ofnts at each base of the codon. For any distribution,there will be a most-favored amino acid (mfaa) withabundance Abun(mfaa) and a least-favored amino acid 97 (lfaa) with abundance Abun(lfaa) . We seek the ntdistribution that allows all twenty amino acids andthat yields the largest ratio Abun(lfaa)/Abun(mfaa)subject to two constraints: equal abundances of acidicand basic amino acids and the least possible number ofstop codons. Thus only nt distributions that yieldAbun(E)+Abun(D) = Abun(R)+Abun(K) are considered, andthe function maximized is: ((l-Abun(stop)) (Abun(lfaa)/Abun(mfaa))}.
We have simplified the search for an optimal ntdistribution by limiting the third base to T or G (C orG is equivalent). All amino acids are possible and thenumber of accessible stop codons is reduced because TGAand TAA codons are eliminated. The amino acids F, Y,C, Η, N, I, and D require T at the third base while W,M, Q, K, and E require G. Thus we use an equimolarmixture of T and G at the third base. A computer program, written as part of the presentinvention and named "Find Optimum vgCodon" (See Table 9), varies the composition at bases 1 and 2, in stepsof 0.05, and reports the composition that gives thelargest value of the quantity {(Abun(lfaa)/Abun(mfaa)(1-Abun(stop)))}. A vg codon is symbolically definedby the nt distribution at each base: T C A G base #1 = tl cl al gi base #2 = t2 c2 a2 g2 base #3 = t3 C3 a3 g3 tl + cl + al + gl = 1.0 t2 + c2 + a2 + g2 = 1.0 98 t3 = g3 = 0.5, c3 = a3 = 0.
The variation of the quantities tl, cl, al, gl, t2, c2,a2, and g2 is subject to the constraint thatAbun(E)+Abun(D) equals Abun(K)+Abun(R);
Abun(E)+Abun(D) = gl*a2
Abun(K)+Abun(R) = al*a2/2 + cl*g2 + al*g2/2 gl*a2 = al*a2/2 + cl*g2 + al*g2/2
Solving for g2, we obtain g2 = (gl*a2 - 0.5*al*a2)/(cl + 0.5*al)
In addition, tl = 1 - al - cl - glt2 = 1 - a2 - c2 - g2
We vary al, cl, gl, a2, and c2 and then calculate tl,g2, and t2. Initially, variation is in steps of 5%.Once an approximately optimum distribution of nts isdetermined, the region is further explored with stepsof 1%. The logic of this program is shown in Table 9.The optimum distribution is:
Optimum vqCodon T C A G base #1 = 0.26 0.18 0.26 0.30 base #2 = 0.22 0.16 0.40 0.22 base #3 = 0.5 0.0 0.0 0.5 99 and yields DNA molecules encoding each type amino acidwith the abundances shown in Table 10.
The computer that controls a DNA synthesizer, suchas the Milligen 7500, can be programmed to synthesizeany base of an oligo-nt with any distribution of nts bytaking some nt substrates (e.g. nt phosphoramidites)from each of two or more reservoirs. Alternatively, ntsubstrates can be mixed in any ratios and placed in oneof the extra reservoir for so called "dirty bottle"synthesis.
The actual nt distribution obtained will differfrom the specified nt distribution due to severalcauses, including: a) differential inherent reactivityof nt substrates, and b) differential deterioration ofreagents. It is possible to compensate partially forthese effects, but some residual error will occur. Wedenote the average discrepancy between specified andobserved nt fraction as Serr,
Serr = square root ( average[ (fobs - fspec)/fspec J ) were fobs is the amount of one type of nt found at abase and fSpec is the amount of that type of nt thatwas specified at the same base. The average is overall specified types of nts and over a number fe.g. 10or 20) different variegated bases. By hypothesis, theactual nt distribution at a variegated base will bewithin 5% of the specified distribution. Actual DNAsynthesizers and DNA synthetic chemistry may havedifferent error levels. It is the user'sresponsibility to determine Serr for the DNAsynthesizer and chemistry employed. 100
To determine the possible effects of errors in ntcomposition on the amino-acid distribution, we modifiedthe program "Find Optimum vgCodon" in four ways: 5 1) the fraction of each nt in the first two bases is allowed to vary from its optimum value times (1- Serr) to the optimum value times (1 + Serr) inseven equal steps (Serr is the hypotheticalfractional error level entered by the user) ,· the 10 sum of nt fractions at one base always equals 1.0, 2) g2 is varied in the same manner as a2, i.e. we dropped the restriction that Abun(D)+Abun(E) =
Abun(K)+Abun(R), 15 3) t3 and g3 are varied from 0.5 times (1 - Serr)to 0.5 times (1 + Serr) in three equal steps, 4) the smallest ratio Abun(Ifaa)/Abun(mfaa) is 20 sought.
In actual experiments, we will direct the synthesizerto produce the optimum DNA distribution "OptimumvgCodon" given above. Incomplete control over DNA 25 chemistry may, however, cause us to actually obtain thefollowing distribution that is the worst that can beobtained if all nt fractions are within 5% of theamounts specified in "Optimum vgCodon". Acorresponding table can be calculated for any given 30 serr usin,3 the program "Find worst vgCodon within Serrof given distribution." given in Table 11.
Optimum vgCodon, worst 5% errors 35 101 base #1 = 0.251 0.189 0.273 0.287 base #2 = 0.209 0.160 0.400 0.231 base #3 = 0.475 0.0 0.0 0.525
This distribution yields DNA encoding differentamino acids at the abundances shown in Table 12.
If five codons are synthesized with reagents mixedso as to produce the nt-distribution "Optimum vgCodon",and if we actually obtained the nt-distribution"Optimum vgCodon, worst 5% errors", then DNA sequencesencoding the mfaa at all of the five codons are about277 times as likely as DNA sequences encoding the lfaaat all of the five codons; about 24% of the DNAsequences will have a stop codon in one or more of thefive codons.
When five codons are synthesized using equimolarmixtures at bases 1 and 2, (Abun(mfaa)/Abun(lfaa))5 =7776. If we program the optimum nt distribution andcome within 5%, then (Abun(mfaa)/Abun(lfaa)) 5 = 277.The total number of different PBDs is unchanged, butthe least-favored sequence is about 28 times moreabundant. Detecting the least-favored amino-acidsequence when varying four residues with equimolar ntsat each varied base requires as sensitive a separationsystem as does detecting the least-favored amino-acidsequence when varying five residues with the optimizednt distribution.
By hypothesis, the distribution "Optimal vgCodon"is used in the second version of the second variegationof hypothetical example 2. The abundance of the DNAencoding each type of amino acid is, however, taken 102 from the Table 12. The abundance of DNA encoding theparental amino acid sequence is:
Amount(parental seq.) E42 * Abun(E) * X .0602 X T47 Abun(T) .0437 F24 = Abun(F) = .0249 = 2.4 X 1( G30 * Abun(G) x .0663 )-7 D34 * Abun(D) x .0545 Therefore, DNA encoding the PPBD sequence as well as very many related sequences will be present insufficient quantity to be detected and we are assuredthat the process will be progressive. A level of variegation that allows recovery of thePPBD has two properties: 1) we cannot regress because the PPBD isavailable, 2) an enormous number of multiple changes relatedto the PPBD are available for selection and we areable to detect and benefit from these changes.
The user must adjust the list of residues to bevaried and levels of variegation at each residue untilthe calculated variegation is within the bounds set by^ntv an<^ csensi·
Preferably, we also consider the interactionsbetween the sites of variegation and the surroundingDNA. If the method of mutagenesis to be used isreplacement of a cassette, we consider whether thevariegation will generate gratuitous restriction sitesand whether they seriously interfere with the intended •·Τ>, 103 introduction of diversity. We reduce or eliminategratuitous restriction sites by appropriate choice ofvariegation pattern and silent alteration of codonsneighboring the sites of variegation. See the DetailedExample.
Sec. 14.1:_Insertion of synthetic vgDNA into a
Plasmids:
For cassette mutagenesis, restriction sites weredesigned and synthesized, and are used to introduce thesynthetic vgDNA into the OCV. Restriction digestionsand ligations are performed by standard methods(AUSU87). In the case of single-stranded-oligonucleotide-directed mutagenesis, synthetic vgDNAis used to create diversity in.the vector (BOTS85).
Sec. 14.2: Transformation of cells:
The present invention is not limited to any onemethod of transforming cells with DNA. Standardmethods, such as thos described in MANI82, may beoptimized for the particular host cells and OCV. Thegoal is to produce a large number of independenttransformants, preferably 107 of more. It is notnecessary to isolate transformed cells betweentransformation and affinity separation. We prefer tohave transformed cells at high concentration so thatthey can be plated densely on relatively few plates.
Sec. 14.3: Growth of the GP(vqPBD) population:
The transformed cells are grown first under non-selective conditions that allow expression of plasmidgenes and then selected to kill untransformed cells. 104
Transformed cells are then induced to express the osp-pbd gene at the appropriate level of induction, asdetermined in Sec. 10.1. The GPs carrying the IPBD areharvested by a method appropriate to the package. 5 A high level of diversity can be generated by invitro variegated synthesis of DNA and this diversitycan be maintained passively through several generationsin an organism without positive selective pressure. 10 Loss or reduction in frequency of deleterious mutationsis advantageous for the purposes of the presentinvention. It is preferable that the selection is mustbe performed before more than a few generations elapse.Moreover, subdividing the variegated population before 15 amplification in an organism by removing a small sample(less than 10%) for further work would result in lossof diversity; therefore, one should use all or most ofthe synthetic DNA and most or all of the transformedcells. 20
Sec. 15.:_Isolation of GP(PBD)s with bindinq-to- tarqet phenotypes :
The harvested packages are enriched for the 25 binding-to-target phenotype by use of affinityseparation involving target material immobilized on amatrix. Packages that fail to bind to target materialare washed away. If the packages are bacteriophage orendospores, it may be desirable to include a 30 bacteriocidal agent, such as azide, in the buffer toprevent bacterial growth.
Sec. 15.1: Attaching the target material to a column: 105
Affinity column chromatography is the preferredmethod of affinity separation, but other affinityseparation methods may be used. A variety ofcommercially available support materials for affinitychromatography are used. These include derivatizedbeads to which the target material is covalentlylinked, or non-derivatized material to which the targetmaterial adheres irreversibly.
Suppliers of support material for affinitychromatography include: Applied Protein TechnologiesCambridge, MA; Bio-Rad Laboratories, Rockville Center,NY; Pierce Chemical Company, Rockford, IL. Targetmaterials are attached to the matrix in accord with thedirections of the manufacturer of each matrixpreparation with consideration of good presentation ofthe target.
Sec. 15.2: Reducing selection due to non-specific binding:
We reduce non-specific binding of GP(PBD)s to thematrix that bears the target in two ways: 1) we treat the column with blocking agents suchas genetically defective GPs or a solution ofprotein before the population of GP(vgPBD)s ischromatographed, and 2) we pass the population of GP(vgPBD)s over amatrix containing no target or a different targetfrom the same class as the actual target prior toaffinity chromatography.
<img img-format="tif" img-content="drawing" file="IL120940AD00028.tif" id="idf0008" />
106
Step (1) above saturates any non-specific binding thatthe affinity matrix might show toward wild-type GPs orproteins in general; step (2) removes components of our population that exhibit non-specific binding to the 5 matrix or to molecules of the same class as the target.If the target were horse heart myoglobin, for example,a column supporting bovine serum albumin could be usedto trap GPs exhibiting PBDs with strong non-specificbinding to proteins. If cholesterol were the target, 10 then a hydrophobic compound, such as p-tertiarybutylbenzyl alcohol, could be used to removeGPs displaying PBDs having strong non-specific bindingto hydrophobic compounds. It is anticipated that PBDsthat fail to fold or that are prematurely terminated 15 will be non-specifically sticky. The capacity of theinitial column that removes indiscriminately adhesivePBDs should be greater (e.g. 5 fold greater) than thecolumn that supports the target molecule. 20 Variation in the support material (polystyrene, glass, agarose, etc.) in analysis of clones carryingSBDs is used to eliminate enrichment for packages thatbind to the support material rather than the target. 25 Sec. 15.3: Eluting the column:
The population of GPs is applied to an affinitymatrix under conditions compatible with the intendeduse of the binding protein and the population is 30 fractionated by passage of a gradient of some soluteover the column. The process enriches for PBDs havingaffinity for the target and for which the affinity forthe target is least affected by the eluants used. 107
Ions or cofactors needed for stability of PBDs(derived from IPBD) or target must be included inbuffers at appropriate levels. We first removeGP(PBD)s that do not bind the target by washing the 5 matrix with the volume of the initial buffer requiredto bring the optical density (at 260 nm or 280 nm) backto base line plus one to five void volumes (Vv) . Thecolumn is then eluted with a gradient of increasing: a)salt, b) [H+] (decreasing pH) , c) neutral solutes, d) 10 temperature (increasing or decreasing), or e) somecombination of these factors. Salt is the mostpreferred solute for gradient formation. Other solutesthat generally weaken non-covalent interaction may alsobe used. "Salt" includes solutions containing any of 15 the following ionic species: Na+ K+ Ca++ Mg++ 20 nh4+ Li+ Sr++ Ba++ Rb+ Cs+ Cl- Br- 25 S04 — hso4- PO4--- hpo4— h2po4- co3— hco3- Acetate Citrate Standard 1- StandardAmino Acids nucleotides Guanidinium Cl 30
Other ionic or neutral solutes may be used. Allsolutes are subject to the necessity that they not killthe genetic packages. Neutral solutes, such as 35 ethanol, acetone, ether, or urea, are frequently usedin protein purification, however, many of these arevery harmful to bacteria and bacteriophage above lowconcentrations. Bacterial spores, on the other hand,are impervious to most neutral solutes. Several passes 40 may be made through the steps in Sec. 15. Different 108 solutes may be used in different analyses, salt in one,pH in the next, etc. 10 15 20 25
Sec. 15.4: Recovery of packages:
Recovery of packages that display binding to anaffinity column may be achieved in several ways,including from: 1) fractions eluted with a gradient as describedabove; 2) fractions eluted with soluble target material, 3) cells grown in situ on the matrix, 4) cells incubated with parts of the matrix, 5) fractions eluted after chemically orenzymatically degrading the linkage holding thetarget to the matrix, and 6) regeneration of GPs after degrading thepackages and recovering OCV DNA.
It is possible to utilize combinations of thesemethods. It should be remembered that what we want torecover from the affinity matrix is not the GPs per se.but the information in them. Recovery of viable GPs isvery strongly preferred, but recovery of geneticmaterial is essential.
Inadvertent inactivation of the GPs is verydeleterious. It is preferred that maximum limits for 30 solutes that do not inactivate the GPs or denature thetarget or the column are determined. One may useconditions that denature the column to elute GPs;before the target is denatured, a portion of theaffinity matrix should be removed for possible use as 35 an inoculum. As the GPs are held together by protein- 109 protein interactions and other non-covalent molecularinteractions, there will be cases in which themolecular package will bind so tightly to the targetmolecules on the affinity matrix that the GPs can notbe washed off in viable form. This will only occurwhen very tight binding has been obtained. In thesecases, methods (3) through (5) above can be used toobtain the bound packages or the genetic messages fromthe affinity matrix.
It is possible, by manipulation of the elutionconditions, to isolate SBDs that bind to the target atone pH (pHb) but not at another pH (pHo) . Thepopulation is applied at pHb and the column is washedthoroughly at pHb. The column is then eluted withbuffer at pHo and GPs that come off at the new pH arecollected and cultured. Similar procedures may be usedfor other solution parameters, such as temperature.For example, GP(vgPBD)s could be applied to a columnsupporting insulin. After eluting with salt to removeGPs with little or no binding to insulin, we elute withsalt and glucose to liberate GPs that display PBDs thatbind insulin or glucose in a competitive manner.
Sec. 15.5: Amplifying the Enriched Packages
Viable GPs having the selected binding trait areamplified by culture in a suitable medium, or, in thecase of phage, infection into a host so cultivated. Ifthe GPs have been inactivated by the chromatography,the OCV carrying the osp-pbd gene must be recoveredfrom the GP, and introduced into a new, viable host.
Sec. 15.6: Determining whether further enrichment is needed: 110
The probability of isolating a GP with improvedbinding increases by Ceff with each separation cycle.Let N be the number of distinct amino-acid sequencesproduced by the variegation. We want to perform Kseparation cycles before attempting to isolate an SBD,where K is such that the probability of isolating asingle SBD is 0.10 or higher. K = the smallest integer>= log10(0.10 N)/log10(Ceff)
For example, if N were 1.0 x 107 and Ceff = 6.31 x 102then log10(1.0 x 106)/log10(6.31 x 102) = 6.0000/2.8000= 2.14. Therefore we would attempt to isolate SBDsafter the third separation cycle. After only twoseparation cycles, the probability of finding an SBD is(6.31 x 102)2/(1.0 x 107) = .04 and attempting toisolate SBDs might be profitable.
Clonal isolates from the last fraction eluted inSec. 15.3 containing any viable GPs, as well as clonalisolates obtained by culturing an inoculum taken fromthe affinity matrix, are cultured. If K separationcycles have been completed, samples from a number, e.q.32, of these clonal isolates are tested for elutionproperties on the (target) column. If none of theisolated, genetically pure GPs show improved binding totarget, or if K cycles have not yet been completed,then we pool and culture, in a manner similar to themanner set forth in Sec. 14.3, the GPs from the lastfew fractions eluted (see Sec. 15.4) that containedviable GPs and from the GPs obtained by culturing aninoculum taken from the column matrix. We then repeatthe enrichment procedure described in Sec. 15. This 111 cyclic enrichment may continue Nchroni passes or untilan SBD is isolated.
If one or more of the isolated GPs has improvedretention on the {target} column, we determine whetherthe retention of the candidate SBDs is due to affinityfor the target material. Target material is attachedto a different support matrix at optimal density andthe elution volumes of candidate GP(SBD)s are measured.We pick the candidate that either has the highestelution volume or that is retained on the column afterelution. If none of the candidate GP(SBD)s has higherelution volume than GP(PPBD of this round), then wepool and culture the GPs from the last few fractionsthat contained viable GPs and the GPs obtained byculturing an inoculum taken from the column matrix. Wethen repeat the enrichment procedure of Sec. 15.
If all of the SBDs show binding that is superiorto PPBD of this round, we pool and culture the GPs fromthe last fraction that contains viable GPs and from theinoculum taken from the column. This population is re-chromatographed at least one pass to fractionatefurther the GPs based on K^.
If an RNA phage were used as GP, the RNA wouldeither be cultured with the assistance of a helperphage or be reverse transcribed and the DNA amplified.The amplified DNA could then be sequenced or subclonedinto suitable plasmids.
Sec. 15.7: Characterizing the Population:
We characterize members of the population showingdesired binding properties by genetic and biochemical 112 methods. We obtain clonal isolates and test thesestrains by genetic and affinity methods to determinegenotype and phenotype with respect to binding totarget. For several genetically pure isolates thatshow binding, we demonstrate that the binding is causedby the artificial chimeric gene by excising the osp-sbdgene and crossing it into the parental GP. We alsoligate the deleted backbone of each GP from which theosp-sbd is removed and demonstrate that each backbonealone cannot confer binding to the target on the GP.We sequence the osp-sbd gene from several clonalisolates.
Sec. 15.8: Testing of binding affinity:
For one or more clonal isolates, we subclone thesbd gene fragment, without the osp fragment, into anexpression vector such that each SBD can be produced asa free protein. Each SBD protein is purified by normalmeans, including affinity chromatography. Physicalmeasurements of the strength of binding are then madeon each free SBD protein by one of the followingmethods: 1) alteration of the Stokes radius as afunction of binding of the target material, measured bycharacteristics of elution from a molecular sizingcolumn such as agarose, 2) retention of radiolabeledSBD on a spun affinity column to which has been affixedthe target material, or 3) retention of radiolabeledtarget material on a spun affinity column to which hasbeen affixed the SBD. The measurements of binding foreach free SBD are compared to the correspondingmeasurements of binding for the PPBD.
In each assay, we measure the extent of binding as 113 a function of concentration of each protein, and otherrelevant physical and chemical parameters.
In addition, the SBD with highest affinity for thetarget from each round is compared to the best SBD ofthe previous round (IPBD for the first round) and tothe IPBD with respect to affinity for the targetmaterial. Successive rounds of mutagenesis andselection-through-binding yield increasing affinityuntil desired levels are achieved.
If binding is not yet sufficient, we must decidewhich residues to vary next (see Sec. 16.0).
Sec. 15.9: Other Affinity Separation Means: FACs may be used to separate GPs that bindfluorescent labeled target with the optimizedparameters determined in Part II. We discriminateagainst artifactual binding to the fluorescent lable byusing two or more different dyes, chosen to bestructurally different.
Electrophoretic affinity separation uses unalteredtarget so that only other ions in the buffer can giverise to artifactual binding. Artifactual binding tothe gel material gives rise to retardation independentof field direction and so is easily eliminated. Avariegated population of GPs will have a variety ofcharges.
First the variegated population of GPs iselectrophoresed in a gel that contains no targetmaterial. The electrophoresis continues until the GPsare distributed along the length of the lane. The 114 target-free lane in which the initial electrophoresisis conducted is separated by a removable baffle from asquare of gel that contains target material. Thebaffle is removed and a second electrophoresis isconducted at right angles to the first. GPs that donot bind target migrate with unaltered mobility whileGPs that do bind target will separate from the majoritythat do not bind target. A diagonal line of non-binding GPs will form. This line is excised anddiscarded. Other parts of the gel are dissolved andthe GPs cultured.
Sec. 16.0: The Next Variegation Cycle:
Which residues of the PBD should be varied in thenext variegation cycle? The general rule is topreserve as much accumulated information as possible.The amino acids just varied are the ones bestdetermined. The environment of other residues haschanged, so that it is appropriate to vary them again.Because there are always more residues in the principaland secondary sets than can be varied simultaneously,we start by picking residues that either have neverbeen varied (highest priority) or that have not beenvaried for one or more cycles. If we find that varyingall the residues except those varied in the previouscycle does not allow a high enough level of diversity,then residues varied in the previous cycle might bevaried again. For example, if the number ofindependent transformants that can be produced and thesensitivity of the affinity separation were such thatseven residues could be varied, and if the principaland secondary sets contained 13 residues, we wouldalways vary seven residues, even though that impliesvarying some residue twice in a row. In such cases, we 115 would pick the residues just varied that contain theamino acids of highest abundance in the variegatedcodons used.
It is the accumulation of information that allowsthe process to select those protein sequences thatproduce binding between the SBD and the target. Someinterfaces between proteins and other molecules involvetwenty or more residues. Complete variation of twentyresidues would generate 1026 different proteins. Bydividing the residues that lie close together in spaceinto overlapping groups of five to seven residues, wecan vary a large surface but never need to test morethan 107 to 109 candidates at once, a savings of 1019to 1017 fold.
Having picked the residues to vary, we again setthe range of variegation for each residue according tothe principles set forth in 13.2, design the vgDNAencoding the desired mutants (Sec. 13.3), clone thevgDNA into GPs (Sec. 14) , and select-by-binding-to-target those GPs bearing SBDs (Sec. 15).
Sec. 17.0: OTHER CONSIDERATIONS:
Sec. 17.1: Joint selections:
One may modify the affinity separation of themethod described to select a molecule that binds tomaterial A but not to material B. One needs to preparetwo selection columns, one with material A and theother with material B. The population of geneticpackages is prepared in the manner described, butbefore applying the population to A, one passes thepopulation over the B column so as to remove those •· 116
members of the population that have high affinity forB. It may be necessary to amplify the population thatdoes not bind to B before passing it over A.Amplification would most likely be needed if A and B 5 were in some ways similar and the PPBD has beenselected for having affinity for A.
For example, to obtain an SBD that binds A but notB, three columns could be connected in series: a) a 10 column supporting some compound, neither A nor B, oronly the matrix material, b) a column supporting B, andc) a column supporting A. A population of GP(vgPBD)sis applied to the series of columns and the columns arewashed with the buffer of constant ionic strength that 15 is used in the application. The columns are uncoupled,and the third column is eluted with a gradient toisolate GP(PBD)s that bind A but not B.
One can also generate molecules that bind to both 20 A and B. In this case we use a 3D model and mutate oneface of the molecule in question to get binding to A.We then mutate a different face to produce binding to B. 25 The materials A and B could be proteins that differ at only one or a few residues. For example, Acould be a natural protein for which the gene has beencloned and B could be a mutant of A that retains theoverall 3D structure of A. SBDs selected to bind A but 30 not B must bind to A near the residues that are mutatedin B. If the mutations were picked to be in the activesite of A (assuming A has an active site), then an SBDthat binds A but not B will bind to the active site ofA and is likely to be an inhibitor of A. 35
<img img-format="tif" img-content="drawing" file="IL120940AD00029.tif" id="idf0009" />
117
To obtain a protein that will bind to both A andB, we can, alternatively, first obtain an SBD thatbinds A and a different SBD that binds B. We can thencombine the genes encoding these domains so that a two- 5 domain single-polypeptide protein is produced. Thefusion protein will have affinity for both A and B.
One can also generate binding proteins withaffinity for both A and B, such that these materials 10 compete for the same site on the binding protein. Weguarantee competition by overlapping the sites for Aand B. We first create a molecule that binds to targetmaterial A. We then vary a set of residues defined as:a) those residues that were varied to obtain binding to 15 A, plus b) those residues close in 3D space to theresidues of set (a) but that are internal and so areunlikely to bind directly to either A or B. Residuesin set (b) are likely to make small changes in thepositioning of the residues in set (a) Such that the 20 affinities for A and B will be changed by smallamounts. Members of these populations are selected foraffinity to both A and B.
Sec. 17.2: Selection for non-binding: 25
The method of the present invention can be used toselect proteins that do not bind to selected targets.Consider a protein of pharmacological importance, suchas streptokinase, that is antigenic to an undesirable 30 extent. We can take the pharmacologically importantprotein as IPBD and antibodies against it as target.Residues on the surface of the pharmacologicallyimportant protein would be variegated and GP(PBD)s thatdo not bind to an antibody column would be collected 35 and cultured. Surface residues may be identified in 118 several ways, including: a) from a 3D structure, b)from hydrophobicity considerations, or c) chemicallabeling. The 3D structure of the pharmacologicallyimportant protein remains the preferred guide to 5 picking residues to vary, except now we pick residuesthat are widely spaced so that we leave as little aspossible of the original surface unaltered.
Destroying binding frequently requires only that a10 single amino acid in the binding interface be changed.If polyclonal antibodies are used, we face the problemthat all or most of the strong epitopes must be alteredin a single molecule. Preferably, one would have a setof monoclonal antibodies, or a narrow range of antibody 15 species. If we had a series of monoclonal antibodycolumns, we could obtain one or more mutations thatabolish binding to each monoclonal antibody. We couldthen combine some or all of these mutations in onemolecule to produce a pharmacologically important 20 protein recognized by none of the monoclonalantibodies. Such mutants must be tested to verify thatthe pharmacologically interesting properties have notbe altered to an unacceptable degree by the mutations. 25 Typically, polyclonal antibodies display a range of binding constants for antigen. Even if we have onlypolyclonal antibodies that bind to thepharmacologically important protein, we may proceed asfollows. We engineer the pharmacologically important 30 protein to appear on the surface of a replicable GP.We introduce mutations into residues that are on thesurface of the pharmacologically important protein orinto residues thought to be on the surface of thepharmacologically important protein so that a 35 population of GPs is obtained. Polyclonal antibodies 119 are attached to a column and the population of GPs isapplied to the column at low salt. The column iseluted with a salt gradient. The GPs that elute at thelowest concentration of salt are those which bearpharmacologically important proteins that have beenmutated in a way that eliminates binding to theantibodies having maximum affinity for thepharmacologically important protein. The GPs elutingat the lowest salt are isolated and cultured. Theisolated SBD becomes the PPBD to further rounds ofvariegation so that the antigenic determinants aresuccessively eliminated.
Sec. 17.3:_Selection of PBDs for retention of structure:
We can select for insertions or deletions thatpreserve the 3D structure of known binding proteins.Consider on GP that express BPTI on its surface. Inthe bpti-osp gene, we can replace the codons for K26and A27 with five variegated codons (3.2 x 106sequences). K26 and A27 are in a turn and are far fromthe trypsin binding surface. We use selection-through-binding to isolate GPs expressing mutants of BPTI thatretain high, specific affinity for trypsin.
Sec. 17.4: Created binding proteins not unique:
For each target, there are a large number of SBDsthat may be found by the method of the presentinvention. To increase the probability that some PBDin the population will bind to the target, we generateas large a population as we can conveniently subject toselection-through-binding. Key questions in managementof the method are "How many transformants can we 120 produce?", and "How small a component can we findthrough selection-through-binding?". Geneticists routinely find mutations with frequencies of one in1010 using simple, powerful selections. The optimum 5 level of variegation is determined by the maximumnumber of transformants and the selection sensitivity,so that for any reasonable sensitivity we may use aprogressive process to obtain a series of proteins withhigher and higher affinity for the chosen target 10 material. Enrichments of 1000-fold by a single pass ofelution from an affinity plate have been demonstrated(SMIT85).
Use of different variation schemes can yield 15 different binding proteins. For any given target, alarge plurality of proteins will bind to it. Thus, ifone binding protein turns out to be unsuitable for somereason (e.g. too antigenic), the procedure can berepeated with different variation parameters. For 20 example, one might choose different residues to vary orpick a different nt distribution at variegated codonsso that a new distribution of amino acids is tested atthe same residues. Even if the same principal set ofresidues is used, one might obtain a different SBD if 25 the order in which one picks subsets to be varied isaltered.
Sec. 17.5: Other modes of mutagenesis possible: 30 The modes of creating diversity in the population of GPs discussed herein are not the only modespossible. Any method of mutagenesis that preserves atleast a large fraction of the information obtained fromone selection and then introduces other mutations in 35 the same domain will work. The limiting factors are 121 the number of independent transformants that can beproduced and the amount of enrichment one can achievethrough affinity separation. Therefore the preferredembodiment uses a method of mutagenesis that focuses 5 mutations into those residues that are most likely toaffect the binding properties of the PBD and are leastlikely to destroy the underlying structure of the IPBD.
Other modes of mutagenesis might allow other GPs10 to be considered. For example, the bacteriophagelambda is not a useful cloning vehicle for cassettemutagenesis because of the plethora of restrictionsites. One can, however, use single-stranded-oligo-nt-directed mutagenesis on lambda without the need for 15 unique restriction sites. No one has used single-strarided-oligo-nt-directed mutagenesis to introduce thehigh level of diversity called for in the presentinvention, but if it is possible, such a method wouldallow use of phage with large genomes. 20 122
Example 1 BPTI-Derived Binding Protein for HHMb; Displayed by M13
Phage
Presented below is a hypothetical example of aprotocol for developing a new binding molecule derivedfrom BPTI with affinity for horse heart myoglobin(HHMb) using the common E. coli bacteriophage M13 asgenetic package. It will be understood that somefurther optimization, in accordance with the teachingsherein, may be necessary to obtain the desired results.Possible modifications in the preferred method arediscussed immediately following various steps of thehypothetical example.
By hypothesis, we set the following technicalcapabilities:
Yqq 500 ng/synthesis of ssDNA 100 bases long, 10 ug/synthesis of ssDNA 60 bases long,1 mg/synthesis of ssDNA 20 bases long.
mDNA 100 bases
<img img-format="tif" img-content="drawing" file="IL120940AD000210.tif" id="idf0010" />
Lef 1 mg/1 0.1 % for blunt-blunt, 4 % for sticky-blunt, 11 % for sticky-sticky.
Mntv 5 x 108 123
Ceff 900-fold enrichment csensi 1 4 x 10®
Nchrom 10 passes
Sg£-£· 0.05
Example 1, Part I
In this example, we will use M13 as a replicableGP and BPTI as IPBD. In Part I, we are concerned onlywith getting BPTI displayed on the outer surface of anM13 derivative. Variable DNA may be introduced in theosp-ipbd gene, but not within the region that codes forthe trypsin-binding region of BPTI. Once BPTI isdisplayed on the M13 outer surface of an M13derivative, we proceed to Part II to optimize theaffinity separation procedures.
For this example, we choose a filamentousbacteriophage of E. coli. M13. We prefer phage overvegetative bacterial cells because phage are much lessmetabolically active. We prefer phage over sporesbecause the molecular mechanisms of the virionformation and 3D structure of the virion are muchbetter understood than are the corresponding processesof spore formation and structures of spores. M13 is a very well studied bacteriophage, widelyused for DNA sequencing and as a genetic vector; it isa typical member of the class of filamentous phages.The relevant facts about M13 and other phages that willallow us to choose among phages are cited in Sec. 124 1.3.1.
Compared to other bacteriophage, filamentous phagein general are attractive and M13 in particular isespecially attractive because: 1) the 3D structure of the virion is known, 2) the processing of the coat protein is wellunderstood, 3) the genome is expandable, 4) the genome is small, 5) the sequence of the genome is known, 6) the virion is physically resistant to shear,heat, cold, guanidinium Cl, low pH, and high salt, 7) the phage is a sequencing vector so thatsequencing is especially easy, and 8) antibiotic-resistance genes have been clonedinto the genome with predictable results (HINE80).
Other criteria listed in Sec. 1.0 and 1.3 of the arealso satisfied: M13 is easily cultured and stored(FRIT85) , each infected cell yielding 100 to 1000 M13progeny after infection. M13 has no unusual orexpensive media requirements and is easily harvestedand concentrated (SALI64, YAMA70, FRIT85). M13 is stable toward physical agents: temperature (10% ofphage survive 30 minutes at 85°C), shear (Waringblender does not kill), desiccation (not applicable), 125 radiation (not applicable), age (stable for years). M13 is stable toward chemicals: pH (< 2.2(SMIT85)), surface active agents: not applicable,chaotropes (guanidinium HCI = 6.0 M), ions (no specificsensitivities), organic solvents (ether and otherorganic solvents are lethal (MARV78)), proteases (notapplicable, HHMb not a protease) . M13 is not known tobe sensitive to other enzymes. M13 genome is 6423 b.p. and the sequence is known(SCHA78). Because the genome is small, cassettemutagenesis is practical on RF M13 (AUSU87), as issingle-stranded oligo-nt directed mutagenesis (FRIT85).M13 is a plasmid and transformation system in itself,and an ideal sequencing vector. M13 can be grown onRec“ strains of E. coli. The M13 genome is expandable(MESS78, FRIT85). M13 confers no advantage, butdoesn't lyse cells. The sequence of gene VIII isknown, and the amino acid sequence can be encoded on asynthetic gene, using lacUV5 promoter and used inconjunction with the Lacl^ repressor. The lacUV5promoter is induced by IPTG. Gene VIII protein issecreted by a well studied process and is cleavedbetween A2 3 and A24. Residues 18, 21, 22, and 23 ofgene VIII protein control cleavage. Mature gene VIIIprotein makes up the sheath around the circular ssDNA.The 3D structure of fl virion is known at mediumresolution; the amino terminus of gene VIII protein ison surface of the virion. No fusions to M13 gene VIIIprotein have been reported. The 2D structure of M13coat protein is implicit in the 3D structure. MatureM13 gene VIII protein has only one domain. There arefour minor proteins: gene III, VI, VII, and IX. Eachof these minor proteins is present in about 5 copies 126 per virion and is related to morphogenesis orinfection. The major coat protein is present in morethan 2500 copies per virion.
Although no fusions of M13 gene VIII to othergenes have been reported, knowledge of the virion 3Dstructure (BANN810) makes attachment of IPBD to theamino terminus of mature M13 coat protein (M13 CP)quite attractive. Should direct fusion of BPTI to M13CP fail to cause BPTI to be displayed on the surface ofM13, we will vary part of the BPTI sequence and/orinsert short random DNA sequences between BPTI and M13CP.
Smith (SMIT85) and de la Cruz et al. (CRUZ88) haveshown that insertions into gene III cause novel proteindomains to appear on the virion outer surface. If BPTIcan not be made to appear on the virion outer surfaceby fusing the bpti gene to the m!3cp gene, we will fusebpti to gene III either at the site used by Smith andby de la Cruz et al. or to one of the termini. We willuse a second, synthetic copy of gene III so that someunaltered gene III protein will be present.
The gene VIII protein is chosen as OSP because itis present in many copies and because its location andorientation in the virion are known. Note that anyuncertainty about the azimuth of the coat protein aboutits own alpha helical axis is unimportant.
The 3D model of fl indicates strongly that fusingBPTI to the amino terminus of M13 CP is more likely toyield a functional protein than any other fusion site.(See Sec. 1.3.3). 127
The amino-acid sequence of M13 pre-coat (SCHA78),called AA_seql, is AA_seql 1 1 2 I 12 3 3 4 4 5 5 0 5 0 V5 0.5 0 5 0
MKKSLVLKASVAVATLVPMLSFAAEGDDPAKAAFNSLQASATEYIGYAWA 5 6 6 7 75 0 5 0 3
MVWIVGATIGIKLFKKFTSKAS
The single-letter codes for amino acids and the codesfor ambiguous DNA are internationally recognized(GEOR87). The best site for inserting a novel proteindomain into M13 CP is after A23 because SP-I cleavesthe precoat protein after A23, as indicated by thearrow. Proteins that can be secreted will appearconnected to mature M13 CP at its amino terminus.Because the amino terminus of mature M13 CP is locatedon the outer surface of the virion, the introduceddomain will be displayed on the outside of the virion. BPTI is chosen as IPBD of this example (See Sec,.2.1) because it meets or exceeds all the criteria: itis a small, very stable protein with a well known 3Dstructure. Marks et al. (MARK86) have shown that afusion of the phoA signal peptide gene fragment and DNAcoding for the mature form of BPTI caused native BPTIto appear in the periplasm of E. coli. demonstratingthat there is nothing in the structure of BPTI toprevent its being secreted.
Marks et al. (MARK87) also showed that the* structure of BPTI is stable even to the removal of oneof the cystine bridges. They did this by replacingboth C14 and C38 with either two alanines or two 128 threonines. The C14/C38 cystine bridge that Marks etal. removed is the one very close to the scissile bondin BPTI; surprisingly, both mutant moleculesfunctioned as trypsin inhibitors. This indicates thatBPTI is redundantly stable and so is likely to foldinto approximately the same structure despite numeroussurface mutations. Using the knowledge of homologues,vide infra, we can infer which residues must not bevaried if the basic BPTI structure is to be maintained.
The 3D structure of BPTI has been determined athigh resolution by X-ray diffraction (HUBE77, MARQ83,WLOD84, WLOD87a, WLOD87b), neutron diffraction(WLOD84) , and by NMR (WAGN87) . In one of the X-raystructures deposited in the Brookhaven Protein DataBank, "6PTI’', there was no electron density for A58,indicating that A58 has no uniquely definedconformation. Thus we know that the carboxy group doesnot make any essential interaction in the foldedstructure. The amino terminus of BPTI is very near tothe carboxy terminus. Goldenberg and Creightonreported on circularized BPTI and circularly permutedBPTI (GOLD83). Some proteins homologous to BPTI havemore or fewer residues at either terminus. BPTI has been called "the hydrogen atom of proteinfolding” and has been the subject of numerousexperimental and theoretical studies (STAT87, SCHW87,GOLD83, CHAZ83). BPTI has the added advantage that at least 32homologous proteins are known, as shown in Table 13. Atally of ionizable groups is shown in Table 14 and thecomposite of amino acid types occurring at each residueis shown in Table 15. 129 BPTI is freely soluble and is not known to bindmetal ions. BPTI has no known enzymatic activity.BPTI binds to trypsin, Kd = 6.0 x 10-14 M (TSCH87) .BPTI is not toxic. If K15 of BPTI is changed to L,there is no measurable binding between the mutant BPTIand trypsin (TSCH87).
All of the conserved residues are buried; of theseven fully conserved residues only G37 has noticeableexposure. The solvent accessibility of each residue inBPTI is given in Table 16 which was calculated from theentry "6PTI” in the Brookhaven Protein Data Bank with asolvent radius of 1.4 A, the atomic radii given inTable 7, and the method of Lee and Richards (LEEB71) .Each of the 51 non-conserved residues can accommodatetwo or more kinds of amino acids. By independentlysubstituting at each residue only those amino acidsalready observed at that residue, we could obtainapproximately 7 x 1042 different amino acid sequences,most of which will fold into structures very similar toBPTI. BPTI will be useful as a IPBD for macromolecules.(See Sec. 2.1.1) BPTI and BPTI homologues bind tightlyand with high specificity to a number of enzymes. BPTI is strongly positively charged except at veryhigh pH, thus BPTI is useful as IPBD for targets thatare not also strongly positive under the conditions ofintended use (see Sec. 2.1.2). There exist homologuesof BPTI, however, having quite different charges (viz.SCI-III from Bombvx mori at -7 and the trypsininhibitor from bovine colostrum at -1). Once aderivative of M13 is found that displays BPTI on its 130 surface, the sequence of the BPTI domain can bereplaced by one of the homologous sequences to produceacidic or neutral IPBDs. BPTI is not an enzyme (See Sec. 2.1.3). BPTI isquite small; if this should cause a pharmacologicalproblem, two or more BPTI-derived domains may be joinedas in the human BPTI homologue that has two domains. A derivative of M13 is the preferred OCV. (SeeSec. 3). A "phagemid” is a hybrid between a phage anda plasmid, and is used in this invention. Double-stranded plasmid DNA isolated from phagemid-bearingcells is denoted by the standard convention, e.g.pXY24. Phage prepared from these cells would bedesignated XY24. Phagemids such as Bluescript K/S(sold by Stratagene) are not suitable for our purposesbecause Bluescript does not contain the full genome ofM13 and must be rescued by coinfection with helperphage. Such coinfections could lead to geneticrecombination yielding heterogeneous phage unsuitablefor the purposes of the present invention.
The bacteriophage M13 bla 61 (ATCC 37039) isderived from wild-type M13 through the insertion of thebeta lactamase gene (HINE80). This phage contains 8.13kb of DNA. M13 bla cat 1 (ATCC 37040) is derived fromM13 bla 61 through the additional insertion of thechloramphenicol resistance gene (HINE80); M13 bla cat 1contains 9.88 kb of DNA. Although neither of thesevariants of M13 contains the ColEl origin ofreplication, either could be used as a starting pointto construct a usable cloning vector for the presentexample. 131
The OCV for the current example is constructed bya process illustrated in Figure 4. A brief descriptionof all the plasmids and phagemids constructed for thisExample is found in Table 17.
For ss oligo-nt site-directed mutagenesis,multiple primers lead to higher efficiency. Three non-mutagenic primers are used: bases 2326-2352 of wt M13,bases 4854-4875 of wt M13, and the complement of bases3431-3451 of pBR322. Note that pLG2 and itsderivatives carry the anti-sense strand of the ampRgene in the + DNA strand. The segments are picked tobe high in GC content and to divide the pLG7 genomeinto several segments of approximately equal length.
The genetic engineering procedures needed toconstruct the OCV are standard, using commerciallyavailable restriction enzymes under recommendedconditions. All restriction fragments of DNA arepurified by electrophoresis or HPLC. M13 and itsengineered derivatives are infected into E. coli strainPE384 (F+,Rec", Sup+,Amps) . Plasmid DNA of M13derivatives is transformed into E. coli strain PE383(F“,Rec",Sup+,Amps) so that we avoid multiple rounds ofinfection in the culture. Isolation of M13 phage is bythe procedure of Salivar et al. (SALI64); isolation ofreplicative fora (RF) M13 is by the procedure ofJazwinski et al. (JAZW73a and JAZW73b). Isolation ofplasmids containing the ColEl origin of replication isby the method of Maniatis (MANI82).
We pick the ampR gene from pBR322 as a convenientantibiotic resistance gene. Another resistance gene,such as kanamycin, could be used. The Acc I-to-Aat IIfragment of pBR322 is a conveniently obtained source of •Γ’χ 132 anvR and the Col El origin. M13mpl8 (New England BioLabs) contains neither AatII nor Acc I sites. Therefore we insert an adaptorthat allows us to insert the Aat II-to-Acc I fragmentof pBR322 that carries the ampR gene and the ColElorigin of replication into a desirable place inM13mpl8. M13mpl8 contains a lacUV5 promoter and a lacZgene that are not useful to the purposes of the presentinvention. By cutting M13mpl8 with Avail and Bsu36Iand discarding the approximately 600 intervening basepairs, we eliminate all recognition sites of severalenzymes useful for engineering the bpti-gene VIII gene.
The following adaptor is synthesized, 5’ GACCGACGTCtgcctcGTATACCGGACCGcatagctCC 3' olig#l3' GCTGCAGacggagCATATGGCCTGGCgtatcgaGGACT 5' olig#2
Avail IAatll| [AccijRsrII I lBsu36I
The annealed adaptor is ligated with RF M13mpl8that has been cut with both Avail and Bsu36I andpurified by PAGE or HPLC. Transformed cells areselected for plasmid uptake with ampicillin. Theresulting construct is called pLGl. DNA from pLGl is cut with both Aat II and Acc I.Aatll-to-AccI fragment of pBR322 is ligated to thebackbone of LG1. The correct construct is named pLG2.
The Acc I restriction site is no longer needed forvector construction. To eliminate this site, RF pLG2dsDNA is cut with Acc I, treated with Klenow fragmentand dATP and dTTP to make it blunt and then religated.The cloning vector, named pLG3, is now ready forstepwise insertion of the osp-ipbd gene. 133
We are now ready to design a gene (See Sec. 4)that will cause BPTI-domains to appear on the outersurface of an M13 derivative: LG7.
To obtain a novel protein domain attached to theoutside of M13, we insert DNA that codes for matureBPTI after A23 of the precoat protein of M13. MatureBPTI begins with an arginine residue, which is charged;cleavage by signal peptidase I is normal in such cases.Signal peptidase I (SP-I) cuts a chimera of M13 coatprotein and BPTI after A23 leaving mature BPTI attachedat its carboxy end to the amino terminus of M13 CP.
The following amino-acid sequence, called AA_seq2,is constructed, by inserting the sequence for matureBPTI (shown underscored) immediately after the signalsequence of M13 precoat protein (indicated by thearrow) and before the sequence for the M13 CP. AA_seq2 1 1 2112 3 3 4 4 55050 V5 05050
MKKS LVLKASVAVATLVPMLS FARPPFCLEPPYTGPCKARIIRYFYNAKA 5 6 6 7 7 8 8 9 9 10 5 0 5 0 5 0 5 0 5 0
GLCOTFVYGGCRAKRNNFKSAEDCMRTCGGAAEGDDPAKAAFNSLQASAT 10 11 11 12 12 13 5 0 5 0 5 0
EYIGYAWAMWVIVGATIGIKLFKKFTSKAS
Sequence numbers of fusion proteins refer to thefusion, as coded, unless otherwise noted. Thus thealanine that begins M13 CP is referred to as "number 134 82", "number 1 of M13 CP", or "number 59 of the matureBPTI-M13 CP fusion".
The osp-ipbd gene is regulated by the lacUV5promoter and terminated by the trpA transcriptionterminator. The host strain of E. coli harbors thelaclS gene. The osp-ipbd gene is expressed andprocessed in parallel with the wild-type gene VIII.The novel protein, that consists of BPTI tethered to aM13 CP domain, constitutes only a fraction of the coat.Affinity separation is able to separate phage carryingonly five or six copies of a molecule that has highaffinity for an affinity matrix (SMIT85) ; 1%incorporation of the chimeric protein results in about30 copies of the protein exposed on the surface. Ifthis is insufficient, additional copies may be providedby, for example, increasing IPTG. A model comprising M13 coat, after the model forfl of Marvin and colleagues (BANN81) , and a BPTIdomain, taken from the Brookhaven Protein Data Bankentry "6PTI", was constructed by standard modelbuilding methods that insure that covalent bond lengthsand angles are close to acceptable values. The modelshows that the fusion protein could fit into thesupramolecular structure in a stereochemicallyacceptable fashion without disturbing the internalstructure of either the M13 CP or BPTI domain.
The ambiguous DNA sequence coding for AA_seq2, isexamined by a computer program for places whererecognition sites for restriction enzymes could becreated without altering the amino-acid sequence. (SeeSec. 4.3). A master table of enzymes is compiled fromthe catalogues of enzyme suppliers. The enzymes that 135 do not cut the OCV. (Preferably constructed asdescribed above).
Using the procedure given in Sec. 4.3, we design aipbd gene, such as that shown in Table 25. Somerestriction enzymes (e.g. Ban I or Hph I) cut the OCVtoo often to be of value.
The entire DNA sequence of the m!3cp-bpti fusionwith annotation appears in Table 25 showing the usefulrestriction sites and biologically important features,viz. the lacUV5 promoter, the lacO operator, the Shine-Dalgarno sequence, the amino acid sequence, the stopcodons, and the transcriptional terminator.
The ipbd gene is synthesized in several stepsusing the method described in Sec. 5.1, generatingdsDNA fragments of 150 to 190 base pairs.
The four steps (See Sec. 6.1) by which we clonesynthetic fragments of the m!3cp-bpti gene (the osp-ipbd gene of the present example) into pLG3 and itsderivatives are illustrated in Figure 5.
The sequence to be introduced into pLG3 comprisesa) the segment from RsrII to Avril (Table 25), b) aspacer sequence (gccgctcc), and c) the segment fromAsuII to Saul. The segment is 158 bases long and issynthesized from two shorter synthetic oligo-nts asdescribed in Sec. 5.1 of the generic specification.
Table 27 shows the antisense strand of thesequence to be inserted. The 99 base fragment shown inupper case letters and underscored (5' —CCGTCC....CCTTCG-3’ = olig#3) is synthesized in the 136 standard manner. Similarly, the 100 base long fragmentof the sense strand shown in lower case (5'-cgctca....aattg-3' = olig#4) is synthesized. Afterannealing, the double-stranded region is extended withKlenow fragment by the procedure given above to makethe entire 176 bases double stranded. The overlapregion is 23 base pairs long and contains 14 CG pairsand 9 AT pairs. The DNA between Avril and AsuII doesnot code for anything in the final pbd gene; it isthere so that the DNA can be cut by both Avril andAsuII at the same time in the next step. Eight baseshave been added to the left of RsrII and nine baseshave been added to the left of Saul (same specificityand' cutting pattern as Bsu36I). These bases at theends are not part of the final product; they must bepresent so that the restriction enzymes can bind andcut the synthetic DNA to produce specific sticky ends.
The synthetic DNA is cut with both Saul and RsrIIand is ligated to similarly cut dsDNA of pLG3. Theconstruct with the correct insert is called pLG4.
The second step of the construction of the OCV isillustrated in Table 28. As in the construction ofpLG4, two pieces of single-stranded DNA aresynthesized: a 99 base long fragment of the anti-sensestrand ending with p25 and a 99 base long fragment(starting with pl8). Both the synthetic dsDNA and dsRFpLG4 DNA are cut with both Avril and AsuII and areligated and used to transform E. coli. The constructcarrying this second insert is called pLG5.
Construction of pLG6 proceeds similarly to theconstruction of pLG5. The sequence is shown in Table30. The two single stranded segments (one from the 137 anti-sense strand ending with N66 and the other fromthe sense strand starting with the third base of thecodon for Y58) are synthesized, annealed, and extendedwith Klenow fragment. Both the synthetic DNA and RFpLG5 are cut with both BssHI and AsuII. purified, andthe appropriate pieces are ligated and used totransform E. coli.
The construction of pLG7 is illustrated in Table32 and proceeds similarly to the constructions of pLG4,pLG5, and pLG6. The two single stranded segments (onefrom the anti-sense strand ending with the first baseof the codon for V110 and the other beginning withE101) are synthesized, annealed, and extended withKlenow fragment. Both the synthetic DNA and RF pLG6are cut with both Bbel and AsuII. purified, and theappropriate pieces are ligated and used to transform E.coli. The construct with the correct fourth insert iscalled pLG7; the display of BPTI on the outer surfaceof LG7 is verified by the methods of Sec. 8. M13am429 is an amber mutation of M13 used toreduce non-specific binding by the affinity matrix forphages derived from M13. M13am429 is derived bystandard genetic methods (MILL72) from wtM13.
Phage LG7 is grown on E. coli strain PE384 in LBbroth with various concentrations of IPTG added to themedium to induce the osp-ipbd gene. Phage LG7 isobtained from cells grown with 0.0, 0.1, 1.0, 10.0 or100.0 uM, or 1.0 mM IPTG, harvested (See Sec. 7) by themethod of Salivar (SALI64), and concentrated to obtaina titre of 1012 pfu/ml by the method of Messing(MESS83). 138
The preferred method of determining whether LG7displays BPTI on its surface (See Sec. 8) is todetermine whether these phage can retain a labeledderivative of trypsin (trp) or anhydrotrypsin (AHTrp)on a filter that allows passage of unbound trp orAHTrp. Trypsin contains 10 tyrosine residues and canbe iodinated with 125I by standard methods; we denotethe labeled trypsin as "trp*". Labeled anhydrotrypsinis denoted as "AHTrp*". Other types of labels can beused on trp or AHTrp, e.g. biotin or a fluorescentlabel. AHTrp* or trp* is labeled to an activity of 0.3uCi/ug. A sample of 1012 LG7(10 mM IPTG) is mixed with1.0 ug of trp* or AHTrp* in 1.0 ml of a buffer of 10 mMKCl, adjusted to pH 8.0 with 1 mM K2HPO4 / KH2PO4. Themixture is passed through an Amicon MSP1 system fittedwith a membrane filter that allows passage of proteinssmaller that Mr = 300,000. Filters are soaked inbuffer containing trp or AHTrp prior to the analysis.The filter is washed twice with 0.5 ml of buffercontaining trp or AHTrp. The radioactivity retained onthe filter is quantitated with a scintillation counteror other suitable device. If each virion displays onecopy of BPTI, then .05 ug of protein can be bound thatwould give rise to 3 χ 104 disintegrations / minute onthe filter.
An alternative way to quantitate display of BPTIon the surface of LG7 is to use the stoichiometricbinding between trypsin and BPTI to titrate the BPTI.A solution that titers 1012 pfu/ml of a phage isapproximately 1.6 χ 10-9 M in phage if each virion isinfective. The ratio of pfu to total phage can bedetermined spectrophotometrically using the molarextinction coefficients at 260 nm and 280 nm correctedfor the increased length of LG7 as compared to wtM13. 139
For example, if a 1.0 ml solution that contains 1012pfu of LG7 phage grown with 1.0 mM IPTG inhibitstrypsin solutions up to 4.8 x 10”7 M, we calculate thatthere are approximately 300 BPTIs/GP (i.e. (4.8 x 1O“7molecules of BPTI/1)/(1.6 x 10“9 phage/1)). Inhibitionof a specified concentration of trypsin is most easilymeasured spectrophotometrically using a peptide-linkeddye, such as Naip^a-benzoyl-Arg-Nan (TSCH87).
Alternatively, binding to an affinity column maybe used to demonstrate the presence of BPTI on thesurface of phage LG7. An affinity column of 2.0 mltotal volume having BioRad Affi-Gel 10(™) matrix and30 mg of AHTrp as affinity material is prepared by themethod of BioRad. The void volume (Vy) of this columnis, by hypothesis, 1.0 . ml. This affinity column isdenoted {AHTrp}. A sample of 1012 M13am429 is applied to {AHTrp} in1.0 ml of 10 mM KCI buffered to pH 8.0 with KH2PO4 /K2HPO4. The column is then washed with the same bufferuntil the optical density at 280 nm of the effluentreturns to base line or 4 x Vy have been passed throughthe column, whichever comes first. Samples of LG7 orLG10 are then applied to the blocked {AHTrp} column atIO3·2 pfu/ml in 1.0 ml of the same buffer. The columnis then washed again with the same buffer until theoptical density at 280 nm of the effluent returns tobase line or 4 x Vy have been passed through, whichevercomes first. Following this wash, a gradient of KCIfrom 10 mM to 2 M in 3 x Vy, buffered to pH 8.0 withphosphate is passed over the column. The first KCIgradient is followed by a KCI gradient running from 2 Mto 5 M in 3 x Vv. The second KCI gradient is followedby a gradient of guanidinium Cl. from 0.0 M to 2.0 M in 140 2 x Vy in 5 M KCl and buffered to pH 8.0 withphosphate. Fractions of 50 ul are collected andassayed for phage by plating 4 ul of each fraction atsuitable dilutions on sensitive cells. Retention ofphage on the column is indicated by appearance of LG7phage in fractions that elute significantly later fromthe column than control phage LG10 or wtM13. Asuccessful isolate of LG7 that displays BPTI isidentified, the bpti insert and junctions aresequenced, and this isolate is used for further workdescribed below.
If vgDNA is used to obtain a functional fusionbetween a BPTI mutant and M13 CP (vide infra), then DNAfrom a clonal isolate is sequenced in the regions thatwere variegated. Then gratuitous restriction sites foruseful restriction enzymes are removed if possible bysilent codon changes. The sequence numbers of residuesin OSP-IPBD will be changed by any insertions;hereinafter, we will, however, denote residues insertedafter residue 23 as 23a, 23b, etc. Insertions afterresidue 81 will be denoted as 81a, 81b, etc. Thispreserves the numbering of residues between C5 and C55of BPTI. Residue C5 of BPTI is always denoted as 28 inthe fusion; residue C55 of BPTI is always denoted as 78in the fusion, and the intervening residues haveconstant numbers.
Should LG7 phage from cells grown with 10 mM IPTGfail to display BPTI on its surface, we have severaloptions. We might try to determine why theconstruction failed to work as expected. There arevarious possible modes of failure, including : a) BPTIis not cleaved from the M13 signal sequence, b) BPTI iscleaved from the M13 CP, and c) the chimeric protein is 141 made and cleaved after the signal sequence, but theprocessed protein is not incorporated into the M13coat. BPTI has been secreted from E. coli (MARKS6) ;however the M13 coat-protein signal sequence was notused. Therefore problems stemming from the signalsequence are unlikely, but possible. We coulddetermine whether BPTI was present in the periplasm orbound to the inner membrane of LG7-infected cells byassays using try* or Antry*.
Proteins in the periplasm can be freed throughspheroplast formation using lysozyme and EDTA in aconcentrated sucrose solution (BIRD67, MALA64). IfBPTI were free in the periplasm, it would be found inthe supernatant. Try* would be mixed with supernatantand passed over a non-denaturing molecular sizingcolumn and the radioactive fractions collected. Theradioactive fractions would then be analyzed by SDS-PAGE and examined for BPTI-sized bands by silverstaining.
Spheroplast formation exposes proteins anchored inthe inner membrane. Spheroplasts are mixed with AHTrp*and then either filtered or centrifuged to separatethem from unbound AHTrp*. After washing withhypertonic buffer, the spheroplasts are analyzed forextent of AHTrp* binding alternatively, membraneproteins are analyzed by western blot analysis.
If BPTI is found free in the periplasm, then wewould expect that the chimeric protein was beingcleaved both between BPTI and the M13 mature coatsequence and between BPTI and the signal sequence. Inthat case, we should alter the BPTI/M13 CP junction byinserting vgDNA at codons for residues 78-82 of 142 AA_seq2.
If BPTI is found attached to the inner membrane,then there are two likely explanations. The first isthat the chimeric protein is being cut after the signalsequence, but is not being incorporated into LG7virion? the treatment would also be to insert vgDNAbetween residues 78 and 82 of AA_seg2. The alternativehypothesis is that BPTI could fold and react withtrypsin even if signal sequence is not cleaved. N-terminal amino acid sequencing of trypsin-bindingmaterial isolated from cell homogenate determines whatprocessing is occurring. If signal sequence were beingcleaved, we would use the procedure above to varyresidues between C78 and A82; subsequent passes wouldadd residues after residue 81. If signal sequence werenot being cleaved, we would vary residues between 23and 27 of AA_seq2. Subsequent passes through thatprocess would add residues after 23.
If BPTI were found neither in the periplasm nor onthe inner membrane, then we would expect that the faultwas in the signal sequence or the signal-sequence-to-BPTI junction. The treatment in this case would be tovary residues between 23 and 27.
Several experiments that introduce variegationinto the bpti-gene VIII fusion are possible, including: 1) 3 variegated codons between residues 78 and 82using olig#12 and olig#13, 2) 3 variegated codons between residues 23 and 27using olig#14 and olig#15, 143 3) 5 variegated codons between residues 78 and 82using olig#13 and olig#12a, 4) 5 variegated codons between residues 23 and 27 5 using olig#15 and olig#14a, 5) 7 variegated codons between residues 78 and 82using olig#13 and olig#12b, and 10 6) 7 variegated codons between residues 23 and 27 using olig#15 and olig#14b.
To alter the BPTI-M13 CP junction, we introduceDNA variegated at codons for residues between 78 and 82 15 into the Sph I and Sfi I sites of pLG7. The residuesafter the last cysteine are highly variable in aminoacid sequences homologous to BPTI, both in compositionand length; in Table 25 these residues are denoted asG79, G80, and A81. The first part of the M13 CP is 20 denoted as A82, E83, and G84. One of the oligo-nts olig#12, olig#12a, or olig#12b and the primer olig#13are synthesized by standard methods. The oligo-ntsare: 25 residue 75 76 77 78 79 80 81 82 83 5’ gc|gag|cGC|ATG|CGT|ACC|TGC|qfk|qfk|qfk|GCT|GAA|- 84 85 86 87 88 89 90 91 30 GGT|GAT ( GAT|CCG|GCC|AAA|GCG(GCC|gcg|cc 3' olig#12. residue 75 76 77 78 79 80 81 81a 81b 5' gcI gag]cGC|ATG|CGT|ACC|TGC|qfk|qfk|qfk|qfk|qfk|- 35 82 83 84 85 86 87 GCT|GAA|GGT|GAT|GAT|CCG|- 88 89 90 91 GCCI AAAIGCGIGCCIgcg(cc 3' olig#l2a 40 144 residue 75 76 77 78 79 80 81 81a 81b 51 gc|gag|cGC|ATG|CGT|ACC|TGC|qfk j qfk|qfk|qfk|qfk|- 81c 81d 82 83 84 85 86 87 qfk|qfk|GCT|GAA|GGT|GAT|GAT|CCG|- 88 89 90 91 GCC|AAA|GCG|GCC|gcg|cc 3' olig#12b residue 91 90 89 88 87 865' ggIcgcIGGCICGCITTTIGGCICGGIATC 3' olig#13 where q is a mixture of (0.26 T, 0.18C, 0.26 A, and0.30 G), f is a mixture of (0.22 T, 0.16 C, 0.40 A, and0.22 G), and k is a mixture of equal parts of T and G.The bases shown in lower case at either end are spacersand are not incorporated into the cloned gene. Theprimer is complementary to the 3' end of each of thelonger oligo-nts. One of the variegated oligo-nts andthe primer olig#13 are combined in equimolar amountsand annealed. The dsDNA is completed with all four(nt)TPs and Klenow fragment. The resulting dsDNA andRF pLG7 are cut with both Sfi I and Sph I, purified,mixed, and ligated. This ligation mixture goes throughthe process described in Sec. 15 in which we select atransformed clone that, when induced with IPTG, bindsAHTrp.
To vary the junction between M13 signal sequenceand BPTI, we introduce DNA variegated at codons forresidues between 23 and 27 into the Kpn I and Xho Isites of pLG7. The first three residues are highlyvariable in amino acid sequences homologous to BPTI.Homologous sequences also vary in length at the aminoterminus. One of the oligo-nts olig#14, olig#14a, orolig#14b and the primer olig#15 are synthesized by η
<img img-format="tif" img-content="drawing" file="IL120940AD000211.tif" id="idf0011" />
145 standard methods. The oligo-nts are: residue : 17 18 19 20 21 22 23 24 25 5 5' g|gcc|gcG|GTA|CCG|ATG|CTG|TCT|TTT|GCT|qfk|qfk|- 26 27 28 29 30 |qfk|TTC|TGT|CTC|GAG|cgc|ccg|cga | 3' olig#14 residue 17 18 19 20 21 22 23 24 25 26 5'gIgccIgcGIGTAICCGIATGICTGITCTITTTIGCTIqfk|qfk|qfk|- 15 26a 26b 27 .28 29 30 |qfk|qfk|TTC|TGT|CTC|GAG|cgc|ccg|cga| 3' olig#14a, 20 residue 17 18 19 20 21 22 23 24 25 26 5'gIgccIgcG|GTA|CCG|ATG|CTGITCTITTTIGCT|qfk|qfk|qfk|- 26a 26b 26c 26d 27 28 29 30 |qfkIqfk|qfk]qfk|TTC|TGT|CTC|GAG|cgc|ccg|cga|3’olig#14b 5' |teg|egg|geg|CTC|GAG|ACA|GAA| 3' olig#15 3 0 where q is a mixture of (0.26 T, 0.18 C, 0.26 A, and0.30 G), f is a mixture of (0.22 T, 0.16 C, 0.40 A, and0.22 G) , and k is a mixture of equal parts of T and G.The bases shown in lower case at either end arespacers. One of the variegated oligo-nts and the 35 primer are combined in equimolar amounts and annealed.The ds DNA is completed with all four (nt) TPs andKlenow fragment. The resulting dsDNA and RF pLG7 arecut with both Kpn I and Xho I, purified, mixed, andligated. This ligation mixture goes through the 40 process described in Sec. 15 in which we select a transformed clone that, when induced with IPTG, binds AHTrp or trp. 146
If none of these approaches produces a workingchimeric protein, we may try a different signalsequence, or a different OSP in M13 (e.g., the gene IIIprotein for which there is fusion data (SMIT85,CRUZ88)), or another genetic package.
Example 1, Part II BPTI binds very tightly to trypsin(1¾ = 6.0 x 10"14 M) and to anhydrotrypsin, so thatthese molecules are not preferred for optimizing theamount of BPTI to display on LG7 or the amount ofaffinity molecule to attach to the column. Tschescheet al. reported on the binding of several BPTIderivatives to various proteases:
Dissociation constants for BPTI : derivatives, Molar. Residue Trypsin Chymotrypsin Elastase Elastase #15 (bovine pancreas) (bovine pancreas) (porcine pancreas) (human leukocytes) lysine 6.0 x 10" 14 9.0 x 10"9 - 3.5 x 10"6 glycine - - + 7.0 x 10"9 alanine + - 2.8 X 10"8 2.5 x 10"9 valine - - 5.7 x 10"8 1.1 xlO"10 leucine — 1.9 X 10"8 2.9 x 10"9
From the report of Tschesche et al. we infer thatmolecular pairs marked '·+" have K^s greater than 3.5 x IO"6 M and that molecular pairs marked have K^s much greater than 3.5 x 10"6 M. Because of thewealth of data about the binding of BPTI and variousmutants to trypsin and other proteases (TSCH87), we canproceed in various ways. (For other PBDs we can obtain 147 two different monoclonal antibodies, one with a highaffinity having 1¾ of order 10-11 M, and one with amoderate affinity having 1¾ on the order of 10“6 M.)In this example, we may use: a) the moderate bindingbetween BPTI and human leukocyte elastase (HuLEl), b)the moderately strong binding of porcine elastase toBPTI(V15), or c) the binding of BPTI(A15) (residue 38in the pbd gene) for trypsin (weak but detectable) orfor porcine pancreatic elastase.
We compare the retention of LG7 virions to theretention of wild-type M13 on (AHTrp). M13 derivativeshaving more DNA than wild-type M13 have correspondinglonger virions. Thus we will create pLG8 that differsfrom pLG7 only in having stop codons at codons 2 and3, and an altered L codon at codon 7 of the osp-ipbdgene. Phage LG8 will have exactly as much DNA as LG7;therefore the LG8 virion is exactly as long as the LG7virion. LG8 can not, however, display BPTI on itssurface.
To expedite identification of different MIS-derived phage, we replace the ampR gene of LG8 with thetetR gene from pBR322 by standard methods. The BSMI-to-Aatll tetR bearing fragment of pBR322 is ligatedinto DNA from pLG8 cut with Xbal and Aatll. Thecorrect construction, having 9.2 kb, is easilydistinguished from pBR322 and is called LG10.
The phage LG7 is grown at various levels of IPTGin the medium and harvested in the way previouslydescribed. An affinity column having bed volume of 2.0ml and supporting an amount of HuLEl picked from therange 0.1 mg to 30.0 mg on 1 ml of BioRad Affi-Gel 10(™) or Affi-Gel 15is designated (HuLEl). 148
An appropriate set of densities of HuLEl on the columnis (0.1 mg/ml, 0.5 mg/ml, 2.0 mg/ml, 8.0 mg/ml, 15.0mg/ml, and 30.0 mg/ml). The Vv of (HuLEl) is, byhypothesis, 1.0 ml. The elution of LG7 phage iscompared to the elution of LG10 on (HuLEl) havingvarying amounts of HuLEl affixed. The columns areeluted in a standard way: 1) 10 mM KC1 buffered to pH 8.0 with phosphate,until optical density at 280nm falls to base lineor 4 x Vv, whichever is first, 2) a gradient of 10 mM to 2 M KC1 in 3 x Vv, pHheld at 8.0 with phosphate, 3) a gradient of 2 M to 5 M KC1 in 3 x Vv,phosphate buffer to pH 8.0, 4) constant 5 M KC1 plus 0 to 0.8 M guanidinium Clin 2 x Vv, with phosphate buffer to pH 8.0.
The preferred level of induction (IPTGOptimal) andamount of affinity molecule on the matrix(DoAMoMOp^-|maj) are those settings that give thesharpest LG7 elution peak that shows significantretardation as compared to LG8, which carries no BPTI.By hypothesis, the best separation occurs for theamount of BPTI/GP produced when the cells are inducedwith 10.0 uM IPTG and when 4.0 mg HuLEl/ml is appliedto BioRad Affi-Gel 10 (™).
When the amount of BPTI/GP and the amount ofHuLEl/volume of support have been optimized, we turn tooptimization of elution rate, initial ionic strength,and the amount of GP/ (volume of support) . These 149 parameters can be optimized separately.
Using optimal BPTI/GP and HuLEl/volume of support,we measure the elution volume of LG7 and LG8 fordifferent elution rates, viz. 1, 1/2, 1/4, 1/8 and 1/16times the maximum flow rate. By hypothesis, 1/4 ofmaximum elution rate is better than 1/2, but 1/8 isabout the same as 1/4. Therefore 1/4 maximum elutionrate will be used.
Elution volumes of LG7 obtained from cells grownon media that is 2.0 mM in IPTG are measured at optimalDoAMoM and elution rate for loadings of 109, 1010,1011, and 1012 pfu. By hypothesis, 1012 pfu of pureLG7 overloads the column and significant number ofphage elute before their characteristic position in theKC1 gradient. We also find that 1011 pfu overloads thecolumn only slightly, and that 1010 pfu does notoverload the column. Because the use of the affinityseparation in Sec. 15 will involve a population inwhich no single member is more than one part in 104, weconclude that 1012 pfu of a variegated population couldbe applied to a column of 1.0 ml matrix volume withoutoverloading with respect any one species. Theoverloading of a 1.0 ml column by 1012 pfu alsoindicates that the initial column that capturesindiscriminately adhesive phage should be 5 to 10 timesas large as the column that supports the targetmaterial.
Elution volumes of LG7 and LG10 obtained fromcells grown on media that is 2.0 mM in IPTG aremeasured at optimal conditions and for a loading of1010 pfu for various initial ionic strengths: 1.0 mM,5.0 mM, 10.0 mM, 20.0 mM, and 50.0 mM. We may find, 150 for example, that LG10 is slightly retarded by thecolumn when loaded at 1.0 mM KCl, but that LG7 alwayscomes off the column at its characteristic place in thegradient. We use 10.0 mM as initial ionic strength inall remaining affinity separations.
To determine the sensitivity of chromatography ofphage that display variants of BPTI on their surfaces(Sec. 10.1), we prepare artificial mixtures of twoclosely-related phage that differ only at one residuein the BPTI domain. One variety of phage has strongaffinity for the column used in this step, while theother phage has no affinity for the column. Wechromatograph these mixtures to discover how little ofthe phage that binds to the column can be detectedwithin a large majority of phage that do not bind thecolumn.
For these tests we choose AHTrp as AfM(BPTI). Acolumn having 2 ml bed volume is prepared with(DoAMoMoptimal mg of AHTrp)/(ml of Affi-Gel 1θ(™)).The column is called (AHTrp) and has Vy = 1.0 ml. A new phage, LG9, is prepared that displaysBPTI(V15) as IPBD in contrast to LG7 that displaysBPTI(K15, wild-type) as IPBD. Residue 15 of BPTI isresidue 38 of the osp-ipbd gene. We introduce thechange K38 to V by replacement of a short segment ofthe osp-ipbd gene between Apa I &amp; Stu I. The correctconstruction is called pLG9. To expeditedifferentiation between LG7 and an LG9-derivativephage, we replace the ampR gene of LG9 with the tetRgene from pBR322. DNA from pBR322 between BsmI (1353,blunted) and Aatll (1428) is ligated to dsDNA from pLG9cut with Xbal (blunted) and Aatll. The correct 151 construction, having 9.2 kb, is easily distinguishedfrom pBR322 and is called LG11. DNA from phage LG11 issequenced in the vicinity the junctions of the newlyinserted tetR gene to confirm the construction. LG7 and LG11 are grown with optimum IPTG (2.0 mM)and harvested. Mixtures are prepared in the ratios LG7:LG11 :: l:Vlim where ranges from 1010 to 105 by factors of 10.Large values of are tested first; once a isfound that allows recovery of LG7, smaller values ofvlim are not be tested.
The column {AHTrp} is first blocked by treatmentwith 1011 virions of M13am429 in 100 ul of 10 mM KC1buffered to pH 8.0 with phosphate; the column is washedwith the same buffer until OD260 returns to base lineor 4 x Vy have passed through the column, whichevercomes first. One of the mixtures of LG7 and LG11containing 1012 pfu in 1 ml of the same buffer isapplied to {AHTrp}. The column is eluted in a standardway : 1) 10 mM KC1 buffered to pH 8.0 with phosphate,until optical density at 280nm falls to base lineor 4 x Vv, whichever is first, (discard effluent), 2) a gradient of 10 mM to 2 M KC1 in 3 x Vv, pHheld at 8.0 with phosphate, (30 x 100 ulfractions), 3) a gradient of 2 M to 5 M KC1 in 3 x Vv,phosphate buffer to pH 8.0, (30 x 100 ul 152 fractions), 4) constant 5 M KCl plus 0 to 0.8 M guanidinium Clin 2 x Vy, with phosphate buffer to pH 8.0, (20 x100 ul fractions), 5) constant 5 M KCl plus 0.8 M guanidinium Cl in 1.2 x Vy, with phosphate buffer to pH 8.0, (12 x 100 ul fractions).
Samples of 4 ul from each fraction are plated atsuitable dilution on phage-sensitive Sup+ cells (sothat M13am429 will not grow) . A sample of the columnmatrix is also used as inoculum for phage-sensitiveSup"*" cells. Plaques are transferred to ampicillin-containing LB agar, and AmpR colonies are tested fordisplay of BPTI(K15) by use of trp* or AHTrp*.
By hypothesis, V]_j_m = 4.0 x 108 is the largestvalue for which LG7 can be recovered. Thus Csens^ -4.0 x 108. Three cycles of chromatography are requiredto isolate LG7, so the first approximation to Ceff is740 ( = exp( loge(4.0 x 10θ)/3 ) ).
We now determine the efficiency of the affinityseparation (Sec. 10.2) . This is done by: a) preparingmixtures of LG7 and LG11 in the ratio 1:Q, b) enrichingthe population for LG7 for one separation cycle, and c)determining the fraction of LG7 in the last phage-bearing fraction. When Q is 1.5 x 104, 3% of coloniesare BPTI positive. When Q is 1.5 x 103, 60% of thecolonies are BPTI positive. Thus we calculate Ceff =.60 x 1.5 x 103 = 900. 153
Our hypothetical LG7 should display one or moreBPTI domains on each virion. The osp-ipbd gene isunder control of the lacUV5 promoter so that expressionlevels of BPTI-M13 CP can be manipulated via [IPTG].This construct may be used to develop many differentbinding proteins, all based on BPTI. An optimum levelof induction and amount of AfM(PBD) (= DoAMoMnpf· ·; = 2.0 mg/(ml of support)) should have been determined;target molecules will be applied to columns in thisamount in the process disclosed in Sec. 15.1. Theseoptimum levels may be adequate for all targets and allvariegations of BPTI displayed on derivatives of M13based on LG7, but some further optimization may beneeded if other values of pH or temperatures are used.
Other obd gene fragments may be substituted forthe boti gene fragment in pLG7 with a high likelihoodthat PBD will appear on the surface of the new LG7derivative.
Example 1, Part III HHMb is chosen as a typical protein target; another protein could be used. HHMb satisfies all of thecriteria for a target; 1) it is large enough to beapplied to an affinity matrix, 2) after attachment itis not reactive, and 3) after attachment there issufficient unaltered surface to allow specific bindingby PBDs.
The essential information for HHMb is known; 1)HHMb is stable at least up to 70°C, between pH 4.4 and9.3, 2) HHMb is stable up to 1.6 M Guanidinium Cl, 3)the pi of HHMb is 7.0, 4) for HHMb, Mr = 16,000, 5) HHMb requires haem, 6) HHMb has no proteolytic 154 activity.
In addition, the following information about HHMband other myoglobins is available: 1) the sequence ofHHMb, 2) the 3D structure of sperm whale myoglobin(HHMb has 19 amino acid differences and it is generallyassumed that the 3D structures are almost identical),3) its lack of enzymatic activity, 4) its lack oftoxicity.
We set the specifications of an SBD as :
1) T = 25°C 2) pH = 8.0 3) Acceptable solutes : A ) for binding : i) phosphate, as buffer, 0 to 20 mM, and ii) KC1, 10 mM, B ) for column elution : i) phosphate, as buffer, 0 to 30 mM, ii) KC1, up to 5 M, and iii) Guanidinium Cl, up to 0.8 M. 4) Acceptable K<j < 1.0 x 10~8 M.
We choose LG7 as GP(IPBD).
Residues to be varied are picked, in part, throughthe use of interactive computer graphics to visualizethe structures. In this section, all residue numbersrefer to BPTI. We pick a set of residues that forms asurface such that all residues can contact one targetmolecule. Information relevant to choosing BPTI 155 residues to vary includes: 1) the 3D structure, 2)solvent accessibility of each residue (LEEB71), 3) acompilation of sequences of other proteins homologousto BPTI, and 4) knowledge of the structural nature ofdifferent amino acid types.
Tables 16 and 34 indicate which residues of BPTI:a) have substantial surface exposure, and b) are knownto tolerate other amino acids in other closely relatedproteins. We use interactive computer graphics to picksets of eight to twenty residues that are exposed andvariable and such that all members of one set can toucha molecule of the target material at one time. If BPTIhas a small amino acid at a given residue, that aminoacid may not be able to contact the targetsimultaneously with all the other residues in theinteraction set, but a larger amino acid might wellmake contact. A charged amino acid might affectbinding without making direct contact. In such cases,the residue should be included in the interaction set,with a notation that larger residues might be useful.In a similar way, large amino acids near the geometriccenter of the interaction set may prevent residues oneither side of the large central residue from makingsimultaneous contact. If a small amino acid, however,were substituted for the large amino acid, then thesurface would become flatter and residues on eitherside could make simultaneous contact. Such a residueshould be included in the interaction set with anotation that small amino acids may be useful.
Table 35 was prepared from standard model partsand shows the maximum span between Cbeba and the tip ofeach type of side group. Cbeta is used because it isrigidly attached to the protein main-chain; rotation 156 about the Caipba-Cbeta bond is the most importantdegree of freedom for determining the location of theside group.
Table 34 indicates five surfaces that meet thegiven criteria. The first surface comprises the set ofresidues that contacts trypsin in the complex oftrypsin with BPTI as reported in the Brookhaven ProteinData Bank entry "1TPA". This set is indicated by thenumber "I". The exposed surface of the residues inthis set (taken from Table 16) totals 1148 A2 and theapproximates the area of contact between BPTI andtrypsin.
Other surfaces, numbered 2 to 5, were picked byfirst picking one exposed, variable residue and thenpicking neighboring residues until a surface wasdefined. The choice of sets of residues shown in Table34 is in no way exhaustive or unique; other sets ofvariable, surface residues can be picked. Hereinafterwe refer to K15 as being at the top of the molecule,while the carboxy and amino termini are at the bottom.
Solvent accessibilities are useful, easilytabulated indicators of a residue's exposure. Solventaccessibilities must be used with some caution; smallamino acids are under-represented and large amino acidsover-represented. The user must consider what thesolvent accessibility of a different amino acid wouldbe when substituted into the structure of BPTI.
To create specific binding between a derivative ofBPTI and HHMb, we will vary the residues in set #2.This set includes the twelve principal residues 17(R),19(1), 21(Y), 27(A), 28(G), 29(L), 31(Q), 32(T), 34(V), 157 48(A), 49(E), and 52 (M) (Sec. 13.1.1). None of theresidues in set #2 is completely conserved in thesample of sequences reported in Table 34; thus we canvary them with a high probability of retaining theunderlying structure. Independent substitution at eachof these twelve residues of the amino acid typesobserved at that residue would produce approximately 4.4 x 109 amino acid sequences and the same number ofsurfaces. BPTI is a very basic protein. This property hasbeen used in isolating and purifying BPTI and itshomologues so that the high frequency of arginine andlysine residues may reflect bias in isolation and isnot necessarily required by the structure. Indeed,SCI-III from Bombvx mori contains seven more acidicthan basic groups (SASA84).
Residue 17 is highly variable and fully exposedand can contain R, K, A, Y, H, F, L, Μ, T, G, Y, P, or S. All types of amino acids are seen: large, small,charged, neutral, and hydrophobic. That no acidicgroups are observed may be due to bias in the sample.
Residue 19 is also variable and fully exposed,containing P, R, I, S, K, Q, and L.
Residue 21 is not very variable, containing F or Yin 31 of 33 cases and I and W in the remaining cases.The side group of Y21 fills the space between T32 andthe main chain of residues 47 and 48. The OH at thetip of the Y side group projects into the solvent.Clearly one can vary the surface by substituting Y or Fso that the surface is either hydrophobic orhydrophilic in that region. It is also possible that 158 the other aromatic amino acid (viz. H) or the otherhydrophobics (L, M, or V) might be tolerated.
Residue 27 most often contains A, but S, K, L, andT are also observed. On structural grounds, thisresidue will probably tolerate any hydrophilic aminoacid and perhaps any amino acid.
Residue 28 is G in BPTI. This residue is in aturn, but is not in a conformation peculiar to glycine.Six other types of amino acids have been observed atthis residue: K, N, Q, R, H, and N. Small side groupsat this residue might not contact HHMb simultaneouslywith residues 17 and 34. Large side groups couldinteract with HHMb at the same time as residues 17 and34. Charged side groups at this residue could affectbinding of HHMb on the surface defined by the otherresidues of the principal set. Any amino acid, exceptperhaps P, should be tolerated.
Residue 29 is highly variable, most oftencontaining L. This fully exposed position willprobably tolerate almost any amino acid except,perhaps, P.
Residues 31, 32, and 34 are highly variable,exposed, and in extended conformations; any amino acidshould be tolerated.
Residues 48 and 49 are also highly variable andfully exposed, any amino acid should be tolerated.
Residue 52 is in an alpha helix. Any amino acid,except perhaps P, might be tolerated. 159
Now we consider possible variation of thesecondary set (Sec. 13.1.2) of residues that are in theneighborhood of the principal set. Neighboringresidues that might be varied at later stages include9(P), 11 (T), 15 (K), 16(A), 18(1), 20(R), 22(F), 24 (N) ,26(K), 35 (Y), 47(S), 50(D), and 53(R).
Residue 9 is highly variable, extended, andexposed. Residue 9 and residues 48 and 49 areseparated by a bulge caused by the ascending chain fromresidue 31 to 34. For residue 9 and residues 48 and 49to contribute simultaneously to binding, either thetarget must have a groove into which the chain from 31to 34 can fit, or all three residues (9, 48, and 49)must have large amino acids that effectively reduce theradius of curvature of the BPTI derivative.
Residue 11 is highly variable, extended, andexposed. Residue 11, like residue 9, is slightly farfrom the surface defined by the principal residues andwill contribute to binding in the same circumstances.
Residue 15 is highly varied. The side group ofresidue 15 points away form the face defined by set #2.Changes of charge at residue 15 could affect binding onthe surface defined by residue set #2.
Residue 16 is varied but points away from thesurface defined by the principal set. Changes incharge at this residue could affect binding on the facedefined by set #2.
Residue 18 is I in BPTI. This residue is in anextended conformation and is exposed. Five other aminoacids have been observed at this residue: M, F, L, V, 160 and T. Only T is hydrophilic. The side group pointsdirectly away from the surface defined by residue set#2. Substitution of charged amino acids at thisresidue could affect binding at surface defined byresidue set #2.
Residue 20 is R in BPTI. This residue is in anextended conformation and is exposed. Four other aminoacids have been observed at this residue: A, S, L, andQ. The side group points directly away from thesurface defined by residue set #2. Alteration of thecharge at this residue could affect binding at surfacedefined by residue set #2.
Residue 22 is only slightly varied, being Y, F, orH in 30 of 33 cases. Nevertheless, A, N, and S havebeen observed at this residue. Amino acids such as L,Μ, I, or Q could be tried here. Alterations at residue22 may affect the mobility of residue 21; changes incharge at residue 22 could affect binding at thesurface defined by residue set #2.
Residue 24 shows some variation, but probably cannot interact with one molecule of the targetsimultaneously with all the residues in the principalset. Variation in charge at this residue might have aneffect on binding at the surface defined by theprincipal set.
Residue 26 is highly varied and exposed. Changesin charge may affect binding at the surface defined byresidue set #2; substitutions may affect the mobilityof residue 27 that is in the principal set.
Residue 35 is most often Y, W has been observed. 161
The side group of 35 is buried, but substitution of For W could affect the mobility of residue 34.
Residue 47 is always T or S in the sequence sampleused. The Ogamma probably accepts a hydrogen bond fromthe NH of residue 50 in the alpha helix. Nevertheless,there is no overwhelming steric reason to precludeother amino acid types at this residue. In particular,other amino acids the side groups of which can accepthydrogen bonds, viz. N, D, Q, and E, may be acceptablehere.
Residue 50 is often an acidic amino acid, butother amino acids are possible.
Residue 53 is often R, but other amino acids havebeen observed at this residue. Changes of charge mayaffect binding to the amino acids in interaction set#2.
From published models (HUBE77, WLOD84) one can seethat R39 is on the opposite side of BPTI from thesurface defined by the residues in set #2. Therefore,variation at residue 39 at the same time as variationof some residues in set #2 is much less likely toimprove binding that occurs along surface #2 than isvariation of the other residues in set #2.
In addition to the twelve principal residues and13 secondary residues, there are two other residues,30(C) and 33(F), involved in surface #2 that we willprobably not vary, at least not until late in theprocedure. These residues have their side groupsburied inside BPTI and are conserved. Changing theseresidues does not change the surface nearly so much as 162 does changing residues in the principal set. Theseburied, conserved residues do, however, contribute tothe surface area of surface #2. The surface of residueset #2 is comparable to the area of the trypsin-bindingsurface. Principal residues 17, 19, 21, 27, 28, 29,31, 32, 34, 48, 49, and 52 have a combined solvent-accessible area of 946.9 82. Secondary residues 9, 11,15, 16, 18, 20, 22, 24, 26, 35, 47, 50, and 53 havecombined surface of 1041.7 82. Residues 30 and 33 haveexposed surface totaling 38.2 82. Thus the threegroups' combined surface is 2026.8 82.
Residue 3 0 is C in BPTI and is conserved in allhomologous sequences. It should be noted, however,that C14/C38 is conserved in all natural sequences, yetMarks et al. (MARK87) showed that changing both C14 andC3 8 to A, A or T,T yields a functional trypsininhibitor. Thus it is possible that BPTI-likemolecules will fold if C30 is replaced.
Residue 33 is F in BPTI and in all homologoussequences. Visual inspection of the BPTI structuresuggests that substitution of Y, Μ, H, or L might betolerated.
Given our hypothetical affinity separationsensitivity, Csens£, we decide to vary six residuesleaving some margin for errors in the actual basecomposition of variegated bases. To obtain maximalrecognition, we choose residues from the principal setthat are as far apart as possible. Table 36 shows thedistances between the beta carbons of residues in theprincipal and peripheral set. R17 and V34 are at oneend of the principal surface. Residues A27, G28, L29,A48, E49, and M52 are at the other end, about twenty 163
Angstroms away; of these, we will vary residues 17, 27,29, 34, and 48. Residues 28, 49, and 52 will be variedat later rounds.
Of the remaining principal residues, 21 is left tolater variations. Among residues 19, 31, and 32, wearbitrarily pick 19 to vary.
Unlimited variation of six residues produces 6.4 x107 amino acid sequences. By hypothesis, Csens^ is 1in 4 x 10®. Table 37 shows the programmed variegationat the chosen residues. The parental sequence ispresent as 1 part in 5.5 x 107, but the least favoredsequences are present at only 1 part in 4.2 χ 109.Among single-amino-acid substitutions from the PPBD,the least favored is F17-I19-A27-L29-V34-A48 and has acalculated abundance of 1 part in 1.6 x 10®. Using theoptimal qfk codon, we can recover the parental sequenceand all one-amino-acid substitutions to the PPBD ifactual nt compositions come within 5% of programmedcompositions. The number of transformants is Mntv =1.0 χ 109 (also by hypothesis), thus we will producemost of the programmed sequences.
The residue numbers above refer to mature BPTI.Since Table 25 refers to the pre-M13CP-BPTI protein,all mature BPTI sequence numbers have been increased bythe length of the signal sequence, 23. Thus, we wishto vary residues 40, 42, 50, 52, 57, and 71. A DNAsubsequence containing all these codons is foundbetween the (Apal) sites at base 191 and the SphI siteat base 309 of the osp-pbd gene. Among Apal. Drall.and Pssl. Apal is preferred because it recognizes sixbases without any ambiguity and will cut fewersequences in the vgDNA. Gratuitous restriction sites ///-:1 164 can be avoided in some cases by use of codon ambiguity:changing the codon for g51 from GGC to GGT makes itimpossible to generate an Apal site at codons 50, 51,and 6=52.
Each piece of dsDNA to be synthesized needs six toeight bases added at either end to allow cutting withrestriction enzymes and is shown in Table 37. Thefirst synthetic base (before cutting with Apal andSphI) is 184 and the last is 322. There are 142 basesto be synthesized. The center of the piece to thesynthesized lies between Q54 and V57. The overlap cannot include varied bases, so we choose bases 245 to 256as the overlap that is 12 bases long. Note that thecodon for F56 has been changed to TTC to increase theGC content of the overlap. The amino acids that arebeing varied are marked as X with a plus over them.Codons 57 and 71 are synthesized on the sense (bottom)strand. The design calls for "qfk" in the antisensestrand, so that the sense strand contains (from 5' to3') a) equal part C and A (i.e. the complement of k) ,b) (0.40 T, 0.22 A, 0.22 C, and 0.16 G) (i.e. thecomplement of f) , and c) (0.26 T, 0.26 A, 0.30 C, and0.18 G).
Each residue that is encoded by "qfk" has 21possible outcomes, each of the amino acids plus stop.Table 12 gives the distribution of amino acids encodedby "qfk", assuming 5% errors. The abundance of theparental sequence is the product of the abundances of RxIxAxLxVxA. The abundance of the least-favored sequence is 1 in 4.2 x 109.
Olig#27 and olig#28 are annealed and extended withKlenow fragment and all four (nt)TPs. Both the ds 165 synthetic DNA and RF pLG7 DNA are cut with both Apa Iand Sph I. The cut DNA is purified and the appropriatepieces ligated (See Sec. 14.1) and used to transformcompetent PE383. (Sec. 14.2). In order to generate asufficient number of transformants, we start with 5.0 1of cells • 1) culture E. coli in 5.0 1 of LB broth at 37°Cuntil cell density reaches 5 x 107 to 7 x 107cells/ml, 2) chill on ice for 65 minutes, centrifuge thecell suspension at 4000g for 5 minutes at 4°C, 3) discard supernatant; resuspend the cells in1667 ml of an ice-cold, sterile solution of 60mM CaCl2, 4) chill on ice for 15 minutes, and thencentrifuge at 4000g for 5 minutes at 4°C, 5) resuspend cells in 2 x 400 ml of ice-cold,sterile 60 mM CaC^; store cells at 4°C for 24hours, 6) add DNA (100 gg) in 20 ml of litigation orTE buffer; mix, inculafe on ice for minutes, 7) distribute into 200 μΐ aliquots and heatshock cells at 42°C for 20 seconds, 8) add 200 ml LB broth and incubate at 37°C for1 hour, 9) add the culture to 2.0 1 of LB broth 166 containing ampicillin at 35-100 ug/ml andculture overnight at 37°C, 10) after 6 hours, remove 200 ml and plate 0.5ml portions with log phase JM 107 on LB agar,using the soft-agar overlay technique. Phageare prepared from the soft agar, 11) centrifuge the overnight culture to removecells, and pellet phage (MESS83), 12) harvest virions by method of Salivar, etal. (SALI64).
It is important to: a) use all or nearly all thevgDNA synthesized in ligation, b) use all or nearly allthe ligation mixture to transform cells, and c) cultureall or nearly all the transformants. These measuresare directed at maintaining diversity.
It is important to collect virions in a way thatsamples all or nearly all the transformants. BecauseF” cells are used in the transformation, multipleinfections do not pose a problem in the overnight phageproduction. F' cells are used for phage production inagar. HHMb has a pi of 7.0 and we carry outchromatography at pH 8.0 so that HHMb is slightlynegative while BPTI and most of its mutants arepositive. HHMb is fixed (Sec. 15.1) to a 2.0 ml columnon Affi-Gel io(™) or Affi-Gel 15(™) at 4.0 mg/mlsupport matrix, the same density that is optimal for acolumn supporting trp. 167
To remove variants of BPTI with strong,indiscriminate binding for any protein or for thesupport matrix (Sec. 15.2), we pass the variegatedpopulation of virions over a column that supportsbovine serum albumin (BSA) before loading thepopulation onto the (HHMb) column. Affi-Gel io(™) orAffi-Gel 15(™) is used to immobilize BSA at thehighest level the matrix will support. A 10.0 mlcolumn is loaded with 5.0 ml of Affi-Gel-linked-BSA;this column, called {BSA}, has Vy = 5.0 ml. Thevariegated population of virions containing 1012 pfu in1 ml (0.2 x Vv) of 10 mM KCl, 1 mM phosphate, pH 8.0buffer is applied to {BSA}. We wash {BSA} with 4.5 ml(0.9 x Vy) of 50 mM KCl, 1 mM phosphate, pH 8.0 buffer.The wash with 50 mM salt will elute virions that adhereslightly to BSA but not virions with strong binding.
The pooled effluent of the {BSA} column is 5.5 ml ofapproximately 13 mM KCl.
The column {HHMb} is first blocked by treatmentwith 10^1 virions of M13(am429) in 100 ul of 10 mM KClbuffered to pH 8.0 with phosphate; the column is washedwith the same buffer until OD2g0 returns to,base lineor 2 x Vy have passed through the column, whichevercomes first. The pooled effluent from {BSA} is addedto {HHMb} in 5.5 ml of 13 mM KCl, 1 mM phosphate, pH8.0 buffer. The column is eluted (Sec. 15.3) in thefollowing way: 1) 10 mM KCl buffered to pH 8.0 with phosphate,until optical density at 280nm falls to base lineor 2 x Vy, whichever is first, (effluentdiscarded), 168 2) a gradient of 10 mM to 2 M KC1 in 3 x Vy, pHheld at 8.0 with phosphate, (30 x 100 μΐfractions), 3) a gradient of 2 M to 5 M KC1 in 3 x Vy,phosphate buffer to pH 8.0 (30 x 100 μΐfractions), 4) constant 5 M KC1 plus 0 to 0.8 M guanidinium Clin 2 x Vy, with phosphate buffer to pH 8.0, (20 x100 μΐ fractions), and 5) constant 5 M KC1 plus 0.8 M guanidinium Cl in 1x Vy, with phosphate buffer to pH 8.0, (10 x 100μΐ fractions).
In addition to the elution fractions, a sample isremoved from the column and used as an inoculum forphage-sensitive Sup+ cells (Sec. 15.4). A sample of 4μΐ from each fraction is plated on phage-sensitive Sup+cells. Fractions that yield too many colonies to countare replated at lower dilution. An approximate titreof each fraction is calculated. Starting with the lastfraction and working toward the first fraction that wastitered, we pool fractions until approximately 109phage are in the pool, i.e. about 1 part in 1000 of thephage applied to the column. This population isinfected into 3 x 1011 phage-sensitive PE384 in 300 mlof LB broth. The low multiplicity of infection ischosen to reduce the possibility of multiple infection.After thirty minutes, viable phage have enteredrecipient cells but have not yet begun to produce newphage. Phage-born genes are expressed at this phase,and we can add ampicillin that will kill uninfectedcells. These cells still carry F-pili and will absorb 169 phage helping to prevent multiple infections.
If multiple infection should pose a problem thatcannot be solved by growth at low multiple-of-infection on F+ cells, the following procedure can beemployed to obviate the problem. Virions obtained fromthe affinity separation are infected into F+ E. coliand cultured to amplify the genetic messages (Sec. 15.5). CCC DNA is obtained either by harvesting RF DNAor by in vitro extension of primers annealed to ssphage DNA. The CCC DNA is used to transform F“ cellsat a high ratio of cells to DNA. Individual virionsobtained in this way should bear proteins encoded onlyby the DNA within.
The variegation produces as many as 6.4 x 107different amino-acid sequences. Ceff is 900. Thus,after two separation cycles, the probability ofisolating a single SBD is less than 0.10; after threecycles, the probability rises above 0.10.
The phagemid population is grown andchromatographed three times and then examined for SBDs(Sec. 15.7). In each separation cycle, phage from thelast three fractions that contain viable phage arepooled with phage obtained by removing some of thesupport matrix as an inoculum. At each cycle, about1012 phage are loaded onto the column and about 109phage are cultured for the next separation cycle.
After the third separation cycle, 32 colonies arepicked from the last fraction that contained viablephage; phage from these colonies are denoted SBD1, SBD2,..., and SBD32.
Each of the SBDs is cultured and tested for 170 retention on a Pep-Tie column supporting HHMb (Sec.15.8). Phage LG7(SBD11) shows the greatest retentionon the Pep-Tie (HHMb) column, eluting at 367 mM KC1while wtM13 elutes at 20 mM KCl. SBD11 becomes theparental amino-acid sequence to the second variegationcycle.
The result of this hypothetical experiment isshown in Table 38. R40 changed to D, 142 changed to Q, A50 changed to E, L52 remained L, and A71 changed to W.
The next round of variegation (Sec. 16) isillustrated in Table 39. The residues to be varied arechosen by: a) choosing some of the residues in theprincipal set that were not varied in the first round(viz. residues 42, 44, 51, 54, 55, 72, or 75 of thefusion), and b) choosing some residues in the secondaryset. Residues 51, 54, 55, and 72 are varied throughall twenty amino acids and, unavoidably, stop. Residue44 is only varied between Y and F. Some residues inthe secondary set are varied through a restrictedrange; primarily to allow different charges (+, 0, -)to appear. Residue 38 is varied through K, R, E, or G.Residue 41 is varied through I, V, K, or E. Residue 43is varied through R, S, G, N, K, D, E, T, or A.
Olig#29 and olig#30 are synthesized, annealed,extended and cloned into pLG7 at the Apa I/Sph I sites.The ligation mixture is used to transform 5 1 ofcompetent PE383 cells so that 109 transformants areobtained. A new (HHMb) is constructed using the samesupport matrix as was used in round 1. A sample of1012 of the harvested LG7 are applied to {HHMb} andaffinity separated. The last 109 phage off the column 171 and an inoculum are pooled and cultured. The culturedphagemids are re-chromatographed for three separationcycles. Thirty-two clonal isolates (denoted SBD11-1,SBD11-2,..., SBD11-32) are obtained from the effluentof the third separation cycle and tested for binding ona Pep-Tie {HHMb} column. Of this set, SBD11-23 showsthe greatest retention on the Pep-Tie {HHMb} column,eluting at 692 mM KC1.
The results of this hypothetical selection isshown in Table 40. Residue 38 (K15 of BPTI) changed to E, 41 becomes V, 43 goes to N, 44 goes to F, 51 goes to F, 54 goes to S', 55 goes to A, and 72 goes to Q.
The sbdll-23 portion of the osp-pbd gene is clonedinto an expression vector and BPTI(E15, D17, V18, Q19,N20, F21, E27, F28, L29, S31, A32, S34, W71, Q72) isexpressed in the periplasm. This protein is isolatedby standard methods and its binding to HHMb is tested.Kd is found to be 4.5 x 10”7 M. A third round of variation, using SBD11-23 asPPBD, is illustrated in Table 41; eight amino acids arevaried. Those in the principal set, residues 40, 55,and 57, are varied through all twenty amino acids.Residue 32 is varied through P, Q, T, K, A, or E.
Residue 34 is varied through T, P, Q, K, A, or E.
Residue 44 is varied through F, L, Y, C, W, or stop.
Residue 50 is varied through E, K, or Q. Residue 52 is varied through L, F, I, M, or V.
The result of this variation is shown in Table 42.The selected SBD is denoted SBD11-23-5 and elutes froma Pep-Tie {HHMb} column at 980 mM KC1. The sbdll-23-5 segment is cloned into an expression vector and 172 BPTI(E9, Qll, E15, A17, V18, Q19, Ν20, W21, Q27, F28,M29, S31, L32, H34, W71, Q72) is produced. This timethe Kd is 7.3 x 10"9 M. 5 This example is hypothetical. It is anticipated
that more variegation cycles will be needed to achievedissociation constants of 10“8 M. It is also possiblethat more than three separation cycles will be neededin some variegation cycles. Real DNA chemistry and DNA 10 synthesizers may have larger errors than our hypothetical 5%. If Serr > 0.05, then we may not beable to vary six residues at once. Variation of 5residues at once is certainly possible. 173
Citations : ACHT78:
Achtman, M, G Morelli, S Schwuchow, J Bacteriol (1978), 135 (3) pl053-61. AKOH72:
Ako, H, RJ Foster, and CA Ryan,
Biochem Biophys Res Commun (USA)(1972), 47(6) pl402-7 ANFI73:
Anfinsen, CB,
Science (1973), 181(96)223-30. ARGO87:
Argos, P, J. Mol. Biol. (1987), 197:331-348. AUDI84a:
Auditore-Hargreaves, K,
United States Patent 4,470,925, September 11, 1984. AUDI84b:
Auditore-Hargreaves, K,
United States Patent 4,479,895, October 30, 1984. AUER87:
Auerswald, E-A, W Schroeder, and M Kotick,
Biol. Chem. Hoppe-Seyler (1987), 368:1413-1425. 174 AUSU87:
Ausubel, FM, R Brent, RE Kingston, DD Moore, JGSeidman, JA Smith, and K Struhl, EditorsCurrent Protocols in Molecular Biology.
Greene Publishing Associates and Wiley-Interscience,Publishers
John Wiley &amp; Sons, New York, 1987. BANN81:
Banner, DW, C Nave, and DA Marvin,
Nature (19811.289:814-816. BASH87:
Bash, PA, UC Singh, R Langridge, and PA Kollman,Science (1987), 236 (4801) p564-8. BECK83:
Beckwith, J, and TJ Silhavy,
Methods in Enzymology (1983), 97:3-11. BECK88:
Beckwith, J, D Boyd, K McGovern, C. Manoil, JL SanMilan, S Froshauer, and N Green
Talk presented at "The Protein Folding Problem", aseries of lectures and posters presented at the 1988annual meeting of AAAS in Boston. BENS84:
Benson, SA, E Bremer, and TJ Silhavy,
Proc Natl Acad Sci USA (1984), 81:3830-3834. 175 BENS86:
Benson, N, P Sugiono, S Bass, LV Mandelman, PYouderian,
Genetics (1986) 114(1)1-14. BETT88:
Better, M, CP Chang, RR Robinson, and AH Horwitz,Science (1988), 240:1041-1043. BIRD67:
Birdsell, DC, and EH Cota-Robles, J Bacteriol (1967), 93:427-437. BLUN88:
Blundell, T, D Carney, S Gardner, F Hayes, B Howlin, THubbard, J Overington, DA Singh, BL Sibanda, and MSutcliffe,
Eur J Biochem (15 March 1988), 172 (3) p513-20. BOEK80:
Boeke, JD, M Russel, and P Model, J. Mol. Biol. (1980), 144:103-116. BONN85:
Bonnafous, JC, J Fornand, J Favero, and J-C Mani,Chapter 8 in Affinity Chromatography, a practicalapproach..
Edited by PDG Dean, WSJohnson and FA Middle, IRL Press, Oxford, UK 1985 BONO85:
Bonomi, F, S Pagani, DM Kurtz Jr,
Eur J Biochem (1985), 148(1)67-73. 176 BOQU87:
Boquet, PL, C Manoil, and J Beckwith, J. Bacteriol. (1987), 169:1663-1669. BOTS85:
Botstein, D, and D Shortle,
Science (1985), 229:1193-1201. BRIG87:
Briggs, MR, JT Kadonaga, SP Bell, and R Tjian,
Science (Oct 3 1986), 234 (4772) 47-52. CANT87:
Canters, GW, FEBS Letters (1987), 212(1)168-172. CARU83:
Caruthers, MH, SL Beaucage, JW Efcavitch, EF Fisher,RA Goldman, PL DeHaseth, W Mandecki, MD Matteucci, MS Rosendahl, and Y Stabinski,
Cold Spr. Harb. Symp. Quant. Biol. (1983), 47:411-418. CARU85:
Caruthers, MH,
Science (1985), 230:281-285. CARU87:
Caruthers, ΜΗ, P Gottlieb, LP Bracco, and L Cummings,in Protein Structure, Folding, and Design 2. 1987.
Ed. D Oxender (New York, AR Liss Inc.) p.9ff. CHAM82:
Chambers, RW, I Kucan, and Z Kucan,
Nucleic Acids Res. (1982), 10(20)6465-73♦ 177 CHAN79:
Chang, CN, P Model, and G Blobel,
Proc. Natl. Acad. Sci. USA (1979), 76:1251-1255. CHAR84:
Charbit, A, J-M Clement, and M Hofnung, J. Mol. Biol. (1984), 175:395-401. CHAR87:
Charbit, A, E Sobczak, ML Michel, A Molla, P Tiollais,M Hofnung, J Immunol (1987), 139:1658-64. CHAZ85:
Chazin, WJ, DP Goldenberg, TE Creighton, and KWuthrich,
Eur J Biochem (1985), 152:(2)429-37. CHEN88:
Chen, W, and K Struhl,
Proc Natl Acad Sci USA (1988), 85:2691-2695. CHOT75:
Chothia, C, and J Janin,
Nature (1975), 256:705-708. CHOT76:
Chothia, C, S Wodak, and J Janin,
Proc. Natl. Acad. Sci. USA (1976), 73:3793-7. CHOT86:
Chothia, C, and AM Lesk, EMBO J (1986), 5:823-826. 178 CHOU74:
Chou, PY, and GD Fasman,
Biochemistry (1974), 13:(2)222-45. 5 CHOU78a:
Chou, PY, and GD Fasman,
Adv Enzymol (1978), 47:45-148. CHOU78b: 10 Chou, PY, and GD Fasman,
Annu Rev Biochem (1978), 47:251-76. CHUN86:
Chung, DW, K Fujikawa, BA McMullen, and EW Davie, 15 Biochemistry (1986), 25:2410-2417. CLEM81:
Clement, JM, and M Hofnung,
Cell (1981), 27:507-514. 20 CLEM83:
Clement JM, E Lepouce, C Marchal, and M Hofnung,EMBO J (1983), 2:77-80. 25 CLOR87:
Clore, GM, AM Gronenborn, M Kjaer, and FM Poulsen,Protein Engineering (1987), 1:305-311. CLUN84: 30 Clune, A, K-S Lee, and T Ferenci,
Biochem. and Biophys. Res. Comm. (1984), 121:34-40. 179 CRAI85:
Craik, CS, C Largman, T Fletcher, S Roczniak, PJ Barr, R Fletterick, and WJ Rutter,
Science (1985), 228:291-7. CRAW87:
Crawford, IP, M Clarke, M van Cleemput, and C Yanofsky,J Biol Chem (1987), 262(1)239-244. CREI84:
Creighton, TE,
Proteins: Structures and Molecular Principles.. W. H. Freeman &amp; Co., New York, 1984. CRUZ88: de la Cruz, VF, AA Lal, and TF McCutchan, J Biol Chem (1988), 263(9)4318-4322. DAIR80:
Dairs, RW, D Botstein, and JR Roth,
Advanced Bacterial Genetics.
Cold Spring Harbor Laboratory Press, 1980. DAWK86:
Dawkins, R,
The Blind Watchmaker W. W. Norton &amp; Co., New York, 1986. DAYR86:
Dayringer, H, A Tramantano, and R Fletterick,
Computer Graphics and Molecular Modeling.
Cold Spring Harbor Laboratory, Cold Spring Harbor, NY, 1986 180 DEBR86:
Debro, L, PC Fitz-James, and A Aronson, J Bacteriol (1986), 165:258-68. DENH78:
Denhardt, DT, D Dressier, and DS Ray editors,
The Spring-Stranded DNA Phages. Cold Spring HarborLaboratory, 1978. DEVO78:
DeVore, DP, and RJ Gruebel,
Biochem Biophys Res Commun (1978), 80(4)993-9. DICK83:
Dickerson, RE, and I Geis,
Hemoglobin: Structure, Function, Evolution, and
Pathology..
The Bejamin/Cummings Publishing Co., Menlo Park, CA1983. DILL87:
Dill, KA,
Protein Engineering (1987), 1:369-371. DONO87
Donovan, W, Z Liangbiao, K Sandman, and R Losick, J Mol Biol (1987), 196:1-10. DUFT85:
Dufton, MJ,
Eur J Biochem (1985), 153:647-654. 181 EISE85:
Eisenbeis, SJ, MS Nasoff, SA Noble, LP Bracco, DRDodds, MH Caruthers,
Proc. Natl. Acad. Sci. USA (1985), 82:1084-1088. ENDE78:
Endermann, R, C Kramer, and U Henning, FEBS Letters (1978), 86:21-24. EPST63:
Epstein , CJ, RF Goldberger, and CB Anfinsen,
Cold Spr. Harb. Symp. Quant. Biol. (1963), 28:439ff. ERIC86:
Erickson, BW, SB Daniels, PA Reddy, CG Unson, JSRichardson, and DC Richardson,
Current Communications in Molecular Biology: Computer
Graphics and Molecular Modeling..
Cold Spring Harbor Laboratory, Cold Spring Harbor, NY,1986,
Fletterick, R and M Zoller, Editors. ERRI88:
Errington, J, S Rong, MS Rosenkranz, and AL Sonenshein,J Bacteriology (1988), 170:1162-1167. FERE80a:
Ferenci, T, J Brass, and W Boos,
Biochem Soc Trans (1980), 8:680-1. FERE80b:
Ferenci, T, and W Boos, J Supramol Struct (1980), 13:101-16.
<img img-format="tif" img-content="drawing" file="IL120940AD000212.tif" id="idf0012" />
i .'·.·. ·! 182 FERE80c:
Ferenci, T,
Eur J Biochem (1980), 108:631-6. FERE82a:
Ferenci, T,
Ann. Microbiol. (Inst. Pasteur) (1982), 133A:167-169 FERE82b:
Ferenci, T, and K-S Lee, ,J. Mol. Biol. (1982), 160:431-444. FERE83:
Ferenci, T, and KS Lee, J Bacteriol. (1983), 154:984-987. FERE86a:
Ferenci, T, and K-S Lee, J. Bacteriol. (1986), 166:95-99. FERE86b:
Ferencei, T, and K-S Lee, J. Bacteriol. (1986), 167:1081-1082. FERE86c:
Ferenci, T, M Muir, K-S Lee, and D Maris,
Biochimica et Biophysica Acta (1986), 860:44-50. FERE87a:
Ferenci, T, and KS Lee,
Biochim Biophys Acta (1987), 896:319-22. 183 FERE87b:
Ferenci, T, TJ Silhavy, J Bacteriol (1987), 169:5339-42. FIOR85:
Fioretti, E, G lacopino, M Angeletti, D Barra, F Bossaand F Ascoli, J Biol Chem (1985), 260:11451-11455. FRIT85:
Fritz, H-J, in DNA Cloning. Editor: DM Glover, IRL Press, Oxford, UK,1985.
Volume I, Chapter 8, pl51-163. GABA82:
Gabay, J, and M Schwartz, J Biol chem (1982), 257(12)6627-6630. GARA83:
Garavito, RM, J Jenkins, JN Jonsonius, R Karlsson, andJP Rosenbusch, J Mol Biol (1983), 164:313-327. GEHR87:
Gehring, K, A Charbit, E Brissaud, and M Hofnung, J Bacteriol (1987), 169(5)2103-2106. GOLD83:
Goldenberg, DP, and TE Creighton, J Mol Biol (1983), 165: (2) p407-13. 184 GOLD87:
Gold, L, and G Stormo,
Volume 2, Chapter 78, p. 1302-1307, in
Escherichia coli and Salmonella typhimurium: Cellular and Molecular Biology.
Neidhardt, FC, Editor-in-Chief,
Amer. Soc. for Microbiology, Washington, DC, 1987. GOTT87:
Gottesman, S,
Volume 2, Chapter 79, p. 1308-1312, in
Escherichia coli and Salmonella typhimurium: Cellular and Molecular Biology.
Neidhardt, FC, Editor-in-Chief,
Amer. Soc. for Microbiology, Washington, DC, 1987. HAYA7 6:
Hayashi, K, M Takechi, N Kaneda, and T Sasaki, FEBS Lett (1976), 66(2)210-4. HEIN87:
Heine, HG, J Kyngdon, and T Ferenci,
Gene (1987), 53:287-92. HEIN88:
Heine, HG, G Francis, KS Lee, and T Ferenci, J Bacteriol (April 1988), 170:1730-8. 185 HERR78:
Herrmann, R, K Neugebauer, H Schaller, and H Zentgraf,in The Single-Stranded DNA Phages. Denhardt, DT, D Dressier, and DS Ray editors, Cold Spring Harbor 5 Laboratory, 1978., p473-476. HICK88:
Hickman, RK, and SB Levy, J Bacteriol (1988), 170(4)1715-1720. 10 HINE80:
Hines, JC, and DS Ray,
Gene (1980), 11:(3-4)207-18. 15 HOGL83:
Hogle, J, T Kirchhausen, and SC Harrison, J. Mol. Biol. (1983), 171:95-100. HOLL83: 20' Hollecker, M, and TE Creighton, J. Mol. Biol. (1983), 168:409-437. HOOP87:
Hoopes, BC, and WR McClure, 25 Volume 2, Chapter 75, p 1231-1240, in
Escherichia coli and Salmonella typhimurium: Cellular and Molecular Biology.
Neidhardt, FC, Editor-in-Chief,
Amer. Soc. for Microbiology, Washington, DC, 1987. 30 186 HUBE77:
Huber, R, W Bode, D Kukla, U Kohl, CA Ryan,
Biophys Struct Meeh (1975), 1(3)189-201 5 INOU86:
Inouye, M, and R Sarma, Editors,
Protein Engineering: Applications in Science, Medicine, and Industry..
Academic Press, New York, 1986. 10 ITOK79:
Ito, K, G Mandel, and W Wickner,
Proc. Natl. Acad. Sci. USA (1979), 76:1199-1203. 15 JANI85:
Janin, J, and C Chothia,
Methods in Enzymology (1985), 115(281420-430. JAZW73a: 20 Jazwinski, SM, R Marco, and A Kornberg,
Proc Natl Acad Sci USA (1973), 70(1)205-9. JAZW73b:
Jazwinski, SM, R Marco, and A Kornberg, 25 Virology (1975), 66(1)294-305. JAZW74:
Marco, R, SM Jazwinski, and A Kornberg,
Virology (1974), 62:(1)209-23. 30 JONE85:
Jones, TA,
Methods Enzymol (1985), 115:157-71. 187 JONE87:
Jones, KA, JT Kadonaga, PJ Rosenfeld, TJ Kelly, and R
Tjian,
Cell (Jan 16 1987), £8:79-89. JOUB80:
Joubert, FJ, and N Taljaard,
Hoppe-Seyler's Z. Physiol. Chem. (1980), 361:661-674. KABS84:
Kabsch, W, and C Sander,
Proc Natl Acad Sci USA (1984), 81(4)1075-8. KADO86:
Kadonaga, JT, and R Tjian,
Proc Natl Acad Sci USA (Aug 1986), 83 (16) 5889-93 KAIS87:
Kaiser, CA, D Preuss, P Grisafi, and D Botstein,Science (1987), 235:312-317. KANE76:
Kaneda, N, T Sasaki, and K Hayashi, FEBS Lett (1976), 70(1)217-22. KAPL78:
Kaplan, DA, L Greenfield, and G Wilcox, in The Single-Stranded DNA Phages. Denhardt, DT, D Dressier, and DS Ray editors, Cold Spring HarborLaboratory, 1978., p461-467. KUHN85a:
Kuhn, A, and W Wickner, J. Biol. Chem. (1985), 260:15914-15918. 188 KUHN85b:
Kuhn, A, and W Wickner, J. Biol. Chem. (1985), 260:15907-15913. KUHN87:
Kuhn, A,
Science (1987), 238:1413-1415. LAND87:
Landick, R, and C Yanofsky,
Volume 2, Chapter 77, p 1276-1301,
Escherichia coli and Salmonella typhimurium: Cellular and Molecular Biology.
Neidhardt, FC, Editor-in-Chief,
Amer. Soc. for Microbiology, Washington, DC, 1987. LEEB71:
Lee, B, and FM Richards, J Mol Biol (1971), 55:(3)379-400, LEEC86:
Lee, C, and J Beckwith,
Ann. Rev. Cell Biol. (1986), 2::315-336. LOSI86:
Losick, R, P Youngman, and PJ Piggot,
Ann Rev Genet (1986), 20:625-669. MAKE80:
Makela, 0, H Sarvas, and I Seppala, J. Immunol. Methods (1980), 37:213-223. 189 MAK080:
Makowski, L, DLD Caspar, and DA Marvin, J. Mol. Biol. (1980), 140:149-181. 5 MALA64:
Malamay, MH, and BL Horecker,
Biochem (1964), 3:1889-1893. MANI82: 10 Maniatis, T, EF Fritsch, and J. Sambrook,
Molecular Cloning.
Cold Spring Harbor Laboratory, 1982. MAN086: 15 Manoil, C, and J Beckwith,
Science (1986), 233:1403-1408. MARC83:
Marchal, C, and M Hofnung, 20 EMBO J (1983), 2:81-86. MARK86:
Marks, CB, M Vasser, P Ng, W Henzel, and S Anderson,J. Biol. Chem. (1986), 261:7115-7118. 25 MARK87:
Marks, CB, H Naderi, PA Kosen, ID Kuntz, and SAnderson,
Science (1987), 235:1370-1373. 30 190 MARQ83:
Marquart, M, J Walter, J Deisinhoffer, W Bode, and RHuber,
Acta Cryst, B (1983), 39:480ff. 5 MARV78:
Marvin, DA, in The Single-Stranded DNA Phages. Denhardt, DT, D Dressier, and DS Ray editors, Cold Spring Harbor 10 Laboratory, 1978., p583-603. MCPH86:
McPheeters, DS, A Christensen, ET Young, G Stormo, andL Gold, 15 Nucleic Acids Res (1986), 14:5813-26. MESS77:
Messing, J, B Gronenborn, B Muller-Hill, and PHHofschneider, 20 Proc Natl Acad Sci USA (1977), 74:3642-6. MESS78:
Messing, J, and B Gronenborn, in The Single-Stranded DNA Phages. Denhardt, DT, 25 D Dressier, and DS Ray editors, Cold Spring Harbor
Laboratory, 1978.,p449-453. MICH86:
Michaelis, s, JF Hunt, and J Beckwith, 30 J. Bacteriol. (1986), 167:160-167. 35 191 MILL72:
Miller, JH,
Experiments in Molecular Genetics.
Cold Spring Harbor Laboratory, Cold Spring Harbor, NY. 1972 MILL87a:
Miller, S, J Janin, AM Lesk, and C Chothia, J Mol Biol (1987), 196:641-656. MILL87b:
Miller, ES, J Karam, M Dawson, M Trojanowska, P Gauss,and L Gold, J Mol Biol (1987), 194:397-410. MILL88:
Miller, J, JA Hatch, S Simonis, and SE Cullen,
Proc Natl Acad Sci USA (1988), 85:1359-1363. MOSE83:
Moser, R, RM Thomas, and B Gutte, FEBS Letters (1983), 157:247-251. MOSE85:
Moser, R, S Klauser, T Leist, H Langen, T Epprecht, andB Gutte,
Angew. Chemie, Internatl Eng Ed. (1985), 24:719-798. MOSE87:
Moser, R, S Frey, K Muenger, T Hehlgans, S Klauser, HLangen, E-L Winnacker, R Mertz, and B Gutte,
Protein Engineering (1987), 1:339-343. 192 NAKA86:
Nakae, T, J. Ishii, and T Ferenci, J. Biol. Chem. (1986), 261:622-626. NAKA87:
Nakamura, Τ, T Hirai, F Tokunaga, S Kawabata, and SIwanaga, J Biochem. (1987), 101:1297-1306. NEID87:
Neidhardt, FC, Editor-in-Chief,
Escherichia coli and Salmonella typhimurium: Cellular and Molecular Biology.
Amer. Soc. for Microbiology, Washington, DC, 1987. NEUH65:
Neu, HC, and LA Heppel, J Biol chem (1965), 240:3685-3692. NIKA84:
Nikaido, H, and HCP Wu,
Proc Natl Acad Sci USA (1984), 81:1048-1052. NIKA87:
Nikaido, H, and M Vaara,
Volume 1, Chapter 3, p7-22.
Escherichia coli and Salmonella typhimurium: Cellular and Molecular Biology.
Neidhardt, FC, Editor-in-Chief,
Amer. Soc. for Microbiology, Washington, DC, 1987. NOMU78:
Nomura, N, A Oka, M Takanami, and H Yamagishi,in The Single-Stranded DNA Phages. Denhardt, DT, D Dressier, and DS Ray editors, Cold Spring Harbor 193
Laboratory, 1978., p467-472. OHKA81:
Ohkawa, I, and RE Webster, 5 J. Biol. chem. (1981), 256:9951-9958. OHTA76:
Ohta, Μ, T Sasaki, and K Hayashi, FEBS Lett (1976), 72(1)161-6. 10 OLIP86:
Oliphant, AR, AL Nussbaum, and K Struhl,
Gene (1986), 44:177-183. 15 OLIP87:
Oliphant, AR, and K Struhl
Methods in Enzvmoloqy 155 (1987) p 568-582.
Editor Wu, R; Academic Press, New York. 20 OLIV85:
Oliver, D,
Ann. Rev. Microbiol. (1985), 39:615-648. OLIV87: 25 Oliver, DB,
Volume 1, Chapter 6, p 56-69, in
Escherichia coli and Salmonella typhimurium: Cellular and Molecular Biology.
Neidhardt, FC, Editor-in-Chief, 30 Amer. Soc. for Microbiology, Washington, DC, 1987. .-:¾ 194 PABO79:
Pabo, CO, RT Sauer, JM Sturtevant, and M Ptashne,Proc. Natl. Acad. Sci. USA (1979), 76:1608-1612. PADL85:
Padlan, EA, and WE Love, J Biol Chem (1985), 260 (14) p8272-9. PAKU86:
Pakula, AA, VB Young, and RT Sauer,
Proc. Natl. Acad. Sci. USA (1986), 83:8829-8833. PALV79:
Palva, ET, and P Westermann, FEBS Letters (1979), 99:77-80. PAPA82:
Papamokos, E, E Weber, W Bode, R Huber, MW Empie, IKato, and M Laskowski Jr., J Mol Biol (1982), 158:515. PARD81:
Pardoe, IU, and ATH Burness, J Gen Virol (1981), 57:239-243. POTE83:
Poteete, AR, J Mol Biol (1983), 171:401-418. PRIV86:
Privalov, PL, YV Griko, SY Venyaminov, and VPKutyshenko, J Mol Biol (1986), 190(3)487-98. 195 QUI087:
Quiocho, FA, NK Vyas, JS Sack and MA Storey,in Crystallography in Molecular Biology. Moras, D. etal.. editors, Plenum Press, 1987. 5 RAOS87:
Rao SN, UC Singh, PA Bash, and PA KollmanNature (1987), 328 (6130) p551-4. 10 RASC86:
Rasched, I, and E Oberer,
Microbiol. Rev. (1986) 50:401-427. RASH84: 15 Rashin, A,
Biochemistry (1984), 23:5518. RAYC87:
Ray, C, KM Tatti, CH Jones, and CP Moran Jr, 20 J Baceriol (1987), 169(5)1807-1811. RAYG86:
Ray, GL, and WG Haldenwang, J Bact (1986), 166:472-78. 25 REID88:
Reidhaar-Olson, JF, and RT Sauer,
Science (1988), 241:53-57. 30 RICH81:
Richardson, JS,
Adv. Protein Chemistry (1981), 34:167-339. 196
<img img-format="tif" img-content="drawing" file="IL120940AD000213.tif" id="idf0013" />
RICH86:
Richards, JH,
Nature (1986), 323:187. 5 ROAM80:
Roa, M, and JM Clement, FEBS Letters (1980), 121:127-129. ROBE86: 10 Roberts, S, and AR Rees
Protein Engineering (1986), 1:59-65. RODR82:
Rodriguez, RL, 15 Gene (1982), 2£:305-316. ROSE85:
Rose, GD,
Methods in Enzymololgy (1985), 115(29}430-440. 20 ROSS81:
Rossman, M, and P Argos,
Ann. Rev. Biochem. (1981), 50:497ff. 25 RUSS81:
Russel, M, and P Model,
Proc. Natl. Acad. Sci. USA (1981), 78:1717-1721 SABB88: 30 Subbarao, MN, and D Kennel1, J Bact (1988), 170:2860-2865. 197 SAIK85:
Saiki, RK, S Scharf, F Faloona, KB Mullis, GT Horn, HAErlich, and N Arnheim,
Science (1985), 230:1350-1354. SALI64:
Salivar, WO, H Tzagoloff, and D Pratt,
Virology (1964), 24:359-71. SASA84:
Sasaki, T, FEBS Lett. (1984), 168:227-230. SCHA78:
Schaller, Η, E Beck, and M Takanami,
The Single-Stranded DNA Phages. Denhardt, D.T., D.Dressier, and D.S. Ray editors, Cold Spring HarborLaboratory, 1978., pl39-163. SCHA86:
Scharf, SJ, GT Horn, and HA Erlich,
Science (1986), 233:1076-1078. SCHO84:
Schold, M, A Colombero, AA Reyes, and RB Wallace,DNA (1984), 3(6)469-477. SCHU79:
Schulz, GE, and RH Schirmer,
Principles of Protein structure.
Springer-Verlag, New York, 1979. 198 SCHW87:
Schwarz, H, HJ Hinz, A Mehlich, H Tschesche, and HR
Wenzel,
Biochemistry (1987), 26:(12)p3544-51. 5 SCOT87:
Scott, MJ, CS Huckaby, I Kato, WJ Kohr, M LaskowskiJr., M-J Tsai and BW O'Malley, J Biol chem (1987), 262(12)5899-5907. 10 SERW87:
Serwer, P, J. Chromatography (1987), 418:345-357. 15 SHOR81:
Shortle, D, D DiMaio, and D Nathans,
Ann. Rev. Genet. (1981), 15:265-294. SHOR85: 20 Shortle, D, and B Lin,
Genetics (1985), 110:539-555. SMIT85:
Smith GP, 25 Science (1985), 228:1315-1317. SMIT87a:
Smith M,
Protein Structure, Folding, and Design 2, 1987. 30 Ed. D Oxender (New York, AR Liss Inc.) p.395ff. 199 SMIT87b:
Smith, H, S Bron, J van Ee, and G Venema, J Bacteriol. (1987), 169:3321-3328. STAT87:
States, DJ, TE Creighton, CM Dobson, and M Karplus, J Mol Biol (1987), 195: (3) p731-9. STRY81:
Strydom, DJ, and FJ Joubert,
Hoppe-Seyler's Z. Physiol. Chem. (1981), 362:1377-1384 SUDH85:
Sudhof, TC, JL Goldstein, MS Brown, and DW Russell,Science (1985), 228:815-822. SURE87:
Surewicz, WK, AG Szabo, HH Mantsch,
Eur J Biochem (1987), 167(31519-523. SUTC87a:
Sutcliffe, MJ, I Haneef, D Carney, and TL Blundell,Protein Engineering (1987), 1.:377-384. SUTC87b:
Sutcliffe, MJ, FRF Hayes, and TL Blundell,Protein Engineering (1987), 1:385-392. SUZU83:
Suzuki, T and K Shikama,
Arch Biochem Biophys (1983), 224(21695-9. 200 TAKA74:
Takahashi, H, S Iwanage, T Kitagawa, Y Hokama, and T
Suzuki, J Biochem (1974), 76:721-733. TANK77:
Tan, NH, and ET Kaiser,
Biochemistry (1977), 16:1531-1541. THER88:
Theriault, NY, JB Carter, and SP Pulaski,
BioTechniques (1988), 6(5)470-473. THOR88:
Thornton, JM, BL Sibinda, MS Edwards, and DJ Barlow,Bioessays Feb-Mar 1988, 8(2) 63-9. TOTH86:
Toth MJ, and P Schimmel, J Biol. Chem. (1986), 261:6643-6646. TSCH87:
Tschesch, H, J Beckmann, A Mehlich, E Schnabel, ETruscheit, and HR Wenzel,
Biochimica et Biophysica Acta (1987), 913:97-101. ULME83:
Ulmer, KM
Science (1983), 219(4585)666-71. VITA84:
Vita, C, D Dalzoppo, and A Fontana,
Biochemistry (1984), 23:5512-5519. 201 WACH80:
Wachter, Ε, K Deppner, and K Hochstrasser,FEBS Letters (1980), 119:58-62. WAGN78:
Wagner, G, K Wuthrich, and H Tschesche,
Eur J Biochem (1978), 89:367-377. WANG87:
Wagner, G, D Bruhwiler, and K Wuthrich, J,Mol Biol (1987), 196:(1) p227-31. WAIT83:
Waite, JH, J Biol chem (1983), 258(5)2911-5. WAIT85:
Waite, JH, TJ Housley, and ML Tanzer,Biochemistry (1985), 24(19)5010-4. WAIT86:
Waite, JH, J Comp Physiol [B] (1986), 156(4)491-6. WAND79:
Wandersman, C, M Schwartz, and T Ferenci,J Bact (1979), 140 (1) pl-13. WARD86:
Ward, WH, DH Jones, and AR Fersht,J Biol Chem (1986), 261(21)9576-8. 202 WATS87:
Molecular Biology of the Gene, Fourth Edition.
Watson, JD, NH Hopkins, JW Roberts, JA Steitz, and AMWeiner,
Benjamin/Cummings Publishing Company, Inc., Menlo Park,CA., 1987. WEBS78:
Webster, RE, and JS Cashman,
The Single-Stranded DNA Phages. Denhardt, DT, D Dressier, and DS Ray editors, Cold Spring HarborLaboratory, 1978., p557-569. WELL87a:
Wells, JA, BC Cunningham, TP Graycar, and DA Estell,Proc. Natl. Acad. Sci. USA (1987), 84:5167-5171. WELL87b:
Wells, JA, DB Powers, RR Bott, TP Graycar, and DAEstell,
Proc. Natl. Acad. Sci. USA (1987), 84:1219-1223. WETZ86:
Wetzel, R,
Protein Engineering (1986), 1:3-6. WHAR86:
Wharton, RP,
The Binding Specificity Determinants of 434 Repressor..
Harvard U. PhD Thesis, 1986,
University Microfilms, Ann Arbor, Michigan. WILK84:
Wilkinson, AJ, AR Fersht, DM Blow, P Carter, and GWinter, 203
Nature (1984), 307:187-188. WINT87:
Winter, RB, L Morrissey, P Gauss, L Gold, T Hsu, and JKaram,
Proc Natl Acad Sci USA (1987), 84:7822-6. WISH75:
Wishner, BC, KB Ward, EE Lattman, and WE Love, J Mol Biol (1975), 98:179-194. WISH76:
Wishner, BC, JC Hanson, WM Ringle, and WE Love,
Proc. of the Svmp. on Molecular Cellular Aspects of
Sickle Cell Disease. DHEW Publication 76-1007, NatlInst Health, Bethesda, Md., pl-31. WLOD84:
Wlodawer, A, J Walter, R Huber, and L Sjolin, J Mol Biol (1984), 180: (2) p301-29. WLOD87a:
Wlodawer, A, J Nachman, GL Gilliland, W Gallagher, andC Woodward, J Mol Biol (1987), 198 (3) p469-80. WLOD87b:
Wlodawer, A, J Deisenhofer, and R Huber, J Mol Biol (1987), 193: (1) pl45-56. 204 YAGE87:
Yager, TD, and PH von Hippel,
Volume 2, Chapter 76, p 1241-1275,
Escherichia coli and Salmonella typhimurium: Cellular 5 and Molecular Biology.
Neidhardt, FC, Editor-in-Chief,
Amer. Soc. for Microbiology, Washington, DC, 1987. ZIMM82: 10 Zimmermann, R, C Watts, and W Wickner, J. Biol. chem. (1982), 257:6529-6536. ZOLL84:
Zoller, MJ, and M Smith, 15 DNA (1984), 3(6)479-488. MESS83:
Messing, J,
Methods in Enzymology (1983), 101:20-78 20 YAMA70:
Yamamoto, KR, BM Alberts, R Benzinger, L Lawhorne, andG Treiber,
Virology (1970), £0:734-744 ♦ Β 205
Table 2: Preferred Outer-Surface Proteins
Preferred
Genetic Outer-Surface
Package Protein Reason for preference Ml 3 coat protein a) exposed amino terminus, (gpVIII) b) predictable post- translational processing, c) numerous copies in virion. gp III a) fusion data available. PhiX174 G protein a) known to be on virionexterior, b) small enough that the G-ipbd gene can replace H gene. E. coli LamB a) fusion data available, b) non-essential. B. subtilis spores CotC a) no post-translationalprocessing, b) distinctive sdequencethat causes protein tolocalize in spore coat, c) non-essential.
CotD
Same as for CotC. 206
Table 7: Atomic radiiAngstroms calpha°carbonylNamideOther atoms 1.70 1.52 1.55 1.80
Table 8
Fraction of DNA molecules havingn non-parental bases when M .9965 reagents that have fraction .97716 M of parental nt. .92612 .8577 .79433 .63096 fO .9000 .5000 .1000 .0100 .0010 .000001 fl .09499 .35061 .2393 .04977 .00777 .0000175 f2 .00485 .1188 .2768 .1197 .0292 .000149 f3 .00016 .0259 .2061 .1854 .0705 .000812 f4 . 000004 .00409 .1110 .2077 .1232 .003207 f8 0. 2XlO~7 .00096 .0336 .1182 .080165 fl6 0. 0. 0. 5X10-7 .00006 .027281 f23 0. 0. 0. 0. 0. .0000089 most 0 0 2 5 7 12 "most" is the value of n having theprobability. highest 207
Table 9: best vgCodon
Program "Find Optimum vgCodon."
5 INITIALIZE-MEMORY-OF-ABUNDANCES DO ( tl = 0.21 to 0.31 in steps of 0.01 ) . DO ( cl = 0.13 to 0.23 in steps of 0.01 ) . . DO ( al = 0.23 to 0.33 in steps of 0.01 )
Comment calculate gl from other concentrations10 . . . gl = 1.0 - tl - cl - al . . . IF( gl .ge. 0.15 ) . . . . DO ( a2 = 0.37 to 0.50 in steps of 0.01 ) ..... DO ( c2 = 0.12 to 0.20 in steps of 0.01 )
Comment Force D+E = R + K15 ...... g2 = (gl*a2 -.5*al*a2)/(cl+0.5*al)
Comment Calc t2 from other concentrations. ......t2 = 1. - a2 - c2 - g2 ...... IF(g2.gt. 0.1.and. t2.gt.0.1)
...... . CALCULATE-ABUNDANCES
20 ....... COMPARE-ABUNDANCES-TO-PREVIOUS-ONES ........end_IF_block .......end_DO_loop ! c2 ......end_DO_loop ! a2 .....end_IF_block i if gl big enough 25 . . ..end_DO_loop ! al . ..end_DO_loop ! cl..end_DO_loop 1 tl WRITE the best distribution and the abundances. 207a-
Table 10: Abundances obtained from optimum vgCodon Amino Amino acid Abundance acid Abundance A 4.80% C 2.86% D 6.00% E 6.00% F 2.86% G 6.60% H 3.60% I 2.86% K 5.20% L 6.82% M 2.86% N 5.20% P 2.88% Q 3.60% R 6.82% s 7.02% mfaa T 4.16% V 6.60% w 2.86% lfaa Y 5.20% stop 5.20% ratio = Abun(W)/Abun(S) = 0.4074 i (1/ratio)j (ratio)3 stop-free 1 2.454 .4074 .9480 2 6.025 .1660 .8987 3 14.788 .0676 .8520 4 36.298 .0275 .8077 5 89.095 .0112 .7657 6 218.7 4.57 X 10“3 .7258 7 536.8 1.86 x 10“3 .6881 lfaa = least - favored amino-acid mfaa = most - favored amino-acid 207b
Table 11: Calculate worst codon.
Program "Find worst vgCodon within Serr of given 5 distribution." INITIALIZE-MEMORY-OF-ABUNDANCES Comment Serr is % error level. READ Serr Comment Tli,Cli,Ali,Gli, T2i,C2i,A2i,G2i, T3i,G3 10 Comment are the intended nt-distribution. READ Tli, Cli, Ali, Gli READ T2i, C2i, A2i, G2i READ T3i, G3i Fdwn = l.-Serr 15 Fup = l.+serr DO ( tl = Tli*Fdwn to Tli*Fup in 7 steps) . DO ( cl = Cli*Fdwn to Cli*Fup in 7 steps) . . DO ( al = Ali*Fdwn to Ali*Fup in 7 steps) • · · gl = 1. - tl - cl - al 20 • · · IF( (gl-Gli)/Gli .It. -Serr)
Comment 25
Comment 30 gl too far below Gli, push it back. gl = Gli*Fdwn . factor = (l.-gl)/(tl + cl + al) . tl = tl*factor . cl = cl*factor . al = al*factor ..end_IF_block IF( (gl-Gli)/Gli . gt. Serr) gl too far above Gli, push it back . gl = Gli*Fup . factor = (l.-gl)/(tl + cl + al) . tl = tl*factor. cl = cl*factor. al = al*factor.. end IF block 35
<img img-format="tif" img-content="drawing" file="IL120940AD000214.tif" id="idf0014" />
207c
Table 11, continued. 5 . . . DO ( a2 = A2i*Fdwn to A2i*Fup in 7 steps) Table 11, continued. • · · DO ( c2 = C2i*Fdwn to C2i*Fup in 7 steps) Comment . DO (g2=G2i*Fdwn to G2i*Fup in 7 steps)Calc t2 from other concentrations. 10 . . t2 = 1. - a2 - c2 - g2 . . IF( (t2-T2i)/T2i .It. -Serr) Comment t2 too far below T2i, push it back . . . t2 = T2i*Fdwn . . . factor = (l.-t2)/(a2 + c2 + g2) 15 . . . a2 = a2*factor . . . c2 = c2*factor . . . g2 = g2*factor . . ..end IF block . . IF( (t2-T2i)/T2i .gt. Serr) 20 Comment t2 too far above T2i, push it back . . . t2 = T2i*Fup . . . factor = (l.-t2)/(a2 + c2 + g2)Table 11, continued. 25 . . . a2 = a2*factor . . . c2 = c2*factor . . . g2 = g2*factor . . ..end IF block . . IF(g2.gt. 0.0 .and. t2.gt.0.0) 30 . . . t3 = 0.5*(1.-Serr) . . . g3 = 1. - t3 . . . CALCULATE-ABUNDANCES . . . COMPARE-ABUNDANCES-TO-PREVIOUS-ONES . . . t3 = 0.5 35 . . . g3 = 1. - t3 208
Table 11, continued.
....... CALCULATE-ABUNDANCES
....... COMPARE-ABUNDANCES-TO-PREVIOUS-ONES 5 .......t3 = 0.5*(l.+Serr) ...... . g3 = 1. - t3
....... CALCULATE-ABUNDANCES
Table 11, continued.
10 ....... COMPARE-ABUNDANCES-TO-PREVIOUS-ONES ........end_IF_block .......end_DO_loop I g2 ......end_DO_loop ! c2 .....end_DO_loop I a2 15 . . ..end_DO_loop ! al . ..end_DO_loop ! cl..end_DO_loop ! tl WRITE the WORST distribution and the abundances. 209
Table 12: Abundances obtainedusing optimum vgCodon assuming5% errors
Amino acid Abundance Amino acid Abundanc A 4.59% C 2.76% D 5.45% E 6.02% F 2.49% lfaa G 6.63% H 3.59% I 2.71% K 5.73% L 6.71% M 3.00% N 5.19% P 3.02% Q 3.97% R 7.68% mfaa S 7.01% T 4.37% V 6.00% W stop 3.05% 5.27% Y 4.77% ratio = Abun(F)/Abun(R) = 0.3248 (1/ratio)j (ratio)3 stop-free 3.079 .3248 .9473 9.481 .1055 .8973 29.193 .03425 .8500 89.888 .01112 .8052 276.78 3.61 X 10"3 .7627 852.22 1.17 x 10"3 .7225 624.1 3.81 X 10"4 .6844 210
Table 13: BPTI Homologues
R # 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 -3 F Z -2 - - - Q T - - - - - - Q - - - H G Z - -1 - - - T E - - - - - - P - - - D D G - 1 R R R P R R R R R R R L A R R R K R A 2 P P P P P P P P P P P R A P P P R P A 3 D D D D D D D D D D D K K D R T D S K 4 - F F F L F F F F F F F L Y F F F I F Y 5 C C C C C C C C C C C C C C C C c C C 6 L L L Q L L L L L L L I K E E N R N K 7 E E E L E E E E E E E L L L L L L L L 8 P P P P P P P P P P P H P P P P P P P 9 P P P Q P P P P P P P R L A A P P A V 10 Y Y Y A Y Y Y Y Y Y Y N R E E E E E R 11 T T T R T T T T T T T P I T T S Q T Y 12 G G G G G G G G G G G G G G G G G G G 13 P P P P P P P P P P P R P L L R P P P 14 C T A C C C C C C C C C C C C C C C C 15 K K K K K V G A L I K Y K K K R K K K 16 A A A A A A A A A A A Q R A A G G A K 17 R R R A A R R R R R R K K Y R H R S K 18 I I I L M I I I I I I I I I I I L I F 19 I I I L I I I I I I I P P R R R P R P 20 R R R R R R R R R R R A s S S R R Q S 21 Y Y Y Y Y Y Y Y Y Y Y F F F F I Y Y F 22 F F F F F F F F F F F Y Y H H Y F Y Y 23 Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y 24 N N N N N N N N N N N N K N N N N N N 25 A A A S A A A A A A A Q W L R L P S W 26 K K K T K K K K K K K K K A A E A K K 27 A A A s A A A A A A A K A A A S S S A 28 G G G N G G G G G G G K K Q Q N R G K 29 L L L A F L L L L L L Q Q Q Q K M G Q 30 C C C C C C C C C C C C C c C c C C C 31 Q Q Q E E Q Q Q Q Q Q E L L L K E Q L 32 T T T P T T T T T T T G P Q E V S Q P 33 F F F F F F F F F F F F F F F F F F F 34 V V V T V V V V V V V T D I I F I I N 35 Y Y Y Y Y Y Y Y Y Y Y W Y Y Y Y Y Y Y 36 G G G G G G G G G G G S S G G G G G S 37 G G G G G G G G G G G G G G G G G G G 38 C T A C C C C C C C C C C C C C C C C 39 R R R Q R R R R R R R G G G G G K R G 40 A A A G A A A A A A A G G G G G G G G 41 K K K N K K K K K K K N N N N N N N N 42 R R R N S R R R R R R S A A A A K Q A 43 N N N N N N N N N N N N N N N N N N N 211
Table 13, continued. 1
N
F
K
S
A
E
D
C
M
R
T
C
G
G
A 2
N
F
K
S
A
E
D
C
M
R
T
C
G
G
A 3 4 5
N N N
F F F
K E K
STSA T A
E E E
D M D
C C C
M L M
R R R
TITC C C
G E G
G P G
APA - Q - - Q “ - T - - D - - K - - S -
6 7 8N N NF F FK K Ks s sAAAE E ED D DC C CΜ Μ MR R RT T TC C CG G GG G GAAA 9 10 11 12 13N N N R RF F F F FK K K K KS S S T TA A A I IE E E E ED D D E EC C C C CΜ Μ E R RR R R R RT T T T Tc c c c cG G G I VG G G R GA A A K -
14 15R RF FK KT TI ID DE EC CR HR RT TC CV VG G 16
N
F
E
T
R
D
E
C
R
E
T
C
G
G
K 17 18N RF FK DT TK TA QE QC Cv QR GA VC CR VP -P -E -R -P - 19
R
F
K
T
I
E
E
C
R
R
T
C
V
G = residue number
1 BPTI 2 Engineered BPTI From MARKS7 3 Engineered BPTI From MARKS7 4 Bovine Colostrum (DUFT85) 5 Bovine Serum (DUFT85) 6 Semisynthetic BPTI, TSCH87 7 Semisynthetic BPTI, TSCH87 8 Semisynthetic BPTI, TSCH87 9 Semisynthetic BPTI, TSCH87 10 Semisynthetic BPTI, TSCH87 11 Engineered BPTI, AUER87 12 Dendroaspis polvlepis polvlepis (Black mamba) venom I(DUFT85) 13 Dendroaspis polvlepis polvlepis (Black Mamba) venom K(DUFT85) 14 Hemachatus hemachates (Ringhals Cobra) HHV II(DUFT85) 15 Na~ia nivea (Cape cobra) NNV II (DUFT85) 16 Vipera russelli (Russel's viper) RW II (TAKA74) 17 Red sea turtle egg white (DUFT85) 18 Snail mucus (Helix pomania) (WAGN78) 19 Dendroaspis angusticeps (Eastern green mamba) C13 SI C3 toxin (DUFT85) 212
Table 13, continued. R # 20 21 22 23 24 25 26 27 28 29 30 31 32 33
-5 ---------- - - -D -4 _3
-2 Z - L Z
-1 P - Q D
1 R R H H
2 R P R P
3 K Y T K
4 L A F F
5 C C C C
6 I Ε K Y
7 L L L L
8 HIPP
9 R V A A
10 N A E D
11 P A P P
12 G G G G
13 R P P R
14 C C C C
15 Y Μ K K
16 D F A A
17 K F S H
18 I I I I
19 P S P P
20 A A A R
21 FFFF
22 Y Y Y Y
23 Y Y Y Y
24 N S N D
25 Q K W S
26 K G A A
27 K A A S
28 K N K N
29 Q K K K
30 C C C C
31 E Y Q N
32 R P L K
33 FFFF
34 D T Η I
35 W Y Y Y
36 S S G G
37 G G G G
38 C C C C
39 G R K P
40 G G G G
41 N N N N
42 S A A A
43 N N N N
RK---D N - - -R R I K TP P N Ε VK T G D AF F D S AC C C C C
Y N E Q NL L L L LP L P G PA P K Y VD Ε V S IP T V A RG G G G GR R P P PC C C C CL N R M RA A A G A
Y L R M FΜ I F T IP P P S QR A R R LF F Y Y W
Y Y Y F A
Y Y Y Y FN N N N DP S S G AA H S T VS L S S KN Η K M GK K R A KC C C C CE Q Ε Ε VK K K T LF F F F FI N I Q P
Y Y Y Y YG G G G GG G G G GC C C C CR G G M QG G G G GN N N N NA A A G GN N N N N
R R - Ε T
Q K - R T
R R R G D
Η Η P F L
R P D L P
D D F D I
C C C C C
D D L T E
K K E S Q
P P P P A
P P P P FG
D D Y V D
K T T T A
G K G G G
N I P P L
C C C C C
- - K R F
G Q A A G
P T K G Y
Y V M F M
R R I K K
A A R R L
F F Y Y Y
Y Y F N S
Y Y Y Y Y
D K N N N
T P A T Q
R S K R E
L A A T T
K K G K K
T R F Q N
C C C C C
K V Ε Ε E
A Q T P E
F F F. F F
Q R V K I
Y Y Y Y Y
R G G G G
G G G G G
C C C C C
D D K K Q
G G A G G
D D K N N
Η H S G D
G G N N N 213
Table 13, continued. R # 20 21 22 23 24 25 26 27 28 29 30 31 32 33 44 R R R N N N N N K N N N R R 45 F F F F F F F F F F F F Y F 46 K K S K K K H V Y K K R K S 47 T T T T T T T T S T S S S T 48 I I I W W I L E E E D A E L 49 E E E D D D E K K T H E Q A 50 E E K E E E E E E L L D D E 51 C C C C C C C C C C C C C C 52 R R R R R Q E L R R R M L E 53 R R H Q H R K Q E C C R D Q 54 T T A T T T V T Y E E T A K 55 C C C c C c c c C C C C C C 56 I V V G V A G R G L E G S I 57 G V G A A A V - V V L G G N 58 - - - S S K R - P Y Y A F - 59 - - - A G Y S - G P R - - - 60 - - - - I G - - D - — — — — 20 Dendroaspis anqusticeps (Eastern GreenMamba) C13 S2 C3 toxin (DUFT85) 21 Dendroaspis polvlepis polvlepes (Blackmamba) B toxin (DUFT85) 22 Dendroaspis polvlepis polvlepes (BlackMamba) E toxin (DUFT85) 23 Vipera ammodvtes TI toxin (DUFT85) 24 Vipera ammodvtes CTI toxin (DUFT85) 25 Bungarus fasciatus VIII B toxin (DUFT85) 26 Anemonia sulcata (sea anemone) 5 II(DUFT85) 27 Homo sapiens HI-14 "inactive" domain(DUFT85) 28 Homo sapiens HI-14 "active" domain(DUFT85) 29 beta bungarotoxin BI (DUFT85) 30 beta bungarotoxin B2 (DUFT85) 31 Bovine spleen ΤΙ II (FIOR85) 32 Tachypleus tridentatus (Horseshoe crab)hemocyte inhibitor (NAKA87) 33 Bombvx mori (silkworm) SCI-III (SASA84)
Notes : a) both beta bungarotoxins have residue 15 deleted. b) B. mori has an extra residue between C5 and C14; wehave assigned F and G to residue 9. c) all natural proteins have C at 5, 14, 30, 38, 50, &amp; 55. d) all homologues have F33 and G37. e) extra C's in bungarotoxins form interchain cystinebridges 214
Table 14: Tally of Ionizable Groups.BPTI homologues.
Sequence
Identifier 1 2 3 4 5 6 7 89 10 11 12 13 14 15 16 17 18 19 202122 23 24 25 26 27 28 29 30 31 32 33 D E K2 2 42 2 42 2 42 4 22 4 42 2 32 2 32 2 32 2 32 2 32 3 40 3 712 82 3 2 14 22 5 32 4 61120 2 9 2 3 60 3 30 2 64 15 3 2 412 5 15 414 22 3 46 2 56 2 6 2 3 5 3 3 5 4 7 3
R Y H 6 4 0 6 4 0 6 4 0 3 3 0 4 4 0 6 4 0 6 4 0 6 4 0 6 4 0 6 4 0 6 4 0 7 3 1 5 4 0 5 3 1 7 2 2 7 3 2 7 3 0 4 4 0 4 4 0 7 3 1 5 5 0 3 3 2 3 4 2 6 5 1 3 3 1 4 4 1 2 4 0 3 3 0 7 4 2 7 4 2 4 4 0 5 4 0 14 0 NH CO21 11 11 11 11 11 11 11 11 11 11 11 11 11 11 11 11 11 11 11 11 11 1 1 11 11 11 11 11 11 11 11 11 11 1 + #6 166 166 16 -1 13 2 165 15 5 15 5 15 5 15 5 15 5 19 11 19 10 18 2 14 4 16 3 19 7 21 4 8 11 17 8 20 5 13 7 13 3 15 5 17 5 13 2 16 -1 112 14 4 22 5 23 4 16 4 18 -7 17
Sequences given in Table 10. + is sum ofK+R+NH-D-E- C02, approximate charge onmolecule at pH 7.0 # is sum ofK+R+NH+D+E+ CO2, i.e. number of ionizedgroups at pH 7.0. 215
Table 15: Amino acids observed at each ResidueBPTI homologues
Number
Different
Res. # AAs Contents -5 2 D -32 -4 2 E -32 -3 5 T P F Z -29 -2 10 Z3 R3 Q2 T2 H G L K E -18 -1 10 D4 T2 P2 Q2 E G N K R -18 1 10 R21 A2 K2 H2 P L I T G D 2 9 P20 R4 A2 H2 N Ε V F L 3 10 D15 K6 T3 R2 P2 S Y G A L 4 7 F19 D4 L3 Y2 12 A2 S 5 1 C33 6 10 Lll E5 N4 K3 Q2 12 Y2 D2 T R 7 5 L18 Ell K2 S Q 8 7 P26 H2 A2 I L G F 9 9 P17 A6 V3 R2 Q L K Y F 10 10 Yll E7 D4 A2 N2 R2 V2 S I D 11 10 T17 P5 A3 R2 I S Q Y V K 12 2 G32 K 13 5 P22 R6 L3 N I 14 3 C31 T A 15 12 K15 R4 Y2 M2 L2 -2 V G A I N 16 7 A22 G5 Q2 R K D F 17 12 R12 K5 A2 Y3 H2 S2 F2 L Μ T < 18 6 121 M4 F3 L2 V2 T 19 7 Ill P10 R6 S2 K2 L Q 20 5 R19 A7 S4 L2 Q 21 4 Y18 F13 W I 22 6 F14 Y14 H2 A N S 23 2 Y32 F 24 4 N26 K3 D3 S 25 10 A12 S5 Q3 P3 W3 L2 T2 K G R 26 9 K16 A6 T2 E2 S2 R2 G Η V 27 5 A18 S8 K3 L2 T2 28 7 G13 K10 N5 Q2 R Η M 29 10 L9 Q7 K7 A2 F2 R2 M G T N 30 1 C33 31 7 Q12 Ell L4 K2 V2 Y N 32 11 T12 P5 K4 Q3 E2 L2 G V S R A 33 1 F33 34 11 Vll 18 T3 D2 N2 Q2 F Η P R K 35 2 Y31 W2 36 3 G27 S5 R 37 1 G33 38 3 C31 T A 39 7 R13 G9 K4 Q3 D2 P M
BPTI 216
Table 15: continued.
Number
Different
Res. # AAs Contents 40 2 G22 All 41 3 N20 Kll D2 42 9 All R9 S4 G3 H2 D Q K N 43 2 N31 G2 44 3 N21 Rll K 45 2 F32 Y 46 8 K24 E2 S2 D Η V Y R 47 2 T19 S14 48 9 All 19 E4 T2 W2 L2 R K D 49 7 E19 D6 A2 Q2 K2 T H 50 6 E16 D12 L2 M Q K 51 1 C33 52 7 R13 MIO L3 E3 Q2 Η V 53 8 R21 Q3 E2 H2 C2 G K D 54 7 T23 A3 V2 E2 I Y K 55 1 C33 56 8 G15 V8 13 E2 R2 A L S 57 8 G19 V4 A3 P2 -2 R L N 58 8 All -10 P3 K3 S2 Y2 R F 59 9 -24 G2QEAYSPR 60 6 -28 Q R I G D 61 3 -31 T P 62 2 -32 D 63 2 -32 K 64 2 -32 S 217
Table 16: Exposure in BPTICoordinates taken from
Brookhaven Protein Data Bank entry 6PTI.
HEADER
COMPND
COMPND
AUTHOR
PROTEINASE INHIBITOR (TRYPSIN) 13-MAY-87BOVINE PANCREATIC TRYPSIN INHIBITOR 2(/BPTI$,CRYSTAL FORM /111$)
A.WLODAWER
6PTI
Solvent radius = 1.40
Atomic radii given in Table 7Areas in Angstroms-squared.
Residue Total area Not Not coveredat all fraction Coveredby M/C fraction ARG 1 342.45 205.09 0.5989 152.49 0.4453 PRO 2 239.12 92.65 0.3875 47.56 0.1989 ASP 3 272.39 158.77 0.5829 143.23 0.5258 PHE 4 311.33 137.82 0.4427 43.21 0.1388 CYS 5 241.06 48.36 0.2006 0.23 0.0010 LEU 6 280.98 151.45 0.5390 115.87 0.4124 GLU 7 291.39 128.91 0.4424 90.39 0.3102 PRO 8 236.12 128.71 0.5451 99.98 0.4234 PRO 9 236.09 109.82 0.4652 45.80 0.1940 TYR 10 330.97 153.63 0.4642 79.49 0.2402 THR 11 249.20 80.10 0.3214 64.99 0.2608 GLY 12 184.21 56.75 0.3081 23.05 0.1252 PRO 13 240.07 130.25 0.5426 75.27 0.3136 CYS 14 237.10 75.55 0.3186 53.52 0.2257 LYS 15 310.77 200.25 0.6444 192.00 0.6178 ALA 16 209.41 66.63 0.3182 45.59 0.2177 ARG 17 351.09 243.67 0.6940 201.48 0.5739 ILE 18 277.10 100.51 0.3627 58.95 0.2127 ILE 19 278.03 146.06 0.5254 96.05 0.3455 ARG 20 339.11 144.65 0.4266 43.81 0.1292 TYR 21 333.60 102.24 0.3065 69.67 0.2089 PHE 22 306.08 70.64 0.2308 23.01 0.0752 TYR 23 338.66 77.05 0.2275 17.34 0.0512 ASN 24 264.88 99.03 0.3739 38.69 0.1461 ALA 25 211.15 85.13 0.4032 48.20 0.2283 LYS 26 313.29 216.14 0.6899 202.84 0.6474 ALA 27 210.66 96.05 0.4560 54.78 0.2601 GLY 28 186.83 71.52 0.3828 32.09 0.1718 LEU 29 280.70 132.42 0.4718 93.61 0.3335 CYS 30 238.15 57.27 0.2405 19.33 0.0812 GLN 31 301.15 141.80 0.4709 82.64 0.2744 THR 32 251.26 138.17 0.5499 76.47 0.3043 218
Table 16, continued. PHE 33 304.27 59.79 0.1965 18.91 0.0622 VAL 34 251.56 109.78 0.4364 42.36 0.1684 TYR 35 332.64 80.52 0.2421 15.05 0.0452 GLY 36 187.06 11.90 0.0636 1.97 0.0105 GLY 37 185.28 84.26 0.4548 39.17 0.2114 CYS 38 234.56 73.64 0.3139 26.40 0.1125 ARG 39 417.13 304.62 0.7303 250.73 0.6011 ALA 40 209.53 94.01 0.4487 52.95 0.2527 LYS 41 314.60 166.23 0.5284 108.77 0.3457 ARG 42 349.06 232.83 0.6670 179.59 0.5145 ASN 43 266.47 38.53 0.1446 5.32 0.0200 ASN 44 269.65 91.08 0.3378 23.39 0.0867 PHE 45 313.22 69.73 0.2226 14.79 0.0472 LYS 46 309.83 217.18 0.7010 155.73 0.5026 SER 47 224.78 69.11 0.3075 24.80 0.1103 ALA 48 211.01 82.06 0.3889 31.07 0.1473 GLU 49 286.62 161.00 0.5617 100.01 0.3489 ASP 50 299.53 156.42 0.5222 95.96 0.3204 CYS 51 238.68 24.51 0.1027 0.00 0.0000 MET 52 293.05 89.48 0.3054 66.70 0.2276 ARG 53 356.20 224.61 0.6306 189.75 0.5327 THR 54 251.53 116.43 0.4629 51.64 0.2053 CYS 55 240.40 69.95 0.2910 0.00 0.0000 GLY 56 184.66 60.79 0.3292 32.78 0.1775 GLY 57 106.58 49.71 0.4664 38.28 0.3592 ALA 58 no position given in Protein Data Bank "Total area" "Not coveredby M/C" "Not coveredat all" is the area measured by a rolling sphereof radius 1.4 A, where only the atomswithin the residue are considered. Thistakes account of conformation. is the area measured by a rolling sphereof radius 1.4 A where all main-chain atomsare considered, fraction is the exposedarea divided by the total area. Surfaceburied by main-chain atoms is moredefinitely covered than is surface coveredby side group atoms. is the area measured by a rolling sphereof radius 1.4 A where all atoms of theprotein are considered. 219
Phage LG1 pLG2 pLG3 pLG4 pLG5 pLG6 pLG7 pLG8 pLG9
pLGlO pLGll
Table 17: Plasmids used in Detailed Example
Contents M13mpl8 with Ava II/Aat II/Acc I/RsrΙΙ/Sau I adaptor LG1 with ampR and ColEl of pBR322 clonedinto Aat II/Acc I sitespLG2 with Acc I site removedpLG3 with first part of osp-pbd genecloned into Rsr ΙΙ/Sau I sites,
Avr II/Asu II sites createdpLG4 with second part of osp-pbd genecloned into Avr II/Asu II sites, BssH Isite created pLG5 with third part of osp-pbd genecloned into Asu II/BssH I sites, Bbe Isite created pLG6 with last part of osp-pbd genecloned into Bbe I/Asu II sitespLG7 with disabled osp-pbd gene, samelength DNA. pLG7 mutated to display BPTI(V15BpTI)pLG8 + tetR gene - ampR genepLG9 + tetR gene - ampR gene 220
Table 25: Annotated Sequence of ipbd gene 5'- C|GGA|CCG|TAT|CCA|GGC|TTT|ACA|CTT|TAT| 28
I Rsr II I I -35 I | GCT|TCC|GGC|TCG|TAT|AAT|GTG|TGG| 52 1-.-10 | I AATITGTI GAG ICGGI ΑΤΑ IACAI ATT I 73 | lac operator_|_ I CCTIAGGIAGGICTCI ACT I 88
| Avr III
| s. D. I |m|k|k|s|l|v|ljk|a|s| I 1 I 2 I 3 I 4 I 5 I 6 I 7 I 8 I 9 I 101 I ATGIAAGI AAA|TCT|CTG|GTT|CTT|AAG|GCT|AGC| 118
I Afl II, Nhe I I |v|a|v|a|t|l|v|p|m|l| I 111 12 I 13 I 14 I 151 161 17| 181 191 201 I GTTIGCTIGTCIGCG|ACC|CTG|GTA|CCG|ATG|CTG| 148
I Nru II I Kpn II |s|f|a|r|p|d|f|c|l|e| I 21, 221 231 24, 251 26, 271 28, 291 30, | TCT|TTT,GCT|CGT|CCG|GAT|TTC|TGT|CTC|GAG| 178 221
Table 25, continued.|AccIIII I Ava I I
| Xho I I |p|p|y|t|g|p|c|k|a|r|I 31 j 321 33 I 341 35| 36| 371 381 391 40||CCG|CCA|TAT|ACT|GGG|CCC|TGC|AAA|GCG|CGC| | PflM I_|_ IBSSH II | J-...Apa I j. | Dra II |
I Pss I I |i|i|r|y|f|y|n|a|k| I 4l| 42 I 431 441 45j 46| 471 481 491|ATC|ATC|CGT|TAT|TTC|TAC|AAC|GCT|AAA| 208 235 222
Table 25, continued. |a|g|l|c|q|t|f|v|y|g|g|| 50| 51| 52| 53| 54| 55| 56| 57| 58| 59| 60||GCA|GGC|CTG|TGC|CAG|ACC|TTT|GTA|TAC|GGT|GGT| [ Stu II I Acc I | | Xca I | 268 |c|r|a|k|r|n|n|f|k|| 61) 62| 63 J 641 651 661 671 68| 691|TGC|CGT)GCT|AAG|CGT|AAC|AAC|TTT]AAA| I Esp I[ 295 |s|a|e|d(clm|r|t|c|g|I 70| 71| 72 I 73 I 74| 751 76) 771 78| 79||TCG|GCC|GAA|GAT|TGC|ATG|CGT|ACC|TGC[GGT|
IXmalllI I Sph II 325
|g|a|a|e|g|d|d|| 80| 81| 82| 83| 84| 85| 86)|GGC|GCC|GCT|GAA|GGT|GAT|GAT|I Bbe I I I Nar I [ 346 I P I a | k | a | a | | 87| 88) 89| 90| 9l| |CCG|GCC|AAA|GCG|GCC| 1 Sfi.I-1 361 223 |f|n|s|l|q|a|s|ajt|Table 25, continued. I 92 I 93 I 941 95| 96| 971 981 99|l00|| TTT|AAC|TCT|CTG|CAA|GCT|TCT|GCT|ACC|
I Hind 3 I |e|y|i|g|y|a|w|I 1011 102 I 103 I 10411051106|107|| GAA|TAT|ATC|GGT|TAC|GCG|TGG| I Mlu I| 388 409 J a | m | v | v | v |1108|109|110|111|112|| GCC|ATG|GTG|GTG|GTT|| BstX I_j.
| Neo II 424 224
Table 25, continued. |i|v|g|a|t|i|g|i( 1113 1114 1115 11161117111811191120| | ATC|GTT|GGT|GCT|ACC|ATC|GGT|ATC| 448 jk|l|f|k|k|f|t|s|k|a|I 1211122 I 123|124|125|126|127|128|129|130|| AAA|CTG|TTT|AAG|AAA|TTT|ACT|TCG|AAA|GCG|
IAsu III 478 I s | . | . j . | 1131|132 I 133|134|
| TCT|TAA|TAG|TGA|GGT|TAC|CAG|TCT|| BstE III 502 | AAG|CCC|GCC|TAA|TGA|GCG|GGC|TTT|TTT|TTT| | Trp terminator_[ 532 | CCT|GAG|G -3'I Sau I 1 539
Note the following en Xma III = Eaq I Acc III = BspM II Dra II = ECO0109 I Asu II = BstB I Sau I = BSU36 I 225
Table 27: DNA_synthl
5Z ICCGITCC1GTCIGGAICCGI TAT 1 CCA IGGCITTTIACA|CTT|TAT I
|GCTITCC|GGC|TCGI TAT|AATIGTG|TGGI
|AAT|TGT1 GAG ICGG|ATA|ACAI ATT I olig#4 = 3Z- gt taa ICCT1AGG| gga tcc / 3Z = olig#3IGCCIGCTICCTITCG( AAA|GCG|egg ega gga age ttt ege |TCTITAA|TAG|TGA|GGT|TAC(CAG(TCT|aga att ate act cca atg gtc aga J AAG|CCC|GCC|TAA|TGA|GCG|GGC|TTT|TTT|TTT|ttc ggg egg att act ege ccg aaa aaa aaa |CCT(GAG|GCA|GGT|GAG|CGgga etc cgt cca etc ge - 5z
<img img-format="tif" img-content="drawing" file="IL120940AD000215.tif" id="idf0015" />
226
Table 27, continued "Top" strand 99 "Bottom" strand 100 Overlap 23 (14 c/g and 9 a/t) Net length 158
<img img-format="tif" img-content="drawing" file="IL120940AD000216.tif" id="idf0016" />
227
Table 28: DNA_seq2 5'- |gca|cca|acg|| spacer j.
ICCTIAGGIAGG|CTC|ACT|I Avr III
| S. D. I lm|k|k|s|l|v|l|k|a|s|| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 ( 10|IATGIAAGI AAA|TCT|CTG|GTT|CTT|AAG|GCT[AGC|
I Afl II ( Nhe I I
|v|a|v|a|t|l|v|p|m|l|I 111 12 I 13 I 14 I 151 161 171 18| 191 201IGTTIGCTIGTCIGCG|ACC|CTG|GTA|CCG|ATG|CTG(I Nru II I Kpn II |s|f|a|r|p|d|f|c|l|e|| 21| 22| 23| 24| 25| 26| 27| 2S| 29| 30j1TCT|TTT|GCT|CGT|CCG|GAT|TTC|TGT j CTC|GAG|
|AccIII, [ Ava I II Xho I I |p|p|y|t|g|p|c|k|a|r|| 3l| 32| 33( 34| 35| 36| 37| 38| 39| 40||CCG|CCA|TAT|ACT|GGG |CCC|TGC|AAA|GCG|CGC|
I PflM I I iBssH III 228
Table 28, continued. | Apa I.J. | Dra II | | Pss I | | 4l| 42| 43||ate|ate|cgt( | t | s | k |
1127 I 128 1129 I |ACT|TCG|AAa|gcg|get|gcg| - 3'|Asu II | spacer_|_
<img img-format="tif" img-content="drawing" file="IL120940AD000217.tif" id="idf0017" />
229
Table 30: DNA_seq3 | 39| 40|5' - |ccc|tgc|aca|GCG|CGC|| spacer |BssH II| |iji|r|y|f|y|n|a|k|| 41| 42 J 43| 44, 45| 46, 471 481 491|ATC|ATC|CGT|TAT|TTC|TAC|AAC|GCT|AAA| |a|g|l,c|q|t|f,v|y|g|g|| 50, 51, 52| 53| 54, 55, 56, 57| 58| 59| 60,|GCA,GGC|CTG|TGC|CAG|ACC|TTT|GTA|TAC|GGT|GGT| I Stu II 1 Acc I II Xca I , |c|r|a|k|r|n,n|f|k|, 61, 62, 63 I 64, 65, 661 671 68| 691|TGC|CGT|GCT|AAG|CGT,AAC|AAC|TTT|AAA|
I gsp I L ,s|a|e|d,c|m|r|t|c|g|I 70| 71| 72 I 73 I 74| 75, 76, 77| 78| 79||TCG|GCC|GAA|GAT|TGC|ATG|CGT|ACC|TGC,GGT| ,Xmalll[ | Sph 11 I α I a |I 80 | 81,
<img img-format="tif" img-content="drawing" file="IL120940AD000218.tif" id="idf0018" />
230
Table 3¾ continued |GGC|GCC|get|gaa| | Bbe I | spacer
I Nar I I | t | s | k | |127|128|129| |ttt|acT|TCG|AAa|geg|teg|ccg| - 3|Asu II|
<img img-format="tif" img-content="drawing" file="IL120940AD000219.tif" id="idf0019" />
231
Table 32: DNA_seq4 |g|a|a|e|g|d|dl5' I 80| 81| 821 S31 841 85J 86|
IcctIcgcIcctIGGCIGCCIGCTIGAAIGGT) GAT|GAT|| spacer | Bbe I | | Nar I 1 | p | a | k j a | a |I 871 881 89| 90| 9l||CCG|GCC|AAA|GCG|GCC| -1 |f|n|s|l|q|a|s|a|t|| 92) 93) 94| 95) 96| 97| 98| 99|l00||TTT J AAC|TCT|CTG|CAA|GCT|TCT|GCT|ACC| |Hind 3| |e|y[i|g|y|a|w|I 1011102 I 103 I 104 I 1051106|107|IGAAI TAT IATC|GGT|TAC|GCG|TGG|
| Mlu II 1 a I m I v 1 V I v ]1108 1109 1110 I 111I 112 IIGCCIATG1GTG|GTG|GTT|| BstX I_|_ | Neo I|
<img img-format="tif" img-content="drawing" file="IL120940AD000220.tif" id="idf0020" />
232
Table 32, continued.|i|v|g|a|t|i|g|i| | 113|11411151116|117|118|1191120| | ATC|GTT|GGT|GCT|ACC|ATC|GGT|ATC| |k|l|f|k|k|f|t|s|k| I 1211 122 I 123|124|125|126|127|128|129| | AAA|CTG|TTT|AAG|AAA|TTT|ACT|TCG|AAa|gcg|teg|ggc| - 3'|Asu III spacer_|_
Number
Res. Diff.
<img img-format="tif" img-content="drawing" file="IL120940AD000221.tif" id="idf0021" />
233
Table 34: Some interaction sets in BPTI
# AAs Contents BPTI 1 2 3 4 5 -5 2 D -32 - -4 2 E -32 - -3 5 T P F Z -29 - -2 10 Z3 R3 Q2 T2 H G L K E -18 - -1 10 D4 T2 P2 Q2 E G N K R -18 - 1 10 R21 A2 K2 H2 P L I T G D R 5 2 9 P20 R4 A2 H2 N Ε V F L P s 5 3 10 D15 K6 T3 R2 P2 S Y G A L D 4 s 4 7 F19 D4 L3 Y2 12 A2 S F s 5 5 1 C33 C X X 6 10 Lll E5 N4 K3 Q2 12 Y2 D2 T R L 4 7 5 L18 Ell K2 S Q E S 4 8 7 P26 H2 A2 I L G F P 3 4 9 9 P17 A6 V3 R2 Q L K Y F P s 3 4 10 10 Yll E7 D4 A2 N2 R2 V2 S I D Y s s 4 11 10 T17 P5 A3 R2 I S Q Y V K T 1 s 3 4 12 2 G32 K G X X X
<img img-format="tif" img-content="drawing" file="IL120940AD000222.tif" id="idf0022" />
<img img-format="tif" img-content="drawing" file="IL120940AD000223.tif" id="idf0023" />
<img img-format="tif" img-content="drawing" file="IL120940AD000224.tif" id="idf0024" />
234
Table 34, continued. 5 P22 R6 L3 N I P 1 S 4 S 3 C31 T A C 1 S s 5 12 K15 R4 Y2 M2 L2 -2 V G A I N F K 1 s 3 4 s 7 A22 G5 Q2 R K D F A 1 s s s 5 12 R12 K5 A2 Y3 H2 S2 F2 L Μ T G P R 12 3 s 6 121 M4 F3 L2 V2 T I 1 s s 5 7 Ill P10 R6 > S2 K2 L Q I 1 : > 3 s 5 R19 A7 S4 L2 Q R sss 5 4 Y18 F13 W I Y 2 s s s 6 F14 Y14 H2 : A N S F s 3 4 2 Y32 F Y s s 4 N26 K3 D3 S N s 3 10 A12 S5 Q3 P3 W3 L2 T2 K G R A s s 9 K16 A6 T2 E2 S2 R2 G Η V K s 3 4 5 A18 S8 K3 L2 T2 A 2 3 4 7 G13 K10 N5 i Q2 R Η M G 2 s s 10 L9 Q7 K7 A2 : F2 R2 M G T N l ; > 3 1 C33 C X X X 7 Q12 Ell L4 K2 V2 Y N Q 2 3 4 11 T12 P5 K4 Q3 E2 L2 G V S R A T 2 3 S 1 F33 F XX X X 11 Vll 18 T3 D2 N2 Q2 F Η P R K V 1 2 3 s 2 Y31 W2 Y S S s 5 3 G27 S5 R G 1 1 G33 G X X 3 C31 T A C 1 S 5 7 R13 G9 K4 Q3 D2 P M R 1 4 s 235
Table 34, continued. 2 G22 All A s s 5 3 N20 Kll D2 K 4 s 9 All R9 S4 G3 H2 D Q K N R s 5 2 N31 G2 N s 3 N21 Rll K N s 2 F32 Y F s 8 K24 E2 S2 D Η V Y R K 5 2 ΤΪ9 S14 S s 5 9 All 19 E4 T2 W2 L2 R K D A 2 s s 7 E19 D6 A2 Q2 K2 T H E 2 s 6 E16 D12 L2 M Q K D s 5 1 C33 C X X 7 R13 MIO L3 E3 Q2 Η V M 2 s 8 R21 Q3 E2 H2 C2 G K D R s 5 7 T23 A3 V2 E2 I Y K T 5 1 C33 C X 8 G15 V8 13 E2 R2 A L S G 8 G19 V4 A3 P2 -2 R L N G 8 All -10 P3 K3 S2 Y2 R F A 9 -24 G2 Q E A Y S P R - 6 -28 Q R I G D - 3 -31 T P - 2 -32 D - 2 -32 K - 2 -32 S — indicates secondary set indicates in or close to surface but buried and/or highlyconserved. 236
Table 35:
Distances from Cbeta toTip of Side Groupin Angstroms
Amino Acid typeA C (reduced)
D
E
F
G
H
I
K
L
M
N
P
Q
R
S
T
VW
Y
Distance 0.0 1.8 2.4 3.5 4.3 4.0 2.55.1 2.63.82-4 2.4 3.56.0 1.51.51.55.35.7
Notes: These distances were calculated for standard model partswith all side groups fully extended. 237
Table 36: Distances, BPTI residue set #2Distances in Angstroms between Cbetas·Hypothetical Cbeta was added to each Glycine. R17 119 Y21 A27 G28 L29 Q31 T32 V34 A48 119 7.7 Y21 15.1 8.4 A27 22.6 17.1 12.2 G28 26.6 20.4 13.8 5.3 L29 22.5 15.8 9.6 5.1 5.2 Q31 16.1 10.4 6.8 6.8 10.6 6.8 T32 11.7 5.2 6.1 12.0 15.5 10.9 5.4 V34 5.6 6.5 11.6 17.6 21.7 18.0 11.4 8.2 A48 18.5 11.0 5.4 12.6 13.3 8.4 8.8 8.3 15.7 E49 22.0 14.7 8.9 16.9 16.1 12.2 13.9 13.3 19.8 5.5 M52 23.6 16.3 8.6 12.2 10.3 7.6 11.3 13.2 20.0 6.2 P9 14.0 11.3 9.0 12.2 15.4 13.3 7.9 9.2 8.7 13.9 Til 9.5 11.2 13.5 18.8 22.5 19.8 13.5 12.1 5.7 18.5 K15 7.9 14.6 20.1 27.4 31.3 27.9 21.4 18.1 10.3 24.6 A16 5.5 10.1 15.9 25.2 28.5 24.6 18.6 14.5 8.6 19.8 118 6.1 6.0 11.2 21.3 24.4 20.2 14.7 10.4 7.0 15.0 R20 10.6 5.9 5.4 16.0 18.5 14.6 9.8 6.9 7.8 10.2 F22 15.6 10.9 5.6 10.5 12.8 10.3 6.2 8.1 10.8 10.3 N24 19.9 14.7 9.4 4.1 7.3 6.1 4.8 10.0 14.7 11.4 K26 24.4 20.1 15.2 5.4 7.7 9.8 10.1 15.3 19.0 17.0 C30 18.9 12.1 4.6 8.8 9.5 5.3 5.9 8.2 14.9 4.9 F33 10.8 7.4 7.7 12.6 16.4 13.0 6.6 5.6 5.5 12.2 Y35 8.4 7.4 9.4 18.4 21.4 17.9 12.2 9.5 5.8 14.4 S47 17.6 10.6 6.6 17.3 17.9 13.4 12.6 10.4 15.9 5.3 D50 20.0 13.6 7.2 17.2 16.8 13.5 13.5 12.9 17.6 7.6 C51 18.9 12.2 4.0 12.1 12.2 8.8 8.8 9.7 15.3 5.4 R53 25.4 18.6 11.0 17.2 15.0 13.0 15.7 16.7 22.3 9.7 R39 15.4 16.9 17.1 24.9 27.2 24.9 20.1 18.7 13.8 22.3
<img img-format="tif" img-content="drawing" file="IL120940AD000225.tif" id="idf0025" />
238
Table 36, continued.
Distances in Angstroms between Cj3etas.Hypothetical Cketa was added to each Glycine. N24 E49 M52 P9 Til K15 A16 118 R20 F22 M52 6.1 P9 17.7 15.5 Til 22.1 21.5 7.2 K15 27.5 28.7 16.4 9.5 A16 22.2 24.2 14.9 9.8 6.2 118 17.4 19.5 12.2 9.5 10.4 4.9 R20 13.0 13.8 8.0 9.4 14.9 10.6 6.2 F22 13.8 11.4 4.1 10.6 19.1 16.3 12.7 6.9 N24 15.6 11.2 8.4 15.3 24.1 21.9 18.2 12.7 6.6 K26 20.9 15.7 12.1 18.6 27.9 26.6 23.3 18.1 11.6 5.9 C30 8.7 5.6 10.6 16.6 24.1 20.2 15.7 9.8 6.8 6.9 F33 16.5 15.4 4.2 7.1 15.0 12.8 9.6 6.1 5.6 9.3 Y35 17.2 17.8 7.8 5.8 11.0 7.6 4.9 4.3 8.8 14.8 S47 4.7 9.1 15.3 18.5 23.1 17.6 12.8 9.1 12.0 15.3 D50 5.5 7.7 14.7 18.6 24.2 19.2 14.7 9.9 11.0 14.7 C51 7.1 5.4 11.0 16.4 23.5 19.2 14.6 8.7 6.9 9.6 R53 6.3 5.6 17.9 23.1 29.6 24.8 20.3 15.0 13.8 15.5 R39 23.9 24.0 13.0 9.5 12.0 11.8 12.5 12.8 14.7 20.8 K26 C3 0 F33 Y35 S47 D50 C51 R53 C30 12.4 F33 13.9 10.1 Y35 19.5 13.5 6.4 S47 21.0 8.8 13.5 13.2 D50 20.1 8.6 14.3 13.7 5.0 C51 15.0 3.7 10.9 12.5 6.9 5.2 R53 19.9 9.9 18.2 18.8 9.4 5.8 7.4 R39 24.3 20.6 14.4 9.6 20.4 19.0 18.8 23.4 239
Table 37: vgDNA to vary BPTI set #2.1+ 5'- CAC1CCT g 35 GGG P 36 CCC C 37 TGC k 38 AAA a 39 GCG X 40 qfk 208 spacer Apa I + i X r y f y n a k 41 42 43 44 45 46 47 48 49 ATC qfk CGT TAT TTC TAC AAC GCT AAA 235 + ! + X 50 gfk g 51 GGt X 52 qfk c 53 TGC q 54 CAG t 55 ACC oligj 28= 3'- acg gtc tgg 78 nts
Overlap 12 (7 CG, 5 AT) / 3' = olig#27 72 nts I +
f X y g g 56 57 58 59 60 TTC qfk TAC GGT GGT aag **m atg cca cca 268
C r a k r n n f k 61 62 63 64 65 66 67 68 69 TGC CGT GCT AAG CGT AAC AAC TTT AAA
acg gca cga ttc gca ttg ttg aaa tttl-Esp I I 295
s X e d c m 70 71 72 73 74 75 TCT qfk GAG GAT TGC ATG
Li 4ΙΛ. ΟΛΟ ΟΛΧ XOO ΛΧΟ V gc **m etc eta acg tac gca ccc accI Rnh TI snarer 322 -5 k = equal parts of T and G; m = ei qual parts of C and A; q = (.26 T, .18 C, .2 6 A, and .30 G) ? f = (.22 T, .16 C, .40 A, and .22 G); * = complement of symbol above Residue 40 42 50 52 57 71 Possibilities 21 X 21 X 21 x 21 : X 21 X 21 = 8.6 X 107 Abundance x 10: of PPBD .768 .271 .459 . 671 .600 1 .459
Produce = 1.77 x 10“8
Parent = 1/(5.5 x 107) least favored = 1/(4.2 x 109)
Least favored one-amino-acid substitution from PPBD presentat 1 in 1.6 x 107 240
Table 38: Result of varying set#2 of BPTI 2.1
P P y 31 32 33 CCG CCA TAT Pi flM I
i Q 41 42 ATC CAG E g 50 51 GAG GGC c r 61 62 TGC CGT s W 70 71 TCG TGG g a 80 81 GGC GCC Bbe I Nar I r 43
CGT
L 52
CTG a 63
GCT
Esp e 72
GAA
1 e 29 30 CTC GAG Ava I Xho I
t g P c k a D 34 35 36 37 38 39 40 ACT GGG CCC TGC AAA GCG GAT
Apa I
Dra II
Pss I 178 208 y 44
TAT
C 53
TGC k 64
AAG
I d 73
GAT f 45
TTC q 54
CAG r 65
CGT
C 74
TGC y 46
TAC t 55
ACC n 66
AAC n 47
AAC f 56
TTT n 67
AAC
m75ATGSph I r 76
CGT a 48
GCT
S 57
TCG f 68
TTT t 77
ACC k 49
AAA y 58
TAC k 69
AAA
C 78
TGC 235 g 59
GGT g 79
GGT g 60
GGT 268 295 325 241 5'
Table 39: vgDNA to vary set#2 BPTI 2.2+
co aca cqc g 35 GGG P 36 CCC C 37 TGC X 38 mrA a 39 GCG D 40 GAT 1 spacer Apa I 208 + + + X 41 rwA Q 42 CAG X 43 rvk X 44 TwT f 45 TTC y 46 TAC n 47 AAC a 48 GCT k 49 AAA g 59 GGT g 60 GGT E 50 GAG + X 51 qfk L 52 CTG C 53 TGC + X 54 qfk + X 55 qfk f 56 TTT S 57 TCG y 58 TAC 91 ni :s o'. Lig#30 3' g cca cca
Overlap = 15 (11 CG, 4 AT) 235 268 /- 3' olig#29 94 nts
c r a k r n n f k 61 62 63 64 65 66 67 68 69 TGC CGT GCT AAG CGT AAC AAC TTT AAA
aeg gca ega ttc gca ttg ttg aaa tttt-Esp. I_L 295 +
s W X d C m 70 71 72 73 74 75 TCG TGG qfk GAT TGC ATG agc acc **m eta aeg tac gcg acc tgc -5'[ Sph I| spacer | k = equal parts of T and G; v = equal parts of C, A, and G; m = equal parts of C and A; r = equal parts of A and G; w = equal parts of A and T; q = (.26 T, .18 C,f = (.22 T, .16 C,* = complement of .26 A, and .30 G); .40 A, and .22 G); symbol above Residue 38 41 43 44 51 54 55 72 Possibilities 4 X 4 X 9 X 2 X 21 x 21 x 21 X 21 = 6.2 x 107
Abundance x 10 2.5 2.5 .833 5. .663 .397 .437 .602
Product = 2.3 x 10“8
Parent = 1/(4.4 x 107) least favored = 1/(1.25 x 109)Least favored one-amino-acid substitution from PPBD presentat 1 in 1.2 x 107 242
Table 40: Result of varying set#2 of BPTI
P 31 CCG P 32 CCA P: V Q 41 42 GTT CAG E F 50 51 GAG TTT C r 61 62 TGC CGT s W 70 71 TCG TGG g a 80 81 GGC GCC Bbe I Nar I y 33
TAT
N 43
AAT
L 52
CTG a 63
GCT
Esp
Q 72
CAG t 34
ACT
F 44
TTT
C 53
TGC k 64
AAG
I d 73
GAT g 35
GGG
P 36
CCC
Apa I f 45
TTC
S 54
TCT r 65
CGT c 74
TGC y 46
TAC
A 55
GCT n 66
AAC c 37
TGC n 47
AAC f 56
TTT n 67
AAC
m75ATGSph I r 76
CGT
E 38
GAG a 48
GCT
S 57
TCG f 68
TTT t 77
ACC
1 e 29 30 CTC GAG Xho I a D 39 40 GCG GAT k 49 AAA y g 58 59 TAC GGT k 69 AAA C g 78 79 TGC GGT g 60
GGT 178 208 235 268 295 325
<img img-format="tif" img-content="drawing" file="IL120940AD000226.tif" id="idf0026" />
243
Table 41: vg DNA set#2 of BPTI 2.3 5Z- ccr acre eta 1 29 CTC e 30 GAG 1 spacer Xhc 3 I P + X Y + X g P c E a + X 31 32 33 34 35 36 37 38 39 40 CCG vmcr TAT vmq GGG CCC TGC GAG GCG qfk V Q N + X f Y n a k 41 42 43 44 45 46 47 48 49 GTT CAG AAT Tdk TTC TAC AAC GCc AAq -3 7 67 nt ;s o: Lig#34 3' g atg ttg egg ttc olig#33
Overlap 178 208 71 nts 13 (7 CG, 6 AT) + X F + X c S + X f + X y g g 50 51 52 53 54 55 56 57 58 59 60 VAG TTT nTk TGC TCT qfk TTT qfk TAC GGT GGT btc aaa nam acg aga **m aaa **m atg cca cca 268
c r a k 61 62 63 64 TGC CGT GCT AAG C acg
gca cga ttc| Esp I gcg acc ggc| spacer | k = equal parts of T and G; m = equal parts of C and A; w = equal parts of A and T; n = equal parts of A,C,G,T; d = equal parts A,G,T; v = equal parts A,C,G; q = (.26 T, .18 C, .26 A, and .30 G); f = (.22 T, .16 c, .40 A, and .22 G); * = complement of symbol above
Residue 32 34 40 44 50 52 55 57 Possibilities 6 x 6 x 21 x 6 X 3 x 5 x 21 X 21 = 3 x 10 Abundance x 10 of PPBD 10/6 10/6 .545 10/6 10/3 30/8 .459 .701 product = 1.01 x 10 parent = 1/(1 x 107) least favored = 1/(4 x 108)
Least favored one-amino-acid substitution from PPBD presentat 1 in 3 x 107 244
Table 42: Result of varying set#2 of BPTI 2
1 e 29 30 CTC GAG Ava I Xho I 178
P E Y Q g P c E a A 31 32 33 34 35 36 37 38 39 40 CCG GAG TAT CAG GGG CCC TGC GAG GCG GCT Apa I V Q N W f Y n a k 41 42 43 44 45 46 47 48 49 GTT CAG AAT TGG TTC TAC AAC GCT AAA Q F M C S L f H Y g g 50 51 52 53 54 55 56 57 58 59 60 CAG TTT ATG TGC TCT CTT TTT CAT TAC GGT GGT c r a k r n n f k 61 62 63 64 65 66 67 68 69 TGC CGT GCT AAG CGT AAC AAC TTT AAA Esp I I s W Q d c m r t c g 70 71 72 73 74 75 76 77 78 79 TCG TGG CAG GAT TGC ATG CGT ACC TGC GGT l-Sph I| 208 235 268 295 325
g a 80 81 GGC GCC Bbe I Nar I
Contents231
140 members in 11 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 24016088 | United States of America | A | |
| 9150189 | Israel | A | |
| 24016088 | – | – | – |
| 91501 | – | – | – |
| IL19890091501 | – | – | – |
| US19880240160 | – | – | – |
Members140
| Document | Office | Kind | |
|---|---|---|---|
| WO9002809A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU4308689A | Australia | A | |
| IL91501D0 | Israel | D0 | |
| EP0436597A1 | European Patent Office (EPO) | A1 | |
| WO9206191A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU8740491A | Australia | A | |
| EP0436597A4 | European Patent Office (EPO) | A4 | |
| JPH04502700A | Japan | A | |
| CA2105300A1 | Canada | A1 | |
| CA2105303A1 | Canada | A1 | |
| CA2105304A1 | Canada | A1 | |
| WO9215605A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO9215677A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO9215679A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU1545692A | Australia | A | |
| AU1578792A | Australia | A | |
| AU1581692A | Australia | A | |
| WO9215605A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US5223409A | United States of America | A | |
| EP0573603A1 | European Patent Office (EPO) | A1 | |
| EP0573611A1 | European Patent Office (EPO) | A1 | |
| EP0575485A1 | European Patent Office (EPO) | A1 | |
| JPH06510522A | Japan | A | |
| JPH07501203A | Japan | A | |
| JPH07501923A | Japan | A | |
| US5403484A | United States of America | A | |
| CA2207820A1 | Canada | A1 | |
| WO9620278A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO9620278A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US5571698A | United States of America | A | |
| EP0436597B1 | European Patent Office (EPO) | B1 | |
| AT151110T | Austria | T | |
| ATE151110T1 | Austria | T1 | |
| EP0768377A1 | European Patent Office (EPO) | A1 | |
| DE68927933D1 | Germany | D1 | |
| DE68927933T2 | Germany | T2 | |
| US5663143A | United States of America | A | |
| IL120939D0 | Israel | D0 | |
| IL120940D0 | Israel | D0 | |
| IL120941D0 | Israel | D0 | |
| EP0797666A2 | European Patent Office (EPO) | A2 | |
| DE768377T1 | Germany | T1 | |
| IL91501A | Israel | A | |
| JPH10510996A | Japan | A | |
| US5837500A | United States of America | A | |
| CA1340288C | Canada | C | |
| ES2124203T1 | Spain | T1 | |
| DE573603T1 | Germany | T1 | |
| EP1026240A2 | European Patent Office (EPO) | A2 | |
| IL120939A | Israel | A | |
| US2002150881A1 | United States of America | A1 | |
| EP1279731A1 | European Patent Office (EPO) | A1 | |
| JP2003159086A | Japan | A | |
| US2003113717A1 | United States of America | A1 | |
| EP0573603B1 | European Patent Office (EPO) | B1 | |
| EP1325931A1 | European Patent Office (EPO) | A1 | |
| AT243710T | Austria | T | |
| ATE243710T1 | Austria | T1 | |
| DE69233108D1 | Germany | D1 | |
| JP3447731B2 | Japan | B2 | |
| US2003175919A1 | United States of America | A1 | |
| DK0573603T3 | Denmark | T3 | |
| US2003219722A1 | United States of America | A1 | |
| US2003219886A1 | United States of America | A1 | |
| US2003223977A1 | United States of America | A1 | |
| JP2004000221A | Japan | A | |
| US2004005539A1 | United States of America | A1 | |
| US2004023205A1 | United States of America | A1 | |
| EP0573611B1 | European Patent Office (EPO) | B1 | |
| EP1026240A3 | European Patent Office (EPO) | A3 | |
| AT262036T | Austria | T | |
| ATE262036T1 | Austria | T1 | |
| ES2124203T3 | Spain | T3 | |
| DE69233325D1 | Germany | D1 | |
| DE69233108T2 | Germany | T2 | |
| DK0573611T3 | Denmark | T3 | |
| EP1452599A1 | European Patent Office (EPO) | A1 | |
| EP0573611B9 | European Patent Office (EPO) | B9 | |
| ES2219638T3 | Spain | T3 | |
| DE69233325T2 | Germany | T2 | |
| EP1541682A2 | European Patent Office (EPO) | A2 | |
| IL120940AThis record | Israel | A | |
| IL120941A | Israel | A | |
| EP1541682A3 | European Patent Office (EPO) | A3 | |
| CA2105304C | Canada | C | |
| EP0797666B1 | European Patent Office (EPO) | B1 | |
| AT311452T | Austria | T | |
| ATE311452T1 | Austria | T1 | |
| US6979538B2 | United States of America | B2 | |
| DE69534656D1 | Germany | D1 | |
| DK0797666T3 | Denmark | T3 | |
| US2006084113A1 | United States of America | A1 | |
| JP3771253B2 | Japan | B2 | |
| ES2255066T3 | Spain | T3 | |
| US2006134087A1 | United States of America | A1 | |
| US7078383B2 | United States of America | B2 | |
| DE69534656T2 | Germany | T2 | |
| JP3819931B2 | Japan | B2 | |
| US7118879B2 | United States of America | B2 | |
| EP1734121A2 | European Patent Office (EPO) | A2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent renewed for 20 yearsKB20 | KB20 | |
| Patent grantedGrantedFF | FF |
Numbers
- Publication, DOCDB
- 120940
- Publication, EPODOC
- IL120940
- Application
- 12094089
- Application, DOCDB
- 12094089
- Application, EPODOC
- IL19890120940
Titles
- English
- FUSION PROTEINS DISPLAYABLE ON THE SURFACE OF FILAMENTOUS PHAGE AND A RECOMBINANT, FILAMENTOUS PHAGE BEARING THE SAME
Classification
- IPC, 7
- C07K14 005
- C07K17 00
- C07K19 00
- C12N7 01
- C12N15 10
- C12N15 62
- C12P21 00