Method for producing virtual chromosomes
Abstract
This record has no abstract on file.
Term
Term ended
Projected expiry passed 16 September 2023, 3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
22 claims: 1 independent, 21 dependent
- 1Claims of equivalent WO 2004029747 A2 Claims 1. A method for producing a virtual chromosome representing a respective natural chromosome, characterised in that it comprises the steps — dividing sequence data of the natural chromosome into fractions, — determining the CG content in every fraction, — calculating for every fraction a value between a minimum and a maximum value according to the CG content and — producing the virtual chromosome by representing each fraction with the value.
94 paragraphs, as filed
Description of equivalent WO 2004029747 A2
Method for producing virtual chromosomes
The present invention relates to a method for producing a virtual chromosome representing a respective natural chromosome as well as a virtual chromosome or a part thereof represented by values according to its CG content, a set of virtual chromosomes or parts thereof and the use of a set of virtual chromosomes .
The human genome is ordered in a highly structured hierarchical fashion. The nucleus of a diploid cell contains approximately 2-3,17'10<sup>9</sup> nucleotide bases, which are arranged into 46 intricately packed DNA threads that become visible in the form of separate chromosomes during cell division. The specific features of these chromosomes, such as number, form, structure and banding patterns, provide the basis for their microscopic assessment through conventional cytogenetic means . This technique is still the most important screening tool for the identification of constitutional as well as acquired karyotype abnormalities .
At present, karyotype abnormalities are described according to the "International System for Human Cytogenetic Nomenclature (ISCN)", a system that is based on the diagrammatic representation of chromosomes and their banding patterns . The band sizes and their distribution in these so-called ideograms are deduced from the measurement of trypsin/Giemsa-stained chromosome images, and their relative staining intensity is symbolized by five different shades. Since the number of discernable bands also depends on the variable length of the respective chromosomes, ideograms with 400, 550 and 850 band resolution represent different stages of condensation.
Compared to the highly sophisticated computer algorithms and software tools that are available for the analysis and evaluation of molecular genetic data, those for depicting and processing cytogenetic data have hardly changed within the last 20 years . On the one hand, the ISCN nomenclature sufficed the limited spatial resolution of the morphological analyses, the resulting inherent subjective band assignment and interpretation of the ensuing abnormalities. On the other hand, this descriptive nature has so far also prevented a more precise and objective representation of cytogenetic and fluorescence in situ hybridization (FISH) data and, therefore, also their seamless integration into existing DNA databases. Over the last two decades numerous meta- and interphase FISH technologies, such as those utilizing heterogeneous types of sequence-specific probes, multicolor chromosome and region-specific painting probes, comparative genomic hybridization (CGH) and comparative expressed sequence hybridization (CESH) , became available and were further developed into purely DNA- and RNA-based microarray- techniques . The description of the ensuing rapidly accumulating molecular genetic results, according to the presently available cytogenetic and particular FISH nomenclature, is cumbersome and error- prone. The data are difficult to process and, at present, their integration into one joint molecular cytogenetic database is virtually impossible. One recent approach to overcome these obstacles to some degree, is the definition of cytogenetic landmarks by means of homogeneously spaced FISH anchor probes along chromosomes and their respective ideograms .
It has been accepted since a long time that the chromosomal banding pattern reflects alternating CG-rich and CG-poor sequence compartments. Nevertheless, the prevalent notion was that the correlation between chromosome bands and CG content could only be considered on the whole as a rather weak approximation. In this context, it seemed highly unlikely that the banding phenomenon could simply result from long-range variations in the linear base pair composition alone. Instead, it was presumed that the banding pattern is significantly co-determined and modified by structural factors, such as the folding, protein coverage, packaging and condensation of the DNA as well as by the accessibility of dyes to DNA.
Finally, a direct computational comparison between the sequence- specific CG content and the particular staining pattern that recently became possible, further corroborated this connection (Niimura and Gojobori ("In silico chromosome staining: Reconstruction of Giemsa bands from the whole human genome sequence", PNAS, vol. 99, no. 2 (797-802))). Their "in silico chromosome staining" was achieved by a method of two windows, one local window of 2.5 mb and a regional window of 9.3 mb, whereby the relationship between the CG content in the local window with respect to the GC content in the regional window was calculated. According to Niimura and Gojobori this two-window- method would better produce an in silico staining than explaining the Giemsa banding patterns only by the difference in base composition. Furthermore, it is assumed that the performance would improve by taking into account the difference of a compaction ratio between G and R bands whereby G bands are more condensed than R bands . By computing 10 kb fragments of the human DNA sequence and various ways of statistical analyses, they found that, at the 850 band level, the CG content and Giemsa bands along a chromosome correlated best at a local window of 2.5 Mb and a regional window of 9.3 Mb. The sizes of these windows were chosen to optimize the correspondence between in silico and Giemsa bands, but the authors also stated that their approach might not be appropriate for fine bands that are smaller than the local window size. However, they anticipate that the correlation between Giemsa and in silico bands might further be improved by integrating the genome-wide FISH mapping data. In this publication, however, no chromosomes were constructed but ideograms were calculated and compared. According to Niimura and Gojobori their results indicate that Giemsa banding patterns cannot be explained only by the difference in base composition. Thus, the relationship between the nucleotide sequence and cytogenic bands would still remain illusive.
Previous attempts to link the cytogenetic map with the sequence of the human genome focused on a top-down approach, by either demarcating band borders or setting band-independent cytogenetic landmarks with specific FISH probes. For example, the BAC Resource Consortium placed 7,600 of such cytogenetically defined landmarks on the draft sequence of the human genome. Although these markers should, amongst other things, "also allow a stringent assessment of sequence differences between the dark and light bands of chromosomes", their location is nevertheless still depicted on ideograms that are not linked to the DNA sequence. Similarly, the "Cancer Chromosome Aberration Project (CCAP) " devised an ideogram-based cytogenetic coordinate system with arbitrarily defined intervals to indicate the location of sequence-anchored FISH clones. In Jingwei et al . (PNAS 94 (1997), pp. 6862-6867) the GC content of chromosome parts is generally examined.
In US 6,136,540 A, a computer program for locating genetic abnormalities is described with which the subjective analysis of selectively stained chromosomes can be avoided. According to this document, the chromosomal abnormalities are determined by specific hybridizing probes (which are labelled with fluoro- phores) and analyzed, wherein chromosomal additions, deletions, amplifications, translocations and re-arrangements can be recognized. Although according to this document the disadvantages of experimental chromosomal stainings, or the problems thereof regarding the reproducibility, respectively, are said to be avoided, also in this instance the experimental work is complex, i.e. with fluorophore-labelled hybridizing probes, whereby again difficulties in terms of reproduction and the inaccuracies inherent in experimental proof likewise exist.
According to Daigo et al . (DNA Res. 6(4) (1999), pp. 227-233), the GC content on certain portions on chromosomes 9 and 3 is determined, this being done according to conventional methods .
The article by Hraber et al . (Genome Biology 2(9) (2001), research 0037.1-0037.14) relates to a possible analysis for examining inter-specific interactions in sequences which are expressed during the interaction between two symbionts , e.g. as regards their GC content. In this instance, however, neither complete larger areas of the genome are compared with each other, nor concretely illustrated as chromosomes.
Therefore, an aim of the present application is to provide a method for producing a virtual chromosome, which chromosome provides an accurate banding pattern in a scale independent and highly region specific fashion with high resolution. Furthermore, these produced virtual chromosomes should comprise not only morphological information as the conventional ideograms or ISCN strings, but also the corresponding genetical information, e.g. the sequence data. Such a virtual chromosome comprising the complete sequence data, being comparable to the conventional representations of chromosomes and having a high resolution has not been produced so far. Therefore, a further aim of the present application is to provide a set of chromosomes which can be used as an interface between conventional representations of chromosomes and genetical information and sequence data, in order to directly compare DNA sequence derived data and natural chromosomal binding patterns .
The aim of the present application is solved by the method according to the present invention as defined above, which is characterised in that it comprises the steps
— dividing sequence data of the natural chromosome into fractions,
— determining the CG content in every fraction,
— calculating for every fraction a value between a minimum and a maximum value according to the CG content and
— producing the virtual chromosome by representing each fraction with the value.
Chromosomes produced by this method were found to provide an excellent correlation between their banding pattern and that of their corresponding natural counterparts . This astonishing concordance not only indicates that the chromosomal banding pattern is to a large extent directly determined by the underlying DNA sequence, but can also provide a unique basis for the joint procession of morphological and molecular genetic data within a single DNA sequence based framework. To the contrary of recent publications, which explicitly state that the Giemsa banding patterns cannot be explained only by the different base composition, it is shown with the present method that a sequence based display of chromosomes according to CG content is possible and results in high resolution virtual chromosomes.
In the scope of the present application the term "a method for producing a virtual chromosome representing a respective natural chromosome" refers not only to complete chromosomes but also to parts thereof, for example separate chromosome arms or ends.
The sequence data of the natural chromosome can be retrieved for example from any electronically available data, in the case of human chromosomes this may be for example sequences from the Hu- man Genome Project Working Draft (http://genome.ucsc.edu/). In the scope of the present application the term "chromosome" refers to any chromosomes of any organism. The organism is for example the human being, however also any animal, in particular mammal, may be the organism from which the chromosome is derived. Virtual chromosomes from mammals and human beings are particularly preferred since they can be used for evolutionary studies .
The step of dividing the sequence data into fractions and determining the CG content in every fraction will preferably be carried out electronically.
The term "value between a minimum and a maximum value" refers to any parameter which is suitable to define and preferably virtually represent a specific CG content. This may be for example a percentage between 0% and 100% or a value between 0 and 1 which can for example be visualised in a two- or three-dimensional picture. The values can furthermore be represented by light values or colour values . Any value which can be visualised is useful for representing a given CG content.
The step of calculating the value for every fraction according to its CG content may be carried out with any suitable table, algorithm, formula or programme, whereby for example a minimum value is assigned to a minimum amount of CG in a fraction and a maximum value is assigned to a maximum amount of CG in a fraction. The values in between are then assigned as a linear function between the two extremes to the varying CG content in the fractions .
Preferably the value is a light value. The advantage of a light value is that the visualisation can be interpreted very rapidly and simply and can furthermore be compared to conventionally produced, for example microscopically taken or ISCN chromosomes.
Still preferred the maximum value is represented by white, the minimum value is represented by black and values in between are represented by grey shades. This representation of a virtual chromosome is directly comparable to the conventionally scanned chromosomes, however the resolution is very high and the virtual chromosome further comprises the sequence data information which is missing in the chromosome representations according to the state of the art.
According to a preferred embodiment the natural chromosome is divided into fractions of a minimum length of 1000, preferably 5000, more preferred 10,000, especially 20,000, and/or into 10 fractions of a length of 10,000 to 1,000,000 bp, preferably a length of 50,000 to 500,000 bp, still preferred a length of 100,000 to 300,000 bp (or combination of these values). These fractions are sufficiently small, in order to provide high resolutions and a maximum amount of information. Preferably, a fraction corresponds to the estimated average size of a DNA loop as well as to that of an isochore. An optimal length of a fraction is for example 200,000 base pairs.
Advantageously, the fraction with a CG content of 30 to 35%, preferably 33%, is assigned to a minimum value and the fraction with a CG content of 60 to 65%, preferably 62%, is assigned to a maximum value. It was found that a variation between these percentages produces chromosomes with grey values of bands corresponding to the conventionally represented chromosomes . Therefore this visualisation of the virtual chromosome can be directly compared to conventional chromosome representations, as for example perceived through a microscope.
Preferably to fractions with unknown sequence the value is assigned according to their morphological appearance. Even though the amount of sequence missing fractions has become very low in particular with human chromosomes due to the nearly complete human genome and will generally decrease rapidly, the few missing sequence fractions can be supplemented with data derived from the morphological appearance as for example taken from an ideogram in order to provide a complete chromosome.
Still preferred after the production of the virtual chromosome a filter for smoothing the appearance is applied, preferably a Gaussian convolution filter. With this also the resulting shades were used to gradually fill the last few pixels at the chromo- some boundaries.
According to a still preferred method, for the production of the virtual chromosome a scale correction filter is applied. This may be a normalisation and non-linear gamma-like grey scale correction filter. With this contrast enhancement is achieved and the image of chromosomes as they are perceived through a microscope is optimally mimicked.
A further aspect of the present application relates to a virtual chromosome or a part thereof, represented by values according to its CG content which is characterised in that it is produced according to the inventive method as defined above. According to the present invention a virtual chromosome is provided which does not only comprise morphological information and can be compared to for example conventional ideograms but which also comprises sequence data. This sequence based visualised chromosome according to the present invention shows an excellent correlation of the banding pattern of virtual chromosomes and that of the corresponding natural counter-parts . With respect to this aspect of the present invention the same definitions and preferred embodiments as mentioned above apply.
Preferably the value is a light value, still preferred the maximum value is represented by white, the minimum value is represented by black and values in between are represented by grey shades. As mentioned above this allows a representation of the chromosome as seen through a microscope and therefore is optimal for comparing with conventional chromosome representations, as for example ideograms or microscopically perceived chromosomes .
According to a further aspect of the present application a set of virtual chromosomes or parts thereof is provided which is characterised that it comprises two or more chromosomes or parts thereof according to the present invention as defined above. Preferably the set comprises a maximum number of chromosomes whereby this set can continuously be supplemented with additional newly found or newly identified chromosomes .
Preferably the set comprises chromosomes or parts thereof spe- cific for one or more organism(s) . The advantage of a set which is specific for one organism lies in the fact that this set is useful for comparing any newly identified modifications or rearrangements of chromosomes of that organism. However, it is of course possible to provide a set comprising chromosomes of various, preferably defined, organisms.
Still preferred the set comprises 24 human chromosomes or parts thereof. This set will be a standard for the normal human chromosomes and can be used for comparing chromosomes of a patient with normal chromosomes in order to detect any modifications or rearrangements .
Still preferred the set further comprises additional modified chromosomes or parts thereof, preferably chromosomes with trans- locations. This is especially advantageous for modified chromosomes which are related to a specific illness, for example a specific tumor. By providing a classification of such modified virtual chromosomes, preferably every modification comprising a reference to a specific illness, it is possible to easily refer to a set of chromosomes isolated from a patient an illness or the risk of developing an illness by comparing the set of virtual chromosomes with the patient's chromosomes. Due to constant detection of new modifications in chromosomes the set can be rapidly and continuously completed with the newest medical information.
Within the scope of the present application the term "chromosome modification" refers to any sequence modification, e.g. any mutation or translocation of a chromosome fragment.
A further aspect of the present invention is the use of the set of virtual chromosomes according to the present invention as mentioned above for cataloguing chromosome modifications. As mentioned above the inventive set is particularly useful for providing electronic information on chromosome modifications and their connection with any illness or risk of illness . In the scope of the present application the term "chromosomal modifications" relates to any modification in the chromosome. This may be on the level of a sequence mutation or on the level of com- plete translocation of a chromosome fragment. Due to the high resolution any chromosome modification can be detected and catalogued. The inventive set allows the description of chromosome abnormalities with a hitherto unknown molecular precision, whereas on the other hand it remains possible to interpret more fuzzy large scale events on the plain chromosomal level such as also obtained with conventional cytogenetic analysis and with chromosome painting multicolour FISH comparative genome hybridisation and comparative expressed sequence hybridisation.
A further aspect of the present application relates to the use of an inventive set of virtual chromosomes as defined above for virtually mapping the chromosomal position of a sequence. Since the inventive set of virtual chromosomes derives from and represents the complete human DNA sequence, it is possible to map and display the chromosomal position of any given, known or unknown sequence or set of sequences that are contained in a data base used to produce the virtual chromosomes.
A further aspect of the present application relates to the use of an inventive set of virtual chromosomes as defined above as an interface between morphological and molecular genetic data. Preferably the morphological data is derived from information based on the International System for Human Cytogenetic Nomenclature (ISCN) .
This can be used to superimpose a graphic interface on any molecular genetic database. Thus, a virtual chromosome-based system will ensure that previously collected data remain accessible and analyzable. For example, a combination of ISCN nomenclature- and virtual chromosome-based graphic interface tools that may be superimposed onto existing cytogenetic databases as mentioned above is extremely useful. Such an interface enables the transformation of ISCN information into the corresponding karyotype image. Conversely, a karyotype picture that is generated with such a virtual chromosome tool can be translated into an ISCN string. Such a graphic interface is extremely valuable for visually crosschecking the ISCN description, by comparing the karyotype picture with the virtual chromosome image. As a valuable by-product, such an approach also significantly improves the quality of cytogenetic data. Moreover, it also eases a seamless intra- and inter-laboratory exchange and communication of cytogenetic data in a standardized form, not only with a remote central facility, but also with FISH and molecular genetic databases. For example, cytogenetic data that are prepared for publication can then be easily reviewed and conveniently transmitted to a central database.
The superimposition of virtual chromosomes as a graphic interface on molecular genetic databases aids in the visualization of any type of FISH-, DNA- and RNA-derived data sets as well as gene expression profiles in a standardized "chromosomal" fashion. The advantages of such a chromosomal representation is that it is independent of the probe distribution on the diverse arrays, and also that its "natural" appearance makes it easier to comprehend and compare by visual inspection. Moreover, the representation of gene expression profile in such a style is of growing interest, since there is accumulating evidence that even functionally unrelated genes are expressed in transcriptional territories in Drosophila, as well as in the human genome. The resulting distribution pattern will resemble those that derive from CGH and CESH analyses, in which differentially labeled DNA or cDNA from a tissue of interest and a control sample are simultaneously hybridized directly onto chromosomes. Consequently, such data sets can be directly correlated and cross-analyzed with other karyotype patterns, for example with the associated karyotype abnormalities.
Because the position of FISH or any other DNA probes (from oli- gonucleotides up to region-specific painting probes) can be conveniently displayed on virtual chromosomes, and cross-checked with the actually obtained hybridization patterns, such graphic interfaces are also of potential interest for resource centers. Moreover, the position of cytogenetic landmarks in the form of evenly distributed FISH probes can be directly integrated into such virtual chromosomes for breakpoint-mapping purposes. Finally, even submicroscopic events that otherwise are not detectable through conventional cytogenetic means, such as microdele- tions, and interphase FISH data, are able to can be mapped and included in such a universal platform.
Preferably the set of chromosomes serves as reference for classifying a phenotype to a sequence arrangement. The term "sequence arrangement" refers to any sequence modification, e.g. sequence mutations or translocations of chromosome fragments. The phenotype may refer to normal or abnormal phenotypes, e.g. various illnesses as tumors. Especially if the chromosomes are classified according to modifications and resulting phenotypes, any chromosome isolated from a patient and analysed with conventional microscopic methods can be compared to the inventive set of chromosomes. Similarities between the modifications of the chromosomes would imply also similar phenotypes or at least the likelihood or risk of developing a similar phenotype.
Advantageously the set of chromosomes serves as a tool for carrying out structural and functional analyses, respectively, of a sequence arrangement. As mentioned above the analysis can be carried out by gene mapping or virtual hybridisation due to the sequence data comprised in the virtual chromosome.
Still preferred the set of chromosomes serves as a tool for determining the influence of a given factor on a sequence arrangement. For example an external factor, as a chemical substance, energy with various wavelengths or even the influence of microorganisms can be analysed on a cytogenetical basis and transferred into or compared to the inventive set of chromosomes, whereby the implications or resulting phenotypes can be retrieved or foreseen.
The chromosomes or set of chromosomes may also be stored on a computer program product (a CD, DUD, diskette or on a web server) as well as the method according to the present invention (as computer readable program means for causing a computer to control execution of the method according to the present invention) .
The present invention is described in more detail with the help of the following examples and figures to which it is, however, not limited, whereby Fig.l shows images of DNA sequence derived human chromosomes compared to Trypsin/Giemsa banded chromosome images;
Fig.2 represents a modelling of uneven condensation of G bands depending on their CG content;
Fig.3 shows virtual chromosomes compared to the cytogenetic and molecular genetic maps;
Fig.4 shows a virtual in situ hybridisation;
Fig.5 shows the construction of virtual chromosome abnormalities; and
Fig.6 represents a graphic interface between cytogenic and molecular genetic data sets .
E x amp l e s
E x a m l e 1 :
Production of virtual chromosomes on the basis of genetic data
For the construction of virtual chromosomes we downloaded the sequence data of the August and December 2001 as well as April and June 2002 releases of the human genome working draft. For the analysis and processing of the data and the resulting images, we used Perl, Mathematica (Wolfram Scientific) and Photoshop (Adobe) . Using the Scripting Language Perl, we first divided the sequence data of each individual chromosome into 200,000 base long fractions. This size corresponds to the estimated average size of a DNA loop and also approximately to that of an isochore. We then determined the percentage of the CG content of all strips whose sequence was at least 70 percent complete. This was the case in virtually all instances commencing with the December 2001 release. Gaps arising from unsequenced nucleotides (N's) were not taken into consideration. The CG content of the individual strips ranged from 33% to 62%, with a mean value of 41%. The tables with these data were stored in temporary files together with the information about the segment and band coordinates of the particular chromosome, and were fur<sub>¬</sub> ther computed and assembled with Mathematica as shown below:
StaticNormalize[L_List] := Block [{Lx = 0, 625 - Li = 0, 33} f<sup>#1 " i</sup><sub>/&</sub> ^ .
~ . L - Li
StatιcNormalize[L_Real] := Block [{Lx = 0. 625<sub>.</sub> Li = 0<sub>.</sub> 33}, ' ]
LJC - Li
Depending on its individual CG content, we assigned to each strip a grey value in a linear normalized fashion, i.e. strips with a CG content of 33% became black and those with a <sub>CG</sub> con<sub>¬</sub> tent of 62% white. The transfer function in Mathematica from percentage of CG to grey value is the sum of the (statistical<sub>)</sub> normalization and contrast enhancement:
<sup>B</sup>an<sup>d</sup>s<sup>A</sup>v<sup>g</sup> = <sup>L</sup>istConvolve[FoldMask, StaticNormalizefChrShades],
{CenterElement, -CenterEle ent} . .41] / (Plus @@ <sub>F</sub>old<sub>M</sub>ask<sub>)</sub>
<sup>T</sup>he derived bars were then integrated into the particular chro<sub>¬</sub> mosome boundaries, as defined by their length and centromere po<sub>¬</sub> sition. %he centromeric, heterochromatic and satellite regions, for which no appropriate sequence information is as yet available, were artificially supplemented according to their morphological appearance. To smoothen the appearance of the virtual chromosomes, we applied a Gaussian convolution filter (N<sub>(</sub>0.1<sub>)</sub>, 22 stripes in length) :
<img file="WO2004029747A2_D0001.tif" /> We furthermore used the resulting shades to gradually fill the last few pixels at the chromosome boundaries. For contrast enhancement and to mimic the images of the chromosomes as they are perceived through a microscope, we applied a normalization and non-linear, gamma like grey-scale correction filter<sub>:</sub>
ColorCorrection [c_ : = (* l+ (c-l)<sup>Λ</sup> 3*)
Interpolatio [ { {0, 0} , { .1, .1} , { .2, .5} , { .5, .8} , {1, 1<sub>}</sub> ,
InterpolationOrder -> 1] [c]
In order to bring the data into the form of chromosomes <sub>:</sub>
<img file="WO2004029747A2_D0002.tif" />
<sup>A</sup>s a final step, the respective images were then imported, assembled and arranged accordingly in Photoshop.
E x a mp l e 2:
Images of DNA sequence-derived human chromosomes
<sup>I</sup>n Fig. 1 Trypsin/Giemsa-banded chromosome images are represented, whereby for each chromosome (a) shows an 850 band stage I<sup>SCN</sup> reference picture, (b) shows their derived, straightened grey scaling pattern, (c) is the comparison with their computed virtual counterparts of the August 2001 and (d) the December 2001, <sup>(</sup>e<sup>)</sup> the April 2002 and (f) the June 2002 release (see ht- tp : //genome . ucsc . edu/ ) .
<sup>D</sup>espite the excellent general overall concordance between the matched sets of chromosome homologues, some local variations and differences become evident, particularly between virtual chromosomes that derive from different sequence releases. The originally excellent concordance between the grey scale-banding pattern of the natural chromosomes and the virtual chromosomes of the August 2001 release deteriorated when we used the <sub>D</sub>ecem- ber 2001 release for comparison. This intriguing observation may be explained by the fact that the later assembly was produced at the NCBI rather than at the UCSC. When compared to the UCSC assembly the NCBI assembly shows a slightly better local order and orientation, but somewhat worse tracking of the chromosome level maps. Thus, such misplacements of sequence portions may change the local banding pattern in a noticeable way, as becomes particularly obvious when comparing the long arms of the virtual chromosomes 1 and 11 from various releases . The location of the centromeres of chromosomes 5 (December 2001 release) , 7 and 12 (both April 2002 release) moved to odd positions. Of note, however, is that the advances and continuous corrections in the sequence assembly significantly improved, as well as the concordance between the natural and virtual banding patterns . Such a comparison of virtual chromosomes that derive from different sequence releases can thus also provide an independent validation of the sequence map .
E x a mp l e 3:
Modeling uneven condensation of G-bands depending on their CG content
Virtual chromosomes allow structural and functional analyses of the genome and provide the means to study the influence of different factors on the large-scale chromosomal banding pattern. The example shown here relates to the potential effects of the unequal contraction of light and dark bands during chromosome condensation. The analysis is based on the notion that Giemsa dark bands may contain up to 11 times more DNA than the light bands and that the DNA compaction ratio is in the order of the cubic root of the respective DNA length (s. Fig. 2).
We first transformed the images of the shortest (ISCNS, 500 band stage) and longest (ISCNL, 850 band stage) ISCN chromosome 7 into a gray scale pattern, by measuring the gray values along the blue path (a) . After these two chromosome images were brought to the same length, their banding pattern was compared with that of virtual homologues, which were modified in different ways. Depending on the respective CG content and as explained in the graphic (s. Fig. 2 b) , the length of the light and dark bands of the virtual chromosomes were stretched or condensed in a linear fashion by applying factors of 0.3, 0.5, and 0.8, respectively, which correspond approximately to a 2.2-, 3.4-, and 5.8-fold difference in their DNA length (s. Fig. 2 c) . The comparison of the resulting banding patterns corroborates previous experimental evidence that chromosome condensation most likely does not only take place in a CG content-independent linear fashion. However, it cannot provide a good explanation for the intriguing pattern obtained by stretching GTG-banded chromosomes . We envision that by determining the distances between light and dark bands of chromosomes at different stages of contraction, it will eventually be possible to deduce a factor or formula, whose plausibility can be subsequently checked by comparing the images of natural chromosomes with the corresponding virtual chromosomes . Although the uneven extension of condensed and decon- densed dark and light chromosome bands may even be visually noticeable, it is of no practical concern for analytical purposes, since the preparation-dependent variation within the chromosome class itself is broader than that which may result from the distorted artificial stretching by computer.
E x a mp l e 4:
Virtual chromosomes link the cytogenetic and the molecular genetic maps
The ISCN chromosome 7 (850 band-stage) is shown together with its virtual counterpart and three different ideograms in Fig. 3a. The banding pattern of the left ideogram is based on the location of the turning points between CG-richer and CG-poorer regions in the sequence-based virtual chromosome. The curve follows the mean CG content. The UCSC ideogram (August 2001 release) is placed in the middle and the according ISCN (850 band- stage) one on the right. The left half of the virtual chromosome displays the raw, unenhanced grey values of the respective CG content, whereas in the right half the contrast has been enhanced according to the curve shown in the graph (b) . The thin horizontal lines provide an absolute 10 Mb scale. However, as explained above for Fig. 2, the DNA might not be distributed along the bands in such a linear fashion as this scale implies. It also becomes evident that the width and distribution of bands may vary considerably on ideogrammatic representations, although their number and designation usually concords. It is therefore impossible to accurately position any absolute or relative chromosomal occurrence on any type of ideogram. Virtual chromosomes solve this problem by linking the absolute precision of DNA sequence position with the arbitrary location indicators of any type of ideogram, which is indicated by the lines combining the band borders of the three examples shown in this Figure.
E x a mp l e 5:
Virtual in situ hybridization
Because virtual chromosomes symbolize the DNA sequence in a very condensed form, it is now possible to display the precise position of any type of DNA sequence or set of sequences, independent of their numbers and DNA sequence length, by virtue of their particular nucleotide coordinates with a hitherto unknown cytogenetic precision. A comparative example of conventional and virtual FISH mapping is shown on the left side of the Fig. 4 for the MLL partner gene GRAF at 5(q31), whose original, cytogenet- ically determined location was diagrammatically delimited with a CGH software (Vysis, Doners , Grove, USA). The metaphase image is shown on the top, the CGH mapping image in the middle and the virtual chromosome 5 with the enlarged natural one on the bottom.
Previous attempts to integrate the cytogenetic map location with existing genomic databases had to rely on such a display of chromosomal events on ideogrammatic coordinate systems, since it was not possible to directly link these two datasets . We show here the distribution of 77 of 82 chromosome 7 CCAP BAC clones as an example for the achievable improvement in assigning the absolute and relative position of a whole set of clones .
E x amp l e 6:
Construction of virtual chromosome abnormalities The definition of a particular chromosomal occurrence with a molecular precision now also facilitates the accurate reconstruction of any chromosome rearrangement, with a known molecular breakpoint location. As exemplified here with the translocation t(4;ll) (q21;q23) (s. Fig. 5), this is an important prerequisite for the potential use of such virtual chromosome rearrangements in pattern recognition systems and automated karyotyping. With the current cytogenetic terminology and precision, the location of the breakpoints of a particular translocation can only arbitrarily be defined by the location of its affected bands . Depending on the sequence release, band 4(q21) encompasses between 11.5 and 14.2 Mb (6.0% - 7.4% of chromosome 4) and band ll(q23) between 10.9 and 11.7 Mb (7.9% - 8.5% of chromosome 11). Without knowing the exact within-band position of the two genes, AF4 and MLL, that are disrupted and fused as a result of this translocation, the breakpoints might be located anywhere within these bands. To demonstrate this point, we have assigned them to the outer boundaries of the bands in question. Most likely, this is also already one of the highest resolutions that can be achieved with an average morphological chromosome analysis. However, compared to the length and banding patterns of derivative chromosomes that originate from precisely positioned molecular break points (indicated by *), those resulting from erratic breakpoint allocations may look rather different. They would therefore certainly be unusable for comparative and redetection purposes in pattern recognition.
E x amp l e 7:
A graphic interface between cytogenetic and molecular genetic data sets
As the top-level entities of the human sequence, virtual chromosomes cover the nine orders of magnitude of the whole genome in a highly condensed, easily expandable and the most natural imaginable "morphological" fashion. Therefore, they can be utilized as a unique front-end tool for the visualization of the information contained in any sequence database; in effect basically from a single base pair up to whole chromosomes in a cytogenetic manner. As an example, we show here in a one Mb scale the distribution of the approximately 15,000 genes and CpGs from the UCSC database along virtual chromosomes and the according UCSC color ideograms. For practical reasons, the scale of the bars for the genes is only shown in half the height of the CpG bars, and the height of these CpG bars on chromosome 19 are cut off.
The vertical bars on the left of the chromosome indicate the size and position of heterochromatic and satellite regions that were artificially supplemented, since their sequence is not yet available. The fine horizontal bars on the left of the chromosomes indicate sequence gaps.
9 priority claims, no other members on record
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 14302002 | Austria | A | |
| 14302002 | Austria | A | |
| 14302002 | Austria | – | |
| 0310254 | European Patent Office (EPO) | W | |
| 0310254 | European Patent Office (EPO) | W | |
| 14302002 | – | – | – |
| AT20020001430 | – | – | – |
| EP2003010254 | – | – | – |
| WO2003EP10254 | – | – | – |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Application deemed to be withdrawnWithdrawn18D | 18D | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWNSTAA | STAA | |
| Request for examination filed17P | 17P | |
| Designated contracting statesAK | AK | |
| Request for extension of the european patentAX | AX | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI |
Numbers
- Publication
- 1563444
- Publication, DOCDB
- 1563444
- Publication, EPODOC
- EP1563444
- Application
- 3798163
- Application, DOCDB
- 03798163
- Application, EPODOC
- EP20030798163
Titles3
- German
- VERFAHREN ZUR ERZEUGUNG VON VIRTUELLEN CHROMOSOMEN
- English
- METHOD FOR PRODUCING VIRTUAL CHROMOSOMES
- French
- PROCEDE PERMETTANT DE PRODUIRE DES CHROMOSOMES VIRTUELS
Classification
- CPC, 4
- G06F19/26
- G16B30/00
- G16B45/00
- G06F19/22
- IPC, 5
- C12N15 00
- C12Q1 68
- G06F
- G16B30 00
- G16B45 00
Designated states31
- Contracting states, 27
- Austria
- Belgium
- Bulgaria
- Switzerland
- Cyprus
- Czechia
- Germany
- Denmark
- Estonia
- Spain
- Finland
- France
- United Kingdom
- Greece
- Hungary
- Ireland
- Italy
- Liechtenstein
- Luxembourg
- Monaco
- Netherlands (Kingdom of the)
- Portugal
- Romania
- Sweden
and 3 moreShow fewer
- Slovenia
- Slovakia
- Türkiye
- Extension states, 4
- Albania
- Lithuania
- Latvia
- North Macedonia