Method for producing virtual chromosomes
Abstract
A method for producing a virtual chromosome representing a respective natural chromosome is provided, which comprises the steps - dividing sequence data of the natural chromosome into fractions, - determining the CG content in every fraction, - calculating for every fraction a value between a minimum and a maximum value according to the CG content and - producing the virtual chromosome by representing each fraction with the value.

Term
No projected expiry on record.
- Priority and filed
- Granted
- Today
22 claims: 22 independent, 0 dependent
- 1PATENT CLAIMS:PATENTANSPRÜCHE: 1. Process for the production of a virtual chromosome which represents a corresponding natural chromosome, characterized in that it comprises the following steps: 1. Verfahren zur Herstellung eines virtuellen Chromosoms, das ein entsprechendes natürliches Chromosom repräsentiert, dadurch gekennzeichnet, dass es die folgenden Schritte umfasst: - Subdivide sequence data of the natural chromosome into fractions with a length of at least 10,000 bp, - Unterteilen von Sequenzdaten des natürlichen Chromosoms in Fraktionen mit einer Länge von mindestens 10.000 bp, - Bestimmen des CG-Gehalts in jeder Fraktion, - determining the CG content in each fraction, - Berechnen eines Werts zwischen einem Minimalwert und einem Maximalwert für jede Fraktion gemäß dem CG-Gehalt, und - calculating a value between a minimum value and a maximum value for each fraction according to the CG content, and - Creation of the virtual chromosome by representing each fraction with the value. - Herstellen des virtuellen Chromosoms durch Darstellen jeder Fraktion mit dem Wert.
- 2Method according to Claim 1, characterized in that the value is a light value. 2. Verfahren nach Anspruch 1, dadurch gekennzeichnet, dass der Wert ein Lichtwert ist.
- 3Method according to Claim 2, characterized in that the maximum value is shown in white, the minimum value in black and values in between are shown in shades of gray. 3. Verfahren nach Anspruch 2, dadurch gekennzeichnet, dass der Maximalwert weiß, der Minimalwert schwarz und Werte dazwischen in Grauschattierungen dargestellt werden.
- 4Verfahren nach einem der Ansprüche 1 bis 3, dadurch gekennzeichnet, dass das natürliche Chromosom in Fraktionen einer Länge von 10.000 bis 1.000.000 bp, vorzugsweise einer Länge von 50.000 bis 500.000 bp, noch bevorzugter einer Länge von 100.000 bis 300.000 bp unterteilt wird. 4th Method according to one of Claims 1 to 3, characterized in that the natural chromosome is divided into fractions with a length of 10,000 to 1,000,000 bp, preferably a length of 50,000 to 500,000 bp, more preferably a length of 100,000 to 300,000 bp.
- 5Method according to one of claims 1 to 4, characterized in that the fraction with a CG content of 30 to 35%, preferably 33%, a minimum value and the fraction with a CG content of 60 to 65%, preferably 62%, is assigned a maximum value. 5. Verfahren nach einem der Ansprüche 1 bis 4, dadurch gekennzeichnet, dass die Fraktion mit einem CG-Gehalt von 30 bis 35 %, vorzugsweise 33 %, einem Minimalwert und die Fraktion mit einem CG-Gehalt von 60 bis 65 %, vorzugsweise 62 %, einem Maximalwert zugeordnet wird.
- 6Verfahren nach einem der Ansprüche 1 bis 5, dadurch gekennzeichnet, dass Fraktionen mit unbekannter Sequenz der Wert gemäß ihrer morphologischen Erscheinung zugeordnet wird. 6th Method according to one of Claims 1 to 5, characterized in that fractions with an unknown sequence are assigned the value according to their morphological appearance.
- 7Verfahren nach einem der Ansprüche 1 bis 6, dadurch gekennzeichnet, dass nach der Herstellung des virtuellen Chromosoms ein Filter zur Glättung der Erscheinung, vorzugsweise ein Gaußsches Faltungsfilter, angewendet wird. 7th Method according to one of Claims 1 to 6, characterized in that after the production of the virtual chromosome, a filter for smoothing the appearance, preferably a Gaussian convolution filter, is used.
- 9Virtual chromosome or part thereof, which is represented by values according to its CG content, characterized in that the production takes place according to one of Claims 1 to 6. 9. Virtuelles Chromosom oder Teil davon, das bzw. der durch Werte gemäß seinem CGGehalt dargestellt wird, dadurch gekennzeichnet, dass die Herstellung nach einem der Ansprüche 1 bis 6 erfolgt.
- 10Virtual chromosome or part thereof according to Claim 9, characterized in that the value is a light value. 10. Virtuelles Chromosom oder Teil davon nach Anspruch 9, dadurch gekennzeichnet, dass der Wert ein Lichtwert ist.
- 11Virtual chromosome or part thereof according to Claim 10, characterized in that the maximum value is shown in white, the minimum value is shown in black and values in between are shown in shades of gray. 11. Virtuelles Chromosom oder Teil davon nach Anspruch 10, dadurch gekennzeichnet, dass der Maximalwert weiß, der Minimalwert schwarz und Werte dazwischen in Grauschattierungen dargestellt sind.
- 12Satz von virtuellen Chromosomen oder Teilen davon, dadurch gekennzeichnet, dass er zwei oder mehr Chromosome oder Teile davon nach einem der Ansprüche 9 bis 11 umfasst. 12th Set of virtual chromosomes or parts thereof, characterized in that it comprises two or more chromosomes or parts thereof according to one of Claims 9 to 11.
- 13Satz nach Anspruch 12, dadurch gekennzeichnet, dass er Chromosome oder Teile davon umfasst, die für einen oder mehrere Organismen spezifisch sind. 13th Set according to claim 12, characterized in that it comprises chromosomes or parts thereof which are specific for one or more organisms.
- 14Satz nach Anspruch 13, dadurch gekennzeichnet, dass er 24 menschliche Chromosome oder Teile davon umfasst. 14th Set according to claim 13, characterized in that it comprises 24 human chromosomes or parts thereof.
- 15Satz nach Anspruch 14, dadurch gekennzeichnet, dass er weiters zusätzliche modifizierte Chromosome oder Teile davon, vorzugsweise Chromosome mit Translokationen, umfasst. 15th The set according to claim 14, characterized in that it further comprises additional modified chromosomes or parts thereof, preferably chromosomes with translocations.
- 17Verwendung eines Satzes von virtuellen Chromosomen nach einem der Ansprüche 12 bis 17th Use of a set of virtual chromosomes according to any one of claims 12 to 13 AT 412 476 Β AT 412 476 Β 15 zur virtuellen Kartierung der chromosomalen Position einer Sequenz. 15th for virtual mapping of the chromosomal position of a sequence.
- 18Verwendung eines Satzes von virtuellen Chromosomen nach einem der Ansprüche 12 bis 15 als Schnittstelle zwischen morphologischen und molekulargenetischen Daten. 18th Use of a set of virtual chromosomes according to one of Claims 12 to 15 as an interface between morphological and molecular genetic data.
- 19Verwendung nach Anspruch 18, dadurch gekennzeichnet, dass die morphologischen 19th Use according to claim 18, characterized in that the morphological 5 Data is obtained from information based on the International System of Human Cytogenetic Nomenclature (ISCN). 5 Daten von Informationen stammen, die auf dem Internationalen System für Humane Zytogenetische Nomenklatur (ISCN) basieren.
- 20Verwendung nach einem der Ansprüche 16 bis 19, dadurch gekennzeichnet, dass der Chromosomensatz als Referenz zur Klassifizierung eines Phänotyps zu einer Sequenzanordnung dient. 20th Use according to one of Claims 16 to 19, characterized in that the set of chromosomes serves as a reference for classifying a phenotype into a sequence arrangement. 10 10
- 21Use according to one of claims 16 to 20, characterized in that the 21. Verwendung nach einem der Ansprüche 16 bis 20, dadurch gekennzeichnet, dass der Chromosome set serves as a tool for carrying out structural or functional analyzes of a sequence arrangement. Chromosomensatz als Werkzeug zur Durchführung von Struktur- bzw. Funktionsanalysen einer Sequenzanordnung dient.
- 22Verwendung nach einem der Ansprüche 16 bis 21, dadurch gekennzeichnet, dass der Chromosomensatz ais Werkzeug zur Bestimmung des Einflusses eines bestimmten Fak15 tors auf eine Sequenzanordnung dient. 22nd Use according to one of Claims 16 to 21, characterized in that the chromosome set serves as a tool for determining the influence of a certain factor on a sequence arrangement.
Independent claims22
111 paragraphs in 1 section, as filed
The present invention relates to a method for producing a virtual chromosome which represents a corresponding natural chromosome, as well as a virtual chromosome or a part thereof, which is represented by values corresponding to its CG content, a set of virtual chromosomes or parts thereof and the use of a set of 5 virtual chromosomes.
The human genome is organized in a strictly hierarchical structure. The nucleus of a diploid cell contains about 2-3,17-10<sup>9</sup> Nucleotide bases, which are arranged in 46 intimately entangled DNA threads, which become visible in the form of separate chromosomes when cells divide. The specific features of these chromosomes, such as number, shape, structure and banding pattern, create the basis for their microscopic determination by conventional cytogenetic means. This method is still the most important screening tool for the identification of constitutional as well as acquired karyotype anomalies.
Currently, karyotype abnormalities are described according to the "International System for Human Cytogenetic Nomenclature (ISCN)", a system based on the diagrammatic representation of 15 chromosomes and their banding patterns. The band sizes and their distribution in these so-called ideograms are derived from the measured value of trypsin / Giemsa-stained chromosis images, and their relative intensity of coloration is symbolically represented by five different shades. Since the number of perceptible bands also depends on the variable length of the respective chromosomes, Ideo20 grams represent different levels of condensation with a band resolution of 400, 550 and 850.
Compared to the highly sophisticated computer algorithms and software tools that are available for the analysis and evaluation of molecular genetic data, those for displaying and processing cytogenetic data have hardly changed in the last 20 years. On the one hand, the ISCN nomenclature was sufficient for the limited spatial resolution of the morphological analyzes, the resulting inherent subjective band assignment and the interpretation of the resulting anomalies. On the other hand, this descriptive nature has hitherto prevented a more precise and objective representation of cytogenetic data and data from fluorescent in situ hybridization (FISH) and, as a result, their seamless integration into existing DNA databases. Over the past two decades, numerous meta- and interphase FISH technologies such as those using heterogeneous types of sequence-specific probes, chromosomal multi-color and region-specific staining probes, comparative genomic hybridization (CGH) and comparative expressed sequence hybridization (CESH) have been developed. available and have evolved into microarray techniques based purely on DNA and RNA. The description of the resulting 35 rapidly accumulating molecular genetic results is tedious and flawed in accordance with the currently available cytogenetic and in particular FISH nomenclature. The data are difficult to process and at present it is practically impossible to integrate them into a common molecular cytogenetic database. A more recent approach to overcoming these obstacles, at least to some extent, has been to define cytogenetic landmarks using homogeneously spaced FISH anchor probes along the chromosomes and the corresponding ideograms.
It has long been recognized that the chromosomal banding pattern alternately reflects CG-rich and CG-poor sequence sections. Nonetheless, the prevailing view was that the overall correlation between chromosome bands and the CG gene45 can only be viewed as a rather weak approximation. In this context, it seemed highly unlikely that the banding phenomenon should simply be the result of long-range changes in linear base pair composition alone. Rather, it was assumed that the banding pattern is significantly influenced and modified by structural factors such as folding, protein coverage, DNA packing and condensation as well as the accessibility of the DNA 50 through colors.
Finally, a direct computational comparison between the sequence-specific CG content and the special staining pattern that recently became possible further confirmed this relationship (Niimura and Gojobori ("In silico chromosome staining: Reconstruction of Giemsa bands from the whole human genome sequence, PNAS, Volume 99 , No. 2 (797-802)). Your “in-silico55 chromosome staining” was carried out using a procedure with two windows, one local window
AT 41 2 476 Β of 2.5 mb and a regional window of 9.3 mb, where the relationship between the CG content in the local window with respect to the GC content in the regional window was calculated. According to Niimura and Gojobori, this two-window method would produce an in-silico coloration better than explaining the Giemsa banding pattern simply by the difference in the base composition. Further, it is believed that the efficiency would be improved if the difference in compression ratio between G- and R-bands were taken into account, G-bands being more condensed than R-bands. By calculating 10 kb fragments of the human DNA sequence and various statistical analysis options, they found that at a band level of 850, the CG content and the Giemsa bands along a chromosome are best at a local window of 2, 5 Mb and a regional window of 9.3 Mb were correlated. These windows were sized to optimize the match between in silico and Giemsa bands, but the authors also found that their approach might not be suitable for fine bands smaller than the local window. However, they assumed that the correlation between Giemsa and in silico bands could be further improved by integrating the genome-wide FISH mapping data. In this publication, however, no chromosomes were condensed, but ideograms were calculated and compared. According to Niimura and Gojobori, their results show that Giemsa banding patterns cannot be explained simply by the difference in the base composition. The relationship between the nucleotide sequence and cytogenic bands would thus remain illusory.
Previous attempts to link the cytogenetic map to the sequence of the human genome have focused on a top-down approach either by delineating band boundaries or by setting band-independent cytogenetic landmarks with specific FISH probes. For example, the “BAC Resource Consortium” placed 7,600 such cytogenetically defined landmarks on the sequence design of the human genome. Although these markers should, among other things, “also allow a strict assessment of sequence differences between the dark and light bands of chromosomes, their position is still illustrated on ideograms that are not linked to the DNA sequence. Likewise, the Cancer Chromosome Aberration Project (CCAP) sets up a cytogenetic coordinate system with arbitrarily defined intervals based on an ideogram to show the location of sequence-anchored FISH clones.
In Jingwei et al. (PNAS 94 (1997), pp. 6862-6867) the GC content of chromosome parts is generally examined.
US Pat. No. 6,136,540 A describes a computer program for finding genetic abnormalities with which the subjective analysis of selectively colored chromosomes can be avoided. According to this document, the chromosomal abnormalities are determined and analyzed by specific hybridization probes (which are marked with fluorophores), whereby chromosomal additions, deletions, amplifications, translocations and rearrangements can be recognized. Although, according to this document, the disadvantages of experimental chromosome staining or their problems with reproducibility are to be avoided, experimentally complex work is also carried out here, namely with fluorophore-labeled hybridization probes, which again results in reproductive difficulties and the inaccuracies inherent in experimental evidence.
According to Daigo et al. (DNA Res. 6 (4) (1999), pp. 227-233) the GC content is determined on certain sections of chromosome 9 and 3 using conventional methods.
The article by Hraber et al. (Genome Biology 2 (9) (2001), research 0037.1-0037.14) relates to an analysis option for examining interspecific interactions in sequences that are expressed during the interaction between two symbionts, for example with regard to their GC content. However, neither complete, larger areas of the genome are compared with one another, nor are they specifically represented as chromosomes.
It is therefore an object of the present invention to provide a method for producing a virtual chromosome which enables an accurate banding pattern in a scale-independent and highly area-specific manner with high resolution. Furthermore, these created virtual chromosomes should not only contain morphological information like the conventional ideogram3
AT 41 2 476 Β me or ISCN bands, but also the corresponding genetic information, for example the sequence data. Such a virtual chromosome, which has the complete sequence data, is comparable with the conventional representations of chromosomes and has a high resolution, has not yet been created. It is therefore another object of the present invention to provide a set of chromosomes which can be used as an interface between conventional representations of chromosomes and genetic information and sequence data to directly compare DNA sequence derived data and natural chromosomal banding patterns.
The aim of the present application is achieved by a method according to the invention as defined above, which is characterized in that it comprises the following steps:
- Subdivide sequence data of the natural chromosome into fractions with a length of at least 10,000 bp,
- determining the CG content in each fraction,
- calculating a value between a minimum value and a maximum value for each fraction according to the CG content, and
- Creation of the virtual chromosome by representing each fraction with the value.
It turned out that chromosomes produced with this method provided an excellent correlation between one's own banding pattern and that of their corresponding natural counterparts. This surprising agreement not only shows that the chromosomal banding pattern is largely determined directly by the underlying DNA sequence, but can also provide a unique basis for the joint processing of morphological and molecular genetic data within a single framework based on the DNA sequence . In contrast to current publications, which expressly state that Giemsa banding patterns cannot be explained by the different base composition alone, the present method shows that a representation of chromosomes on a sequence basis according to the CG content is very possible and virtual Leads to high resolution chromosomes.
In the context of the present application, the expression “method for producing a virtual chromosome which represents a corresponding natural chromosome” relates not only to complete chromosomes, but also to parts thereof, for example separate chromosome arms or ends. For the sequence data of the natural chromosome, for example, any electronically available data can be used, in the case of human chromosomes, for example, sequences from the working draft of the human genome project [Human Genome Project Working Draft (http://genome.ucsc.edu/)] . In the context of the present application, the term “chromosome” refers to any chromosome of any organism. The organism is, for example, humans, but any living being, in particular a mammal, can also represent the organism from which the chromosome originates. Mammalian and human virtual chromosomes are particularly preferred because they can be used for evolutionary studies.
The step of dividing the sequence data into fractions and determining the CG content in each fraction is preferably carried out electronically.
The expression “value between a minimum value and a maximum value” relates to any parameter that is suitable for defining a specific CG content and, preferably, for representing it virtually. This can be, for example, a percentage value between 0% and 100% or a value between 0 and 1 that is made visible in a two- or three-dimensional image, for example. The values can also be represented by light values or color values. Any value that can be made visible is suitable to represent a certain CG content.
The step of calculating the value for each fraction according to its CG content can be carried out with any suitable table or formula, algorithm or program, for example a minimum value being assigned to a minimum amount of CG in a fraction and a maximum value being assigned to a maximum amount of CG is assigned to a faction. The values in between are then assigned as a linear function between the two extreme values to the changing CG content in the fractions.
The value is preferably a light value. The advantage of a light value is that the
AT 41 2 476 B
Visualization can be interpreted very quickly and easily and can also be compared with conventionally generated, for example microscopically removed or ISCN chromosomes.
In a further preferred manner, the maximum value is shown in white, the minimum value in black and values in between in shades of gray. This representation of a virtual chromosome is directly comparable to the conventionally scanned chromosomes, but the resolution is very high and the virtual chromosome further comprises the sequence data information that is missing in the chromosome representations of the prior art.
According to a preferred embodiment, the natural chromosome is divided into fractions with a length of up to 1,000,000 bp, preferably a length of 50,000 to 500,000 bp, more preferably a length of 100,000 to 300,000 bp. These fractions are sufficiently small to allow high resolutions and the greatest possible wealth of information. Preferably a fraction corresponds to the estimated average size of a DNA loop as well as that of an isochore. An optimal fraction length is, for example, 200,000 base pairs.
The fraction with a CG content of 30 to 35%, preferably 33%, is advantageously assigned to a minimum value and the fraction with a CG content of 60 to 65%, preferably 62%, is assigned to a maximum value. A change between these percentages has been found to produce chromosomes with gray band values that correspond to the conventionally displayed chromosomes. This visualization of the virtual chromosome can therefore be compared directly with conventional chromosome representations, such as those perceived through a microscope, for example.
Preferably, fractions with an unknown sequence are assigned the value according to their morphological appearance. Even if the amount of fractions that are missing a sequence, especially in the case of human chromosomes, has become very small due to the almost complete human genome and will generally decrease rapidly, the few fractions with a missing sequence can be supplemented with data from the morphological appearance derived, e.g. taken from an ideogram, to provide a complete chromosome.
In a further preferred manner, after the production of the virtual chromosome, a filter for smoothing the appearance, preferably a Gaussian convolution filter, is used. The resulting shading was also used to gradually fill in the last few pixels at the chromosome boundaries.
According to another preferred method, a scale correction filter is used to produce the virtual chromosome. This can be a normalization and a non-linear gray scale correction filter of the gamma type. This results in a contrast enhancement and the image of chromosomes as perceived through the microscope is optimally mimicked.
Another aspect of the present application relates to a virtual chromosome or a part thereof, which is represented by values according to its CG content, which is characterized in that the production takes place according to the above-defined method according to the invention. According to the present invention, a virtual chromosome is created which not only comprises morphological information and which is compared, for example, with conventional ideograms can be , but which also contains sequence data. This sequence-based visualized chromosome according to the present invention shows an excellent correlation of the banding pattern of virtual chromosomes and that of the corresponding natural counterparts. With regard to this aspect of the present invention, the same definitions and preferred embodiments apply as above.
The value is preferably a light value, more preferably the maximum value is white, the minimum value black and values in between are shown in shades of gray. As stated above, this enables a representation of the chromosome as it is seen through the microscope and is therefore ideally suited for a comparison with conventional chromosome representations, such as ideograms or microscopically perceived chromosomes.
According to a further aspect of the present application, a set of virtual chromosomes or parts thereof is provided, which is characterized in that it comprises two or more chromosomes according to the invention or parts thereof as defined above.
AT 412 476 Β
The set preferably comprises a maximum number of chromosomes, it being possible for this set to be continuously supplemented by further newly found or newly identified chromosomes.
The set preferably comprises chromosomes or parts thereof which are specific for one or more organisms. The advantage of a set that is specific to an organism is that that set is useful for comparing any newly identified modifications or rearrangements of chromosomes in that organism. However, it is of course possible to provide a set with chromosomes from different, preferably defined, organisms.
More preferably, the set 24 comprises human chromosomes or parts thereof. This set is a standard set for normal human chromosomes and can be used to compare chromosomes from a patient with normal chromosomes to detect any modifications or rearrangements.
In a further preferred manner, the set further comprises additional modified chromosomes or parts thereof, preferably chromosomes with translocations. This is particularly advantageous for modified chromosomes that are related to a specific disease, for example a specific tumor. By providing a classification of such modified virtual chromosomes, preferably any modification with an indication of a specific disease, it is very easily possible to assign a disease or the risk of a disease breaking out to a set of chromosomes isolated from a patient by using the Set of virtual chromosomes is compared with the patient's chromosomes. Due to the constant detection of new modifications in chromosomes, the set can be completed quickly and permanently with the latest medical information.
In the context of the present application, the term “chromosome modification” relates to any sequence modification, for example any mutation or translocation of a chromosome fragment.
Another aspect of the present invention is the use of the set of virtual chromosomes according to the invention as mentioned above for cataloging chromosome modifications. As stated above, the kit according to the invention is particularly useful for providing electronic information about chromosome modifications and their connection to any disease or risk of disease. In the context of the present application, the term chromosome modification refers to any modification in the chromosome. It can be a sequence mutation or the complete translocation of a chromosome fragment. Due to the high resolution, every chromosome modification can be detected and cataloged.
The sentence according to the invention allows the description of chromosomal anomalies with a previously unknown molecular precision, while on the other hand there is still the possibility of interpreting more blurred major events on the pure chromosome level, as it is also by means of conventional cytogenetic analysis and by means of comparative genome hybridization and comparative expressed sequence hybridization on the basis of chromosomal multi-color staining FISH is possible.
Another aspect of the present application relates to the use of a set according to the invention of virtual chromosomes as defined above for virtual mapping of the chromosomal position of a sequence. Since the set of virtual chromosomes according to the invention is derived from the complete human DNA sequence or this aarsteiit, it is possible to map and display the chromosomal position of any given, known or unknown sequence or group of sequences contained in a database for the production of the virtual chromosomes.
Another aspect of the present application relates to the use of a set of virtual chromosomes according to the invention as defined above as an interface between morphological and molecular genetic data. Preferably, the morphological data are derived from information based on the International System for Human Cytogenetic Nomenclature (ISCN).
A graphic interface can be superimposed over any molecular genetic database. A virtual system based on chromosomes ensures that previously recorded data remains accessible and analyzable. For example, a combination of graphic cut6
AT 41 2 476 B, tools based on ISCN nomenclature and virtual chromosomes, which can overlay existing cytogenetic databases as explained above, are extremely useful. Such an interface enables the transformation of ISCN information into the corresponding karyotype image. Conversely, a karyotype image generated with such a virtual chromosome tool can be translated into an ISCN tape. Such a graphical interface is extremely valuable for visually cross-checking the ISCN description by comparing the karyotype image with the virtual chromosome image. As a valuable by-product, such an approach also significantly improves the quality of cytogenetic data. In addition, it also facilitates the seamless exchange and transfer of cytogenetic data in standardized form in a laboratory or between laboratories not only with a remote control center, but also with FISH and molecular genetic databases. For example, cytogenetic data that are being prepared for publication can then be easily checked and conveniently transmitted to a central database.
ιιλπ »/ ir- + i ole» nCrbni + fotollo i'ihor mrilal / i full VIILUGIIÜII VI II VI I IVOUI I lül I yiUllWIIV Wl 11 IlllUkVIlV IIIVIVIM.
ΓΊϊ II z- »r · I z ·« · ”ι ir ·» / - »l / ic uuciiayoi ui iy tische databases supports the visualization of all types of FISH-, DNA- and RNA-derived data sets as well as gene expression profiles on standardized "Chromosomal" way. The advantages of such a chromosomal representation are that it is independent of the probe distribution on the various arrays, and also that its “natural” appearance facilitates understanding and comparison by visual examination. In addition, such a representation of gene expression profiles is becoming more and more important because there is increasing evidence that functionally unrelated genes are also expressed in transcription territories in Drosophila and in the human genome. The resulting distribution pattern is similar to that obtained from CGH and CESH analyzes, in which differently labeled DNA or cDNA from a tissue of interest and a control sample are hybridized directly to chromosomes at the same time. As a result, such data sets can be directly correlated and cross-analyzed with other karyotype patterns, for example the associated karyotype anomalies.
Since the position of FISH or other DNA probes (from oligonucleotides to region-specific staining probes) can be conveniently displayed on virtual chromosomes and cross-checked with the hybridization patterns actually obtained, such graphic interfaces are also of potential interest for resource centers. In addition, the position of cytogenetic landmarks in the form of evenly distributed FISH probes can be integrated directly into such chromosomes for the purpose of mapping defects. Finally, even submicroscopic events that are otherwise undetectable with conventional cytogenetic means, such as microdeletions and interphase FISH data, can be mapped and included in such a universal platform.
The chromosome set is preferably used as a reference for classifying a phenotype into a sequence arrangement. The term “sequence arrangement” refers to any modification, e.g. sequence mutations or translocations of chromosome fragments. The phenotype can refer to normal or abnormal phenotypes, e.g. various diseases such as tumors. In particular, when the chromosomes are classified according to modifications and resulting phenotypes, each chromosome isolated from a patient and analyzed using conventional microscopic methods can be compared with the chromosome set according to the invention. Similarities between the modifications of the chromosomes would also imply similar phenotypes, or at least the likelihood or risk of developing a similar phenotype.
The chromosome set advantageously serves as a tool for carrying out structural or functional analyzes of a sequence arrangement. As stated above, the analysis can be carried out by means of gene mapping or virtual hybridization on the basis of the sequence data contained in the virtual chromosome.
In a further preferred way, the set of chromosomes serves as a tool for determining the influence of a certain factor on a sequence arrangement. For example, an external factor such as a chemical substance, energy with different wavelengths or the influence of microorganisms can be analyzed on a cytogenetic basis and transferred to the chromosome set according to the invention or compared with it, thereby creating implications or
AT 412 476 Β resulting phenotypes can be derived or foreseen.
The present invention is described in more detail with reference to, but not limited to, the following examples and figures, in which:
1 shows images of a DNA sequence derived from human chromosomes in comparison with trypsin / Giemsa banded chromosome images;
Fig. 2 shows a model of uneven condensation of G-bands as a function of their
Represents CG content;
Figure 3 shows virtual chromosomes compared to the cytogenetic and molecular genetic maps;
Figure 4 shows virtual in situ hybridization;
Fig. 5 shows the construction of virtual chromosomal abnormalities; and
Figure 6 illustrates a graphical interface between cytogenic and molecular genetic data sets.
Examples Example 1:
Production of virtual chromosomes based on genetic data
To create virtual chromosomes, the sequence data from the August and December 2001 and April and June 2002 editions of the human genome working draft were downloaded. Perl, Mathematica (Wolfram Scientific) and Photoshop (Adobe) were used to analyze and process the data and the resulting images. Using the script language Perl, the sequence data of each individual chromosome was first divided into 200,000 base long fractions. This size corresponds to the estimated average size of a DNA loop and also approximately that of an isochore. The CG content of all strips whose sequence was at least 70% complete was then determined as a percentage. This was the case in practically all cases beginning with the December 2001 issue. Gaps due to unsequenced nucleotides (N's) were not considered. The CG content of the individual strips ranged from 33% to 62% with a mean value of 41%. The tables with these data were saved together with the information about the segment and band coordinates of the respective chromosome in temporary files and further calculated and assembled with Mathematica as shown below:
StaticNormalize [L_List]: = Block [{Lx = 0.625, Li = 0.33}, # 1-Li
Lx-Li & / @ L]
StaticNormalizefL Real]: = Block [{Lx = 0.625, Li = 0.33}, ———]
Lx-Li
Corresponding to the individual CG content, a gray value was assigned to each stripe in a linear normalized manner, ie stripes with a CG content of 33% became black and those with a CG content of 62% became white. The transfer function in Mathematica from percent CG to the gray value is the sum of (statistical) normalization and contrast enhancement:
BandsAvg = ListConvolve [FoldMask, StaticNormalize [ChrShades], {CenterElement, -CenterElement}, .41] / (Plus @@ FoldMask)
The derived bars were then integrated within the respective chromosome boundaries as defined by their length and centromeric position. The centromeric, heterochromatic and satellite regions, for which no suitable sequence information is yet available, have been artificially supplemented according to their morphological appearance. To smooth out the appearance of the virtual chromosomes, a Gaussian convolution filter (N (0.1), 22 strips in length) was applied:
NiceUnitBand [Pos_, Width_, Stain_, BitFields_, Y_, H_, ChrNames_]: = Block [{CL},
CL = Select [{ChrStartPos [ChrNames], ChrSatelitePos [ChrNames]} ~ Join ~
ChrCentroPosList [ChrNames] - Join - {ChrEndPos [ChrNames]}, NumberQ];
AT 412 476 Β
Raster [Table [lf [False, {Max [O, Min [1, Section] [x]]]}, {Max [O, Min [1, Section [ColorCorrection [Stain]] [x]]]}], (x, 0, 1, -¼].
wU {{Pos - Width / 2, Y + H * (1 - Boundary [CL, Pos])}, {Pos + Width / 2, Y + Η - H * (1 - Boundary [CL, Pos])}} ,
ColorFunction —► GrayLevel],
Maplndexed [
Rectangle [{Pos - Width / 2, - (# 2 [[1] J * 10 + 1)}, {Pos + Width / 2, - (# 2 [[1]] * 10 + 9)}, 10 Graphics [{}, Background If [# 1> = 0, Hue [# 1], GrayLevel [1]]]] &, BitFields]}] jltisrsndcn shades are used to gradually fill in the last few pixels at the chromosome borders . A normalizing and non-linear gray-scale correction filter of the gamma type was used to enhance the contrast and mimic the chromosome images as seen through a microscope:
ColorCorrection [c_]: = (* 1 + (c-1)<sup>Λ</sup> 3 *) Interpolation [{{0,0), {.1, .1), {.2, .5}, {.5, .8}, {1, 1}, 20 InterpolationOrder -> 1] [c ]
To bring the data into chromosome form:
Boundary [CL_, p_]: = Block [{r = 5000000, d}, 25 d = Min [Abs [CL-p]];
lf [d <r, l ^ 1- (l- £ j<sup>2</sup> + N [2/3], 1]]
As a last step, the respective images were then imported into Photoshop, assembled and arranged accordingly.
Example 2:
Images of DNA sequence-derived human chromosomes
In Fig. 1 Trpysin / Giemsa-banded chromosome images are shown, wherein for each chromosome (a) shows an ISCN reference image with 850 band levels, (b) shows its derived straight gray shading pattern, (c) a comparison with its calculated virtual counterparts of the August 2001 and (d) December 2001, (e) April 2002 and (f) June 2002 Issues 40 (see http://genome.ucsc.edu/).
Despite the excellent overall overall concordance between the matched sets of chromosome homologues, some local changes and differences became apparent, particularly between virtual chromosomes derived from different sequence outputs. The originally excellent correspondence between the gray scale45 banding pattern of the natural chromosomes and the virtual chromosomes of the August 2001 edition deteriorated when the December 2001 edition was used for comparison purposes. This amazing finding can be explained by the fact that the later compilation was produced by CNBI and not by USCS. When compared with the UCSC compilation, the NCBI compilation shows slightly better local order 50 and orientation, but slightly poorer tracking of the chromosome level maps. Such shifts of sequence sections can thus change the local banding pattern noticeably, which becomes particularly clear when comparing the long arms of the virtual chromosomes 1 and 11 from different editions. The position of the centromeres of chromosomes 5 (December 2001 edition), 7 and 12 (both 2002 edition) shifted to an odd 55 positions. It is noteworthy, however, that the pre-storage and constant corrections in
AT 412 476 B significantly improved the sequence composition as did the concordance between the natural and virtual banding patterns. Such a comparison of virtual chromosomes that originate from different sequence outputs can thus also enable an independent validation of the sequence map.
Example 3:
Modeling of uneven condensation of G-bands as a function of their CG content
Virtual chromosomes allow structural and functional analyzes of the genome and create opportunities to investigate the influence of various factors on the large-scale chromosomal banding pattern. The example shown herein relates to the potential effects of the unequal contraction of light and dark bands during chromosome condensation. The analysis is based on the idea that dark Giemsa bands can contain up to eleven times more DNA than the light bands and that the DNA compression ratio is in the order of magnitude of the cube root of the respective DNA length (see FIG. 2).
First, the images of the shortest (ISCNS, band level 500) and longest (ISCNL, band level 850) ISCN chromosome 7 were transformed into a gray scale pattern by measuring the gray values along the blue path (a). After these two chromosome images were brought to the same length, their banding pattern was compared with that of virtual homologues, which were modified in different ways. Depending on the respective CG content and as explained in the graphic (see Fig. 2b) the length of the light and dark bands of the virtual chromosomes was linearly stretched or condensed by applying the factors 0.3, 0.5 and 0.8, which approximate a 2.2-, 3.4- and 5.8 times the difference in their DNA length correspond (see Fig. 2c). A comparison of the resulting banding patterns reinforces previous experimental evidence that chromosome condensation is in all likelihood not just in a CG content-dependent linear fashion. However, it cannot provide a good explanation for the startling pattern obtained by stretching GTG-banded chromosomes. It is conceivable that by determining the distances between light and dark chromosome bands in different stages of contraction it will one day be possible to derive a factor or a formula, the plausibility of which is subsequently checked by comparing the images of natural chromosomes with the corresponding virtual chromosomes can. Even if the uneven elongation of condensed and decondensed dark and light chromosome bands can even be perceived visually, there is no practical need for analytical purposes, since the preparation-dependent change within the chromosome class itself is wider than that caused by the distorted artificial elongation by means of a computer.
Example 4:
Virtual chromosomes link the cytogenetic and the molecular genetic map
In Fig. 3a the ISCN chromosome 7 (band level 850) is shown together with its virtual counterpart and three different ideograms. The banding pattern of the left ideogram is based on the position of the turning points between CG-rich and CG-poor regions in the sequence-based virtual chromosome. The curve follows the mean CG content. The UCSC ideogram (August 2001 edition) is placed in the middle and the corresponding ISCN (band level 850) on the right. The left half of the virtual chromosome shows the raw, non-amplified gray values of the respective CG content, whereas in the right half the contrast is increased according to the curve shown in diagram (b). The thin horizontal lines provide an absolute 10 Mb scale. However, as explained above for Figure 2, the DNA may not be distributed along the bands in such a linear fashion as this scale suggests. It will also be clear that the width and distribution of ribbons in ideograph representations can vary considerably, although their number and names are usually the same. It is therefore not possible to precisely position every absolute or relative chromosomal occurrence on any ideogram. Virtual chromosomes solve this problem by linking the absolute precision of the DNA sequence positioning with the arbitrary position indicators of some kind of ideogram, which is illustrated by the lines that the Band10
AT 412 476 Β combine the limits of the three examples shown in this figure.
Example 5:
Virtual in-situ hybridization
Since virtual chromosomes symbolize the DNA sequence in a very condensed form, it is now possible to determine the exact position of any type of DNA sequence or sequence set regardless of the number and DNA sequence length using the respective nucleotide coordinates with a previously unknown cytogenetic precision to display. A comparative example of conventional and virtual FISH mapping is shown in Fig. 4th shown on the left for the MLL partner gene GRAF at 5 (q31), whose original, cytogenetically determined position was graphically restricted to CGH software (Vysis, Döners, Grove, USA). The metaphase image is shown above, the CGH mapping image in the middle, and the virtual chromosome 5 with the elongated natural one below.
Earlier attempts to integrate the cytogenetic map position into existing genomic databases had to rely on such a display of chromosomal events on ideogram coordinate systems, since it was not possible to link these two data sets directly. The distribution of 77 of 82 chromosome 7 CCAP-BAC clones is shown here as an example of the improvement that can be achieved in the assignment of the absolute and relative position of an entire set of clones.
Example 6:
Construction of virtual chromosome abnormalities
The definition of a special chromosomal event with molecular precision now facilitates the exact reconstruction of every chromosome rearrangement with a known molecular defect location. As illustrated here by the example of the translocation t (4; 11) (q21; q23) (see FIG. 5), this is an important prerequisite for the potential use of such virtual chromosome rearrangements in pattern recognition systems and in automatic karyotyping. With current cytogenetic terminology and precision, the location of the defect sites of a particular translocation can only be arbitrarily defined by the location of the ligaments in question. Depending on the sequence release, volume 4 (q21) comprises between 11.5 and 14.2 Mb (6.0% - 7.4% of chromosome 4) and volume 11 (q23) between 10.9 and 11.7 Mb ( 7.9% - 8.5% of chromosome 11). Without knowing the exact position of the two genes AF4 and MLL, which are destroyed and fused as a result of the translocation, within the ligaments, the defect sites could lie somewhere within these ligaments. To demonstrate this point of view, they have been assigned to the outer limits of the bands in question. It is very likely that this is already one of the highest resolutions that can be achieved with an average morphological chromosome analysis. Compared to the length and banding patterns of derivative chromosomes derived from precisely positioned molecular defects (indicated by *), those resulting from undefined defect assignments may look quite different. They would therefore certainly be useless for the purposes of comparison and renewed evidence in the case of pattern recognition.
Example 7:
Graphical interface between cytogenetic and molecular genetic data sets
As the top units of the human sequence, virtual chromosomes cover the nine orders of magnitude of the complete genome in a highly condensed, easily expandable and most naturally imaginable “morphological” way. They can therefore be used as a unique front-end tool for visualizing the information contained in each sequence database; namely in principle from a single base pair to whole chromosomes in a cytogenetic way. As an example, the distribution of approximately 15,000 genes and CpGs from the UCSC database along virtual chromosomes and the corresponding UCSC color ideograms are shown here on a 1 Mb scale. For practical reasons, the scale of the bars for the genes is only shown halfway up the CpG bars and the height of these CpG bars on chromosome 19 is cut off.
The vertical bars on the left side of the chromosomes indicate the size and location of heterochromatic and satellite regions that have been artificially added as their sequence
AT 41 2 476 B is currently not yet available. The fine horizontal bars on the left side of the chromosomes indicate gaps in the sequence.
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both waysCites: the store holds 1 of 2
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US6136540A | Cites | United States of America | Search report |
7 members in 4 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 14302002 | Austria | A | |
| AT20020001430 | – | – | – |
Members7
| Document | Office | Kind | |
|---|---|---|---|
| WO2004029747A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2003275968A1 | Australia | A1 | |
| AU2003275968A8 | Australia | A8 | |
| ATA14302002A | Austria | A | |
| AT412476BThis record | Austria | B | |
| WO2004029747A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1563444A2 | European Patent Office (EPO) | A2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Ceased due to non-payment of the annual feeCeasedELJ | ELJ |
Numbers
- Publication, DOCDB
- 412476
- Publication, EPODOC
- AT412476B
- Application
- 143002
- Application, DOCDB
- 14302002
- Application, EPODOC
- AT20020001430
Titles2
- English
- METHOD FOR PRODUCING A VIRTUAL CHROMOSOME
- German
- VERFAHREN ZUR HERSTELLUNG EINES VIRTUELLEN CHROMOSOMS
Classification
- CPC, 4
- G06F19/26
- G16B30/00
- G16B45/00
- G06F19/22
- IPC, 5
- C12N15 00
- C12Q1 68
- G06F
- G16B30 00
- G16B45 00