Stain-based optimized compression of digital pathology slides
17 claims: 10 independent, 7 dependent
- 1デジタル 染色 画像を圧縮する方法であって、各色変換が、特定の染色 方法 に対応し た サンプル標本を含 む トレーニング 画 像にしたがって算出され 、各色変換が特定の染色方法に対応してRGB成分を脱相関する、 複数の色変換を事前に計算するステップと、 特定の染色方法でマッピングされた 入力デジタル画像スキャンを 該入力デジタル画像スキャンの染色方法に対応する色変換を使用して 圧縮するステップと、を含む方法。
- 2前記色変換が主成分分析(PCA)を使用して算出される、請求項1に記載の方法。
- 3前記複数の色変換がデータベースに記憶される、請求項1 または2 に記載の方法。
- 4各色変換の主成分に従って時間-周波数変換用の量子化ステップを選択することをさらに含む、請求項1 乃至3のいずれか に記載の方法。
- 5最小の量子化ステップが第1の主成分に適用され、次に大きな量子化ステップが第2の主成分に適用され、最大の量子化ステップが第3の主成分に適用される、請求項 4 に記載の方法。
- 6前記デジタル染色画像がデジタル病理染色画像であり、 前記染色方法が、組織化学染色方法である、請求項1乃至5のいずれかに記載の方法。
- 7入 力デジタル画 像 スキャンを特定の染色にマッピングするステップ を含む、請求項1乃至5のいずれかに記載の方法。
- 8前記入力デジタル画 像 スキャンを特定の染色にマッピングするステップが、ラボラトリインフォメーションシステム(LIS)から読みこんだ情報に基づく、請求項7に記載の方法。
- 9前記入力デジタル画像スキャンを受け取るステップ(120)と、 前記入力デジタル画像スキャンに適用された染色方法を判定するステップ(122)と、 前記複数の色変換の中から前記対応する色変換を検索するステップ(124)と、 を含む、請求項1乃至8のいずれかに記載の方法。
- 10前記複数の色変換を事前に計算するステップが、 特定の組織種の特定の組織化学染色方法を代表する、サンプル病理標本を含む1組のトレーニング 画像 からサンプル染色データを受け取り、サンプル染色データから入力ベクトルを形成するステップと、各トレーニング 画像 ごとに、前記入力ベクトルを利用して色変換を算出し、結果として得られた行列係数をデータベースに記憶するステップと、を含む 、請求項1乃至9のいずれかに記載の 方法。
- 11前記事前に 算出する前記ステップは、前記トレーニング 画像 の共分散行列を算出するステップと、前記共分散行列の固有ベクトルおよび固有値を算出するステップと、前記固有値によって前記固有ベクトルを正規化するステップと、前記正規化された固有値を前記固有値によってソートするステップと、を含む、請求項1 乃至10のいずれか に記載の方法。
- 12デジタル 染色 画像 の 圧縮を実行するサーバコンピュータであって、 各色変換が、特定の染色方法に対応したサンプル標本を含むトレーニング画像にしたがって事前器算出され、各色変換が、特定の染色方法に対応してRGB成分を脱相関する、複数の色変換を記憶するデータベースに接続し、 複数の デジタル画像 スキャンを記憶するように構成された画像記憶デバイスと、 特定の染色方法でマッピングされた 前記入力デジタル画像 スキャン該入力デジタル画像スキャンの染色方法に対応する色変換を使用して を圧縮する画像圧縮モジュールと、を備え る サーバコンピュータ。
- 13前記色変換が主成分分析(PCA)を使用して算出される、請求項12に記載の サーバコンピュータ 。
- 14前記画像圧縮モジュールが、各色変換の主成分にしたがって時間-周波数変換用の量子化ステップを選択する手段をさらに備える、請求項12 または13 に記載の サーバコンピュータ 。
- 15最小の量子化ステップが第1の主成分に適用され、次に大きな量子化ステップが第2の主成分に適用され、最大の量子化ステップが第3の主成分に適用される、請求項14に記載の サーバコンピュータ 。
- 16コンピュータメモリにロードされたときに 請求項1乃至11のいずれかに記載の方法 が実行されるコンピュータプログラム 。
- 17デジタル染色画像の圧縮を実行するシステムであって、 各色変換が、特定の染色方法に対応したサンプル標本を含むトレーニング画像にしたがって事前器算出され、各色変換が、特定の染色方法に対応してRGB成分を脱相関する、複数の色変換を記憶するデータベースと、 特定の染色方法でマッピングされた前記入力デジタル画像スキャン該入力デジタル画像スキャンの染色方法に対応する色変換を使用してを圧縮する画像圧縮モジュールと、 を備えるシステム。
Independent claims17
77 paragraphs, as filed
The subject matter disclosed herein relates to the field of digital imaging, and in particular to the mechanism for optimized compression by staining of digital pathological slide scans.
Pathology is the examination and diagnosis of a disease, usually by examining the body tissue under a microscope. Currently, pathologists make diagnoses by manually inspecting stained tissue samples on glass slides under a light microscope. Tissue samples are usually stained by a specialist called a hist technician. Currently, pathologists are observing slides of tissue samples using a light microscope. This process hasn't changed much over 100 years. Because the process is thus manual, the correct slide must be physically delivered to the appropriate pathologist, which can delay the initial diagnosis and subsequent second opinion.
Digitizing tissue sample images makes evaluation easier and faster without organizing, transporting, and managing glass slides. Using digital pathology technology reduces time and improves the pathologist's overall diagnostic process. The demand for such technologies and solutions is high due to increasing pressure to increase medical costs and the widespread need to digitize patient medical records. This field of digital pathology is known as Hall Slide Imaging (WSI), which digitally scans the entire slide so that it can be viewed on a computer.
The technique involves scanning the tissue-prepared glass slides. Since scanning of slides is done at very high resolutions, the uncompressed digital output of slides typically has a very large size, for example 10GB to 30GB representing an image of about 40,000 pixels x 40,000 pixels.
The next step in the whole slide imaging scheme is to compress the digital slide. To effectively store and stream digital images, lossy compression techniques must be used to compress the digital slides. The compression algorithm used preferably exhibits high speed distortion performance, i.e., achieves strong compression with high visual quality. After compression, the digital slide image is stored on the image server and streamed to a client viewer located at an arbitrary location.
The problem arises in that digital pathology slide images contain a significant amount of visual content. In this case, it is difficult to compress the slide image sufficiently and at the same time maintain high visual quality.
<p><patcit num="1"><text>U.S. Pat. No. 7,187798</text></patcit></p>
<p> Therefore, there is a need for an optimized image compression mechanism that can compress large digital pathological slide images with significant visual content while maintaining high visual quality.</p>
<p> Therefore, according to the present invention, it is an image compression method, in which each color conversion is calculated according to a set of training images, a step of pre-calculating a plurality of color conversions, and digital using one color conversion. A method is provided that includes a step of compressing the image.</p><p> According to the present invention, there is a method of compressing a digital pathologically stained image, wherein each color conversion is calculated according to a set of training slide images corresponding to a specific staining, and a step of pre-calculating a plurality of color conversions. Also provided is a method comprising mapping an input digital image to a particular stain and compressing the input digital image using a pre-calculated color conversion corresponding to the mapped stain.</p><p> According to the present invention, a server computer that performs image compression of digital pathological slide scans, an image storage device configured to store multiple pathological slide scans, and an input slide scan digital image for specific staining. It features an image compression module that maps and compresses the input digital image using pre-calculated color conversions that correspond to the mapped stains, each calculated according to a set of training images that correspond to a particular stain. Further provided is a server computer in which the color conversion is calculated in advance.</p><p> According to the present invention, a set of training slide scans representing a particular histochemical staining method for a particular tissue species, a method of calculating the optimized color conversion used when compressing a digital pathology slide scan. The step of forming an input vector from the sample staining data, calculating the color conversion using the input vector for each training slide set, and storing the resulting matrix coefficient in the database. Also provided is a method in which the color transformation is calculated using Principal Component Analysis (PCA), including steps.</p><p> According to the present invention, a computer program product that performs a digital pathological staining image compression process when loaded into computer memory, and the computer program product has program code that can be used on the computer performed thereby. Computer-usable program code with computer-usable media to pre-calculate multiple color conversions, with each color conversion calculated according to a set of training slide images corresponding to a particular dyeing. A computer-enabled code that is configured to map an input digital image to a particular stain, and a pre-calculated color conversion that corresponds to the mapped stain. Further provided are computer program products characterized in that they are configured to compress input digital images using.</p><p> In the present specification, the present invention will be described only as an example with reference to the accompanying drawings.</p>
<figref num="1">It is a block diagram which shows the digital pathology system constructed by this invention.</figref><figref num="2">It is a block diagram which shows an example of the computer processing system which realizes the mechanism of this invention.</figref><figref num="3">It is a figure of a part of an example of a pathology slide image stained by H & E.</figref><figref num="4">It is a flow chart which shows an example of the optimized image compression method of this invention.</figref><figref num="5">It is a figure which shows the RGB color component of a sample pathology slide.</figref><figref num="6">It is a figure which shows an example of the 3D histogram of the sample pathology slide in RGB color space.</figref><figref num="7">It is a figure which shows the YCbCr component of the sample pathology slide.</figref><figref num="8">It is a figure which shows the 1st, 2nd, and 3rd PCA color components of a sample pathology slide.</figref><figref num="9">It is a figure of an example of a red-green-blue component histogram of a sample pathology slide.</figref><figref num="10">It is a figure of an example of the Y-Cb-Cr color component histogram of the sample pathology slide.</figref><figref num="11">It is a figure of an example of the PCA component histogram of a sample pathology slide.</figref><figref num="12">It is a flow chart which shows an example of the method of calculating the color conversion for various dyeing methods in advance.</figref><figref num="13">It is a flow chart which shows an example of the method of calculating the PCA matrix for a training set.</figref><figref num="14">It is a figure which shows the PCA matrix operation of the training set of the H & E stained pathological image.</figref><figref num="15">It is a figure which shows the PCA of an example of a 2D set multivariate Gaussian distribution.</figref><figref num="16">It is a flow chart which shows an example of the received image compression method.</figref><figref num="17">It is a table comparing the speed distortion performance of YCbCr conversion and PCA conversion.</figref>
Notation used throughout The following notation is used throughout this specification. Term definition ASCII American Standard Code for Information Interchange ASIC Application Specific Integrated Circuit (ASIC) CAD Computer Aided Design CDROM Compact Disc Read Only Memory CPU Central Processing Unit DCT Discreet Cosine Transform DICOM Digital Imaging and Communications in Medicine DNA Deoxyribonucleic Acid DSP Digital Signal Processor DVD Digital Versatile Disc DWT Discrete Wavelet Transform EPROM Erasable Programmable Read-Only Memory FIR Finite Impulse Response FPGA Field Programmable Gate Array FTP File Transfer Protocol FWT Forward Wavelet Transform GUI Graphical User Interface HTTP Hyper-Text Transport Protocol I / F Interface I / O Input / Output IP Internet Protocol IWT Inverse Subband / Wavelet Transform JPEG Joint Photographic Experts Group (Japeg) KLT Karhunen-Loeve Transform LAN Local Area Network LIS Laboratory Information System MAC Media Access Control NIC Network Interface Card PC Personal Computer PCA Principle Component Analysis PSNR Peak Signal-To-Noise Ratio RAM Random Access Memory RF Radio Frequency RGB Red, Green, Blue (red, green, blue) ROI Region of Interest ROM Read Only Memory SAN Storage Area Network SMTP Simple Mail Transfer Protocol TCP Transmission Control Protocol URL Uniform Resource Locator WAN Wide Area Network WSI Whole Slide Imaging WWAN Wireless Wide Area Network The present invention is a method and system for performing optimized image compression of digital pathology slide images. The optimized image compression mechanism of the present invention operates to realize an image compression algorithm with improved velocity distortion performance by utilizing the special color characteristics of stained tissue represented by digital pathological slides. The optimized color conversion is pre-calculated using a training set of pathological slide image scan data for each staining type. By using optimized color conversion to compress input slide image scans, image streaming becomes more efficient (acting as an image streaming platform), thereby allowing users to use hospitals, satellite centers, homes, or Very large digital slide scans obtained from any connected location, such as mobile phones, can be considered. Pathology slide observation client / server system A block diagram showing the digital pathology system constructed by the present invention is shown in FIG. The system is referred to by 10 overall and includes an observer station 12, a backend 16, and a scanner 14. The scanner 14 includes a laboratory information system (LIS) broker 26, an image acquisition unit 28 including an optimized image compression module 30, a DICOM library 32, a compression library 34, and a calculated color conversion database 36. The backend 16 includes a storage management module 38, a database 40, a storage 42 for storing image files, a streaming server 44, and a CAD / analysis block 46. The observer station 12 includes a workflow GUI18 and an observation GUI20 including a streaming client 22 and a color management unit 24. Observer stations, backends, and scanners are any suitable such as internet, intranet, wide area network (WAN), wireless wide area network (WWAN), local area network (LAN), storage area network (SAN). Communicate on the means of communication. Those skilled in the art will recognize that the present invention can be practiced using any of the various communication networks. The streaming server 44, streaming client 22, storage management device 38, and storage 42 communicate with each other using any suitable language / protocol such as SMTP, HTTP, TCP / IP.
In one embodiment, the observer station 12 and the backend 16 may include a MAC or PC computer powered by an Intel® or AMD® microprocessor or equivalent. The observer station 12 and the backend 16 may include a cache and a suitable storage device (eg, 42) such as a large capacity disc, CDROM, or DVD.
The streaming client operates at the observer station to communicate with the streaming server in the back end on the network and capture the imaging data stored in the storage 42.
Note that in one embodiment, the optimized image compression module 30 is implemented within the scanner and the calculated color conversion 36 is stored in the optimized image compression module 30. In other embodiments, compression can be performed on the back end. In this embodiment, the compressed image is generated, stored on the backend, streamed to the observer station and displayed to the user.
Also note that the optimized image compression and observation client functionality can be implemented as a plug-in on a standard web browser. In this embodiment, the web browser comprises imaging client software and optimized image compression software that are loaded into the browser. The web browser may be any suitable browser such as Mozilla Firefox®, Apple Safari®, Microsoft Internet Explorer®, Google Chrome®.
In other embodiments, there is no backend. Instead, the observer station captures image data directly from an image storage device (such as a hard drive) on the scanner, and color conversion and image compression calculations and processing are performed on the client computer. Computer processing system As will be appreciated by those skilled in the art, the present invention can be implemented as a system, method, computer program product, or a combination thereof. Accordingly, the present invention takes the form of a fully hardware embodiment, a fully software embodiment (including firmware, resident software, microcode, etc.), or a combination of software and hardware aspects. Often, all of these embodiments are commonly referred to herein as "circuits," "modules," or "systems." Further, the present invention may take the form of a computer program product realized in any tangible representation medium having a computer-usable program code realized in the medium.
Any combination of computer-enabled or computer-readable media may be utilized. Computer-enabled or computer-readable media can be, but are not limited to, for example, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, devices, or propagation media. More specific examples (limited list) of computer-readable media include the following media: electrical connections with one or more wires, portable computer diskettes, hard disks, random access memory (RAM): ), Read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), fiber optics, portable compact disk read-only memory (CDROM), optical storage devices, transmissions that support the Internet or intranet A medium or a magnetic storage device may be included. Note that the computer-usable or computer-readable medium may, in some cases, be paper, or any other suitable medium on which the program is printed. This is because, in this case, the program is electrically captured and then compiled or interpreted, as needed, for example by optical scanning of paper or other medium, or the program is subjected to other processing in an appropriate manner and then This is because it can be stored in the computer memory. As used herein, computer-enabled or computer-readable media includes, stores or communicates, or stores or communicates programs so that they can be used by or in connection with instruction execution systems, devices, or devices. It can be propagated or moved. Computer-enabled media may include propagated data signals in the baseband or as part of a carrier wave that implement computer-usable program code. Computer-enabled program code wireless, wireline, optical
The computer program code that performs the operations of the present invention is one or more programming languages, including object-oriented programming languages such as Java®, Smalltalk®, C ++®, and "C" programming. It can be written in any combination with traditional procedural programming languages such as languages and similar programming languages. The program code can be run entirely on the user's computer as a stand-alone software package, or partly on the user's computer, partly on the user's computer and partly on the remote computer. It can be run on or as a whole on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer or to an external computer through any type of network, including local area networks (LANs) or wide area networks (WANs) (eg, the Internet). Through the internet using a service provider).
The present invention will be described below with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It will be appreciated that each block of the flowchart and / or block diagram, and the combination of blocks within the flowchart and / or block diagram, can be realized by computer program instructions. By feeding computer program instructions to the processor of a general purpose computer, special purpose computer, or other programmable data processor, the instructions are executed through the processor of the computer or other programmable data processor. It is possible to generate a machine such that a means for performing a specified function / operation is formed in a block of a flowchart and / or a block diagram.
These computer program instructions can also be stored on a computer-readable medium that can instruct a computer or other programmable data processor to perform certain functions, and thus on a computer-readable medium. The stored instructions generate a product containing command means that perform the function / operation specified in the block of the flowchart and / or block diagram.
A computer program instruction is loaded onto a computer or other programmable data processing device to perform a series of operating steps on the computer or other programmable device, and the instruction is performed on the computer or other programmable device. By being executed, it is possible to generate a computer-executed process such that a process including instruction means for performing a function / operation specified in a block of a flowchart and / or a block diagram is generated.
A block diagram showing an example of a computer processing system that realizes the optimized image compression mechanism of the present invention is shown in FIG. The computer system is referred to as a whole at 60 and includes a processor 62 that may include a digital signal processor (DSP), central processing unit (CPU), microphone controller, microprocessor, microcontroller, ASIC or FPGA core. .. The system also has static read-only memory 68 and dynamic main memory 70, all of which communicate with the processor. The processor also communicates via bus 64 with some peripheral devices that are also included in the computer system. Peripheral devices coupled to the bus include a display device 78 (eg, monitor), an alphanumeric input device 80 (eg, keyboard), and a pointing device 82 (eg, mouse, tablet, etc.).
The computer system is connected to one or more external networks, such as LAN / WAN / SAN76, via a communication line connected to the system via a data I / O communication interface 72 (eg, a network interface card or NIC). ing. A network adapter 72 connected to the system allows the data processing system to be coupled to another data processing system or remote printer or storage device through an intervening dedicated or public network. Modems, cable modems, and Ethernet® cards are just a few of the types of network adapters currently available. The system also includes a magnetic or semiconductor storage device 74 that stores application programs and data. The system may include any suitable memory means including, but not limited to, magnetic storage, optical storage, semiconductor volatile or non-volatile memory, biological storage devices, or any other memory storage device. It has a possible storage medium.
The software configured to implement the optimized image compression mechanism of the present invention is configured to reside on a computer-readable medium such as a magnetic disk in a disk drive unit. Alternatively, computer-readable media include floppy disks, removable hard disks, flash memory 66, EEROM-based memory, bubble memory storage, ROM storage, distribution media, intermediate storage media, computer execution memory, and later books. Any other medium or device that can sort the computer programs that implement the mechanism of the invention so that they can be read by a computer may be provided. The software configured to implement the optimized image compression mechanism of the present invention is integrated into static or dynamic main memory or firmware in the processor of a computer system (ie, in the microcontroller, microprocessor or internal memory of the microcomputer). Alternatively, it can be partially resident.
To the extent that other digital computer system configurations can be used to implement the optimized image compression mechanism of the present invention, and specific system configurations can implement the systems and methods of the present invention, this system configuration is It is equivalent to the typical digital computer system of FIG. 2 and is within the gist and scope of the present invention.
When such a digital computer system is programmed to perform a particular function in accordance with instructions from program software that implements the systems and methods of the invention, it effectively becomes a special purpose computer specific to the mechanism of the invention. Become. The technology required for this is known to computer system vendors.
In general, computer programs that implement the systems and methods of the invention are distributed to users on distribution media such as floppy disks and CDROMs, or on networks such as the Internet using FTP, HTTP, or other suitable protocols. Please note that you can download it. Computer programs are often copied from there to a hard disk or similar intermediate storage medium. At run time, the program is loaded from its distribution medium or intermediate storage medium into the computer's execution memory and is configured to operate the computer according to the methods of the invention. All of these operations are known to computer system vendors.
The flowcharts and block diagrams of each diagram show the architecture, function, and operation of possible embodiments of systems, methods, and computer program products according to various embodiments of the present invention. In this case, each block in the flowchart or block diagram can represent a module, segment, or portion of code that contains one or more executable instructions that perform a specified logical function. It should also be noted that in some other embodiments, the functions shown in the blocks are performed in an order other than that shown in the figure. For example, depending on the function involved, two blocks shown in succession can actually be executed at about the same time, or in some cases each block can be executed in reverse order. Each block of the block diagram and / or flowchart diagram and a combination of blocks of the block diagram and / or flowchart diagram are realized by a special purpose hardware-based system or a combination of special purpose hardware and computer instructions that perform a specified function. Also note that you can. Optimized image compression The present invention can be applied to compression of pathological images, including image scans of stained tissue samples. Staining is an auxiliary technique that has long been used in the field of microscopy to improve the contrast of microscopic images. In biochemistry, staining involves adding class-specific dyes (eg, DNA, proteins, lipids, carbohydrates) to the substrate to qualify or quantify the presence of a particular compound, as in the case of fluorescent labeling. including.
Stainings and dyes are often used in biology and medicine to highlight structures within biological tissues under different microscopes so that they can be observed. Dyes can be used to identify and test individual intracellular tissues (eg, highlight muscle fibers or connective tissue), cell populations (eg, classify various blood cells), or organelles. ..
Known hematoxylin and eosin dyes (called H & E dyes or HE dyes) are common dyeing methods in the field of histology. This is the most widely used dye in medical diagnosis, for example, when a pathologist performs a biopsy of a suspected cancerous tissue section, the tissue section is stained with H & E, H & E section, H + E section, Or it is often called an HE section.
Staining methods include the addition of (1) the basic dye hematoxylin, which colors the eosinophil structure to a blue-purple shade, and (2) the alcohol-based acidic eosin Y, which colors the eosinophil structure to a bright pink. .. One example is an RGB pathology slide image of a HE dye tissue sample. It is clear from the slide images that digital slides containing images of tissue stained by a method have the same color characteristics. This is because the histological process is the same for all slides that use the same staining method.
A flow chart showing an example of the optimized image compression method of the present invention is shown in FIG. The optimized image compression mechanism uses a transformational image compression algorithm to compress the large original image of the slide scan. Before performing this method, calculate one or more color conversions for various dyeing techniques. As described in more detail below, several sets of training images are used to pre-calculate color transformations that optimize compression results taking into account specific color aspects of a particular dyeing technique.
The first step is to obtain a dyeing method for a given slide from a calculated color conversion database. In one embodiment, this data is extracted from LIS26 (Figure 1). The first step applies one of the pre-calculated optimized color transformations to the input image scan (step 90). Then perform a Discrete Wavelet Transform (DWT) or Discrete Cosine Transform (DCT) on the result of the previous step (step 92). The resulting conversion factor is then quantized (step 94) and coded (step 96). The resulting compressed image is stored in the image file storage (step 98).
The image restoration method is a method based on the operation in which the order of the operations in FIG. 4 is reversed. First, the compressed image is decoded, and then the conversion factor is dequantized. Then apply the reverse DWT or reverse DCT, then apply the reverse color transformation. The result is an image that is close to the original image (assuming a lossy compression algorithm is used). Color conversion and YCbCr color space: The three components of a basic color digital image are red, green, and blue (RGB), as shown in FIG. An example of a 3D histogram of sample pathology slides in RGB color space is shown in FIG. Note that there is a significant visual correlation between the three RGB components shown in Figure 6, which shows that the RGB pixel values are fairly well approximated by the lower dimensional surface. Therefore, in order to improve the performance of the compression algorithm, a color conversion that decorrelates the components is used.
Any linear color transformation can be represented as a regular 3x3 matrix.
<maths num="1"><img file="JP5111583B2_D0001.tif" /></maths>
One such example is the YCbCr standard compressed color space, which is widely used as part of the JPEG and JPEG2000 image compression standards and the MPEG video compression standard. The Y component is the luminance component, and Cb and Cr are the blue difference and red difference saturation components, respectively. The YCbCr component of the sample pathology slide is shown in Figure 7. As shown, most of the energy in the image is contained in the luminance component. The transformation matrix related to YCbCr is as follows.
<maths num="2"><img file="JP5111583B2_D0002.tif" /></maths>
Time-frequency conversion: The two common transforms used in image compression are (1) the Discrete Cosine Transform (DCT), which represents a series of data points as the sum of cosine functions that oscillate at different frequencies, and (2) the wavelet. Includes the Discrete Wavelet Transform (DWT), which is an arbitrary wavelet transform that is sampled discretely.
During operation, the transformation is applied separately to the components of the color space. In most cases, the number of output time-frequency coefficients is approximately equal to the number of input data samples. However, the transformation produces a "coarse representation", i.e., only a small portion of the coefficients are significant, while the rest have absolute values less than a certain threshold. The smoother the input data, the smaller the number of significant coefficients. Therefore, the image compression algorithm by conversion preferably uses color conversion to create three components in which the second and / or third components are generally smoother when applied to RGB data. Quantization: Quantization is a lossy compression technique realized by compressing a range of values into a single quantum value. Reducing the number of discrete symbols in a given stream increases the compression of the stream. For example, reducing the number of colors needed to represent a digital image can reduce the file size of the digital image. Specific applications include DCT data quantization in JPEG and DWT data quantization in JPEG 2000. Coding of quantization coefficient In one example of an embodiment, arithmetic coding, a known technique for reversible data compression, is used. Character strings are usually represented using a fixed number of bits per character, similar to ASCII code. Like Huffman coding, arithmetic coding makes a string another form, that is, a character that is frequently used with fewer bits and a character that is rarely used with more bits. , It is a variable-length entropy coding of a form that converts into a form in which the number of bits used as a whole is expected to be reduced.
The mechanism of the present invention utilizes the special color characteristics of digital pathology slides and operates to improve the speed distortion performance of image compression algorithms. In addition, the compression algorithm acts as a platform for image streaming, allowing users to review large numbers of digital slide image files from anywhere, such as hospitals, satellite centers, homes, and even mobile phones. .. This is achieved by combining a training sample set of digital pathology slides with principal component analysis (PCA) and adaptive selection of quantization steps.
In one embodiment, the optimized color conversion is calculated in a pretreatment step for a training sample set representing a particular histochemical staining method for a particular tissue species (eg, skin, liver, etc.). The corresponding matrix coefficients of this optimized color transformation are stored in the database for future use. In one embodiment, different optimized color conversions are calculated for each dyeing method and for each individual laboratory. Color conversions are calculated for each staining method and for each laboratory because the colors of tissues stained at different locations by the same method are often slightly different.
In the examples of embodiments shown herein, the optimized color transformation is calculated by applying Principal Component Analysis (PCA) to several sets of digital training slide images. Here, the PCA method will be described schematically. PCA is a known mathematical procedure used to transform some potentially correlated variables into a small number of uncorrelated variables called principal components. The first principal component considers as much data variation as possible, and each subsequent component considers the remaining variation as much as possible. Depending on the field of application, PCA is also referred to as discrete carunen-roubaix transformation (KLT), hoteling transformation, or eigen-orthogonal decomposition (POD).
PCA is often used as a means of explanatory data analysis and to create predictive models. In PCA, the eigenvalue decomposition of the data covariance matrix or the singular value decomposition of the data matrix is usually calculated after the mean value and median value of the data are obtained for each attribute. The PCA is defined as an orthogonal linear transformation P: X Y that transforms the data into a new coordinate system, so the maximum variance of any projection of the data is located in the first coordinate (called the first principal component). Then, the second largest variance is located at the second coordinate, and the coordinates of each variance are determined in the same manner. The transformation P is called the PCA matrix.
PCA is theoretically the optimal transformation of a least squares term for given data. PCA is used to reduce the number of dimensions in a dataset by preserving the characteristics of the dataset that contributes most to the variance. This is achieved by preserving the lower principal components and ignoring (ie discarding) the higher principal components. Such subcomponents often contain the most important aspects of the data. However, this may not be the case depending on the application. It is preferred to minimize redundancy by maximizing the variance of the first output component and minimizing the final variance.
By definition, the covariance must be non-negative, so the minimum covariance is zero. Optimized covariance matrix C<sub>Y</sub>In, all off-diagonal terms are zero, so C<sub>Y</sub>Must be diagonal. In multidimensional, this is done as follows: FIG. 14 shows a flow chart showing an example of how to calculate the PCA matrix for the training set. First, the training set data is collected to form n-dimensional m measurement data represented as follows. X: = (X<sub>m</sub>), Dim (X<sub>m</sub>) = n (3) The covariance matrix is calculated from this training set vector (step 110). The nxn covariance matrix is then calculated (step 112). C<sub>X</sub>= Cov (X) (4) Next, the covariance matrix C<sub>X</sub>Eigenvector of P = {P<sub>1</sub>, ..., P<sub>n</sub>} And eigenvalues {λ<sub>i</sub>} Is calculated (step 112) and normalized (step 114). The eigenvector is λ by the eigenvalue<sub>i</sub> λ<sub>i + 1</sub>And create a PCA matrix feature vector (ie component) Y, Y: = PX (step 118). FIG. 8 shows the first, second, and third PCA color components of the sample portion of the pathology slide. An example of a red-green-blue component histogram of a sample pathology slide is shown in FIG. An example of a Y-Cb-Cr color component histogram on a sample pathology slide is shown in FIG. An example of a PCA component histogram of a sample pathology slide is shown in FIG. Note that most of the visual activity of the image is contained in the first component compared to the RGB component histogram in Figure 9.
In one embodiment, PCA techniques are used to generate color transformations of pathological images that utilize the same staining method. FIG. 12 shows a flow chart showing an example of a method of pre-calculating the color conversion for each different staining method. The PCA algorithm is used to create an optimized color transformation of the pathological image to improve the decorrelation of the output color components and thereby improve the overall velocity distortion performance of the compression algorithm. In one embodiment, the method of FIG. 12 is applied in offline mode, i.e. in advance, using one or more training sample sets of pathological images representing a particular staining method for a particular laboratory.
First, obtain one or more training sets of stained pathological images (step 100). Many pathological images are collected from the same stain. In one embodiment, 10 H & E stained images are used. It should be understood that any number of stained images can be used.
Format the training set image pixels to form a single large input vector X (step 102). Each element of vector X is an RGB pixel obtained from a training set of images. The input vector represents the input data to which PCA is applied. Note that the order of the data is not important as the PCA operates per pixel. The set of RGB pixels forms a training set of images represented by the input vector X, which is represented by:
<maths num="3"><img file="JP5111583B2_D0003.tif" /></maths>
The training set input vector X is then used to calculate the 3x3PCA matrix (step 104). Apply the PCA algorithm to the input vector X. The result is a 3x3 transformation matrix that is the basis of the optimized color transformation by dyeing.
The matrix coefficients are then stored in the calculated color conversion database or other storage depending on the staining type (step 106). If there is another training set (step 108), repeat steps 100, 102, 104, 106.
A flow diagram showing an example of a method for calculating the PCA matrix for the training set is shown in FIG. Next, the step of calculating the PCA matrix (step 104 in FIG. 12) will be described in more detail. In the first step, the covariance matrix C for training set X<sub>X</sub>: = Calculate Cov (X) (step 110). This covariance matrix is a symmetric 3x3 matrix in which the main diagonal components are the variances of each color component (RGB). The off-diagonal component is the covariance of two different components, that is, the correlation between the distributions of the two components.
<maths num="4"><img file="JP5111583B2_D0004.tif" /></maths>
Matrix C<sub>X</sub>Represents the dependency between the red-green-blue components of the pixels obtained from the training set image.
Next, the covariance matrix C<sub>X</sub>Eigenvalues λ<sub>i</sub>And the eigenvector P<sub>i</sub>Is calculated (step 112). The eigenvalues and eigenvectors can be expressed as: C<sub>X</sub>P<sub>i</sub>= λ<sub>i</sub>P<sub>i</sub>, I = 1,2,3 (7) Eigenvector {P<sub>1</sub>, P<sub>2</sub>, P<sub>3</sub>Note that} is an orthonormal basis in 3D space. Eigenvector {P<sub>1</sub>, P<sub>2</sub>, P<sub>3</sub>} Can represent any point in 3D real space, for example any RGB pixel x = (r, g, b) to a new pixel Px: = (P<sub>1</sub>x, P<sub>2</sub>x, P<sub>3</sub>Can be converted to x).
Next, the eigenvector is normalized by the special factor σ as follows (step 114).
<maths num="5"><img file="JP5111583B2_D0005.tif" /></maths>
The output value of the optimized color conversion is limited to the precision range of 1 byte, that is,<maths num="6"><img file="JP5111583B2_D0006.tif" /></maths>Normalize the conversion so that
Then the normalized eigenvector
<maths num="7"><img file="JP5111583B2_D0007.tif" /></maths>The eigenvalue λ<sub>i</sub>Sort by (step 116). Next, λ<sub>1</sub>> λ<sub>2</sub>> λ<sub>3</sub>Create an optimized color transformation matrix assuming> 0 (step 118). This matrix is defined as follows.
<maths num="8"><img file="JP5111583B2_D0008.tif" /></maths>
Note that the eigenvalues are the variance of the new color components. A diagram showing an example of PCA matrix calculation for a training set of HE-stained pathological images is shown in FIG. A diagram showing a PCA of an example of a 2D set multivariate Gaussian distribution is shown in FIG.
In order to show the principle of the present invention, the color conversion by the optimized dyeing by H & E dyeing is as follows.
<maths num="9"><img file="JP5111583B2_D0009.tif" /></maths>
For comparison, the standard YCbCr transformation matrix is as follows.
<maths num="10"><img file="JP5111583B2_D0010.tif" /></maths>
The optimized color conversion can be used to compress the input image after being pre-calculated and stored in the database. A flow chart showing an example of the received image compression method is shown in FIG. Each time a new scanned slide is received (step 120), information about the tissue species and the staining method applied when the specimen was processed (eg, read from the specimen tag or label) (step 122). .. Next, the calculated color conversion database 36 (Fig. 1) is used to search for the color conversion (ie, matrix) corresponding to the staining type of the specimen slide. If a color conversion is found, the calculated calculated color conversion coefficients optimized for this staining method are read from the database. In one embodiment, the optimized color conversion will vary depending on the location of the laboratory performing the process.
Once the input image is mapped to some sort of stain (such parameters are known at the time of image acquisition), the pre-processed step of the compression algorithm is the pre-calculated corresponding color PCA matrix for that particular stain. Is captured and applied to the input image. As shown below, the pixel data of the received image is transformed using the optimized color transformation matrix.
<maths num="11"><img file="JP5111583B2_D0011.tif" /></maths>
Optimized image compression is then performed using the preprocessed input image (step 126). The resulting compressed image is stored in the image file storage (step 128).
Note that for each optimized color transformation by dyeing, the selection of the quantization step of the transformation coefficients is based on the characteristics of the color transformation matrix, such as the expected variance of the color channel values. The quantization step for the time-frequency conversion step of the image compression process is based on the fact that the energy (ie, information) in the first PCA principal component is greater than or equal to the energy in the second PCA component, which is greater than or equal to the third PCA component. Is judged. Therefore, for best velocity strain performance, select the minimum quantization step for the first PCA component, the larger quantization step for the second PCA component, and the third PCA component. Select the maximum quantization step.
Therefore, the finest quantization step is used to quantize the first PCA component and the coarsest step is used to quantize the third PCA component. This significantly improves the overall speed distortion performance of the image compression process.
To show the advantages of the present invention, the results of the YCbCr conversion and PCA conversion speed strain performance of the examples are shown in Table 1 of FIG. This table incorporates the results of experiments performed by us using the known Kakadu® JPEG2000 toolkit, which has YCbCr color conversion and color conversion by dyeing. In this experiment, two relatively large H & E digitally stained images were selected. As shown in the table, in both cases, that is, for two different bit rates or the same bit rate, the PSNR quality goes from 0.33 dB to 0.41 dB using the optimized image compression mechanism. It has improved significantly.
Corresponding structures, materials, actions, and equivalents of all means or steps and functional members within the claims are any that perform a function in combination with other specifically claimed members. It includes structure, material, or operation. The description of the present invention is presented for purposes of illustration and description, and is neither exhaustive nor limited to the disclosed forms of the invention. Many modified and modified embodiments will be apparent to those skilled in the art without departing from the scope and gist of one or more embodiments of the invention. Each embodiment best describes the principles and practical uses, and one of the present inventions relates to various embodiments, including various modified embodiments suitable for a particular use considered by a person skilled in the art. Selected and described so that one or more embodiments can be understood.
The appended claims cover all features and advantages of the invention, such as those within the spirit and scope of the invention. The present invention is not limited to the limited number of embodiments described herein, as many modified and modified embodiments are readily envisioned by those skilled in the art. Therefore, it will be appreciated that all suitable modified embodiments, modified embodiments, and equivalents within the spirit and scope of the invention can be claimed.
10, 60 systems 12 Observer Station 14 scanner 16 back end 18, 20 GUI 22 Streaming client 24 color management department 26 Laboratories Information System Broker 28 Image acquisition department 30 Optimized image compression module 32 DICOM library 34 compression library 36 Calculated color conversion database 38 Memory Management Module 40 database 42 storage 44 Streaming server 46 CAD / analysis block 62 processor 64 bus 66 Flash memory 68 Static read-only memory 70 dynamic main memory 72 Data I / O communication interface 74 Storage device 76 LAN / WAN / SAN 78 Display device 80 alphanumeric input device 82 Pointing device
32 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2006526357A | Cites | Japan |
| US20080075360A1 | Cites | United States of America |
| JP2005130308A | Cites | Japan |
| JP2005509140A | Cites | Japan |
6 members in 3 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 12570290 | United States of America | – | |
| 57029009 | United States of America | A | |
| 57029009 | United States of America | A | |
| 2009570290 | – | – | – |
| US20090570290 | – | – | – |
Members6
| Document | Office | Kind | |
|---|---|---|---|
| DE102010037855A1 | Germany | A1 | |
| US2011075897A1 | United States of America | A1 | |
| JP2011078096A | Japan | A | |
| DE102010037855A8 | Germany | A8 | |
| US8077959B2 | United States of America | B2 | |
| JP5111583B2This record | Japan | B2 |
19 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on accelerated examinationJAPANESE INTERMEDIATE CODE: A971005A975 | A975 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Explanation of circumstances concerning accelerated examinationJAPANESE INTERMEDIATE CODE: A871A871 | A871 |
Numbers
- Publication
- 5111583
- Publication, DOCDB
- 5111583
- Publication, EPODOC
- JP5111583B
- Application
- 217988
- Application, DOCDB
- 2010217988
- Application, EPODOC
- JP20100217988
Titles2
- Japanese
- デジタル病理スライドの染色による最適化圧縮
- English
- Optimized compression by staining digital pathology slides
Classification
- CPC, 5
- H04N19/136
- H04N19/63
- H04N19/60
- H04N19/186
- H04N19/85
- IPC, 1
- H04N1 41
