Video decoding system supporting multiple standards
Claim Score by NHIP
Abstract
System and method for decoding digital video data. The decoding system employs hardware accelerators that assist a core processor in performing selected decoding tasks. The hardware accelerators are configurable to support a plurality of existing and future encoding/decoding formats. The accelerators are configurable to support substantially any existing or future encoding/decoding formats that fall into the general class of DCT-based, entropy decoded, block-motion-compensated compression algorithms. The hardware accelerators illustratively comprise a programmable entropy decoder, an inverse quantization module, a inverse discrete cosine transform module, a pixel filter, a motion compensation module and a de-blocking filter. The hardware accelerators function in a decoding pipeline wherein at any given stage in the pipeline, while a given function is being performed on a given macroblock, the next macroblock in the data stream is being worked on by the previous function in the pipeline.

Term
Term ended
Expired 1 April 2022, 4.5 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
31 claims: 6 independent, 25 dependent
- 1A method of decoding a digital media data stream, comprising:(a) receiving media data of a first encoding/decoding format;(b) configuring at least one external decoding function based on the first encoding/decoding format;(c) decoding media data of the first encoding/decoding format using the at least one external decoding function;(d) receiving media data of a second encoding/decoding format;(e) configuring the at least one external decoding function based on the second encoding/decoding format;and (f) decoding media data of the second encoding/decoding format using the at least one external decoding function , wherein decoding media data using the at least one external decoding function in operations (c) and (f) comprises at least one configurable hardware module performing the at least one external decoding function, and wherein configuring the at least one external decoding function in operations (b) and (e) comprises configuring the at least one configurable hardware module, wherein the at least one configurable hardware module is a plurality of configurable hardware modules, and wherein each of the plurality of configurable hardware modules performs at least one decoding function, wherein at least one of the plurality of configurable hardware modules does not include a processor .
- 9Broadest claimClaim Score 82, broad(NHIP)A video decoding method, comprising:receiving a first video macroblock encoded in a first encoding format;configuring a first external decoding function based on the first encoding format;decoding the first video macroblock using the first external decoding function;receiving a second video macroblock encoded in a second encoding format;configuring a second external decoding function based on the second encoding format;and decoding the second video macroblock using the second external decoding function.
- 17A video decoding method, comprising:determining a data format for a video data stream;configuring a programmable entropy decoder to perform entropy decoding based on the determined data format;configuring an inverse quantizer to perform an inverse quantization based on the determined data format;configuring an inverse transform accelerator to perform an inverse transform operations based on the determined data format;configuring a pixel filter to perform a pixel filtering based on the determined data format;configuring a motion compensator to perform a motion compensation based on the determined data format;and configuring a de-blocking filter to perform a de-blocking operation based on the determined data format.
- 25A method of decoding a digital media data stream, comprising:(a) receiving media data of a first encoding/decoding format;(b) configuring at least one external decoding function based on the first encoding/decoding format;(c) decoding media data of the first encoding/decoding format using the at least one external decoding function;(d) receiving media data of a second encoding/decoding format;(e) configuring the at least one external decoding function based on the second encoding/decoding format;and (f) decoding media data of the second encoding/decoding format using the at least one external decoding function, wherein decoding media data using the at least one external decoding function in operations (c) and (f) comprises at least one configurable hardware module performing the at least one external decoding function, and wherein configuring the at least one external decoding function in operations (b) and (e) comprises configuring the at least one configurable hardware module, wherein the at least one configurable hardware module is a plurality of configurable hardware modules, and wherein each of the plurality of configurable hardware modules performs at least one decoding function, wherein none of the plurality of configurable hardware modules includes a processor.
- 27A method of decoding a digital media data stream, comprising:(a) receiving media data of a first encoding/decoding format;(b) configuring at least one external decoding function based on the first encoding/decoding format;(c) decoding media data of the first encoding/decoding format using the at least one external decoding function;(d) receiving media data of a second encoding/decoding format;(e) configuring the at least one external decoding function based on the second encoding/decoding format;and (f) decoding media data of the second encoding/decoding format using the at least one external decoding function, wherein decoding media data using the at least one external decoding function in operations (c) and (f) comprises at least one configurable hardware module performing the at least one external decoding function, and wherein configuring the at least one external decoding function in operations (b) and (e) comprises configuring the at least one configurable hardware module, wherein the at least one configurable hardware module is a plurality of configurable hardware modules, and wherein each of the plurality of configurable hardware modules performs at least one decoding function, wherein each of the configurable hardware modules is separate from others of the plurality of configurable hardware modules.
- 31A method of decoding a digital media data stream, comprising:(a) receiving media data of a first encoding/decoding format;(b) configuring at least one external decoding function based on the first encoding/decoding format;(c) decoding media data of the first encoding/decoding format using the at least one external decoding function;(d) receiving media data of a second encoding/decoding format;(e) configuring the at least one external decoding function based on the second encoding/decoding format;and (f) decoding media data of the second encoding/decoding format using the at least one external decoding function, wherein decoding media data using the at least one external decoding function in operations (c) and (f) comprises at least one configurable hardware module performing the at least one external decoding function, and wherein configuring the at least one external decoding function in operations (b) and (e) comprises configuring the at least one configurable hardware module, wherein the at least one configurable hardware module is a plurality of configurable hardware modules, and wherein each of the plurality of configurable hardware modules performs at least one decoding function, wherein each of the plurality of configurable hardware modules is independently controlled by a core decoding processor, wherein the core decoding processor independently controls each of the plurality of configurable hardware modules by programming a register for each of the plurality of configurable hardware modules.
Independent claims6
88 paragraphs in 8 sections, as filed
id="REI-00001" date="20211207"
CROSS-REFERENCE
id="REI-00001"
id="REI-00002" date="20211207"
CROSS REFERENCE
id="REI-00002"
TO RELATED APPLICATIONS
0001This application is a reissue of U.S. patent application Ser. No. 13/608,221 filed Sep. 10, 2012 (now U.S. Pat. No. 9,417,883) entitled, “Video Decoding System Supporting Multiple Standards,” which is a divisional application of and claims priority to U.S. patent application Ser. No. 10/114,798, filed on Apr. 1, 2002, having the title “VIDEO DECODING SYSTEM SUPPORTING MULTIPLE STANDARDS,” and issued as U.S. Pat. No. 8,284,844 on Oct. 9, 2012, which is incorporated by reference herein as if expressly set forth in its entirety.
FIELD OF THE INVENTION
0002The present invention relates generally to video decoding systems, and more particularly to a video decoding system supporting multiple standards.
BACKGROUND
0003Digital video decoders decode compressed digital data that represent video images in order to reconstruct the video images. A relatively wide variety of encoding /decoding algorithms and encoding/decoding standards presently exist, and many additional algorithms and standards are sure to be developed in the future. The various algorithms and standards produce compressed video bit streams of a variety of formats. Some existing public format standards include MPEG -1, MPEG-2 (SD/HD), MPEG-4, H.263, H.263+ and H.26LIJVT. Also, private standards have been developed by Microsoft Corporation (Windows Media), RealNetworks, Inc., Apple Computer, Inc. (QuickTime), and others. It would be desirable to have a multi-format decoding system that can accommodate a variety of encoded bit stream formats, including existing and future standards, and to do so in a cost-effective manner.
0004A highly optimized hardware architecture can be created to address a specific video decoding standard, but this kind of solution is typically limited to a single format. On the other hand, a fully software based solution is often flexible enough to handle any encoding format, but such solutions tend not to have adequate performance for real time operation with complex algorithms, and also the cost tends to be too high for high volume consumer products. Currently a common software based solution is to use a general-purpose processor running in a personal computer, or to use a similar processor in a slightly different system. Sometimes the general-purpose processor includes special instructions to accelerate digital signal processor (DSP) operations such as multiply-accumulate (MAC); these extensions are intimately tied to the particular internal processor architecture. For example, in one existing implementation, an Intel Pentium processor includes an MMX instruction set extension. Such a solution is limited in performance, despite very high clock rates, and does not lend itself to creating mass market, commercially attractive systems.
0005Others in the industry have addressed the problem of accommodating different encoding/decocting algorithms by designing special purpose DSPs in a variety of architectures. Some companies have implemented Very Long Instruction Word (VLIW) architectures more suitable to video processing and able to process several instructions in parallel. In. these cases, the processors are difficult to program when compared to a general-purpose processor. Despite the fact that the DSP and VLIW architectures are intended for high performance, they still tend not to have enough performance for the present purpose of real time decoding of complex video algorithms. In special cases, where the processors are dedicated for decoding compressed video, special processing accelerators are tightly coupled to the instruction pipeline and are part of the core of the main processor.
0006Yet others in the industry have addressed the problem of accommodating different encoding/decoding algorithms by simply providing, multiple instances of hardware, each dedicated to a single algorithm. This solution is inefficient and is not cost-effective.
0007Further limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art through comparison of such systems with the present invention as set forth in the remainder of the present application with reference to the drawings.
SUMMARY
0008One aspect of the present invention is directed to a digital media decoding system having a processor and a hardware accelerator. The processor is adapted to control a decoding process. The hardware accelerator is coupled to the processor and performs a decoding function on a digital media data stream. The accelerator is configurable to perform the decoding function according to a plurality of decoding methods.
0009Another aspect of the present invention is directed to a method of decoding a digital media data stream. Pursuant to the method, in a first stage, a first decoding function is performed on an i<sup>th </sup>data element of the data stream with a first decoding accelerator. In a second stage, after the first stage, a second decoding function is performed on the i<sup>th </sup>data element with a second decoding accelerator, while the first decoding function is performed on an i+1<sup>st </sup>data element in the data stream with the first decoding accelerator.
0010Another aspect of the present invention is directed to a method of decoding a digital video data stream. Pursuant to the method, in a first stage, entropy decoding is performed on an i<sup>th </sup>data element of the data stream. In a second stage, after the first stage, inverse quantization is performed on a product of the entropy decoding of the i<sup>th </sup>data element, while entropy decoding is performed on an i+1<sup>st </sup>data element in the data stream.
0011Still another aspect of the present invention is directed to a method of decoding a digital media data stream. Pursuant to this method, media data of a first encoding/decoding format is received. At least one external decoding function is configured based on the first encoding/decoding format. Media data of the first encoding/decoding format is decoded using the at least one external decoding function. Media data of a second encoding/decoding format is received. The at least one external decoding function is configured based on the second encoding/decoding format. Then media data of the second encoding/decoding format is decoded using the at least one external decoding function.
0012It is understood that other embodiments of the present invention will become readily apparent to those skilled in the art from the following detailed description, wherein embodiments of the invention are shown and described only by way of illustration of the best modes contemplated for carrying out the invention. As will be realized, the invention is capable of other and different embodiments and its several details are capable of modification in various other respects, all without departing from the spirit and scope of the present invention. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not as restrictive.
BRIEF DESCRIPTION OF THE DRAWINGS
0013These and other features, aspects, and advantages of the present invention will become better understood with regard to the following description, appended claims, and accompanying drawings where:
0014<figref idref="DRAWINGS">FIG. 1</figref> is a functional block diagram of a digital media system in which the present invention may be illustratively employed.
0015<figref idref="DRAWINGS">FIG. 2</figref> is a functional block diagram demonstrating a video decode data flow according to an illustrative embodiment of the present invention.
0016<figref idref="DRAWINGS">FIG. 3</figref> is a high-level functional block diagram of a digital video decoding system according to an illustrative embodiment of the present invention.
0017<figref idref="DRAWINGS">FIG. 4a</figref> is a functional block diagram of a digital video decoding system according to an illustrative embodiment of the present invention.
0018<figref idref="DRAWINGS">FIG. 4b</figref> is a functional block diagram of a motion compensation filter engine according to an illustrative embodiment of the present invention.
0019<figref idref="DRAWINGS">FIG. 5</figref> is a block diagram depicting a clocking scheme for a decoding system according to an illustrative embodiment of the present invention.
0020<figref idref="DRAWINGS">FIG. 6</figref> is a chart representing a decoding pipeline according to an illustrative embodiment of the present invention.
0021<figref idref="DRAWINGS">FIG. 7</figref> is a flowchart representing a macroblock decoding loop according to an illustrative embodiment of the present invention.
0022<figref idref="DRAWINGS">FIG. 8</figref> is a flowchart representing a method of decoding a digital video data stream containing more than one video data format, according to an illustrative embodiment of the present invention.
DETAILED DESCRIPTION OF THE EMBODIMENTS
0023The present invention forms an integral part of a complete digital media system and provides flexible and programmable decoding resources. <figref idref="DRAWINGS">FIG. 1</figref> is a functional block diagram of a digital media system in which the present invention may be illustratively employed. It will be noted, however, that the present invention can be employed in systems of widely varying architectures and widely varying designs.
0024The digital media system of <figref idref="DRAWINGS">FIG. 1</figref> includes transport processor <b>102</b>, audio decoder <b>104</b>, direct memory access (DMA) controller <b>106</b>, system memory controller <b>108</b>, system memory <b>110</b>, host CPU interface <b>112</b>, host CPU <b>114</b>, digital video decoder <b>116</b>, display feeder <b>118</b>, display engine <b>120</b>, graphics engine <b>122</b>, display encoders <b>124</b> and analog video decoder <b>126</b>. The transport processor <b>102</b> receives and processes a digital media data stream. The transport processor <b>102</b> provides the audio portion of the data stream to the audio decoder <b>104</b> and provides the video portion of the data stream to the digital video decoder <b>116</b>. In one embodiment, the audio and video data is stored in main memory <b>110</b> prior to being provided to the audio decoder <b>104</b> and the digital video decoder <b>116</b>. The audio decoder <b>104</b> receives the audio data stream and produces a decoded audio signal. DMA controller <b>106</b> controls data transfer amongst main memory <b>110</b> and memory units contained in elements such as the audio decoder <b>104</b> and the digital video decoder <b>116</b>. The system memory controller <b>108</b> controls data transfer to and from system memory <b>110</b>. In an illustrative embodiment, system memory <b>110</b> is a dynamic random access memory (DRAM) unit. The digital video decoder <b>116</b> receives the video data stream, decodes the video data and provides the decoded data to the display engine <b>120</b> via the display feeder <b>118</b>. The analog video decoder <b>126</b> digitizes and decodes an analog video signal (NTSC or PAL) and provides the decoded data to the display engine <b>120</b>. The graphics engine <b>122</b> processes graphics data in the data stream and provides the processed graphics data to the display engine <b>120</b>. The display engine <b>120</b> prepares decoded video and graphics data for display and provides the data to display encoders <b>124</b>, which provide an encoded video signal to a display device.
0025<figref idref="DRAWINGS">FIG. 2</figref> is a functional block diagram demonstrating a video decode data flow according to an illustrative embodiment of the present invention. Transport streams are parsed by the transport processor <b>102</b> and written to main memory <b>110</b> along with access index tables. The video decoder <b>116</b> retrieves the compressed video data for decoding, and the resulting decoded frames are written back to main memory <b>110</b>. Decoded frames are accessed by the display feeder interface <b>118</b> of the video decoder for proper display by a display unit. In <figref idref="DRAWINGS">FIG. 2</figref>, two video streams are shown flowing to the display engine <b>120</b>, suggesting that, in an illustrative embodiment, the architecture allows multiple display streams by means of multiple display feeders.
0026Aspects of the present invention relate to the architecture of digital video decoder <b>116</b>. In accordance with the present invention, a moderately capable general purpose CPU with widely available development tools is used to decode a variety of coded streams using hardware accelerators designed as integral parts of the decoding process.
0027Specifically, the most widely-used compressed video formats fall into a general class of DCT-based, variable-length coded, block-motion-compensated compression algorithms. As mentioned above, these types of algorithms encompass a wide class of international, public and private standards, including MPEG-1, MPEG-2 (SD/HD), MPEG-4, H.263, H.263-F, H.26LINT, Microsoft Corp, Real Networks, QuickTime, and others. Fundamental functions exist that are common to most or all of these formats. Such functions include, for example, programmable variable-length decoding (VLD), arithmetic decoding (AC), inverse quantization (IQ), inverse discrete cosine transform (IDCT), pixel filtering (PF), motion compensation (MC), and deblocking/de-ringing (loop filtering or post-processing). The term “entropy decoding” may be used generically to refer to variable length decoding, arithmetic decoding, or variations on either of these. According to the present invention, these functions are accelerated by hardware accelerators.
0028However, each of the algorithms mentioned above implement some or all of these functions in different ways that prevent fixed hardware implementations from addressing all requirements without duplication of resources. In accordance with one aspect of the present invention, these hardware modules are provided with sufficient flexibility or programmability enabling a decoding system that decodes a variety of standards efficiently and flexibly.
0029The decoding system of the present invention employs high-level granularity acceleration with internal programmability or configurability to achieve the requirements above by implementation of very fundamental processing structures that can be configured dynamically by the core decoder processor. This contrasts with a system employing fine-granularity acceleration, such as multiply-accumulate (MAC), adders, multipliers, FFT functions, DCT functions, etc. In a fine-granularity acceleration system, the decompression algorithm has to be implemented with firmware that uses individual low-level instructions (such as MAC) to implement a high-level function, and each instruction runs on the core processor. In the high-level granularity system of the present invention, the firmware configures each hardware accelerator, which in turn represent high-level functions (such as motion compensation) that run (using a well-defined specification of input data) without intervention from the main core processor. Therefore, each hardware accelerator runs in parallel according to a processing pipeline dictated by the firmware in the core processor. Upon completion of the high-level functions, each accelerator notifies the main core processor, which in turn decides what the next processing pipeline step should be.
0030The software control typically consists of a simple pipeline that orchestrates decoding by issuing commands to each hardware accelerator module for each pipeline stage, and a status reporting mechanism that makes sure that all modules have completed their pipeline tasks before issuing the start of the next pipeline stage.
0031<figref idref="DRAWINGS">FIG. 3</figref> is a high-level functional block diagram of a digital video decoding system <b>300</b> according to an illustrative embodiment of the present invention. The digital video decoding system <b>300</b> of <figref idref="DRAWINGS">FIG. 3</figref> can illustratively be employed to implement the digital video decoder <b>116</b> of <figref idref="DRAWINGS">FIGS. 1 and 2</figref>. The core processor <b>302</b> is the central control unit of the decoding system <b>300</b>. The core processor <b>302</b> prepares the data for decoding. The core processor <b>302</b> also orchestrates the macroblock (MB) processing pipeline for all modules and fetches the required data from main memory via the bridge <b>304</b>. The core processor <b>302</b> also handles some data processing tasks. Picture level processing, including sequence headers, GOP headers, picture headers, time stamps, macroblock-level information except the block coefficients, and buffer management, are performed directly and sequentially by the core processor <b>302</b>, without using the accelerators <b>304</b>, <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b> and <b>314</b> other than the PVLD <b>306</b> (which accelerates general bitstream parsing). Picture level processing does not overlap with slice level/macroblock decoding in this embodiment.
0032Programmable variable length decoder (PVLD) <b>306</b>, inverse quantizer <b>308</b>, inverse transform module <b>309</b>, pixel filter <b>310</b>, motion compensation module <b>312</b> and loop/post filter <b>314</b> are hardware accelerators that accelerate special decoding tasks that would otherwise be bottlenecks for real-time video decoding if these tasks were handled by the core processor <b>302</b> alone. Each hardware module <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b> and <b>314</b> is internally configurable or programmable to allow changes according to various processing algorithms. In an alternative embodiment, modules <b>308</b> and <b>309</b> are implemented in the form of a transform engine <b>307</b> that handles all functionality, but which is conceptually equivalent to the union of <b>308</b> and <b>309</b>. In a further alternative embodiment, modules <b>310</b> and <b>312</b> are implemented in the form of a filter engine <b>311</b> which consists of an internal SIMD (single instruction multiple data) processor and a general purpose controller to interface to the rest of the system, but which is conceptually equivalent to the union of <b>310</b> and <b>312</b>. In a further alternative embodiment, module <b>314</b> is implemented in the form of another filter engine similar to <b>311</b> which consists of an internal SIMD (single instruction multiple data) processor and a general purpose controller to interface to the rest of the system, but which is conceptually equivalent to <b>314</b>. In a further alternative embodiment, module <b>314</b> is implemented in the form of the same filter engine <b>311</b> that can also implement the equivalent function of the combination of <b>310</b> and <b>311</b>. Each hardware module <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b> and <b>314</b> performs its task after being so instructed by the core processor <b>302</b>. In-an illustrative embodiment of the present invention, each hardware module includes a status register that indicates whether the module has completed its assigned tasks. The ore processor <b>302</b> polls the status register to determine whether the hardware module has completed its task. In an alternative embodiment, the hardware accelerators share a status register.
0033In an illustrative embodiment, the PVLD engine <b>306</b> performs variable-length code (VLD) decoding of the block DCT coefficients. It also helps the core processor <b>302</b> to decode the header information in the compressed bitstream. In an illustrative embodiment of the present invention, the PVLD module <b>306</b> is designed as a coprocessor to the core processor <b>302</b>, while the rest of the modules <b>308</b>,<b>309</b>,<b>310</b>,<b>312</b> and <b>314</b> are designed as hardware accelerators. Also, in an illustrative embodiment, the PVLD module <b>306</b> includes two variable length decoders. Each of the two programmable variable-length decoders can be hardwired to efficiently perform decoding according to a particular video compression standard, such as MPEG2 HD. One of them can be optionally set as a programmable VLD engine, with a code RAM to hold VLC tables for media coding formats other than MPEG2. The two VLD engines are controlled independently by the core processor <b>302</b>, and either one or both of them will be employed at any given time, depending on the application.
0034The IQ engine <b>308</b> performs run-level pair decoding, inverse scan and quantization. The inverse transform engine <b>309</b> performs IDCT operations or other inverse transform operations like the Integer Transform of the H.26x standards. In an illustrative embodiment of the present invention, the IQ module <b>308</b> and the inverse transform module <b>309</b> are part of a common hardware module and use a similar interface to the core processor <b>302</b>.
0035The pixel filter <b>310</b> performs pixel filtering and interpolation. The motion compensation module <b>312</b> performs motion compensation. The pixel filter <b>310</b> and motion compensation module <b>312</b> are shown as one module in the diagram to emphasize a certain degree of direct cooperation between them. In an illustrative embodiment of the present invention, the PF module <b>310</b> and the MC module <b>312</b> are part of a common programmable module <b>311</b> designated as a filter engine capable of performing internal SIMD instructions to process data in parallel with an internal control processor.
0036The filter module <b>314</b> performs the de-blocking operation common in many low bit-rate coding standards. In one embodiment of the present invention, the filter module comprises a loop filter that performs de-blocking within the decoding loop. In another embodiment, the filter module comprises a post filter that performs de-blocking outside the decoding loop. In yet another embodiment, the filter module comprises a de-ringing filter, which may function as either a loop filter or a post filter, depending on the standard of the video being processed. In yet another embodiment, the filter module <b>314</b> includes both a loop filter and a post filter. Furthermore, in yet another embodiment, the filter module <b>314</b> is implemented using the same filter engine <b>311</b> implementation as for <b>310</b> and <b>312</b>, except that module <b>311</b> is programmed to produce deblocked or deringed data as the case may be.
0037The bridge module <b>304</b> arbitrates and moves picture data between decoder memory <b>316</b> and main memory. The bridge interface <b>304</b> includes an internal bus network that includes arbiters and a direct memory access (DMA) engine. The bridge <b>304</b> serves as an interface to the system buses.
0038In an illustrative embodiment of the present invention, the display feeder module <b>318</b> reads decoded frames from main memory and manages the horizontal scaling and displaying of picture data. The display feeder <b>318</b> interfaces directly to a display module. In an illustrative embodiment, the display feeder <b>318</b> converts from <b>420</b> to <b>422</b> color space. Also, in an illustrative embodiment, the display feeder <b>318</b> includes multiple feeder interfaces, each including its own independent color space converter and horizontal scaler. The display feeder <b>318</b> handles its own memory requests via the bridge module <b>304</b>.
0039Decoder memory <b>316</b> is used to store macroblock data and other time-critical data used during the decode process. Each hardware block <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b>, <b>314</b> accesses decoder memory <b>316</b> to either read the data to be processed or write processed data back. In an illustrative embodiment of the present invention, all currently used data is stored in decoder memory <b>316</b> to minimize accesses to main memory. Each hardware module <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b>, <b>314</b> is assigned one or more buffers in decoder memory <b>316</b> for data processing. Each module accesses the data in decoder memory <b>316</b> as the macro blocks are processed through the system. In an exemplary embodiment, decoder memory <b>316</b> also includes parameter buffers that are adapted to hold parameters that are needed by the hardware modules to do their job at a later macroblock pipeline stage. The buffer addresses are passed to the hardware modules by the core processor <b>302</b>. In an illustrative embodiment, decoder memory <b>316</b> is a static random access memory (SRAM) unit.
0040<figref idref="DRAWINGS">FIG. 4a</figref> is a functional block diagram of digital video decoding system <b>300</b> according to an illustrative embodiment of the present invention. In <figref idref="DRAWINGS">FIG. 4a</figref>, elements that are common to <figref idref="DRAWINGS">FIG. 3</figref> are given like reference numbers. In <figref idref="DRAWINGS">FIG. 4a</figref>, various elements are grouped together to illustrate a particular embodiment where <b>308</b> and <b>309</b> form part of a transform engine <b>307</b>, <b>310</b> and <b>312</b> form part of a filter engine <b>311</b> that is a programmable module that implements the functionality of PF and MC, <b>313</b> and <b>315</b> form part of another filter engine <b>314</b> which is another instance of the same programmable module except that it is programmed to implement the functionality of a loop filter <b>313</b> and a post filter <b>315</b>. In addition to the elements shown in <figref idref="DRAWINGS">FIG. 3</figref>, <figref idref="DRAWINGS">FIG. 4a</figref> shows, phase-locked loop (PLL) element <b>320</b>, internal data bus <b>322</b>, register bus <b>324</b> and separate loop and post filters <b>313</b> and <b>315</b> embodied in a filter engine module <b>314</b> which implements the functionality of <b>313</b> and <b>315</b>.
0041The core processor <b>302</b> is the master of the decoding system <b>300</b>. It controls the data flow of decoding processing. All video decode processing, except where otherwise noted, is performed in the core processor. The PVLD <b>306</b>, IQ <b>308</b>, inverse transform <b>309</b>, PF <b>310</b> and MC <b>312</b>, and filter <b>314</b> are hardware accelerators to help the core processor achieve the required performance. In an illustrative embodiment of the present invention, the core processor <b>302</b> is a MIPS processor, such as a MIPS32 implementation, for example. The core processor <b>302</b> incorporates a D cache and an I cache. The cache sizes are chosen to ensure that time critical operations are not impacted by cache misses. For example, instructions for macroblock-level processing of MPEG-2 video runs from cache. For other algorithms, time-critical code and data also reside in cache. The determination of exactly which functions are stored in cache involves a trade-off between cache size, main memory access time, and the degree of certainty of the firmware implementation for the various algorithms. The cache behavior with proprietary algorithms depends in part in the specific software design. In an illustrative embodiment, the cache sizes are 16 kB for instructions and 4 kB for data. These can be readily expanded if necessary.
0042At the macroblock level, the core processor <b>302</b> interprets the decoded bits for the appropriate headers and decides and coordinates the actions of the hardware blocks <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b> and <b>314</b>. Specifically, all macroblock header information, from the macroblock address increment (MBAinc) to motion vectors (MV s) and to the cbp pattern in the case of MPEG2 decoding, for example, is derived by the core processor <b>302</b>. The core processor <b>302</b> stores related information in a particular format or data structure (determined by the hardware module specifications) in the appropriate buffers in the decoder memory <b>316</b>. For example, the quantization scale is passed to the buffer the IQ engine <b>308</b>; macroblock type, motion type and pixel precision are stored in the parameter buffer for the pixel filter engine <b>310</b>. The core processor keeps track of certain information in order to maintain the correct pipeline, and it may store some such information in its D cache, some in main system memory and some in the decoder memory <b>316</b>, as required by the specific algorithm being performed. For example, for some standards, motion vectors of the macroblock are kept as the predictors for future motion vector derivation.
0043In an illustrative embodiment the programmable variable length decoder <b>306</b> performs decoding of variable length codes (VLC) in the compressed bit stream to extract values, such as DCT coefficients, from the compressed data stream. Different coding formats generally have their own unique VLC tables. The PVLD <b>306</b> is completely configurable in terms of the VLC tables it can process. The PVLD <b>306</b> can accommodate a dynamically changing set of VLC tables, for example they may change on a macroblock-to-macroblock basis. In an illustrative embodiment of the present invention, the PVLD <b>306</b> includes a register that the core processor can program to guide the PVLD <b>306</b> to search for the VLC table of the appropriate encoding/decoding algorithm. The PVLD <b>306</b> decodes variable length codes in as little as one clock, depending on the specific code table in use and the specific code being decoded.
0044The PVLD <b>306</b> is designed to support the worst-case requirement for VLD operation with MPEG-2 HDTV (MP@HL), while retaining its full programmability. The PVLD <b>306</b> includes a code table random access memory (RAM) for fastest performance. Code tables such a MPEG-2 video can fit entirely within the code RAM. Some formats, such as proprietary formats, may require larger code tables that do not fit entirely within the code RAM in the PVLD <b>306</b>. For such cases, the PVLD <b>306</b> can make use of both the decoder memory <b>316</b> and the main memory as needed. Performance of VLC decoding is reduced somewhat when codes are searched in video memory <b>316</b> and main memory. Therefore, for formats that require large tables of VLC codes, the most common codes are typically stored in the PVLD code RAM, the next most common codes are stored in decoder memory, and the least common codes are stored in main memory. Also, such codes are stored in decoder memory <b>316</b> and main memory such that even when extended look-ups in decoder memory <b>316</b> and main memory are required, the most commonly occurring codes are found more quickly. This allows the overall performance to remain exceptionally high.
0045In an illustrative embodiment of the present invention, the PVLD <b>306</b> is architected as a coprocessor of the core processor <b>302</b>. That is, it can operate on a single-command basis where the core processor issues a command (via a coprocessor instruction) and waits (via a Move From Coprocessor instruction) until it is executed by the PVLD <b>306</b>, without polling to determine completion of the command. This increases performance when a large number of VLC codes are parsed under software control. Additionally, the PVLD <b>306</b> can operate on a block-command basis where the core processor <b>302</b> commands the PVLD <b>306</b> to decode a complete block of VLC codes, such as DCT coefficients, and the core processor <b>302</b> continues to perform other tasks in parallel. In this case, the core processor <b>302</b> verifies the completion of the block operation by checking a status bit in the PVLD <b>306</b>. The PVLD produces results (tokens) that are stored in decoder memory <b>316</b>.
0046The PVLD <b>306</b> checks for invalid codes and recovers gracefully from them. Invalid codes may occur in the coded bit stream for a variety of reasons, including errors in the video encoding, errors in transmission, and improper discontinuities in the stream.
0047The inverse quantizer module <b>308</b> performs run-level code (RLC) decoding, inverse scanning (also called zig-zag scanning), inverse quantization and mismatch control. The coefficients, such as DCT coefficients, extracted by the PVLD <b>306</b> are processed by the inverse quantizer <b>308</b> to bring the coefficients from the quantized domain to the DCT domain. In an exemplary embodiment of the present invention, the IQ module <b>308</b> obtains its input data (run-level values) from the decoder memory <b>316</b>, as the result of the PVLD module <b>306</b> decoding operation. In an alternative embodiment, the IQ module <b>308</b> obtains its input data directly from the PVLD <b>306</b>. This alternative embodiment is illustratively employed in conjunction with encoding/decoding algorithms that are relatively more involved, such as MPEG-2 HD decoding, for best performance. The run-length, value and end-of-block codes read by the IQ module <b>308</b> are compatible with the format created by the PVLD module when it decodes blocks of coefficient VLCs, and this format is not dependent on the specific video coding format being decoded. In an exemplary embodiment, the IQ <b>308</b> and inverse transform <b>309</b> modules form part of a tightly coupled module labeled transform engine <b>307</b>. This embodiment has the advantage of providing fast communication between modules <b>308</b> and <b>309</b> by virtue of being implemented in the same hardware block.
0048The scan pattern of the IQ module <b>308</b> is programmable in order to be compatible with any required pattern. The quantization format is also programmable, and mismatch control supports a variety of methods, including those specified in MPEG-2 and MPEG-4. In an exemplary embodiment, the IQ module <b>308</b> can accommodate block sizes of 16×16, 8×8, 8×4, 4×8 and 4×4. In an illustrative embodiment of the present invention, the IQ module <b>308</b> includes one or more registers that are used to program the scan pattern, quantization matrix and mismatch control method. These registers are programmed by the core processor <b>302</b> to dictate the mode of operation of the IQ module. The IQ module <b>306</b> is designed in such a way that the core processor <b>302</b> can intervene at any point in the process, in case a particular decoding algorithm requires software processing of some aspect of the algorithmic steps performed by the IQ module <b>308</b>. For example, there may be cases where an unknown algorithm could require a different form of rounding; this can be performed in the core processor <b>302</b>. The IQ module <b>308</b> has specific support for AC prediction as specified in MPEG-4 Advanced Simple Profile. In an exemplary embodiment, the IQ module <b>308</b> also has specific support for the inverse quantization functions of the ISO-ITU NT (Joint Video Team) standard under development.
0049The inverse transform module <b>309</b> performs the inverse transform to convert the coefficients produced by the IQ module <b>308</b> from the frequency domain to the spatial domain. The primary transform supported is the IDCT, as specified in MPEG-2, MPEG-4, IEEE, and several other standards. The coefficients are programmable, and it can support alternative related transforms, such as the “linear” transform in H.26L (also known as JVT), which is not quite the same as IDCT. The inverse transform module <b>309</b> supports a plurality of matrix sizes, including 8×8, 4×8, 8×4 and 4×4 blocks. In an illustrative embodiment of the present invention, the inverse transform module <b>309</b> includes a register that is used to program the matrix size. This register is programmed by the core processor <b>302</b> according to the appropriate matrix size for the encoding/decoding format of the data stream being decoded.
0050In an illustrative embodiment of the present invention, the coefficient input to the inverse transform module <b>309</b> is read from decoder memory <b>316</b>, where it was placed after inverse quantization by the IQ module <b>308</b>. The transform result is written back to decoder memory <b>316</b>. In an exemplary embodiment, the inverse transform module <b>309</b> uses the same memory location in decoder memory <b>316</b> for both its input and output, allowing a savings in on-chip memory usage. In an alternative embodiment, the coefficients produced by the IQ module are provided directly to the inverse transform module <b>309</b>, without first depositing them in decoder memory <b>316</b>. To accommodate this direct transfer of coefficients, in one embodiment of the present invention, the IQ module <b>308</b> and inverse transform module <b>309</b> use a common interface directly between them for this purpose. In an exemplary embodiment, the transfer of coefficients from the IQ module <b>308</b> to the inverse transform module <b>309</b> can be either direct or via decoder memory <b>316</b>. For encoding/decoding algorithms that require very high rates of throughput, such as MPEG-2 HD decoding, the transfer is direct in order to save time and improve performance.
0051In an illustrative embodiment, the functionality of the PF <b>310</b> and MC <b>312</b> are implemented by means of a filter engine (FE) <b>311</b>. The FE is the combination of an 8-way SIMD processor <b>2002</b> and a 32-bit RISC processor <b>2004</b>, illustrated in <figref idref="DRAWINGS">FIG. 4b</figref>. Both processors operate at the same clock frequency. The SIMD engine <b>2002</b> is architected to be very efficient as a coprocessor to the RISC processor (internal MIPS) <b>2004</b>, performing specialized filtering and decision-making tasks. The SIMD <b>2002</b> includes: a split X-memory <b>2006</b> (allowing simultaneous operations), a Y-memory, a Z-register input with byte shift capability, 16 bit per element inputs, and no branch or jump functions. The SIMD processor <b>2002</b> has hardware for three-level looping, and it has a hardware function call and return mechanism for use as a coprocessor. All of these help to improve performance and minimize the area. The RISC processor <b>2004</b> controls the operations of the FE <b>311</b>. Its functions include the control of the data flow and scheduling tasks. It also takes care of part of the decision-making functions. The FE <b>311</b> operates like the other modules on a macro block basis under the control of the mum core processor <b>302</b>.
0052Referring again to <figref idref="DRAWINGS">FIG. 4a</figref>, the pixel filter <b>310</b> performs pixel filtering and interpolation as part of the motion compensation process. Motion compensation uses a small piece of an image from a previous frame to predict a piece of the current image; typically the reference image segment is in a different location within the reference frame. Rather than recreate the image anew from scratch, the previous image is used and the appropriate region of the image moved to the proper location within the frame; this may represent the image accurately, or more generally there may still be a need for coding the residual difference between this prediction and the actual current image. The new location is indicated by motion vectors that denote the spatial displacement in the frame with respect to the reference frame.
0053The pixel filter <b>310</b> performs the interpolation necessary when a reference block is translated (motion-compensated) by a vector that cannot be represented by an integer number of whole-pixel locations. For example, a hypothetical motion vector may indicate to move a particular block 10.5 pixels to the right and 0.25 pixels down for the motion-compensated prediction. In an illustrative embodiment of the present invention, the motion vectors are decoded by the PVLD 3D6 in a previous processing pipeline stage and are further processed in the core processor <b>302</b> before being passed to the pixel filter, typically via the decoder memory <b>316</b>. Thus, the pixel filter <b>310</b> gets the motion information as vectors and not just bits from the bitstream. In an illustrative embodiment, the reference block data that is used by the motion compensation process is read by the pixel filter <b>310</b> from the decoder memory <b>316</b>, the required data having been moved to decoder memory <b>316</b> from system memory <b>110</b>; alternatively the pixel filter obtains the reference block data from system memory <b>110</b>. Typically the pixel filter obtains the processed motion vectors from decode memory <b>316</b>. The pixel data that results from motion compensation of a given macroblock is stored in memory after decoding of said macroblock is complete. In an illustrative embodiment, the decoded macroblock data is written to decoder memory <b>316</b> and then transferred to system memory <b>110</b>; alternatively, the decoded macro block data may be written directly to system memory <b>110</b>. If and when that decoded macroblock data is needed for additional motion compensation of another macroblock, the pixel filter <b>310</b> retrieves the reference macroblock pixel information from memory, as above, and again the reconstructed macroblock pixel information is written to memory, as above.
0054The pixel filter <b>310</b> supports a variety of filter algorithms, including ½ pixel and ¼ pixel interpolations in either or both of the horizontal and vertical axes; each of these can have many various definitions, and the pixel filter can be configured or programmed to support a wide variety of filters, thereby supporting a wide range of video formats, including proprietary formats. The PF module can process block sizes of 4, 8 or 16 pixels per dimension (horizontal and vertical), or even other sizes if needed. The pixel filter <b>310</b> is also programmable to support different interpolation algorithms with different numbers of filter taps, such as 2, 4, or 6 taps per filter, per dimension. In an illustrative embodiment of the present invention, the pixel filter <b>309</b> includes one or more registers that are used to program the filter algorithm and the block size. These registers are programmed by the core processor <b>302</b> according to the motion compensation technique employed with the encoding/decoding format of the data stream being decoded. In another illustrative embodiment, the pixel filter is implemented using the filter engine (FE) architecture, which is programmable to support any of a wide variety of filter algorithms. As such, in either type of embodiment, it supports a very wide variety of motion compensation schemes.
0055The motion compensation module <b>312</b> reconstructs the macroblock being decoded by performing the addition of the decoded difference (or residual or “error”) pixel information from the inverse transform module <b>309</b> to the pixel prediction data from the output of the pixel filter <b>310</b>. The motion compensation module <b>312</b> is programmable to support a wide variety of block sizes, including 16×16, 16×8, 8×16, 8×8, 8×4, 4×8 and 4×4. The motion compensation module <b>312</b> is also programmable to support different transform block types, such as field-type and frame-type transform blocks. The motion compensation module <b>312</b> is further programmable to support different matrix formats. Furthermore, MC module <b>312</b> supports all the intra and inter prediction modes in the H.26L/JVT proposed standard. In an illustrative embodiment of the present invention, the motion compensation module <b>312</b> includes one or more registers that are configurable to select the block size and format. These registers are programmed by the core processor <b>302</b> according to the motion compensation technique employed with the encoding/decoding format of the data stream being decoded. In another illustrative embodiment, the motion compensation module is a function of a filter engine (FE) that is serving as the pixel filter and motion compensation modules, and it is programmable to perform any of the motion compensation functions and variations that are required by the format being decoded.
0056The loop filter <b>313</b> and post filter <b>315</b> perform de-blocking filter operations. In an illustrative embodiment of the present invention, the loop filter <b>313</b> and post filter <b>315</b> are combined in one filter module <b>314</b>, as shown in <figref idref="DRAWINGS">FIG. 3</figref>. The filter module <b>314</b> in an illustrative embodiment is the same processing structure as described for <b>311</b>, except that it is programmed to perform the functionality of <b>313</b> and <b>315</b>. Some decoding algorithms employ a loop filter and others employ a post filter. Therefore, the filter module <b>314</b> (or loop filter <b>313</b> and post filter <b>315</b> independently) is programmable to turn on either the loop filter <b>313</b> or the post filter <b>315</b> or both. In an illustrative embodiment, the filter module <b>314</b> (or loop filter <b>313</b> and post filter <b>315</b>) has a register that controls whether a loop filter or post filter scheme is employed. The core processor <b>302</b> programs the filter module register(s) according to the bit-stream semantics. The loop filter <b>313</b> and post filter <b>315</b> each have programmable coefficients and thresholds for performing a variety of de-blocking algorithms in either the horizontal or vertical directions. Deblocking is required in some low bit-rate algorithms. De-blocking is not required in MPEG-2. However, in one embodiment of the present invention, de-blocking is used to advantage with MPEG-2 at low bit rates.
0057In one embodiment of the present invention, the input data to the loop filter <b>313</b> and post filter <b>315</b> comes from decoder memory <b>316</b>, the input pixel data having been transferred from system memory <b>110</b> as appropriate, typically at the direction of the core processor <b>302</b>. This data includes pixel and block/macroblock parameter data generated by other modules in the decoding system <b>300</b>. The output data from the loop filter <b>313</b> and post filter <b>315</b> is written into decoder memory <b>316</b>. The core processor <b>302</b> then causes the processed data to be put in its correct location in system memory <b>110</b>. The core processor <b>302</b> can program operational parameters into loop filter <b>313</b> and post filter <b>315</b> registers at any time. In an illustrative embodiment, all parameter registers are double buffered. In another illustrative embodiment the loop filter <b>313</b> and post filter <b>315</b> obtain input pixel data from system memory <b>110</b>, and the results may be written to system memory <b>110</b>.
0058The loop filter <b>313</b> and post filter <b>315</b> are both programmable to operate according to any of a plurality of different encoding/decoding algorithms. In the embodiment wherein loop filter <b>313</b> and post filter <b>315</b> are separate hardware units, the loop filter <b>313</b> and post filter <b>315</b> can be programmed similarly to one another. The difference is where in the processing pipeline each filter <b>313</b>, <b>315</b> does its work. The loop filter <b>313</b> processes data within the reconstruction loop and the results of the filter are used in the actual reconstruction of the data. The post filter <b>315</b> processes data that has already been reconstructed and is fully decoded in the two-dimensional picture domain. In an illustrative embodiment of the present invention, the coefficients, thresholds and other parameters employed by the loop filter <b>313</b> and the post filter <b>315</b> (or, in the alternative embodiment, filter module <b>314</b>) are programmed by the core processor <b>302</b> according to the de-blocking technique employed with the encoding/decoding format of the data stream being decoded.
0059The core processor <b>302</b>, bridge <b>304</b>, PVLD <b>306</b>, IQ <b>308</b>, inverse transform module <b>309</b>, pixel filter <b>310</b>, motion compensation module <b>312</b>, loop filter <b>313</b> and post filter <b>315</b> have access to decoder memory <b>316</b> via the internal bus <b>322</b> or via equivalent functionality in the bridge <b>304</b>. In an exemplary embodiment of the present invention, the PVLD <b>306</b>, IQ <b>308</b>, inverse transform module <b>309</b>, pixel filter <b>310</b>, motion compensation module <b>312</b>, loop filter <b>313</b> and post filter <b>315</b> use the decoder memory <b>316</b> as the source and destination memory for their normal operation. In another embodiment, the PL VD <b>306</b> uses the system memory <b>110</b> as the source of its data in normal operation. In another embodiment, the pixel filter <b>310</b> and motion compensation module <b>312</b>, or the equivalent function in the filter module <b>314</b>, use the decoder memory <b>316</b> as the source for residual pixel information and they use system memory <b>110</b> as the source for reference pixel data and as the destination for reconstructed pixel data. In another embodiment, the loop filter <b>313</b> and post processor <b>315</b>, or the equivalent function in the filter module <b>314</b>, use system memory <b>110</b> as the source and destination for pixel data in normal operation. The CPU has access to decoder memory <b>316</b>, and the DMA engine <b>304</b> can transfer data between decoder memory <b>316</b> and the main system memory <b>110</b>. The arbiter for decoder memory <b>316</b> is in the bridge module <b>304</b>. In an illustrative embodiment, decoder memory <b>316</b> is a static random access memory (SRAM) unit.
0060The bridge module <b>304</b> performs several functions. In an illustrative embodiment, the bridge module <b>304</b> includes an interconnection network to connect all the other modules of the MVP as shown schematically as internal bus <b>322</b> and register bus <b>324</b>. It is the bridge between the various modules of decoding system <b>300</b> and the system memory. It is the bridge between the register bus <b>324</b>, the core processor <b>302</b>, and the main chip-level register bus. It also includes a DMA engine to service the memories within the decoder system <b>300</b>, including decoder memory <b>316</b> and local memory units within individual modules such as PVLD <b>306</b>. The bridge module illustratively includes an asynchronous interface capability and it supports different clock rates in the decoding system <b>300</b> and the main memory bus, with either clock frequency being greater than the other.
0061The bridge module <b>304</b> implements a consistent interface to all of the modules of the decoding system <b>300</b> where practical. Logical register bus <b>324</b> connects all the modules and serves the purpose of accessing control and status registers by the main core processor <b>302</b>. Coordination of processing by the main core processor <b>302</b> is accomplished by a combination of accessing memory, control and status registers for all modules.
0062In an illustrative embodiment of the present invention, the display feeder <b>318</b> module reads decoded pictures (frames or fields, as appropriate) from main memory in their native decoded format (4:2:0, for example), converts the video into 4:2:2 format, and performs horizontal scaling using a polyphase filter. According to an illustrative embodiment of the present invention, the coefficients, scale factor, and the number of active phases of the polyphase filter are programmable. In an illustrative embodiment of the present invention, the display feeder <b>318</b> includes one or more registers that are used to program these parameters. These registers are programmed by the core processor <b>302</b> according to the desired display format. In an exemplary embodiment the polyphase filter is an 8 tap, 11 phase filter. The output is illustratively standard 4:2:2 format YCrCb video, in the native color space of the coded video (for example, ITU-T 709-2 or ITU-T 601-B color space), and with a horizontal size that ranges, for example, from 160 to 1920 pixels. The horizontal scaler corrects for coded picture sizes that differ from the display size, and it also provides the ability to scale the video to arbitrary smaller or larger sizes, for use in conjunction with subsequent 2-dimensional scaling where required for displaying video in a window, for example. In one embodiment, the display feeder <b>318</b> is adapted to supply two video scan lines concurrently, in which case the horizontal scaler in the feeder <b>318</b> is adapted to scale two lines concurrently, using identical parameters.
0063<figref idref="DRAWINGS">FIG. 5</figref> is a block diagram depicting a clocking scheme for decoding system <b>300</b> according to an illustrative embodiment of the present invention. In <figref idref="DRAWINGS">FIG. 5</figref>, elements that are common to <figref idref="DRAWINGS">FIGS. 3 and 4</figref> are given like reference numbers. Hardware accelerators block <b>330</b> includes PVLD <b>306</b>, IQ <b>308</b>, inverse transform module <b>309</b>, pixel filter <b>310</b>, motion compensation module <b>312</b> and filter engine <b>314</b>. In an illustrative embodiment of the present invention, the core processor <b>302</b> runs at twice the frequency of the other processing modules. In another related illustrative embodiment, hardware accelerator block <b>330</b> includes PVLD <b>306</b>, IQ <b>308</b>, and inverse transform module <b>309</b>, while one instance of the filter engine module <b>311</b> implements pixel filter <b>310</b> and motion compensation <b>312</b>, and yet another instance of the filter module <b>314</b> implements loop filter <b>313</b> and post filter <b>315</b>, noting that FE <b>311</b> and FE <b>314</b> receive both 243 MHz and 121.5 MHz clocks. In an exemplary embodiment, the core processor runs at 243 MHz and the individual modules at half this rate, i.e., 121.5 MHz. An elegant, flexible and efficient clock strategy is achieved by generating two internal clocks in an exact 2:1 relationship to each other. The system clock signal (CLK_IN) <b>332</b> is used as input to the phase locked loop element (PLL) <b>320</b>, which is a closed-loop feedback control system that locks to a particular phase of the system clock to produce a stable signal with little jitter. The PLL element <b>320</b> generates a IX clock (targeting, e.g., 121.5 MHz) for the hardware accelerators <b>330</b>, bridge <b>304</b> and the core processor bus interface <b>303</b>, while generating a 2X clock (targeting, e.g., 243 MHz) for the core processor <b>302</b> and the core processor bus interface <b>303</b>.
0064Referring again to <figref idref="DRAWINGS">FIGS. 3 and 4</figref>, for typical video formats such as MPEG-2, picture level processing, from the sequence level down to the slice level, including the sequence headers, picture headers, time stamps, and buffer management, are performed directly and sequentially by the core processor <b>302</b>. The PVLD <b>306</b> assists the core processor when a bit-field in a header is to be decoded. Picture level processing does not overlap with macroblock level decoding.
0065The macroblock level decoding is the main video decoding process. It occurs within a direct execution loop. In an illustrative embodiment of the present invention, hardware blocks PVLD <b>306</b>, IQ <b>308</b>, inverse transform module <b>309</b>, pixel filter <b>310</b>, motion compensation module <b>312</b> (and, depending on which decoding algorithm is being executed, possibly loop filter <b>313</b>) are all involved in the decoding loop. The core processor <b>302</b> controls the loop by polling the status of each of the hardware blocks involved.
0066Still another aspect of the present invention is directed to a method of decoding a digital media data stream. Pursuant to this method, media data of a first encoding/decoding format is received. At least one external decoding function, such as variable-length decoding or inverse quantization, e.g., is configured based on the first encoding/decoding format. Media data of the first encoding/decoding format is decoded using the at least one external decoding function. Media data of a second encoding/decoding, format is received. The at least one external decoding function is configured based on the second encoding/decoding format. Then media data of the second encoding/decoding format is decoded using the at least one external decoding function.
0067In an illustrative embodiment of the present invention, the actions of the various hardware blocks are arranged in an execution pipeline comprising a plurality of stages. As used in the present application, the term “stage” can refer to all of the decoding functions performed during a given time slot, or it can refer to a functional step, or group of functional steps, in the decoding process. The pipeline scheme aims to achieve maximum throughput in defined worst case decoding scenarios. Pursuant to this objective, it is important to utilize the core processor efficiently. <figref idref="DRAWINGS">FIG. 6</figref> is a chart representing a decoding pipeline according to an illustrative embodiment of the present invention. The number of decoding functions in the pipeline may vary depending on the target applications. Due to the selection of hardware elements that comprise the pipeline, the pipeline architecture of the present invention can accommodate, at least, substantially any existing or future compression algorithms that fall into the general class of block-oriented algorithms.
0068The rows of <figref idref="DRAWINGS">FIG. 6</figref> represent the decoding functions performed as part of the pipeline according to an exemplary embodiment. Variable length decoding <b>600</b> is performed by PVLD <b>306</b>. Run length/inverse scan/IQ/mismatch <b>602</b> are functions performed by IQ module <b>308</b>. Inverse transform operations <b>604</b> are performed by the inverse transform module <b>309</b>. Pixel filter reference fetch <b>606</b> and pixel filter reconstruction <b>608</b> are performed by pixel filter <b>310</b>. Motion compensation reconstruction <b>610</b> is performed by motion compensation module <b>312</b>. The columns of <figref idref="DRAWINGS">FIG. 6</figref> represent the pipeline stages. The designations MB<sub>i</sub>, MB<sub>i+2</sub>, etc. represent the i<sup>th </sup>macroblock in a data stream, the i+1<sup>st </sup>macroblock in the data stream, the i+2<sup>nd </sup>macroblock, and so on. The pipeline scheme supports one pipeline stage per module, wherein any hardware module that depends on the result of another module is arranged in a following MB pipeline stage. In an illustrative embodiment, the pipeline scheme can support more than one pipeline stage per module.
0069At any given stage in the pipeline, while a given function is being performed on a given macroblock, the next macroblock in the data stream is being worked on by the previous function in the pipeline. Thus, at stage x <b>612</b> in the pipeline represented in <figref idref="DRAWINGS">FIG. 6</figref>, variable length decoding <b>600</b> is performed on MBi. Exploded view <b>620</b> of the variable length decoding function <b>600</b> demonstrates how functions are divided between the core processor <b>302</b> and the PVLD <b>306</b> during this stage, according to one embodiment of the present invention. Exploded view <b>620</b> shows that during stage x <b>612</b>, the core processor <b>302</b> decodes the macmblock header of MB<sub>i</sub>. The PVLD <b>306</b> assists the core processor <b>302</b> in the decoding of macroblock headers. The core processor <b>302</b> also reconstructs the motion vectors of MB<sub>i</sub>, calculates the address of the pixel filter reference fetch for MB<sub>i</sub>, performs pipeline flow control and checks the status of IQ module <b>308</b>, inverse transform module <b>309</b>, pixel filter <b>310</b> and motion compensator <b>312</b> during stage x <b>612</b>. The hardware blocks operate concurrently with the core processor <b>302</b> while decoding a series of macroblocks. The core processor <b>302</b> controls the pipeline, initiates the decoding of each macroblock, and controls the operation of each of the hardware accelerators. The core processor firmware checks the status of each of the hardware blocks to determine completion of previously assigned tasks and checks the buffer availability before advancing the pipeline. Each block will then process the corresponding next macroblock. The PVLD <b>306</b> also decodes the macro block coefficients of Mbi during stage x. Block coefficient VLC decoding is not started until the core processor <b>302</b> decodes the whole macro block header. Note that the functions listed in exploded view <b>620</b> are performed during each stage of the pipeline of <figref idref="DRAWINGS">FIG. 6</figref>, even though, for simplicity's sake, they are only exploded out with respect to stage x <b>612</b>.
0070At the next stage x+1 <b>614</b>, the inverse quantizer <b>308</b> works on MB<sub>i </sub>(function <b>602</b>) while variable length decoding <b>600</b> is performed on the next macroblock, MB<sub>i+1</sub>. In stage x+1 <b>614</b>, the data that the inverse quantizer <b>308</b> works on are the quantized transform coefficients of MB<sub>i </sub>extracted from the data stream by the PVLD <b>306</b> during stage x <b>612</b>. In an exemplary embodiment of the present invention, also during stage x+1 <b>614</b>, the pixel filter reference data is fetched for MB<sub>i </sub>(function <b>606</b>) using the pixel filter reference fetch address calculated by the core processor <b>302</b> during stage x <b>612</b>.
0071Then, at stage x+2 <b>616</b>, the inverse transform module <b>309</b> performs inverse transform operations <b>604</b> on the MB<sub>i </sub>transform coefficients that were output by the inverse quantizer <b>308</b> during stage x+1. Also during stage x+2, the pixel filter <b>310</b> performs pixel filtering <b>608</b> for MB<sub>i </sub>using the pixel filter reference data fetched in stage x+1 <b>614</b> and the motion vectors reconstructed by the core processor <b>302</b> in stage x <b>612</b>. Additionally at stage x+2 <b>616</b>, the inverse quantizer <b>308</b> works on MB<sub>i+1 </sub>(function <b>602</b>), the pixel filter reference data is fetched for MB<sub>i+1 </sub>(function <b>606</b>), and variable length decoding <b>600</b> is performed on MB<sub>i+2</sub>.
0072At stage x+3 <b>618</b>, the motion compensation module <b>312</b> performs motion compensation reconstruction <b>610</b> on MB<sub>i </sub>using decoded difference pixel information produced by the inverse transform module <b>309</b> (function <b>604</b>) and pixel prediction data produced by the pixel filter <b>310</b> (function <b>608</b>) in stage x+2 <b>616</b>. Also during stage x+3 <b>618</b>, the inverse transform module <b>309</b> performs inverse transform operations <b>604</b> on MB<sub>i+h </sub>the pixel filter <b>310</b> performs pixel filtering <b>608</b> for MB<sub>i+1</sub>, the inverse quantizer <b>308</b> works on MBi+2 (function <b>602</b>), the pixel filter reference data is fetched for MB<sub>i+2 </sub>(function <b>606</b>), and variable length decoding <b>600</b> is performed on MB<sub>i+3</sub>. While the pipeline of <figref idref="DRAWINGS">FIG. 6</figref> shows just four pipeline stages, in an illustrative embodiment of the present invention, the pipeline includes as many stages as is needed to decode a complete incoming data stream.
0073In an alternative embodiment of the present invention, the functions of two or more hardware modules are combined into one pipeline stage and the macroblock data is processed by all the modules in that stage sequentially. For example, in an exemplary embodiment, inverse transform operations for a given macroblock are performed during the same pipeline stage as IQ operations. In this embodiment, the inverse transform module <b>309</b> waits idle until the inverse quantizer <b>308</b> finishes and the inverse quantizer <b>308</b> becomes idle when the inverse transform operations start. This embodiment will have a longer processing time for the “packed” pipeline stage, and therefore such embodiments may have lower throughput. The benefits of the packed stage embodiment include fewer pipeline stages, fewer buffers and possibly simpler control for the pipeline.
0074The above-described macroblock-level pipeline advances stage-by-stage. Conceptually, the pipeline advances after all the tasks in the current stage are completed. The time elapsed in one macroblock pipeline stage will be referred to herein as the macroblock (MB) time. In the general case of decoding, the MB time is not a constant and varies from stage to stage according to various factors, such as the amount of processing time required by a given acceleration module to complete processing of a given block of data in a given stage. It depends on the encoded bitstream characteristics and is determined by the bottleneck module, which is the one that finishes last in that stage. Any module, including the core processor <b>302</b> itself, could be the bottleneck from stage to stage and it is not pre-determined at the beginning of each stage.
0075However, for a given encoding/decoding algorithm, each module, including the core processor <b>302</b>, has a defined and predetermined task or group of tasks to complete. The macroblock time for each module is substantially constant for a given decoding standard. Therefore, in an illustrative embodiment of the present invention, the hardware acceleration pipeline is optimized by hardware balancing each module in the pipeline according to the compression format of the data stream.
0076The main video decoding operations occur within a direct execution loop that also includes polling of the accelerator functions. The coprocessor/accelerators operate concurrently with the core processor while decoding a series of macro blocks. The core processor <b>302</b> controls the pipeline, initiates the decoding of each macro block, and controls the operation of each of the accelerators. The core processor also does a lot of actual decoding, as described in previous paragraphs. Upon completion of each macroblock processing stage in the core processor, firmware checks the status of each of the accelerators to determine completion of previously assigned tasks. In the event that the firmware gets to this point before an accelerator module has completed its required tasks, the firmware polls for completion. This is appropriate, since the pipeline cannot proceed efficiently until all of the pipeline elements have completed the current stage, and an interrupt driven scheme would be less efficient for this purpose. In an alternative embodiment, the core processor <b>302</b> is interrupted by the coprocessor or hardware accelerators when an exceptional occurrence is detected, such as an error in the processing task. In another alternative embodiment, the coprocessor or hardware accelerators interrupt the core processor when they complete their assigned tasks.
0077Each hardware module <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b>, <b>313</b>, <b>315</b> is independently controllable by the core processor <b>302</b>. The core processor <b>302</b> drives a hardware module by issuing a certain start command after checking the module's status. In one embodiment, the core processor <b>302</b> issues the start command by setting up a register in the hardware module.
0078<figref idref="DRAWINGS">FIG. 7</figref> is a flowchart representing a macroblock decoding loop according to an illustrative embodiment of the present invention. <figref idref="DRAWINGS">FIG. 7</figref> depicts the decoding of one video picture, starting at the macro block level. In an illustrative embodiment of the present invention, the loop of macroblock level decoding pipeline control is fully synchronous. At step <b>700</b>, the core processor <b>302</b> retrieves a macroblock to be decoded from system memory <b>110</b>. At step <b>710</b>, the core processor starts all the hardware modules for which input data is available. The criteria for starting all modules depends on an exemplary pipeline control mechanism illustrated in <figref idref="DRAWINGS">FIG. 6</figref>. At step <b>720</b>, the core processor <b>302</b> decodes the macroblock header with the help of the PVLD <b>306</b>. At step <b>730</b>, when the macroblock header is decoded, the core processor <b>302</b> commands the PVLD <b>306</b> for block coefficient decoding. At step <b>740</b>, the core processor <b>302</b> calculates motion vectors and memory addresses, such as the pixel filter reference fetch address, controls buffer rotation and performs other housekeeping tasks. At step <b>750</b>, the core processor <b>302</b> checks to see whether the acceleration modules have completed their respective tasks. At decision box <b>760</b>, if all of the acceleration modules have completed their respective tasks, control passes to decision box <b>770</b>. If, at decision box <b>760</b>, one or more of the acceleration modules have not finished their tasks, the core processor <b>302</b> continues polling the acceleration modules until they have all completed their tasks, as shown by step <b>750</b> and decision box <b>760</b>. At decision box <b>770</b>, if the picture is decoded, the process is complete. If the picture is not decoded, the core processor <b>302</b> retrieves the next macroblock and the process continues as shown by step <b>700</b>. In an illustrative embodiment of the present invention, when the current picture has been decoded, the incoming macroblock data of the next picture in the video sequence is decoded according to the process of <figref idref="DRAWINGS">FIG. 7</figref>.
0079In general, the core processor <b>302</b> interprets the bits decoded (with the help of the PVLD <b>306</b>) for the appropriate headers and sets up and coordinates the actions of the hardware modules. More specifically, all header information, from the sequence level down to the macroblock level, is requested by the core processor <b>302</b>. The core processor <b>302</b> also controls and coordinates the actions of each hardware module. The core processor configures the hardware modules to operate in accordance with the encoding/decoding format of the data stream being decoded by providing operating parameters to the hardware modules. The parameters include but are not limited to (using MPEG2 as an example) the cbp (coded block pattern) used by the PVLD <b>306</b> to control the decoding of the transform block coefficients, the quantization scale used by the IQ module <b>308</b> to perform inverse quantization, motion vectors used by the pixel filter <b>309</b> and motion compensation module <b>310</b> to reconstruct the macroblocks, and the working buffer address(es) in decoder memory <b>316</b>.
0080Each hardware module <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b>, <b>313</b>, <b>315</b> performs the specific processing as instructed by the core processor <b>302</b> and sets up its status properly in a status register as the task is being executed and when it is done. Each of the modules has or shares a status register that is polled by the core processor to determine the module's status. In an alternative embodiment, each module issues an interrupt signal to the core processor so that in addition to polling the status registers, the core processor can be informed asynchronously of exceptional events like errors in the bitstream. Each hardware module is assigned a set of macroblock buffers in decoder memory <b>316</b> for processing purposes. In an illustrative embodiment, each hardware module signals the busy/available status of the working buffer(s) associated with it so that the core processor <b>302</b> can properly coordinate the processing pipeline.
0081In an exemplary embodiment of the present invention, the hardware accelerator modules <b>306</b>, <b>308</b>, <b>309</b>, <b>319</b>, <b>312</b>, <b>313</b>, <b>314</b>, <b>315</b> generally do not communicate with each other directly. The accelerators work on assigned areas of decoder memory <b>316</b> and produce results that are written back to decoder memory <b>316</b>, in some cases to the same area of decoder memory <b>316</b> as the input to the accelerator, or results are written back to main memory. In one embodiment of the present invention, when the incoming bitstream is of a format that includes a relatively large amount of data, or of a relatively complex encoding/decoding format, the accelerators in some cases may bypass the decoder memory <b>316</b> and pass data between themselves directly.
0082Software codecs from other sources, such as proprietary codecs, are ported to the decoding system <b>300</b> by analyzing the code to isolate those functions that are amenable to acceleration, such as variable-length decoding, run-length coding, inverse scanning, inverse quantization, transform, pixel filter, motion compensation, de-blocking filter, and display format conversion, and replacing those functions with equivalent functions that use the hardware accelerators in the decoding system <b>300</b>. In an exemplary embodiment of the present invention, modules <b>310</b>, <b>312</b> and <b>313</b>, <b>315</b> are implemented in a programmable SIMD/RISC filter engine module (<b>311</b> and <b>314</b> respectively) that allows execution of a wide range of decoding algorithms, even ones that have not yet been specified in by any standards body. Software representing all other video decoding tasks is compiled to run directly on the core processor.
0083In an illustrative embodiment of the present invention, some functions are interrupt driven, particularly the management of the display, i.e., telling the display module which picture buffer to display from at each field time, setting display parameters that depend on the picture type (e.g. field or frame), and performing synchronization functions. The decoding system <b>300</b> of the present invention provides flexible configurability and programmability to handle different video stream formats. <figref idref="DRAWINGS">FIG. 8</figref> is a flowchart representing a method of decoding a digital video data stream or set of streams containing more than one video data format, according to an illustrative embodiment of the present invention. At step <b>800</b>, video data of a first encoding/decoding format is received. At step <b>810</b>, at least one external decoding function, such as variable-length decoding or inverse quantization: is configured based on the first encoding/decoding format. At step <b>820</b>, video data of the first encoding/decoding format is decoded using the at least one external decoding function. In an illustrative embodiment of the present invention, a full picture, or a least a full row, is processed before changing formats and before changing streams. At step <b>830</b>, video data of a second encoding/decoding format is received. At step <b>840</b>, at least one external decoding function is configured based on the second encoding/decoding format. Then, at step <b>850</b>, video data of the second encoding/decoding format is decoded using the at least one external decoding function. In an exemplary embodiment, the at least one decoding function is performed by one or more of hardware accelerators <b>306</b>, <b>308</b>, <b>309</b>, <b>310</b>, <b>312</b>, <b>313</b>, <b>314</b> and <b>315</b>. The hardware accelerators are programmed or configured by the core processor <b>302</b> to operate according to the appropriate encoding/decoding format. As is described above with respect to the individual hardware accelerators of <figref idref="DRAWINGS">FIGS. 3 and 4</figref>, in one illustrative embodiment the programming for different decoding formats is done through register read/write. The core processor programs registers in each module to modify the operational behavior of the module.
0084In another illustrative embodiment, some or all of the hardware accelerators comprise programmable processors which are configured to operate according to different encoding/decoding formats by changing the software executed by those processors, in addition to programming registers as appropriate to the design. Although a preferred embodiment of the present invention has been described, it should not be construed to limit the scope of the appended claims. For example, the present invention is applicable to any type of media, including audio, in addition to the video media illustratively described herein. Those skilled in the art will understand that various modifications may be made to the described embodiment. Moreover, to those skilled in the various arts, the invention itself herein will suggest solutions to other tasks and adaptations for other applications. It is therefore desired that the present embodiments be considered in all respects as illustrative and not restrictive, reference being made to the appended claims rather than the foregoing description to indicate the scope of the invention.
Contents8
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO0028430A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO0030040A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO0145426A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO02087248A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| EP0367182A2 | Cites | European Patent Office (EPO) | Applicant |
| EP0537932A2 | Cites | European Patent Office (EPO) | Applicant |
| EP0572262A2 | Cites | European Patent Office (EPO) | Applicant |
| EP0615206A1 | Cites | European Patent Office (EPO) | Applicant |
| EP0663762A2 | Cites | European Patent Office (EPO) | Applicant |
| EP0884910A1 | Cites | European Patent Office (EPO) | Applicant |
| EP0945001B1 | Cites | European Patent Office (EPO) | Applicant |
| EP0947092B1 | Cites | European Patent Office (EPO) | Applicant |
| EP1124181A1 | Cites | European Patent Office (EPO) | Applicant |
| JP2000041262A | Cites | Japan | Applicant |
| JP2000059234A | Cites | Japan | Applicant |
| JP2000059769A | Cites | Japan | Applicant |
| JP2000066894A | Cites | Japan | Applicant |
| JP2000149445A | Cites | Japan | Applicant |
| JP2000215057A | Cites | Japan | Applicant |
| JP2000235644A | Cites | Japan | Applicant |
| JP2000331150A | Cites | Japan | Applicant |
| JP2000513523A | Cites | Japan | Applicant |
| US2001005432A1 | Cites | United States of America | Applicant |
| US2001014166A1 | Cites | United States of America | Applicant |
| US2001022816A1 | Cites | United States of America | Applicant |
| US2001026587A1 | Cites | United States of America | Applicant |
| US2001028680A1 | Cites | United States of America | Applicant |
| US2001033617A1 | Cites | United States of America | Applicant |
| US2001046260A1 | Cites | United States of America | Applicant |
| US2001046264A1 | Cites | United States of America | Applicant |
| JP2001094437A | Cites | Japan | Applicant |
| JP2001167057A | Cites | Japan | Applicant |
| JP2001168728A | Cites | Japan | Applicant |
| JP2001184500A | Cites | Japan | Applicant |
| JP2001256048A | Cites | Japan | Applicant |
| JP2001309386A | Cites | Japan | Applicant |
| JP2001507555A | Cites | Japan | Applicant |
| JP2001508631A | Cites | Japan | Applicant |
| JP2002010277A | Cites | Japan | Applicant |
| US2002034252A1 | Cites | United States of America | Applicant |
| US2002047919A1 | Cites | United States of America | Applicant |
| US2002057446A1 | Cites | United States of America | Applicant |
| US2002057739A1 | Cites | United States of America | Applicant |
| US2002065952A1 | Cites | United States of America | Applicant |
| US2002071490A1 | Cites | United States of America | Applicant |
| US2002085021A1 | Cites | United States of America | Search report |
| US2002085768A1 | Cites | United States of America | Applicant |
| US2002095689A1 | Cites | United States of America | Applicant |
| US2002118758A1 | Cites | United States of America | Applicant |
| JP2002135778A | Cites | Japan | Applicant |
| US2002150159A1 | Cites | United States of America | Applicant |
| US2002172288A1 | Cites | United States of America | Applicant |
| US2003021344A1 | Cites | United States of America | Applicant |
| US2003079035A1 | Cites | United States of America | Applicant |
| US2003112864A1 | Cites | United States of America | Applicant |
| US2003156652A1 | Cites | United States of America | Applicant |
| US2003174774A1 | Cites | United States of America | Applicant |
| US2003187662A1 | Cites | United States of America | Applicant |
| US2003222998A1 | Cites | United States of America | Applicant |
| US2003235251A1 | Cites | United States of America | Applicant |
| US2004008788A1 | Cites | United States of America | Applicant |
| US2004028141A1 | Cites | United States of America | Applicant |
| US2004039903A1 | Cites | United States of America | Applicant |
| US2004120404A1 | Cites | United States of America | Applicant |
| US2004233332A1 | Cites | United States of America | Applicant |
| US2005094729A1 | Cites | United States of America | Applicant |
| US2006126962A1 | Cites | United States of America | Applicant |
| JP2010246130A | Cites | Japan | Applicant |
| US2015195535A1 | Cites | United States of America | Applicant |
| JP2638613B2 | Cites | Japan | Applicant |
| JP2759896B2 | Cites | Japan | Applicant |
| JP2761449B2 | Cites | Japan | Applicant |
| JP2767847B2 | Cites | Japan | Applicant |
| JP2810896B2 | Cites | Japan | Applicant |
| JP2858602B2 | Cites | Japan | Applicant |
| EP3007419A1 | Cites | European Patent Office (EPO) | Applicant |
| JP3546437B2 | Cites | Japan | Applicant |
| JP3889069B2 | Cites | Japan | Applicant |
| JP4558409B2 | Cites | Japan | Applicant |
| US5212777A | Cites | United States of America | Applicant |
| US5239654A | Cites | United States of America | Applicant |
| US5269001A | Cites | United States of America | Applicant |
| US5379351A | Cites | United States of America | Applicant |
| US5379356A | Cites | United States of America | Applicant |
| US5386233A | Cites | United States of America | Applicant |
| US5432900A | Cites | United States of America | Applicant |
| US5488419A | Cites | United States of America | Applicant |
| US5506604A | Cites | United States of America | Applicant |
| US5508746A | Cites | United States of America | Applicant |
| US5512962A | Cites | United States of America | Applicant |
| US5528528A | Cites | United States of America | Applicant |
| US5568167A | Cites | United States of America | Applicant |
| US5576765A | Cites | United States of America | Applicant |
| US5579052A | Cites | United States of America | Applicant |
| US5589886A | Cites | United States of America | Applicant |
| US5592399A | Cites | United States of America | Applicant |
| US5594679A | Cites | United States of America | Applicant |
| US5594813A | Cites | United States of America | Applicant |
| US5598483A | Cites | United States of America | Applicant |
| US5598514A | Cites | United States of America | Applicant |
326 members in 9 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 11479802 | United States of America | A | |
| 11479802 | United States of America | A | |
| 201213608221 | United States of America | A | |
| 201213608221 | United States of America | A | |
| 201816103107 | United States of America | A | |
| 10114798 | – | – | – |
| 13608221 | – | – | – |
| US20020114798 | – | – | – |
| US201213608221 | – | – | – |
| US201816103107 | – | – | – |
Members326
| Document | Office | Kind | |
|---|---|---|---|
| US4258423A | United States of America | A | |
| CA1118057A | Canada | A | |
| WO0028518A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU1910800A | Australia | A | |
| US6189064B1 | United States of America | B1 | |
| WO0145426A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2260601A | Australia | A | |
| EP1145218A2 | European Patent Office (EPO) | A2 | |
| WO0028518A8 | World Intellectual Property Organization (WIPO) | A8 | |
| US6380945B1 | United States of America | B1 | |
| US2002093517A1 | United States of America | A1 | |
| US2002106018A1 | United States of America | A1 | |
| EP1238541A1 | European Patent Office (EPO) | A1 | |
| EP1239667A2 | European Patent Office (EPO) | A2 | |
| US2002145613A1 | United States of America | A1 | |
| US6501480B1 | United States of America | B1 | |
| US6529935B1 | United States of America | B1 | |
| US6538656B1 | United States of America | B1 | |
| US6570579B1 | United States of America | B1 | |
| US6573905B1 | United States of America | B1 | |
| US2003117406A1 | United States of America | A1 | |
| US6608630B1 | United States of America | B1 | |
| US2003158987A1 | United States of America | A1 | |
| US2003184457A1 | United States of America | A1 | |
| US2003185298A1 | United States of America | A1 | |
| US2003185305A1 | United States of America | A1 | |
| US2003185306A1 | United States of America | A1 | |
| US2003187824A1 | United States of America | A1 | |
| US2003187895A1 | United States of America | A1 | |
| US2003188127A1 | United States of America | A1 | |
| US6630945B1 | United States of America | B1 | |
| EP1351511A2 | European Patent Office (EPO) | A2 | |
| EP1351512A2 | European Patent Office (EPO) | A2 | |
| EP1351513A2 | European Patent Office (EPO) | A2 | |
| EP1351514A2 | European Patent Office (EPO) | A2 | |
| EP1351515A2 | European Patent Office (EPO) | A2 | |
| EP1351516A2 | European Patent Office (EPO) | A2 | |
| US2003189571A1 | United States of America | A1 | |
| US2003189982A1 | United States of America | A1 | |
| WO03085494A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO03085981A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US6636222B1 | United States of America | B1 | |
| EP1355499A2 | European Patent Office (EPO) | A2 | |
| US2003206174A1 | United States of America | A1 | |
| EP1365319A1 | European Patent Office (EPO) | A1 | |
| EP1365385A2 | European Patent Office (EPO) | A2 | |
| US6661422B1 | United States of America | B1 | |
| US6661427B1 | United States of America | B1 | |
| WO03085494A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2003235251A1 | United States of America | A1 | |
| EP1376379A2 | European Patent Office (EPO) | A2 | |
| US2004017398A1 | United States of America | A1 | |
| US2004028141A1 | United States of America | A1 | |
| US6700588B1 | United States of America | B1 | |
| US2004047194A1 | United States of America | A1 | |
| EP1238541B1 | European Patent Office (EPO) | B1 | |
| US2004056864A1 | United States of America | A1 | |
| US2004056874A1 | United States of America | A1 | |
| US6721837B2 | United States of America | B2 | |
| AT262253T | Austria | T | |
| ATE262253T1 | Austria | T1 | |
| DE60009140D1 | Germany | D1 | |
| US6731295B1 | United States of America | B1 | |
| US6738072B1 | United States of America | B1 | |
| EP1145218B1 | European Patent Office (EPO) | B1 | |
| US6744472B1 | United States of America | B1 | |
| AT267439T | Austria | T | |
| ATE267439T1 | Austria | T1 | |
| DE69917489D1 | Germany | D1 | |
| US2004130558A1 | United States of America | A1 | |
| US6762762B2 | United States of America | B2 | |
| US6768774B1 | United States of America | B1 | |
| US6771196B2 | United States of America | B2 | |
| US2004150652A1 | United States of America | A1 | |
| EP1376379A3 | European Patent Office (EPO) | A3 | |
| US6781601B2 | United States of America | B2 | |
| US2004169660A1 | United States of America | A1 | |
| US2004177190A1 | United States of America | A1 | |
| US2004177191A1 | United States of America | A1 | |
| US6798420B1 | United States of America | B1 | |
| US2004207644A1 | United States of America | A1 | |
| US2004208245A1 | United States of America | A1 | |
| US2004212730A1 | United States of America | A1 | |
| US2004212734A1 | United States of America | A1 | |
| US6819330B2 | United States of America | B2 | |
| US2004246257A1 | United States of America | A1 | |
| US2005007264A1 | United States of America | A1 | |
| US2005012759A1 | United States of America | A1 | |
| DE60009140T2 | Germany | T2 | |
| US2005024369A1 | United States of America | A1 | |
| US6853385B1 | United States of America | B1 | |
| US2005044175A1 | United States of America | A1 | |
| EP1239667A3 | European Patent Office (EPO) | A3 | |
| US6870538B2 | United States of America | B2 | |
| US6879330B2 | United States of America | B2 | |
| DE69917489T2 | Germany | T2 | |
| US2005122335A1 | United States of America | A1 | |
| US2005122341A1 | United States of America | A1 | |
| US2005123057A1 | United States of America | A1 | |
| EP1351514A3 | European Patent Office (EPO) | A3 |
91 transactions on the USPTO file
Allowed after 2 non-final rejections, 2 final rejections and 1 RCE.
- Non-final rejections
- 2
- Final rejections
- 2
- RCEs
- 1
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Email NotificationEML_NTR | EML_NTR | |
| Application ready for PDX access by participating foreign officesCCRDY | CCRDY | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Reasons for AllowanceEX.R | EX.R | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Oath or Declaration Filed (Including Supplemental)C602 | C602 | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Advisory Action (PTOL - 303)MCTAV | MCTAV | |
| Advisory Action (PTOL-303)CTAV | CTAV | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Oath or Declaration Filed (Including Supplemental)C602 | C602 | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Oath or Declaration Filed (Including Supplemental)C602 | C602 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Paralegal or electronic terminal disclaimer approvedP574 | P574 | |
| Terminal Disclaimer FiledDIST | DIST | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Preliminary AmendmentA.PE | A.PE | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Notice of Reissue Published in Official GazetteNRE. | NRE. | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Email NotificationEML_NTR | EML_NTR | |
| Application Is Now CompleteCOMP | COMP | |
| Filing Receipt - UpdatedFLRCPT.U | FLRCPT.U | |
| Application Dispatched from OIPEOIPE | OIPE | |
| FITF set to NO - revise initial settingFTFI | FTFI | |
| Additional Application Filing FeesADDFLFEE | ADDFLFEE | |
| A statement by one or more inventors satisfying the requirement under 35 USC 115, Oath of the ApplicOATHDECL | OATHDECL | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTR | EML_NTR | |
| Email NotificationEML_NTF | EML_NTF | |
| Filing ReceiptFLRCPT.O | FLRCPT.O | |
| Notice Mailed--Application Incomplete--Filing Date AssignedINCD | INCD | |
| PTO/SB/69-Authorize EPO Access to Search ResultsSREXR141 | SREXR141 | |
| Applicants have given acceptable permission for participating foreignAPPERMS | APPERMS | |
| Cleared by OIPE CSRL194 | L194 | |
| Entity Status Set To Undiscounted (Initial Default Setting or Status Change)BIG. | BIG. | |
| Initial Exam Team nnIEXX | IEXX |
1 recorded assignment at the USPTO, latest first
- Now
Now: Held by
BROADCOM CORP - 2019-03-26
Assignment of assignors interest.
- From
- MACINNIS, ALEXANDER G.ALVAREZ, JOSE' R.ZHONG, SHENG
and 2 moreShow fewer
XIE, XIAODONGHSIUN, VIVIAN - To
- BROADCOM CORPORATION
Recorded 2019-03-26, Signed 2002-03-30
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| AssignmentAS | AS | |
| Fee payment procedureENTITY STATUS SET TO UNDISCOUNTED (ORIGINAL EVENT CODE: BIG.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP |
Numbers
- Publication
- RE048845
- Publication, DOCDB
- RE48845
- Publication, EPODOC
- USRE48845E
- Application
- 16103107
- Application, DOCDB
- 201816103107
- Application, EPODOC
- US201816103107
Titles
- English
- Video decoding system supporting multiple standards
Classification
- CPC, 16
- G06F9/3861
- G06F9/3877
- H04N19/176
- H04N19/70
- H04N19/12
- H04N19/122
- H04N19/129
- H04N19/157
- H04N19/61
- H04N19/60
- H04N19/423
- H04N19/91
- H04N19/44
- H04N19/625
- H04N19/82
- H04N19/90
- IPC, 19
- G06F9 38
- H04N19 12
- H04N19 122
- H04N19 129
- H04N19 157
- H04N19 176
- H04N19 423
- H04N19 44
- H04N19 60
- H04N19 61
- H04N19 625
- H04N19 70
- H04N19 82
- H04N19 90
- H04N19 91
- G06T9 00
- H04N7 26
- H04N7 30
- H04N7 50