System, method and format for scalable encoded media delivery
Abstract
Problem to be solved.To distribute scalable encoded media data to different types of reception bases.
Solution.A scalable coded media source data is formatted into a format including a first part and a second part. The first part corresponds to the media type non-specific scalability attributes of the encoded media source data and the data structure information of the second part. The second part corresponds to scalable encoded media source data arranged in a one-dimensional media type non-unique indexable data structure. Preformatted scalable encoded media source using data structure information before delivery to media destination Data is code-transformed and formatted scalable encoded media source based on matching scalability and receive attributes. Generate a scaled version of the data. The code-converted media data matches the receive attribute of the media destination. [Selection diagram] Fig. 10
Term
Term ended
Projected expiry passed 15 July 2023, 3.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
4 claims: 4 independent, 0 dependent
- 1符号化されたスケーラブルなメディアデータを配信する方法であって、 少なくとも符号化メディア元データのメディアタイプ非固有のスケーラビリティ属性と第2の部分のデータ構造情報とに対応する第1の部分と、少なくとも1つの次元を有するメディアタイプ非固有の索引付け可能なデータ構造に配列された前記スケーラブルな符号化メディア元データに対応する第2の部分とを含むフォーマットに、スケーラブルな符号化メディア元データをフォーマットし、 スケーラブルな符号化メディアの少なくとも1つのタイプのメディア宛先の受信属性に対応する情報を提供し、 前記フォーマット済のメディアデータを前記メディア宛先に配信する前にコード変換して、前記スケーラビリティ属性と前記受信属性とのマッチングに基づいて、前記メディア宛先の前記受信属性に適合するように前記フォーマット済メディアデータのスケーリング版を生成することを含む方法。
- 2スケーラブルな符号化メディアデータを配信するシステムであって、 少なくとも符号化メディア元データのメディアタイプ非固有のスケーラビリティ属性と第2の部分のデータ構造情報とに対応する第1の部分と、少なくとも1つの次元を有するメディアタイプ非固有の索引付け可能なデータ構造に配列された前記スケーラブルな符号化メディア元データに対応する第2の部分とを含むフォーマットで、スケーラブルな符号化メディア元データを提供するメディア発信元と、 スケーラブルな符号化メディアの少なくとも1つのタイプのメディア宛先の受信属性に対応する情報を提供するメディア宛先と、 前記スケーラビリティ属性と前記受信属性とのマッチングに基づいて、前記フォーマット済のスケーラブルな符号化メディア元データを前記メディア宛先に配信する前にコード変換して、前記フォーマット済のスケーラブルな符号化メディア元データのスケーリング版を生成するトランスコーダと、 を備えるシステム。
- 3少なくとも符号化メディア元データのメディアタイプ非固有スケーラビリティ属性と第2の部分のデータ構造情報とに対応する第1の部分と、少なくとも1つの次元を有するメディアタイプ非固有の索引付け可能なデータ構造に配列された前記スケーラブルな符号化メディア元データに対応する前記第2の部分とを含む、フォーマット済のスケーラブルなメディアビットストリームを受信し、 スケーラブルな符号化メディアの少なくとも1つのタイプのメディア宛先の受信属性に対応する情報を受信し、 前記スケーラビリティ属性と前記受信情報とを比較し、 該比較に応じて、前記ビットストリーム部分集合の打ち切り、ドロップ、および再配列のうちの1つを行って、前記メディア宛先に適合された前記フォーマット済のスケーラブルな符号化メディア元データのスケーリング版を前記データ構造情報を使用して生成することを含む、コード変換方法。
- 4スケーラブルな符号化メディアビットストリームのフォーマットであって、 少なくとも符号化メディア元データのメディアタイプ非固有スケーラビリティ属性と第2の部分のデータ構造情報とに対応する第1の部分と、 少なくとも1つの次元を有するメディアタイプ非固有の索引付け可能なデータ構造に配列された前記スケーラブルな符号化メディアデータに対応する第2の部分と、 を含むフォーマット。
Independent claims4
65 paragraphs, as filed
The present invention relates to the delivery of media data, in particular to systems, methods and formats thereof for delivering all types of scalable encoded media to media destinations with various receiving attributes.
Today, users access the Internet with connections ranging from 56Kbps modems to high-speed 100Mb / s Ethernet, using devices ranging from handhelds to powerful workstations. The available bandwidth, display and processing powers continue to grow, but the breadth of diversity and capabilities is always present. On the other hand, as bandwidth and other factors continue to grow, so does the abundance of media delivered to users.
In such situations, a strict media representation format that produces decompressed content with only fixed resolution and quality is clearly unsuitable. A distribution system based on such a compression method can satisfactorily distribute content to only a small part of users who are interested in the content. The rest of the users will either not be able to receive anything at all, or will receive at a lower quality and / or resolution compared to the capabilities of their respective network connections and / or access devices. This inability to respond to diversity is a determinant of the growth of new and abundant media. This is because only power users, who are only a small part of the whole, can provide such rich content. Without proper focus on seamless content fitting, media accessibility and availability will always be severely restricted.
One of the known methods of providing media content to users with different abilities and preferences is to provide multiple versions of media suitable for different abilities and preferences. This technique works well for delivery models where the receiver connects directly to the media originator, but other multi-hop multi-recipient delivery scenarios introduce significant redundancy and inefficiency, wasting bandwidth and storage. Connect. This is especially true when media authors want to offer a wide range of choices to a large consumer base and therefore need to maintain a large number of versions that differ in many ways.
To overcome the above problems, a scalable compression format has been proposed. The scalable compressed representation can be adapted to all users by automatically maximizing the multimedia experience with respect to the computing power and connection speed of a given user. By adapting the rich media content written for high-end machines with high-speed connections to poorly performing machines with slower connections, we decided to create another version for another scenario as described above. The accompanying overhead can be virtually eliminated. In addition, content created in the highest quality possible today remains "timeless" when expressed in a scalable format, and the experience it provides continues to increase in machine capabilities and connection speeds. Gradually increase.
An example of a scalable compressed representation is JPEG2000. JPEG2000 is a scalable standard for still images that combines quality scalability and resolution scalability in a format unique to JPEG2000 compressed data, enabling distribution and viewing on a variety of connections and devices. However, to benefit from the scalability of the format, it is necessary to develop and deploy infrastructures that specifically support the code conversion of JPEG2000 content and the distribution to heterogeneous recipient bases.
In recent years, there has been a great deal of interest in the distribution of streaming video via the Internet and wireless. For this reason, MPEG-X (mostly MPEG-4) and H.26X family video standards have been developed that incorporate various forms of scalability in order to deliver media content such as streaming video to different reception bases. However, this type of scalable video over the Internet is limited to maintaining multiple versions for a few different types of connections. This is because there is no complete social infrastructure to support the transmission of scalable video formats.
The laying of any social infrastructure is expensive and requires large amounts of financial support from supporting companies or joint ventures. It is also desirable to standardize the format that represents the content in order to guarantee the stability of the format. On the other hand, standardization will take years to take effect, much longer than the usual pace of change in the multimedia industry. As new types of media evolve beyond traditional image, video and audio, it becomes increasingly difficult to predict the standards that support their respective representations. Even when evolving scalable formats for every new type of media, the inevitable differences in content structure require the use of different social infrastructures for the scalable delivery of different types of media. The associated costs are a huge obstacle to the adoption of these new media and the supportability of scalability features.
There are various types of bitstream scalability that can be devised depending on the type of media. For example, signal-to-noise (SNR) scalability refers to quality that gradually increases as the number of bitstreams contained increases, and applies to most types of media. Resolution scalability refers to the fineness of spatial data sampling and is applied to visual media such as images, videos, and 3D. Time scalability refers to the particle size of sampling in the time domain and applies to video and other image sequences. There are several types of audio scalability, such as the number of channels and sampling frequency. In the future, the evolution of newer, more abundant, and more interactive types of media will create newer types of scalability that are currently unknown. A scalable bitstream does not always have a single type of scalability. Different types of scalability can exist at the same time to provide a range of matching choices.
In new and abundant media, different media elements are often bundled together to provide a composite media experience. According to one known technique, audio annotated images and certain animations use three media elements (images, audio clips, and some animation data) to provide a complex presentation experience. Such a rich composite media model leads to a new type of scalability that is unique to media. This is because certain non-essential elements of the complex may be scraped to address other, more important elements within the limited resources of the receiver.
<p> The present invention is a system, method, and format for delivering scalable encoded media data to heterogeneous reception bases, which is non-specific to the type of media being delivered, thereby all types of present. It provides systems, methods, and formats thereof that require a single delivery infrastructure to deliver known and future media.</p>
<p> Systems, methods and formats for delivering encoded and scalable media data are provided. According to the first embodiment of the distribution method, the scalable encoded media source data is formatted into a format including the first part and the second part. The first part corresponds to at least the media type non-specific scalability attributes of the encoded media source data and the data structure information of the second part. The second part corresponds to scalable encoded media source data arranged in a media type non-unique indexable data structure with at least one dimension. Further, it is possible to provide information corresponding to the reception attribute of the scalable encoded media by the media destination. Preformatted scalable encoded media source using data structure information before delivery to media destination Data is code-transformed and formatted scalable encoded media source based on matching scalability and receive attributes. Generate a scaled version of the data. The code-converted media data matches the receive attribute of the media destination.</p><p> According to the second embodiment of the delivery method, it has at least one dimension and a first part corresponding to the media type non-unique scalability attribute of the encoded media source data and the data structure information of the second part. Receives a formatted scalable media bitstream containing multiple bitstream subsets, including a second part corresponding to the scalable encoded media source data arranged in a media type non-unique indexable data structure. A code conversion method is provided. In addition, information corresponding to at least one type of receive attribute of the scalable encoded media by the media destination is received. Compares scalability attributes and received information, and depending on the comparison, either truncate, drop, or rearrange the bitstream subset of the bitstream and use the data structure information. Generates a scaled version of the formatted, scalable coded media source data adapted to the media destination.</p><p> One embodiment of a system that delivers coded scalable media data comprises a media source, a media destination, and a transcoder. The media source provides scalable encoded media data in a format that includes a first part and a second part. The first part corresponds to at least the media type non-specific scalability attributes of the encoded media source data and the data structure information of the second part. The second part corresponds to scalable encoded media source data arranged in an indexable data structure that is unique to the media type and has at least one dimension. The media destination provides information corresponding to at least one type of receive attribute of the scalable encoded media by the media destination. In another embodiment, the receive attribute information includes information corresponding to the preference of the media destination. The transcoder translates the formatted, scalable coded media source data before delivering it to the media destination, and uses the data structure information to match the scalability and receive attributes to the formatted scalable code. Generate a scaled version of the original data.</p><p> A scalable coded media bitstream format that includes a first part and a second part is provided. The first part corresponds to at least the media type non-unique scalability attributes of the encoded media source data and the data structure information of the second part, and the second part is media type non-unique with at least one dimension. Corresponds to scalable encoded media source data arranged in an indexable data structure.</p>
In general, the present invention is a system and method for delivering a scalable encoded media bitstream and providing a general purpose (ie, media type non-specific) format for the scalable encoded media bitstream. The system, methods and formats offer a variety of reception capabilities, including tailoring tailored based on readability of media destinations, and media-specific and non-media-specific scalability attributes of formatted, scalable encoded media bitstreams. Provides seamless and flexible delivery to attributed media destinations. The system and methods are extensible to support the adaptation and delivery of all new types of scalable media that will evolve in the future.
Scalable Encoded Media Bitstreams, when grouped, are smaller multiple codes of the bitstream that can generate media representations with different scales of a particular scalable attribute of the media bitstream, such as quality, resolution, etc. Generally defined as a coded bitstream consisting of a modified subset. For example, if a scalable coded bitstream containing multiple coded subsets is resolution related, then all coded subsets provide the highest resolution representation. Lower resolutions can be obtained by dropping the subset (in the intermediate bitstream) or by truncate the subset (from the end of the bitstream), by dropping / breaking the subset again. Even lower resolutions can be obtained. After dropping / censoring the subset, repack the remaining bitstream subset and adjust the position of the dropped subset. In general, the action of dropping, censoring, and rearranging subsets is commonly referred to as a code conversion operation.
It should be understood that according to the present invention, a scalable bitstream can have more than one type of scalability. In addition, different types of scalability (eg, signal-to-noise ratio (SNR)), resolution, time, and interactivity can be applied to different types of media. In addition, the scalable encoded bitstream contains a nested (nested) scalability hierarchy.
The generation of scalable encoded media bitstreams is known in the field of media distribution. For example, in JPEG2000 image compression, a scalable encoded media bitstream can be generated by wavelet-decomposing the original media bitstream to obtain a block of coefficients (called a subband). The subbands of the coefficients are scanned in such a way as to obtain bitplane-by-bitplane encoding of the original media bitstream, and each encoded bitplane is represented by a plurality of bitstream subsets. Other known techniques for producing scalable encoded media bitstreams include video compression and audio compression. It should be understood that the systems, methods, and formats of the present invention are applicable to any scalable encoded bitstream generated by any technique.
FIG. 1 shows a first embodiment of a media type non-unique format of scalable encoded media data according to the invention, including a first portion 10 and a second portion 11. The first part 10 corresponds to at least media type non-specific scalability attributes. Media type non-specific scalability attributes generally include attributes that are common to all media types. For example, media type non-specific scalability attributes include size (corresponding to the size of the bitstream), display resolution (required to display the content obtained from the bitstream), and SNR (compression of the content obtained from the bitstream). It can include, but is not limited to, a unit of fidelity to a non-version), and processing power (required for a media experience). In one embodiment, each attribute is associated with an n-byte code that uniquely identifies the attribute. Reservation codes can be used for standardized attributes that have universal meaning across multiple media types, and other bytes can be reserved for future attribute type codes. Attributes can be represented by standardized values to maintain uniformity across all media types and capabilities. In one embodiment, the attribute can be quantized with a "decrease" code value or an "increase" code value. The first part 10 also contains the media type non-unique data structure information of the second part 11. In one embodiment, the data structure information relates to the dimension of the multidimensional representation of the scalable encoded media bitstream.
The second part 11 corresponds to scalable encoded media data arranged in a content-independent indexable data structure. Specifically, whatever the content of the coded media data is, it is arranged in the general-purpose format according to the present invention. Arranging the coded media data in this way allows for general purpose code conversion and is single because the code conversion operation is performed without knowing the actual media content and without decoding or decoding the media data. It will be possible to deliver many types of media content with the foundation / transcoder of. In addition, the format facilitates code conversion operations such as truncation, bitstream skipping, and repacking of the encoded bitstream without knowing the compression scheme previously applied to the actual content or encoded bitstream. Generate a scaled version. In addition, content-independent indexable nested formats are not unique to any type of media and can be used for both current and future media types.
In one embodiment, the scalable encoded media data is arranged as shown in FIG. 2A, and each layer corresponds to a different type of scalability. Data can be indexed using multiple TOCs, and each hierarchy can be indexed by the corresponding TOC. In another embodiment, the second part is indexed using a single TOC. For example, as shown in FIG. 2A, the first layer of the bitstream contains a first bitstream-encoded subset (subset 0) and a second bitstream-encoded subset (subset 1). As mentioned above, scalability can be achieved by grouping subsets of bitstreams to provide scalability to a particular hierarchy. For example, the first scalability can only be obtained from subset 0, while the second scalability can be provided from a subset that combines subset 0 and subset 1. The type of scalability provided by the first tier depends on the actual content of the first and second subsets. The first and second bitstream subsets of the first layer can be further subdivided into the first and second bitstream subsets of the second layer (subset 0 and subset 1), respectively. Again, the content of the subset of the second tier determines what type of scalability the tier provides. The third hierarchy is divided in the same way. An example of this type of multi-layer scalable bitstream is a JPEG 2000 bitstream. In one of JPEG 2000's progressive modes, the highest layer corresponds to resolution scalability, and a second layer, the signal-to-noise ratio (SNR) subset, is nested within the resolution-scalable subset. Note that the example shown in Figure 2A provides a TOC (table of contents), in part for random access and for fast identification of subsets to drop or censor during code conversion operations.
FIG. 2B shows an alternative representation of the stacked scaleable encoded bitstream shown in FIG. 2A. In this representation, each dimension of the cube contains a plurality of bitstream subsets B (x, y, z) arranged in the cube corresponding to each layer of FIG. 2A. A given attribute can decrease or increase along the dimension. For example, if layer 1 corresponds to resolution, the resolution increases along the x-dimensional. In this representation, the code conversion can be easily performed by dropping the layer and updating the TOC from time to time. In other words, the coding conversion / scaling of the coded bitstream can be achieved by truncating the rows or columns of the cube. For example, if layer 1 corresponds to resolution, layer 2 corresponds to SNR, layer 3 corresponds to interactivity, and if the subset shown as 10 is truncated, the encoded bitstream will have an increased SNR. But the resolution and interactivity are not scaled. In one embodiment, causality is maintained during the encoding and encryption of media data.
According to the second embodiment of the media non-unique format, the media content passed during each transmission instance is called a parcel. Each parcel in the generic case can be configured with multiple media components to provide a complex experience. For example, one component may be an image and the next component may be audio annotations associated with the image, both components packed together in a single parcel and viewed with audio annotations. Provide an experience. Each media component, along with a header containing a description, is a coded data unit that can be represented in a scalable, non-media-specific format. The overall media description of the parcel consists of a description of the individual components in the header, while the overall parcel data consists of (scalable) coded data of the individual components.
In general, according to a second embodiment of a non-media-specific format, each parcel contains two parts: a parcel header and parcel data (FIG. 3A). In general, the parcel header portion specifically includes the number of media components, as well as the individual headers of each component. The parcel data portion contains coded data for the individual components.
An embodiment of the format of each media component header is shown in FIG. 3B. The header begins with a flag that specifies whether the media component is in a media type non-specific format according to the invention. If unformatted, no code conversion is done and the entire media parcel is transferred directly to the outbound connection. In this case, there is no component description in the header. However, if the flag indicates that the parcel is scalable and conforms to a media type non-specific format, the header is followed by a component description.
The component description includes the number of nested scalability layers L corresponding to the number of dimensions of the cube shown in FIG. 2B, followed by the number of layers in each layer i corresponding to the number of rows of the cube. List l<sub>i</sub>including. Next comes a list called a consistency list, which consists of a subset of the hierarchy that is important for maintaining consistency across the same type of parcel. This will be described in detail later.
The consistency list is followed by a single-bit flag called the Scalability_Flag, which describes whether the data portion is in a scalable format or whether the bitstream is packed with multiple independent versions. .. In other words, the same media component header can be applied to both incrementally scalable bitstreams and multiple versions of scalable bitstreams. In general, all code conversion operations for incremental and multi-version scalable bitstreams are the same. However, in certain cases, it is possible to increase the bandwidth efficiency of the transcoder by knowing that the scalable bitstream contains multiple independent versions.
The next field is the number of attributes N associated with the media component, followed by a list of data required for each. In general, an attribute is quantitatively represented by a non-negative number called an attribute value. For reserved attributes, quantification is also standardized with the code. For example, the size can be expressed in K bytes, the display resolution (display_resolution) can be expressed as the diagonal width of the screen in terms of the number of pixels, and the processing power (processing_power) is CPU speed x number of processors (CPU_speed). × Number_of_processors), and so on. The methods used to quantify reserved attributes are standardized to preserve uniformity and capacity transfer methods across different types of media. For the best known attributes, the value does not decrease or increase with the layer. Therefore, as more layers are added to the scalable media, the attribute values usually change monotonically.
The data for each attribute first contains a unique attribute code (Attribute_code) field that identifies this attribute. In one embodiment, the attribute code consists of two fields: an attribute ID (Attribute_ID) and an attribute combination (Attribute_combination). The attribute ID is a unique identifier, and the attribute join is a field that describes how an attribute value changes when combined with another media component that has the same attributes. Possible values are addition, maximum value, minimum value, and the like. For example, size is always additive in combination, while display resolution is the maximum of individual components after combination. That is, when two or more media components are combined, the required size is the sum of the required sizes for all. On the other hand, the required display resolution is the highest of all. Overall, a unique attribute code not only identifies an attribute, but also defines its behavior when combined with another component.
Next, the field is an attribute monotonous type (Attribute_Monotone_Type), and shows how the attribute value changes as the number of layers increases. Possible types are monotonic non-decrease, monotonic non-increase, and non-monotonic with the number of layers.
The next field in the attribute data is the reference attribute value. This is a numeric reference value for the attribute, which, when multiplied by the distribution value that follows, gives the attribute value for the various layer drop options.
The reference attribute value field is followed by a layer / hierarchical distribution field, which specifies how the attribute value changes when the layer is dropped. This designation is called a distribution because of its concurrency with the cumulative distribution of random vectors. The distribution specified can be accurate or approximate.
In one embodiment, the exact distribution is similar to the multidimensional cumulative distribution. As an example, there are L nested hierarchies, l in the i-th hierarchy.<sub>i</sub>If there are layers, the distribution field will be l<sub>0</sub>× l<sub>1</sub>× ... × l<sub>L-1</sub>Is an L-dimensional matrix of the size of. However, C (j<sub>0</sub>, j<sub>1</sub>, ..., j<sub>L-1</sub>) Is shown (j<sub>0</sub>, j<sub>1</sub>, ..., j<sub>L-1</sub>) Third element (j<sub>0</sub>= 0,1, ..., l<sub>0</sub>-1; j<sub>1</sub>= 0,1, ..., l<sub>1</sub>-1; j<sub>i</sub>= 0,1, ..., l<sub>1</sub>-1; ...; j<sub>L-1</sub>= 0,1, ..., l<sub>L-1</sub>-1) is the maximum (j<sub>0</sub>, j<sub>1</sub>, ..., j<sub>L-1</sub>) If only the layer is sent, it is the number that specifies the multiplier of the reference attribute value to get the attribute value of the component. Optional empty multiplier C<sub>ψ</sub>Specifies a multiplier for the reference attribute value to get the attribute value of the component when the entire component is dropped, i.e. no layer is sent. The default empty multiplier is 0. Therefore, the total number of multipliers that need to be sent is 1 + l<sub>0</sub>× l<sub>1</sub>× ... × l<sub>L-1</sub>Is. Last multiplier C (l<sub>0</sub>-1, l<sub>1</sub>-1, ..., l<sub>L-1</sub>Note that multiplying the reference attribute value (Reference_Attribute_value) by -1) gives the full attribute value, that is, the attribute value that the media would have if it were transmitted without dropping any layers. Also, for monotonous non-decreasing type attributes, the reference attribute value is equal to the attribute value of the complete media where the layer is not dropped, that is, C (l).<sub>0</sub>-1, l<sub>1</sub>-1, ..., l<sub>L-1</sub>If -1) = 1, the multiplier C (j)<sub>0</sub>, j<sub>1</sub>, ..., j<sub>L-1</sub>) Is similar to the cumulative distribution of multidimensional discrete random vectors.
In an exemplary embodiment for the first two layers of JPEG2000RLCP progressive mode, the size and display resolution attribute distribution specifications are as shown in Figures 3C and 3D. Both are monotonous and non-decreasing. Here we have four nested spatial scalability layers, each with three SNR scalable layers. Note that in Figure 3D, the display resolution attribute does not change with the SNR scalable layer. If the SNR layer and two spatial layers are dropped as a result of the code conversion, the size attribute of the code conversion bitstream shown in shadows in Figures 3C and 3D is 0.18 times the reference size value and the display resolution attribute is. It is 0.25 times the reference display resolution value.
In one embodiment, the cumulative distribution is specified accurately or approximately using the product of one or more lower dimensional marginal distributions. In this case, element C (j) uses the product connection of the marginal distribution.<sub>0</sub>, j<sub>1</sub>, ..., j<sub>L-1</sub>) Approximately C ^ (j<sub>0</sub>, j<sub>1</sub>, ..., j<sub>L-1</sub>). That is, the specification is P low-dimensional cumulative distribution C that covers the L dimension together.<sub>i</sub>(.)including. That is, C ^ (j<sub>0</sub>, j<sub>1</sub>, ..., j<sub>L-1</sub>) = C<sub>0</sub>() × C<sub>1</sub>() × ... × C<sub>P-1</sub>(). Empty part C<sub>ψ</sub>Is sent separately.
Whether the distribution is accurate or approximate, the distribution description is first empty part C<sub>ψ</sub>Is followed by a number P indicating the number of specified product distributions, one for each L layer, followed by L P-adic numbers (P-) indicating which layer is mapped to which distribution. A list of ary) elements follows. This is followed by the actual specifications of the P distribution.
In one exemplary embodiment for JPEG 2000, an approximate specification using two one-dimensional perimeters and the resulting final approximate distribution can be expressed as shown in FIGS. 3E and 3F. As shown in Figure 3F, the display resolution is accurately represented using the approximation technique, and the size is represented only approximately.
FIG. 4A shows another embodiment of a media type non-unique format with a component dependency matrix D that defines the mode in which the components depend. In particular, components may or may not be excluded during code conversion. Certain components in the media must be included after code conversion, even if they are only the lowest scalable layer B (0,0), and certain other components must be dropped completely. Can be done. In addition, if one component is included or excluded, depending on the media, then certain other components must also be included or excluded. All this information at the component level is communicated with respect to the component dependency matrix.
FIG. 4B shows an example of matrix D. In one embodiment, if there are M components in the media parcel, the component dependency rule is specified for the M × M matrix D. Diagonal element d<sub>ii</sub>Can be a binary number and can specify whether the i-th component must be included even if it is only the lowest layer after code conversion. d<sub>ii</sub>= 1 indicates that the i-th component must be included, d<sub>ii</sub>= 0 indicates that the i-th component can be dropped if necessary. Off-diagonal element d<sub>ij</sub>(i j) is a pentadecimal number (5-ary), and if the i-th component is included or excluded, does it have to include or exclude the j-th component? To specify. d<sub>ij</sub>= 0 indicates that there is no dependency between the i-th component and the j-th component, and d<sub>ij</sub>= 1 indicates that if the i-th component is included, then the j-th component must also be included, d<sub>ij</sub>= 2 indicates that if the i-th component is included, the j-th component must be excluded, d<sub>ij</sub>= 3 indicates that if the i-th component is excluded, the j-th component must be included, d<sub>ij</sub>= 4 indicates that if the i-th component is excluded, the j-th component must also be excluded.
Figure 4A also shows a media description type (TYPE) field that can take one of three types defined by the value of the type field. Type = I (integration) refers to a parcel integrated with media description and data. Figure 5A shows the type = I format. Figure 5B shows a type = D (data only) format showing a data-only parcel with no description. Figure 5C shows a type = H (header only) format that points to a parcel that has only a description and no data.
The signature field (SIG. In Figures 5A-5C) uniquely identifies the parcel class (type) and follows the type field. The transcoder stores in internal memory all header information indexed by its signature as well as layer drop decisions made for the parcel for future reference. Once the signature is registered with the transcoder, a Type D parcel can be sent. In this case, the media description (header information) corresponding to the signature in the parcel is searched in the internal memory of the transcoder. The description and judgment information stored for each signature is updated each time a new parcel with the same signature (class) is routed. For Type I and H Purcells, the new media description in the current Purcell replaces the description stored inside the transcoder, and for Type I and D Purcells, with the code conversion decisions made for the current Purcell. , Replaces the decisions stored inside the transcoder for that class. The stored information makes it possible to use Type D parcels and maintain the consistency of code conversion described below.
For Type I and H parcels with header data, the signature field in the parcel header is followed by the specification of the number of media components, followed by the dependency data for the components called component dependencies, followed by the figure. Followed by a list of individual media component headers in the format shown in 3B. For Type I parcels, this parcel header is followed by a list of actual encoded scalable data for each component in the metabitstream format of Figure 2A. For Type H Purcell, the parcel ends at the end of the header. For Type D parcels, there is no header and only the list of each scalable data component in the format of Figure 2A is included.
Note that given the attributes of the individual components and their respective values, the attribute values for the entire parcel can be obtained. The parcel-wide attribute list contains a concatenation of all the attributes specified for all of its components. In addition, if the same attribute occurs in one or more components, the join type defined in the attribute join field of the attribute code (the "COMBINE" field in Figure 3B) determines the overall value. For example, if attribute join = addition, the total attribute value is the sum of the attribute values of the individual components, and if attribute join = maximum, the total attribute value is the largest of the individual component attribute values. .. The whole attribute value of the code-transformed parcel is used in the code conversion operation to determine which layer from which component to drop in order to satisfy what is imposed by the outbound constraint.
FIG. 6 shows a first embodiment of the method of delivering scalable encoded media data according to the present invention. According to this embodiment, the scalable encoded media source data is formatted into a format including the first and second parts shown in FIG. Specifically, the media data is in the first part, which corresponds to the non-media type scalability attributes and the data structure information in the second part, and the media type non-specific indexable data structure (Figures 2A and 2B). It is formatted to include a second part that corresponds to the arrayed scalable encoded media source data (60). In addition, information corresponding to the receive attributes of the media destination of any type of scalable encoded media is provided (61). Then, before being delivered to the media destination, the formatted scalable encoded media source data is code-transformed based on the matching of the scalability attribute and the receive attribute to obtain a scaled version of the formatted scalable encoded media source data. Generate and match the receive attributes of the media destination (62).
The receiving attributes of the receiving destination and any intermediate link (also called outbound constraints) are standardized so that they can be clearly communicated to the transcoder (similar to the scalability attributes contained within the media type non-unique format of the present invention). It should be understood that it is possible to compare the scalability attribute with the receive attribute. In one embodiment, the designation of a received attribute is based on a constraint on a definable multivariate function called the measure of the attribute. A definable measure is essentially a linear combination of the products of simple univariate functions of attribute values. According to an example of a multivariate function, the following is defined: That is, (i) the number of product terms in the combination N, and (ii) the number of elements in each product term n.<sub>i</sub>, (Iii) Attribute a in each product term<sub>ij</sub>Attribute code of, (iv) Function code of a specific simple univariate function for the attribute value f<sub>ij</sub>(.) And (v) Linear combination multiplier λ<sub>i</sub>Is. Given the definition parameters of the function, the measure can be expressed as shown in Equation 1.
<maths num="1"><img file="JP2004046879A_D0001.tif" /></maths>
In the formula, f<sub>ij</sub>(x) is the code corresponding to what should be included in the standard specification, x, x<sup>2</sup>, X<sup>-1</sup>, Log (x), e<sup>x</sup>It is a simple univariate function such as. Next, constraints are imposed on the measures defined above. Constraints can be of two types:
Limit constraints Outbound constraints most often consist of specific limits for an attribute measure called limit constraints. These constraints are specified as the maximum and / or minimum supportable values for the recipient of the measure. If both maximum and minimum are specified for the attribute measure, there is a range of supported values. For example, an example of a limit constraint is size / latency <300KB / s. Here, size is an attribute, but 1 / latency is specified as a multiplier in the outbound constraint. Overall, this indicates a bandwidth limitation on the receiving media by the receiving destination. Another example is a display resolution <800 diagonal pixels.
Optimization constraints It is also possible to specify constraints on the minimization or maximization of the required attribute measure. In this case, the description consists of whether it is desired to minimize or maximize the measure. The most important example of such a constraint occurs in rate distortion optimization where the measure is minimized, such as mean_squared_error + λ.size. Here, the size attribute corresponds to the rate (R), and the mean squared error attribute (mean_squared_error) corresponds to the distortion (D).
In general, code conversion (62 in Figure 6) can be performed as a simple delimiter of the bitstream subset, bitstream repacking and TOC updates as appropriate, depending on the comparison between the scalability and receive attributes. Also note that by arranging the scalable encoded media data in a data structure that is not unique to the media type, there is no need to decode or decode the content for code conversion. Subsets are dropped from the outer edge of each hierarchy (Figure 2A). Referencing the alternative representation shown in Figure 2B, the outer rows and columns are dropped.
In one embodiment, the code conversion is performed according to the method shown in FIG. As shown, it receives media data in a format that includes the first and second parts described above (70) and receives receive attributes (71). Compares scalability and receive attributes (72) and, depending on the comparison, aborts, drops, and repacks the bitstream subset, and a formatted scalable encoded media source that matches the media destination. Generate a scaled version of the data.
In an alternative embodiment, each receive attribute measure is compared to the first part of the formatted media data (eg, the media component description) to see if there is a corresponding scalability attribute. If one of the attributes is not in the description of any media component, then code conversion using this attribute is not possible, so the receive attribute measure is simply discarded as invalid.
For each valid receive attribute measure specified by the limit constraint (that is, there is a matching scalability attribute in the first part of the formatted media data), compare the full measure value of the entire packet to the limit constraint. , Check if it is within the limit constraint. The full measure value of the formatted media data is derived from the full attribute value of the formatted media data, and the formatted media data combines the attributes of the media components using the attribute join type field of the attribute code (Figure 3B). Obtained by If the full measure value does not exceed the outbound limit constraint, the formatted media data is transferred or transmitted without code conversion. If at least one of the complete measures is outside the limits constraint, the subset (ie, the outer row or column as shown in Figure 2B) is truncated, removed, or removed from one or more media components. Code conversion is performed by repacking.
It should be understood that the determination of which row or column to drop from which component can be made in a variety of ways, from simple methods to methods involving complex optimizations. For example, if the attribute monotonous type field contained in the component header indicates that the attribute is monotonous (non-decreasing or non-increasing), then a simple method of dropping rows or columns can be used. Alternatively, complex relationships can be created between the components to determine which subset to drop.
If an optimization constraint in the receive attribute is specified, its priority is lower than the limit constraint. From the options that do not violate the limit constraint, the transcoder selects the one that maximizes or minimizes the measure value. This is when choosing the best layer based on rate distortion criteria (ie traditional D + λR), or based on the relative preference of the user to choose one attribute over the other. It can be particularly beneficial when choosing.
In one embodiment, once a decision has been made as to which subset to drop from which component, the transcoder drops the subset in the bitstream before sending the code-converted media data, as appropriate. Update the TOC and break the attribute distribution matrix based on the dropped subset. If the data is of a multi-version type and is the last in the chain before the transcoder reaches the receiving destination, the code conversion operation extracts only the desired atom and discards the rest. including.
If multiple packets are directed to the same receiving destination, it may not be practical to include the media destination in each packet and expect the transcoder to drop layers accordingly. For example, if a consumer receives one presentation slide at a different resolution than the next slide, the media experience will be diminished. Therefore, according to an alternative embodiment, media descriptions common to classes of packets of the same type are usually used. In particular, during code conversion, media description data and code conversion decisions are stored for registered classes indexed by the identification signature (SIG. Field in Figure 4A). When first receiving formatted media data containing descriptive data about a class (Type I or Type H packets in Figures 5A-5C), an entry corresponding to a given signature is created in the buffer. If a given signature already exists in memory, it will be overwritten. Then, when a Type D packet belonging to the same class is sent, the description is checked, a layer drop decision is made, and the new decision is stored in the memory of that class, using only the signature instead of the media description. To. When a Type H packet is sent, the stored description for the class is simply updated. When a Type I packet is sent, the packet description in memory corresponding to a given signature is first updated, then a layer drop decision is made using the new descriptor, and finally a new decision is made. Stored in class memory. For classes of type D and type I packets, decisions are stored for future consistency.
Consistency means that the layer drop profile of each component remains unchanged from one packet to another for the list of hierarchies shown in the consistency list in the component header (Figure 3B). It is a constraint. In one embodiment, the consistency list includes subsets of all hierarchies. For a consistent hierarchy of components, the number of subsets dropped is the same as the judgment made for the previous packet, stored in memory for the class. This is an additional constraint that the subset drop determination mechanism can follow. At the code conversion decision stage, the hierarchy in the currently stored consistency list for the class is maintained the same as the pre-stored decision for the class. Therefore, for Type I packets, the new consistency list is used in place of the old list at the decision stage, based on the order of operations described above. This is because the description is updated before the decision is made, even if the decision of the previous formatted media data is used as a reference.
The consistency mechanism guarantees consistency when delivering media data that belongs to the same class, while allowing changes in layer drops in hierarchies that are not on the consistency list, thereby allowing changes in the same type of formatted media. Allows adaptation based on changes in the description of the data and changes in receive attributes (such as bandwidth).
According to an alternative embodiment of the method described above, each signature remains in the storage unit until dropped as a result of being unused. In one embodiment, the circular buffer maintains an ordered list of recently used signatures. If a particular signature has not been used for some time, it can eventually be replaced by a new signature.
FIG. 8 shows a first embodiment of a system of the invention that delivers scalable encoded media data including a media source 80, a transcoder 81, and a media destination 82. The media source 80 is a media type non-unique index having at least one dimension and a first part corresponding to the media type non-specific scalability attributes and the data structure information of the second part of the encoded media source data. It provides scalable encoded media data 80A in a format that includes a second portion corresponding to the scalable encoded media source data arranged in an attachable data structure. The media destination 82 provides information corresponding to the receive attribute 82A of the media destination of at least one type of scalable encoded media. The transcoder 81 codes the formatted scalable encoded media source data based on the matching of the scalability attribute and the reception attribute before delivering it to the media destination 82, and the formatted scalable encoded media source. Generate a scaled version of data 81A.
In general, the transcoder can connect directly to the media, in which case the media destination provides the receive attributes directly to the transcoder (or the transcoder detects the receive attributes) so that the transcoder scales the formatted data. It will be possible to provide a version. Alternatively, the transcoder can receive or detect the aggregation capabilities of all downstream media destinations. In this case, the scalable encoded media data is delivered to the media destination based on their respective integration capabilities. For example, FIG. 9 shows a network containing a plurality of transcoders that perform code conversion on each of the formatted media data according to the present invention and depend on the integrated reception attribute of the reception attribute (white arrow) of the downstream media destination. .. A single bitstream of formatted media data generated by transcoders 90 and 91 provides formatted media data that matches the receive attributes of both recipients 93 and 94, and transcoder 92 provides the recipients 93 and Note that it produces individual formatted media data bitstreams, each adapted to one of the 94 capabilities.
In one embodiment, the transcoder can be embodied in any one of a media server, a midstream router, or an edge server and can be implemented in any combination of hardware, software, and firmware. it can.
FIG. 10 shows an embodiment of the transcoder 100 that receives the formatted media data 100A and the media destination receive attribute 100B and generates a scaled version of the formatted media data 100C. The transcoder comprises a first parser 101 that receives and parses the first portion 20 (FIG. 1) of the formatted media data. The transcoder further comprises a second parser 102 that receives and parses the media destination receive attribute 100B. Parsers 101 and 102, respectively, parse the desired attribute data and information and provide it to the optimizer / determiner 103. The transcoder 100 further includes a first portion subtranscoder 104 and a second portion subtranscoder 105. The optimizer / determiner 103 gives control to both subtranscoders and gives the transcoder one of the code conversions (ie, bitstream subset truncation) for each of the first and second parts of the formatted media data. , Removal, repacking) to generate scaled formatted media data 100C, including scaled versions of the first and second parts respectively.
In this way, independent of media type and content type, scalable encoded media provides a universal delivery system, method, and format for all current types of media and any future media. Described the system, method, and format for delivering data.
In the above description, many specific details have been given for a complete understanding of the present invention. However, it will be apparent to those skilled in the art that it is not necessary to use these specific details in the practice of the present invention. Further, it should be understood that the particular embodiments described as examples should not be considered limiting. Reference to the details of the embodiments is not intended to limit the scope of the claims.
<figref num="1">It is a figure which shows the 1st Embodiment of the media type non-specific format of the scalable media data by this invention.</figref><figref num="2A">It is a figure which shows one Embodiment of the media data which was formatted in the data structure which is not peculiar to the media type which has the multilayer scalability by this invention.</figref><figref num="2B">It is a figure which shows the alternative representation of the scalable coded media data corresponding to the representation shown in FIG. 2A according to this invention.</figref><figref num="3A">FIG. 5 illustrates a second embodiment of a media type non-unique format for scalable media data including parcel components and parcel data information according to the present invention.</figref><figref num="3B">It is a figure which shows one Embodiment of the component header adopted in the media type non-specific format of the scalable media data of this invention.</figref><figref num="3C">It is a figure which shows the example of the attribute distribution specification adopted in the media type non-specific format of the scalable media data of this invention.</figref><figref num="3D">It is a figure which shows the example of the attribute distribution specification adopted in the media type non-specific format of the scalable media data of this invention.</figref><figref num="3E">It is a figure which shows the example of the attribute distribution specification adopted in the media type non-specific format of the scalable media data of this invention.</figref><figref num="3F">It is a figure which shows the example of the attribute distribution specification adopted in the media type non-specific format of the scalable media data of this invention.</figref><figref num="4A">It is a figure which shows the 3rd embodiment of the media type non-unique format which has the component dependency matrix D which defines the mode in which a component depends.</figref><figref num="4B">It is a figure which shows an example of the dependency matrix D by this invention.</figref><figref num="5A">It is a figure which shows the media type non-unique format which has a different type field.</figref><figref num="5B">It is a figure which shows the media type non-unique format which has a different type field.</figref><figref num="5C">It is a figure which shows the media type non-unique format which has a different type field.</figref><figref num="6">It is a figure which shows the 1st Embodiment of the method of delivering a scalable coded medium.</figref><figref num="7">It is a figure which shows the 1st Embodiment of the code conversion method.</figref><figref num="8">It is a figure which shows the 1st Embodiment of the system which distributes a scalable coded media.</figref><figref num="9">It is a figure which shows an example of the structure which carries out the system and method of this invention.</figref><figref num="10">It is a figure which shows the 1st Embodiment of the transcoder by this invention.</figref>
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US7930433B2 | Cited by | United States of America | Applicant |
| JP4849479B2 | Cited by | Japan | Examiner |
| JP2008535098A | Cited by | Japan | Search report |
| WO2006126260A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO02052730A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JP2000092485A | Cites | Japan | Examiner |
| JP2001117809A | Cites | Japan | Examiner |
| JP2002176359A | Cites | Japan | Examiner |
7 members in 3 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 10196506 | United States of America | – | |
| 19650602 | United States of America | A | |
| 19650602 | United States of America | A | |
| 2002196506 | – | – | – |
| US20020196506 | – | – | – |
Members7
| Document | Office | Kind | |
|---|---|---|---|
| US2004010614A1 | United States of America | A1 | |
| JP2004046879AThis record | Japan | A | |
| US2004139212A1 | United States of America | A1 | |
| WO2005055610A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US7133925B2 | United States of America | B2 | |
| JP4744794B2 | Japan | B2 | |
| US8244895B2 | United States of America | B2 |
20 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of acceptance of power of attorneyJAPANESE INTERMEDIATE CODE: A7422RD02 | RD02 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2004046879
- Publication, DOCDB
- 2004046879
- Publication, EPODOC
- JP2004046879
- Application
- 274622
- Application, DOCDB
- 2003274622
- Application, EPODOC
- JP20030274622
Titles2
- Japanese
- スケーラブルな符号化メディアを配信するシステム、方法及びフォーマット
- English
- Systems, methods and formats for delivering scalable coded media
Classification
- CPC, 9
- H04N7/165
- H04N21/234327
- H04N21/23439
- H04N21/25833
- H04N21/2662
- H04L65/612
- H04L65/70
- Y10S707/99942
- H04L65/1101
- IPC, 11
- G06F13 00
- H04L29 06
- H04N1 41
- H04N7 16
- H04N7 173
- H04N19 00
- H04N19 147
- H04N19 33
- H04N21 2343
- H04N21 258
- H04N21 845