Audio encoder and decoder
Abstract
Problem to be solved.To provide a method, an apparatus and a computer program product for encoding and decoding a vector of parameters in an audio coding system. The present disclosure further relates to methods and devices for reconstructing audio objects in audio decoding systems. According to the present disclosure, a modulo difference approach for coding and encoding aperiodic quantities of vectors can improve coding efficiency and provide encoders and decoders with reduced memory requirements. .. In addition, it provides an efficient way to encode and decode sparse matrices. [Selection diagram] Fig. 7

Term
10.4 yearsto projected expiry
Projected expiry 1 March 2037, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1オーディオ・エンコード・システムにおいてパラメータのベクトルをエンコードする方法であって、各パラメータは非周期的な量に対応し、前記ベクトルは、第一の要素および少なくとも一つの第二の要素をもち、当該方法は:N通りの値を取り得るインデックス値によって前記ベクトル中の各パラメータを表現する段階と;前記少なくとも一つの第二の要素のそれぞれをシンボルに関連付ける段階であって、前記シンボルは: 前記第二の要素のインデックス値と前記ベクトル中でその先行する要素のインデックス値との間の差を計算し;該差にモジュロNを適用することによって計算される、段階と;前記少なくとも一つの第二の要素に関連付けられた前記シンボルを、シンボルの確率を含む確率テーブルに基づいてエントロピー符号化することによって、前記少なくとも一つの第二の要素のそれぞれをエンコードする段階とを含む、方法。
56 paragraphs, as filed
0001Cross-reference to related applications This application claims the benefit of the filing date of US Provisional Patent Application No. 61 / 827,264 filed on May 24, 2013. The content of the application is incorporated herein by reference.
0002Technical field The disclosure of this paper generally relates to audio coding. More specifically, it relates to the encoding and decoding of a vector of parameters in an audio coding system. The present disclosure further relates to methods and devices for reconstructing audio objects in audio decoding systems.
0003In a typical audio system, a channel-based approach is used. Each channel may represent, for example, the content of one speaker or one speaker array. Possible coding schemes for such systems include discrete multi-channel coding or parametric coding such as MPEG surround.
0004More recently, new approaches have been developed. This approach is object-based. In systems that use an object-based approach, a 3D audio scene is represented by an audio object with associated position metadata. These audio objects move around in the 3D audio scene during playback of the audio signal. The system may further include so-called bed channels. Bed channels may be described as static audio objects that map directly to speaker locations in a typical audio system, for example as described above.
0005A possible problem in object-based audio systems is how to efficiently encode and decode audio signals and maintain the quality of the encoded signals. One possible coding scheme allows the encoder side to regenerate the downmix signal containing several channels from the audio object and bed channel, and the decoder side to regenerate the audio object and bed channel. Includes generating side information and.
0006MPEG Spatial Audio Object Coding (MPEG SAOC) describes a system for parametric coding of audio objects. The system sends side information, upmix matrix references, that describe the attributes of the object, with parameters such as object level differences and cross-correlation. These parameters are then used on the decoder side to control the regeneration of the audio object. This process is mathematically complex and often requires reliance on assumptions about the attributes of audio objects that are not explicitly described by the parameters. The methods presented in MPEG SAOC can reduce the required bitrate for object-based audio systems, but may require further improvements to further increase efficiency and quality as described above. is there.
0007An exemplary embodiment will be described below with reference to the accompanying drawings.<figref num="1">FIG. 6 is a generalized block diagram of an audio encoding system based on an exemplary embodiment.</figref><figref num="2">It is a generalized block diagram of the exemplary upmix matrix encoder shown in FIG.</figref><figref num="3">FIG. 5 shows an exemplary probability distribution for the first element in the vector of parameters corresponding to the elements in the upmix matrix determined by the audio encoding system of FIG.</figref><figref num="4">FIG. 5 shows an exemplary probability distribution for at least one modulo differentially encoded second element in a vector of parameters corresponding to the elements in the upmix matrix determined by the audio encoding system of FIG. ..</figref><figref num="5">FIG. 6 is a generalized block diagram of an audio decoding system based on an exemplary embodiment.</figref><figref num="6">It is a generalized block diagram of the upmix matrix decoder shown in FIG.</figref><figref num="7">It is a figure which shows the encoding method about the said 2nd element in the vector of the parameter corresponding to the element in the upmix matrix determined by the audio encoding system of FIG.</figref><figref num="8">It is a figure which shows the encoding method about the 1st element in the vector of the parameter corresponding to the element in the upmix matrix determined by the audio encoding system of FIG.</figref><figref num="9">It is a figure which shows the parts of the encoding method of FIG. 7 about the said 2nd element in the vector of an exemplary parameter.</figref><figref num="10">It is a figure which shows various parts of the encoding method of FIG. 8 about the said 1st element in a vector of exemplary parameters.</figref><figref num="11">It is a generalized block diagram of the second exemplary upmix matrix encoder shown in FIG.</figref><figref num="12">FIG. 6 is a generalized block diagram of an audio decoding system based on an exemplary embodiment.</figref><figref num="13">It is a figure which shows the encoding method for sparse encoding of the row of an upmix matrix.</figref><figref num="14">It is a figure which shows the parts of the encoding method of FIG. 10 about the exemplary row of an upmix matrix.</figref><figref num="15">It is a figure which shows the parts of the encoding method of FIG. 10 about the exemplary row of an upmix matrix. All drawings are schematic and generally only show the parts necessary to clarify this disclosure. On the other hand, other parts may be omitted or only suggested. Unless otherwise noted, similar reference numerals refer to similar parts in different drawings.</figref>
0008In view of the above, it is an object of the present invention to provide encoders and decoders and related methods that provide increased efficiency and quality of encoded audio signals.
0009<u style="single"><I. Overview-Encoder></u> According to the first aspect, exemplary embodiments propose encoding methods, encoders and computer program products for encoding. The proposed methods, encoders and computer program products may generally have the same features and advantages.
0010An exemplary embodiment provides a method of encoding a vector of parameters in an audio encoding system. Each parameter corresponds to an aperiodic quantity. A vector has a first element and at least one second element. The method includes: representing each parameter in the vector by an index value that can take N possible values; and associating each of the at least one second element with a symbol. Calculate the difference between the index value of the second element and the index value of its preceding element in the vector; calculated by applying the modulo N to the difference. The method further encodes each of the at least one second element by entropy encoding the symbols associated with the at least one second element based on a probability table containing the probabilities of the symbols. Including stages.
0011The advantage of this method is that the number of possible symbols is reduced by about half compared to the usual differential coding strategy where the modulo N is not applied to the difference. As a result, the size of the probability table is reduced by about half. As a result, the encoder can be made cheaper in this way, as less memory is required to store the probability table and the probability table is often stored in expensive memory in the encoder. In addition, the speed of searching for symbols in the probability table can be increased. A further advantage is that the coding efficiency can be increased because all the symbols in the probability table are possible candidates to be associated with a particular second element. This can be compared to the usual differential coding strategies where only about half of the symbols in the probability table are candidates for being associated with a particular second element.
0012According to embodiments, the method further comprises associating the first element in the vector with a symbol. The symbol is calculated by: shifting the index value representing the first element in the vector by an offset value; applying modulo N to the shifted index value. The method further encodes the first element by entropy encoding the symbol associated with the first element using the same probability table used to encode the at least one second element. Including stages.
0013This embodiment is similar to the fact that the probability distribution of the index value of the first element and the probability distribution of the symbols of the at least one second element are similar, although they are shifted relative to each other by some offset value. use. As a result, instead of a dedicated probability table, the same probability table can be used for the first element in the vector. As a result, as mentioned above, it can lead to reduced memory requirements and cheaper encoders.
0014According to one embodiment, the offset value is equal to the difference between the most probable index value for the first element and the most probable symbol for the at least one second element in the probability table. .. This means that the peaks of those probability distributions are aligned. As a result, substantially the same coding efficiency is maintained for the first element as compared to the case where a dedicated probability table is used for the first element.
0015According to embodiments, the first element and the at least one second element of the vector of the parameters correspond to different frequency bands used in the audio encoding system in a particular time frame. That is, data corresponding to a plurality of frequency bands can be encoded with the same operation. For example, the vector of the parameters may correspond to upmixes or reconstruction coefficients that vary over multiple frequency bands.
0016According to certain embodiments, the first element and the at least one second element of the vector of the parameters correspond to different time frames used in the audio encoding system in a particular frequency band. That is, data corresponding to a plurality of time frames can be encoded with the same operation. For example, the vector of the parameters may correspond to upmixes or reconstruction coefficients that vary over multiple time frames.
0017According to embodiments, the probability table is translated into a Huffman codebook. Here, the symbol associated with an element in the vector is used as a codebook index, and the encoding step is to place the second element by the codebook index associated with the second element. It involves encoding each of the at least one second element by representing it in codewords in the indexed codebook. By using the symbol as a codebook index, the speed of searching for codewords representing the element can be improved.
0018According to embodiments, the encoding step represents the first element by a codeword in the Huffman codebook indexed by a codebook index associated with the first element. It involves encoding the first element in the vector using the same Huffman codebook used to encode the at least one second element. As a result, only one Huffman codebook needs to be stored in the encoder's memory, which can lead to cheaper encoders as described above.
0019According to one further embodiment, the vector of the parameters corresponds to the elements in the upmix matrix determined by the audio encoding system. This can reduce the bit rate required in audio encoding / decoding systems because the upmix matrix can be coded efficiently.
0020According to an exemplary embodiment, a computer-readable medium having computer code instructions adapted to perform any method of the first aspect when executed on a device having processing capabilities is provided.
0021According to an exemplary embodiment, an encoder is provided that encodes a vector of parameters in an audio encoding system. Each parameter corresponds to an aperiodic quantity. A vector has a first element and at least one second element. The encoder: The receiving component adapted to receive the vector; the indexing component adapted to represent each parameter in the vector by an index value that can take N possible values; the at least one first. It has an association component adapted to associate each of the two elements with a symbol. The symbol: Calculates the difference between the index value of the second element and the index value of its preceding element in the vector; it is calculated by applying modulo N to the difference. The encoder further encodes each of the at least one second element by entropy encoding the symbols associated with the at least one second element based on a probability table containing the probabilities of the symbols. It has an encoding component.
0022<u style="single"><II. Overview-Decoder></u> According to the second aspect, exemplary embodiments propose decoding methods, decoders and computer program products for decoding. The proposed methods, decoders and computer program products may generally have the same features and advantages.
0023The features and setup benefits presented in the above encoder overview may also generally be valid for the corresponding features and setup for the decoder.
0024An exemplary embodiment provides a method of decoding a vector of entropy-encoded symbols in an audio decoding system into a vector of parameters related to aperiodic quantities. The vector of entropy-encoded symbols has a first entropy-encoded symbol and at least one second entropy-encoded symbol, and the vector of the parameters is the first element and at least the second element. Have. The method: By using a probability table, a step of expressing each entropy-encoded symbol in the vector of the entropy-encoded symbol by a symbol that can take N integer values; the first entropy The step of associating the encoded symbol with the index value; the step of associating each of the at least one second entropy-encoded symbol with the index value, and the step of associating the at least one second entropy-encoded symbol with the index value. The index values of the symbols are: the index value associated with the entropy-encoded symbol preceding the second entropy-encoded symbol in the vector of the entropy-encoded symbol, and the second entropy-encoding. Calculate the sum with the symbol representing the symbol; calculated by applying the modulo N to the sum. The method further comprises expressing the at least one second element of the vector of the parameter by a parameter value corresponding to the index value associated with the at least one second entropy-encoded symbol. ..
0025According to an exemplary embodiment, the step of representing each entropy-encoded symbol in the vector of the entropy-encoded symbol by the symbol is all entropy-encoded in the vector of the entropy-encoded symbol. It is executed using the same probability table for the symbols. The index value associated with the first entropy-encoded symbol is: Shift the symbol representing the first entropy-encoded symbol in the vector of the entropy-encoded symbol by an offset value; Calculated by applying Modulo N to the shifted symbols. The method further comprises expressing the first element of the vector of the parameters with a parameter value corresponding to the index value associated with the first entropy-encoded symbol.
0026According to one embodiment, the probability table is translated into a Huffman codebook, and each entropy-coded symbol corresponds to a codeword in the Huffman codebook.
0027According to a further embodiment, each codeword in the Huffman codebook is associated with a codebook index, and the symbol represents each entropy-encoded symbol in said vector of the entropy-encoded symbol. It comprises representing the entropy-encoded symbol by the codebook index associated with the codeword corresponding to the entropy-encoded symbol.
0028According to embodiments, each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different frequency band used in the audio decoding system at a particular time frame.
0029According to one embodiment, each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different time frame used in the audio decoding system in a particular frequency band.
0030According to embodiments, the vector of the parameters corresponds to an element in the upmix matrix used by the audio decoding system.
0031According to an exemplary embodiment, a computer-readable medium having computer code instructions adapted to perform any method of the second aspect when executed on a device having processing capabilities is provided.
0032According to an exemplary embodiment, a decoder is provided that decodes a vector of entropy-encoded symbols in an audio decoding system into a vector of parameters related to aperiodic quantities. The vector of entropy-encoded symbols has a first entropy-encoded symbol and at least one second entropy-encoded symbol, and the vector of the parameters is the first element and at least the second element. Have. The decoder: With a receiving component configured to receive a vector of entropy-encoded symbols; said vector of entropy-encoded symbols by a symbol that can take N integer values by using a probability table. Includes an indexing component configured to represent each entropy-encoded symbol in; and an association component configured to associate the first entropy-encoded symbol with an index value; Each of the at least one second entropy-encoded symbol is further configured to associate with an index value, and the index value of the at least one second entropy-encoded symbol is: entropy-encoded. Calculate the sum of the index value of the entropy-encoded symbol preceding the second entropy-encoded symbol in the vector of the symbol and the symbol representing the second entropy-encoded symbol; Calculated by applying the modulo N to the sum. The decoder is further configured to represent said at least one second element of the vector of said parameter by a parameter value corresponding to the index value associated with said said at least one second entropy encoded symbol. Has a decoding component.
0033<u style="single"><III. Overview-Sparse Matrix Encoder></u> According to a third aspect, exemplary embodiments propose encoding methods, encoders and computer program products for encoding. The proposed methods, encoders and computer program products may generally have the same features and advantages.
0034An exemplary embodiment provides a method of encoding an upmix matrix in an audio encoding system. Each row of the upmix matrix contains M elements that allow the reconstruction of the time / frequency tiles of the audio object from the downmix signal containing M channels. The method is for each row of the upmix matrix: select a subset of the elements from the M elements in that row of the upmix matrix; each element in the selected subset of the elements, the value and the position in the upmix matrix. Represented by; involves encoding the value and position of each element in the upmix matrix in the selected subset of elements.
0035In the usage herein, by the term downmix signal containing M channels, a signal containing M signals or channels, each channel containing a plurality of audios including said audio object to be reconstructed. -It means a combination of objects. The number of channels is typically greater than 1, often greater than or equal to 5.
0036As used herein, the term upmix matrix refers to a matrix with N rows and M columns that allows N audio objects to be reconstructed from a downmix signal containing M channels. The elements in each row of the upmix matrix correspond to an audio object and give the coefficients to be multiplied by the M channels of the downmix to reconstruct the audio object.
0037As used herein, position in an upmix matrix means a row and column index that points to the rows and columns of a matrix element. The term position can also mean a column index in a given row of an upmix matrix.
0038In some cases, sending all the elements of the upmix matrix for each time / frequency tile requires an undesirably high bit rate in an audio encoding / decoding system. The advantage of this method is that only a subset of the upmix matrix elements need to be encoded and transmitted to the decoder. Since less data is transmitted, it may reduce the required bit rate of the audio encoding / decoding system and the data can be encoded more efficiently.
0039Audio encoding / decoding systems typically divide the time-frequency space into time / frequency tiles, for example by applying a suitable filter bank to the input audio signal. A time / frequency tile generally means a portion of the time-frequency space that corresponds to a time interval and frequency subband. The time interval typically corresponds to the duration of the time frame used in the audio encoding / decoding system. Frequency subbands typically correspond to one or several neighboring frequency subbands defined by the filter bank used in the encoding / decoding system. If the frequency subband corresponds to several neighboring frequency subbands defined by the filter bank, this is a non-uniform frequency subband in the audio signal decoding process, eg, the higher frequency of the audio signal. Allows to have a wider frequency subband. For broadband where the audio encoding / decoding system acts over the entire frequency range, the frequency subbands of the time / frequency tile may cover the entire frequency range. The above method discloses various encoding steps for encoding an upmix matrix in an audio encoding system to allow the reconstruction of audio objects between one such time / frequency tile. .. However, it is understood that the method may be repeated for each time / frequency tile of the audio encoding system. It is also understood that several time / frequency tiles may be encoded at the same time. Typically, neighboring time / frequency tiles may overlap slightly in time and / or frequency. For example, overlap in time can be equivalent to temporal interpolation of the elements of the reconstructed matrix, i.e., linear interpolation from one time interval to the next. However, this disclosure also targets other parts of the encoding / decoding system.
0040According to embodiments, for each row in the upmix matrix, the position of the selected subset of elements in the upmix matrix varies across multiple frequency bands and / or across multiple time frames. .. Therefore, the selection of those elements may depend on a particular time / frequency tile, and thus different elements may be selected for different time / frequency tiles. This provides a more flexible encoding method, which enhances the quality of the encoded signal.
0041According to embodiments, the selected subset of elements contains the same number of elements for each row of the upmix matrix. In a further embodiment, the number of elements selected may be exactly one. This adds to the complexity of the encoder, as the algorithm only has to select the same number of elements (s) for each row, that is, the most important elements (s) when performing an upmix on the decoder side. Reduce.
0042According to embodiments, for each row in the upmix matrix and for multiple frequency bands or multiple time frames, the element values of the selected subset of elements form one or more vectors of parameters. Each parameter in the vector of parameters corresponds to one of the plurality of frequency bands or the plurality of time frames, and the vector of the parameters is encoded using a method based on the first aspect. To. In other words, the values of the selected elements can be efficiently coded. The features and setup advantages presented in the first aspect of the above overview may generally be valid for this embodiment as well.
0043According to embodiments, for each row in the upmix matrix and for multiple frequency bands or multiple time frames, the position of the elements in the selected subset of elements forms one or more vectors of the parameters. Each parameter in the vector of parameters corresponds to one of the plurality of frequency bands or the plurality of time frames, and the vector of the parameters is encoded using a method based on the first aspect. To. In other words, the positions of the selected elements can be efficiently coded. The features and setup advantages presented in the first aspect of the above overview may generally be valid for this embodiment as well.
0044According to an exemplary embodiment, a computer-readable medium having computer code instructions adapted to perform any method of the third aspect when executed on a device having processing capabilities is provided.
0045According to an exemplary embodiment, an encoder that encodes an upmix matrix in an audio encoding system is provided. Each row of the upmix matrix contains M elements that allow the reconstruction of the time / frequency tiles of the audio object from the downmix signal containing M channels. The encoder: With the receiving component adapted to receive each row in the upmix matrix; with the selection component adapted to select a subset of the elements from the M elements of the row in the upmix matrix; It has an encode component adapted to represent each element in the subset of the elements by value and position in the upmix matrix, said encode component further of each element in the selected subset of elements. Adapted to encode values and positions in upmix matrices.
0046<u style="single"><IV. Overview-Sparse Matrix Decoder></u> According to a fourth aspect, exemplary embodiments propose decoding methods, decoders and computer program products for decoding. The proposed methods, decoders and computer program products may generally have the same features and advantages.
0047The features and setup benefits presented in the sparse matrix encoder overview above may generally also be valid for the corresponding features and setup for the decoder.
0048An exemplary embodiment provides a method of reconstructing the time / frequency tiles of an audio object in an audio decoding system. The method is: receiving a downmix signal containing M channels; and receiving at least one encoded element representing a subset of the M elements of a row in the upmix matrix. Each encoded element contains a value and a position in that row in the upmix matrix, the position indicating one of the M channels of the downmix signal to which the encoded element corresponds. , And; reconstructing the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded element. Including. In the linear combination, each downmix channel is multiplied by the value of its corresponding encoded element.
0049Thus, according to this method, the time / frequency tiles of the audio object are reconstructed by forming a linear combination of subsets of the downmix channel. The subset of downmix channels corresponds to the channel from which the upmix coefficients encoded for it have been received. Thus, the method allows the audio object to be reconstructed despite the fact that only a subset of the upmix matrix, eg, a sparse subset, is received. By forming a linear combination of only the downmix channels corresponding to the at least one encoded element, the complexity of the decoding process can be reduced. An alternative would be to form a linear combination of all the downmix signals and then multiply some of them (those that do not correspond to at least one encoded element) by the value 0.
0050According to embodiments, the position of the at least one encoded element varies across multiple frequency bands and / or across multiple time frames. Thus, in other words, different elements of the upmix matrix may be encoded for different time / frequency tiles.
0051According to embodiments, the number of elements in at least one encoded element is equal to one. That is, the audio object is reconstructed from one downmix channel at each time / frequency tile. However, that one downmix channel used to reconstruct an audio object can vary between different time / frequency tiles.
0052According to embodiments, for multiple frequency bands or multiple time frames, the values of the at least one encoded element form one or more vectors, each value represented by an entropy-encoded symbol. Each symbol in each vector of entropy-encoded symbols corresponds to one of the plurality of frequency bands or one of the plurality of time frames, and the one or more of the entropy-encoded symbols. The vector of is decoded using a method based on the second aspect. In this way, the values of the elements of the upmix matrix can be efficiently coded.
0053According to embodiments, for multiple frequency bands or multiple time frames, the positions of the at least one encoded element form one or more vectors, each position represented by an entropy-encoded symbol. Each symbol in each vector of entropy-encoded symbols corresponds to one of the plurality of frequency bands or the plurality of time frames, and the one or more vectors of the entropy-encoded symbol correspond to the plurality of frequency bands or the plurality of time frames. , Decoded using a method based on the second aspect. In this way, the positions of the elements of the upmix matrix can be efficiently coded.
0054According to an exemplary embodiment, a computer-readable medium having computer code instructions adapted to perform any method of the third aspect when executed on a device having processing capabilities is provided.
0055According to an exemplary embodiment, a decoder is provided that reconstructs the time / frequency tiles of an audio object. The decoder is a receiving component configured to receive at least one encoded element that represents a subset of the M elements of a row in a downmix signal containing M channels and an upmix matrix. Each encoded element contains a value and a position in that row in the upmix matrix, the position indicating one of the M channels of the downmix signal to which the encoded element corresponds. Configured to reconstruct the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded component. It has a reconstructed component. In the linear combination, each downmix channel is multiplied by the value of its corresponding encoded element.
<p num="0056"><u style="single"><V. Illustrative Embodiment></u> FIG. 1 shows a generalized block diagram of an audio encoding system 100 for encoding an audio object 104. The audio encoding system has a downmix component 106 that produces a downmix signal 110 from various audio objects 104. The downmix signal 110 may be, for example, a 5.1 or 7.1 surround signal backwards compatible with a Dolby Digital Plus or MPEG standard, such as an established sound decoding system such as AAC, USAC or MP3. In a further embodiment, the downmix signal is not backward compatible.</p><p num="0057"> The upmix parameters are determined from the downmix signal 110 and the audio object 104 in the upmix parameter analysis component 112 so that the audio object 104 can be reconstructed from the downmix signal 110. For example, the upmix parameters may correspond to elements of the upmix matrix that allow the reconstruction of the audio object 104 from the downmix signal 110. The upmix parameter analysis component 112 processes the downmix signal 110 and the audio object 104 for individual time / frequency tiles. Thus, the upmix parameters are determined for each time / frequency tile. For example, an upmix matrix for each time / frequency tile may be determined. For example, the upmix parameter analysis component 112 is a quadrature mirror filter (QMF:) that allows frequency-selective processing. It may operate in a frequency domain such as the Quadrature Mirror Filters) region. For this reason, the downmix signal 110 and the audio object 104 may be converted into the frequency domain by applying the downmix signal 110 and the audio object 104 to the filter bank 108. This may be done, for example, by applying a QMF transformation or any other suitable transformation.</p><p num="0058"> The upmix parameter 114 may be organized in vector format. The vector may represent an upmix parameter for reconstructing a particular audio object from the audio object 104 in different frequency bands in a particular time frame. For example, a vector may correspond to a matrix element in an upmix matrix. Here, the vector contains the values of the matrix elements for a series of frequency bands. In a further embodiment, the vector may represent an upmix parameter for reconstructing a particular audio object from the audio object 104 in various time frames in a particular frequency band. For example, the vector may correspond to a matrix element of an upmix matrix, which contains the values of the matrix element for a series of time frames, but in the same frequency band.</p><p num="0059"> Each parameter in the vector corresponds to an aperiodic quantity, eg, an quantity that takes a value between -9.6 and 9.4. A non-periodic quantity generally means a non-periodic quantity in a value that the quantity can take. This is in contrast to periodic quantities such as angles where there is a clear periodic correspondence between the possible values of that quantity. For example, for angles, there is a periodicity of 2π, for example, angle 0 corresponds to angle 2π.</p><p num="0060"> The upmix parameter 114 is then received by the upmix matrix encoder 102 in vector format. The upmix matrix encoder will be described in detail here in the context of FIG. The vector is received by the receiving component 202 and has a first element and at least one second element. The number of elements depends, for example, on the number of frequency bands in the audio signal. The number of elements may depend on the number of time frames of the audio signal encoded in one encoding operation.</p><p num="0061"> The vector is then indexed by the indexing component 204. The indexing component is adapted to represent each parameter in the vector by an index value that can take a predefined number of values. This expression can be made in two steps. First, the parameters are quantized, then the quantized values are indexed by the index value. As an example, if each parameter in the vector can take a value between -9.6 and 9.4, this can be done by using a quantization step of 0.2. The quantized values may then be indexed by indexes 0-95, ie 96 different values. In the example below, the index value is in the range 0-95, but this is of course just an example, and other ranges of index values, such as 0-191 and 0-63, are equally possible. Smaller quantization steps can produce a less distorted decoded audio signal on the decoder side, but at a higher bit rate required for data transmission between the audio encoding system 100 and the decoder. Can also occur.</p><p num="0062"> The indexed value is then sent to the association component 206. The association component 206 uses a modulo differential encoding strategy to associate each of the at least one second element with a symbol. The association component 206 is adapted to calculate the difference between the index value of the second element and the index value of the previous element in the vector. The difference can be anywhere in the range -95 to 95, simply by using the usual diff encoding strategy. That is, there are 191 possible values. This means that when the difference is encoded using entropy encoding, a probability table containing 191 probabilities is needed. That is, there is one probability for each of the 191 possible values for the difference. Moreover, for each difference, about half of the 191 probabilities are impossible, which reduces the efficiency of encoding. For example, if the second element to be diff-encoded has an index value of 90, the possible difference is in the range -5 to +90. Typically, having an entropy encoding strategy in which some of the probabilities are not possible for each value to be encoded reduces the efficiency of encoding. The differential coding strategy in the present disclosure overcomes this problem by applying a modulo 96 operation to the difference, while at the same time reducing the number of codes required to 96. Therefore, the association algorithm can be expressed as follows.</p><p num="0063"> Δ<sub>idx</sub>(b) = (idx (b)-idx (b-1)) mod N<sub>Q</sub> (Equation 1) Where b is an element in the differentially encoded vector, N<sub>Q</sub>Is the number of possible index values, Δ<sub>idx</sub>(b) is the symbol associated with element b.</p><p num="0064"> According to some embodiments, the probability table is converted into a Huffman codebook. In this case, the symbol associated with an element in the vector is used as the codebook index. Encoding component 208 then represents at least one of the second elements by representing the second element with codewords in the Huffman codebook indexed by the codebook index associated with the second element. Each of the two second elements can be encoded.</p><p num="0065"> Any other suitable entropy coding strategy may be implemented by encoding component 208. For example, such an encoding strategy may be a range coding strategy or an arithmetic coding strategy.</p><p num="0066"> Below, we show that the entropy of the modulo approach is always less than or equal to the entropy of the usual differential approach. Entropy E of the usual difference approach<sub>p</sub>Is<maths num="1"><img id="000003" he="19" wi="123" file="JP2017102484A_D0001.tif" img-format="tif" img-content="drawing" /></maths>Is. Here, p (n) is the probability of a simple difference index value n.</p><p num="0067"> Modulo approach entropy E<sub>q</sub>Is<maths num="2"><img id="000004" he="18" wi="123" file="JP2017102484A_D0001.tif" img-format="tif" img-content="drawing" /></maths>Is. Where q (n) is the probability of the modulo difference index value n, q (0) = p (0) (Equation 4) q (n) = p (n) + p (nN)<sub>Q</sub>) n = 1 ... N<sub>Q</sub>-1 (Equation 5) Given by.</p><p num="0068"> Therefore, it becomes as follows.<maths num="3"><img id="000005" he="22" wi="153" file="JP2017102484A_D0001.tif" img-format="tif" img-content="drawing" /></maths> In the last sum n = jN<sub>Q</sub>Substituting for, we get:</p><p num="0069"><maths num="4"><img id="000006" he="55" wi="166" file="JP2017102484A_D0001.tif" img-format="tif" img-content="drawing" /></maths> Comparing the sum by term<maths num="5"><img id="000007" he="33" wi="123" file="JP2017102484A_D0001.tif" img-format="tif" img-content="drawing" /></maths>So E<sub>p</sub> E<sub>q</sub>Will be.</p><p num="0070"> As shown above, the entropy for the modulo approach is always less than or equal to the entropy for the usual differential approach. If the entropies are equal, it is a rare case where the encoded data is pathological, that is, poorly behaved data, which is often not the case, for example, in upmix matrices.</p><p num="0071"> Since the entropy for the modular approach is always less than or equal to the entropy of the normal differential approach, the entropy coding of the symbols calculated by the modular approach is compared to the entropy coding of the symbols calculated by the regular differential approach. , Lower or at least the same bit rate. In other words, the entropy coding of symbols calculated by the modulo approach is often more efficient than the entropy coding of symbols calculated by the usual differential approach.</p><p num="0072"> A further advantage is that, as mentioned above, the number of probabilities required in the probability table in the modulo approach is approximately half the number of probabilities required in the conventional non-modulo approach.</p><p num="0073"> In the above, a modulo approach for encoding the at least one second element in a vector of parameters has been described. The first element may be encoded with an index value representing the first element. Since the probability distributions of the index value of the first element and the modulo difference value of at least one of the second elements can be very different (see Figure 3 for the probability distribution of the indexed first element, said. Modulo differential values, i.e. the probability distribution of symbols for at least one second element (see Figure 4), may require a dedicated probability table for the first element. This requires that both the audio encoding system 100 and the corresponding decoder have such a dedicated probability table in memory.</p><p num="0074"> However, we have observed that in some cases the shapes of the probability distributions are very similar, albeit shifted relative to each other. This observation can be used to approximate the probability distribution of the indexed first element by a shifted version of the probability distribution of the symbols for at least one second element. In such a shift, the association component 206 associates the first element in the vector with a symbol by shifting the index value representing the first element in the vector by an offset value, and then the shifted index. It may be implemented by adapting the value to apply Modulo 96 (or the corresponding value).</p><p num="0075"> Therefore, the calculation of the symbol associated with the first element is idx<sub>excessive</sub>(1) = (idx (1) -abs_offset) mod N<sub>Q</sub> (Equation 11) May be expressed as.</p><p num="0076"> The symbol thus achieved is used by encoding component 208. The encoding component 208 entropy encodes the symbol associated with the first element using the same probability table used to encode the at least one second element. Encode one element. The offset value may be equal to or at least close to the difference between the most probable index value for the first element and the most probable symbol for at least one second element in the probability table. In FIG. 3, the most probable index value for the first element is represented by arrow 302. Assuming that the most probable symbol for at least one second element is 0, the value represented by arrow 302 is the offset value used. By using the offset approach, the peaks in the distributions in Figures 3 and 4 are aligned. This approach avoids the need for a dedicated probability table for the first element, thus saving memory in the audio encoding system 100 and the corresponding decoder. On the other hand, it often maintains almost the same coding efficiency as a dedicated probability table gives.</p><p num="0077"> If the entropy coding of the at least one second element is done using the Huffman codebook, the encoding component 208 encodes the first element in the vector and the at least one second element. It may be encoded using the same Huffman codebook used for. It is by representing the first element with codewords in the Huffman codebook indexed by the codebook index associated with the first element.</p><p num="0078"> The memory in which the codebook is stored is advantageously fast memory and therefore expensive, as search speed can be important when encoding parameters in an audio decoding system. Thus, by using only one probability table, the encoder can be cheaper than if two probability tables were used.</p><p num="0079"> It should be noted that the probability distributions shown in Figures 3 and 4 are often pre-computed for the training dataset and thus not during vector encoding. But, of course, it is also possible to calculate the distribution "on the fly" during encoding.</p><p num="0080"> It should be noted that the above description of the audio encoding system 100, which uses the vector from the upmix matrix as the vector of the parameters to be encoded, is for illustrative purposes only. The method of encoding a vector of parameters based on the present disclosure may be used in other applications in an audio encoding system. For example, when encoding other internal parameters in a downmix encoding system, such as the parameters used in a parametric bandwidth expansion system such as spectral band replication (SBR).</p><p num="0081"> FIG. 5 is a generalized block diagram of an audio decoding system 500 for regenerating an audio object encoded from an encoded downmix signal 510 and an encoded upmix matrix 512. The encoded downmix signal 510 is received by the downmix receiving component 506, where the signal is decoded and converted into the preferred frequency domain, unless it is already in the preferred frequency domain. The decoded downmix signal 516 is then sent to the upmix component 508. The upmix component 508 uses the decoded downmix signal 516 and the decoded upmix matrix 504 to regenerate the encoded audio object. More specifically, the upmix component 508 may perform a matrix operation in which the decoded upmix matrix 504 is multiplied by a vector containing the decoded downmix signal 516. The decoding process of the upmix matrix is described below. The audio decoding system 500 also has a rendering component 514 that outputs an audio signal based on the reconstructed audio object 518, depending on the type of playback unit connected to the audio decoding system 500. ..</p><p num="0082"> The encoded upmix matrix 512 is received by the upmix matrix decoder 502. This upmix matrix decoder 502 will be described in detail here in relation to FIG. The upmix matrix decoder 502 is configured in an audio decoding system to decode a vector of entropy-encoded symbols into a vector of parameters related to aperiodic quantities. The vector of entropy-encoded symbols contains a first entropy-encoded symbol and at least one second entropy-encoded symbol, and the vector of parameters contains the first element and at least the second element. Including. Thus, the encoded upmix matrix 512 is received by the receiving component 602 in vector format. The decoder 502 further has an indexing component 604 configured to represent each entropy-encoded symbol in the vector by a symbol that can take N possible values by using a probability table. N may be 96, for example. The association component 606 converts the first entropy-encoded symbol into an index value by any suitable means, depending on the encoding method used to encode the first element in the parameter vector. It is configured to be associated. The symbol for each of the second signs and the index value for the first sign are then used by the association component 606. The association component 606 associates each of the at least one second entropy-encoded symbol with the index value. The index value of the at least one entropy-encoded symbol is first the index associated with the entropy-encoded symbol preceding the second entropy-encoded symbol in the vector of the entropy-encoded symbol. The value and the second entropy -Calculated by calculating the sum with the symbol representing the coded symbol. Then modulo N is applied to the sum. Without loss of generality, suppose the minimum index value is 0 and the maximum index value is N-1, for example 95. Then the association algorithm is: idx (b) = (idx (b-1) + Δ<sub>idx</sub>(b)) mod N<sub>Q</sub> (Equation 12) May be expressed as. Where b is the element in the vector being decoded and N<sub>Q</sub>Is the number of possible index values.</p><p num="0083"> The upmix matrix decoder 502 further expresses the at least one second element of the parameter vector by a parameter value corresponding to the index value associated with the at least one second entropy-encoded symbol. It has a configured decoding component 608. Thus, this representation is, for example, a decoded version of the parameter encoded by the audio encoding system shown in FIG. In other words, this representation is equal to the quantized parameters encoded by the audio encoding system shown in Figure 1.</p><p num="0084"> According to one embodiment of the invention, each entropy-encoded symbol in the vector of entropy-encoded symbols has the same probability table for all entropy-encoded symbols in the vector of entropy-encoded symbols. Represented by a symbol using. The advantage of this is that only one probability table needs to be stored in the decoder's memory. In an audio decoding system, the memory in which the probability table is stored is advantageously fast and therefore expensive, as search speed can be important when decoding entropy-encoded symbols. Therefore, by using only one probability table, the decoder can be cheaper than if two probability tables were used. According to this embodiment, the association component 606 first encodes the first entropy code by first shifting the symbol representing the first entropy coded symbol in the vector of the entropy coded symbol by an offset value. It may be configured to associate the symbol with the index value. Modulo N is then applied to the shifted symbol. Therefore, the association algorithm idx (1) = (idx<sub>excessive</sub>(1) + abs_offset) mod N<sub>Q</sub> (Equation 13) It may be expressed as.</p><p num="0085"> The decoding component 608 is configured to represent the first element of the parameter vector by the parameter value corresponding to the index value associated with the first entropy-encoded symbol. Thus, this representation is, for example, a decoded version of the parameter encoded by the audio encoding system 100 shown in FIG.</p><p num="0086"> The method of differentially encoding aperiodic quantities will be further described in relation to FIGS. 7 to 10.</p><p num="0087"> Figures 7 and 9 describe how to encode the four second elements in the parameter vector. Therefore, the input vector 902 contains five parameters. These parameters can take any value between a minimum and a maximum. In this example, the minimum value is -9.6 and the maximum value is 9.4. The first stage S702 of the encoding method expresses each parameter in the vector 902 by an index value which can take N kinds of values. In this case, N is chosen as 96. That is, the quantization step size is 0.2. This gives the vector 904. The next step, S704, calculates the difference between the second element, each of the four upper parameters in vector 904, and its predecessor element. Thus, the resulting vector 906 contains four difference values-the four upper values in the vector 906. As can be seen in FIG. 9, these diffs can be negative, 0 or positive. As explained above, it is advantageous to have N different values, in this case 96 different values. To achieve this, in the next step S706 of this method, modulo 96 is applied to the second element in vector 906. The resulting vector 908 does not contain any negative values. The thus achieved symbol shown in vector 908 is then used to encode the second element of the vector in the final stage S708 of the method shown in FIG. It is by entropy encoding the symbol associated with the at least one second element based on a probability table containing the probabilities of the symbols shown in vector 908.</p><p num="0088"> As can be seen in Figure 9, the first element is not processed after the indexing stage S702. 8 and 10 describe how to encode the first element in the input vector. The same assumptions made in the above description of FIGS. 7 and 9 regarding the minimum and maximum values of the parameters and the number of possible index values are valid when describing FIGS. 8 and 10. The first element 1002 is received by the encoder. In the first stage of the encoding method, S802, the parameters of the first element are represented by the index value 1004. In the next stage S804, the indexed value 1004 is shifted by an offset value. In this example, the offset value is 49. This value is calculated as described above. In the next stage, S806, modulo 96 is applied to the shifted index value 1006. The resulting value 1008 is then used to encode the first element by entropy encoding symbol 1008 using the same probability table used in FIG. 7 to encode at least one of the elements. used.</p><p num="0089"> FIG. 11 shows an embodiment 102'with the upmix matrix encoding component 102 in FIG. The upmix matrix encoder 102'may be used to encode the upmix matrix in an audio encoding system, such as the audio encoding system 100 shown in FIG. As mentioned above, each row of the upmix matrix contains M elements that allow the reconstruction of the audio object from the downmix signal containing M channels.</p><p num="0090"> At a low overall target bitrate, sending all M upmix matrix elements per object and T / F tile encoded one by one for each downmix channel is an undesirably high bitrate. May be required. This can be reduced by "sparsening" the upmix matrix, that is, by trying to reduce the number of non-zero elements. In some cases, four of the five elements are 0, and a single downmix channel is used as the basis for the reconstruction of the audio object. A sparse matrix has a coded index (absolute or differential) probability distribution that differs from a non-sparse matrix. If the upmix matrix contains a large percentage of 0s, the value 0 is more certain than 0.5, and Huffman coding is used, the coding efficiency is reduced. This is because the Huffman coding algorithm is inefficient when certain values, such as 0, have a probability greater than 0.5. Moreover, since many of the elements in the upmix matrix have a value of 0, they contain no information at all. Therefore, one strategy could be to select a subset of upmix matrix elements, encode only that, and transmit it to the decoder. This can reduce the required bit rate of the audio encoding / decoding system because less data is transmitted.</p><p num="0091"> A dedicated coding mode for sparse matrices may be used to increase the efficiency of coding the upmix matrix. This will be described in detail below.</p><p num="0092"> Encoder 102'has a receiving component 1102 adapted to receive each row in the upmix matrix. The encoder 102'also has a selection component 1104 adapted to select a subset of elements from the M elements of a row in the upmix matrix. In most cases, the subset contains all elements that do not have a value of 0. However, in certain embodiments, the selection component may choose not to select elements with non-zero values, such as elements with values close to zero. According to embodiments, the selected subset of elements may contain the same number of elements for each row of the upmix matrix. The number of elements selected may be one to further reduce the required bit rate.</p><p num="0093"> The encoder 102'also has an encoding component 1106 that is adapted to represent each element in the selected subset of elements by value and position in the upmix matrix. Encoding component 1106 is further adapted to encode the value of each element in the selected subset of elements and their position in the upmix matrix. Encoding component 1106 may be adapted to encode values using, for example, modulo differential encoding as described above. In this case, for each row in the upmix matrix and for multiple frequency bands or multiple time frames, the element values of the selected subset of elements form one or more vectors of parameters. Each parameter in the parameter vector corresponds to one of the plurality of frequency bands or the plurality of time frames. The parameter vector may be encoded using the modulo difference encoding described above. In a further embodiment, the parameter vector may be encoded using conventional differential encoding. In yet another embodiment, the encoding component 1106 is adapted to encode each value separately using the true quantization value of each value, i.e., the fixed rate coding of the non-differential quantized quantized value. To.</p><p num="0094"> The following example of average bit rate was observed for typical content. Their bitrates are M = 5, the number of audio objects to be reconstructed on the decoder side is 11, the number of frequency bands is 12, and the parameter quantizer step size is 0.1. , 192 levels were measured. The following average bitrates were observed when all five elements were encoded for each row in the upmix matrix.</p><p num="0095"> Fixed rate coding: 165kb / sec Delta encoding: 51 kb / sec Modulo differential coding: 51 kb / sec, but half the size of the probability table or codebook as described above.</p><p num="0096"> For each row in the upmix matrix, only one element is selected by the selection component 1104, i.e. for sparse encoding, the following average bitrates were observed.</p><p num="0097"> Fixed rate coding (8 bits for value, 3 bits for position): 45 kb / sec Modulo differential encoding for both element values and element positions: 20 kb / sec.</p><p num="0098"> Encoding component 1106 may be adapted to encode the position of each element in the upmix matrix in the subset of elements in the same way as the value. Encoding component 1106 may be adapted to encode the position of each element in the upmix matrix in a subset of the elements in a different way than encoding the value. When encoding positions using differential or modulo differential coding, the position of the elements in the selected subset of elements for each row in the upmix matrix and for multiple frequency bands or multiple time frames Form one or more vectors of parameters. Each parameter in the parameter vector corresponds to one of the plurality of frequency bands or the plurality of time frames. The parameter vector is encoded using the above differential coding or modulo differential coding.</p><p num="0099"> It may be noted that the encoder 102'may be combined with the encoder 102 of FIG. 2 to achieve the modulo differential coding of the sparse upmix matrix described above.</p><p num="0100"> Further, the method of encoding rows in a sparse matrix is illustrated above for encoding rows in a sparse upmix matrix, but this method is of another type of sparse matrix well known to those of skill in the art. It may be noted that it may be used to encode.</p><p num="0101"> The method of encoding the sparse upmix matrix will be further described in relation to FIGS. 13 to 15.</p><p num="0102"> The upmix matrix is received, for example, by the receiving component 1102 in FIG. For each row 1402, 1502 in the upmix matrix, the method involves selecting a subset from the M of that row of the upmix matrix, eg, five elements (S1302). Each element in the selected subset of elements is then represented by a value and a position in the upmix matrix (S1304). In FIG. 14, one element is selected as the subset (S1302). For example, element number 3 with a value of 2.34. Thus, the representation may be a vector 1404 with two fields. The first field in the vector 1404 represents a value, eg 2.34, and the second field in the vector 1404 represents a position, eg 3. In FIG. 15, two elements are selected as the subset (S1302). For example, element number 3 with a value of 2.34 and element number 5 with a value of -1.81. Therefore, the representation may be a vector 1504 with four fields. The first field in vector 1504 represents the value of the first element, eg 2.34, and the second field in vector 1504 represents the position of the first element, eg 3. The third field in vector 1504 represents the value of the second element, eg -1.81, and the fourth field in vector 1504 represents the position of the second element, eg 5. Representations 1404, 1504 are then encoded according to the above (S1306).</p><p num="0103"> FIG. 12 is a generalized block diagram of an audio decoding system 1200 based on an exemplary embodiment. The decoder 1200 is configured to receive a downmix signal 1210 containing M channels and at least one encoded element 1204 representing a subset of the M elements in a row in the upmix matrix. It has component 1206. Each of the encoded elements contains a value and a position in that row in the upmix matrix. The position indicates which of the M channels of the downmix signal 1210 corresponds to the encoded element. The at least one encoded element 1204 is decoded by the upmix matrix element decoding component 1202. The upmix matrix element decoding component 1202 is configured to decode the at least one encoded element 1204 according to the encoding strategy used to encode the at least one encoded element 1204. Examples of such encoding strategies are disclosed above. The at least one decoded element 1214 is then sent to the reconstruction component 1208. This reconstruction component 1208 will reconstruct the time / frequency tile of the audio object from the downmix signal 1210 by forming a linear combination of the downmix channels corresponding to at least one encoded element 1204. It is configured. When forming a linear combination, each downmix channel is multiplied by its corresponding encoded element 1204.</p><p num="0104"> For example, if the decoded element 1214 contains the value 1.1 and position 2, the time / frequency tile of the second downmix channel is multiplied by 1.1, which is then used to reconstruct the audio object.</p><p num="0105"> The audio decoding system 500 also has a rendering component 1216 that outputs an audio signal based on the reconstructed audio object 1218. The type of the audio signal depends on what type of playback unit is connected to the audio decoding system 1200. For example, if a pair of headphones are connected to the audio decoding system 1200, the rendering component 1216 may output a stereo signal.</p><p num="0106"><u style="single"><Equivalent, extension, alternative, etc.></u> Examination of the above description will reveal to those skilled in the art further embodiments of the present disclosure. Although this article and the drawings disclose embodiments and examples, this disclosure is not limited to these individual examples. Numerous modifications and modifications can be made without departing from the scope of the present disclosure as defined by the appended claims. Even if there is a reference code appearing in the claims, it is not understood to limit the scope thereof.</p><p num="0107"> Further, from the examination of the drawings, the present disclosure and the accompanying claims, variations to the disclosed embodiments can be understood and implemented by those skilled in the art who implement the present disclosure. In the claims, the word "have / include" does not exclude other elements or steps, and the singular representation does not exclude plurals. The fact that certain measures are listed in different dependent claims does not indicate that the combination of these measures cannot be used in an advantageous manner.</p><p num="0108"> The systems and methods disclosed above may be implemented as software, firmware, hardware or a combination thereof. In a hardware implementation, the division of tasks between functional units mentioned in the above description does not necessarily correspond to the division into physical units. Conversely, one physical component may have multiple functions, or one task may be jointly performed by several physical components. Certain components or all components may be implemented as software executed by a digital signal processor or microprocessor, or as hardware or as a purpose-built integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-temporary media) and communication media (or temporary media). As is well known to those of skill in the art, the term computer storage medium is implemented in any method or technique for storing information such as computer readable instructions, data structures, program modules or other data. Includes volatile and non-volatile, removable and non-removable media. Computer storage media are, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROMs, digital versatile discs (DVDs) or other optical disc storage, magnetic cassettes, magnetic tapes, magnetics. Includes disk storage or other magnetic storage devices or any other medium that can be used to store desired information and can be accessed by a computer. In addition, the communication medium typically embodies computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transfer mechanism, including any information delivery medium. That is well known to those in the art.</p><p num="0109"> Some aspects are described. [Aspect 1] A method of encoding a vector of parameters in an audio encoding system, where each parameter corresponds to an aperiodic quantity, said vector having a first element and at least one second element. Is: At the stage of expressing each parameter in the vector by an index value that can take N kinds of values; At the stage of associating each of the at least one second element with a symbol, the symbol is: Calculate the difference between the index value of the second element and the index value of its preceding element in the vector; Calculated by applying modulo N to the difference, with the steps; Includes a step of encoding each of the at least one second element by entropy encoding the symbol associated with the at least one second element based on a probability table containing the probabilities of the symbol. , Method. [Aspect 2] At the stage of associating the first element in the vector with the symbol, the symbol is: The index value representing the first element in the vector is shifted by an offset value; Calculated by applying modulo N to the shifted index value, with steps; It further comprises the step of encoding the first element by entropy encoding the symbol associated with the first element using the same probability table used to encode the at least one second element. , The method according to aspect 1. [Aspect 3] The method of aspect 2, wherein the offset value is equal to the difference between the most probable index value for the first element and the most probable symbol for the at least one second element in the probability table. [Aspect 4] The first element and the at least one second element of the vector of the parameters correspond to any of the different frequency bands used in the audio encoding system at a particular time frame. Or the method described in item 1. [Aspect 5] The first element and the at least one second element of the vector of the parameters correspond to different time frames used in the audio encoding system in a particular frequency band, any of aspects 1 to 3. Or the method described in item 1. [Aspect 6] The probability table is converted into a Huffman codebook, the symbol associated with an element in the vector is used as a codebook index, and the encoding step is to each of the at least one second element. Of aspects 1-5, comprising encoding by representing the second element in codewords in a codebook indexed by the codebook index associated with the second element. The method described in any one of the items. [Aspect 7] The encoding step is at least one second by representing the first element with a codeword in the Huffman codebook indexed by a codebook index associated with the first element. The method of aspect 6 in quoting aspect 2, comprising encoding the first element in the vector using the same Huffman codebook used to encode the elements of. [Aspect 8] The method of any one of aspects 1-7, wherein the parameter vector corresponds to an element in the upmix matrix determined by the audio encoding system. [Aspect 9] A computer-readable storage medium having computer code instructions adapted to perform the method according to any one of aspects 1 to 8 when executed on a device having processing capabilities. [Aspect 10] An encoder that encodes a vector of parameters in an audio encoding system, where each parameter corresponds to an aperiodic quantity, the vector having a first element and at least one second element. Ha: With a receiving component adapted to receive the vector; With an indexing component adapted to represent each parameter in the vector by an index value that can take N different values; An association component adapted to associate each of the at least one second element with a symbol, wherein the symbol is: Calculate the difference between the index value of the second element and the index value of its preceding element in the vector; With the associative component calculated by applying the modulo N to the difference; An encoding component that encodes each of the at least one second element by entropy encoding the symbol associated with the at least one second element based on a probability table containing the probabilities of the symbol. Have, Encoder. [Aspect 11] A method of decoding a vector of entropy-encoded symbols in an audio decoding system into a vector of parameters related to aperiodic quantities, wherein the vector of entropy-encoded symbols is the first entropy. It has a coded symbol and at least one second entropy coded symbol, the vector of said parameters has a first element and at least one second element, the method is: The stage of expressing each entropy-encoded symbol in the vector of the entropy-encoded symbol by a symbol that can take N integer values by using a probability table; With the step of associating the first entropy-encoded symbol with the index value; The index value of the at least one second entropy-encoded symbol includes the step of associating each of the at least one second entropy-encoded symbol with the index value. An index value associated with an entropy-encoded symbol preceding the second entropy-encoded symbol in the vector of the entropy-encoded symbol, and a symbol representing the second entropy-encoded symbol. Calculate the sum with; Calculated by applying modulo N to the sum, with steps; The step of expressing the at least one second element of the vector of the parameter by a parameter value corresponding to the index value associated with the at least one second entropy-encoded symbol. Method. [Aspect 12] The step of representing each entropy-encoded symbol in the vector of the entropy-encoded symbol by a symbol has the same probability table for all entropy-encoded symbols in the vector of the entropy-encoded symbol. The index value associated with the first entropy-encoded symbol that was executed using: The symbol representing the first entropy-encoded symbol in the vector of the entropy-encoded symbol is shifted by an offset value; Calculated by applying modulo N to shifted symbols, The method is further: The first element of the vector of the parameters is represented by a parameter value corresponding to the index value associated with the first entropy-encoded symbol. The method according to aspect 11. [Aspect 13] The method according to aspect 11 or 12, wherein the probability table is converted into a Huffman codebook, and each entropy-coded symbol corresponds to a codeword in the Huffman codebook. [Aspect 14] Each codeword in the Huffman Codebook is associated with a codebook index, and the symbol represents each entropy-encoded symbol in the vector of the entropy-encoded symbol. 13. The method of aspect 13, comprising representing the entropy-encoded symbol by a codebook index associated with the codeword corresponding to the symbol. [Aspect 15] Each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different frequency band used in the audio decoding system at a particular time frame, any one of aspects 11-14. Item description method. [Aspect 16] Each entropy-encoded symbol in the vector of entropy-encoded symbols corresponds to a different time frame used in the audio decoding system in a particular frequency band, any one of aspects 11-14. Item description method. [Aspect 17] The method of any one of aspects 11-16, wherein the parameter vector corresponds to an element in the upmix matrix used by the audio decoding system. [Aspect 18] A computer-readable storage medium having computer code instructions adapted to perform the method according to any one of aspects 11 to 17 when executed on a device having processing capabilities. [Aspect 19] A decoder that decodes a vector of entropy-encoded symbols in an audio decoding system into a vector of parameters related to aperiodic quantities, said vector of entropy-encoded symbols being the first entropy. It has a coded symbol and at least one second entropy coded symbol, the vector of the parameter has a first element and at least a second element, and the decoder said: With a receiving component configured to receive said vector of entropy-encoded symbols; With an indexing component configured to represent each entropy-encoded symbol in said vector of entropy-encoded symbols by symbols that can take N integer values by using a probability table; An association component configured to associate the first entropy-encoded symbol with an index value. The association component is further configured to associate each of the at least one second entropy-encoded symbol with an index value, and the index value of the at least one second entropy-encoded symbol is: The sum of the index value of the entropy-encoded symbol preceding the second entropy-encoded symbol in the vector of the entropy-encoded symbol and the symbol representing the second entropy-encoded symbol. Calculate; Calculated by applying modulo N to the sum, With associated components; With a decoding component configured to represent said at least one second element of the vector of said parameters with a parameter value corresponding to the index value associated with said said at least one second entropy-encoded symbol. Have, decoder. [Aspect 20] A method of encoding an upmix matrix in an audio encoding system, where each row of the upmix matrix allows the reconstruction of time / frequency tiles of audio objects from a downmix signal containing M channels. It contains M elements and the method is: For each row in the upmix matrix: Select a subset of the elements from the M elements in that row in the upmix matrix; Each element in the selected subset of elements is represented by a value and a position in the upmix matrix; Includes encoding the value of each element in the selected subset of elements and their position in the upmix matrix. Method. [Aspect 21] 20. Aspect 20, wherein for each row in the upmix matrix, the positions of the elements of the selected subset in the upmix matrix vary across multiple frequency bands and / or across multiple time frames. the method of. [Aspect 22] 20 or 21. The method of aspect 20 or 21, wherein the selected subset of elements comprises the same number of elements for each row of the upmix matrix. [Aspect 23] For each row of the upmix matrix, the selected subset of elements comprises exactly one element from the M elements of that row in the upmix matrix, any one of aspects 20-22. The method described. [Aspect 24] For each row in the upmix matrix and for multiple frequency bands or multiple time frames, the element values of the selected subset of elements form one or more vectors of the parameter, each in the vector of the parameter. The parameter corresponds to one of the plurality of frequency bands or the plurality of time frames, and the one or more vectors of the parameter use the method according to any one of aspects 1 to 8. The method according to any one of aspects 20 to 23, which is encoded. [Aspect 25] For each row in the upmix matrix and for multiple frequency bands or multiple time frames, the position of the element in the selected subset of elements forms one or more vectors of the parameter, each in the vector of the parameter. The parameter corresponds to one of the plurality of frequency bands or the plurality of time frames, and the vector of the parameter is encoded using the method according to any one of aspects 1 to 8. The method according to any one of aspects 20 to 24. [Aspect 26] A computer-readable storage medium having computer code instructions adapted to perform the method according to any one of aspects 20 to 25 when executed on a device having processing capabilities. [Aspect 27] An encoder that encodes an upmix matrix in an audio encoding system, where each row of the upmix matrix allows the reconstruction of time / frequency tiles of audio objects from a downmix signal containing M channels. The encoder contains M elements: With a receiving component adapted to receive each row in the upmix matrix; With a selection component adapted to select a subset of elements from the M elements in the row in the upmix matrix; Each element in the selected subset of elements has an encode component adapted to be represented by a value and a position in the upmix matrix, and the encode component is further in the selected subset of elements. Adapted to encode the value of each element and its position in the upmix matrix, Encoder. [Aspect 28] A way to reconstruct the time / frequency tiles of an audio object in an audio decoding system: At the stage of receiving a downmix signal containing M channels; At the stage of receiving at least one encoded element representing a subset of M elements in a row in the upmix matrix, each encoded element has a value and a position in that row in the upmix matrix. Including, the position indicates one of the M channels of the downmix signal to which the encoded element corresponds, with a step; A step of reconstructing the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded element, said linear. In a combination, each downmix channel is multiplied by the value of its corresponding encoded element, including a step. Method. [Aspect 29] 28. The method of aspect 28, wherein the position of the at least one encoded element varies across multiple frequency bands and / or across multiple time frames. [Aspect 30] 28 or 29. The method of aspect 28 or 29, wherein the number of elements of at least one encoded element is equal to one. [Aspect 31] For multiple frequency bands or multiple time frames, the values of the at least one encoded element form one or more vectors, each value being represented by an entropy-encoded symbol and entropy-encoded. Each entropy-encoded symbol in each vector of the symbol corresponds to one of the plurality of frequency bands or one of the plurality of time frames, and the one or more vectors of the entropy-encoded symbol. Is the method according to any one of aspects 28 to 30, wherein is decoded using the method according to any one of aspects 11 to 17. [Aspect 32] For multiple frequency bands or multiple time frames, the positions of the at least one encoded element form one or more vectors, each position being represented by an entropy-encoded symbol and entropy-encoded. Each symbol in each vector of the symbol corresponds to one of the plurality of frequency bands or the plurality of time frames, and the one or more vectors of the entropy-encoded symbol correspond to the plurality of modes 11 to 17. The method according to any one of aspects 28 to 31, which is decoded using the method according to any one. [Aspect 33] A computer-readable storage medium having computer code instructions adapted to perform the method according to any one of aspects 28-32 when executed on a device having processing capabilities. [Aspect 34] A decoder that reconstructs the time / frequency tiles of an audio object: A receiving component configured to receive at least one encoded element that represents a subset of the M elements of a row in a downmix signal containing M channels and an upmix matrix, each encoded. The element contains a value and a position in that row in the upmix matrix, the position indicating one of the M channels of the downmix signal to which the encoded element corresponds. When; A reconstruction component configured to reconstruct the time / frequency tile of the audio object from the downmix signal by forming a linear combination of the downmix channels corresponding to the at least one encoded element. In the linear combination, each downmix channel is multiplied by the value of its corresponding encoded element. decoder.</p>
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| WO2012058229A1 | Cites | World Intellectual Property Organization (WIPO) | Y | Search report | 1-9 |
| WO2013064957A1 | Cites | World Intellectual Property Organization (WIPO) | Y | Search report | 1-9 |
| JP6105159B2 | Cites | Japan | EX | Search report | 5,9 |
99 members in 19 offices
Members99
| Document | Office | Kind | |
|---|---|---|---|
| CA2911746A1 | Canada | A1 | |
| CA2990261A1 | Canada | A1 | |
| CA3077876A1 | Canada | A1 | |
| CA3163664A1 | Canada | A1 | |
| WO2014187988A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2014187988A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AU2014270301A1 | Australia | A1 | |
| SG11201509001YA | Singapore | A | |
| CN105229729A | China | A | |
| KR20160013154A | Republic of Korea | A | |
| MX2015015926A | Mexico | A | |
| EP3005350A2 | European Patent Office (EPO) | A2 | |
| US2016111098A1 | United States of America | A1 | |
| JP2016526186A | Japan | A | |
| UA112833C2 | Ukraine | C2 | |
| HK1217246A | Hong Kong, China | A | |
| HK1217246A1 | Hong Kong, China | A1 | |
| JP6105159B2 | Japan | B2 | |
| EP3005350B1 | European Patent Office (EPO) | B1 | |
| JP2017102484AThis record | Japan | A | |
| RU2015155311A | Russian Federation | A | |
| US9704493B2 | United States of America | B2 | |
| DK3005350T3 | Denmark | T3 | |
| BR112015029031A2 | Brazil | A2 | |
| KR101763131B1 | Republic of Korea | B1 | |
| KR20170087971A | Republic of Korea | A | |
| AU2014270301B2 | Australia | B2 | |
| ES2629025T3 | Spain | T3 | |
| MX350117B | Mexico | B | |
| PL3005350T3 | Poland | T3 | |
| US2017309279A1 | United States of America | A1 | |
| EP3252757A1 | European Patent Office (EPO) | A1 | |
| SG10201710019SA | Singapore | A | |
| RU2643489C2 | Russian Federation | C2 | |
| CA2911746C | Canada | C | |
| US9940939B2 | United States of America | B2 | |
| US2018240465A1 | United States of America | A1 | |
| KR20180099942A | Republic of Korea | A | |
| KR101895198B1 | Republic of Korea | B1 | |
| IL242410A | Israel | A | |
| IL242410B | Israel | B | |
| RU2676041C1 | Russian Federation | C1 | |
| CN105229729B | China | B | |
| CN110085238A | China | A | |
| JP6573640B2 | Japan | B2 | |
| US10418038B2 | United States of America | B2 | |
| EP3252757B1 | European Patent Office (EPO) | B1 | |
| US2020013415A1 | United States of America | A1 | |
| RU2710909C1 | Russian Federation | C1 | |
| JP2020016884A | Japan | A | |
| KR102072777B1 | Republic of Korea | B1 | |
| EP3605532A1 | European Patent Office (EPO) | A1 | |
| KR20200013091A | Republic of Korea | A | |
| MY173644A | Malaysia | A | |
| CA2990261C | Canada | C | |
| US10714104B2 | United States of America | B2 | |
| MX375380B | Mexico | B | |
| MX2020010038A | Mexico | A | |
| KR102192245B1 | Republic of Korea | B1 | |
| KR20200145837A | Republic of Korea | A | |
| US2020411017A1 | United States of America | A1 | |
| BR112015029031B1 | Brazil | B1 | |
| KR20210060660A | Republic of Korea | A | |
| US11024320B2 | United States of America | B2 | |
| RU2019141091A | Russian Federation | A | |
| KR102280461B1 | Republic of Korea | B1 | |
| JP6920382B2 | Japan | B2 | |
| EP3605532B1 | European Patent Office (EPO) | B1 | |
| JP2021179627A | Japan | A | |
| US2021390963A1 | United States of America | A1 | |
| EP3961622A1 | European Patent Office (EPO) | A1 | |
| ES2902518T3 | Spain | T3 | |
| KR102384348B1 | Republic of Korea | B1 | |
| KR20220045259A | Republic of Korea | A | |
| CA3077876C | Canada | C | |
| KR102459010B1 | Republic of Korea | B1 | |
| KR20220148314A | Republic of Korea | A | |
| US11594233B2 | United States of America | B2 | |
| JP7258086B2 | Japan | B2 | |
| JP2023076575A | Japan | A | |
| CN110085238B | China | B | |
| KR102572382B1 | Republic of Korea | B1 | |
| US2023282219A1 | United States of America | A1 | |
| KR20230129576A | Republic of Korea | A | |
| MY199032A | Malaysia | A | |
| EP3961622B1 | European Patent Office (EPO) | B1 | |
| EP4290510A2 | European Patent Office (EPO) | A2 | |
| EP4290510A3 | European Patent Office (EPO) | A3 | |
| ES2965423T3 | Spain | T3 | |
| KR102715092B1 | Republic of Korea | B1 | |
| KR20240151867A | Republic of Korea | A | |
| JP7585379B2 | Japan | B2 | |
| US12236961B2 | United States of America | B2 | |
| JP2025028867A | Japan | A | |
| MX375380B | Mexico | B | |
| JP2025100674A | Japan | A | |
| CA3163664C | Canada | C | |
| EP4290510B1 | European Patent Office (EPO) | B1 | |
| US2025356860A1 | United States of America | A1 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A132A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2017102484
- Application
- 38524
Titles2
- Japanese
- オーディオ・エンコーダおよびデコーダ
- English
- Audio encoder and decoder
Classification
- CPC, 7
- G10L19/0017
- G10L19/008
- G10L19/038
- G10L19/032
- H04S3/02
- H04S2400/01
- H04S2420/03
- IPC, 2
- G10L19 038
- H03M7 40