Device and program for fade-in/fade-out processing for audio frame
Abstract
[Subject] The fade-in / fade-out processing unit, and the program which can add a temporal response to volume are offered without decrypting audio information completely, even if it is equipment of low computing speed and low memory quantity like a cellular phone. [Solution means] The bit stream of an audio frame based on an MPEG audio standard is decomposed into a header and a live-data part, The bit stream decomposition part 10 which outputs global_gain, It has the profit parameter change part 12 which increases or decreases global_gain in the predetermined time range, and the bit stream composition part 14 which compounds the headers and live-data parts including changed global_gain which were decomposed. Moreover, bs_data_env for every subband is increased or decreased about the audio information containing SBR data. [Selection figure] Fig. 2
Term
No projected expiry on record.
- Priority and filed
- Published
- Today
12 claims: 5 independent, 7 dependent
- 1It is changed to a bitstream decomposition means that decomposes the bitstream of an audio frame into a header element and an actual data part and outputs a gain parameter value, and a gain parameter changing means that increases or decreases the gain parameter value within a predetermined time range. A fade-in / fade-out processing apparatus comprising:a bitstream synthesizing means for synthesizing the header element and the actual data unit including the gain parameter value. オーディオフレームのビットストリームをヘッダ要素及び実データ部に分解して、利得パラメータ値を出力するビットストリーム分解手段と、 前記利得パラメータ値を所定時間範囲で増加又は減少させる利得パラメータ変更手段と、 変更された前記利得パラメータ値を含めて、前記ヘッダ要素及び実データ部を合成するビットストリーム合成手段とを有することを特徴とするフェードイン/フェードアウト処理装置。
- 4The bitstream decomposition means for decomposing the bitstream of the audio frame into the header element and the actual data part, the Huffman decoding means for restoring the code string of the gain parameter of the envelope included in the actual data part to the gain parameter value, and the above. A bit stream that synthesizes the header element and the actual data part with the gain parameter changing means that increases or decreases the gain parameter value within a predetermined time range and Huffman-encodes the changed gain parameter value to be included in the actual data part. A fade-in / fade-out processing apparatus comprising a synthetic means. オーディオフレームのビットストリームをヘッダ要素及び実データ部に分解するビットストリーム分解手段と、 前記実データ部に含まれる包絡線の利得パラメータの符号列を利得パラメータ値に復元するハフマン復号化手段と、 前記利得パラメータ値を所定時間範囲で増加又は減少させて、変更された前記利得パラメータ値をハフマン符号化して当該実データ部に含める利得パラメータ変更手段と、 前記ヘッダ要素及び実データ部を合成するビットストリーム合成手段とを有することを特徴とするフェードイン/フェードアウト処理装置。
- 6The gain parameter changing means increases or decreases the gain parameter value monotonically, exponentially, or in a characteristic curve specified in advance with respect to time progress. The fade-in / fade-out processor according to any one of 5 to 5. 前記利得パラメータ変更手段は、前記利得パラメータ値を、時間進行に対して、単調的に、指数曲線的に又は予め指定された特徴ある曲線的に、増加又は減少させることを特徴とする請求項1から5のいずれか1項に記載のフェードイン/フェードアウト処理装置。
- 7It is changed to a bitstream decomposition means that decomposes the bitstream of an audio frame into a header element and an actual data part and outputs a gain parameter value, and a gain parameter changing means that increases or decreases the gain parameter value within a predetermined time range. A fade-in / fade-out processing program characterized by functioning as a bitstream synthesizing means for synthesizing the header element and the actual data portion including the gain parameter value. オーディオフレームのビットストリームをヘッダ要素及び実データ部に分解して、利得パラメータ値を出力するビットストリーム分解手段と、 前記利得パラメータ値を所定時間範囲で増加又は減少させる利得パラメータ変更手段と、 変更された前記利得パラメータ値を含めて、前記ヘッダ要素及び実データ部を合成するビットストリーム合成手段として機能させることを特徴とするフェードイン/フェードアウト処理プログラム。
- 10The bitstream decomposition means for decomposing the bitstream of the audio frame into the header element and the actual data part, the Huffman decoding means for restoring the code string of the gain parameter of the envelope included in the actual data part to the gain parameter value, and the above. A bit stream that synthesizes the header element and the actual data part with the gain parameter changing means that increases or decreases the gain parameter value within a predetermined time range and Huffman-encodes the changed gain parameter value to be included in the actual data part. A fade-in / fade-out processing program characterized by functioning as a synthesis means. オーディオフレームのビットストリームをヘッダ要素及び実データ部に分解するビットストリーム分解手段と、 前記実データ部に含まれる包絡線の利得パラメータの符号列を利得パラメータ値に復元するハフマン復号化手段と、 前記利得パラメータ値を所定時間範囲で増加又は減少させて、変更された前記利得パラメータ値をハフマン符号化して当該実データ部に含める利得パラメータ変更手段と、 前記ヘッダ要素及び実データ部を合成するビットストリーム合成手段として機能させることを特徴とするフェードイン/フェードアウト処理プログラム。
Independent claims5
37 paragraphs, as filed
The present invention relates to a fade-in / fade-out processor and program for an audio frame.
In music distribution via the Internet, PCM code data obtained from the original sound is usually distributed in a compressed form. As a typical audio compression method, there is MP3 (ISO / IEC 11172-3, JIS X 4323) based on the MPEG1 audio layer III standard. In addition, the MPEG2 Audio Layer III standard, which has been greatly extended from the MPEG1 Audio Layer III while maintaining compatibility, has a coding efficiency of up to 20% to 50% compared to the MPEG1 Audio Layer III, although it is not compatible. AAC (Advanced Audio Coding) to be achieved is standardized. AAC, which realizes high sound quality with such a very small amount of code, has been attracting attention as a code for music distribution to mobile phones.
In recent years, audio data can be reproduced in various scenes according to the taste of the user. The user can not only listen to music as a hobby, but also, for example, in a mobile phone, ring the music as a ringtone or ring instead of an alarm. At this time, there is a demand for the user to fade in (monotonically increase) or fade out (monotonically decrease) the volume to make the music comfortable. However, in order to change the volume, the user usually has to manually change the volume of the speaker. In addition, once the music was played, the sound whose volume was changed by itself had to be recorded again and stored in the memory of the device.
On the other hand, there is a method of realizing fade-in by decoding only the sample of the front part of the audio data, gradually increasing the gain, and further encoding to regenerate the audio data (for example, Patent Document). 1). According to this method, fade-out is realized by decoding only the sample of the rear part of the audio data, gradually lowering the gain, and further encoding to regenerate the audio data.
<patcit num="1"><text>Japanese Unexamined Patent Publication No. 7-220394</text></patcit>
<p> However, according to the conventional method, the gain of the decoded audio data must be controlled and further encoded to regenerate the audio data in order to only change the volume with time, which is relatively high. It requires a calculation speed and a memory amount. On the other hand, there is a problem that it is difficult to realize it with a mobile phone having a low calculation speed and a low memory amount.</p><p> Therefore, according to the present invention, even in a device having a low calculation speed and a low memory amount such as a mobile phone, the volume can be changed with time without completely decoding the audio data. It is an object of the present invention to provide a processing device and a program.</p>
<p> According to the fade-in / fade-out processing device of the present invention, a bitstream decomposition means that decomposes a bitstream of an audio frame into a header element and an actual data part and outputs a gain parameter value, and a gain parameter value within a predetermined time range. It is characterized by having a gain parameter changing means for increasing or decreasing, and a bitstream synthesizing means for synthesizing a header element and an actual data part including the changed gain parameter value.</p><p> According to another embodiment of the fade-in / fade-out processing apparatus of the present invention, it is also preferable that the audio frame includes AAC data based on the MPEG audio standard and the gain parameter value is global_gain.</p><p> Further, according to another embodiment of the fade-in / fade-out processing apparatus of the present invention, the bitstream decomposition means is configured to output a scale factor, and is set to global_gain so that the quantization step width is not negative. It further has an operable area monitoring means for notifying the gain parameter changing means of the minimum value in the quantization step width calculated from the scale factor of the difference value, and the gain parameter changing means has the quantization step width from global_gain. It is also preferable that the global_gain is configured so as not to decrease more than the value obtained by subtracting the minimum value.</p><p> According to the fade-in / fade-out processing apparatus of the present invention, the bitstream decomposition means for decomposing the bitstream of the audio frame into the header element and the actual data portion, and the code string of the gain parameter of the envelope included in the actual data portion are gained. A Huffman decoding means that restores to a parameter value, a gain parameter changing means that increases or decreases the gain parameter value within a predetermined time range, Huffman-encodes the changed gain parameter value, and includes the changed gain parameter value in the actual data part, and a header element. It is characterized by having a bitstream synthesizing means for synthesizing an actual data unit.</p><p> According to another embodiment of the fade-in / fade-out processing apparatus of the present invention, it is also preferable that the audio frame includes SBR data based on the MPEG audio standard and the gain parameter value is bs_data_env.</p><p> Further, according to another embodiment of the fade-in / fade-out processing apparatus of the present invention, the gain parameter changing means specifies the gain parameter value monotonically, exponentially, or in advance with respect to the time progress. It is also preferable to increase or decrease in a characteristic curve.</p><p> According to the fade-in / fade-out processing program of the present invention, a bitstream decomposition means that decomposes a bitstream of an audio frame into a header element and an actual data part and outputs a gain parameter value, and a gain parameter value within a predetermined time range. It is characterized in that it functions as a bitstream synthesizing means for synthesizing a header element and an actual data part including a gain parameter changing means for increasing or decreasing and a changed gain parameter value.</p><p> Further, according to another embodiment in the fade-in / fade-out processing program of the present invention, the audio frame includes AAC data based on the MPEG audio standard, and the gain parameter value is made to function as global_gain. Is also preferable.</p><p> Further, according to another embodiment in the fade-in / fade-out processing program of the present invention, the bitstream decomposition means is configured to output a scale factor, and is set to global_gain so that the quantization step width is not negative. It further has an operable area monitoring means for notifying the gain parameter changing means of the minimum value in the quantization step width calculated from the scale factor of the difference value, and the gain parameter changing means has the quantization step width from global_gain. It is also preferable to make the function so that global_gain does not decrease more than the value obtained by subtracting the minimum value.</p><p> According to the fade-in / fade-out processing program of the present invention, the bitstream decomposition means for decomposing the bitstream of the audio frame into the header element and the actual data part, and the code string of the gain parameter of the envelope included in the actual data part are gained. A Huffman decoding means that restores to a parameter value, a gain parameter changing means that increases or decreases the gain parameter value within a predetermined time range, Huffman-encodes the changed gain parameter value, and includes the changed gain parameter value in the actual data part, and a header element. It is characterized in that it functions as a bitstream synthesizing means for synthesizing the actual data part.</p><p> Further, according to another embodiment in the fade-in / fade-out processing program of the present invention, the audio frame includes SBR data based on the MPEG audio standard, and the gain parameter value is made to function as bs_data_env. Is also preferable.</p><p> Furthermore, according to another embodiment of the fade-in / fade-out processing program of the present invention, the gain parameter changing means specifies the gain parameter value monotonically, exponentially or in advance with respect to time progression. It is also preferable to make it function to increase or decrease in a characteristic curve.</p>
<p> According to the fade-in / fade-out processing device and program of the present invention, by changing only the gain parameter (global_gain) of the audio frame, even a device having a low calculation speed and a low memory amount such as a mobile phone can perform audio. The data can be changed to audio data that can be played by changing the volume over time without completely decoding the data.</p><p> Further, according to the present invention, it is possible to add fade-in / fade-out processing not only to AAC data in the low frequency region but also to SBR data in the high frequency region based on the MPEG audio standard.</p><p> Further, according to the present invention, various patterns of changing the volume to be faded in / faded out can be specified according to the user's wishes.</p>
Hereinafter, the best embodiments of the present invention will be described in detail with reference to the drawings.
FIG. 1 is a configuration diagram of one audio frame.
According to the frame based on the MPEG audio standard, the AAC section, which consists of channels Ch1 and Ch2 (for example, right channel and left channel) and covers the low frequency region, and SBR (Spectral Band Replication) that covers the high frequency region. The department is separated by a tag.
The AAC part contains global_gain for each channel. global_gain stores the value actually used for decoding. Further, the AAC unit includes a scale factor (difference value) corresponding to audio data and coded data, which are subband-decomposed for each channel. The scale factor is in the form of a predicted difference value, and different values for each subband are stored in one place in an array format. Since the scale factor is stored in Huffman coding, it needs to be Huffman-decoded.
SBR (Spectral Band Replication) is a technique for improving sound quality by replicating a high frequency region using a low frequency region on the decoding side. Since SBR only needs to transmit the low frequency domain and a small amount of side information, it is possible to realize sound quality equivalent to that of AAC with a high bit rate with low bit rate information. The SBR part consists of a header part and an actual data part. In the actual data part, bs_data_env (envelope gain parameter), which is different for each subband, is consolidated in one place in an array format, and noise data for synthesis. And are included. Since bs_data_env is stored in Huffman coding, it needs to be Huffman-decoded.
FIG. 2 is a functional configuration diagram of the AAC fade-in / fade-out processing device 1 according to the present invention. These functions are preferably realized programmatically.
An AAC audio frame is input to the AAC fade-in / fade-out processing device 1, and an audio frame to which the fade-in / fade-out processing is applied is output. The bitstream decomposition unit 10 decomposes the bitstream into a header element and an actual data unit. Then, the global_gain included in the header element is notified to the gain parameter changing unit 12, and the code string of the scale factor for each subband is notified to the Huffman decoding unit 11. The Huffman decoding unit 11 decodes the code string of the scale factor, and the extracted scale factor is notified to the operable area monitoring unit 13.
In the gain parameter changing unit 12, control information such as whether to perform fade-in or fade-out and how long the time range is to be performed is specified in advance by the user. Then, the gain parameter changing unit 12 gradually increases or decreases global_gain within a predetermined time range. global_gain is an initial value, and the quantization step size is determined by calculating from that value and the scale factor. When the code string is shortened by changing global_gain, the code length is not changed by inserting the stuffing data so that the audio frame length becomes a predetermined fixed length in the bitstream synthesizing unit 14. Can be done.
The Huffman decoding unit 11 decodes the code string of the scale factor for each decomposed subband. The decoded array of scale factors is notified to the operable area monitoring unit 13.
In the operable monitoring unit 13, global_gain is input from the bitstream decomposition unit 10, and a scale_factor [] array is input from the Huffman decoding unit 11. Then, the operable monitoring unit 13 notifies the gain parameter changing unit 12 of the minimum value in the quantization step width so that the quantization step width calculated from the scale_factor [] does not become negative. The gain parameter changing unit 12 operates so that global_gain does not decrease more than the value obtained by subtracting the minimum value of the quantization step width from global_gain. This makes it possible to prevent the minimum value of the quantization step width calculated from scale_factor [] from becoming negative.
global_gain, scale_factor [] and quantization step size take the following relational values, for example. global_gain = 15 scale_factor [] = 0, -2, -1, -2, +4, Quantization step width = 15, 13, 12, 10, 14,
At this time, it is assumed that global_gain = 15-> 3 is changed. Then, the values of the relationship are as follows. global_gain = 3 scale_factor [] = 0, -2, -1, -2, +4, Quantization step size = 3, 1, 0, -2, 2,
In this case, there will be a negative value part where the quantization step width is "-2". In order to prevent the quantization step width from becoming negative in this way, it is necessary to prevent the global_gain from being reduced below the minimum value "10" of the quantization step width when global_gain = 15. Therefore, the following relationship is the minimum value of global_gain. global_gain = 15-> 5 scale_factor [] = 0, -2, -1, -2, +4, Quantization step width = 5, 3, 2, 0, 4,
In the case of the above example, the operable monitoring unit 13 notifies the gain parameter changing unit 12 of the minimum value "10" of the quantization step width when global_gain = 15. The gain parameter changing unit 12 operates so as not to reduce global_gain below the minimum value "10" of the quantization step width.
The bitstream synthesizing unit 14 synthesizes the decomposed header element and the actual data unit into the bitstream, including the gain parameter output from the gain parameter changing unit 12. As a result, the fade-in / fade-out processing device 1 outputs AAC data to which the fade-in / fade-out processing has been added.
FIG. 3 is a graph showing standard quantization characteristics. Further, FIG. 4 is a graph showing the quantization characteristics in which the volume is reduced by one step. Further, FIG. 5 is a graph showing the quantization characteristics in which the volume is reduced by two steps.
Each graph is represented with the horizontal axis as the input and the vertical axis as the output, and the result of dividing the input signal by the quantization step width Δ is truncated. Changing the step size to Figure 3-> Figure 4-> Figure 5 means fading out, and changing the step size to Figure 5-> Figure 4-> Figure 3 means fading in. In this way, by increasing or decreasing global_gain, the quantization step width is changed, and it becomes possible to control the volume in a pseudo manner.
FIG. 6 is a graph of fade-out change patterns.
In this graph, the vertical axis represents the ratio of global_gain, and the horizontal axis represents the passage of time. Pattern 1 changes from 100% of global_gain with a monotonous decrease. Pattern 2 changes exponentially with a decrease. Pattern 3 is changing after decreasing, increasing, and decreasing again. Such a pattern can be made in any way by changing the global_gain of the gain parameter changing unit 12. How to change it is a design matter.
FIG. 7 is an explanatory diagram of the SBR and bs_data_env parameters.
According to FIG. 7, the low frequency region is encoded by AAC, and that portion is used to duplicate the high frequency region. The envelope in the high frequency domain is represented as the bs_data_env parameter. Fade-in / fade-out can be achieved even in the high frequency range by increasing or decreasing the bs_data_env parameter in the same way as AAC's global_gain and scale factor.
FIG. 8 is a functional configuration diagram of the fade-in / fade-out processing device 2 of the SBR. These functions are preferably realized programmatically.
The SBR fade-in / fade-out processing device 2 includes a bitstream decomposition unit 20, a Huffman decoding unit 21, a gain parameter changing unit 22, and a bitstream synthesizing unit 23.
The bitstream decomposition unit 20 decomposes the bitstream into a header unit and an actual data unit, and notifies the Huffman decoding unit 21 of the Huffman code of the envelope gain parameter included in the actual data unit. The Huffman decoding unit 21 decodes and extracts the code string of bs_data_env (envelope gain parameter) for each subband. The gain parameter changing unit 22 increases or decreases bs_data_env for each subband. Then, the gain parameter changing unit 22 replaces the changed bs_data_env with the Huffman codeword corresponding to the changed bs_data_env and notifies the bitstream synthesizing unit 23. At this time, if the code string is shortened by changing bs_data_env, it is possible to prevent the code length from being changed by inserting stuffing data. The bitstream synthesizing unit 23 synthesizes the header part and the actual data part, and outputs the bitstream. At this time, when replacing the Huffman codeword, when the codeword length becomes long in the case of the Huffman codeword corresponding to the one-step sound reduction, the Huffman codeword having the same or shorter codeword length. It can also be replaced with a quieter sound. As a result, it is possible to prevent an increase in the data length of the entire SBR.
Note that FIG. 2 shows a fade-in / fade-out processing device for AAC, while FIG. 8 shows a fade-in / fade-out processing device for SBR. Therefore, in order to fade in / fade out the low frequency region of AAC and the high frequency region of SBR at the same time, it can be realized by merging the functional configurations of FIGS. 2 and 6.
According to the various embodiments of the present invention described above, various changes, modifications and omissions within the scope of the technical idea and viewpoint of the present invention can be easily made by those skilled in the art. The above explanation is just an example and does not attempt to restrict anything. The present invention is limited only to the scope of claims and their equivalents.
<figref num="1">It is a block diagram of one audio frame.</figref><figref num="2">It is a functional block diagram of the fade-in / fade-out processing apparatus of AAC in this invention.</figref><figref num="3">This is an example of a graph showing standard quantization characteristics.</figref><figref num="4">This is an example of a graph showing the quantization characteristics with the volume reduced by one step.</figref><figref num="5">This is an example of a graph showing the quantization characteristics with the volume reduced by two steps.</figref><figref num="6">It is a graph of the change pattern of fade-out.</figref><figref num="7">It is explanatory drawing of SBR and bs_data_env parameter.</figref><figref num="8">It is a functional block diagram of the fade-in / fade-out processing apparatus of SBR in this invention.</figref>
Code description
1 AAC fade-in / fade-out processing device 10 Bitstream decomposition unit 11 Huffman decoding unit 12 Gain parameter change unit 13 Operable area monitoring unit 14 Bitstream synthesis unit 2 SBR fade-in / fade-out processing unit 20 Bitstream decomposition unit 21 Huffman decoding unit 22 Gain parameter change unit 23 Bitstream synthesis unit 4 Audio data storage unit
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2012118462A | Cited by | Japan | Search report |
| JP2012118462A | Cited by | Japan | Examiner |
| US8364474B2 | Cited by | United States of America | Applicant |
| JP2007187905A | Cited by | Japan | Search report |
| JP2008203739A | Cited by | Japan | Search report |
| JP2007171821A | Cited by | Japan | Examiner |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004111028 | Japan | A | |
| JP20040111028 | – | – | – |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Decision of refusalA02 | A02 | |
| Written amendmentA521 | A521 | |
| Notification of reasons for refusalA131 | A131 | |
| Written amendmentA521 | A521 | |
| Notification of reasons for refusalA131 | A131 | |
| Written request for application examinationA621 | A621 |
Numbers
- Publication
- 2005292702
- Publication, DOCDB
- 2005292702
- Publication, EPODOC
- JP2005292702
- Application
- 111028
- Application, DOCDB
- 2004111028
- Application, EPODOC
- JP20040111028
Titles2
- Japanese
- オーディオフレームに対するフェードイン/フェードアウト処理装置及びプログラム
- English
- Fade-in / fade-out processor and program for audio frames
Classification
- CPC, 3
- G10L19/167
- G10L19/083
- G10L21/0364
- IPC, 3
- G10L21 02
- G10L19 00
- G10L19 14