Audio apparatus
Abstract
A method comprising receiving at a user equipment encrypted content. The content is stored in said user equipment in an encrypted form. At least one key for decryption of said stored encrypted content is stored in the user equipment.
Term
1.6 yearsto projected expiry
Projected expiry 9 May 2028, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
17 claims: 7 independent, 10 dependent
- 1Claims Zastrzeżenia patentowe 1. A device for coding an audible signal configured to:1. Urządzenie do kodowania sygnału dźwiękowego skonfigurowane do: otrzymywania większej ilości elementów dźwiękowych ze źródła dźwięku z co najmniej jednego mikrofonu umieszczonego lub skierowanego na źródło dźwięku;receiving more sound elements from the sound source from at least one microphone placed or directed to the sound source;generating a first sound signal containing a greater number of sound elements from the sound source;generowania pierwszego sygnału dźwiękowego zawierającego większą ilość elementów dźwiękowych ze źródła dźwięku;otrzymywania mniejszej ilość elementów dźwiękowych ze źródła dźwięku z co najmniej jednego dodatkowego mikrofonu umieszczonego lub skierowanego z dala od źródła dźwięku;oraz generowania drugiego sygnału dźwiękowego zawierającego mniejszą ilość elementów dźwiękowych ze źródła dźwięku. receiving a smaller number of sound elements from a sound source from at least one additional microphone positioned or directed away from the sound source;and generating a second audio signal including fewer sound elements from the sound source.
- 5An apparatus for decoding a scalable encoded audio signal configured to:5. Urządzenie do dekodowania skalowalnego kodowanego sygnału dźwiękowego skonfigurowane do: dividing the scalable encoded audio signal into at least a first scalable encoded audio signal and a second scalable encoded audio signal;podziału skalowalnego kodowanego sygnału dźwiękowego na co najmniej pierwszy skalowalny kodowany sygnał dźwiękowy i drugi skalowalny kodowany sygnał dźwiękowy;decoding the first scalable encoded audio signal from at least one microphone positioned or directed towards the sound source to generate a first sound signal containing a greater amount of sound components from the sound source;and decoding the second scalable encoded audio signal from at least one downstream microphone positioned or directed away from the sound source to generate a second audio signal containing a lesser amount of sound components from the audio source. dekodowania pierwszego skalowalnego kodowanego sygnału dźwiękowego z co najmniej jednego mikrofonu umieszczonego lub skierowanego w kierunku źródła dźwięku dla generowania pierwszego sygnału dźwiękowego zawierającego większą ilość składników dźwiękowych ze źródła dźwięku;oraz dekodowania drugiego skalowalnego kodowanego sygnału dźwiękowego z co najmniej jednego dalszego mikrofonu umieszczonego lub skierowanego z dala od źródła dźwięku dla generowania drugiego sygnału dźwiękowego zawierającego mniejszą ilość składników dźwiękowych ze źródła dźwięku.
- 7Urządzenie według któregokolwiek z zastrz.5 do 6, ponadto skonfigurowane do generowania co najmniej pierwszej kombinacji pierwszego sygnału dźwiękowego i drugiego sygnału dźwiękowego i przesyłania pierwszej kombinacji do pierwszego głośnika. The apparatus of any of claims 5 to 6, further configured to generate at least a first combination of the first audio signal and the second audio signal and transferring the first combination to the first loudspeaker.
- 9A device according to any one of claims 1-8. 5 to 8, wherein at least one of the first scalable encoded audio signal and the second scalable coded audio signal includes at least one of:9. Urządzenie według któregokolwiek z zastrz. 5 do 8, przy czym co najmniej jeden z pierwszego skalowalnego kodowanego sygnału dźwiękowego i drugiego skalowalnego kodowanego sygnału dźwiękowego obejmuje co najmniej jedno z: advanced audio coding, AAC;zaawansowane kodowanie dźwięku, AAC;MPEG-1 layer 3, MP3;MPEG-1 warstwa 3, MP3;base line coding by ITU -T speech codec with built-in variable bit rate, EV-VBR;kodowanie linii bazowej przez ITU -T kodek mowy o wbudowanej zmiennej przepływności, EV-VBR ;coding by adaptive multi-pass bandwidth, AMR-WB;kodowanie przez adaptacyjne wieloprzepustowe szerokie pasmo, AMR-WB;ITU-T G.729.1 (G.722.1, G.722.1 C);ITU-T G.729.1 (G.722.1, G.722.1 C);coding by generation of comfort noise, CNG;and coding by adaptive multi-pass wide band, AMR-WB +. kodowanie przez generację szumu komfortu, CNG;oraz kodowanie przez adaptacyjne wieloprzepustowe szerokie pasmo, AMR-WB+.
- 10A method of coding a sound signal comprising:10. Sposób kodowania sygnału dźwiękowego obejmujący: - having more sound elements from the sound source from at least one microphone placed or directed to the sound source;-19otrzymywanie większej ilości elementów dźwiękowych ze źródła dźwięku z co najmniej jednego mikrofonu umieszczonego lub skierowanego na źródło dźwięku;generating a first sound signal containing a greater number of sound elements from the sound source;generowanie pierwszego sygnału dźwiękowego zawierającego większą ilość elementów dźwiękowych ze źródła dźwięku;otrzymywanie mniejszej ilości elementów dźwiękowych ze źródła dźwięku z co najmniej jednego dodatkowego mikrofonu umieszczonego lub skierowanego z dala od źródła dźwięku;oraz generowanie drugiego sygnału dźwiękowego zawierającego mniejszą ilość elementów dźwiękowych ze źródła dźwięku. receiving fewer sound elements from the sound source from at least one additional microphone positioned or directed away from the sound source;and generating a second audio signal including fewer sound elements from the sound source.
- 1516. Sposób według któregokolwiek z zastrz. 14 do 15, ponadto obejmujący generowanie co najmniej pierwszej kombinacji pierwszego sygnału dźwiękowego i drugiego sygnału dźwiękowego i przesyłanie pierwszej kombinacji do pierwszego głośnika. A method according to any one of claims 1-16. 14 to 15, further comprising generating at least a first combination of the first audio signal and a second audio signal and transmitting the first combination to the first loudspeaker.
- 1718. Sposób według któregokolwiek z zastrz. 14 do 17, przy czym co najmniej jeden z pierwszego skalowalnego kodowanego sygnału dźwiękowego i drugiego skalowalnego kodowanego sygnału dźwiękowego obejmuje co najmniej jedno z:18. The method according to any one of claims 1-18. The system of any one of claims 14 to 17, wherein at least one of the first scalable encoded audio signal and the second scalable encoded audio signal includes at least one of: advanced audio coding, AAC;zaawansowane kodowanie dźwięku, AAC;MPEG-1 layer 3, MP3;MPEG-1 warstwa 3, MP3;base line coding by ITU -T speech codec with built-in variable bit rate, EV-VBR;kodowanie linii bazowej przez ITU -T kodek mowy o wbudowanej zmiennej przepływności, EV-VBR ;coding by adaptive multi-pass bandwidth, AMR-WB;kodowanie przez adaptacyjne wieloprzepustowe szerokie pasmo, AMR-WB;ITU-T G.729.1 (G.722.1, G.722.1 C);ITU-T G.729.1 (G.722.1, G.722.1 C);coding by generation of comfort noise, CNG;and coding by adaptive multi-pass wide band, AMR-WB +. kodowanie przez generację szumu komfortu, CNG;oraz kodowanie przez adaptacyjne wieloprzepustowe szerokie pasmo, AMR-WB+. -21 □ governance of the microne speaker t-lektronif soup -21□rządzenie głośnik mikruluny t-lektronif zup OAC OAC RKiTa procesor [Mmięc dane programu _ za kodowan? darte ;dane j warstwy 'rdzeniowej dane rozszerzonej wa rstwy dekodowanie i warstwy rdzeniowej , db tworzenia i „bliskiego” sygnału. RKiTa processor [Mmw program data _ to be coded? darte;data 'core layer data of the extended decoding layer and core layer, db creation and' close 'signal. • dźwiękowego j ’ dekodowanie i rozszerzonej warstwy+ ‘ syntetyzowanej « warstwy rdzeniowej dla tworzenia • sound j 'decoding and expanded layer +' synthesized 'core layer for creation And, "distant" audio signal adder r selektor'bliskich <distant sound signals mixed sound signal ik I , „odległego" sygnału dźwiękowego sumator r selektor‘bliskich < odległych sygnałów dźwiękowych . mieszany sygnał dźwiękowy i-k
Independent claims7
137 paragraphs, as filed
The invention relates to a device and method for encoding sound and reproduction, and in particular, but not exclusively, devices for coded signals and speech and auditory signals.
BACKGROUND OF THE INVENTION [0002] An audio signal, such as speech or music, is encoded, for example, to allow efficient transmission or efficient storage of audio signals.
[0003] Sound encoders and decoders are used to represent audio signals, such as music or background noise. These types of encoders usually do not use the speech model in the encoding process, but rather use processes to represent all types of audio signals, including speech.
[0004] Speech encoders and decoders (codecs) are typically optimized for speech signals, and can operate either in a fixed or variable bit rate.
[0005] The audio codec can also be configured to operate at a variable bit rate. At lower bit rates, such audio codecs can operate with speech signals in the equivalent of a compression variable for encoders designed exclusively for speech coding. At higher bit rates, the audio codec can encode any signal, including music, background noise and speech, with higher quality and performance.
[0006] In some audio codecs, the input signal is divided into a limited number of bands. Each of the bandwidth signals can be quantized. It is known from the psychoacoustic theory that higher frequencies in the spectrum are noticeably less important than low frequencies. In some audio codecs, this is reflected by the allocation of bits, with fewer bits being assigned to high frequency signals than to low frequency signals.
[0007] One of the emerging trends in the field of carrier encoding is so-called layered codecs, for example ITU -T speech / sound codec with embedded variable bit rate (EV-VBR) and ITU-T scalable video codec (SVC - Scalable Video Codec). Scalable data carriers consist of a core that is always necessary to allow recovery at the receiving end, and one of several extended layers that can be used to improve the quality of the data being restored (e.g., improving the quality or increasing the fault tolerance of transmitting, etc. ) [0008] Scalability of these codecs can be used at the transmission level e.g. for controlling network bandwidth or broadcasting shaping. multicast) bit streams to facilitate operation with participants behind access lines with different bandwidth. At the application level, scalability can be used to control variables such as computational complexity, coding delay or desired level
-2jakości. It should be noted that although in some scenarios scalability may be used at the transmitting endpoint, there are also operational scenarios in which it is more preferable that the intermediate network element is able to perform scaling.
[0009] Most real-time speech coding takes place with respect to mono signals, but for some high-end video and audio systems, stereo coding has been used to produce speech reproduction better for the listener. Traditional stereo speech coding involves encoding a separate left and right signal channel that positions the source at some point in the scene. Common use of speech stereo is binaural encoding, while a sound source (such as a voice from a speaker) is detected by two microphones that are placed on the left and right ear on a simulated reference head.
[0010] Decoding and transmission (or storage) of the generated signal requires more bandwidth and processing power because there are more signals to encode and decode compared to a conventional monophonic recording source. One way to reduce the amount of bandwidth (storage) used in the stereo coding methods is to require the encoder to mix the left and right channels together, and then encode the generated (combined) mono signal as the core layer. Information about the differences between left and right channels can be coded as separate bit streams or extended layer (enhancement layer). This type of coding, however, produces a mono signal on the decoder with a sound quality worse than the traditional mono signaling from a single microphone (located near the mouth, for example) because two microphone signals connected together receive much more background or background interference than a single microphone placed near the sound source ( for example, paragraph). This makes the quality of the backward compatible 'mono' output using the existing playback system worse than the original mono recording and playback process.
[0011] Furthermore, the binaural positioning of the stereo microphone in which the microphones are placed in simulated positions on the ears of the simulated head may produce a disturbing sound signal for the listener, especially when the sound source is moving suddenly or abruptly. For example, in a system in which the microphone is placed near the loudspeaker source, poor listening quality can be generated simply when the speaker turns his head causing sudden and sudden switching of the left and right output signals.
An example of a known binaural encoding system is disclosed by Faller et al .: "Binaural coding system - Part II: Schemes and applications", IEEE Transactions on Speech and Audio Processing, New York, vol. 11, no. 6, November 1, 2003, pages 520-531, XP011104739
Summary of the invention
[0013] The application proposes a mechanism that enables efficient imaging of stereo image for environments such as conference activities and the use of mobile devices by users.
[0014] Embodiments of the invention are intended to solve or at least partially alleviate the above problem.
[0015] According to a first object of the invention, an apparatus for coding an audible signal according to claim 1 is disclosed.
[0016] According to a second object of the invention, a decoding device according to claim 5 may be provided.
[0017] According to a third aspect of the invention, a method for coding an audio signal according to claim 10 is disclosed.
[0018] According to a fourth aspect of the invention, a decoding method of a scalable encoded audio signal according to claim 14 is disclosed.
[0019] Particular embodiments are defined in the dependent claims.
[0020] The encoder may comprise a device as described above.
[0021] The decoder may comprise a device as described above.
[0022] The electronic device may comprise a device as described above.
[0023] A set of integrated circuits may comprise a device as described above.
Brief Description of the Drawings [0024] For a better understanding of the invention, reference will now be made, by way of example, of the accompanying drawings, in which:
Fig. 1 schematically shows an electronic device using embodiments of the invention;
Fig. 2 schematically shows an audio codec system using embodiments of the invention;
Fig. 3 schematically shows a coding part of the sound coding system shown in Fig. 2;
4 is a schematic block diagram illustrating the operation of an embodiment of an audio encoder as shown in FIG. 3 according to the invention;
Fig. 5 schematically shows a decoding part of the sound coding system shown in Fig. 2;
Fig. 6 is a block diagram illustrating the operation of an embodiment of an audio encoder as shown in Fig. 5 according to the invention; and Fig. 7a-7h show the possible locations of a microphone / loudspeaker according to an embodiment of the invention.
Advantages of preferred embodiments of the invention [0025] In the following, possible mechanisms for providing a scalable audio coding system are described in more detail. Accordingly, reference is made to Fig. 1, which is a block diagram of an exemplary electrical apparatus 10, which may include a codec according to an embodiment of the invention.
[0026] The electronic device 10 may be, for example, a mobile terminal or user equipment of a wireless communication system.
The electronic device 10 comprises a microphone 11 which is connected by an analog-to-digital converter 14 to the processor 21. The processor 21 is further connected via an analog-digital converter 32 to the loudspeakers 33. The processor 21 is further connected to the transceiver ( TX / RX) 13, with a user interface (UI) 15 and memory 22.
[0028] The processor 21 may be configured to support different program codes. The program codes used include a code coding for sound for encoding a combined sound signal and a code for extracting and encoding auxiliary information relating to spatial information from a plurality of channels. The program codes 23 used further include a sound decoding code. The program codes 23 used may be stored, for example, on the storage medium 22 and, if necessary, downloaded by the processor 21. The storage medium 22 may further provide a section 24 for storing data, e.g. data, which have been coded in accordance with the invention.
[0029] The coding and decoding code in the embodiments of the invention may be implemented in the form of computer hardware or firmware.
[0030] The user interface allows the user to enter commands into the electronic device 10, for example by means of a keypad, and / or receive information from the electronic device 10, for example by means of a display. The transceiver 13 allows communication with other electronic devices, e.g. via a wireless communication network.
[0031] It should be understood that the structure of the electronic device 10 can be supplemented and varied in many ways.
[0032] The user of the electronic device 10 may use microphones 11 to input a speech signal to be transmitted to another electronic device and to be stored in a section 24 of data memory 22. Therefore, the respective application has been activated by the user via the interface 15. This application , which can be run by the processor 21 causes the processor 21 to decode the code stored in the memory 22.
[0033] The analog-to-digital converter 14 converts the input analog audio signal into a digital audio signal and provides a digital audio signal to the processor 21.
[0034] The processor 21 may then process the digital audio signal in the manner described with reference to Figures 3 and 4.
[0035] The resulting bit stream is delivered to the transceiver 13 for transmission to another electronic device. Alternatively, the encoded data could be stored in the memory data section 24, e.g. for subsequent transmission or for subsequent presentation by the same electronic device 10.
[0036] The electronic device 10 may also receive a bit stream with correspondingly decoded data from another electronic device via its transceiver 13. In this case, the processor 21 may execute the decoding program code stored in the memory 22. The processor 21 decodes the received data, and provides decoded data to the digital-to-analogue converter 32. The digital / analogue converter 32 converts digitally coded data into analogue audio data and emits them via loudspeakers 33. The execution of the decoder program code could also be called by an application that has been selected by the user via the user interface.
[0037] The received coded data may also be stored instead of instant presentation by the loudspeaker (i) 33 in section 24 of the memory data 22, for example to allow subsequent presentation or transfer to yet another electronic device.
[0038] It should be noted that the schematic structures described in Figures 3 and 5 and the method steps in Figs. 4 and 6 show only a part of the operation of a complete audio codec, as exemplified for example in the case of the electronic device shown in Figure 1.
[0039] With reference to Figures 7a and 7b, examples of a microphone arrangement suitable for embodiments of the invention are shown. In Fig. 7a, an exemplary arrangement of the first and second microphones 11a and 11b is shown. The first microphone 11 a is located near the first sound source, for example a conference speaker 701a. The audio signals received from the first microphone 11a may be a designated "near" signal. The second microphone 11b is also shown remote from the sound source 701a. The audio signal received from the second microphone 11b may be defined as a "distant" audio signal.
[0040] As will be understood by those skilled in the art, the difference between placing the microphone to produce a "near" and "distant" sound signal is a relative difference from the sound source 701a. Thus, for the second audio source, the further conference speaker 701b, the audio signal from the second microphone 11b is a "near" sound signal while the sound signal from the first microphone 11a is considered a "remote" signal.
[0041] With reference to Fig. 7b, an example of placing a microphone to generate auditory "near" and "distant" signals for a typical mobile communication device may be shown. In such an arrangement, the microphone 11a generating a "near" sound signal is located close to the sound source 703, which could be for example in
- a microphone-like position in conventional mobile communication devices, and thus be located close to the user's mouth 705 of the mobile communication device, while the second microphone 11b generating a "remote" sound signal is located on the opposite side of the mobile communication device 707 and is configured to receive audio signals from the environment, being protected from capturing direct sound tracks from sound source 703 through the mobile communications device 707 itself.
[0042] Although FIG. 7a shows the first microphone 11a and the second microphone 11b, it will be apparent to those skilled in the art that "near" and "distant" audio signals may be generated by any number of microphones.
[0043] For example, "near" and "distant" audible signals can be generated using a single microphone with directional elements. In this embodiment, it may be possible to create a close signal using the directional elements of the microphone directed to the sound signal and the generation of a "distant" sound signal from the directional elements of the microphone directed away from the sound source.
[0044] Furthermore, in other embodiments of the invention, it is possible to use multiple microphones for generation of "near" and "distant" audio signals. In these embodiments, there may be pre-treatment of microphone signals for generation of a "near" sound signal by mixing audio signals received from a microphone (microphones) near the sound source and a "distant" audio signal by mixing audio signals received from the microphone (s) placed or away from the sound source.
[0045] Although above and below "near" and "distant" signals are discussed as being either directly generated by microphones or generated by signals generated by pre-processed signals generated by the microphone, it is noteworthy that "near" and "distant" signals "signals may be pre-recorded / stored or received signals other than directly from the microphone / pre-processor.
[0046] Furthermore, although both above and below the coding and decoding of "near" and "distant" audio signals are discussed, it is noteworthy that in some embodiments of the invention more than two audio signals may be coded. For example, in one embodiment, multiple audio "close" signals or multiple "remote" signals may be included. In other embodiments of the invention, there may be a primary & quot; near & quot; sound signal and complex sub-original & quot; near & quot; sound signals wherein the signal originates from a position between & quot; close & quot; and & quot; distant & quot; sound signals.
[0047] For discussion of the remainder of the invention, coding and decoding for two microphones / close and distant channels of the encoding and decoding process will be discussed.
[0048] Referring to Fig. 7c and 7d, examples of a loudspeaker arrangement suitable for embodiments of the invention are shown. Figure 7c shows a traditional or older speaker layout.
- The user 705 has a loudspeaker 709 located near one of the ears of the user 705. In the arrangement shown in Fig. 7c, a single speaker 709 may provide a signal "close" to the selected ear. In some embodiments of the invention, a single speaker 709 may provide a "near" signal and a processed or filtered component of a "remote" signal to add space to the output signal.
[0049] In Fig. 7d, a user 705 is provided with a headset 711 including a pair of loudspeakers 711a and 711b. In this arrangement, the first loudspeaker 711a may emit a "near" signal and the second loudspeaker 711b may emit a "remote" signal.
[0050] In other embodiments of the invention, both the first loudspeaker 711a and the second loudspeaker 711b provide a combination of "near" and "distant" signals.
In some embodiments of the invention, the first loudspeaker 711a has a provided combination of "near" and "distant" audio signals so that the first loudspeaker 711a receives a "near" signal and a modified "distant" audio signal. The second loudspeaker 711b receives a "distant" sound signal and a modified "near" sound signal. In this embodiment, the terms α and β mean that filtering or processing has been performed on the audio signal.
[0052] With reference to Fig. 7e, a further example of both a microphone and a loudspeaker arrangement is provided, suitable for embodiments of the invention. In such an embodiment, the user 705 is provided with a first headset / headset unit comprising a loudspeaker 713a and a microphone 713b which is located near the preferred ear and mouth, respectively. The user 705 is further provided with another separate Bluetooth device 715 having a separate Bluetooth speaker 715a and a separate microphone 715b of the Bluetooth device. The microphone 715b of the separate Bluetooth device 715 is configured such that it does not directly receive signals from the user's audio source 705, in other words from the user 705.
[0053] Referring to FIG. 7f, a further example of a microphone and loudspeaker arrangement suitable for embodiments of the invention is shown. FIG. 7f shows a cable that can connect or may not be directly connected to an electronic device. Cable 717 includes a 729 speaker and several separate microphones. Microphones are arranged along the length of the wire and form a microne array. Thus, the first microphone 727 is located near the loudspeaker 729, the second microphone 725 is positioned along the conduit 717 from the first microphone 727. The third microphone 723 is located downstream of the conduit 717 from the second microphone 725. The fourth microphone 721 is located further down the conduit 717 from the third microphone 723. The fifth microphone 719 is located further down the conduit 717 from the fourth microphone 721. The arrangement of the microphones may be in a linear or non-linear configuration depending on the embodiments of the invention. In such an arrangement, the "near" signal can be produced by mixing the combination
Audio signals received by microphones closest to user 705. A & quot; distant & quot; sound signal may be generated by mixing a combination of an audio signal received from microphones most distant from user 705. As described above in some embodiments, each of the microphones may be used for generating a separate audio signal, which is then processed as described in detail below.
[0054] In these embodiments, those skilled in the art will appreciate that the actual number of microphones is negligible. Thus, many microphones in any arrangement can be used in embodiments of the invention to capture the sound field, and sound processing methods can be used to recover "near" and "distant" signals.
[0055] With reference to Fig. 7g, a further example of a microphone and loudspeaker arrangement suitable for embodiments of the invention is shown. Figure 7g shows a Bluetooth device connected to the preferred user ear 705. The Bluetooth device 735 includes a "near" microphone 731 located near the user 705. The Bluetooth device 735 further includes a & quot; distant & quot; microphone 733 spaced from the nearest (near) position of the microphone 731.
[0056] Further, with reference to Fig. 7h an example of a microphone / loudspeaker arrangement suitable for embodiments of the invention is provided. In Fig. 7h, user 705 supports a headset 751. The headset includes a binaural headset stereo with the first speaker 737 and the second speaker 739. The headset is then shown with a pair of microphones. The first microphone 741, which is shown in Fig. 7h as placed 100 millimeters from the speaker 739 and the second microphone 743 placed 200 millimeters from the speaker 739. In this arrangement, the first speaker 737 and the second speaker 739 can be configured depending on the reproduction system described in with reference to Fig.7d.
[0057] Furthermore, the microphone arrangement of the first microphone 741 and the second microphone 743 may be configured such that the first microphone 741 is configured to receive or generate a "close" component of the audio signal and the second microphone 743 is configured to generate a "distant" audio signal.
The general operation of audio codecs used by embodiments of the invention is shown in Figure 2. General audio coding / decoding systems include encoders and decoders, as shown schematically in Figure 2. System 102 is shown with encoder 104, storage or carrier channel 106 and decoder 108.
[0059] The encoder 104 compresses the input audio signal 110 to form a 112 bit stream that is either stored or transmitted over the carrier channel 106. The 112 bits stream can be received inside the decoder 108. The decoder 108 decompresses the 112 bit stream and produces the output audio signal 114. The bit stream rate of 112 bits and the quality of the audio output signal 114 with respect to the input signal 110 are the main features that determine the system coding 102.
[0060] Fig. 3 schematically illustrates an encoder 104 according to an embodiment of the invention.
[0061] The encoder 104 includes a core codec processor 301 that is configured to receive a "near" sound signal, for example, as shown in FIG. 3, a sound signal from the microphone 11a. The core codec processor is further adapted to be connected to the multiplexer 305 and the processor 303 of the extended layer.
[0062] The advanced layer processor 303 is further configured to receive a & quot; distant & quot; sound signal, which is shown in Fig. 3 as a sound signal received from the microphone 11b. The enhanced layer processor is further configured to be connected to the multiplexer 305. The multiplexer 305 is configured to output a bit stream such as the bitstream of 112 bits shown in FIG. 2.
[0063] The operation of these components is described in more detail with reference to the block diagram of fig.4 showing the operation of the encoder 104.
[0064] "Near and" distant "sound signals are obtained by the encoder 104. In a first embodiment of the invention, the" near "and" distant "audio signals are digitally sampled signals. In another embodiment of the invention," near "and" distant "signals The sound signals may be an analogue sound signal received from microphones 11a and 11b which are converted from analogue to digital (A / D). In further embodiments of the invention, the audio signals are converted from a digital signal (PCM - pulse code modulation) to the digital signal (AM) The reception of sound signals from microphones is shown in FIG. 4 in step 401.
[0065] As shown above, in some embodiments of the invention, "near" and "distant" audio signals may be processed from a microphone system (which may include more than 2 microphones). The audio signals received from the microphone system, such as the arrangement shown in FIG. 7f, can generate "near" and "distant" audio signals using signal processing methods such as beamforming, speech enhancement, source tracking, and noise suppression. Thus, in embodiments of the invention, a "near" sound signal is selected and determined to include preferably (clean) speech signals (in other words an acoustic signal without excessive interference), and a "distant" generated audio signal is selected and determined so
[0066] The core codec processor 301 receives a "near" sound signal to be encoded and emits coding parameters that represent the encoded core level signals. Processor 301 of the core codec may also generate for internal use a synthesized "near" sound signal (in other words, a "near" sound signal is coded to parameters and then the parameters are decoded in a mutual process to produce a synthesized "near" sound signal).
[0067] The core codec processor 301 may use any suitable coding technique to produce the core layer.
[0068] In a first embodiment of the invention, the core codec processor 301 produces a core layer using an encoder of the embedded variable bit rate (EB-VBR).
- embedded variable bit rate codec).
[0069] In other embodiments of the invention, the core codec processor may be the coding of an algebraic code excited linear prediction (ACELP) and is configured to emit a bitstream of typical ACELP parameters.
[0070] It should be understood that embodiments of the invention may equally use sound and speech based codecs to represent the core layer.
[0071] The production of signals encoded by the core layer is shown in Fig. 4 in step 403. The signal encoded by the core layer is transferred from the core codec 301 to the mux 305.
The processor 303 of the expanded layer receives a & quot; distant & quot; sound signal and generates the output signals of the expanded layer from a & quot; distant & quot; sound signal. In some embodiments of the invention, the processor of the extended layer performs a similar encoding on the "remote" sound signal that is performed by the core code processor 301 on the "near" sound signal. In other embodiments of the invention, a & quot; distant & quot; sound signal is coded using the appropriate coding method. For example, a "distant" audio signal may be coded using similar schemes as those used in discontinuous transmission (DTX), with the comfort noise generation (CNG) coding being used in low bit rates, coding of the excited prediction of linear algebraic code (ACELP) and modified discrete cosine transformation (MDCT), the remaining coding methods can be applied to coders with medium and high bit rate efficiency. In some embodiments of the invention, the quantization of a "distant" audio signal may also be specifically selected to match the type of signal.
[0073] In some embodiments of the invention, the enhanced layer processor is configured to receive a synthesized & quot; near & quot; sound signal and & quot; distant & quot; sound signal. An advanced layer processor 303 may, in embodiments of the present invention, generate an encoded bit stream, also known as an extended layer, depending on the & quot; distant & quot; sound signal and the synthesized & quot; near & quot; sound signal. For example, in one embodiment of the invention, the enhanced layer processor subtracts the synthesized "near" signal from the "distant" audio signal, and then encodes the difference of audio signals, e.g. converting the time domain into a frequency domain and encoding the frequency domain result as an expanded layer.
[0074] In other embodiments of the invention, the extended layer processor 303 is configured to receive a & quot; distant & quot; sound signal, a & quot; near & quot;
- a sound signal and a "near" sound signal and generate a stream of bits of the extended layer depending on the combination of these three input signals.
[0075] Thus, the audio signal coding apparatus according to the invention is configured to generate the first scaled coded signal layer from the first audio signal, generate the second scaled coded signal layer from the second audio signal, and combine the first and second scaled coded signal layer to create a third scaled coded signal signal layer.
[0076] The device in the embodiments may further be configured to generate a first audio signal including a greater amount of sound components from the sound source and generate a second audio signal including fewer sound components by the sound sources.
[0077] The apparatus may further in embodiments of the invention be configured to receive more sound components from the sound source from at least one microphone positioned or directed to the sound source and receive a smaller amount of audio components from the audio source from at least one further microphone positioned or directed away from the sound source.
[0078] For example, in some embodiments of the invention, at least a portion of the bit stream signals of the extended layer is generated depending on the synthesis of a "near" sound signal and a "near" sound signal and a portion of the bitstream signal of the expanded layer is dependent only on the "distant" audio signal . In this embodiment of the invention, the extended layer processor 303 performs a similar processing of a core codec of a "distant" audio signal to produce a "remote" encoded layer, which is generated by the core layer processor 301 on a "near" sound signal, but for the "remote" portion of the audio signal. .
[0079] In further embodiments of the invention, a "close" synthesized signal and a "distant" audio signal are transformed into a frequency domain and the difference between two frequency domain signals is then coded to produce the extended layer data.
[0080] In embodiments of the invention, the use of a frequency-coding frequency band to be transformed in the frequency domain may take place using any suitable converter such as discrete cosine transform (DCT), discrete Fourier transform (DFT), fast Fourier transform (FFT).
[0081] In some embodiments of the invention, extended layers of ITU-T speech / sound codecs with an embedded bit rate variable (EV-VBR) and extended ITU-T layers of a scalable video codec (SVC) may be generated.
[0082] Furthermore, embodiments of the invention may include, but are not limited to, creating an expanded layer using a variable multi-band wide band (VMR-WB), ITU-T G.729, ITU-TG. 729.1, ITU-T G.722.1, ITU G.722.1C, adaptive multi-pass bandwidth (AMR-WB)
-12adaptive multirate wideband) and adaptive multi-pass wide band + (AMRWB +).
[0083] In other embodiments, to extract a compound between a synthesized & quot; near & quot; signal and a & quot; distant & quot; signal to generate a preferred encoded data signal of the extended layer, any suitable codec may be used.
[0084] The generation of the expanded layer is shown in Fig. 4 in step 405.
[0085] The extended layer data is transferred from the 303 processor of the expanded layer to the multiplexer 305.
[0086] The multiplexer 305 then multiplexes the core layer obtained from the core codifier processor 301 and the expanded layer or extended layers from the extended layer processor 303 to produce an encoded stream of 112 bits. Multiplexing for core and extended layers to produce a bitstream is shown in FIG. 4 in step 407.
[0087] To further assist in understanding the present invention, the decoder 108 operation is shown with respect to embodiments of the invention with respect to the decoder schematically shown in fig. 5 and a block diagram illustrating the decoding operation in fig. 6.
[0088] The decoder 108 includes an input 502 through which an encoded stream of 112 bits can be received. Input 502 is connected to a 1401bit receiver / demultiplexer. The demultiplexer 1401 is configured to strip the core layer and the extended stream of 112 bits. The core layer data is transferred from the demultiplexer 1401 to the core coder decoder processor 1403 and the extended layer data is transferred from the demultiplexer 1401 to the decoder processor 1405 of the extended layer.
Furthermore, the core coder decoder processor 1403 is connected to an adder and a sound mixer 1407, and then to the processor 1405 of the expanded layer.
[0090] Processor 1405 of the decoder of the extended layer is connected to the adder and sound mixer 1407. The output of the adder and the sound mixer 1407 is connected to the audio signal 114 of the output.
[0091] The reception of the encoded bitstream of the multiplexer is shown in FIG. 6 in step 501.
[0092] The decoding of the bit stream and the separation into core layer data and expanded layer data is shown in FIG. 6 in step 503.
[0093] A core coder decoder processor 1403 performs a mutual process in the core coding processor 301 as shown on the encoder 104 to produce a synthesized "near" sound signal. It is transmitted from processor 1403 to the core encoder decoder to the adder and sound mixer 1407.
[0094] Furthermore, in some embodiments of the invention a synthesized "near" sound signal is also transmitted to the decoder of the enhanced layer decoder processor 1405.
[0095] The decoding of the core layer to produce a synthesized "near" sound signal is shown in FIG. 6 in step 505.
[0096] The decoder of the enhanced layer decoder processor 1405 receives at least the signals of the expanded layer from the demultiplexer 1401. Furthermore, in some embodiments of the decoder of the enhanced layer decoder processor 1405 receives a synthesized "near" sound signal from the core codec decoder processor 1403. Furthermore, in some embodiments of the invention, the processor of the enhanced layer decoder processor 1405 receives a synthesized "near" sound signal from the core codec decoder processor 1403 and some decoded core layer parameters.
[0097] The decoder of the enhanced layer decoder processor 1405 then performs a reciprocal process for the extended encoder layer 104 generated within the processor 303 to produce at least a "distant" audio signal.
[0098] In some embodiments of the decoder of the enhanced layer decoder processor 1405, it may further produce additional components of the audio signal for a "near" sound signal. The production of a "distant" audio signal from decoding the expanded layer (and in some embodiments of the synthesized core layer) is shown in FIG. 6 in step 507.
[0099] A "remote" audio signal from the decoder processor of the extended layer is forwarded to the adder and sound mixer 1407.
[0100] The adder and sound mixer 1407 upon receipt of the synthesized "near" sound signal and the decoded "distant" audio signal then generates a combined and / or selected combination of the two received signals and the signals are mixed sound signals at the output audio output.
[0101] In some embodiments of the invention, the combiner and audio signal mixer receives further information from either the input bit stream via the demultiplexer 1401 or already knows the placement of the microphones used to generate "near" and "distant" audio signals to digitally process the synthesized "close" signals. and decoded & quot; distant & quot; sound signals with respect to the position of the loudspeakers or the placement of headphones for the listener to produce a correct or advantageous combination of sounds from "near" and "distant" audio signals.
[0102] In some embodiments of the invention, the combiner and audio signal mixer may emit only a "near" sound signal. In such an embodiment, it produces an acoustic signal similar to the previous one
- encoding / decoding, and therefore produces results that are backward compatible with the present audio signals.
[0103] In some embodiments of the invention, "near" and "distant" signals are decoded from the bitstream and the amount of "remote" signal is mixed with a "near" signal to obtain a pleasantly sounding monaural auditory background. In such an embodiment of the invention, it is possible for the listener to be aware of the environment of the sound source without interfering with the understanding of the sound source. It will also enable the receiving person to adapt the amount of "surroundings" to their preferences.
[0104] The use of "near" and "distant" signals produces a signal that is more stable than a conventional binaural process and is less susceptible to movement of the sound source. Furthermore, embodiments of the invention have the advantage that they do not require an encoder to be connected to multiple microphones to produce a pleasing sounding sound.
Thus, of course, it follows from the foregoing that, in embodiments of the invention, the decoding scalable coded audio signal device is configured to divide the scalable coded audio signal into at least a first scalable encoded audio signal and a second scalable encoded audio signal. The apparatus is further configured to decode the first scalable encoded audio signal to generate the first audio signal. The device is also configured to decode a second scalable encoded audio signal to produce a second audio signal.
[0106] Furthermore, in embodiments of the invention, the apparatus may further be configured to: emit at least a first audio signal to the first loudspeaker.
[0107] As described above, in some embodiments, the apparatus may further be configured to generate at least one combination of the first audio signal and the second audio signal and transmit the first combination to the first loudspeaker.
[0108] The apparatus may further be configured in other embodiments to generate a further combination of the first audio signal and the second audio signal and transfer the second combination to the second loudspeaker.
[0109] It should be understood that although the invention has been described, for example, with reference to the core layer and the expanded layer, it should be understood that the invention may refer to further expanded layers.
[0110] The embodiments of the invention described above describe the codec with respect to the separate coding devices 104 and decoding 108 to facilitate understanding of the processes involved. However, it should be understood that the device, structures and operations can be used as single devices / structures / coded coded operations. Furthermore, in some embodiments of the invention, the coder and the decoder may share some or all of the common elements.
[0111] As mentioned earlier, although the above process describes a single-core coded audio signal and an encoded audio signal of a single expanded layer, the same method can be used to synchronize two bit streams using the same or a similar transmission protocol packet.
[0112] Although the above examples describe embodiments of the invention operating in a codec within the electronic device 610, it is noteworthy that the invention, as described below, can be used as part of any codec with variable or adaptive speed of sound (or speech). . Thus, embodiments of the invention may be implemented in the form of a codec that may use audio coding via fixed or wired communication paths.
[0113] Thus, the user equipment may comprise a sound codec such as that described in the above embodiment of the invention.
[0114] It should be noted that the concept of user equipment is intended to include all appropriate types of wireless user equipment, such as mobile phones, portable data processing devices or portable internet browsers.
[0115] Furthermore, the elements of a public land mobile network (PLMN) may also include audio codecs as described above.
[0116] In general, various embodiments of the invention may be applied to a hardware platform or to target, software, logic circuits or any combination thereof. For example, some embodiments may be implemented in hardware, while other embodiments may be implemented on firmware or software that may be implemented as controllers, microprocessors, or other computing devices, although the invention is not limited thereto. While various embodiments of the invention may be represented and described as block diagrams, block diagrams, or other pictorial representations, it should be understood that these blocks, devices, systems, techniques or methods described herein may be used in, as non-limiting examples, hardware, software, firmware,
[0117] For example, embodiments of the invention may be used as a chipset, in other words, as a series of integrated circuits that communicate with each other. The chipset may include microprocessors adapted to run code, dedicated integrated circuits (ASICs), or programmable signal processors to perform the operations described above.
[0118] Embodiments of the invention may be implemented by a computer program executable by a data processor from a mobile device, such as a processor unit, or by hardware, or by a combination of software and hardware. Furthermore, in this regard, it should be noted that all blocks of the logic stream as shown in the figures
They may represent program steps, or combined logic circuits, blocks and functions, or combinations of program steps and logic circuits, blocks and functions.
[0119] The memory may be of any type suitable for the local technical environment and may be implemented using all appropriate data storage technologies such as semiconductor-based storage media, magnetic storage media, and systems, optical storage media and systems, fixed memory and removable memory. . Data processors may be of any type suitable for the local technical environment, and may include one or more general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), and processors based on a multi-core processor architecture, as not limiting examples.
[0120] Embodiments of the inventions may be used in various components, such as integrated circuit modules. Designing integrated circuits is basically a highly automated process. Complicated and powerful software tools are available to convert a logic level design to a semiconductor circuit pattern ready for pickling and forming on a semiconductor substrate.
[0121] Programs such as those provided by Synopsys, Inc. from Mountain View, Calif. and Cadence Design, from San Jose, California, they automatically guide and place components on a semiconductor chip using well-known design principles and a library of ready-made pattern modules. When the design of the semiconductor circuit pattern circuit has been completed, the resulting pattern, in a standardized electronic format (e.g., Opus, GDSII, or the like) can be sent to the semiconductor production plant or "factory" for production.
The above description is presented as a standard and non-limiting example of a complete and informative description of an exemplary embodiment of the invention. However, various modifications and adaptations may become apparent to those skilled in the art in the light of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications to embodiments of the invention still fall within the scope of the invention as defined in the appended claims.
15 members in 9 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008055776 | European Patent Office (EPO) | W | |
| 087502431 | – | – | – |
| WO2008EP55776 | – | – | – |
Members15
| Document | Office | Kind | |
|---|---|---|---|
| CA2721702A1 | Canada | A1 | |
| WO2009135532A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20110002086A | Republic of Korea | A | |
| EP2301017A1 | European Patent Office (EPO) | A1 | |
| US2011093276A1 | United States of America | A1 | |
| CN102067210A | China | A | |
| RU2010149667A | Russian Federation | A | |
| RU2477532C2 | Russian Federation | C2 | |
| CN102067210B | China | B | |
| KR101414412B1 | Republic of Korea | B1 | |
| US8930197B2 | United States of America | B2 | |
| CA2721702C | Canada | C | |
| EP2301017B1 | European Patent Office (EPO) | B1 | |
| ES2613693T3 | Spain | T3 | |
| PL2301017T3This record | Poland | T3 |
Numbers
- Publication
- 2301017
- Publication, DOCDB
- 2301017
- Publication, EPODOC
- PL2301017T
- Application
- 8750243
- Application, DOCDB
- 08750243
- Application, EPODOC
- PL08750243T
Titles2
- English
- AUDIO APPARATUS
- Polish
- Urządzenie akustyczne
Classification
- CPC, 1
- G10L19/008
- IPC, 4
- G10L19 008
- G10L19 24
- G10L19 00
- H03M7 30