Real-time objective voice analyzer
Abstract
Problem to be solved.To provide a method and an apparatus for real-time objective voice analysis.
Solution.The apparatus includes a sound quality analyzer for receiving at least one first signal and providing at least one second signal indicative of at least one non-intrusive estimate of a sound quality based on the at least one first signal.
Copyright (C)2006,JPO&NCIPI
Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
10 claims: 2 independent, 8 dependent
- 1A voice quality analyzer that receives at least one first signal and provides at least one second signal that indicates at least one non-intrusive evaluation of voice quality based on the at least one first signal. Equipment including. 少なくとも1つの第1の信号を受信し、前記少なくとも1つの第1の信号に基づき、音声品質についての少なくとも1つの非侵入型の評価を示す少なくとも1つの第2の信号を提供する音声品質アナライザを含む装置。
- 7Non-intrusive audio quality of the at least one processed audio signal based on the step of receiving at least one first signal indicating at least one processed audio signal and the at least one first signal. A method comprising the step of determining in, and the step of providing at least one second signal indicating the voice quality of the at least one processed voice signal. 少なくとも1つの処理された音声信号を示す少なくとも1つの第1の信号を受信する工程と、 前記少なくとも1つの第1の信号に基づき、前記少なくとも1つの処理された音声信号の音声品質を非侵入型で決定する工程と、 前記少なくとも1つの処理された音声信号の前記音声品質を示す少なくとも1つの第2の信号を提供する工程とを含む方法。
Independent claims2
25 paragraphs, as filed
The present invention generally relates to network systems, and more particularly to audio signals within a network.
Voice signals are transmitted by various network systems such as POTS (Plain Old Telephone System), Internet-based networks using VoIP (Voice over Internet Protocol), and wireless communication systems. Generally, when the original voice signal, eg, an acoustic signal generated by the voice of a first user, is transmitted to the ear of a second user via a network system, the signal is processed by a number of devices. For example, in a wireless communication network, the original voice signal is processed by a first mobile unit, a first base station, a network hub, a second base station, a second mobile unit, and other intermediate devices. Only then can the second user hear the processed audio signal.
Each device in the network, as well as wired and / or wireless channels that carry the processed audio signal, may modify the processed audio signal. Some of the modifications are desirable. For example, you can use various filters to remove unwanted noise from the processed audio signal, add comfort noise to the processed audio signal to remove unnatural silence, and remove the processed audio signal. For example, compressing to reduce the total amount of data transmitted. Some modifications of the processed audio signal are undesirable. For example, transmission errors can occur when the processed audio signal traverses the network. This error may cause gaps, unwanted noise, etc. in the processed audio signal. The processing of the original audio signal by the network system can reduce the quality of the processed audio signal, whether desirable or undesirable. Subjective techniques based on human perception can be used to evaluate the quality of processed audio signals. For example, a network system processes a database of original audio samples, provides the processed audio signal to a set of listeners, and the listener evaluates the processed audio signal based on ranks 1-5. can do. However, subjective techniques are time consuming and costly. In situations where subjective testing methods are costly and / or time consuming, for example, collecting voice databases, recruiting large listening teams, and paying for a statistically significant assessment of voice quality. Includes getting, preparing soundproof rooms and other equipment.
Objective methods can also be used to evaluate the quality of the processed audio signal. In a typical objective assessment of the quality of processed speech, commonly referred to as the intrusive method, the original speech signal is processed by a network system and then a sample of the original speech. Both the and processed audio samples are provided to the computer. The computer then compares the original audio signal with the processed audio signal to evaluate the quality of the processed audio signal. However, if the original audio signal is not available, traditional intrusive objective methods cannot be used to evaluate the quality of the processed audio signal. The estimated original audio signal can be used instead of the lost original audio signal, but the greater the distortion of the processed audio signal, the lower the quality of the estimated original audio signal.<patcit num="1"><text>U.S. Patent Application No. 10/186840</text></patcit>
<p> To provide effective remedies for one or more of the above issues.</p>
<p> In one embodiment of the present invention, an apparatus for real-time objective speech analysis is provided. This device includes a voice quality analyzer. The voice quality analyzer receives at least one first signal and, based on the received first signal, at least one second that indicates at least one non-intrusive evaluation of voice quality. Provides a signal for.</p><p> Other embodiments of the invention provide methods for real-time objective speech analysis. This method is based on the step of receiving at least one first signal indicating at least one processed audio signal and the audio of the at least one processed audio signal based on the received at least one first signal. It includes a step of determining the quality non-invasively and a step of providing at least one second signal indicating the determined audio quality of the at least one processed audio signal. The present invention should be understood by reading the following description in conjunction with the accompanying drawings. In the drawings, similar reference numbers indicate similar elements.</p><p> Although various modifications and alternatives are feasible with respect to the present invention, specific embodiments of the present invention will be illustrated and described in detail herein by way of example. However, the description of the present specification for a particular embodiment is not intended to limit the invention to the disclosed particular embodiments, and conversely, the invention is made by the appended claims. It should be understood that it is intended to include all modifications, equivalents, and alternative embodiments within the defined purpose and scope of the invention.</p>
An exemplary embodiment of the present invention is described below. For the sake of clarity, this detailed description does not necessarily describe all features of the actual implementation. Not surprisingly, in the development of actual embodiments, many implementation-specific decisions are made to achieve each implementation-specific goal, such as compliance with system constraints and business constraints. Will be understood to be necessary. Moreover, it will be appreciated that such development work, even if it is a complex and time-consuming work, is only routine work for those skilled in the art who will enjoy the benefits of the present disclosure.
FIG. 1 shows an exemplary embodiment of the wireless communication network 100. Although the present invention describes an exemplary embodiment of the wireless communication network 100, those skilled in the art will understand that the present invention is not limited to the wireless communication network as shown in FIG. I want to. In an alternative embodiment, the invention can be implemented in other networks, such as Internet-based networks using POTS (General Telephone System), VoIP (Voice over Internet Protocol), and the like. Further, since the structure and operation of the wireless communication network 100 are generally well known to those skilled in the art, in order to make the explanation easy to understand, this specification is useful for understanding the present invention of the wireless communication network 100. Only aspects related to structure and operation will be described.
The wireless communication network 100 includes a first mobile unit 105 capable of transmitting a signal to and receiving a signal from the base station 110 via the wireless communication channel 115. The base station 110 is communicatively coupled to the network 120. In various alternative embodiments, the base station 110 can be communicatively coupled to the network 120 by any desired method, such as a wireless communication link, a wired communication link, or the like. The network 120 can include devices such as routers, switches, filters, signal processors, etc. that can be interconnected in any desired way. The network 120 is also communicated with at least one base station 125. The base station can transmit and / or receive signals to and / or receive signals to mobile unit 130 via wireless communication channel 135.
Upon operation, the original audio signal 140 is provided to the mobile unit 105. For example, the first user can speak to a microphone (not shown) built into the mobile unit 105. The mobile unit 105 processes the original audio signal 140 to form the processed audio signal 145, and the audio signal is transmitted to the base station 110. The processed audio signal 145 can be transmitted from the base station 110 to the mobile unit 130 via the network 120, the base station 125, the wireless communication channel 135, other intermediate devices and / or channels, and the like. The mobile unit 130 can then provide an acoustic signal to the second user based on the processed audio signal 145.
The processed audio signal 145 may be modified by mobile units 105, 130, base stations 110, 125, network 120, wireless communication channels 115, 135, other intermediates and / or channels, and the like. As a result, the processed audio signal 145 may differ from the original audio signal 140. In general, modifications of the original audio signal 140 tend to reduce the audio quality of the processed audio signal 145. For example, the processed audio signal 145 may contain spike noise 150 that is not present in the original audio signal 140. However, when the degradation in voice quality of the processed voice signal 145 is relatively small, it may not be easily perceptible to the human ear and therefore may not be a concern.
Therefore, by providing the voice quality analyzer 155, the voice quality of the processed voice signal 145 is evaluated using a non-intrusive voice quality evaluation technique. According to common use in the art, the term "non-intrusive" is understood herein to mean a voice quality evaluation technique that can be performed without the use of the original voice signal. To. In the embodiment shown in FIG. 1, the voice quality analyzer 155 receives a signal indicating the processed voice signal 145 from the base station 125 and evaluates the voice quality of the processed voice signal 145 based on the received signal. can do. However, at least in part, the voice quality analyzer 155 uses non-intrusive voice quality evaluation technology, so that the voice quality analyzer 155 sends a signal indicating the processed voice signal 145 to the wireless communication network 100. It can be received from any part. For example, in one embodiment, the voice quality analyzer 155 may receive a signal indicating the processed voice signal 145 from a portion of the network 120.
In the exemplary embodiment shown in FIG. 1, the voice quality analyzer 155 is outside the path of the processed voice signal 145. However, the present invention is not limited to the voice quality analyzer 155 outside the path of the processed voice signal 145. In an alternative embodiment, the voice quality analyzer 155 can be substantially placed within the path of the processed voice signal 145. For example, the voice quality analyzer 155 can be installed in series between the base station 125 and the mobile unit 130. In another alternative embodiment, the voice quality analyzer 155 can be placed in parallel on any part of the wireless communication network 100. In addition, assessing the voice quality of the processed voice signal 145 at a selected location within the wireless communication network 100 by deploying two or more voice quality analyzers 155 using non-intrusive technology. You can also.
In one embodiment, the voice quality analyzer 155 can provide feedback to base station 125 based on the non-invasively evaluated voice quality of the processed voice signal 145. For example, the audio quality analyzer 155 determines that the audio quality of the processed audio signal 145 has deteriorated due to the presence of the noise spike 150, and filters the amplitude of the noise spike 150 in the processed audio signal 145. A signal can be provided to base station 125 indicating that it is desirable to apply and suppress. However, those skilled in the art are not limited to the application of filtering, and in alternative embodiments, any desirable device is optional in response to the feedback provided by the audio quality analyzer 155. It should be understood that signal processing techniques can be used to reduce the effects of unwanted parts of the processed audio signal 145.
FIG. 2 shows an exemplary embodiment of the voice quality analyzer 155. The voice quality analyzer 155 receives one or more processed voice signals via one or more input lines 200 (1 to n), such as the processed voice signal 145 shown in FIG. Can be done. In one embodiment, the input lines 200 (1 to n) are T1 lines, and each of these T1 lines is a gateway device, such as an OC3-T1 converter coupled to a Cisco Media Gateway MXG. It can be obtained from a converter coupled to (not shown). Generally, one T1 line runs about 24 calling lines. However, for those skilled in the art, the input line 200 (1 to n) is not limited to the T1 line, and in an alternative embodiment, it is any desired type of line through any desired number of call channels. Please understand that it is good.
Input lines 200 (1 to n) provide the processed audio signal to an interface 205, such as a PCMCIA interface. Interface 205 can provide one or more signals indicating the processed audio signal to one or more digital signal processors (DSPs) 210 (1-m). In an exemplary embodiment, the digital signal processor 210 is formed on a separate chip located on the substrate 215. However, the present invention is not limited to one or more digital signal processors 210 (1 to m) arranged on a single substrate 215. In alternative embodiments, substrate 215 may not be provided. In another alternative embodiment, the digital signal processor 210 (1 to m) can be placed on a plurality of boards 215.
The digital signal processor 210 (1 to m) implements a non-intrusive method for evaluating the voice quality of the processed voice signal 145. In one embodiment, the digital signal processor 210 (1 to m) implements the ANIQUE (Auditory Non-Intrusive Quality Optimization) algorithm. This auditory-articulatory analysis technique evaluates the audio quality of an audio signal by comparing the power in the articulatory frequency range with the power in the non-articulatory frequency range. For example, the ANIQUE algorithm evaluates the audio quality of a processed audio signal by comparing the power in the tuning frequency range of about 2 to 12.5 Hz with the power in the non-tuning frequency range above about 12.5 Hz. An exemplary embodiment of the non-intrusive ANIQUE algorithm is, for example, "Auditory-Articulatory Analysis for Speech Quality" by Kim. It is described in US Patent Application No. 10/186840, filed July 1, 2002, entitled "Assessment," which is incorporated herein by reference in its entirety.
The complexity of the ANIQUE algorithm can be assessed by adopting a WMOPS (Weighted Million Operations Per Second) calculation routine from the selectable mode vocoder to the C source code used to implement the ANIQUE algorithm. According to the evaluation results, the ANIQUE algorithm has a complexity of about 217 WMOPS. However, as one of ordinary skill in the art should understand, this evaluation depends on the individual implementation of the algorithm. For example, assessing the complexity of the ANIQUE algorithm reduces the number of points at the Fast Fourier Transform point from 4096 to 2048, uses 4-element simultaneous multiplication and accumulation operations during the filtering process, and optimizes the source code. By doing so, it can be reduced to 122 WMOPS or less.
In one embodiment, the voice quality analyzer 155 includes 16 digital signal processors 210 (1 to m). If the non-intrusive voice quality evaluation technology implemented in each digital signal processor 210 (1 to m) used a computational speed of approximately 80 MIPS (Million Instructions per Second), this number is discussed above for the ANIQUE algorithm. Although somewhat lower than the 122 WMOPS, this implementation of the Voice Quality Analyzer 155 can handle nearly 64 call channels simultaneously. However, one of ordinary skill in the art should understand that this assessment of the number of call channels that can be processed simultaneously by the Voice Quality Analyzer 155 is exemplary and not intended to limit the invention.
The digital signal processor 210 (1 to m) provides one or more signals indicating the evaluated audio quality of the processed audio signal to interface 217, such as the PCMCIA interface. In one embodiment, the interface 217 may provide the computer 220 with one or more signals indicating the evaluated voice quality for the processed voice signal. For example, interface 217 can provide a signal to laptop computer 220. The computer 220 can then display information indicating the evaluated voice quality of the processed voice signals on one or more communication channels analyzed by the voice quality analyzer 155. For example, computer 220 can display this information using the graphical user interface 225.
FIG. 3A is an exemplary embodiment of the graphical user interface 225. In the illustrated embodiment, the graphical user interface 225 displays information indicating the communication channel in column 300 (eg, channel number, etc.) and information indicating the evaluated voice quality in column 305 (eg, voice 1 to 5). The quality rank, etc.) is displayed in column 310, information indicating the time and / or duration of the processed audio signal (eg, time stamp, etc.), and the user activation button 315 is displayed in column 320. The user activation button 315 can be used to allow the user to see a portion of the waveform of the processed audio signal, eg, the exemplary waveform 330 shown in FIG. 3B. However, one of ordinary skill in the art will appreciate that the present invention is not limited to the information shown in FIG. 3A, and in alternative embodiments, any desired information can be displayed in the graphical user interface 225. ..
Returning to FIG. 2, as mentioned above, the voice quality analyzer 155 can provide feedback based on a non-intrusive assessment of voice quality. Thus, in one embodiment, the computer 220 is communicatively connected to the wireless communication network 100 and can provide a signal indicating a modification applicable to the processed audio signal. This signal can be provided to one or more devices in the wireless communication network 100, and the signal can be used by each device to modify the processed voice signal. Alternatively, the computer 220 can modify the processed audio signal. For example, the computer 220 may allow the user to select and / or utilize various audio editing tools for the processed audio signal. Audio editing tools can include, for example, time and / or frequency filtering, compression, interpolation, fading, normalization, enveloping, and the like.
The audio quality analyzer 155 is in operation because the audio quality analyzer 155 described above can evaluate the audio quality of one or more processed audio signals in a non-intrusive manner, i.e. without using the original audio signal. It can be used to evaluate the audio quality of other systems that cannot use the network or the original audio signal. In addition, the voice quality analyzer 155 does not need to be driven using a predetermined test signal and can objectively evaluate the voice quality, so that the voice quality analyzer 155 is more of a network than traditional subjective methods. The time and cost for assessing voice quality can be reduced.
The particular embodiments disclosed above are provided for illustration purposes only. Accordingly, the present invention can be modified and implemented in different but equal manners, which will be apparent to those skilled in the art who have the benefit of the teachings herein. Moreover, the detailed structure or design presented herein is not intended to be a limitation other than those set forth in the appended claims. Therefore, it is clear that the particular embodiments disclosed above can be modified or modified, and all such modifications are considered to be within the scope and gist of the present invention. Therefore, the protections sought by the present specification are set forth in the appended claims.
<figref num="1">It is a figure which shows the communication network including the voice quality analyzer by one Embodiment of this invention.</figref><figref num="2">It is a figure which shows an exemplary embodiment of the voice quality analyzer, for example, the voice quality analyzer shown in FIG. 1 according to one embodiment of the present invention.</figref><figref num="3A">FIG. 5 illustrates an exemplary embodiment of a graphical user interface that can be used to display information provided by the voice quality analyzer shown in FIG. 2, according to an embodiment of the present invention.</figref><figref num="3B">FIG. 6 illustrates an exemplary portion of the waveform of a processed audio signal that can be viewed using the graphical user interface shown in FIG. 3A, according to an embodiment of the invention.</figref>
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9729602B2 | Cited by | United States of America | Applicant |
4 priority claims, no other members on record
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 10818435 | United States of America | – | |
| 81843504 | United States of America | A | |
| 2004818435 | – | – | – |
| US20040818435 | – | – | – |
Numbers
- Publication
- 2005292841
- Publication, DOCDB
- 2005292841
- Publication, EPODOC
- JP2005292841
- Application
- 108161
- Application, DOCDB
- 2005108161
- Application, EPODOC
- JP20050108161
Titles2
- Japanese
- リアルタイムの客観的音声アナライザ
- English
- Real-time objective voice analyzer
Classification
- CPC, 3
- G10L25/69
- F02B61/045
- B63H21/14
- IPC, 3
- G10L11 00
- G10L19 00
- H04M3 00