Method and apparatus for adjusting volume of user terminal, and terminal
16 claims: 5 independent, 11 dependent
- 1ユーザ端末によって実行される、前記 ユーザ端末の音量を調整する方法であって、 前記ユーザ端末の周囲の音声信号を収集するステップと、 前記音声信号の組成情報を取得するために、前記収集した音声信号の分析を実行するステップであって、前記組成情報は、前記音声信号に含まれる音声タイプ、および様々なタイプの音声の比率を含み、前記音声タイプは、無音、人声、およびノイズを含む、分析を実行するステップと、 前記音声信号の前記組成情報に従って、前記ユーザ端末の現在のシーンモードを決定するステップと、 前記決定したシーンモードに従って、前記ユーザ端末の前記音量を調整するステップとを含む、方法。
- 2前記ユーザ端末の周囲の音声信号を収集する前記ステップが、具体的に、 呼び出し信号の到達が検出されたときに、前記ユーザ端末が置かれている現在の環境の音声信号を収集すること、または、 前記ユーザ端末が置かれている現在の環境の音声信号を定期的に収集することを含む、請求項1に記載の方法。
- 3前記決定したシーンモードに従って、前記ユーザ端末の前記音量を調整する、前記ステップが、 前記決定したシーンモード、ならびにシーンモードと音量調整係数との間の予め記憶した対応関係に従って、音量調整係数を決定することと、前記音量調整係数に従って、前記ユーザ端末の前記音量を調整することとを含む、請求項2に記載の方法。
- 4前記ユーザ端末の前記音量が、着信音量および受話音量を含み、前記音量調整係数が、着信音量調整係数、および受話音量調整係数を含み、 前記音量調整係数に従って、前記ユーザ端末の前記音量を調整する前記ステップが、 前記着信音量調整係数に従って、前記ユーザ端末の前記着信音量を調整することと、前記受話音量調整係数に従って、前記ユーザ端末の前記受話音量を調整することとを含む、請求項3に記載の方法。
- 5前記呼び出し信号の接続が検出された後に、 前記ユーザ端末のマイク音量をリアルタイムで取得するステップと、 前記取得したマイク音量が、予め記憶した基準音量よりも大きいときに、前記ユーザ端末の前記受話音量を大きくするステップと、 前記取得したマイク音量が、前記予め記憶した基準音量よりも小さいときに、前記ユーザ端末の前記受話音量を小さくするステップとをさらに含む、請求項4に記載の方法。
- 6前記呼び出し信号の接続が検出された後に、 前記ユーザ端末の周囲の前記音声信号を定期的に収集するステップ、および前記音声信号の前記組成情報を取得するために、前記収集した音声信号の分析を実行するステップと、 前記音声信号の前記組成情報に従って、前記音声信号が人声を含むかどうか、および1人のみの声を含むかどうかを判定するステップであって、前記音声信号が、人声を含み、かつ1人のみの声を含む場合は、前記音声信号の音量を計算し、前記音声信号の前記計算した音量と前記予め記憶した基準音量との平均値を取り、前記平均値を新たな基準音量として記憶する、判定するステップとをさらに含む、請求項5に記載の方法。
- 7前記音声信号の組成情報を取得するために、前記収集した音声信号の分析を実行する、前記ステップが、具体的に、 前記収集した音声信号を複数の音声データに分割することと、 各音声データの音声周波数を計算すること、ならびに前記計算した音声周波数に従って、前記各音声データを無音、人声、およびノイズに分類することと、 全ての音声データの無音、人声、およびノイズの比率の統計を収集することであって、 人声と識別された音声データについては、人声の音声データのメル周波数ケプストラム係数を計算し、前記人声に含まれる人数についての情報を判定するために、メル周波数ケプストラム係数が同一の音声データを1人の声として統計を収集することとを含む、請求項6に記載の方法。
- 8前記計算した音声周波数に従って、各音声データを無音、人声、およびノイズに分類する、前記ステップが、具体的に、 前記各音声データの前記音声周波数が、20Hz~20000Hzの範囲内にあるかどうかを決定することであって、 前記音声データの前記音声周波数が、20Hz~20000Hzの前記範囲内にあると決定されたときは、前記音声データの基本周波数を計算し、 前記基本周波数が、85Hz~255Hzの範囲内にあると決定されたときは、前記音声データは人声であるとみなし、 前記基本周波数が、85Hz~255Hzの前記範囲外にあると決定されたときは、前記音声データはノイズであるとみなし、 前記音声データの前記音声周波数が、20Hz~20000Hzの前記範囲外にあると決定されたときは、前記音声データは無音であるとみなす、決定することを含む、請求項7に記載の方法。
- 9コンピュータ読み取り可能なプログラムを記憶するように構成されたメモリと、 前記メモリから前記コンピュータ読み取り可能なプログラムを読み取るように構成されたプロセッサであって、 ユーザ端末の周囲の音声信号を収集することと、 前記音声信号の組成情報を得るために、前記収集した音声信号の分析を実行することであって、前記組成情報は、前記音声信号に含まれる音声タイプ、および様々なタイプの音声の比率を含み、前記音声タイプは、無音、人声、およびノイズを含む、分析を実行することと、 前記音声信号の前記組成情報に従って、前記ユーザ端末の現在のシーンモードを決定することと、 前記決定したシーンモードに従って、前記ユーザ端末の音量を調整することとを実行する、 プロセッサとを備える、端末。
- 10前記プロセッサが、呼び出し信号の到達が検出されたときに、前記ユーザ端末が置かれている現在の環境の音声信号を収集するように構成され、または、 前記プロセッサが、前記ユーザ端末が置かれている現在の環境の音声信号を定期的に収集するように構成される、請求項9に記載の端末。
- 11前記プロセッサが、前記決定したシーンモード、およびシーンモードと音量調整係数との間の予め記憶した対応関係に従って音量調整係数を決定するように、かつ前記音量調整係数に従って、前記ユーザ端末の前記音量を調整するように構成される、請求項10に記載の端末。
- 12前記ユーザ端末の前記音量が、着信音量および受話音量を含み、前記音量調整係数が、着信音量調整係数、および受話音量調整係数を含み、 前記プロセッサが、前記着信音量調整係数に従って前記ユーザ端末の前記着信音量を調整し、前記受話音量調整係数に従って前記ユーザ端末の前記受話音量を調整するように構成される、請求項11に記載の端末。
- 13前記呼び出し信号の接続が検出された後に、前記プロセッサが、前記ユーザ端末のマイク音量をリアルタイムで取得し、前記取得したマイク音量が、予め記憶した基準音量よりも大きいときは、前記ユーザ端末の前記受話音量を大きくし、前記取得したマイク音量が、前記予め記憶した基準音量よりも小さいときは、前記ユーザ端末の前記受話音量を小さくするようにさらに構成される、請求項12に記載の端末。
- 14前記呼び出し信号の接続が検出された後に、前記プロセッサが、 前記ユーザ端末の周囲の前記音声信号を定期的に収集して、前記音声信号の前記組成情報を得るために、前記収集した音声信号の分析を実行し、 前記音声信号の前記組成情報に従って、前記音声信号が人声を含み、かつ1人のみの声を含むかどうかを判定し、前記音声信号が人声を含み、かつ1人のみの声を含む場合は、前記音声信号の音量を計算し、前記音声信号の前記計算した音量と、前記予め記憶した基準音量との平均値を取り、前記平均値を新たな基準音量として記憶するようにさらに構成される、請求項13に記載の端末。
- 15前記プロセッサが、 前記収集した音声信号を複数の音声データに分割し、 各音声データの音声周波数を計算し、前記計算した音声周波数に従って、各音声データを無音、人声、およびノイズに分類し、 全ての音声データの無音、人声、およびノイズの比率の統計を収集し、 人声として識別された音声データについては、人声の音声データのメル周波数ケプストラム係数を計算し、前記人声に含まれる人数についての情報を判定するために、メル周波数ケプストラム係数が同一の音声データを1人の声として統計を収集するように構成される、請求項14に記載の端末。
- 16コンピュータ読み取り可能なプログラムであって、プロセッサによって読み取られると、前記プロセッサが、 ユーザ端末の周囲の音声信号を収集し、 前記音声信号の組成情報を得るために、前記収集した音声信号の分析を実行し、前記組成情報は、前記音声信号に含まれる音声タイプ、および様々なタイプの音声の比率を含み、前記音声タイプは、無音、人声、およびノイズを含み、 前記音声信号の前記組成情報に従って、前記ユーザ端末の現在のシーンモードを決定し、 前記決定したシーンモードに従って、前記ユーザ端末の音量を調整する、 コンピュータ読み取り可能なプログラムを含む、非一時的記憶媒体。
Independent claims16
76 paragraphs, as filed
The present invention relates to the field of communication technology, and more particularly to user terminals, as well as methods and devices for adjusting the volume of the terminals.
With the continuous development of communication technology, mobile user terminals such as mobile phones and tablet computers have become indispensable belongings for people's lives and work. People can make and receive calls anytime, anywhere. When receiving a call, it is usually necessary to change the ring volume setting. Mobile phone users typically use relatively low volume ringtones in quiet environments such as offices to avoid making relatively loud sounds that could affect the normal work of other office staff. I want to. However, in noisy public places such as shopping malls and train stations, a relatively loud ringtone is required to ensure that incoming calls can be received without delay.
Currently, the ring volume of many mobile phones and the call volume of the handset are usually manually adjusted by the user, or many mobile phones are simply self-sensing, such as in certain scenes such as indoor mode or outdoor mode. Incorporates the scene mode of.
When the self-sensing scene mode is used to adjust the playback volume of the mobile phone ringtone, the decibels of the ambient sound are generally extracted and determined by using a voice detection module when the call comes in. Then, the ringtone volume and the earpiece volume are adjusted according to the pre-stored correspondence between the voice decibel amount and the ringtone volume of the incoming call. In the volume adjustment method described above, the final volume of the ringtone and the final volume of the handset are determined only according to the decibel of the ambient sound, but the decibel of the ambient sound accurately determines the situation in which the user is placed. Also, if the extracted environmental sound contains a relatively loud human voice such as a conference, debate, or presentation with relatively lively discussion, the extracted environmental sound will be It may be mistakenly regarded as noise. In such situations, the volume should be reduced, but the volume determined according to the decibels of the environmental sound is relatively loud. As a result, the volume adjustment of the mobile phone does not match the actual scene, and its accuracy is not high.
<p num="0005"> In the embodiment of the present invention, the volume of the user terminal and the volume of the terminal are adjusted so that the volume of the mobile phone adapts to the situation in which the user is placed and more accurately matches the volume of the user to enhance the usability of the user. A method and an apparatus for the present invention are provided.</p><p num="0006"> According to the first aspect, a method of adjusting the volume of the user terminal is provided, the method of which is: Steps to collect audio signals around the user terminal, In order to obtain the composition information of the voice signal, the step of performing the analysis of the collected voice signal, the composition information includes the voice type contained in the voice signal and the ratio of various types of voice, and the voice type. Steps to perform the analysis, including silence, human voice, and noise, Steps to determine the current scene mode of the user terminal according to the composition information of the audio signal, It includes a step of adjusting the volume of the user terminal according to the determined scene mode.</p><p num="0007"> In the first possible embodiment, according to the first aspect, the step of collecting audio signals around the user terminal is Collecting the voice signal of the current environment in which the user terminal is located, or collecting the voice signal when the arrival of the call signal is detected. Specifically, it includes periodically collecting audio signals of the current environment in which the user terminal is located.</p><p num="0008"> The step of adjusting the volume of the user terminal according to the determined scene mode in the second possible implementation method, which is related to the first possible embodiment of the first aspect or the first possible embodiment. This includes determining the volume adjustment coefficient according to the determined scene mode and the pre-stored correspondence between the scene mode and the volume adjustment coefficient, and adjusting the volume of the user terminal according to the volume adjustment coefficient.</p><p num="0009"> In a third possible embodiment relating to the second possible embodiment of the first aspect, the volume of the user terminal includes the ringtone volume and the earpiece volume, and the volume adjustment factors are the ringtone volume adjustment factor and the ringtone volume adjustment factor. Includes earpiece volume adjustment factor The step of adjusting the volume of the user terminal according to the volume adjustment coefficient is This includes adjusting the ringtone volume of the user terminal according to the ringtone volume adjustment coefficient, and adjusting the earpiece volume of the user terminal according to the earpiece volume adjustment coefficient.</p><p num="0010"> A fourth aspect, which relates to a first aspect, a first possible implementation of the first aspect, a second possible implementation of the first aspect, or a third possible implementation of the first aspect. In the possible implementation method of After the call signal connection is detected, this method Steps to acquire the microphone volume of the user terminal in real time, When the acquired microphone volume is louder than the pre-stored reference volume, the step of increasing the earpiece volume of the user terminal and When the acquired microphone volume is lower than the reference volume stored in advance, the step of reducing the earpiece volume of the user terminal is further included.</p><p num="0011"> In a fifth possible practice, which relates to a fourth possible practice of the first aspect, after the call signal connection is detected, this method A step of periodically collecting audio signals around the user terminal, and a step of performing analysis of the collected audio signals in order to obtain information on the composition of the audio signals. Further including a step of determining whether the audio signal contains a human voice and whether it contains only one voice according to the composition information of the audio signal, the audio signal includes a human voice and the audio signal contains a human voice. , And only one voice When including only one voice, calculate the volume of the audio signal, take the average value of the calculated volume of the audio signal and the pre-stored reference volume, and set a new average value. Store as the reference volume.</p><p num="0012"> Of the first aspect, the first possible implementation of the first aspect, the second possible implementation of the first aspect, or the third possible implementation of the first aspect, the first aspect. Of the audio signal collected in order to obtain composition information of the audio signal in the sixth possible implementation method, which is related to the fourth possible implementation method or the fifth possible implementation method of the first aspect. The steps to perform the analysis are Dividing the collected voice signal into multiple voice data and To calculate the voice frequency of each voice data, and to classify each voice data into silence, human voice, and noise according to the calculated voice frequency. Specifically, it includes collecting statistics on the ratio of silence, human voice, and noise of all voice data. For voice data identified as human voice, the mel frequency cepstrum coefficient of the human voice voice data is calculated, and in order to determine information about the number of people included in the human voice, voice data having the same mel frequency cepstrum coefficient is used. Collect statistics as one voice.</p><p num="0013"> In connection with the sixth possible practice of the first aspect, in the seventh possible practice, the steps of classifying each voice data into silence, human voice, and noise according to the calculated voice frequency are: Specifically, it includes determining whether or not the voice frequency of each voice data is in the range of 20 Hz to 20000 Hz. When it is determined that the voice frequency of the voice data is within the range of 20 Hz to 20000 Hz, the fundamental frequency of the voice data is calculated. When it is determined that the fundamental frequency is within the range of 85Hz to 255Hz, the voice data is considered to be human voice, and when it is determined that the fundamental frequency is outside the range of 85Hz to 255Hz, the voice data is considered. Is considered noise When it is determined that the voice frequency of the voice data is outside the range of 20 Hz to 20000 Hz, the voice data is considered to be silent.</p><p num="0014"> According to the second aspect, a device for adjusting the volume of the user terminal is provided, and the device is A collection unit configured to collect audio signals around the user terminal, and An analysis unit configured to perform analysis of the collected voice signal in order to obtain the composition information of the voice signal, the composition information of the voice type contained in the voice signal and various types of voice. Includes ratios, voice types include silence, human voice, and noise, with analysis units, A scene mode determination unit configured to determine the current scene mode of the user terminal according to the composition information of the audio signal. It includes a volume adjusting unit configured to adjust the volume of the user terminal according to the determined scene mode.</p><p num="0015"> In the first possible implementation, which relates to the second aspect, the collection unit is to collect the voice signal of the current environment in which the user terminal is located when the arrival of the call signal is detected. , Or specifically configured to periodically collect audio signals in the current environment in which the user terminal is located.</p><p num="0016"> In the second possible embodiment, which relates to the second aspect, or the first possible embodiment of the second aspect, the volume control unit is of the determined scene mode, as well as the scene mode and the volume control factor. It is specifically configured to determine the volume adjustment coefficient according to the correspondence relationship stored in advance between them, and to adjust the volume of the user terminal according to the volume adjustment coefficient.</p><p num="0017"> In the third possible implementation, which relates to the second possible implementation of the second aspect, the volume of the user terminal includes the ring volume and the earpiece volume, and the volume adjustment factor of the volume control unit is the ring volume. Specifically including the adjustment coefficient and the earpiece volume adjustment coefficient, The volume adjustment unit is specifically configured to adjust the ringtone volume of the user terminal according to the ringtone volume adjustment coefficient, and to adjust the earpiece volume of the user terminal according to the earpiece volume adjustment coefficient.</p><p num="0018"> A fourth aspect relating to a second aspect, a first possible implementation of the second aspect, a second possible implementation of the second aspect, or a third possible implementation of the second aspect. In the possible implementation of, this device An acquisition unit configured to acquire the microphone volume of the user terminal in real time after it is detected that a call signal has been connected. Further equipped with a comparison unit configured to compare the magnitude of the acquired microphone volume with the pre-stored reference volume. The volume adjustment unit increases the earpiece volume of the user terminal when the acquired microphone volume is louder than the pre-stored reference volume, and increases the earpiece volume of the user terminal when the acquired microphone volume is lower than the pre-stored reference volume. It is further configured to reduce the earpiece volume of.</p><p num="0019"> In a fifth possible practice, which relates to a fourth possible practice of the second aspect, the collection unit periodically sends a voice signal around the user terminal after the call signal connection is detected. Collect and In order to obtain the composition information of the voice signal, the analysis unit is to perform the analysis of the voice signal periodically collected by the collection module after the connection of the call signal is detected, and according to the composition information of the voice signal. , The audio signal is further configured to determine whether the voice signal contains a human voice, and whether the voice signal contains only one voice, the audio signal contains a human voice, the audio signal contains a human voice, and the audio signal is a person. If the voice is included and only one voice is included, the volume of the audio signal is calculated, and the average value of the calculated volume of the audio signal and the pre-stored reference volume is taken and the average value is taken. Is memorized as a new reference volume.</p><p num="0020"> A second aspect, a first possible implementation method of the second aspect, a second possible implementation method of the second aspect, a third possible implementation method of the second aspect, a second aspect of the second aspect. In the sixth possible implementation method, which is related to the fourth possible implementation method, or the fifth possible implementation method of the second aspect, the analysis unit. A first processing unit configured to divide the collected voice signal into multiple voice data, and A second processing unit configured to calculate the voice frequency of each voice data and classify each voice data into silence, human voice, and noise according to the calculated voice frequency. With a third processing unit configured to collect silence, human voice, and noise ratio statistics for all voice data, For voice data identified as human voice, the mel frequency cepstrum coefficient is calculated, and in order to determine information about the number of people included in the human voice, voice data with the same mel frequency cepstrum coefficient is statistically regarded as one voice. Specifically includes a fourth processing unit configured to collect data.</p><p num="0021"> In the seventh possible implementation, which relates to the sixth possible implementation of the second aspect, the second processing unit determines whether the audio frequency of each audio data is in the range of 20Hz to 20000Hz. When it is determined that the audio frequency of the audio data is within the range of 20Hz to 20000Hz, the basic frequency of the audio data is calculated and it is determined that the basic frequency is within the range of 85Hz to 255Hz. When it is determined that the voice data is a human voice and the basic frequency is outside the range of 85Hz to 255Hz, the voice data is regarded as noise and the voice frequency of the voice data is 20Hz. When it is determined that the voice data is out of the range of ~ 20000 Hz, the voice data is specifically configured to be regarded as silent.</p><p num="0022"> According to a third aspect, there is provided a terminal comprising a speaker and a handset, further comprising the above-mentioned volume control device according to an embodiment of the present invention.</p><p num="0023"> The method and device for adjusting the volume of a user terminal and the terminal provided in the embodiment of the present invention have the following beneficial effects.</p><p num="0024"> Once the current scene mode of the user terminal is determined, a reference for the composition information of the audio signal is added so that the corresponding current scene is closer to the actual scene and more accurately matched to the scene in which the user is located. This makes it possible to significantly reduce the occurrence of a situation caused by an error in determining the scene, that is, the adjustment of the playback volume does not match the scene, and improve the usability of the user.</p>
<figref num="1">FIG. 1 is a schematic flowchart 1 of a volume adjusting method according to an embodiment of the present invention.</figref><figref num="2">FIG. 2 is a schematic flowchart 2 of a volume adjusting method according to an embodiment of the present invention.</figref><figref num="3">FIG. 3 is a schematic flowchart 3 of a volume adjusting method according to an embodiment of the present invention.</figref><figref num="4">FIG. 4 is a schematic flowchart 4 of a volume adjusting method according to an embodiment of the present invention.</figref><figref num="5">FIG. 5 is a schematic flowchart 5 of a volume adjusting method according to an embodiment of the present invention.</figref><figref num="6">It is a schematic structural drawing of the volume control apparatus by embodiment of this invention.</figref><figref num="7">It is a schematic structural drawing of the user terminal by embodiment of this invention.</figref>
To solve the existing problem that the adjusted volume does not match the actual scene due to the inaccurate accuracy of identifying the scene when the user terminal automatically adjusts the volume according to the surrounding environment. , A method and device for adjusting the volume of a user terminal, and a terminal are provided in embodiments of the present invention, which can perform an accurate analysis of environmental sounds, resulting in consistent current scenes. By being closer to the actual scene and further obtaining an appropriate volume by the adjustment, the occurrence of a situation in which the volume adjustment does not match the scene, which is caused by an erroneous judgment of the scene, is greatly reduced. Hereinafter, the technical solution of the embodiment of the present invention will be described more clearly and completely with reference to the accompanying drawings of the embodiment of the present invention. It is clear that the embodiments described are only some, but not all, of the embodiments of the present invention. Other embodiments proposed on the basis of embodiments of the present invention without creative effort by those skilled in the art fall within the scope of protection of the present invention.
The method for adjusting the volume of a user terminal provided in the embodiment of the present invention can be mainly used for communication of a communication terminal such as a mobile phone or a transceiver. For example, when the ringing signal arrives, the mobile phone may not ring immediately, but after determining the volume using the volume adjustment method provided in the embodiment of the present invention, appropriate ring volume control is performed. In addition, it is possible to appropriately control the earpiece volume after the telephone is connected. The volume adjusting method provided in the embodiment of the present invention can also be used, for example, in a mobile television terminal attached to a bus or subway to broadcast a program. For example, a mobile television terminal for a bus can automatically adjust the volume of a program according to the number of passengers and the volume in the bus by using the volume adjusting method provided in the embodiment of the present invention.
Referring to FIG. 1, embodiments of the present invention provide a method of adjusting the volume of a user terminal, which method comprises the following steps.
S101. Collects audio signals around the user terminal.
In a specific embodiment, the time for collecting the audio signal can be controlled according to the processing capacity of the device that adjusts the volume. For example, when the processing power of the device is relatively high, the audio signal of the current environment in which the user terminal is located may be collected only when the arrival of the ringing signal is detected, and then the subsequent steps are performed. However, when the processing power of the device is relatively low, the audio signal of the current environment in which the user terminal is located may be collected periodically, and then the subsequent procedure is performed. When the ringing signal arrives, the ringtone is played directly using the adjusted volume.
In a specific implementation, the audio signal of the current environment can be collected using components such as the microphone of the terminal, or another audio sensor can be used to collect the audio signal of the current environment. It may be configured in, which is not limited herein.
S102. In order to obtain the composition information of the voice signal, the analysis of the collected voice signal is performed, and the composition information includes the voice type contained in the voice signal and the ratio of various types of voice, and the voice type is silence. Includes (blank sound), human voice, and noise.
In a specific embodiment, the voice signal can be divided into silence, human voice, non-human voice (noise) and the like. Silence is a voice signal that cannot be recognized by the human ear. In general, audio signals with audio frequencies outside the range of 20 Hz to 20000 Hz can be considered silent. Noise refers to audio signals other than human voice that can be recognized by the human ear. In general, audio signals with audio frequencies in the range of 20 Hz to 85 Hz and in the range of 255 Hz to 20000 Hz can be considered noise. By performing an analysis of the collected audio signal, the ratio of silence, the ratio of noise, and the ratio of human voice contained in the audio signal are calculated, and as a result, the identification of the subsequent scene mode is performed.
S103. Determine the current scene mode of the user terminal according to the composition information of the audio signal.
The scene mode of the user terminal is used to indicate the environmental conditions in which the user terminal is placed, such as a quiet library, a conference room where discussions are held, a quiet bedroom, or a noisy road. In a specific implementation, a corresponding scene mode correspondence is established for the magnitude relationship between the three types of audio signals, that is, audio signals with different ratios correspond to different scene modes, and each The scene mode corresponds to the corresponding volume. Also, once the correspondence is established, a measure of human voice volume can be added, resulting in volume adjustments caused by environmental misjudgment due to the corresponding scene modes being closer to the actual environment. Significantly reduces the occurrence of situations that do not match the environment.
S104. Adjust the volume of the user terminal according to the determined scene mode.
In a specific implementation, the volume adjustment coefficient can be determined according to the determined scene mode and the pre-stored correspondence between the scene mode and the volume adjustment coefficient, and the volume of the user terminal is adjusted according to the volume adjustment coefficient. Will be done.
In a specific implementation, when applied to a communication terminal such as a mobile phone, the volume of the user terminal may specifically include the volume of a ring tone (also called a ring tone) and the volume of an earpiece. , The volume adjustment coefficient specifically includes two types, that is, the ring tone adjustment coefficient and the earpiece volume adjustment coefficient.
Accordingly, the ringtone volume of the user terminal can be adjusted according to the ringtone volume adjustment coefficient of the user terminal, and the earpiece volume of the user terminal can be adjusted according to the earpiece volume adjustment coefficient.
Table 1 shows the ringtone volume adjustment coefficient and the earpiece volume adjustment coefficient set for different numbers of people in the following multiple scene modes. In each scene, the two values separated by slashes ("/") indicate, from left to right, the ringtone volume adjustment factor and the earpiece volume adjustment factor.
Volume adjustment shall be divided into 10 levels, ie 0.1-1.0. Items with "d" indicate coefficient values that need to be further determined according to the intensity of the environmental sound (decibel), where d indicates the intensity level of the environmental volume. Specifically, for example, a reasonable volume range that can be heard by the human ear is 20 to 120 decibels (all volumes above 120 decibels are calculated as 120), and the environmental volume is also 1 every 10 decibels. It can be divided into 10 levels according to the rule of level increase, that is, the range of values of d is 1, 2, ~ 10. For some items with d, the calculation result does not have to correspond exactly to 10 values from 0.1 to 1.0, close and larger values are selected, and the calculation result is less than 0.1 or 1.0. If greater than, these two boundary values are selected.
<tables num="1"><img id="000002" he="54" wi="159" file="JP6381153B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>
In a specific embodiment, step S102 described above performing an analysis of an audio signal collected to obtain composition information of the audio signal provided in an embodiment of the invention, as shown in FIG. Specifically, it can be carried out by the following method.
S201. The collected voice signal is divided into a plurality of voice data, for example, n pieces of S1, S2, and ~ Sn.
S202. Calculate the voice frequency of each voice data and classify each voice data into silence, human voice, and noise according to the calculated voice frequency.
S203. Collect statistics on silence, human voice, and noise ratios for all voice data. Specifically, the amount of various audio data is calculated individually and compared to the total amount of audio data to obtain a ratio.
S204. For voice data distinguished as human voice, calculate the Mel-Frequency Cepstral Coefficients (MFCC), and then determine the Mel frequency to determine information about the number of people included in the human voice. Collect statistics using voice data with the same cepstrum coefficient as one voice.
Specifically, for voice data distinguished as human voice, the calculation of MFCC features can be performed, and then the similarity between the two MFCC feature matrices is calculated. Multiple MFCC features with similar results can be considered as one voice, but conversely, multiple MFCC features are different voices and therefore the number of people included in the N samples. Information about can be obtained by collecting statistics.
It will be appreciated that this embodiment of the present invention is primarily based on frequency and spectral analysis to determine the composition information of the collected audio signal. Other similar frequency / spectral analysis methods can achieve all of this purpose, but are not listed one by one herein.
Specifically, the above-mentioned step S202, which classifies each voice data into silence, human voice, and noise according to the calculated voice frequency as shown in FIG. 3, can be performed by using the following procedure.
S301. Audio frequencies of the audio data, to determine if it is within a range of 20 Hz ~ 20000 Hz, the sound of each audio data if the voice frequency is within the range of 20 Hz ~ 20000 Hz performs step S302, each If the voice frequency of the voice data is not within the range of 20 Hz to 20000 Hz, step S306 is performed.
S302. Calculate the fundamental frequency of voice data. When a sounding body emits a vibrating sound, the sound can generally be decomposed into multiple pure sine waves. That is, every natural sound is basically formed by many sine waves of different frequencies, the lowest frequency sine wave is the fundamental frequency, and the fundamental frequency is used to distinguish between different sounding bodies. be able to.
S303. Determine if the fundamental frequency is in the range of 85Hz to 255Hz, if the fundamental frequency is in the range of 85Hz to 255Hz, perform step S304, and if the fundamental frequency is not in the range of 85Hz to 255Hz, perform step S304. , Step S305.
S304. Voice data is regarded as human voice.
S305. Voice data is regarded as noise.
S306. The voice data is regarded as silence.
Also, after the above-mentioned step S104 of adjusting the volume of the user terminal according to the determined scene mode provided in this embodiment of the present invention has been performed, and after the user connects the call signal during a call. If the call goes undisturbed (in a quiet environment), the voice intensity is generally fixed. If the caller considers the surrounding environment to be relatively noisy, the caller unknowingly increases the voice intensity of the call, or the caller is in a very quiet current environment (for example, as before). If you consider that multiple people were talking, but one answered the phone and everyone else stopped talking), you don't want the voice of the call to get in the way of others, or you're personal to the conversation. If information is included, the caller does not want other people to hear the content of the call, so the voice intensity of the call is lower than in the normal case. In the above example, the volume adjustment method provided in this embodiment of the present invention further provides a scheme for fine-tuning the playback volume twice after the call signal connection is detected, resulting in The effect is that the earpiece volume is adjusted to match the current situation.
Based on this, the above-mentioned volume adjusting method provided in this embodiment of the present invention, as shown in FIG. 4, further includes the following steps.
S401. Acquires the microphone volume of the user terminal in real time after the call signal connection is detected.
S402. Compare the magnitude of the acquired microphone volume with the pre-stored reference volume, and if the acquired microphone volume is higher than the pre-stored reference volume, execute step S403 and the acquired microphone volume is set in advance. If it is lower than the memorized reference volume, step S404 is executed, and if the acquired microphone volume is equal to the memorized reference volume, the procedure ends.
S403. Increase the earpiece volume of the user terminal.
S404. Decrease the earpiece volume of the user terminal.
In a specific embodiment, when the acquired microphone volume is compared with the pre-stored reference volume in S402, the pre-stored reference volume may be set to a single numerical value or set to a numerical range. You may. As long as the acquired microphone volume is within this value range, the acquired microphone volume can be regarded as equal to the reference volume, and the playback volume of the handset does not need to be adjusted.
The execution of steps S401 to S404 described above is stored in advance.<u style="single">Criteria</u>It is carried out based on the volume.<u style="single">Criteria</u>The volume is generally a fixed value and is determined and stored in the course of the call prior to the current call. Of course<u style="single">Criteria</u>The volume can also be updated as shown in FIG. 5, which can be done using the following steps.
S501. After the call signal connection is detected, the voice signal around the user terminal is collected periodically.
S502. Perform analysis of the collected audio signal to obtain the composition information of the audio signal. Please refer to steps S201 to S204 for the specific steps to be executed in the specific implementation.
S503. According to the composition information of the voice signal, it is determined whether the voice signal contains a human voice and contains only one human voice, the voice signal contains the human voice, and the voice signal contains the human voice, 1 Human-only voice If the voice signal contains only one voice, perform step S504, and if the audio signal does not contain human voice and does not contain only one voice, a scheme for fine-tuning the playback volume is provided. Execute twice, that is, execute steps S401 to S404.
S504. Calculates the volume of the audio signal, takes the average value of the calculated volume of the audio signal and the reference volume stored in advance, and stores the average value as a new reference volume.
In the above-mentioned volume adjustment method provided in this embodiment of the present invention, when the current scene mode of the user terminal is determined, a reference for the composition information of the audio signal is added, so that the corresponding current scene can be obtained. It is closer to the actual scene and can be more accurately matched to the scene in which the user is placed, which greatly increases the occurrence of the situation caused by the mistake of judging the scene that the adjustment of the playback volume does not match the scene. To improve the usability of the user.
Based on the same concept of the invention, the present invention further provides a device for adjusting the volume of a user terminal as shown in FIG. A collection unit 601 configured to collect audio signals around the user terminal, An analysis unit 602 configured to perform analysis of the collected voice signal in order to obtain the composition information of the voice signal, the composition information includes the voice type contained in the voice signal and the ratio of various voices. , Voice types include silence, human voice, and noise, with analysis unit 602, A scene mode determination unit 603 configured to determine the current scene mode of the user terminal according to the composition information of the audio signal, and It includes a volume control unit 604 configured to adjust the volume of the user terminal according to the determined scene mode.
Specifically, in the above-described apparatus provided in this embodiment of the present invention, the collection unit 601 is When the arrival of the call signal is detected, the voice signal of the current environment where the user terminal is located is collected, or the voice signal of the current environment where the user terminal is located is periodically collected. It is specifically configured.
Specifically, in the above-described apparatus provided in this embodiment of the present invention, the volume control unit 604 adjusts the volume according to the determined scene mode and the pre-stored correspondence between the scene mode and the volume control coefficient. It is specifically configured to determine a coefficient and adjust the volume of the user terminal according to the volume adjustment coefficient.
Specifically, in the above-described device provided in this embodiment of the present invention, the volume of the user terminal includes the ringtone volume and the earpiece volume, and the volume adjustment coefficient of the volume control unit 604 is specifically an incoming call. Includes volume adjustment factor and earpiece volume adjustment factor.
Specifically, the volume adjustment unit 604 is configured to adjust the ringtone volume of the user terminal according to the ringtone volume adjustment coefficient, and adjust the earpiece volume of the user terminal according to the earpiece volume adjustment coefficient.
Specifically, the above-described device provided in this embodiment of the present invention, as shown in FIG. An acquisition unit 605 configured to acquire the microphone volume of the user terminal in real time after the call signal connection is detected. Further equipped with a comparison unit 606 configured to compare the magnitude of the acquired microphone volume with the pre-stored reference volume. The volume adjustment unit 604 increases the earpiece volume of the user terminal when the acquired microphone volume is higher than the pre-stored reference volume, and increases the earpiece volume of the user terminal when the acquired microphone volume is lower than the pre-stored reference volume. It is further configured to reduce the earpiece volume.
Specifically, in the above-described apparatus provided in this embodiment of the present invention, the collection unit 601 periodically collects audio signals around the user terminal after the connection of the call signal is detected. It is specifically configured.
In order to acquire the composition information of the voice signal, the analysis unit 602 performs analysis of the voice signal periodically collected by the collection module 601 after the connection of the call signal is detected, and according to the composition information of the voice signal. , Determine if the audio signal contains a human voice, and if it contains only one voice, the audio signal contains a human voice, the audio signal contains a human voice, and only one voice If the voice is included, the volume of the audio signal is calculated, the average value of the calculated volume of the audio signal and the pre-stored reference volume is taken, and the average value is further configured to be stored as a new reference volume. To.
Specifically, in the device described above provided in this embodiment of the present invention, the analysis unit 602 A first processing unit configured to divide the collected voice signal into multiple voice data, and Calculate the voice frequency of each voice data, and the calculated voice<u style="single">frequency</u>A second processing unit configured to classify each voice data into silence, human voice, and noise according to With a third processing unit configured to collect statistics on the ratio of silence, human voice, and noise of all voice data, For voice data identified as human voice, the mel frequency cepstrum coefficient of the human voice voice data is calculated, and in order to determine information about the number of people included in the human voice, voice data having the same mel frequency cepstrum coefficient is used. It specifically includes a fourth processing unit configured to collect statistics as a single voice.
Specifically, in the above-described apparatus provided in this embodiment of the present invention, the second processing unit determines whether the voice frequency of each voice data is in the range of 20 Hz to 20000 Hz. When the voice frequency of the voice data is determined to be in the range of 20Hz to 20000Hz, the basic frequency of the voice data is calculated, and when it is determined that the basic frequency is in the range of 85Hz to 255Hz, When it is determined that the voice data is a human voice and the basic frequency is outside the range of 85Hz to 255Hz, the voice data is regarded as noise and the voice frequency of the voice data is in the range of 20Hz to 20000Hz. When it is determined to be outside, the audio data is specifically configured to be considered silent.
According to the volume control device described above provided in this embodiment of the present invention, when the current scene mode of the user terminal is determined, a reference for the composition information of the audio signal is added, so that the corresponding current volume control device is used. Occurrence of a situation caused by a misjudgment of the scene, where the scene is closer to the actual scene and can be more accurately matched to the scene in which the user is placed, which causes the playback volume adjustment to not match the scene. Is greatly reduced and the usability of the user is improved.
Based on the concept of the same invention, embodiments of the present invention further provide a speaker, a handset, and a terminal with the volume control device described above provided in the embodiment of the invention, where the volume control device is the volume of the speaker. , And are configured to adjust the volume of the handset. Specifically, the terminal may be any product or component having a playback function, such as a mobile phone, transceiver, tablet computer, television, display, or laptop computer. For the implementation at the terminal, it is desired to refer to the above-described embodiment of the device for controlling the playback volume, and the overlapping portion is not described again here. An embodiment of the present invention that provides another terminal, as shown in FIG. A voice sensor 150 configured to collect voice signals around the user terminal 100, configured to emit a ring tone of an incoming call when a ringing signal reaches the user terminal 100. The speaker 130 is provided, and the speaker 130 may be further configured to reproduce audio data such as music. It will be appreciated that the handset 170 is configured to reproduce the other party's voice when the user is talking to the other party using the user terminal 100.
The terminal 100 further comprises a display unit 140, which can be configured to display information entered by the user, or information provided to the user, and various menu interfaces of the terminal 100. The display unit 140 can include a display panel 141, and if necessary, the display panel 141 may be an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like. May be good.
In some embodiments, memory 120 stores an executable module or data structure, or a subset thereof or an extension set thereof.
In this embodiment of the present invention, by calling a program or instruction stored in the memory 120, the processor 160 performs analysis of the voice signal collected by the voice sensor 150 in order to obtain composition information of the voice signal. The composition information includes the voice type contained in the voice signal and the ratio of various types of voice, and the voice type includes silence, human voice, and noise, and the user terminal according to the composition information of the voice signal. Determines the current scene mode of, and adjusts the volume of speaker 130 and / or the volume of handset 170 according to the determined scene mode.
If necessary, in one embodiment, the voice sensor 150 acquires the volume of the microphone 110 after the call signal connection is detected.
Processor 160 When the acquired microphone volume is louder than the pre-stored reference volume, increase the earpiece volume of the user terminal, and when the acquired microphone volume is lower than the pre-stored reference volume, decrease the earpiece volume of the user terminal. Is further configured.
The voice sensor 150 is a unit configured to collect voice signals, and the voice sensor 150 may be specifically integrated with the microphone 110, or is another component not limited in the present invention. Please note that it may be.
In addition, the terminal device 100 can further implement the methods and embodiments of FIGS. 1-5, details of which are not described again herein in this embodiment of the present invention.
According to the user terminal described above provided in this embodiment of the present invention, when the current scene mode of the user terminal is determined, a reference for the composition information of the audio signal is added, so that the corresponding current scene Is closer to the actual scene and can be more accurately matched to the scene in which the user is placed, which causes a situation caused by an erroneous scene judgment that the playback volume adjustment does not match the scene. Significantly reduce and improve user experience.
Based on the above description of the methods of practice, those skilled in the art will clearly understand that embodiments of the present invention can be implemented by hardware or software in addition to the general hardware platforms required. Let's do it. Based on this understanding, the technical solutions in the embodiments of the present invention can be implemented in the form of software products. The software product can be stored in a non-volatile storage medium (which may be a CD-ROM, USB flash drive, removable hard disk, etc.) so as to perform some of the methods described in the embodiments of the present invention. Includes several instructions that direct a computer device (eg, a personal computer, server, or network device).
Those skilled in the art will appreciate that the accompanying drawings are merely schematic representations of the exemplary embodiments, and that the modules or steps of the accompanying drawings are not necessarily necessary for the practice of the present invention.
Modules of the devices provided in this embodiment may be distributed in the device according to the description of the embodiment, or may be placed in one or more devices different from those described in the present embodiment. Those skilled in the art will understand that it is good. The modules of the above embodiments may be combined into one module or divided into a plurality of submodules.
The serial numbers of the above-described embodiments of the present invention are for illustrative purposes only and are not intended to indicate the priority of the embodiments.
It goes without saying that a person skilled in the art can make various modifications and modifications to the present invention without departing from the spirit and scope of the present invention. The present invention is intended to cover such modifications and modifications as long as they are within the scope of the claims below and the protection defined by the equivalent art.
100 user terminal 110 microphone 120 memory 130 speaker 140 display unit 141 Display panel 150 voice sensor 160 processor 170 Handset 601 collection unit 602 Analytical unit 603 Scene mode determination unit 604 Volume control unit 605 Acquisition unit 606 comparison unit
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2003023474A | Cites | Japan |
| JP2001223767A | Cites | Japan |
| JP10136058A | Cites | Japan |
| JP2010258687A | Cites | Japan |
| US20130143543A1 | Cites | United States of America |
| US20050282590A1 | Cites | United States of America |
26 members in 7 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 201410152521 | China | A | |
| 201410152521 | China | A | |
| 2014101525219 | China | – | |
| 2015072906 | China | W | |
| 2015072906 | China | W | |
| 2014101525219 | – | – | – |
| CN201410152521 | – | – | – |
| CN20141152521 | – | – | – |
| CN2015072906 | – | – | – |
| WO2015CN72906 | – | – | – |
Members26
| Document | Office | Kind | |
|---|---|---|---|
| CN103945062A | China | A | |
| WO2015158182A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201541344A | Taiwan Province of China | A | |
| KR20160145730A | Republic of Korea | A | |
| CN103945062B | China | B | |
| US2017034362A1 | United States of America | A1 | |
| TWI570624B | Taiwan Province of China | B | |
| EP3133799A1 | European Patent Office (EPO) | A1 | |
| EP3133799A4 | European Patent Office (EPO) | A4 | |
| JP2017514392A | Japan | A | |
| US9866707B2 | United States of America | B2 | |
| US2018084116A1 | United States of America | A1 | |
| KR101884709B1 | Republic of Korea | B1 | |
| JP6381153B2This record | Japan | B2 | |
| US10200545B2 | United States of America | B2 | |
| US2019132453A1 | United States of America | A1 | |
| US2019149668A1 | United States of America | A1 | |
| EP3133799B1 | European Patent Office (EPO) | B1 | |
| US10516788B2 | United States of America | B2 | |
| US10554826B2 | United States of America | B2 | |
| EP3611910A1 | European Patent Office (EPO) | A1 | |
| US2020145539A1 | United States of America | A1 | |
| US11044369B2 | United States of America | B2 | |
| US2021274049A1 | United States of America | A1 | |
| US11483434B2 | United States of America | B2 | |
| EP3611910B1 | European Patent Office (EPO) | B1 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6381153
- Publication, DOCDB
- 6381153
- Publication, EPODOC
- JP6381153B
- Application
- 2016563102
- Application, DOCDB
- 2016563102
- Application, EPODOC
- JP20160563102
Titles2
- Japanese
- ユーザ端末ならびに端末の音量を調整する方法および装置
- English
- User terminals and methods and devices for adjusting the volume of terminals
Classification
- CPC, 9
- H04M19/044
- H04M1/724
- H04M1/725
- H04M1/60
- H04M2250/12
- H04M1/72454
- G10L21/0232
- H03G3/32
- H04W88/02
- IPC, 6
- H04M1 00
- G10L21 0332
- G10L25 84
- H04M1 72454
- H04M1 725
- H04R3 00
