Volume adjustment method and device, electronic equipment and storage medium
Abstract
The embodiment of the invention relates to the field of speech recognition, and discloses a volume adjustment method and device, electronic equipment and a storage medium. The volume adjustment methodcomprises the steps: acquiring each audio sample in a training set for training a speech recognition model, wherein the speech recognition model is used for speech recognition; determining a volume value of each audio sample in the training set; determining a volume reference value of the training set according to the volume value of each audio sample; and adjusting the volume value of each audiosample according to the volume reference value, wherein the difference between the adjusted volume value of each audio sample and the volume reference value is within a preset difference range. According to the volume adjustment method provided by the embodiment of the invention, volume adjustment can be performed on each piece of audio data based on the whole training set, and the volume value of the audio samples in the training set is properly adjusted, so that the recognition effect of the speech recognition model is improved.

Term
13.9 yearsto projected expiry
Projected expiry 28 August 2040, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
10 claims: 2 independent, 8 dependent
- 11 A method for volume adjustment, comprising:obtaining each audio sample in a training set for training a speech recognition model;wherein the speech recognition model is used for speech recognition;determining each audio sample in the training set According to the volume value of each audio sample, determine the volume reference value of the training set;According to the volume reference value, adjust the volume value of each audio sample;wherein, the adjusted The difference between the volume value of each audio sample and the volume reference value is within a preset difference range. 1 .一种音量调节的方法,其特征在于,包括: 获取用于训练语音识别模型的训练集中的各音频样本;其中,所述语音识别模型用于 语音识别; 确定所述训练集中的各音频样本的音量值; 根据所述各音频样本的音量值,确定所述训练集的音量基准值; 根据所述音量基准值,对所述各音频样本的音量值进行调节;其中,调节后的所述各音 频样本的音量值与所述音量基准值的差值在预设的差值范围内。
- 88 A device for realizing volume adjustment, characterized in that it comprises:8 .一种音量调节的实现装置,其特征在于,包括: An acquisition module for acquiring each audio sample in a training set for training a voice recognition model;wherein the voice recognition model is used for voice recognition;a calculation module for determining the volume value of each audio sample in the training set;The statistics module is used to determine the volume reference value of the training set according to the volume value of each audio sample;the adjustment module is used to adjust the volume value of each audio sample according to the volume reference value;wherein , The difference between the adjusted volume value of each audio sample and the volume reference value is within a preset difference range. 获取模块,用于获取用于训练语音识别模型的训练集中的各音频样本;其中,所述语音 识别模型用于语音识别; 计算模块,用于确定所述训练集中的各音频样本的音量值; 统计模块,用于根据所述各音频样本的音量值,确定所述训练集的音量基准值; 调节模块,用于根据所述音量基准值,对所述各音频样本的音量值进行调节;其中,调 节后的所述各音频样本的音量值与所述音量基准值的差值在预设的差值范围内。
Independent claims2
116 paragraphs, as filed
Volume adjustment method, device, electronic equipment and storage mediumTechnical field
[0001] The embodiments of the present invention relate to the field of speech recognition, and in particular, to a method, device, electronic device, and storage medium for volume adjustment.
Background technique
[0002] With the development of computer technology, voice recognition technology is applied to more and more fields, such as smart home, industrial control, voice interaction systems for terminal devices, and so on. Using speech recognition technology can make information processing and acquisition more convenient, thereby improving work efficiency. The speech recognition model is obtained based on a large amount of audio data, through deep neural network for learning and inference, and iterative training. The quality of the audio data used for training will greatly affect the performance of the speech recognition model.
[0003] The inventor found that there are at least the following problems in the prior art: the recognition effect of the speech recognition model is heavily dependent on the quality of the trained audio data. The prior art first performs high-pass filtering on the training samples, which will filter out part of the effective audio data. Data, the training samples processed by high-pass filtering are processed by automatic gain control. However, for the case where the audio itself is very small or very loud, the automatic gain effect is not good, and the volume information cannot be adjusted well, which ultimately leads to the failure of the speech recognition model. The recognition effect is poor.
Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a volume adjustment method, device, electronic equipment and storage medium, which can adjust the volume of each piece of audio data based on the entire training set, and appropriately adjust the volume value of the audio sample in the training set. , So as to improve the recognition effect of the speech recognition model.
[0005] In order to solve the above technical problems, the embodiments of the present invention provide a method for volume adjustment, including the following steps: acquiring each audio sample in a training set for training a speech recognition model; wherein the speech recognition model is used In speech recognition; determine the volume value of each audio sample in the training set; determine the volume reference value of the training set according to the volume value of each audio sample; determine the volume reference value of the training set according to the volume reference value; The volume value of is adjusted; wherein the difference between the adjusted volume value of each audio sample and the volume reference value is within a preset difference range.
[0006] The embodiment of the present invention also provides a volume adjustment device, including: an acquisition module for acquiring each audio sample in a training set for training a speech recognition model; wherein the speech recognition model is used for speech Recognition; a calculation module for determining the volume value of each audio sample in the training set; a statistics module for determining the volume reference value of the training set according to the volume value of each audio sample; an adjustment module for The volume value of each audio sample is adjusted according to the volume reference value; wherein the difference between the adjusted volume value of each audio sample and the volume reference value is within a preset difference range.
[0007] Embodiments of the present invention also provide an electronic device, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein the memory stores the memory that can be used by the at least one processor; An instruction executed by the processor, the instruction being executed by the at least one processor, so that the at least one processor can execute the above-mentioned volume adjustment method.
[0008] The embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the
When the computer program is executed by the processor, the method for implementing the above-mentioned volume adjustment is realized.
[0009] Compared with the prior art, the embodiment of the present invention obtains each audio sample in a training set for training a speech recognition model; wherein, the speech recognition model is used for speech recognition; determines each of the training set The volume value of the audio sample. Considering that the prior art will perform high-pass filtering processing on the acquired audio samples, but the high-pass filter will filter out part of the effective data in the audio samples, which is not conducive to the training of the speech recognition model, and the embodiments of the present invention directly perform the processing on the acquired audio samples. The audio samples are processed to maximize the integrity of the samples. Further, the volume reference value of the training set is determined according to the volume value of each audio sample, and the volume value of each audio sample is adjusted according to the volume reference value; wherein, each of the adjusted audio samples The difference between the volume value of the audio sample and the volume reference value is within a preset difference range. Considering that the volume value of each audio sample in the training set is different, the volume value of some of the samples may be too large or too small. When adjusting the volume value of the audio sample in the prior art, the audio sample processed by high-pass filtering is automatically performed. Gain control (automatic gain control, abbreviated as: AGC) processing, but for the case where the volume value of the sample itself is too large or too small, the automatic gain control effect is not good, and the embodiment of the present invention determines the volume value of each audio sample in the training set , Determine the volume reference value based on the volume value of each audio sample in the entire training set, and adjust each tone The volume of the frequency samples can make the adjusted volume value of each audio sample in the entire training set not much different from the volume reference value. Use audio samples with appropriate volume values to train the speech recognition model, thereby improving the recognition effect of the speech recognition model.
[0010] In addition, determining the volume reference value of the training set according to the volume value of each audio sample includes: selecting N audio samples in the training set according to the volume value of each audio sample; Wherein, the volume value of the N audio samples is within a preset volume value range, and the N is a natural number greater than 1; the volume average value of the N audio samples is determined; the volume average value is used as the The volume reference value of the training set. In order to reduce the adverse effects of audio samples whose volume values are too large or too small on the determination of the volume reference value of the training set, the embodiment of the present invention selects audio samples with a volume value within a preset volume value range in the training set. To a certain extent, the work efficiency can be improved, and the volume average value of the selected multiple audio samples is calculated as the volume reference value, so that the determined volume reference value is more in line with the needs of training.
[0011] In addition, selecting audio samples in the training set according to the volume values of the audio samples includes: sorting the audio samples in the training set according to the volume values of the audio samples; determining all the audio samples in the training set; The median of the volume value of each audio sample after sorting, and the audio sample corresponding to the median as the target audio sample; taking the sorting position of the target audio sample as the starting point for selection, starting from sorting in the target N audio samples are selected from each audio sample on both sides of the audio sample. Selecting audio samples based on the median of the volume value can better attenuate the adverse effects of samples that are too loud or too small on the entire training set, so that the volume value of the selected audio sample is more in line with the training needs.
[0012] In addition, determining the average volume of the selected audio samples includes: determining the weight coefficient of the N audio samples; and determining the weight of the volume value of the N audio samples according to the weight coefficient Average value; the weighted average value is used as the volume average value. Determine the weight coefficient and perform a weighted average of the selected audio samples to obtain a volume average value that is more in line with the training needs of the speech recognition model.
[0013] In addition, determining the volume value of each audio sample in the training set includes: determining the volume value of each frequency in each of the audio samples in the training set; according to the volume value of each frequency in the audio sample , Determine the average value of the volume value of each frequency; taking the average value as the volume value of the audio sample, the volume value of each frequency in an audio sample can be comprehensively considered, so that the volume value of the obtained audio sample is more accurate .
[0014] In addition, adjusting the volume value of each audio sample according to the volume reference value includes:
The volume reference value adjusts the volume value of each frequency in each audio sample. The volume value of each frequency in an audio sample is adjusted, so that the volume value of the entire audio sample is closer to the volume reference value, and the recognition effect of the speech recognition model is further improved.
Description of the drawings
[0015] One or more embodiments are exemplified by the pictures in the corresponding drawings, and these exemplified descriptions do not constitute a limitation on the embodiments.
[0016] FIG. 1 is a flowchart of a method for volume adjustment according to a first embodiment of the present invention;
[0017] FIG. 2 is a flowchart of the sub-steps of determining the volume reference value of the training set according to the volume value of each audio sample in the first embodiment of the present invention;
[0018] FIG. 3 is a flowchart of a method for volume adjustment according to a second embodiment of the present invention;
[0019] FIG. 4 is a flowchart of the sub-steps of determining the average volume of selected audio samples according to the second embodiment of the present invention;
[0020] FIG. 5 is a flowchart of a method for volume adjustment according to a third embodiment of the present invention;
[0021] FIG. 6 is a block diagram of a volume adjustment device according to a fourth embodiment of the present invention;
[0022] FIG. 7 is a schematic structural diagram of an electronic device according to a fifth embodiment of the present invention.
Detailed ways
[0023] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, a person of ordinary skill in the art can understand that in each embodiment of the present invention, many technical details are proposed for the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solution claimed in this application can be realized. The following division of the various embodiments is for convenience of description, and should not constitute any limitation on the specific implementation of the present invention, and the various embodiments may be combined with each other without contradiction.
[0024] The first embodiment of the present invention relates to a method for volume adjustment, which is applied to an electronic device; wherein, the electronic device may be a terminal or a server. In this embodiment and the following embodiments, the electronic device uses the server as an example Description. The following specifically describes the implementation details of the volume adjustment method of this embodiment. The following content is only provided for ease of understanding and is not necessary for implementing this solution.
[0025] The specific process of the volume adjustment method of this embodiment may be as shown in FIG. 1, and includes:
[0026] Step 101: Obtain each audio sample in a training set for training a speech recognition model;
[0027] Specifically, a speech recognition model is used for speech recognition. When the server trains the speech recognition model, it uses a large number of audio samples to train the speech recognition model based on machine learning methods. The audio samples used for training are collected in advance, that is, the training set of the speech recognition model comes from the actual production and life of human society. Rich and authentic.
[0028] In a specific implementation, the training set used to train the speech recognition model can be obtained from the Internet, and the content of each audio sample includes at least natural language information. After the server obtains the training set used to train the speech recognition model, it can Preprocessing is performed on pieces of audio samples, where the preprocessing operations include but are not limited to: performing echo cancellation processing on each piece of audio samples; performing noise suppression processing on each piece of audio samples, and so on. Performing operations such as echo cancellation and noise suppression can make the obtained audio samples more pure. Using such audio samples for training can improve the recognition effect of the speech recognition model.
[0029] In an example, the server obtains a training set for training a speech recognition model, and the speech recognition model should
Used in the field of smart home, the voice content of each audio sample in the training set can include: "turn off the TV", "switch the air conditioner to cooling mode", "set the bathroom water heater temperature to 62 degrees" and so on. The server uses adaptive filtering technology to process each audio sample to perform echo cancellation, and input the echo-cancelled audio sample into a Recurrent Neural Network (RNN) for noise suppression processing. It should be noted that this example only takes the application of the speech recognition model in the field of smart homes as an example. In actual applications, it is not limited to this. The audio samples in the training set can be targeted according to the application field of the speech recognition model. Select to improve the recognition effect of the speech recognition model obtained after training. In other words, the audio samples in the training set can be obtained according to the application field of the speech recognition model.
[0030] Step 102, determine the volume value of each audio sample in the training set;
[0031] Specifically, the server may obtain the volume value of each audio sample after preprocessing each audio sample in the training set. Taking into account the characteristics of human pronunciation, a person usually has changes in pitch and loudness when speaking, and different pitches and loudness may contain different information. Among them, the loudness when a human speaks can be understood as the volume value of the acquired audio sample. By acquiring the volume value of each audio sample in the training set, combined with the human pronunciation habits, the information contained in each audio sample can be accurately judged. At the same time, grasping the volume value of each audio sample in the training set can maintain the integrity of the sample to the greatest extent.
[0032] In a specific implementation, since the human ears perceive the size of sound in a logarithmic relationship, in the field of acoustics, decibel (abbreviation: dB) is usually used to describe the sound size. In acoustics, 20 micropascals are generally used as a reference value when calculating the decibel value, that is, the minimum human auditory response value, where 20 micropascals corresponds to OdB, and the following formula can be used to calculate the decibel value of sound: Lp = 201 Tao. (Black) For example, where pm is the reference sound pressure, generally 2×1 (Γ5ρΗ, ρ3 is the root mean square value of the target audio sample sound pressure, and Lp represents the sound pressure level, that is, the volume value of the audio sample.
[0033] In an example, the server may determine the volume value of an audio sample by acquiring the volume value of the audio sample at each time. For example, the server obtains the training set of the speech recognition model applied to the smart home field. The training set contains audio samples with a long duration, such as an audio sample with a duration of 4 seconds: "Set the temperature of the bathroom water heater to 62 degrees." The server uses the volumedetect module of ffmpeg technology to obtain the corresponding volume value of the audio sample per second as 55(18, 60(18, 62(18, 59(18, calculate the average value of the volume value at each time to 59(18, then The average value of the volume values at each time, that is, 59 dB, is used as the volume value of the audio sample.
[0034] Step 103: Determine the volume reference value of the training set according to the volume value of each audio sample;
[0035] Specifically, after determining the volume value of each audio sample in the training set, the server comprehensively considers the characteristics of the volume value of each audio sample in the entire training set to determine the volume reference value of the training set. Taking into account the different application scenarios of the speech recognition model, the volume value of the audio sample suitable for training the speech recognition model is also different. Comprehensive consideration of the volume value of each audio sample in the entire training set can make the determined volume reference value more suitable for training target speech recognition Model, thereby improving the recognition effect of the speech recognition model.
[0036] In one example, the server obtains a training set for training a speech recognition model, which is used in the field of industrial control. Considering the roar of machines and noisy sounds in the factory environment, workers need to increase their loudness when talking to each other in order to hear each other clearly. , When using a voice recognition system for human-computer interaction, a higher volume is required to interact. The training set contains an audio sample with a volume value of 81dB. "Turn off the No. 3 engine, 81dB is the volume value suitable for the noisy environment of the factory, and the server can use the volume value of this audio sample, that is, 81dB as the volume reference value of the training set.
[0037] In another example, according to the volume value of each audio sample, the volume reference value of the training set can be determined as follows:
The sub-steps shown in Figure 2 are realized:
[0038] Sub-step 1031, according to the volume value of each audio sample, select N audio samples in the training set;
[0039] Wherein, the volume values of the N audio samples are within a preset volume value range, and N is a natural number greater than 1. The number of selected audio samples N can be set by the developer. Selecting a certain number of audio samples in the training set can improve work efficiency to a certain extent. The preset volume value range can be set by the developer according to different application scenarios.
[0040] Specifically, the training set obtained by the server contains a large number of audio samples with uneven quality and volume values. If the server determines the volume reference value of the training set based on some audio samples that are too large or too small, the volume will be caused The reference value has a large deviation, and the selection of audio samples within the preset volume value range can well meet the training needs of the target speech recognition model.
[0041] In an example, the server obtains a training set for training a speech recognition model, which is used in the industrial control field. The training set contains 2000 audio samples, and the preset volume value range is based on the factorys noisy environment settings. From 70dB to 110dB, there are 900 audio samples with a volume value of 70dB to 110dB in the training set, and the server selects 600 audio samples among them.
[0042] In another example, the server obtains a training set for training a speech recognition model, which is applied to the field of intelligent medical care. The training set contains 1500 audio samples, and the preset volume value range is based on the quietness of the hospital ward. The environment is set to 30dB to 70dB, there are 1400 audio samples with the volume value of 30dB to 70dB in the training set, and the server selects 1300 audio samples among them.
[0043] Sub-step 1032, determine the average volume of N audio samples;
[0044] Specifically, after selecting the required audio sample, the server calculates the average volume according to the volume value of the selected audio sample.
[0045] In an example, the server selects a total of 300 audio samples with a volume value of 70 dB to 110 dB in the training set, and calculates an average volume of 88 dB based on the volume values of the 300 audio samples.
[0046] Sub-step 1033, the average volume is used as the volume reference value of the training set.
[0047] Specifically, after the server calculates the volume average value of the selected audio samples, the volume average value is used as the volume reference value of the training set, which can weaken the determination of the average volume due to the excessive or small volume value of some audio samples. The deviation caused by the value, so as to obtain the volume reference value suitable for training the speech recognition model.
[0048] In an example, in the training set used to train the speech recognition model, the server selects a total of 1350 audio samples with a volume value of 30dB to 70dB. The speech recognition model is applied to the field of intelligent medical care, and the 1350 audio samples are calculated The average volume of is 52dB, and 52dB is used as the volume reference value of the training set. 52dB is also the volume value suitable for training the voice recognition model applied in the field of intelligent medical treatment.
[0049] Step 104: Adjust the volume value of each audio sample according to the volume reference value.
[0050] Wherein, the difference between the adjusted volume value of each audio sample and the volume reference value is within a preset difference range, and the preset difference range can be set by the developer according to actual needs. Considering that when the volume value of each audio sample is actually adjusted, there may be deviations, so it is necessary to set the acceptable error range of the difference range.
[0051] Specifically, after determining the volume reference value of the training set, the server adjusts the volume value of each audio sample according to the volume reference value, and the reference value can be understood as the optimal value of the volume in the training set. Lower the volume value of audio samples higher than the volume reference value, and raise the volume value of audio samples lower than the volume reference value, so that the volume value of each audio sample in the training set is close to the optimal value, and the training is adjusted appropriately The volume value of each audio sample is concentrated to improve the accuracy of the speech recognition model.
[0052] In an example, the server can directly modify the data in the audio sample file to adjust the volume value of the audio sample. For example, the server uses ffmpeg technology to parse the relevant data in the audio sample file, and issue an adjustment instruction to the data representing the volume value. The adjustment instruction can be a piece of code to directly modify the volume of the audio file in decibels, and adjust the volume value of the audio file to The volume reference value or close to the volume reference value.
[0053] In another example, the server may use audio visualization software to adjust the volume value of the audio sample. For example, the server inputs audio files into audio and video editing software such as Premiere, and adjusts the audio "amplitude" through the visual operation interface provided by the audio and video editing software to adjust the volume value of the audio file. The volume reference value can be expressed as a horizontal line in the visualization interface. In the visualization page, the volume value above the horizontal line is lowered by operations such as "peak clipping and valley filling", and the volume value below the horizontal line is increased, that is, the volume of the audio file Adjust the value to the volume reference value or close to the volume reference value.
[0054] Compared with the prior art, the first embodiment of the present invention obtains each audio sample in a training set for training a speech recognition model; wherein the speech recognition model is used for speech recognition; and determines the training set The volume value of each audio sample. Considering that the prior art will perform high-pass filtering processing on the acquired audio samples, but the high-pass filter will filter out part of the effective data in the audio samples, which is not conducive to the training of the speech recognition model, and the embodiments of the present invention directly perform the processing on the acquired audio samples. The audio samples are processed to maximize the integrity of the samples. Further, the volume reference value of the training set is determined according to the volume value of each audio sample, and the volume value of each audio sample is adjusted according to the volume reference value; wherein, each of the adjusted audio samples The difference between the volume value of the audio sample and the volume reference value is within a preset difference range. Considering that the volume value of each audio sample in the training set is different, the volume value of some of the samples may be too large or too small. When the volume value of the audio sample is adjusted in the prior art, the audio sample processed by the high-pass filter is subjected to AGC Processing, but the automatic gain control effect is not good for the case where the volume of the sample itself is too large or too small, and the embodiment of the present invention determines the volume value of each audio sample in the training set based on the volume value of each audio sample in the entire training set Volume reference value, adjust the volume value of each audio sample according to the reference value, which can make each tone in the entire training set The volume value after the frequency sample adjustment is not much different from the volume reference value, and the voice recognition model is trained with the audio sample with the appropriate volume value to improve the recognition effect of the voice recognition model.
[0055] The second embodiment of the present invention relates to a method of volume adjustment. The following is a detailed description of the implementation details of the volume adjustment method of this embodiment. The following content is only provided for ease of understanding and is not necessary to implement this solution. Figure 3 is the volume adjustment method described in the second embodiment. Schematic diagram, including:
[0056] Step 201: Obtain each audio sample in the training set used to train the speech recognition model;
[0057] Step 202: Determine the volume value of each audio sample in the training set;
[0058] Wherein, step 201 to step 202 have been described in the first embodiment, and will not be repeated here.
[0059] Step 203, sort the audio samples in the training set according to the volume value of each audio sample;
[0060] Specifically, after determining the volume value of each audio sample in the training set, the server sorts the audio samples according to the size of the volume value. The sorting method can be sorted by volume value from large to small, or by volume value from small to large, and the position of each audio sample after sorting in the sequence is determined.
[0061] In an example, the server sorts the 1000 audio samples in the acquired training set from small to large according to the volume value, and saves the positions of the sorted 1000 audio samples in the sequence. For example, the volume of an audio sample "turn off the bedside lamp" is 55dB, which ranks 478 out of 1000 audio samples.
[0062] Step 204: Determine the median of the volume value of each audio sample after sorting, and use the audio sample corresponding to the median as the target audio sample;
[0063] Specifically, considering that the acquired training set is used to train a speech recognition model, in the speech recognition process, humans usually use the volume during normal conversation for human-computer interaction, which is about 60 dB. The training set obtained by the server contains a large number of audio samples. These audio samples are basically collected from the actual production and life of human beings. When the training set contains thousands of audio samples, the median of the volume value of the training set is basically normal. Within the volume range during conversation.
[0064] Step 205, using the sorting position of the target audio sample as a starting point for selection, select N audio samples from each audio sample sorted on both sides of the target audio sample;
[0065] Wherein, the number of selected audio samples N can be set by the developer according to actual needs. Taking the sorting position of the target audio sample as the starting point for selection, N audio samples are selected from each audio sample sorted on both sides of the target audio sample. Samples can improve work efficiency to a certain extent. Take the target audio sample as a starting point to select audio samples from both sides, so that the volume of the selected audio file is close to the normal sound range. That is to say, the audio samples are selected from both sides with the sorting position of the audio samples corresponding to the median as the starting point. When the number of selected audio samples reaches the preset ratio of the training set, the selection process is stopped to make the selected audio files The volume is close to the normal sound range. Among them, the preset ratio is the ratio of N to the total number of samples in the training set.
[0066] In an example, the sorted training set contains 1000 audio samples, the server determines that the median of the volume value of the training set is 63dB, and the server takes the audio sample corresponding to 63dB as the starting point and selects from both sides of the median A total of 800 audio samples.
[0067] Step 206, determine the average volume of the selected N audio samples;
[0068] In an example, determining the volume average value of the selected N audio samples is implemented in each sub-step as shown in FIG. 4:
[0069] Sub-step 2061, determining the weight coefficients of the selected N audio samples;
[0070] In a specific implementation, the server can determine the weight coefficients of the selected N audio samples according to the positions in the sequence and actual needs of the selected N audio samples after sorting. Through the setting of different weight coefficients, the final acquisition can be guaranteed. The volume reference value meets the actual needs of training.
[0071] In an example, the server may determine the weight coefficients of the N audio samples according to the sorting positions of the N audio samples relative to the target audio sample; wherein, the weight coefficients of the N audio samples are based on the weight coefficient of the target audio sample Decrease the initial value to decrease to both sides. For example, the server determines that the median of the volume value of the training set is 63dB, takes the audio sample corresponding to 63dB as the target audio sample and selects a total of 100 audio samples from both sides, and the server sets the audio sample corresponding to the median of the volume value. That is, the weight coefficient of the target audio sample is 0.3, and the weight coefficient decreases from the target audio sample to both sides, and it is ensured that the sum of all weight coefficients is equal to 1.
[0072] Sub-step 2062: Determine a weighted average of the volume values of N audio samples according to the weight coefficient, and use the weighted average as the volume average;
[0073] Specifically, after the server determines the weight coefficients of the selected N pieces of audio samples, it performs a weighted average according to the volume values and weight coefficients of the N pieces of audio samples, and uses the weighted average value obtained as the volume average value, using the weighted average value. The obtained volume reference value can be more in line with actual training needs.
[0074] Step 207, use the average volume as the volume reference value of the training set;
[0075] Step 208, adjust the volume value of each audio sample according to the volume reference value;
[0076] Wherein, step 207 to step 208 have been described in the first embodiment, and will not be repeated here.
[0077] Compared with the prior art, in this embodiment, according to the volume value of each audio sample, select
Taking audio samples includes: sorting each audio sample in the training set according to the volume value of each audio sample; determining the median of the volume value of each audio sample after the sorting, and combining The audio sample corresponding to the number of bits is taken as the target audio sample; taking the sorting position of the target audio sample as the starting point for selection, N audio samples are selected from each audio sample sorted on both sides of the target audio sample; wherein, the selected The number of audio samples accounts for a preset ratio of the total number of audio samples in the training set. Selecting audio samples based on the median of the volume value can better attenuate the adverse effects of samples that are too loud or too small on the entire training set, so that the volume value of the selected audio sample is more in line with the training needs. Determining the average volume value of the selected audio samples includes: determining the weight coefficient of the N audio samples; determining the weighted average value of the volume values of the N audio samples according to the weight coefficient; The weighted average value is used as the volume average value. Determine the weight coefficient and perform a weighted average of the selected audio samples to obtain a volume average value that is more in line with the training needs of the speech recognition model.
[0078] The third embodiment of the present invention relates to a method of volume adjustment. The following is a detailed description of the implementation details of the volume adjustment method of this embodiment. The following content is only provided for ease of understanding and is not necessary to implement this solution. Figure 5 is the volume adjustment method described in the third embodiment. Schematic diagram, including:
[0079] Step 301: Obtain each audio sample in the training set used to train the speech recognition model;
[0080] Wherein, step 301 has been described in the first embodiment, and will not be repeated here.
[0081] Step 302, determine the volume value of each frequency in each audio sample in the training set;
[0082] Specifically, considering that the frequency of human occurrence ranges from 85 Hz to 1100 Hz, a persons voice can contain different frequencies when speaking. Correspondingly, an audio sample also contains different frequencies, and sounds at different frequencies are Can correspond to different volume values. Therefore, the server can determine the volume value of each frequency in each audio sample, which can further ensure the integrity of the sample.
[0083] In an example, the server uses pulse code modulation (Pulse Code Modulation, PCM for short) technology to determine the volume value of each frequency in the audio sample. For example: the server converts audio samples into PCM data between -1 and 1, and performs fast Fourier transform (FFT) on the PCM data to obtain a spectrogram of the audio sample, according to the ordinate of the spectrogram , The energy of sound waves, using the formula: 10log10 (a<sup>2</sup>+b<sup>2</sup>) Calculate the volume value of each frequency of the audio sample, where a represents the real part and b represents the imaginary part.
[0084] In another example, the server samples the audio samples, and determines the volume value of each frequency in the audio sample by a method of proportional mapping. For example, the server samples audio samples, records the energy value of each sampling point, and maps the energy value of each sampling point to a ratio of 1-100. Under normal circumstances, the human voice is distributed in a lower energy range, and the quantized value is roughly distributed in the interval of 1-20. The quantized value is amplified by 5 times. For values less than 100, use the formula: 1010g (10X magnification After the quantization value), the volume value of each frequency of the audio sample is calculated. For values greater than 100, directly assign the volume value to 100dB.
[0085] Step 303: Determine the average value of the volume value of each frequency according to the volume value of each frequency in the audio sample, and use the average value as the volume value of the audio sample;
[0086] Specifically, after determining the volume value of each frequency of the audio sample, the server may calculate an average value, and use the average value as the volume value of the audio sample. The volume value of each frequency in an audio sample is comprehensively considered to make the determined volume value of the audio sample more accurate.
[0087] Step 304: Determine the volume reference value of the training set according to the volume value of each audio sample;
[0088] Wherein, step 304 has been described in the first embodiment, and will not be repeated here.
[0089] Step 305, according to the volume reference value, adjust the volume value of each frequency in each audio sample.
[0090] Specifically, when the server adjusts the volume value of each audio sample according to the volume reference value, it may adjust the volume value of each frequency of the audio sample. The volume value of the entire audio sample can be made close to the volume reference value, which further improves the recognition effect of the speech recognition model.
[0091] Compared with the prior art, in this embodiment, determining the volume value of each audio sample in the training set includes: determining the volume value of each frequency in each audio sample in the training set; The volume value of each frequency in the audio sample determines the average value of the volume value of each frequency; taking the average value as the volume value of the audio sample, the volume value of each frequency in an audio sample can be comprehensively considered , So that the volume value of the obtained audio sample is more accurate. Adjusting the volume value of each audio sample according to the volume reference value includes: adjusting the volume value of each frequency in each audio sample according to the volume reference value. The volume value of each frequency in an audio sample is adjusted, so that the volume value of the entire audio sample is closer to the volume reference value, and the recognition effect of the speech recognition model is further improved.
[0092] The division of the steps of the various methods above is just for clarity of description. When implemented, it can be combined into one step or some steps can be split into multiple steps, as long as they include the same logical relationship, they are all in this patent. Within the scope of protection; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are within the scope of protection of the patent.
[0093] The fourth embodiment of the present invention relates to a volume adjustment device. The details of the volume adjustment device in this embodiment will be described in detail below. The following content is only provided for ease of understanding and is not necessary for implementing this solution. FIG. 6 is a schematic diagram of the volume adjustment device in the fourth embodiment. include:
[0094] The acquisition module 401 is used to acquire each audio sample in a training set used for training a speech recognition model; wherein, the speech recognition model is used for speech recognition;
[0095] The calculation module 402 is used to determine the volume value of each audio sample in the training set;
[0096] The statistics module 403 is used to determine the volume reference value of the training set according to the volume value of each audio sample;
[0097] The adjustment module 404 is used to adjust the volume value of each audio sample according to the volume reference value; wherein the difference between the volume value of each audio sample after adjustment and the volume reference value is within a preset difference range .
[0098] It is not difficult to find that this embodiment is an example of a device corresponding to the first to third embodiments, and this embodiment can be implemented in cooperation with the first to third embodiments. The related technical details and technical effects mentioned in the first to third embodiments are still valid in this embodiment, and in order to reduce repetition, they will not be repeated here. Correspondingly, the related technical details mentioned in this embodiment can also be applied to the first to third embodiments.
[0099] It is worth mentioning that the modules involved in this embodiment are all logical modules. In practical applications, a logical unit can be a physical unit, or a part of a physical unit, or The combination of multiple physical units is realized. In addition, in order to highlight the innovative part of the present invention, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by the present invention, but this does not mean that there are no other units in this embodiment.
[0100] The fifth embodiment of the present invention relates to an electronic device, as shown in FIG. 7, comprising: at least one processor 501; and a memory 502 communicatively connected with the at least one processor 501; wherein, the memory 502 stores instructions that can be executed by the at least one processor 501, and the instructions are executed by the at least one processor 501, so that the at least one processor 501 can execute the volume adjustment methods in the foregoing embodiments .
[0101] The memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects one or more processors and various circuits of the memory together. The bus can also connect various other circuits such as peripherals, voltage regulators, power management circuits, etc., all of which are well known in the art
Therefore, this article will not further describe it. The bus interface provides an interface between the bus and the transceiver. The transceiver may be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna, and further, the antenna also receives the data and transmits the data to the processor.
[0102] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when performing operations.
[0103] The sixth embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by the processor, the above method embodiment is realized.
[0104] That is, those skilled in the art can understand that all or part of the steps in the method of the foregoing embodiments can be implemented by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable A device (may be a single-chip microcomputer, a chip, etc.) or a processor (processor) executes all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical disks and other media that can store program codes. .
[0105] A person of ordinary skill in the art can understand that the above-mentioned embodiments are specific examples for realizing the present invention, but in practical applications, various changes can be made in form and details without departing from the present invention. Spirit and scope.
1 sheet
Sheet 1
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN108536420A | Cites | China | A | Search report | 1-10 |
| CN111462761A | Cites | China | A | Search report | 1-10 |
| US2005282590A1 | Cites | United States of America | A | Search report | 1-10 |
| US2020042285A1 | Cites | United States of America | A | Search report | 1-10 |
| CN206558213U | Cites | China | A | Search report | 1-10 |
1 member in 1 office
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 202010886561 | China | A | |
| CN20201886561 | – | – | – |
Members1
| Document | Office | Kind | |
|---|---|---|---|
| CN112037771AThis record | China | A |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent grantGrantedGR01 | GR01 | |
| Entry into force of request for substantive examinationSE01 | SE01 | |
| Entry into force of request for substantive examinationSE01 | SE01 | |
| PublicationPB01 | PB01 | |
| PublicationPB01 | PB01 |
Numbers
- Publication
- 112037771
- Publication, DOCDB
- 112037771
- Publication, EPODOC
- CN112037771
- Application
- 108865611
- Application, DOCDB
- 202010886561
- Application, EPODOC
- CN202010886561
Titles2
- Chinese
- 音量调节的方法、装置、电子设备和存储介质
- English
- Volume adjustment method, device, electronic equipment and storage medium
Classification
- CPC, 5
- G10L15/063
- G10L15/02
- G10L21/003
- G10L21/0208
- G10L2021/02082
- IPC, 4
- G10L15 06
- G10L15 02
- G10L21 003
- G10L21 0208