Audio data processing method and apparatus, and device and storage medium
Abstract
The disclosure provides an audio data processing method, device, device, and storage medium, and relates to the technical field of audio processing, in particular to the technical field of speech synthesis. The specific implementation plan is: decompose the original audio data to obtain human voice audio data and background audio data; perform electronic processing on human voice audio data to obtain electronic human voice data; and synthesize electronic human voice data and background audio data , to get the target audio data.

Term
14.9 yearsleft in the term
Expires 24 August 2041.
- Priority and filed
- Granted
- Today
- Expires
16 claims: 8 independent, 8 dependent
- 1一种音频数据处理方法,包括: 分解原始音频数据,得到人声音频数据和背景音频数据; 对所述人声音频数据进行电音化处理,得到电音人声数据;以及将所述电音人声数据和所述背景音频数据合成,得到目标音频数据; 其中,所述对所述人声音频数据进行电音化处理,得到电音人声数据,包括: 提取所述人声音频数据的原始基频; 对所述原始基频进行修正,得到第一基频; 根据预定电音参数,调整所述第一基频,得到第二基频; 针对所述第二基频进行量化处理,得到第三基频;以及根据所述第三基频,确定所述电音人声数据; 其中,所述对所述原始基频进行修正,得到第一基频,包括: 将所述人声音频数据分为多个音频片段; 针对所述多个音频片段中的每个音频片段,确定所述音频片段的能量和过零率; 根据所述能量和过零率,确定所述音频片段是否为浊音音频片段;以及利用线性插值算法,对所述浊音音频片段的基频进行修正; 其中,所述浊音音频片段的基频是根据算法从所述浊音音频片段中提取得到的,所述浊音音频片段的基频包括多个,多个所述浊音音频片段的基频中包括基频值为0的基频,所述利用线性插值算法,对所述浊音音频片段的基频进行修正包括: 利用根据与所述基频值为0的基频相邻的进行线性插值处理,得到修正后的基频,所述修正后的基频的基频值不等于0。
- 2根据权利要求1所述的方法,其中,所述分解原始音频数据,得到背景音频数据和人声音频数据,包括: 确定与所述原始音频数据对应的原始梅尔频谱数据; 使用神经网络确定与所述原始梅尔频谱数据对应的背景梅尔频谱数据和人声梅尔频谱数据;以及根据所述背景梅尔频谱数据,生成所述背景音频数据,并根据所述人声梅尔频谱数据, 生成所述人声音频数据。
- 3根据权利要求1所述的方法,其中,所述音频片段设置有多个采样点;所述确定所述音频片段的能量包括: 根据所述音频片段中每个采样点的数值,确定所述音频片段的能量。
- 4根据权利要求1所述的方法,其中,所述音频片段包括多个采样点;所述确定所述音频片段的过零率包括: 确定所述音频片段中每两个相邻采样点的数值是否彼此符号相反;以及确定所述音频片段中相邻采样点为异号的次数占所有采样点个数的比值,作为所述过零率。
- 5根据权利要求1所述的方法,其中,所述预定电音参数包括电音程度参数和/或电音音调参数;所述根据预定电音参数,调整所述第一基频,得到第二基频,包括: 根据所述浊音音频片段的基频,确定基频方差和/或基频平均值; 根据所述电音程度参数和所述基频方差,确定修正基频方差,以及/或者,根据所述电 音音调参数和所述基频平均值,确定修正基频平均值;以及根据所述修正基频方差和/或修正基频平均值,调整所述第一基频,得到所述第二基频。
- 6根据权利要求1-5中任一项所述的方法,其中,所述针对所述第二基频进行量化处理,得到第三基频,包括: 根据以下公式确定频率范围: F0' scale = 1 + 12 * log 2 (—-~) 27.5 其中,所述scale为所述频率范围,所述F0/为所述第二基频; 基于所述频率范围,根据以下公式确定所述第三基频: ^scale-1^ F0 = 27,5 * 2( 12) 其中,所述F0”为所述第三基频。
- 7根据权利要求1-5中任一项所述的方法,还包括:根据所述人声音频数据和所述第一基频,确定频谱包络和非周期参数; 其中,所述根据所述第三基频,确定所述电音人声数据,包括: 根据所述第三基频、所述频谱包络和所述非周期参数,确定所述电音人声数据。
- 8一种音频数据处理装置,包括: 分解模块,用于分解原始音频数据,得到人声音频数据和背景音频数据; 电音处理模块,用于对所述人声音频数据进行电音化处理,得到电音人声数据;以及合成模块,用于将所述电音人声数据和所述背景音频数据合成,得到目标音频数据; 其中,所述电音处理模块包括: 提取子模块,用于提取所述人声音频数据的原始基频; 修正子模块,用于对所述原始基频进行修正,得到第一基频; 调整子模块,用于根据预定电音参数,调整所述第一基频,得到第二基频; 量化子模块,用于针对所述第二基频进行量化处理,得到第三基频;以及电音确定子模块,用于根据所述第三基频,确定所述电音人声数据; 其中,所述修正子模块包括: 分段单元,用于将所述人声音频数据分为多个音频片段; 能量确定单元,用于针对所述多个音频片段中的每个音频片段,确定所述音频片段的能量; 过零率确定单元,用于针对所述多个音频片段中的每个音频片段,确定所述音频片段的过零率; 浊音判断单元,用于根据所述能量和过零率,确定所述音频片段的类型是否为浊音音频片段;以及修正单元,用于利用线性插值算法,对所述浊音音频片段的基频进行修正;其中,所述浊音音频片段的基频是根据算法从所述浊音音频片段中提取得到的,所述浊音音频片段的基频包括多个,多个所述浊音音频片段的基频中包括基频值为0的基频,所述修正单元进一步被配置为: CN 113689837 Β 利用根据与所述基频值为0的基频相邻的进行线性插值处理,得到修正后的基频,所述修正后的基频的基频值不等于0。
- 9根据权利要求8所述的装置,其中,所述分解模块包括: 梅尔频谱确定子模块,用于确定与所述原始音频数据对应的原始梅尔频谱数据; 分解子模块,用于使用神经网络确定与所述原始梅尔频谱数据对应的背景梅尔频谱数据和人声梅尔频谱数据;以及生成子模块,用于根据所述背景梅尔频谱数据,生成所述背景音频数据,并根据所述人声梅尔频谱数据,生成所述人声音频数据。
- 10根据权利要求8所述的装置,其中,所述音频片段设置有多个采样点;所述能量确定单元还用于: 根据所述音频片段中每个采样点的数值,确定所述音频片段的能量。
- 11根据权利要求8所述的装置,其中,所述音频片段包括多个采样点;所述过零率确定单元还用于: 确定所述音频片段中每两个相邻采样点的数值是否彼此符号相反;以及确定所述音频片段中相邻采样点为异号的次数占所有采样点个数的比值,作为所述过零率。
- 12根据权利要求8所述的装置,其中,所述预定电音参数包括电音程度参数和/或电音音调参数;所述调整子模块包括: 第一确定单元,用于根据所述浊音音频片段的基频,确定基频方差和/或基频平均值; 第二确定单元,用于根据所述电音程度参数和所述基频方差,确定修正基频方差,以及/或者,根据所述电音程度参数和所述基频平均值,确定修正基频平均值;以及调整单元,用于根据所述修正基频方差和/或修正基频平均值,调整所述第一基频,得到所述第二基频。
- 13根据权利要求872中任一项所述的装置,其中,所述量化子模块包括: 频率范围确定单元,用于根据以下公式确定频率范围: FO' scale = 1 + 12 *20如份行) 其中,所述seale为所述频率范围,所述F0'为所述第二基频; 第三基频确定单元,用于基于所述频率范围,根据以下公式确定所述第三基频: ^scale-1^ FO = 27.5 * 2( 12 ) 其中,所述F0〃为所述第三基频。
- 14根据权利要求8-12中任一项所述的装置,还包括: 确定模块,用于根据所述人声音频数据和所述第一基频,确定频谱包络和非周期参数; 其中,所述电音确定子模块还用于: 根据所述第三基频、所述频谱包络和所述非周期参数,确定所述电音人声数据。
- 15一种电子设备,包括: 至少一个处理器;以及与所述至少一个处理器通信连接的存储器;其中, 所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行权利要求1-7中任一项所述的方法。
- 16一种存储有计算机指令的非瞬时计算机可读存储介质,其中,所述计算机指令用于使所述计算机执行根据权利要求1-7中任一项所述的方法。
Independent claims16
147 paragraphs in 1 section, as filed
Audio data processing method, device, equipment and storage medium technical field
[0001] The present disclosure relates to the technical field of audio processing, in particular to the technical field of speech synthesis.
Background technique
As a kind of sound filter, electronic sound effect can be used to adjust and beautify sound, and it is widely used in scenes such as karaoke works or small video works. High-quality electronic sound effects can improve the sound quality of the work. For application products, if high-quality electronic sound effects can be provided, product competitiveness can be enhanced, product gameplay can be enriched, and user interest can be increased.
Contents of the invention
[0003] The present disclosure provides an audio data processing method, device, device, storage medium and program product.
According to one aspect of the present disclosure, a kind of audio data processing method is provided, comprising: decomposing original audio data, obtain human voice audio data and background audio data; The human voice audio data is carried out electrophonic processing, obtains electronic vocal data; and synthesizing the electronic vocal data with the background audio data to obtain target audio data.
According to another aspect of the present disclosure, a kind of audio data processing device is provided, comprising: decomposition module, for decomposing original audio data, obtain human voice audio data and background audio data; Electronic sound processing module, for The human voice audio data is electronically processed to obtain electronic human voice data; and a synthesis module is used to synthesize the electronic human voice data and the background audio data to obtain target audio data.
Another aspect of the present disclosure provides an electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein, the memory stores information that can be used by the at least one processor An instruction executed by a processor, the instruction is executed by the at least one processor, so that the at least one processor can execute the method shown in the embodiments of the present disclosure.
[0007] According to another aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the method shown in the embodiments of the present disclosure. method.
[0008] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer programs/instructions, characterized in that, when the computer program/instructions are executed by a processor, the methods shown in the embodiments of the present disclosure are implemented. step.
[0009] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following.
Description of drawings
Accompanying drawing is used for better understanding this scheme, does not constitute the limitation of this disclosure. in:
Fig. 1 schematically shows the flow chart of the audio data processing method according to an embodiment of the present disclosure;
Fig. 2 schematically shows a flow chart of a method for decomposing original audio data according to an embodiment of the present disclosure;
Fig. 3 schematically shows the method flow chart that human voice audio data is carried out electrophonic processing according to an embodiment of the present disclosure;
Fig. 4 schematically shows the flowchart of the audio data processing method according to another embodiment of the present disclosure;
[0015] Fig. 5 schematically shows a block diagram of an audio data processing device according to an embodiment of the present disclosure; and
[0016] FIG. 6 schematically illustrates a block diagram of an example electronic device that may be used to implement embodiments of the present disclosure.
Detailed ways
[0017] The exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, which include various details of the embodiments of the present disclosure to facilitate understanding, and they should be considered as exemplary only. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the disclosure. Also, descriptions of well-known functions and constructions are omitted in the following description for clarity and conciseness.
[0018] The audio data processing method of the disclosed embodiment will be described below in conjunction with FIG. 1. It should be noted that in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of audio data and other data involved are all in compliance with relevant laws and regulations, and do not violate public order and good customs.
[0019] FIG. 1 is a flowchart of an audio data processing method according to an embodiment of the present disclosure.
[0020] As shown in Figure 1, the audio data processing method 100 includes at operation S110, decomposing the original audio data to obtain vocal audio data and background audio data.
[0021] In operation S120, the human voice audio data is electronically processed to obtain electronic voice data.
[0022] In operation S130, electronic voice data and background audio data are synthesized to obtain target audio data.
[0023] According to an embodiment of the present disclosure, the original audio data may include, for example, human voice information and background sound information, wherein the human voice may be singing, for example, and the background sound may be accompaniment music, for example. In this embodiment, for example, a sound source separation algorithm can be used to separate the vocal information and background information in the original audio data to obtain vocal audio data containing vocal information and background audio data containing background sound information.
According to the disclosed embodiment, by separating the human voice information in the original audio data from the background sound information, the human voice information is electronically converted, and then the electronically converted human voice information is synthesized with the background voice information, Realized the electrophonicization of audio data with background sound information and human voice information at the same time.
[0025] According to an embodiment of the present disclosure, a neural network can be used to implement a sound source separation algorithm to decompose the original audio data. The input of the neural network may be audio data with background sound information and human voice information, and the output of the neural network may be human voice audio data including human voice information and background audio data including background sound information.
[0026] According to an embodiment of the present disclosure, the music file and the human voice file can be obtained in advance, and the music file and the human voice file can be cut into segments of equal length to obtain a plurality of music segments X and a plurality of vocal segments Y. Each music segment X can be synthesized with a corresponding human voice segment Y to obtain an original audio data Z. Each original audio data Z is used as the input of the neural network, and the music segment X and the human voice segment Y corresponding to the original audio data Z are used as the expected output to train the neural network. In addition, in order to improve the training effect and speed up the network convergence, music clip X, vocal clip Y and original audio data Z can be preprocessed into Mel spectrum. Correspondingly, the output result of the neural network is also based on the Mel spectrum. Exemplarily, the output result in the form of the mel spectrum can be synthesized into corresponding original audio data through an algorithm such as a Griffin-Lim (Griffin-Lin) algorithm.
Based on this, below with reference to Fig. 2, in conjunction with specific embodiment, the method for decomposing original audio data shown above is further described. Those skilled in the art can understand that the following exemplary embodiments are only for understanding the present disclosure, and the present disclosure is not limited thereto.
[0028] FIG. 2 schematically shows a flowchart of a method for decomposing original audio data according to an embodiment of the present disclosure.
[0029] As shown in FIG. 2, the method 210 of decomposing the original audio data includes, in operation S211, determining the original Mel spectrum data corresponding to the original audio data.
Then, in operation S212, use neural network to determine the background mel spectrum corresponding to original mel spectrum data
CN 113689837 Β
data and vocal mel spectrum data.
[0031] According to an embodiment of the present disclosure, the background mel spectrum data may include background sound information in the original mel spectrum data, and the vocal mel spectrum data may include vocal information in the original mel spectrum data.
[0032] In operation S213, generate background audio data according to the background Mel spectrum data, and generate vocal audio data according to the vocal Mel spectrum data.
[0033] According to an embodiment of the present disclosure, the background audio data can be generated according to the background Mel spectrum data through algorithms such as the Griffin-Lim algorithm, and the vocal audio data can be generated according to the vocal Mel spectrum data.
[0034] According to an embodiment of the present disclosure, the electrophonic processing of human voice audio data can be realized by quantizing the fundamental frequency of the human voice data. For example, the fundamental frequency, spectral envelope and aperiodic parameters of vocal data can be determined. Among them, the fundamental frequency represents the vibration frequency of the vocal cords during pronunciation, which is reflected in the audio frequency as the pitch. Next, the fundamental frequency can be quantized, and the human voice data can be re-synthesized according to the quantized fundamental frequency, spectrum envelope and aperiodic parameters, so as to realize the electroacoustic processing of the human voice audio data. Wherein, the resynthesized human voice data is electronic vocal data, which includes human voice information with electronic sound effects.
Below with reference to Fig. 3, in conjunction with specific embodiment, the method that the human voice audio data shown above is carried out to the method for electrophonic processing is further described. Those skilled in the art can understand that the following exemplary embodiments are only for understanding the present disclosure, and the present disclosure is not limited thereto.
[0036] FIG. 3 schematically shows a flowchart of a method for performing electronic processing of vocal audio data according to an embodiment of the present disclosure.
[0037] As shown in FIG. 3, the method 320 for electrophonic processing of human voice audio data may include, in operation S321, extracting the original fundamental frequency of human voice audio data.
[0038] According to an embodiment of the present disclosure, for example, the original fundamental frequency can be extracted from human voice audio data according to algorithms such as DIO and Harvest.
[0039] In operation S322, the original fundamental frequency is corrected to obtain the first fundamental frequency.
[0040] According to an embodiment of the present disclosure, by modifying the fundamental frequency, the electronic sound effect can be improved. For example, in this embodiment, human voice audio data may be divided into multiple audio segments. Then, for each audio segment of the plurality of audio segments, the energy and zero-crossing rate of the audio segment are determined. Based on energy and zero-crossing rate, determine whether an audio segment is a voiced audio segment. Then use the linear interpolation algorithm to modify the fundamental frequency of voiced audio clips.
[0041] According to an embodiment of the present disclosure, the vocal audio data can be divided into a plurality of audio segments with a preset unit length, and the length of each audio segment is a preset unit length. Wherein, the preset unit length can be set according to actual needs. Exemplarily, in this embodiment, the preset unit length may be any value from 10ms to 401ns.
[0042] According to an embodiment of the present disclosure, each audio segment is provided with a plurality of sampling points. The energy of the audio clip can be determined according to the value of each sample point in the audio clip. For example, the energy of an audio clip can be calculated according to the following formula:
<img file="CN113689837B_D0001.tif" />
Wherein, Xj represents the numerical value of the i sampling point, and n is the quantity of sampling point.
[0045] According to an embodiment of the present disclosure, the number n of sampling points may be determined according to the length and sampling rate of the audio segment.
Taking the preset unit length of 10ms as an example, the number n of sampling points can be calculated according to the following formula:
n=10/1000*sr=0.01sr
Wherein, sr represents the sampling rate of audio frequency.
CN 113689837 Β
[0048] According to an embodiment of the present disclosure, it may be determined whether the values of every two adjacent sampling points in the audio segment are opposite in sign to each other. Then determine the ratio of the number of adjacent sampling points with different signs to the number of all sampling points in the audio clip as the zero-crossing rate.
According to an embodiment of the present disclosure, the zero-crossing rate of the audio clip can be calculated according to the following formula: "<sub>Ί</sub> £ren; (yang*yang-1 < 0)
ZCR=n
Wherein, ZCR is the zero-crossing rate of audio clip, and n is the quantity of sampling point in audio clip, and Xj represents the numerical value of the ith sampling point in audio clip, and xai represents the i-1th sampling point in audio clip value.
[0052] According to an embodiment of the present disclosure, the number n of sampling points may be determined according to the length and sampling rate of the audio segment. Taking the preset unit length of 10ms as an example, the number n of sampling points can be calculated according to the following formula:
n=10/1000*sr=0.01sr
Wherein, sr represents the sampling rate of audio frequency.
[0055] When the human body is pronouncing, for the phonation of unvoiced sounds, the vocal cords do not vibrate, so the corresponding fundamental frequency is 0. For voiced sounds, the vocal cords vibrate, and the corresponding fundamental frequency is not 0. Based on this, in this embodiment, the fundamental frequency can be corrected by using the above characteristics.
For example, for each audio clip, if the energy E of this audio clip is less than threshold value e_min, and the zero-crossing rate ZCR of this audio clip is greater than threshold value zcr_max, then this audio clip is an unvoiced audio clip, and its fundamental frequency is 0. Otherwise, the audio clip is a voiced audio clip with a non-zero fundamental frequency. Among them, e_min and zcr_max can be set according to actual needs.
[0057] For each unvoiced audio clip, the base frequency of the audio clip can be set to 0. For voiced audio clips, the fundamental frequency of each voiced audio clip can be extracted according to algorithms such as DI0 and Harvest, and then it can be detected one by one whether the fundamental frequency value of each voiced audio clip is 0. For a voiced sound audio segment with a base frequency value of 0, a linear interpolation algorithm can be used to obtain a base frequency value that is not 0 as the base frequency of the voiced sound audio segment by using linear interpolation based on the value of the voiced sound audio segment near the voiced sound audio segment value.
For example, there are 6 voiced audio segments, and the fundamental frequency values are respectively: 100,100,0,0,160,100. That is, the fundamental frequency values of the third and fourth voiced audio segments are 0. Therefore, linear interpolation can be performed based on non-zero fundamental frequency values near the fundamental frequency values of the third and fourth voiced sound audio clips, that is, linear interpolation can be performed based on the second fundamental frequency value of 100 and the fifth fundamental frequency value of 160, The fundamental frequency values of the 3rd and 4th voiced sound audio clips are obtained as 120 and 140. That is, the corrected 6 fundamental frequency values are 100, 100, 120, 140, 160, 100.
[0059] Then, in operation S323, according to predetermined electronic sound parameters, adjust the first base frequency to obtain the second base frequency.
[0060] According to an embodiment of the present disclosure, the predetermined electronic sound parameters may include, for example, an electronic sound level parameter and/or an electronic sound tone parameter. Wherein, the electronic sound degree parameter can be used to control the degree of the electronic sound. The Electronic Tone parameter can be used to control the tone. Exemplarily, in this embodiment, the electronic sound level parameter may include, for example, 1, 1.2, and 1.4, and the greater the electronic sound level parameter, the more obvious the electronic sound effect. The electronic tone parameters may include -3, -2, -1, +1, +2, +3, for example. Among them, -1, -2, and -3 represent 1, 2, and 3 tone down respectively, and +1, +2, +3 represent 1, 2, and 3 tone up respectively.
[0061] In the related art, the electronic sound effect cannot adjust parameters, and the effect is single. According to the embodiment of the present disclosure, based on the characteristics of the electronic sound, two parameters, the degree parameter of the electronic sound and the pitch parameter of the electronic sound, are set to control the effect of the electronic sound, which can meet different user demands.
[0062] According to an embodiment of the present disclosure, the fundamental frequency variance and/or the fundamental frequency average may be determined according to the fundamental frequencies of all voiced audio segments. Determine the modified fundamental frequency variance according to the electronic sound degree parameter and the fundamental frequency variance, and/or determine the modified fundamental frequency average value according to the electronic sound degree parameter and the fundamental frequency average value. Then according to the modified fundamental frequency variance and/or the corrected fundamental frequency mean value,
The first fundamental frequency is adjusted to obtain the second fundamental frequency.
[0063] Exemplarily, in this embodiment, the variance of the fundamental frequency of all voiced sound audio segments can be calculated as the fundamental frequency variance, and the average value of the fundamental frequency of all voiced sound audio segments can be calculated as the fundamental frequency average value.
Then, can calculate and revise fundamental frequency variance according to following formula:
[0065] new_var=
Wherein, new_var is to revise fundamental frequency variance, and var is fundamental frequency variance, and a is also sound degree parameter.
Can calculate and revise fundamental frequency mean value according to following formula:
[0068] new_mean=mean*2<sup>b/12</sup>
Wherein, new_mean is to revise fundamental frequency mean value, and mean is fundamental frequency mean value, and b is the electronic tone tone parameter.
Next, the second fundamental frequency can be calculated according to the following formula:
[0071] FO =------* new_var+new_mean var - [0072] Wherein, F0' is the second fundamental frequency.
[0073] In operation S324, quantize the second base frequency to obtain a third base frequency.
[0074] In natural audio, the tone of the sound is cadenced and changes gradually, while the electronic tone quantizes the tone to a specific scale, so that the tone changes discontinuously, similar to the tone sent by an electronic musical instrument. Based on this, according to the embodiments of the present disclosure, each key frequency of the piano can be used as the target frequency to quantify the fundamental frequency of the vocal data.
Exemplarily, in the present embodiment, frequency range can be determined according to the following formula:
F0,
[0076] scale=1+12*log<sub>2</sub>(—)
Wherein, scale is frequency range, and F0 ' is the second fundamental frequency;
Then, based on the frequency range, the third fundamental frequency can be determined according to the following formula: ""rci.scale-i.
F0=27.5*
Wherein, F0 " is the 3rd fundamental frequency.
[0081] In operation S325, according to the third fundamental frequency, determine the electronic vocal data.
[0082] According to an embodiment of the present disclosure, the spectrum envelope and the aperiodic parameters may be determined according to the vocal audio data and the first fundamental frequency. Then the electronic vocal data can be determined according to the third fundamental frequency, spectrum envelope and aperiodic parameters.
[0083] Referring to Fig. 4 below, the audio data processing method shown above will be further described in conjunction with specific embodiments. Those skilled in the art can understand that the following exemplary embodiments are only for understanding the present disclosure, and the present disclosure is not limited thereto.
[0084] FIG. 4 schematically shows a flowchart of an audio data processing method according to another embodiment of the present disclosure.
As shown in Figure 4, this audio data processing method 400 comprises in operation S401, whether comprises accompaniment music (accompaniment for short) in judging audio data (abbreviation audio). If accompaniment is included, then execute operation S402. If only human voice is included but no accompaniment is included, perform operation S403.
[0086] In operation S402, a sound source separation algorithm is used to separate the human voice from the accompaniment. Then perform operation S403 for the separated human voice.
[0087] In operation S403, extract the zero-crossing rate, the fundamental frequency f 0 and the energy of the human voice.
In operation S404, based on zero-crossing rate and energy, basic frequency is revised and obtained F0 0
[0089] In operation S405, the spectrum envelope SP and the aperiodic parameter Ap are calculated using the human voice and the corrected fundamental frequency F0.
In operation S406, according to the electric sound degree parameter a and the electric sound tone parameter b of user's setting, adjust fundamental frequency and obtain
Get F0'.
[0091] In operation S407, the fundamental frequency F0' is quantized to obtain F0".
[0092] In operation S408, the base frequency F0 ", the spectrum envelope SP and the aperiodic parameter AP are used to jointly synthesize a human voice with an electronic sound effect.
[0093] In operation S409, if the audio has an accompaniment, then perform operation S410. Otherwise, perform operation S411.
[0094] In operation S410, the accompaniment is also combined into the human voice to generate the final audio with electronic sound effects.
[0095] In operation S411, audio with an electronic sound effect is output.
[0096] According to the audio data processing method of the embodiment of the present disclosure, it is possible to flexibly and efficiently add electronic sound effects to the audio data, so as to enhance the user's entertainment interest.
[0097] FIG. 5 schematically shows a block diagram of an audio data processing device according to an embodiment of the present disclosure.
[0098] As shown in FIG. 5, the audio data processing device 500 includes a decomposition module 510, an electronic sound processing module 520 and a synthesis module 530.
[0099] Decomposition module 510 is used to decompose the original audio data to obtain vocal audio data and background audio data.
[0100] The electronic sound processing module 520 is used to process the human voice audio data electronically to obtain the electronic sound human voice data. [0101] Synthesizing module 530, used for synthesizing electronic voice vocal data and background audio data to obtain target audio data.
[0102] According to an embodiment of the present disclosure, the decomposition module may include a mel spectrum determination submodule, a decomposition submodule and a generation submodule. Wherein, the mel spectrum determination submodule can be used to determine the original mel spectrum data corresponding to the original audio data. The decomposition sub-module can be used to determine background mel spectrum data and vocal mel spectrum data corresponding to the original mel spectrum data by using a neural network. The generation sub-module can be used to generate background audio data according to the background Mel spectrum data, and generate human voice audio data according to the vocal Mel spectrum data.
[0103] According to an embodiment of the present disclosure, the electronic sound processing module may include an extraction sub-module, a correction sub-module, an adjustment sub-module, a quantization sub-module and an electronic sound determination sub-module. Among them, the extraction sub-module can be used to extract the original fundamental frequency of the human voice audio data. The correction sub-module can be used to correct the original fundamental frequency to obtain the first fundamental frequency. The adjustment sub-module can be used to adjust the first fundamental frequency according to predetermined electronic sound parameters to obtain the second fundamental frequency. The quantization sub-module can be used to perform quantization processing on the second fundamental frequency to obtain the third fundamental frequency. The electronic sound determination sub-module can be used to determine electronic sound human voice data according to the third fundamental frequency.
[0104] According to an embodiment of the present disclosure, the correction submodule may include: a segmentation unit, an energy determination unit, a zero-crossing rate determination unit, a voiced sound judgment unit, and a correction unit. Wherein, the segmentation unit can be used to divide the human voice audio data into multiple audio segments. The energy determination unit may be configured to determine the energy of the audio segment for each audio segment in the plurality of audio segments. The zero-crossing rate determination unit can be used for determining the zero-crossing rate of the audio segment for each audio segment in the plurality of audio segments. The voiced sound judging unit can be used to determine whether the type of the audio segment is a voiced sound audio segment according to the energy and the zero-crossing rate. The correction unit can be used to correct the fundamental frequency of the voiced sound audio segment by using a linear interpolation algorithm.
[0105] According to an embodiment of the present disclosure, an audio segment is provided with a plurality of sampling points. The energy determination unit can also be used to determine the energy of the audio segment according to the value of each sampling point in the audio segment.
[0106] According to an embodiment of the present disclosure, the zero-crossing rate determination unit can also be used to determine whether the values of every two adjacent sampling points in the audio segment are opposite to each other, and then determine that the adjacent sampling points in the audio segment are of different signs The ratio of the number of times to the number of all sampling points is used as the zero-crossing rate.
[0107] According to an embodiment of the present disclosure, the predetermined electronic sound parameters may include electronic sound level parameters and/or electronic sound tone parameters. The adjustment submodule may include a first determination unit, a second determination unit and an adjustment unit. Wherein, the first determination unit may be configured to determine the variance of the fundamental frequency and/or the average value of the fundamental frequency according to the fundamental frequency of the voiced audio segment. The second determination unit can be used for
Determine the modified fundamental frequency variance according to the electronic sound degree parameter and the fundamental frequency variance, and/or determine the modified fundamental frequency average value according to the electronic sound degree parameter and the fundamental frequency average value. The adjustment unit can be used to adjust the first fundamental frequency according to the modified fundamental frequency variance and/or the corrected fundamental frequency average value to obtain the second fundamental frequency.
[0108] According to an embodiment of the present disclosure, the quantization submodule may include a frequency range determination unit and a third fundamental frequency determination unit
JL ο
Wherein, frequency range determination unit can be used for determining frequency range according to the following formula: fO/
[oho] scale = 1 + 12 * log<sub>2</sub>{——)
27*5
Wherein, scale is a frequency range, and F0' is the second fundamental frequency.
The 3rd fundamental frequency determines unit, can be used for based on frequency range, determines the 3rd fundamental frequency according to following formula:
[0113] F0=27,5*
Wherein, F0 " is the 3rd fundamental frequency.
[0115] According to an embodiment of the present disclosure, the above-mentioned audio data processing device may further include a determination module, which may be used to determine the spectrum envelope and aperiodic parameters according to the vocal audio data and the first fundamental frequency.
[0116] According to an embodiment of the present disclosure, the electronic sound determining submodule can also be used to determine the electronic sound human voice data according to the third fundamental frequency, spectrum envelope and aperiodic parameters.
[0117] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0118] FIG. 6 schematically illustrates a block diagram of an example electronic device 600 that may be used to implement embodiments of the present disclosure. Electronic device is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processing, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are by way of example only, and are not intended to limit implementations of the disclosure described and/or claimed herein.
[0119] As shown in FIG. 6 , the device 600 includes a computing unit 601 that can be programmed according to a computer program stored in a read-only memory (ROM) 602 or loaded from a storage unit 608 into a random access memory (RAM) 603. , to perform various appropriate actions and processing. In the RAM 603, various programs and data necessary for the operation of the device 600 can also be stored. The computing unit 601 , ROM 602 and RAM 603 are connected to each other through a bus 604 . An input/output (1/0) interface 605 is also connected to the bus 604 .
Multiple components in the device 600 are connected to the I/O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk , an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 609 allows the device 600 to exchange information/data with other devices over a computer network such as the Internet and/or various telecommunication networks.
[0121] The computing unit 601 may be various general-purpose and/or special-purpose processing components having processing and computing capabilities. Some examples of computing units 601 include, but are not limited to, central processing units (CPUs), graphics processing units (GPUs), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processing processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the audio data processing method. For example, in some embodiments, the audio data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium , such as the storage unit 608 . In some embodiments, the computer program's
Part or all may be loaded and/or installed onto the device 600 via the ROM 602 and/or the communication unit 609 . When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the audio data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the audio data processing method in any other suitable manner (for example, by means of firmware).
[0122] Various implementations of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), Implemented in a System-on-Chip (SOC), Load Programmable Logic Device (CPLD), computer hardware, firmware, software, and/or combinations thereof. These various embodiments may include being implemented in one or more computer programs executable and/or interpreted on a programmable system including at least one programmable processor, the programmable processor Can be special-purpose or general-purpose programmable processor, can receive data and instruction from storage system, at least one input device, and at least one output device, and transmit data and instruction to this storage system, this at least one input device, and this at least one output device an output device.
[0123] Program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special purpose computer, or other programmable data processing devices, so that the program codes, when executed by the processor or controller, make the functions/functions specified in the flow diagrams and/or block diagrams Action is implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0124] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include one or more wire-based electrical connections, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or flash memory), fiber optics, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0125] To provide for interaction with the user, the systems and techniques described herein may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) for displaying information to the user. ) monitor); and a keyboard and pointing device (eg, a mouse or trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and may be in any form (including Acoustic input, voice input, or tactile input) to receive input from the user.
[0126] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or including such backend components, intermediary components, or any combination of front-end components in a computing system. The components of the system can be interconnected by any form or medium of digital data communication (eg, a communication network). Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.
[0127] A computer system may include a client and a server. Clients and servers are generally remote from each other and usually
interact through a communication network. The relationship of client and server arises by computer programs running on the respective computers and having a client-server relationship to each other.
[0128] It should be understood that steps may be reordered, added or deleted using the various forms of the flow shown above. For example, each step described in the present disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in the present disclosure can be achieved, no limitation is imposed herein.
[0129] The specific embodiments described above do not constitute a limitation to the protection scope of the present disclosure. It should be apparent to those skilled in the art that various modifications, combinations, sub-combinations and substitutions may be made depending on design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Attached picture
100 Decompose the original audio data to obtain vocal audio data and background audio data. Perform electronic processing on the vocal audio data to obtain electronic vocal data. Synthesize electronic vocal data and background audio data to obtain target audio data TS130
<img file="CN113689837B_D0002.tif" />
figure 2
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| JPH04340600A | Cites | Japan | A | Search report | 1-16 |
| CN114360587A | Cites | China | – | Search report | 1-16 |
| KR19980029993A | Cites | Republic of Korea | A | Search report | 1-16 |
| US2012053933A1 | Cites | United States of America | A | Search report | 1-16 |
| US6078880A | Cites | United States of America | A | Search report | 1-16 |
| US6115684A | Cites | United States of America | A | Search report | 1-16 |
| CN110706679A | Cites | China | X | Search report | 1,10,19-20 |
| CN111243619A | Cites | China | Y | Search report | 2,11 |
| CN113178183A | Cites | China | Y | Search report | 3-9,12-18 |
| CN103440862A | Cites | China | A | Search report | 1-21 |
| CN112086085A | Cites | China | A | Search report | 1-21 |
| CN111724757A | Cites | China | A | Search report | 1-21 |
| CN109166593A | Cites | China | A | Search report | 1-21 |
| CN109346109A | Cites | China | A | Search report | 1-21 |
| Speech analysis and synthesis by linear prediction of the speech wave;BS Atal,等;《 The journal of the acoustical society of America》;全文 | Non-patent | – | – | Search report | – |
8 members in 5 offices
Members8
| Document | Office | Kind | |
|---|---|---|---|
| CN113689837A | China | A | |
| WO2023024501A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP4167226A1 | European Patent Office (EPO) | A1 | |
| CN113689837BThis record | China | B | |
| JP2023542760A | Japan | A | |
| JP7465992B2 | Japan | B2 | |
| US2024212703A1 | United States of America | A1 | |
| EP4167226A4 | European Patent Office (EPO) | A4 |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent grantGrantedGR01 | GR01 | |
| Entry into force of request for substantive examinationSE01 | SE01 | |
| PublicationPB01 | PB01 |
Numbers
- Publication
- 113689837
- Application
- 109780653
Titles2
- Chinese
- 音频数据处理方法、装置、设备以及存储介质
- English
- Audio data processing method, device, device and storage medium
Classification
- CPC, 21
- G10H1/02
- G10L21/0272
- G10L21/013
- G10L21/0308
- G10L13/047
- G10L25/24
- G10L25/30
- G10L25/09
- G10L25/93
- G10H2210/021
- G10H2210/005
- G10H2210/155
- G10H2210/041
- G10H2250/455
- G10H1/366
- G10H2250/311
- G10H2210/331
- Y02D30/70
- G10L13/02
- G10L19/032
- G10L21/0232
- IPC, 8
- G10H1 02
- G10L21 0272
- G10L21 0308
- G10L13 047
- G10L25 24
- G10L25 30
- G10L25 09
- G10L25 93