Audio data processing method and device, equipment and storage medium
Abstract
The present disclosure provides an audio data processing method, device, equipment, and storage medium, and relates to the technical field of audio processing, and in particular to the technical field of speech synthesis. The specific implementation scheme is: decompose the original audio data to obtain human voice audio data and background audio data; perform electro-sound processing on the human voice audio data to obtain electronic voice data; and synthesize the electronic voice data and background audio data , Get the target audio data.

Term
14.9 yearsto projected expiry
Projected expiry 24 August 2041, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
21 claims: 9 independent, 12 dependent
- 1一种音频数据处理方法,包括: 分解原始音频数据,得到人声音频数据和背景音频数据; 对所述人声音频数据进行电音化处理,得到电音人声数据;以及 将所述电音人声数据和所述背景音频数据合成,得到目标音频数据。
- 2根据权利要求1所述的方法,其中,所述分解原始音频数据,得到背景音频数据和人 声音频数据,包括: 确定与所述原始音频数据对应的原始梅尔频谱数据; 使用神经网络确定与所述原始梅尔频谱数据对应的背景梅尔频谱数据和人声梅尔频 谱数据;以及 根据所述背景梅尔频谱数据,生成所述背景音频数据,并根据所述人声梅尔频谱数据, 生成所述人声音频数据。
- 3根据权利要求1所述的方法,其中,所述对所述人声音频数据进行电音化处理,得到 电音人声数据,包括: 提取所述人声音频数据的原始基频; 对所述原始基频进行修正,得到第一基频; 根据预定电音参数,调整所述第一基频,得到第二基频; 针对所述第二基频进行量化处理,得到第三基频;以及 根据所述第三基频,确定所述电音人声数据。
- 4根据权利要求3所述的方法,其中,所述对所述原始基频进行修正,得到第一基频,包 括: 将所述人声音频数据分为多个音频片段; 针对所述多个音频片段中的每个音频片段,确定所述音频片段的能量和过零率; 根据所述能量和过零率,确定所述音频片段是否为浊音音频片段;以及 利用线性插值算法,对所述浊音音频片段的基频进行修正。
- 5根据权利要求4所述的方法,其中,所述音频片段设置有多个采样点;所述确定所述 音频片段的能量包括: 根据所述音频片段中每个采样点的数值,确定所述音频片段的能量。
- 6根据权利要求4所述的方法,其中,所述音频片段包括多个采样点;所述确定所述音 频片段的过零率包括: 确定所述音频片段中每两个相邻采样点的数值是否彼此符号相反;以及 确定所述音频片段中相邻采样点为异号的次数占所有采样点个数的比值,作为所述过 零率。
- 7根据权利要求4所述的方法,其中,所述预定电音参数包括电音程度参数和/或电音 音调参数;所述根据预定电音参数,调整所述第一基频,得到第二基频,包括: 根据所述浊音音频片段的基频,确定基频方差和/或基频平均值; 根据所述电音程度参数和所述基频方差,确定修正基频方差,以及/或者,根据所述电 音音调参数和所述基频平均值,确定修正基频平均值;以及 根据所述修正基频方差和/或修正基频平均值,调整所述第一基频,得到所述第二基 频。
- 8根据权利要求3-7中任一项所述的方法,其中,所述针对所述第二基频进行量化处 理,得到第三基频,包括: 根据以下公式确定频率范围: F0' scale = 1 + 12 * log 2 (——) 乙/·0 其中,所述scale为所述频率范围,所述F0'为所述第二基频; 基于所述频率范围,根据以下公式确定所述第三基频: ^scale—1^ F0 = 27.5 * ) 其中,所述F0〃为所述第三基频。
- 9根据权利要求3-7中任一项所述的方法,还包括:根据所述人声音频数据和所述第一 基频,确定频谱包络和非周期参数; 其中,所述根据所述第三基频,确定所述电音人声数据,包括: 根据所述第三基频、所述频谱包络和所述非周期参数,确定所述电音人声数据。
- 10一种音频数据处理装置,包括: 分解模块,用于分解原始音频数据,得到人声音频数据和背景音频数据; 电音处理模块,用于对所述人声音频数据进行电音化处理,得到电音人声数据;以及 合成模块,用于将所述电音人声数据和所述背景音频数据合成,得到目标音频数据。
- 11根据权利要求10所述的装置,其中,所述分解模块包括: 梅尔频谱确定子模块,用于确定与所述原始音频数据对应的原始梅尔频谱数据; 分解子模块,用于使用神经网络确定与所述原始梅尔频谱数据对应的背景梅尔频谱数 据和人声梅尔频谱数据;以及 生成子模块,用于根据所述背景梅尔频谱数据,生成所述背景音频数据,并根据所述人 声梅尔频谱数据,生成所述人声音频数据。
- 12根据权利要求10所述的装置,其中,所述电音处理模块包括: 提取子模块,用于提取所述人声音频数据的原始基频; 修正子模块,用于对所述原始基频进行修正,得到第一基频; 调整子模块,用于根据预定电音参数,调整所述第一基频,得到第二基频; 量化子模块,用于针对所述第二基频进行量化处理,得到第三基频;以及 电音确定子模块,用于根据所述第三基频,确定所述电音人声数据。
- 13根据权利要求12所述的装置,其中,所述修正子模块包括: 分段单元,用于将所述人声音频数据分为多个音频片段; 能量确定单元,用于针对所述多个音频片段中的每个音频片段,确定所述音频片段的 能量; 过零率确定单元,用于针对所述多个音频片段中的每个音频片段,确定所述音频片段 的过零率; 浊音判断单元,用于根据所述能量和过零率,确定所述音频片段的类型是否为浊音音 频片段;以及 修正单元,用于利用线性插值算法,对所述浊音音频片段的基频进行修正。
- 14根据权利要求13所述的装置,其中,所述音频片段设置有多个采样点;所述能量确 定单元还用于: 根据所述音频片段中每个采样点的数值,确定所述音频片段的能量。
- 15根据权利要求13所述的装置,其中,所述音频片段包括多个采样点;所述过零率确 定单元还用于: 确定所述音频片段中每两个相邻采样点的数值是否彼此符号相反;以及 确定所述音频片段中相邻采样点为异号的次数占所有采样点个数的比值,作为所述过 零率。
- 16根据权利要求13所述的装置,其中,所述预定电音参数包括电音程度参数和/或电 音音调参数;所述调整子模块包括: 第一确定单元,用于根据所述浊音音频片段的基频,确定基频方差和/或基频平均值; 第二确定单元,用于根据所述电音程度参数和所述基频方差,确定修正基频方差,以 及/或者,根据所述电音程度参数和所述基频平均值,确定修正基频平均值;以及 调整单元,用于根据所述修正基频方差和/或修正基频平均值,调整所述第一基频,得 到所述第二基频。
- 17根据权利要求1276中任一项所述的装置,其中,所述量化子模块包括: 频率范围确定单元,用于根据以下公式确定频率范围: scale = 1 + 12 * 其中,所述scale为所述频率范围,所述F0'为所述第二基频; 第三基频确定单元,用于基于所述频率范围,根据以下公式确定所述第三基频: ^scale—1^ F0 = 27.5 * 其中,所述F0〃为所述第三基频。
- 18根据权利要求1276中任一项所述的装置,还包括: 确定模块,用于根据所述人声音频数据和所述第一基频,确定频谱包络和非周期参数; 其中,所述电音确定子模块还用于: 根据所述第三基频、所述频谱包络和所述非周期参数,确定所述电音人声数据。
- 19一种电子设备,包括: 至少一个处理器;以及 与所述至少一个处理器通信连接的存储器;其中, 所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处 理器执行,以使所述至少一个处理器能够执行权利要求1-9中任一项所述的方法。
- 20一种存储有计算机指令的非瞬时计算机可读存储介质,其中,所述计算机指令用于 使所述计算机执行根据权利要求1-9中任一项所述的方法。
- 21一种计算机程序产品,包括计算机程序/指令,其特征在于,该计算机程序/指令被 处理器执行时实现权利要求1-9中任一项所述方法的步骤。
Independent claims21
141 paragraphs in 1 section, as filed
Audio data processing method, device, equipment and storage medium technical field
[0001] The present disclosure relates to the field of audio processing technology, and in particular to the field of speech synthesis technology.
Background technique
[0002] As a sound filter, the electronic sound effect can be used to adjust and beautify the sound, and has a wide range of applications in scenes such as karaoke works or small video works. High-quality electronic sound effects can improve the sound quality of the work. For application products, if you can provide high-quality electronic sound effects, you can enhance product competitiveness, enrich product gameplay, and increase user interest.
Summary of the invention
[0003] The present disclosure provides an audio data processing method, device, equipment, storage medium, and program product.
[0004] According to one aspect of the present disclosure, there is provided an audio data processing method, including: decomposing original audio data to obtain human voice audio data and background audio data; performing electro-sound processing on the human voice audio data to obtain Electronic voice human voice data; and synthesizing the electronic voice human voice data and the background audio data to obtain target audio data.
[0005] According to another aspect of the present disclosure, there is provided an audio data processing device, including: a decomposition module for decomposing original audio data to obtain human voice audio data and background audio data; The human voice audio data is electro-soundized to obtain electronic voice human voice data; and a synthesis module is used to synthesize the electronic voice human voice data and the background audio data to obtain target audio data.
[0006] Another aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected with the at least one processor; An instruction executed by the processor, the instruction being executed by the at least one processor, so that the at least one processor can execute the method shown in the embodiment of the present disclosure.
[0007] According to another aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to make the computer execute the method.
[0008] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program/instruction, characterized in that, when the computer program/instruction is executed by a processor, it implements step.
[0009] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood by the following.
Description of the drawings
[0010] The accompanying drawings are used to better understand the solution, and do not constitute a limitation of the present disclosure. in:
[0011] FIG. 1 schematically shows a flowchart of an audio data processing method according to an embodiment of the present disclosure;
[0012] FIG. 2 schematically shows a flowchart of a method for decomposing original audio data according to an embodiment of the present disclosure;
[0013] FIG. 3 schematically shows a flow chart of a method for electro-sound processing of human voice audio data according to an embodiment of the present disclosure;
[0014] FIG. 4 schematically shows a flowchart of an audio data processing method according to another embodiment of the present disclosure;
[0015] FIG. 5 schematically shows a block diagram of an audio data processing device according to an embodiment of the present disclosure; and
[0016] FIG. 6 schematically shows a block diagram of an example electronic device that may be used to implement embodiments of the present disclosure.
Detailed ways
[0017] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to facilitate understanding, and should be regarded as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Likewise, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] The audio data processing method of an embodiment of the present disclosure will be described below in conjunction with FIG. 1. It should be noted that in the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of audio data and other data involved are in compliance with relevant laws and regulations, and does not violate public order and good customs.
[0019] FIG. 1 is a flowchart of an audio data processing method according to an embodiment of the present disclosure.
[0020] As shown in FIG. 1, the audio data processing method 100 includes in operation S110, decomposing original audio data to obtain human voice audio data and background audio data.
[0021] In operation S120, electro-sound processing is performed on the human voice audio data to obtain electro-sound human voice data.
[0022] In operation S130, the electronic voice and human voice data and background audio data are synthesized to obtain target audio data.
[0023] According to an embodiment of the present disclosure, the original audio data may include, for example, human voice information and background sound information, where the human voice may be, for example, singing voice, and the background sound may be, for example, accompaniment music. In this embodiment, for example, a sound source separation algorithm may be used to separate human voice information and background information in the original audio data to obtain human voice audio data containing human voice information and background audio data containing background sound information.
[0024] According to the disclosed embodiment, by separating the human voice information and the background sound information in the original audio data, the human voice information is electro-soundized, and then the electro-soundized human voice information and the background sound information are synthesized, It realizes the electro-soundization of audio data with both background sound information and human voice information.
[0025] According to an embodiment of the present disclosure, a neural network may be used to implement a sound source separation algorithm to decompose the original audio data. The input of the neural network may be audio data with background sound information and human voice information, and the output of the neural network may include human voice data including human voice information and background audio data including background sound information.
[0026] According to the embodiments of the present disclosure, music files and vocal files can be obtained in advance, and the music files and vocal files can be cut into equal length segments to obtain multiple music fragments X and multiple vocal fragments Y. Each music segment X and a corresponding vocal segment Y can be synthesized separately to obtain an original audio data Z. Take each original audio data Z as the input of the neural network, and use the music segment X and the human voice segment Y corresponding to the original audio data Z as the expected output to train the neural network. In addition, in order to improve the training effect and speed up the network convergence, the music segment X, the vocal segment Y and the original audio data Z can all be preprocessed into the Mel spectrum. Correspondingly, the output result of the neural network is also based on the Mel spectrum. Exemplarily, the output result in the form of the Mel spectrum can be used to synthesize the corresponding original audio data through an algorithm such as the Griffin-Lim algorithm.
[0027] Based on this, the method for decomposing the original audio data shown above will be further described with reference to FIG. 2 in conjunction with specific embodiments. Those skilled in the art can understand that the following exemplary embodiments are only for understanding the present disclosure, and the present disclosure is not limited thereto.
[0028] FIG. 2 schematically shows a flowchart of a method for decomposing original audio data according to an embodiment of the present disclosure.
[0029] As shown in FIG. 2, the method 210 for decomposing original audio data includes in operation S211, determining original Mel spectrum data corresponding to the original audio data.
[0030] Then, in operation S212, a neural network is used to determine a background mel spectrum corresponding to the original mel spectrum data
Data and vocal mel spectrum data.
[0031] According to an embodiment of the present disclosure, the background mel spectrum data may include background sound information in the original mel spectrum data, and the vocal mel spectrum data may include the human voice information in the original mel spectrum data.
[0032] In operation S213, background audio data is generated based on the background Mel spectrum data, and human voice audio data is generated based on the human voice Mel spectrum data.
[0033] According to the embodiments of the present disclosure, the background audio data can be generated based on the background Mel spectrum data by algorithms such as the Griffin-Lim algorithm, and the human voice audio data can be generated based on the human voice Mel spectrum data.
[0034] According to the embodiments of the present disclosure, the electro-sound processing of the human voice audio data can be realized by quantizing the fundamental frequency of the human voice data. For example, the fundamental frequency, spectral envelope, and aperiodic parameters of human voice data can be determined. Among them, the fundamental frequency represents the vibration frequency of the vocal cords during pronunciation, which is reflected in the audio frequency as the pitch. Next, the fundamental frequency can be quantized, and the vocal data can be re-synthesized according to the quantized fundamental frequency, spectrum envelope, and non-periodic parameters, so as to realize the electro-phonic processing of the vocal audio data. Among them, the re-synthesized human voice data is the electronic voice human voice data, which contains human voice information with electronic sound effects.
[0035] In the following, referring to FIG. 3, the method of performing electro-sound processing on human voice audio data shown above will be further described in conjunction with specific embodiments. Those skilled in the art can understand that the following exemplary embodiments are only used to understand the present disclosure, and the present disclosure is not limited thereto.
[0036] FIG. 3 schematically shows a flowchart of a method for electro-sounding processing of human voice audio data according to an embodiment of the present disclosure.
[0037] As shown in FIG. 3, the method 320 for performing electro-sound processing on human voice audio data may include in operation S321, extracting the original fundamental frequency of the human voice audio data.
[0038] According to the embodiments of the present disclosure, for example, the original fundamental frequency can be extracted from the human voice audio data according to algorithms such as DI0 and Harvest.
[0039] In operation S322, the original fundamental frequency is modified to obtain the first fundamental frequency.
[0040] According to the embodiments of the present disclosure, by correcting the fundamental frequency, the electronic sound effect can be improved. For example, in this embodiment, the human voice audio data can be divided into multiple audio segments. Then, for each audio segment in the multiple audio segments, the energy and zero-crossing rate of the audio segment are determined. According to the energy and the zero-crossing rate, it is determined whether the audio segment is a voiced audio segment. Then use linear interpolation algorithm to correct the fundamental frequency of voiced audio clips.
[0041] According to an embodiment of the present disclosure, human voice audio data may be divided into a plurality of audio segments with a preset unit length, and the length of each audio segment is a preset unit length. Among them, the preset unit length can be set according to actual needs. Exemplarily, in this embodiment, the preset unit length may be any value from 10 ms to 401 ns.
[0042] According to an embodiment of the present disclosure, each audio segment is provided with a plurality of sampling points. The energy of the audio segment can be determined according to the value of each sampling point in the audio segment. For example, the energy of the audio segment can be calculated according to the following formula: [0043] E = MarinaΪΖΕ η
[0044] Wherein, Xj represents the value of the i-th sampling point, and n is the number of sampling points.
[0045] According to an embodiment of the present disclosure, the number n of sampling points may be determined according to the length and sampling rate of the audio segment. Taking the default unit length of 10ms as an example, the number of sampling points n can be calculated according to the following formula:
[0046] n=10/1000*sr=0. Olsr
[0047] Wherein, sr represents the sampling rate of the audio.
[0048] According to an embodiment of the present disclosure, it can be determined whether the values of every two adjacent sampling points in the audio segment have opposite signs to each other. Then determine the ratio of the number of adjacent sampling points with different signs in the audio clip to the number of all sampling points, as the zero-crossing rate.
[0049] According to an embodiment of the present disclosure, the zero-crossing rate of an audio segment can be calculated according to the following formula: ""see * sign 7 <0)
[0050] ZCR = B<sup>[</sup>-1 ' <sup>1</sup>_________2 η
[0051] Among them, ZCR is the zero-crossing rate of the audio segment, n is the number of sampling points in the audio segment, Xj represents the value of the i-th sampling point in the audio segment, and xai represents the value of the i-1th sampling point in the audio segment. Numerical value.
[0052] According to an embodiment of the present disclosure, the number n of sampling points may be determined according to the length and sampling rate of the audio segment. Taking the default unit length of 10ms as an example, the number of sampling points n can be calculated according to the following formula:
[0053] n=10/1000*sr=0. Olsr
[0054] Wherein, sr represents the sampling rate of the audio.
[0055] When the human body is speaking, the vocal cords do not vibrate for unvoiced sounds, so the corresponding fundamental frequency is 0. For voiced sounds, the vocal cords vibrate, and the corresponding fundamental frequency is not zero. Based on this, in this embodiment, the above characteristics can be used to correct the fundamental frequency.
[0056] For example, for each audio segment, if the energy E of the audio segment is less than the threshold e_min, and the zero-crossing rate ZCR of the audio segment is greater than the threshold zcr_max, then the audio segment is an unvoiced audio segment with a fundamental frequency of 0. Otherwise, the audio segment is a voiced audio segment, and its fundamental frequency is non-zero. Among them, e_min and zcr_max can be set according to actual needs.
[0057] For each unvoiced audio segment, the fundamental frequency of the audio segment can be set to 0. For voiced audio segments, the fundamental frequency of each voiced audio segment can be extracted according to algorithms such as DIO and Harvest, and then whether the fundamental frequency value of each voiced audio segment is 0 is detected one by one. For a voiced audio segment with a fundamental frequency value of 0, a linear interpolation algorithm can be used to perform linear interpolation based on the voiced audio segment value near the voiced audio segment to obtain a fundamental frequency value other than 0 as the fundamental frequency of the voiced audio segment value.
[0058] For example, there are 6 voiced audio clips, and the fundamental frequency values are: 100, 100, 0, 0, 160, and 100 respectively. That is, the fundamental frequency value of the 3rd and 4th voiced audio segments is 0. Therefore, linear interpolation can be performed based on the non-zero fundamental frequency values near the fundamental frequency values of the third and fourth voiced audio segments, that is, linear interpolation can be performed based on the second fundamental frequency value of 100 and the fifth fundamental frequency value of 160, The fundamental frequency values of the third and fourth voiced audio segments are 120 and 140. That is, the corrected 6 fundamental frequency values are 100, 100,
120,140,160,100ο
[0059] Then, in operation S323, the first fundamental frequency is adjusted according to the predetermined electronic tone parameters to obtain the second fundamental frequency.
[0060] According to an embodiment of the present disclosure, the predetermined electric sound parameters may include, for example, electric sound level parameters and/or electric sound pitch parameters. Among them, the degree of electronic sound parameter can be used to control the degree of electronic sound. The electronic tone tone parameter can be used to control the tone. Exemplarily, in this embodiment, the electric sound level parameter may include, for example, 1, 1.2, 1.4, and the larger the electric sound level parameter, the more obvious the electric sound effect. The electric tone tone parameters may include -3, -2, -1, +1, +2, +3, for example. Among them, -1, -2, and -3 indicate down 1, 2, and 3 tones respectively, and +1, +2, and +3 indicate up 1, 2, and 3 tones respectively.
[0061] In the related art, the parameters of the electronic sound effect cannot be adjusted, and the effect is single. According to the embodiment of the present disclosure, based on the characteristics of the electric sound, two parameters, the electric sound level parameter and the electric sound tone parameter, are set to control the electric sound effect, which can meet different user requirements.
[0062] According to an embodiment of the present disclosure, the fundamental frequency variance and/or the fundamental frequency average value can be determined according to the fundamental frequencies of all voiced audio segments. Determine the corrected fundamental frequency variance according to the electrical tone degree parameter and the fundamental frequency variance, and/or determine the corrected fundamental frequency average value based on the electrical tone degree parameter and the fundamental frequency average value. Then according to the corrected fundamental frequency variance and/or corrected fundamental frequency average value, adjust
Adjust the first fundamental frequency to obtain the second fundamental frequency.
[0063] Illustratively, in this embodiment, the variance of the fundamental frequencies of all voiced audio segments can be calculated as the fundamental frequency variance, and the average value of the fundamental frequencies of all voiced audio segments can be calculated as the fundamental frequency average value.
[0064] Then, the corrected fundamental frequency variance can be calculated according to the following formula:
[0065] new_var =
[0066] Among them, new_var is the modified fundamental frequency variance, vax is the fundamental frequency variance, and a is the pitch parameter.
[0067] The corrected fundamental frequency average value can be calculated according to the following formula:
[0068] new_mean=mean*2<sup>b/12</sup>
[0069] Among them, new_mean is the modified fundamental frequency average value, mean is the fundamental frequency average value, and b is the electric tone pitch parameter.
[0070] Next, the second fundamental frequency can be calculated according to the following formula:
[0071] FO =------* new_var + new_mean var-[0072] Among them, F0' is the second fundamental frequency.
[0073] In operation S324, quantization processing is performed on the second fundamental frequency to obtain the third fundamental frequency.
[0074] In natural audio, the tone of the sound is accented and frustrated, and changes gradually, while the electronic tone quantizes the tone to a specific scale, so that the tone changes discontinuously, similar to the tone emitted by an electronic musical instrument. Based on this, according to the embodiment of the present disclosure, the fundamental frequency of the human voice data can be quantized with each key frequency of the piano as the target frequency.
[0075] Exemplarily, in this embodiment, the frequency range may be determined according to the following formula:
FCP
[0076] scale = 1 + 12 * log<sub>2</sub>(—)
27+5
[0077] Wherein, scale is the frequency range, and F0' is the second fundamental frequency;
[0078] Then, based on the frequency range, the third fundamental frequency can be determined according to the following formula: ""rci, scale-i.
[0079] FO = 27.5 *
[0080] Wherein, F0 " is the third fundamental frequency.
[0081] In operation S325, the electronic voice vocal data is determined according to the third fundamental frequency.
[0082] According to an embodiment of the present disclosure, the spectral envelope and aperiodic parameters can be determined according to the human voice audio data and the first fundamental frequency. Then the electronic voice and human voice data can be determined according to the third fundamental frequency, spectrum envelope and non-periodic parameters.
[0083] The audio data processing method shown above will be further described below with reference to FIG. 4 in conjunction with specific embodiments. Those skilled in the art can understand that the following exemplary embodiments are only for understanding the present disclosure, and the present disclosure is not limited thereto.
[0084] FIG. 4 schematically shows a flowchart of an audio data processing method according to another embodiment of the present disclosure.
[0085] As shown in FIG. 4, the audio data processing method 400 includes in operation S401, determining whether the audio data (referred to as audio) contains accompaniment music (referred to as accompaniment). If accompaniment is included, operation S402 is performed. If only the human voice is included but no accompaniment is included, operation S403 is performed.
[0086] In operation S402, the human voice is separated from the accompaniment using a sound source separation algorithm. Then, operation S403 is performed on the separated human voice.
[0087] In operation S403, the zero-crossing rate, fundamental frequency f0, and energy are extracted from the human voice.
[0088] In operation S404, the fundamental frequency is corrected based on the zero-crossing rate and energy to obtain F0.
[0089] In operation S405, the spectrum envelope SP and the aperiodic parameter Ap are calculated using the human voice and the corrected fundamental frequency F0.
[0090] In operation S406, according to the electronic tone level parameter a and the electronic tone tone parameter b set by the user, the fundamental frequency is adjusted and obtained
Get F0'.
[0091] In operation S407, the fundamental frequency F0is quantized to obtain F0".
[0092] In operation S408, the fundamental frequency F0", the spectral envelope SP, and the aperiodic parameter AP are used to synthesize a human voice with an electronic sound effect.
[0093] In operation S409, if the audio is accompanied by accompaniment, operation S410 is performed. Otherwise, perform operation S411.
[0094] In operation S410, the accompaniment is also incorporated into the human voice to generate the final audio with electronic sound effects.
[0095] In operation S411, audio with electronic sound effects is output.
[0096] According to the audio data processing method of the embodiment of the present disclosure, it is possible to flexibly and efficiently add electro-sound effects to the audio data, and enhance the user's entertainment interest.
[0097] FIG. 5 schematically shows a block diagram of an audio data processing device according to an embodiment of the present disclosure.
[0098] As shown in FIG. 5, the audio data processing device 500 includes a decomposition module 510, an electronic sound processing module 520, and a synthesis module 530.
[0099] The decomposition module 510 is used to decompose the original audio data to obtain human voice audio data and background audio data.
[0100] The electro-sound processing module 520 is used to perform electro-sound processing on the human voice audio data to obtain the electro-sound human voice data. [0101] The synthesis module 530 is used to synthesize the electronic voice and human voice data and background audio data to obtain target audio data.
[0102] According to an embodiment of the present disclosure, the decomposition module may include a Mel spectrum determination sub-module, a decomposition sub-module, and a generation sub-module. Among them, the mel spectrum determination sub-module can be used to determine the original mel spectrum data corresponding to the original audio data. The decomposition sub-module can be used to determine the background Mel spectrum data and the human voice Mel spectrum data corresponding to the original Mel spectrum data using a neural network. The generation sub-module can be used to generate background audio data based on the background Mel spectrum data, and generate human voice audio data based on the human voice Mel spectrum data.
[0103] According to an embodiment of the present disclosure, the electronic sound processing module may include an extraction sub-module, a correction sub-module, an adjustment sub-module, a quantization sub-module, and an electronic sound determination sub-module. Among them, the extraction sub-module can be used to extract the original fundamental frequency of human voice audio data. The correction sub-module can be used to correct the original fundamental frequency to obtain the first fundamental frequency. The adjustment sub-module can be used to adjust the first fundamental frequency according to the predetermined electronic tone parameters to obtain the second fundamental frequency. The quantization sub-module can be used to perform quantization processing on the second fundamental frequency to obtain the third fundamental frequency. The electronic sound determination sub-module can be used to determine the electronic sound human voice data according to the third fundamental frequency.
[0104] According to an embodiment of the present disclosure, the correction sub-module may include: a segmentation unit, an energy determination unit, a zero-crossing rate determination unit, a voiced sound judgment unit, and a correction unit. Among them, the segmentation unit can be used to divide the human voice audio data into multiple audio segments. The energy determining unit may be used to determine the energy of the audio segment for each audio segment of the multiple audio segments. The zero-crossing rate determining unit may be used to determine the zero-crossing rate of the audio segment for each audio segment of the multiple audio segments. The voiced sound judging unit can be used to determine whether the type of the audio segment is a voiced audio segment according to the energy and the zero-crossing rate. The correction unit can be used to correct the fundamental frequency of voiced audio clips by using a linear interpolation algorithm.
[0105] According to an embodiment of the present disclosure, the audio segment is provided with a plurality of sampling points. The energy determining unit may also be used to determine the energy of the audio segment according to the value of each sampling point in the audio segment.
[0106] According to an embodiment of the present disclosure, the zero-crossing rate determination unit may also be used to determine whether the values of every two adjacent sampling points in the audio segment have opposite signs to each other, and then determine whether the adjacent sampling points in the audio segment have different signs. The ratio of the number of times to the number of all sampling points is used as the zero-crossing rate.
[0107] According to an embodiment of the present disclosure, the predetermined electronic sound parameter may include an electronic sound degree parameter and/or an electric sound tone parameter. The adjustment sub-module may include a first determination unit, a second determination unit, and an adjustment unit. The first determining unit may be used to determine the fundamental frequency variance and/or the fundamental frequency average value according to the fundamental frequency of the voiced audio segment. The second determination unit can be used for
Determine the corrected fundamental frequency variance according to the electrical tone degree parameter and the fundamental frequency variance, and/or determine the corrected fundamental frequency average value based on the electrical tone degree parameter and the fundamental frequency average value. The adjustment unit may be used to adjust the first fundamental frequency according to the corrected fundamental frequency variance and/or the corrected fundamental frequency average value to obtain the second fundamental frequency.
[0108] According to an embodiment of the present disclosure, the quantization sub-module may include a frequency range determination unit and a third fundamental frequency determination unit.
JL· ο
[0109] Wherein, the frequency range determining unit may be used to determine the frequency range according to the following formula:
[oho] scale = 1 + 12 * log<sub>2</sub>Q-)
27.5
[0111] Wherein, scale is the frequency range, and F0' is the second fundamental frequency.
[0112] The third fundamental frequency determining unit may be used to determine the third fundamental frequency based on the frequency range according to the following formula:
[0113] F0 = 27,5 *)
[0114] Wherein, F0 " is the third fundamental frequency.
[0115] According to an embodiment of the present disclosure, the above audio data processing device may further include a determining module, which may be used to determine the spectral envelope and aperiodic parameters according to the human voice audio data and the first fundamental frequency.
[0116] According to an embodiment of the present disclosure, the electronic sound determination sub-module may also be used to determine the electronic sound human voice data according to the third fundamental frequency, spectrum envelope, and non-periodic parameters.
[0117] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0118] FIG. 6 schematically shows a block diagram of an example electronic device 600 that may be used to implement embodiments of the present disclosure. Electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the present disclosure described and/or required herein.
[0119] As shown in FIG. 6, the device 600 includes a computing unit 601, which can be based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from the storage unit 608 into a random access memory (RAM) 603 , To perform various appropriate actions and processing. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The calculation unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input/output Q/0) interface 605 is also connected to the bus 604.
[0120] Multiple components in the device 600 are connected to the I/O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; and a storage unit 608, such as a disk , Optical disc, etc.; and communication unit 609, such as network card, modem, wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information/data with other devices through a computer network such as the Internet and/or various telecommunication networks.
[0121] The computing unit 601 may be various general-purpose and/or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 601 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processing DSP, and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as an audio data processing method. For example, in some embodiments, the audio data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium , For example, the storage unit 608. In some embodiments, the computer program
Part or all may be loaded and/or installed on the device 600 via the ROM 602 and/or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the calculation unit 601, one or more steps of the audio data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the audio data processing method in any other suitable manner (for example, by means of firmware).
[0122] The various implementations of the systems and technologies described herein above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application-specific standard products (ASSP), Implemented in a system on chip system (SOC), load programmable logic device (CPLD), computer hardware, firmware, software, and/or a combination thereof. These various embodiments may include: being implemented in one or more computer programs, the one or more computer programs may be executed and/or interpreted on a programmable system including at least one programmable processor, the programmable processor It can be a dedicated or general-purpose programmable processor that can receive data and instructions from the storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device. An output device.
[0123] The program code used to implement the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processors or controllers of general-purpose computers, special-purpose computers, or other programmable data processing devices, so that when the program codes are executed by the processor or controller, the functions specified in the flowcharts and/or block diagrams/ The operation is implemented. The program code can be executed entirely on the machine, partly executed on the machine, partly executed on the machine and partly executed on the remote machine as an independent software package, or entirely executed on the remote machine or server.
[0124] In the context of the present disclosure, a machine-readable medium may be a tangible medium, which may contain or store a program for use by an instruction execution system, apparatus, or device or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
[0125] In order to provide interaction with the user, the systems and techniques described herein can be implemented on a computer that has: a display device for displaying information to the user (for example, a CRT (cathode ray tube) or an LCD (liquid crystal display) ) Monitor); and a keyboard and pointing device (for example, a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (for example, visual feedback, auditory feedback, or tactile feedback); and can be in any form (including Acoustic input, voice input, or tactile input) to receive input from the user.
[0126] The systems and technologies described herein can be implemented in a computing system that includes back-end components (for example, as a data server), or a computing system that includes middleware components (for example, an application server), or a computing system that includes front-end components (For example, a user computer with a graphical user interface or a web browser through which the user can interact with the implementation of the system and technology described herein), or include such background components, intermediate Components, or any combination of front-end components in a computing system. The components of the system can be connected to each other through any form or medium of digital data communication (for example, a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0127] The computer system may include a client and a server. The client and server are generally far away from each other and usually communicate
Interaction through the communication network. The relationship between the client and the server is generated by computer programs that run on the corresponding computers and have a client-server relationship with each other.
[0128] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in the present disclosure can be achieved, this is not limited herein.
[0129] The foregoing specific implementations do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent replacement and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| WO2025092363A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| CN116312431A | Cited by | China | – | Search report | – |
| CN119922371A | Cited by | China | – | Search report | – |
| CN114449339A | Cited by | China | – | Search report | – |
| CN115054915A | Cited by | China | – | Search report | – |
| WO2023024501A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| CN103440862A | Cites | China | A | Search report | 1-21 |
| CN109166593A | Cites | China | A | Search report | 1-21 |
| CN109346109A | Cites | China | A | Search report | 1-21 |
| CN110706679A | Cites | China | YX | Search report | 2-9,11-18 |
| CN111243619A | Cites | China | Y | Search report | 2,11 |
| CN111724757A | Cites | China | A | Search report | 1-21 |
| CN112086085A | Cites | China | A | Search report | 1-21 |
| CN113178183A | Cites | China | Y | Search report | 3-9,12-18 |
| CN114360587A | Cites | China | – | Search report | 1-16 |
| KR19980029993A | Cites | Republic of Korea | A | Search report | 1-16 |
| US2012053933A1 | Cites | United States of America | A | Search report | 1-16 |
| US6078880A | Cites | United States of America | A | Search report | 1-16 |
| US6115684A | Cites | United States of America | A | Search report | 1-16 |
| JPH04340600A | Cites | Japan | A | Search report | 1-16 |
| BS ATAL,等: "Speech analysis and synthesis by linear prediction of the speech wave", 《 THE JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA》 | Non-patent | – | – | Search report | – |
8 members in 5 offices
Members8
| Document | Office | Kind | |
|---|---|---|---|
| CN113689837AThis record | China | A | |
| WO2023024501A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP4167226A1 | European Patent Office (EPO) | A1 | |
| CN113689837B | China | B | |
| JP2023542760A | Japan | A | |
| JP7465992B2 | Japan | B2 | |
| US2024212703A1 | United States of America | A1 | |
| EP4167226A4 | European Patent Office (EPO) | A4 |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent grantGrantedGR01 | GR01 | |
| Entry into force of request for substantive examinationSE01 | SE01 | |
| PublicationPB01 | PB01 |
Numbers
- Publication
- 113689837
- Application
- 109780653
Titles2
- Chinese
- 音频数据处理方法、装置、设备以及存储介质
- English
- Audio data processing method, device, equipment and storage medium
Classification
- CPC, 21
- G10H1/02
- G10L21/0272
- G10L21/013
- G10L21/0308
- G10L13/047
- G10L25/24
- G10L25/30
- G10L25/09
- G10L25/93
- G10H2210/021
- G10H2210/005
- G10H2210/155
- G10H2210/041
- G10H2250/455
- G10H1/366
- G10H2250/311
- G10H2210/331
- Y02D30/70
- G10L13/02
- G10L19/032
- G10L21/0232
- IPC, 8
- G10H1 02
- G10L21 0272
- G10L21 0308
- G10L13 047
- G10L25 24
- G10L25 30
- G10L25 09
- G10L25 93