Audio data processing method and apparatus, and device and storage medium
17 claims: 15 independent, 2 dependent
- 1オーディオデータ処理装置によるオーディオデータ処理方法であって、 オリジナルオーディオデータを分解し、人声オーディオデータ及び背景オーディオデータを取得することと、前記人声オーディオデータに対して電気音響化処理を行い、電気音響人声データを取得することと、前記電気音響人声データと前記背景オーディオデータを合成して、目標オーディオデータを取得することと、を含 み 前記人声オーディオデータに対して電気音響化処理を行い、電気音響人声データを取得することは、 前記人声オーディオデータのオリジナルの基本周波数を抽出することと、 前記オリジナル基本周波数を補正し、第一基本周波数を取得することと、 予め定められた電気音響パラメータに基づいて、前記第一基本周波数を調整し、第二基本周波数を取得することと、 前記第二基本周波数に対して量子化処理を行い、第三基本周波数を取得することと、 前記第三基本周波数に基づいて、前記電気音響人声データを決定することと、を含み、 前記オリジナル基本周波数を補正し、第一基本周波数を取得することは、 前記人声オーディオデータを複数のオーディオセグメントに分けることと、 前記複数のオーディオセグメントにおける各オーディオセグメントに対して、前記オーディオセグメントのエネルギー及びゼロクロスレートを決定することと、 前記エネルギー及びゼロクロスレートに基づいて、前記オーディオセグメントが濁音オーディオセグメントであるか否かを決定することと、 線形補間アルゴリズムを利用して、前記濁音オーディオセグメントの基本周波数を補正することと、を含む オーディオデータ処理方法。
- 2前記オリジナルオーディオデータを分解し、背景オーディオデータ及び人声オーディオデータを取得することは、前記オリジナルオーディオデータに対応するオリジナルメルスペクトルデータを決定することと、ニューラルネットワークを用いて前記オリジナルメルスペクトルデータに対応する背景メルスペクトルデータ及び人声メルスペクトルデータを決定することと、前記背景メルスペクトルデータに基づいて、前記背景オーディオデータを生成し、かつ前記人声メルスペクトルデータに基づいて、前記人声オーディオデータを生成することと、を含む請求項1に記載の方法。
- 3前記オーディオセグメントに複数のサンプリングポイントが設置され、前記オーディオセグメントのエネルギーを決定することは、前記オーディオセグメントにおける各サンプリングポイントの数値に基づいて、前記オーディオセグメントのエネルギーを決定することを含む請求項 1 に記載のオーディオデータ処理方法。
- 4前記オーディオセグメントは複数のサンプリングポイントを含み、前記オーディオセグメントのゼロクロスレートを決定することは、前記オーディオセグメントにおける隣接する二つのサンプリングポイント毎の数値の符号が互いに逆であるか否かを決定することと、前記オーディオセグメントにおける隣接するサンプリングポイントが異なる符号である回数が全てのサンプリングポイントの個数を占める割合を決定し、前記ゼロクロスレートとすることと、を含む請求項 1 に記載の方法。
- 5前記予め定められた電気音響パラメータは、電気音響程度パラメータ及び/又は電気音響トーンパラメータを含み、前記予め定められた電気音響パラメータに基づいて、前記第一基本周波数を調整し、第二基本周波数を取得することは、前記濁音オーディオセグメントの基本周波数に基づいて、基本周波数分散及び/又は基本周波数平均値を決定することと、前記電気音響程度パラメータ及び前記基本周波数分散に基づいて、補正基本周波数分散を決定し、及び/又は、前記電気音響トーンパラメータ及び前記基本周波数平均値に基づいて、補正基本周波数平均値を決定することと、前記補正基本周波数分散及び/又は補正基本周波数平均値に基づいて、前記第一基本周波数を調整し、前記第二基本周波数を取得することと、を含む請求項 1 に記載の方法。
- 6前記第二基本周波数に対して量子化処理を行い、第三基本周波数を取得することは、以下の式に基づいて周波数範囲を決定することを含み、 ここで、前記scale は、前記周波数範囲であり、前記F0’ は、前記第二基本周波数であり、前記周波数範囲に基づいて、以下の式に基づいて前記第三基本周波数を決定し、 ここで、前記F0’’ は、前記第三基本周波数である請求項 1、 3~ 5 のいずれか一項に記載の方法。
- 7前記人声オーディオデータ及び前記第一基本周波数に基づいて、スペクトルエンベロープ及び非周期パラメータを決定することをさらに含み、ここで、前記第三基本周波数に基づいて、前記電気音響人声データを決定することは、前記第三基本周波数、前記スペクトルエンベロープ及び前記非周期パラメータに基づいて、前記電気音響人声データを決定することを含む請求項 1、 3~ 5 のいずれか一項に記載の方法。
- 8オリジナルオーディオデータを分解し、人声オーディオデータ及び背景オーディオデータを取得するための分解モジュールと、前記人声オーディオデータに対して電気音響化処理を行い、電気音響人声データを取得するための電気音響処理モジュールと、前記電気音響人声データと前記背景オーディオデータを合成し、目標オーディオデータを取得するための合成モジュールと、を含 み 前記電気音響処理モジュールは、 前記人声オーディオデータのオリジナル基本周波数を抽出するための抽出サブモジュールと、 前記オリジナル基本周波数を補正し、第一基本周波数を取得するための補正サブモジュールと、 予め定められた電気音響パラメータに基づいて、前記第一基本周波数を調整し、第二基本周波数を取得するための調整サブモジュールと、 前記第二基本周波数に対して量子化処理を行い、第三基本周波数を取得するための量子化サブモジュールと、 前記第三基本周波数に基づいて、前記電気音響人声データを決定するための電気音響決定サブモジュールと、を含み、 前記補正サブモジュールは、 前記人声オーディオデータを複数のオーディオセグメントに分けるためのセグメント化ユニットと、 前記複数のオーディオセグメントにおける各オーディオセグメントに対して、前記オーディオセグメントのエネルギーを決定するためのエネルギー決定ユニットと、 前記複数のオーディオセグメントにおける各オーディオセグメントに対して、前記オーディオセグメントのゼロクロスレートを決定するためのゼロクロスレート決定ユニットと、 前記エネルギー及びゼロクロスレートに基づいて、前記オーディオセグメントのタイプが濁音オーディオセグメントであるか否かを決定するための濁音判断ユニットと、 線形補間アルゴリズムを用いて、前記濁音オーディオセグメントの基本周波数を補正するための補正ユニットと、を含む オーディオデータ処理装置。
- 9前記分解モジュールは、前記オリジナルオーディオデータに対応するオリジナルメルスペクトルデータを決定するためのメルスペクトル決定サブモジュールと、ニューラルネットワークを用いて前記オリジナルメルスペクトルデータに対応する背景メルスペクトルデータ及び人声メルスペクトルデータを決定するための分解サブモジュールと、前記背景メルスペクトルデータに基づいて、前記背景オーディオデータを生成し、前記人声メルスペクトルデータに基づいて、前記人声オーディオデータを生成するための生成サブモジュールと、を含む請求項 8 に記載の装置。
- 10前記オーディオセグメントに複数のサンプリングポイントが設置され、前記エネルギー決定ユニットは、さらに、前記オーディオセグメントにおける各サンプリングポイントの数値に基づいて、前記オーディオセグメントのエネルギーを決定する請求項 8 に記載の装置。
- 11前記オーディオセグメントは複数のサンプリングポイントを含み、前記ゼロクロスレート決定ユニットは、さらに、前記オーディオセグメントにおける隣接する二つのサンプリングポイント毎の数値の符号が互いに逆であるか否かを決定し、前記オーディオセグメントにおける隣接するサンプリングポイントが異なる符号である回数が全てのサンプリングポイントの個数を占める割合を決定し、前記ゼロクロスレートとする請求項 8 に記載の装置。
- 12前記予め定められた電気音響パラメータは、電気音響程度パラメータ及び/又は電気音響トーンパラメータを含み、前記調整サブモジュールは、前記濁音オーディオセグメントの基本周波数に基づいて、基本周波数分散及び/又は基本周波数平均値を決定するための第一決定ユニットと、前記電気音響程度パラメータ及び前記基本周波数分散に基づいて、補正基本周波数分散を決定し、及び/又は、前記電気音響程度パラメータ及び前記基本周波数平均値に基づいて、補正ベース周波数平均値を決定するための第二決定ユニットと、前記補正基本周波数分散及び/又は補正基本周波数平均値に基づいて、前記第一基本周波数を調整し、前記第二基本周波数を取得するための調整ユニットと、を含む請求項 8 に記載の装置。
- 13前記量子化サブモジュールは、周波数範囲決定ユニットおよび第三基本周波数決定ユニットを含み、前記周波数範囲決定ユニットは、以下の式に基づいて周波数範囲を決定するために用いられ、 ここで、前記scale は、前記周波数範囲であり、前記F0’ は、前記第二基本周波数であり、前記第三基本周波数決定ユニットは、前記周波数範囲に基づいて、以下の式に基づいて前記第三基本周波数を決定するために用いられ、 ここで、前記F0’’ は、前記第三基本周波数である請求項 8、10 ~ 12 のいずれか一項に記載の装置。
- 14前記人声オーディオデータ及び前記第一基本周波数に基づいて、スペクトルエンベロープ及び非周期パラメータを決定するための決定モジュールをさらに含み、ここで、前記電気音響決定サブモジュールは、さらに、前記第三基本周波数、前記スペクトルエンベロープ及び前記非周期パラメータに基づいて、前記電気音響人声データを決定する請求項 8、10 ~ 12 のいずれか一項に記載の装置。
- 15少なくとも一つのプロセッサと、前記少なくとも一つのプロセッサと通信接続されたメモリとを含み、前記メモリは、前記少なくとも一つのプロセッサにより実行可能な命令を記憶し、前記少なくとも一つのプロセッサが請求項1- 7 のいずれか一項に記載の方法を実行することができるように、前記命令は前記少なくとも一つのプロセッサにより実行される、電子機器。
- 16コンピュータ命令を記憶した非一時的なコンピュータ可読記憶媒体であって、前記コンピュータ命令は 、コ ンピュータに請求項1- 7 のいずれか一項に記載の方法を実行させるために用いられる記憶媒体。
- 17プロセッサにより実行される時に請求項1- 7 のいずれか一項に記載の方法を実現する命令を含むコンピュータプログラム。
Independent claims17
108 paragraphs, as filed
This application claims priority to the Chinese patent application filed on August 24, 2021 and with application number 202110978065.3, the entire contents of which are incorporated by reference into this application.
TECHNICAL FIELD The present disclosure relates to the field of audio processing technology, and particularly to the field of speech synthesis technology.
Electroacoustic effects are used as audio filters to adjust and beautify audio, and have wide application in scenes such as KTV productions or small video productions. Good electroacoustic effects can improve the audio quality of your work. If high-quality electroacoustic effects can be provided to application products, it will improve the product's competitiveness, enrich the way the product can be played, and increase user interest.
The present disclosure provides an audio data processing method, apparatus, device, storage medium, and program.<u style="Single">Mu</u>provide.
According to one aspect of the present disclosure, an audio data processing method is provided, comprising decomposing original audio data to obtain human voice audio data and background audio data, and performing electroacousticization on the human voice audio data. and obtaining target audio data by synthesizing the electroacoustic human voice data and the background audio data.
According to another aspect of the present disclosure, an audio data processing apparatus is provided, including a decomposition module for decomposing original audio data and obtaining human voice audio data and background audio data; an electroacoustic processing module for performing electroacousticization processing and obtaining electroacoustic human voice data; a synthesis module for synthesizing the electroacoustic human voice data and the background audio data to obtain target audio data; including.
Another aspect of the disclosure provides an electronic device including at least one processor and a memory communicatively coupled to the at least one processor, the memory storing instructions executable by the at least one processor. The instructions are stored and executed by the at least one processor so that the at least one processor can perform the methods described in the embodiments of the present disclosure.
According to another aspect of embodiments of the present disclosure, a non-transitory computer-readable storage medium having computer instructions stored thereon is provided, the computer instructions for causing the computer to perform a method illustrated in an embodiment of the present disclosure. used for.
According to another aspect of embodiments of the present disclosure, a computer program<u style="Single">Mu</u>and computer programs/instructions that, when executed by a processor, implement the methods described in the embodiments of the present disclosure.
It should be understood that the content described in this section is not intended to identify key points or important features of the embodiments of the disclosure and is not intended to limit the scope of the disclosure. Other features of the disclosure will be readily understood from the following description.
The drawings are used for a better understanding of the solution and are not intended to limit the disclosure.
<figref num="1">FIG. 1 schematically shows a flowchart of an audio data processing method according to an embodiment of the present disclosure.</figref><figref num="2">FIG. 2 schematically depicts a flowchart of a method for decomposing original audio data according to an embodiment of the present disclosure.</figref><figref num="3">FIG. 3 schematically shows a flowchart of a method for performing electroacousticization processing on human voice audio data according to an embodiment of the present disclosure.</figref><figref num="4">FIG. 4 schematically shows a flowchart of an audio data processing method according to another embodiment of the present disclosure.</figref><figref num="5">FIG. 5 schematically shows a block diagram of an audio data processing device according to an embodiment of the present disclosure.</figref><figref num="6">FIG. 6 schematically depicts a block diagram of an exemplary electronic device for implementing an embodiment of the invention.</figref>
The following describes exemplary embodiments of the present disclosure with reference to the drawings and includes various details of embodiments of the present disclosure for ease of understanding, and which should be considered as illustrative. be. Accordingly, those skilled in the art will appreciate that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for the sake of clarity and brevity, well-known functions and structures will not be described in the following description.
An audio data processing method according to an embodiment of the present disclosure will be described below using FIG. 1. It should be explained that in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of related data such as audio data all comply with the provisions of the relevant law regulations. , and does not violate public order and morals.
FIG. 1 is a flowchart of an audio data processing method according to an embodiment of the present disclosure.
As shown in FIG. 1, this audio data processing method 100 includes the following.
In operation S110, the original audio data is decomposed to obtain human voice audio data and background audio data.
In operation S120, electroacoustic processing is performed on the human voice audio data to obtain electroacoustic human voice data.
In operation S130, electroacoustic human voice data and background audio data are synthesized to obtain target audio data.
According to embodiments of the present disclosure, the original audio data may include, for example, human voice information and background voice information, where the human voice may be, for example, a singing voice, and the background voice may be, for example, accompaniment music. You can. In this embodiment, for example, a sound source separation algorithm can be used to separate human voice information and background information in the original audio data, and obtain human voice audio data including human voice information and background audio data including background audio information. can.
According to the embodiment of the present disclosure, human voice information is converted into electric voice by separating human voice information and background voice information in original audio data, and the electric voice information and background voice information are synthesized. This realizes electroacoustic conversion of audio data that simultaneously contains background voice information and human voice information.
According to embodiments of the present disclosure, original audio data can be decomposed by implementing a sound source separation algorithm using a neural network. The input of the neural network may be audio data having background speech information and human voice information, and the output of the neural network may be human voice audio data including human voice information and background audio data including background voice information. Something can happen.
According to an embodiment of the present disclosure, a music file and a human voice file are obtained in advance, the music file and the human voice file are cut into segments of equal length, and a plurality of music segments X and a plurality of human voice segments Y are obtained. I can do it. Each music segment X and one corresponding human voice segment Y can be synthesized to obtain original audio data Z. The neural network is trained by using each original audio data Z as an input to the neural network, and using the music segment X and the human voice segment Y corresponding to the original audio data Z as expected outputs. Also, in order to improve the training effect and accelerate the convergence of the network, the music segment X, the human voice segment Y, and the original audio data Z can all be preprocessed into a mel spectrum. Correspondingly, the output result of the neural network is also based on the mel spectrum. Illustratively, the output result in the mel-spectral format can be synthesized with corresponding original audio data by an algorithm such as the Griffin-Lim algorithm.
Based on this, the method of decomposing the above original audio data will be further explained below with reference to FIG. 2 by combining specific embodiments. As will be understood by those skilled in the art, the following illustrative examples are used to understand the present disclosure and are not intended to limit the present disclosure.
FIG. 2 schematically depicts a flowchart of a method for decomposing original audio data according to an embodiment of the present disclosure.
As shown in FIG. 2, a method 210 of decomposing original audio data includes the following.
In operation S211, original mel spectrum data corresponding to the original audio data is determined.
Next, in operation S212, background mel spectral data and human voice mel spectral data corresponding to the original mel spectral data are determined using a neural network.
According to embodiments of the present disclosure, the background mel spectral data may include background audio information in the original mel spectral data, and the human voice mel spectral data may include human voice information in the original mel spectral data.
In operation S213, background audio data is generated based on the background mel spectral data, and human voice audio data is generated based on the human voice mel spectral data.
According to embodiments of the present disclosure, background audio data is generated based on background mel spectral data by an algorithm such as the Griffin-Lim algorithm, and human voice audio data is generated based on human voice mel spectral data. I can do it.
According to the embodiment of the present disclosure, electroacousticization processing for human voice audio data can be realized by quantizing the fundamental frequency of human voice data. For example, the fundamental frequency, spectral envelope, and aperiodic parameters of human voice data can be determined. Here, the fundamental frequency indicates the vibration frequency of the vocal cords during pronunciation, and is the pitch of the tone when embodied in audio. Next, the fundamental frequency is quantized, and the human voice data is resynthesized based on the quantized fundamental frequency, the spectrum envelope, and the aperiodic parameter, thereby realizing electroacousticization processing for the human voice audio data. Here, the resynthesized human voice data is electroacoustic human voice data, and includes human voice information having an electric sound effect.
Hereinafter, with reference to FIG. 3, the method of electroacoustic processing for human voice audio data described above will be further explained by combining specific examples. As will be understood by those skilled in the art, the following illustrative examples may be used to understand the present disclosure, and the present disclosure is not limited thereto.
FIG. 3 schematically shows a flowchart of a method for performing electroacousticization processing on human voice audio data according to an embodiment of the present disclosure.
As shown in FIG. 3, a method 320 for performing electroacousticization processing on this human voice audio data may include the following.
In operation S321, the original fundamental frequency of the human voice audio data is extracted.
According to embodiments of the present disclosure, the original fundamental frequency can be extracted from human voice audio data based on algorithms such as DIO and Harvest.
In operation S322, the original fundamental frequency is corrected to obtain a first fundamental frequency.
According to the embodiments of the present disclosure, the electroacoustic effect can be improved by correcting the fundamental frequency. For example, in this embodiment, human voice audio data can be divided into multiple audio segments. Next, for each audio segment in the plurality of audio segments, the audio segment energy and zero crossing rate are determined. Based on the energy and zero crossing rate, it is determined whether the audio segment is a voiced sound audio segment. A linear interpolation algorithm is then utilized to correct the fundamental frequency of the voiced sound audio segment.
According to an embodiment of the present disclosure, human voice audio data is divided into a plurality of audio segments with a predetermined unit length, and the length of each audio segment is one predetermined unit length. Here, the predetermined unit length can be set according to actual demand. Illustratively, in this embodiment, the predetermined unit length may be any value from 10ms to 40ms.
According to embodiments of the present disclosure, multiple sampling points are provided for each audio segment. Based on the numerical value of each sampling point in the audio segment, the energy of the audio segment can be determined. For example, the energy of an audio segment can be calculated based on the following formula:
<math num="1"><img file="JP7465992B2_D0001.tif" /></math> Here, x<sub>i</sub>indicates the numerical value of the i-th sampling point, and n is the number of sampling points.
According to embodiments of the present disclosure, the number of sampling points n may be determined based on the length of the audio segment and the sampling rate. Taking the predetermined unit length as 10ms as an example, the number of sampling points n can be calculated based on the following formula:<math num="2"><img file="JP7465992B2_D0002.tif" /></math> Here, sr represents the audio sampling rate.
According to embodiments of the present disclosure, it may be determined whether the numerical values of two adjacent sampling points in an audio segment are opposite in sign to each other. Next, the ratio of the number of times that adjacent sampling points of the audio segment have opposite signs to the total number of sampling points is determined as the zero crossing rate.
According to embodiments of the present disclosure, the zero-crossing rate of an audio segment may be calculated based on the following equation:
<math num="3"><img file="JP7465992B2_D0003.tif" /></math> where ZCR is the zero crossing rate of the audio segment, n is the number of sampling points in the audio segment, and x<sub>i</sub>represents the numerical value of the i-th sample point in the audio segment, x<sub>i-1</sub> represents the numerical value of the i-1th sampling point in the audio segment.
According to embodiments of the present disclosure, the number of sampling points n may be determined based on the length of the audio segment and the sampling rate. Taking the predetermined unit length as 10ms as an example, the number of sampling points n can be calculated based on the following formula:<math num="4"><img file="JP7465992B2_D0004.tif" /></math> Here, sr represents the audio sampling rate.
When the human body produces sounds, the corresponding fundamental frequency is 0 because the vocal cords do not vibrate in response to clear sounds. Because the vocal cords vibrate in response to voiced sounds, the corresponding fundamental frequency is not 0. Based on this, in this embodiment, the fundamental frequency can be corrected using the above characteristics.
For example, for each audio segment, if the energy E of the audio segment is less than the threshold e_min and the zero cross rate ZCR of the audio segment is greater than the threshold zcr_max, then the audio segment is a clear sound audio segment and its basic The frequency is 0. Otherwise, the audio segment is a voiced sound audio segment and its fundamental frequency is not zero. Here, e_min and zcr_max can be set according to actual demand.
For each clear audio segment, the fundamental frequency of the audio segment can be set to zero. For the voiced sound audio segment, extract the fundamental frequency of each voiced sound audio segment based on algorithms such as DIO and Harvest, and then detect whether the fundamental frequency value of each voiced sound audio segment is 0 or not one by one. be able to. For a voiced sound audio segment whose fundamental frequency value is 0, a non-zero fundamental frequency value is calculated by performing linear interpolation based on the neighboring voiced sound audio segment values of the voiced sound audio segment based on a linear interpolation algorithm. It can be obtained as the fundamental frequency value of the voiced sound audio segment.
For example, there are six voiced audio segments, and the fundamental frequency values are 100, 100, 0, 0, 160, and 100, respectively. That is, the fundamental frequency value of the third and fourth voiced sound audio segments is 0. Therefore, linear interpolation can be performed based on the non-zero fundamental frequency values in the vicinity of the fundamental frequency values of the third and fourth voiced audio segments, i.e. the second fundamental frequency value 100 and the fifth fundamental frequency value 160, perform linear interpolation and obtain that the fundamental frequency values of the third and fourth voiced sound audio segments are 120 and 140. That is, the six fundamental frequency values after correction are 100, 100, 120, 140, 160, and 100.
Next, in operation S323, the first fundamental frequency is adjusted based on predetermined electroacoustic parameters, and the second fundamental frequency is obtained.
According to embodiments of the present disclosure, the predetermined electroacoustic parameters may include, for example, an electroacoustic degree parameter and/or an electroacoustic tone parameter. Here, the electroacoustic intensity parameter may be used to control the electroacoustic intensity. Electroacoustic tone parameters may be used to control the tone. Illustratively, in this embodiment, the electroacoustic degree parameter may include, for example, 1, 1.2, and 1.4, and the larger the electroacoustic degree parameter is, the more pronounced the electrosound effect is. Electroacoustic tone parameters may include, for example, -3, -2, -1, +1, +2, +3. where -1, -2, -3 indicate lowering by 1, 2, and 3 tones, respectively, and +1, +2, +3 indicate increasing by 1, 2, and 3 tones, respectively. to show that
In the related technology, the electroacoustic effect cannot adjust the parameters and has a single effect. According to embodiments of the present disclosure, two parameters, an electroacoustic degree parameter and an electroacoustic tone parameter, are set based on the characteristics of electroacoustic, and are used to control the electroacoustic effect to meet the needs of different users. can be met.
According to embodiments of the present disclosure, a fundamental frequency variance and/or a fundamental frequency average value may be determined based on the fundamental frequencies of all voiced sound audio segments. A corrected fundamental frequency variance is determined based on the electroacoustic extent parameter and the fundamental frequency variance, and/or a corrected fundamental frequency average value is determined based on the electroacoustic extent parameter and the fundamental frequency average value. Next, the first fundamental frequency is adjusted based on the corrected fundamental frequency variance and/or the corrected fundamental frequency average value to obtain a second fundamental frequency.
Illustratively, in this embodiment, the variance of the fundamental frequencies of all the voiced sound audio segments can be calculated, and the average value of the fundamental frequencies of all the voiced sound audio segments is calculated as the fundamental frequency variance, and the fundamental frequency average value shall be.
The corrected fundamental frequency dispersion can then be calculated based on the following formula:<math num="5"><img file="JP7465992B2_D0005.tif" /></math> Here, new_var is the corrected fundamental frequency variance, var is the fundamental frequency variance, and a is the electroacoustic degree parameter.
The corrected fundamental frequency average value can be calculated based on the following formula:<math num="6"><img file="JP7465992B2_D0006.tif" /></math> Here, new_mean is the corrected fundamental frequency average value, mean is the fundamental frequency average value, and b is the electroacoustic tone parameter.
The second fundamental frequency can then be calculated based on the following formula:<math num="7"><img file="JP7465992B2_D0007.tif" /></math> Here, F0' is the second fundamental frequency.
In operation S324, quantization processing is performed on the second fundamental frequency to obtain a third fundamental frequency.
In natural audio, the vocal tones are intonations, which vary gradually, and electroacoustics quantizes the tones into a specific scale, where the tones vary discontinuously, similar to the tones transmitted from electronic musical instruments. do. Based on this, according to the embodiment of the present disclosure, the fundamental frequency of human voice data can be quantized using the frequency of each key of the piano as the target frequency.
Illustratively, in this embodiment, the frequency range can be determined based on the following formula:<math num="8"><img file="JP7465992B2_D0008.tif" /></math> Here, scale is the frequency range and F0 ́ is the second fundamental frequency.
Then, based on the frequency range, the third fundamental frequency can be determined based on the following formula:<math num="9"><img file="JP7465992B2_D0009.tif" /></math> Here, F0'' is the third fundamental frequency.
In operation S325, electroacoustic human voice data is determined based on the third fundamental frequency.
According to embodiments of the present disclosure, a spectral envelope and aperiodic parameters can be determined based on human voice audio data and a first fundamental frequency. Electroacoustic human voice data can then be determined based on the third fundamental frequency, the spectral envelope and the aperiodic parameters.
The audio data processing method described above will be further explained below by combining specific examples with reference to FIG. As will be understood by those skilled in the art, the following illustrative examples may be used to understand the present disclosure, and the present disclosure is not limited thereto.
FIG. 4 schematically shows a flowchart of an audio data processing method according to another embodiment of the present disclosure.
As shown in FIG. 4, this audio data processing method 400 includes the following. In operation S401, it is determined whether the audio data (abbreviated as audio) includes accompaniment music (abbreviated as accompaniment). If accompaniment is included, operation S402 is executed. If only human voices are included and no accompaniment is included, operation S403 is executed.
In operation S402, human voice and accompaniment are separated using a sound source separation algorithm. Then, operation S403 is executed for the separated human voice.
In operation S403, the zero close rate, fundamental frequency f0, and energy are extracted for the human voice.
In operation S404, the fundamental frequency is corrected based on the zero cross rate and energy to obtain F0.
In operation S405, a spectral envelope SP and an aperiodic parameter AP are calculated using the human voice and the corrected fundamental frequency F0.
In operation S406, the fundamental frequency is adjusted to obtain F0' based on the electroacoustic degree parameter a and the electroacoustic tone parameter b set by the user.
In operation S407, fundamental frequency F0' is quantized to obtain F0''.
In operation S408, a human voice with an electric sound effect is synthesized using the fundamental frequency F0'', the spectral envelope SP, and the aperiodic parameter AP.
In operation S409, if the audio has an accompaniment, operation S410 is performed. Otherwise, operation S411 is performed.
In operation S410, the accompaniment is also matched to the human voice, and final audio with electroacoustic effects is generated.
In operation S411, audio with electroacoustic effects is output.
According to the audio data processing method according to the embodiment of the present disclosure, it is possible to flexibly and efficiently add electroacoustic effects to audio data, thereby improving the entertainment interest of the user.
FIG. 5 schematically shows a block diagram of an audio data processing device according to an embodiment of the present invention.
As shown in FIG. 5, this audio data processing device 500 includes a decomposition module 510, an electroacoustic processing module 520, and a synthesis module 530.
The decomposition module 510 is used to decompose the original audio data and obtain human voice audio data and background audio data.
The electroacoustic processing module 520 is used to perform electroacoustic processing on human voice audio data to obtain electroacoustic human voice data.
The synthesis module 530 is used to synthesize electroacoustic human voice data and background audio data to obtain target audio data.
According to embodiments of the present disclosure, the decomposition module may include a mel spectrum determination submodule, a decomposition submodule, and a generation submodule. Here, the mel spectral determination sub-module is used to determine original mel spectral data corresponding to original audio data. The decomposition sub-module is used to determine background mel spectral data and human voice mel spectral data corresponding to the original mel spectral data using a neural network. The generation sub-module is used to generate background audio data based on the background mel spectral data and to generate human voice audio data based on the human voice mel spectral data.
According to embodiments of the present disclosure, the electroacoustic processing module may include an extraction submodule, a correction submodule, an adjustment submodule, a quantization submodule, and an electroacoustic determination submodule. Here, the extraction sub-module is used to extract the original fundamental frequency of human voice audio data. The correction sub-module is used to correct the original fundamental frequency and obtain the first fundamental frequency. The adjustment sub-module is used to adjust the first fundamental frequency to obtain a second fundamental frequency based on predetermined electroacoustic parameters. The quantization sub-module is used to perform quantization processing on the second fundamental frequency to obtain a third fundamental frequency. The electroacoustic determination sub-module is used to determine electroacoustic human voice data based on the third fundamental frequency.
According to embodiments of the present disclosure, the correction sub-module may include a segmentation unit, an energy determination unit, a zero crossing rate determination unit, a voiced sound determination unit, and a correction unit. Here, the segmentation unit is used to divide human voice audio data into multiple audio segments. The energy determining unit is used for determining the energy of the audio segment for each audio segment in the plurality of audio segments. The zero crossing rate determining unit is used to determine the zero crossing rate of the audio segment for each audio segment in the plurality of audio segments. The voiced sound determination unit is used to determine whether the type of the audio segment is a voiced sound audio segment based on the energy and the zero crossing rate. The correction unit is used to correct the fundamental frequency of the voiced audio segment using a linear interpolation algorithm.
According to embodiments of the present disclosure, multiple sampling points are provided in an audio segment. The energy determining unit is further used to determine the energy of the audio segment based on the numerical value of each sampling point in the audio segment.
According to embodiments of the present disclosure, the zero crossing rate determination unit is further used to determine whether the numerical values of every two adjacent sampling points in the audio segment are opposite in sign to each other; The ratio of the number of times that adjacent sampling points have opposite signs to the total number of sampling points is determined as the zero crossing rate.
According to embodiments of the present disclosure, the predetermined electroacoustic parameters may include an electroacoustic degree parameter and/or an electroacoustic tone parameter. The coordination sub-module may include a first decision unit, a second decision unit and a coordination unit. Here, the first determining unit is used to determine the fundamental frequency variance and/or the fundamental frequency average value based on the fundamental frequency of the voiced sound audio segment. The second determining unit determines a corrected fundamental frequency variance based on the electroacoustic intensity parameter and the fundamental frequency variance, and/or determines a corrected fundamental frequency average value based on the electroacoustic intensity parameter and the fundamental frequency average value. used for The adjustment unit is used to adjust the first fundamental frequency to obtain a second fundamental frequency based on the corrected fundamental frequency variance and/or the corrected fundamental frequency average value.
According to embodiments of the present disclosure, the quantization sub-module may include a frequency range determining unit and a third fundamental frequency determining unit.
Here, the frequency range determination unit is used to determine the frequency range based on the following formula:<math num="10"><img file="JP7465992B2_D0010.tif" /></math> Here, scale is the frequency range and F0' is the second fundamental frequency.
The third fundamental frequency determining unit is used to determine the third fundamental frequency based on the following formula based on the frequency range:<math num="11"><img file="JP7465992B2_D0011.tif" /></math> Here, F0'' is the third fundamental frequency.
According to embodiments of the present disclosure, the audio data processing apparatus may further include a determination module, which is used to determine the spectral envelope and the aperiodic parameter based on the human voice audio data and the first fundamental frequency. It will be done.
According to embodiments of the present disclosure, the electroacoustic human voice determination sub-module is further used to determine electroacoustic human voice data based on the third fundamental frequency, the spectral envelope, and the aperiodic parameter.
According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.<u style="Single">Mu</u>provide.
FIG. 6 is a schematic block diagram illustrating an example of an electronic device 600 that can implement embodiments of the present disclosure. Electronic equipment is intended to refer to various types of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, large format computers, and other suitable computers. Electronic devices can further represent various types of mobile devices, such as personal digital processing, mobile phones, smart phones, wearable devices and other similar computing devices. The components, their connections and relationships, and their functions depicted herein are illustrative only and are not intended to limit implementation of the disclosure as described and/or required herein.
As shown in FIG. 6, the device 600 includes a calculation unit 601, which is based on a computer program stored in a read-only memory (ROM) 602 or loaded into a random access memory (RAM) 603 from a storage unit 608. may perform various appropriate operations and processing. RAM 603 can further store various programs and data necessary for operating device 600. Computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. An input/output interface 605 is also connected to the bus 604.
A plurality of components in the device 600 are connected to an I/O interface 605, and include an input unit 606 such as a keyboard and a mouse, an output unit 607 such as various types of displays and speakers, and an output unit 607 such as a magnetic disk, an optical disk, etc. It includes a storage unit 608 and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 enables device 600 to exchange information/data with other devices via computer networks such as the Internet and/or various telecommunication networks.
Computing unit 601 may be a variety of general purpose and/or special purpose processing modules with processing and computing capabilities. Some examples of the calculation unit 601 include a central processing unit (CPU), a GPU (Graphics Processing Unit), various dedicated artificial intelligence (AI) calculation chips, calculation units for various behavioral machine learning model algorithms, and a DSP (Digital Signal processor), as well as any suitable processor, controller, microcontroller, etc. The calculation unit 601 performs the methods and processes described above, such as, for example, audio data processing methods. For example, in some embodiments, the audio data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and/or installed on the electronic device 600 via the ROM 602 and/or the communication unit 609. When the computer program is loaded into RAM 603 and executed by calculation unit 601, it may perform one or more steps of the audio data processing method described above. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the audio data processing method in any other suitable manner (eg, via firmware).
Various embodiments of the systems and techniques described herein include digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and application specific standard products (ASSPs). ), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and/or combinations thereof. These various embodiments are implemented in one or more computer programs that are executed and/or interpreted on a programmable system that includes at least one programmable processor. The programmable processor may be a special purpose or general purpose programmable processor and receives data and instructions from a storage system, at least one input device, and at least one output device, and receives data and instructions from a storage system, at least one input device, and at least one output device. The method may include being able to transmit instructions to the storage system, the at least one input device, and the at least one output device.
Program code for implementing the methods of this disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program codes may be implemented in a flowchart and/or block format. The functions/operations specified in the diagram are performed. The program code may be executed entirely on the device, partially on the device, partially on the device as a separate software package, and partially on a remote device, or It may be performed entirely on a remote device or server.
In the context of this disclosure, a machine-readable medium may be a tangible medium and includes a program for use in or in combination with an instruction-execution system, device, or electronic device. Or it may be memorized. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or electronic device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connection through one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), fiber optics, portable compact disc read-only memory (CD-ROM), optical storage, magnetic storage, or any suitable combination of the foregoing.
A computer may implement the systems and techniques described herein to provide interaction with a user, and the computer may include a display device (e.g., a CRT (cathode ray tube) or a liquid crystal display (LCD) monitor), a keyboard and a pointing device (eg, a mouse or trackball) through which a user can provide input to the computer. Other types of devices may further provide interaction with the user, for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback). and may receive input from the user in any form, including voice input, audio input, or tactile input.
The systems and techniques described herein may be used in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components. system (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein); The present invention may be implemented in a computing system that includes any combination of background components, middleware components, or front-end components. The components of the system may be connected to each other by any form or medium of digital data communication (eg, a communication network). Examples of communication networks illustratively include local area networks (LANs), wide area networks (WANs), and the Internet.
A computer system may include a client and a server. Clients and servers are generally remote and typically interact via a communications network. The relationship between client and server is created by a computer program running on the relevant computer and having a client-server relationship. The server may be a cloud server, a distributed system server, or a blockchain coupled server.
It should be understood that the various types of flows shown above may be used and steps may be re-sorted, added or removed. For example, each step described in the present invention may be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution of the present disclosure can be achieved. The specification is not limited here.
The specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should appreciate that various modifications, combinations, subcombinations, and substitutions may be made depending on design requirements and other factors. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| WO2019116889A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP2021516786A | Cites | Japan |
| JP2013117556A | Cites | Japan |
| JP2012098318A | Cites | Japan |
| WO2020145353A1 | Cites | World Intellectual Property Organization (WIPO) |
8 members in 5 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 2021109780653 | China | – | |
| 202110978065 | China | A | |
| 2022082305 | China | W |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| CN113689837A | China | A | |
| WO2023024501A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP4167226A1 | European Patent Office (EPO) | A1 | |
| CN113689837B | China | B | |
| JP2023542760A | Japan | A | |
| JP7465992B2This record | Japan | B2 | |
| US2024212703A1 | United States of America | A1 | |
| EP4167226A4 | European Patent Office (EPO) | A4 |
8 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 7465992
- Application
- 560146
Titles2
- Japanese
- オーディオデータ処理方法、装置、機器、記憶媒体及びプログラム
- English
- Audio data processing method, device, equipment, storage medium and program
Classification
- CPC, 21
- G10H1/02
- G10L21/0272
- G10L21/013
- G10L21/0308
- G10L13/047
- G10L25/24
- G10L25/30
- G10L25/09
- G10L25/93
- G10H2210/021
- G10H2210/005
- G10H2210/155
- G10H2210/041
- G10H2250/455
- G10H1/366
- G10H2250/311
- G10H2210/331
- Y02D30/70
- G10L13/02
- G10L19/032
- G10L21/0232
- IPC, 2
- G10L21 013
- G10L21 0272
