Audio file format conversion
Abstract
According to the invention, the manipulation of audio data can be simplified, such as, for example, with relation to the combination of individual audio channels to give multi-channel audio data streams, or for the general manipulation of an audio data stream, whereby a data block is modified (56) in an audio data stream (10), divided into data blocks (10a, 10b) with determining blocks (14, 16) and data block audio data (18) such as, for example, by inclusion in, addition to, or replacement of a part thereof, itself containing a length indicator which expresses a data amount or length of the block audio data, or a data amount or length of the data block, such as to give a second audio data stream with modified data blocks. Alternatively, an audio data stream (10) with pointers in determining blocks (14, 10), which point to the determining block audio data (44, 46), allocated to the determining blocks, but distributed in various data blocks, is converted into an audio stream, whereby the determining block audio data (44, 46) are combined to give coherent determining block audio data (48). The coherent determining block audio data (48) can be contained with the corresponding determining block (14, 16) in a self-contained channel element (52a).

Term
Term ended
Projected expiry passed 13 July 2024, 2.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1Patent claims Zastrzeżenia patentowe 1. The method of converting the first audio data stream, which represents the encoded audio signal, covering time intervals, and has the first file format, in the second audio data stream, which represents the encoded audio signal and has a second file format, where the time range includes a number of audio values, and where, according to the first file format, the first audio data stream is divided into successive data blocks, wherein the data block has a header and audio data of the data block, all headers have identical redundancy for all headers, involving the following step:modifying data blocks, that they contain information about the length, which indicates the amount of data data blocks or the amount of data audio data data blocks, to obtain channel elements from data blocks, which make up the second audio data stream wherein the modification stage involves replacing the identical for all headers of the redundant part with the length information, wherein the method further comprises preceding (60, 62) a second audio data stream with a combined header, and the combined header has the redundant portion identical for all headers or the redundant portion identical for all headers is a synchronization word. 1. Sposób konwersji pierwszego strumienia danych audio, który reprezentuje zakodowany sygnał audio, obejmujący przedziały czasu, i ma pierwszy format pliku, w drugim strumieniu danych audio, który reprezentuje zakodowany sygnał audio oraz ma drugi format pliku, przy czym przedział czasu obejmuje pewną ilość wartości audio, oraz gdzie zgodnie z pierwszym formatem pliku pierwszy strumień danych audio jest podzielony na następujące kolejno po sobie bloki danych, przy czym blok danych posiada nagłówek oraz dane audio bloku danych, przy czym wszystkie nagłówki wykazują identyczną dla wszystkich nagłówków część nadmiarową, obejmujący następujący etap: modyfikowanie bloków danych, aby zawierały one informację o długości, która wskazuje na ilość danych bloków danych lub też ilość danych danych audio bloków danych, w celu uzyskania z bloków danych elementów kanałów, które tworzą drugi strumień danych audio, przy czym etap modyfikacji obejmuje zastępowanie identycznej dla wszystkich nagłówków części nadmiarowej przez informację o długości, przy czym sposób ponadto obejmuje poprzedzanie (60, 62) drugiego strumienia danych audio nagłówkiem łącznym, a nagłówek łączny posiada identyczną dla wszystkich nagłówków część nadmiarową lub też identyczna dla wszystkich nagłówków część nadmiarowa jest słowem synchronizacji. 2. The method according to claim 1, wherein the header (14, 16) is assigned header audio data, which is obtained by encoding the time interval, the header comprising an indicator that indicates the beginning of the header audio data (12a-12c), and where the end of the data (12a -12c) the header audio is before the beginning of the data (12b, 12c) of the header audio in the audio data stream that is assigned to the next data block, comprising the following steps: 2. Sposób według zastrz. 1, przy czym do nagłówka (14, 16) przyporządkowane są dane audio nagłówka, które są uzyskiwane poprzez kodowanie przedziału czasu, przy czym nagłówek zawiera wskaźnik, który wskazuje na początek danych (12a-12c) audio nagłówka, oraz gdzie koniec danych (12a-12c) audio nagłówka znajduje się przed początkiem danych (12b, 12c) audio nagłówka w strumieniu danych audio, które są przyporządkowane do następnego bloku danych, obejmujący następujące etapy: combining (42) data (44, 46) audio header, which are assigned to one header, from at least two data blocks, to obtain related header data (48), which form part of the second audio data stream;attaching (50) related data (48) header audio to header (14, 16) to which audio header data are assigned (44, 46) from which the associated header audio data is obtained to obtain a channel element (52a);arranging channel elements to obtain a second audio data stream;łączenie (42) danych (44, 46) audio nagłówka, które są przyporządkowane do jednego nagłówka, z co najmniej dwóch bloków danych, w celu uzyskania powiązanych ze sobą danych (48) audio nagłówka, które tworzą część drugiego strumienia danych audio;dołączanie (50) powiązanych ze sobą danych (48) audio nagłówka do nagłówka (14, 16), do którego przyporządkowane są dane audio nagłówka (44, 46), z których uzyskiwane są powiązane ze sobą dane audio nagłówka w celu otrzymania elementu (52a) kanału;rozmieszczanie elementów kanałów w celu otrzymania drugiego strumienia danych audio;and modifying (56) the channel element (54a-54c) to include information about the length that indicates the amount of data of the channel element (54a-54c) or the amount of data related to the audio header data, wherein the modification step is to replace (56 ) identical for all headers, the redundant portion by length information. oraz modyfikowanie (56) elementu (54a-54c) kanału, aby zawierał on informację o długości, która wskazuje ilość danych elementu (54a-54c) kanału lub też ilość danych powiązanych ze sobą danych audio nagłówka, przy czym etap modyfikowania ma zastępować (56) identyczną dla wszystkich nagłówków część nadmiarową przez informację o długości. 3. The method according to claim 3. The method of claim 1 or 2, where the joining step comprises the following partial steps: 3. Sposób według zastrz. 1 albo 2, w przypadku którego etap łączenia obejmuje następujące etapy częściowe: odczytywanie wskaźnika w nagłówku;reading the indicator in the header;odczytywanie pierwszej części danych audio nagłówka, która jest zawarta w danych audio bloku danych jednego z co najmniej dwóch bloków danych oraz obejmuje początek danych audio nagłówka, na który wskazuje wskaźnik nagłówka;reading the first portion of the header audio data that is contained in the audio data of the data block of one of the at least two data blocks and includes the beginning of the header audio data indicated by the header indicator;odczytywanie drugiej części danych audio nagłówka, która jest zawarta w danych audio bloku danych drugiego z co najmniej dwóch bloków danych oraz obejmuje koniec danych audio nagłówka;oraz łączenie części pierwszej i drugiej. reading the second portion of the header audio data, which is contained in the audio data of the data block of the second of the at least two data blocks and includes the end of the header audio data;and combining the first and second parts. 4. A method of combining a first audio data stream that represents an encoded first audio signal and a second audio data stream that represents an encoded second audio signal into a multi-channel audio data stream, comprising the following steps: 4. Sposób łączenia pierwszego strumienia danych audio, który reprezentuje zakodowany pierwszy sygnał audio, oraz drugiego strumienia danych audio, który reprezentuje zakodowany drugi sygnał audio, do postaci wielokanałowego strumienia danych audio, obejmujący następujące etapy: converting the first audio data stream to the first partial audio data stream according to the method of one of claims 1 to 3;and converting the second audio data stream to the second partial audio data stream according to the method of one of claims 1 to 3, with the deployment steps being carried out in such a way that both partial audio data streams together form a multi-channel audio data stream, and that in the multi-channel audio data stream the elements (70a) of the channels of the first partial audio data stream and the elements (72a) of the channels of the second partial audio data stream, respectively, which contain related audio header data, obtained by coding simultaneous time intervals, are arranged one after the other in an associated access unit (78). konwersję pierwszego strumienia danych audio na pierwszy częściowy strumień danych audio zgodnie ze sposobem według jednego z zastrzeżeń 1 do 3;oraz konwersję drugiego strumienia danych audio na drugi częściowy strumień danych audio zgodnie ze sposobem według jednego z zastrzeżeń 1 do 3, przy czym etapy rozmieszczania są wykonywane w taki sposób, że oba częściowe strumienie danych audio tworzą razem wielokanałowy strumień danych audio, oraz że w wielokanałowym strumieniu danych audio odpowiednio elementy (70a) kanałów pierwszego częściowego strumienia danych audio oraz elementy (72a) kanałów drugiego częściowego strumienia danych audio, które zawierają powiązane ze sobą dane audio nagłówka, uzyskane poprzez kodowanie równoczesnych przedziałów czasu, są umieszczone jeden za drugim w powiązanej ze sobą jednostce (78) dostępowej. 5. The method according to claim 4, which further includes the following stage: 5. Sposób według zastrz. 4, który ponadto obejmuje następujący etap: poprzedzanie drugiego strumienia danych audio nagłówkiem łącznym, przy czym nagłówek łączny zawiera informację o formacie, która określa, w jakiej kolejności uporządkowane są elementy (70a) kanałów pierwszego częściowego strumienia danych audio oraz drugiego częściowego strumienia (70b) danych audio w jednostkach (78) dostępowych. preceding the second audio data stream with a combined header, wherein the combined header contains format information that specifies the order in which the elements (70a) of the channels of the first partial audio data stream and the second partial stream (70b) of audio data are ordered in access units (78) . 6. The method according to one of the preceding claims, in which the data blocks are data blocks of the same or predefined variable size, which depends on the information about the sampling frequency and the bit rate information contained in their header. 6. Sposób według jednego z powyższych zastrzeżeń, w którym bloki danych są blokami danych o tym samym lub z góry zdefiniowanym zmiennym rozmiarze, który zależy od informacji o częstotliwości próbkowania oraz informacji o przepływności, zawartej w ich nagłówku. 7. The method according to one of the claims 1 to 3, which also has the following stages: 7. Sposób według jednego z zastrz. 1 do 3, który ponadto ma następujące etapy: zeroing (180) the indicators in the headers, and therefore they define as the beginning of the header audio data that the header audio data begins immediately after the header concerned;and changing (182) the bit rate information in the headers, so the bit length dependent data block information according to the first audio data format is sufficient to fit the given header and the assigned header audio data. zerowanie (180) wskaźników w nagłówkach, w związku z czym określają one jako początek danych audio nagłówka, że dane audio nagłówka rozpoczynają się bezpośrednio za danym nagłówkiem;oraz zmienianie (182) informacji o przepływności w nagłówkach, w związku z czym zależna od informacji o przepływności długość bloku danych zgodnie z pierwszym formatem danych audio jest wystarczająca, aby zmieścić dany nagłówek oraz przyporządkowane dane audio nagłówka. 8. Method of decoding the second audio data stream, which represents the encoded audio signal, covering time intervals, and has a second file format, based on the decoder, which is able to decode the first audio data stream, which represents the encoded audio signal and has the first data format, to the form of an audio signal, where the time range includes a number of audio values, and where, according to the first file format, the first audio data stream with the bit reservoir function is divided into successive blocks of data (10a-10c), wherein the data block has a header (14, 16) and audio data block data (18), with the heading (14, 16) audio header data are assigned, obtained by coding the time interval, the header contains a pointer, which points to the beginning of the header audio data (12a-12c), and where the end of the data (12a-12c) of the header audio is before the beginning of the data (12b, 12c) header audio in the audio data stream, which are assigned to the next block of data, and where, according to the second file format, the second audio data stream is divided into channel elements, wherein the channel element contains related data (44, 46) audio header, which are obtained by combining header audio data, assigned to the header, from two data blocks, and contain the assigned header in the form, in which identical for all headers, previously the redundant part is modified, so that it can be replaced by length information, which determines the amount of data for a given channel element or the corresponding amount, related header data, the second audio data stream precedes the combined header, which has a redundant portion identical for all headers, comprising the following stages: 8. Sposób dekodowania drugiego strumienia danych audio, który reprezentuje zakodowany sygnał audio, obejmujący przedziały czasu, i ma drugi format pliku, na podstawie dekodera, który jest w stanie zdekodować pierwszy strumień danych audio, który reprezentuje zakodowany sygnał audio i ma pierwszy format danych, do postaci sygnału audio, przy czym przedział czasu obejmuje pewną ilość wartości audio, oraz gdzie zgodnie z pierwszym formatem pliku pierwszy strumień danych audio z funkcją rezerwuaru bitów jest podzielony na następujące kolejno po sobie bloki (10a-10c) danych, przy czym blok danych ma nagłówek (14, 16) oraz dane audio bloku danych (18), przy czym do nagłówka (14, 16) przyporządkowane są dane audio nagłówka, uzyskiwane poprzez kodowanie przedziału czasu, przy czym nagłówek zawiera wskaźnik, który wskazuje na początek danych (12a-12c) audio nagłówka, oraz gdzie koniec danych (12a-12c) audio nagłówka znajduje się przed początkiem danych (12b, 12c) audio nagłówka w strumieniu danych audio, które są przyporządkowane do następnego bloku danych, oraz gdzie zgodnie z drugim formatem pliku drugi strumień danych audio jest podzielony na elementy kanałów, przy czym element kanału zawiera powiązane ze sobą dane (44, 46) audio nagłówka, które są otrzymywane w wyniku połączenia danych audio nagłówka, przyporządkowanych do nagłówka, z dwóch bloków danych, i zawierają przyporządkowany nagłówek w formie, w jakiej identyczna dla wszystkich nagłówków, uprzednio nadmiarowa część jest zmodyfikowana, aby mogła być zastąpiona przez informację o długości, która określa ilość danych danego elementu kanału lub też ilość odpowiednich, powiązanych ze sobą danych nagłówka, przy czym drugi strumień danych audio poprzedza nagłówek łączny, który posiada identyczną dla wszystkich nagłówków część nadmiarową, obejmujący następujące etapy: creating an input data stream that represents the encoded audio signal and has a first file format from the second audio data stream by verifying that the second audio data stream is a reformatted data stream from the first file format;tworzenie wejściowego strumienia danych, który reprezentuje zakodowany sygnał audio i ma pierwszy format pliku, z drugiego strumienia danych audio poprzez weryfikację, czy w przypadku drugiego strumienia danych audio chodzi o strumień danych przeformatowany z pierwszego formatu pliku;odczytywanie identycznej, uprzednio nadmiarowej części z nagłówka łącznego;reading an identical, previously redundant portion from the combined header;replacing length information in headers with an identical, previously redundant portion;zastępowanie informacji o długości w nagłówkach przez identyczną, uprzednio nadmiarową część;zeroing the indicators in the headers of the channel elements of the second audio data stream, and therefore they define as the beginning of the header audio data that the header audio data begins immediately after the given header to obtain the reset headers;zerowanie wskaźników w nagłówkach elementów kanałów drugiego strumienia danych audio, w związku z czym określają one jako początek danych audio nagłówka, że dane audio nagłówka rozpoczynają się bezpośrednio za danym nagłówkiem, w celu otrzymania nagłówków zerowania;changing information about the bit rate in the headers of channel elements of the second audio data stream, therefore, the data block dependent on the bit rate information in accordance with the first audio data format in all headers is sufficient, to fit a given header and the associated header audio data, to obtain changed headers and zero headers;and inserting bits between each channel element and the next channel element, therefore, the length of each channel element plus inserted bits is adapted to the changed bit rate information, and feeding the input data stream to the decoder according to the changed bit rate information to obtain an audio signal. zmienianie informacji o przepływności w nagłówkach elementów kanałów drugiego strumienia danych audio, w związku z czym zależna od informacji o przepływności długość bloku danych zgodnie z pierwszym formatem danych audio we wszystkich nagłówkach jest wystarczająca, aby zmieścić dany nagłówek oraz przyporządkowane dane audio nagłówka, w celu otrzymania nagłówków o zmienionej przepływności oraz nagłówków zerowania;oraz wstawianie bitów pomiędzy każdym elementem kanału oraz następnym w kolejności elementem kanału, w związku z czym długość każdego elementu kanału plus wstawionych bitów jest dostosowana do zmienionej informacji o przepływności, oraz doprowadzanie wejściowego strumienia danych do dekodera zgodnie ze zmienioną informacją o przepływności w celu otrzymania sygnału audio. 9. Device for converting the first audio data stream, which represents the encoded audio signal, covering time intervals, and has the first file format, to the second audio data stream, which represents the encoded audio signal and has a second file format, where the time range includes a number of audio values, and where, according to the first file format, the first audio data stream is divided into successive data blocks, wherein the data block has a header and audio data of the data block, all headers have identical redundancy for all headers, having the following feature: 9. Urządzenie do konwersji pierwszego strumienia danych audio, który reprezentuje zakodowany sygnał audio, obejmujący przedziały czasu, i ma pierwszy format pliku, na drugi strumień danych audio, który reprezentuje zakodowany sygnał audio i ma drugi format pliku, przy czym przedział czasu obejmuje pewną ilość wartości audio, oraz gdzie zgodnie z pierwszym formatem pliku pierwszy strumień danych audio jest podzielony na następujące kolejno po sobie bloki danych, przy czym blok danych posiada nagłówek oraz dane audio bloku danych, przy czym wszystkie nagłówki posiadają identyczną dla wszystkich nagłówków część nadmiarową, mające następującą cechę: system for modifying data blocks, that they contain information about the length, which indicates the amount of data data blocks or the amount of data audio data data blocks, to obtain channel elements from data blocks, which make up the second audio data stream wherein the modification stage involves replacing the identical for all headers of the redundant part with the length information, the device further has a placement system (60, 62) the combined header before the second audio data stream, and the combined header has the redundant portion identical for all headers or where the redundant identical for all headers is a synchronization word. układ do modyfikowania bloków danych, aby zawierały one informację o długości, która wskazuje na ilość danych bloków danych lub też ilość danych danych audio bloków danych, w celu uzyskania z bloków danych elementów kanałów, które tworzą drugi strumień danych audio, przy czym etap modyfikacji obejmuje zastępowanie identycznej dla wszystkich nagłówków części nadmiarowej przez informację o długości, przy czym urządzenie ponadto ma układ do umieszczania (60, 62) nagłówka łącznego przed drugim strumieniem danych audio, a nagłówek łączny ma identyczną dla wszystkich nagłówków część nadmiarową lub też gdzie identyczna dla wszystkich nagłówków część nadmiarowa jest słowem synchronizacji. 10. The device according to claim 9, wherein the header (14, 16) is assigned header audio data, which is obtained by encoding the time interval, the header comprising an indicator that indicates the beginning of the header audio data (12a-12c), and where the end of the data (12a -12c) the header audio is before the beginning of the data (12b, 12c) of the header audio in the audio data stream that is assigned to the next data block having the following characteristics: 10. Urządzenie według zastrz. 9, przy czym do nagłówka (14, 16) przyporządkowane są dane audio nagłówka, które są uzyskiwane poprzez kodowanie przedziału czasu, przy czym nagłówek zawiera wskaźnik, który wskazuje na początek danych (12a-12c) audio nagłówka, oraz gdzie koniec danych (12a-12c) audio nagłówka znajduje się przed początkiem danych (12b, 12c) audio nagłówka w strumieniu danych audio, które są przyporządkowane do następnego bloku danych, mające następujące cechy: system for combining (42) data (44, 46) audio header, which are assigned to one header, from two data blocks, to obtain related header data (48), which form part of the second audio data stream;a system for attaching (50) related data (48) header audio to header (14, 16) to which the data are assigned (44, 46) audio header, from which the associated header audio data is obtained to obtain a channel element (52a);and an arrangement for arranging channel elements to obtain a second audio data stream;and a system for modifying (56) a channel element (54a-54c), so that it contains information about the length, which indicates the amount of data of the channel element (54a-54c) or the amount of data related to the audio header data, wherein the modifying system (56) is shaped to replace the identical for all headers of the redundant portion by length information. układ do łączenia (42) danych (44, 46) audio nagłówka, które są przyporządkowane do jednego nagłówka, z dwóch bloków danych, w celu uzyskania powiązanych ze sobą danych (48) audio nagłówka, które tworzą część drugiego strumienia danych audio;układ do dołączania (50) powiązanych ze sobą danych (48) audio nagłówka do nagłówka (14, 16), do którego przyporządkowane są dane (44, 46) audio nagłówka, z których uzyskiwane są powiązane ze sobą dane audio nagłówka w celu otrzymania elementu (52a) kanału;oraz układ do rozmieszczania elementów kanałów w celu otrzymania drugiego strumienia danych audio;oraz układ do modyfikowania (56) elementu (54a-54c) kanału, aby zawierał on informację o długości, która wskazuje ilość danych elementu (54a-54c) kanału lub też ilość danych powiązanych ze sobą danych audio nagłówka, przy czym układ do modyfikowania (56) jest ukształtowany w celu zastępowania identycznej dla wszystkich nagłówków części nadmiarowej przez informację o długości. 11. A device for decoding the second audio data stream, which represents the encoded audio signal, covering time intervals and has a second file format, based on the decoder, which is able to decode the first audio data stream, which represents the encoded audio signal and has the first data format, to the form of an audio signal, where the time range includes a number of audio values, and where, according to the first file format, the first audio data stream is divided into successive blocks (10a-10c) of data, wherein the data block has a header (14, 16) and audio data block data (18), with the heading (14, 16) audio header data are assigned, obtained by coding the time interval, the header contains a pointer, which points to the beginning of the header data (12a12c), and where the end of the data (12a-12c) of the header audio is before the beginning of the data (12b, 12c) header audio in the audio data stream, which are assigned to the next block of data, and where, according to the second file format, the second audio data stream is divided into channel elements, wherein the channel element contains related data (44, 46) audio header, which are obtained by combining header audio data, assigned to the header, from two data blocks, and include the assigned header in the form, in which identical for all headers, previously the redundant part is modified, so that it can be replaced by length information, which determines the amount of data for a given channel element or the corresponding amount, related header data, the second audio data stream precedes the combined header, which has a redundant portion identical for all headers, having the following characteristics: 11. Urządzenie do dekodowania drugiego strumienia danych audio, który reprezentuje zakodowany sygnał audio, obejmujący przedziały czasu i ma drugi format pliku, na podstawie dekodera, który jest w stanie zdekodować pierwszy strumień danych audio, który reprezentuje zakodowany sygnał audio oraz ma pierwszy format danych, do postaci sygnału audio, przy czym przedział czasu obejmuje pewną ilość wartości audio, oraz gdzie zgodnie z pierwszym formatem pliku pierwszy strumień danych audio jest podzielony na następujące kolejno po sobie bloki (10a-10c) danych, przy czym blok danych posiada nagłówek (14, 16) oraz dane audio bloku danych (18), przy czym do nagłówka (14, 16) przyporządkowane są dane audio nagłówka, uzyskiwane poprzez kodowanie przedziału czasu, przy czym nagłówek zawiera wskaźnik, który wskazuje na początek danych (12a12c) audio nagłówka, oraz gdzie koniec danych (12a-12c) audio nagłówka znajduje się przed początkiem danych (12b, 12c) audio nagłówka w strumieniu danych audio, które są przyporządkowane do następnego bloku danych, oraz gdzie zgodnie z drugim formatem pliku drugi strumień danych audio jest podzielony na elementy kanałów, przy czym element kanału zawiera powiązane ze sobą dane (44, 46) audio nagłówka, które są otrzymywane w wyniku połączenia danych audio nagłówka, przyporządkowanych do nagłówka, z dwóch bloków danych, i obejmują przyporządkowany nagłówek w formie, w jakiej identyczna dla wszystkich nagłówków, uprzednio nadmiarowa część jest zmodyfikowana, aby mogła być zastąpiona przez informację o długości, która określa ilość danych danego elementu kanału lub też ilość odpowiednich, powiązanych ze sobą danych nagłówka, przy czym drugi strumień danych audio poprzedza nagłówek łączny, który posiada identyczną dla wszystkich nagłówków część nadmiarową, mające następujące cechy: a system for creating an input data stream that represents the encoded audio signal and has a first file format from the second audio data stream by verifying that the second audio data stream is a reformatted data stream from the first file format;układ do tworzenia wejściowego strumienia danych, który reprezentuje zakodowany sygnał audio i ma pierwszy format pliku, z drugiego strumienia danych audio poprzez weryfikację, czy w przypadku drugiego strumienia danych audio chodzi o strumień danych przeformatowany z pierwszego formatu pliku;odczytywanie identycznej, uprzednio nadmiarowej części z nagłówka łącznego;reading an identical, previously redundant portion from the combined header;replacing length information in headers with an identical, previously redundant portion;zastępowanie informacji o długości w nagłówkach przez identyczną, uprzednio nadmiarową część;zeroing the indicators in the headers of the channel elements of the second audio data stream, and therefore they define as the beginning of the header audio data that the header audio data begins immediately after the given header to obtain the reset headers;zerowanie wskaźników w nagłówkach elementów kanałów drugiego strumienia danych audio, w związku z czym określają one jako początek danych audio nagłówka, że dane audio nagłówka rozpoczynają się bezpośrednio za danym nagłówkiem, w celu otrzymania nagłówków zerowania;changing information about the bit rate in the headers of channel elements of the second audio data stream, therefore, the data block dependent on the bit rate information in accordance with the first audio data format in all headers is sufficient, to fit a given header and the associated header audio data, to get headers with changed bit rates and restored;and inserting bits between each channel element and the next channel element, therefore, the length of each channel element plus inserted bits is adapted to the changed bit rate information, and a system for feeding the input data stream to the decoder according to the changed bit rate information to obtain an audio signal. zmienianie informacji o przepływności w nagłówkach elementów kanałów drugiego strumienia danych audio, w związku z czym zależna od informacji o przepływności długość bloku danych zgodnie z pierwszym formatem danych audio we wszystkich nagłówkach jest wystarczająca, aby zmieścić dany nagłówek oraz przyporządkowane dane audio nagłówka, w celu otrzymania nagłówków o zmienionej przepływności oraz przywróconych;oraz wstawianie bitów pomiędzy każdym elementem kanału oraz następnym w kolejności elementem kanału, w związku z czym długość każdego elementu kanału plus wstawionych bitów jest dostosowana do zmienionej informacji o przepływności, oraz układ do doprowadzania wejściowego strumienia danych do dekodera zgodnie ze zmienioną informacją o przepływności w celu otrzymania sygnału audio. 12. The device according to claim 11, wherein a redundant portion identical for all headers means a sync word. 12. Urządzenie według zastrz. 11, przy czym identyczna dla wszystkich nagłówków część nadmiarowa oznacza słowo synchronizacji. 13. The device according to claim 11 or 12, wherein the system for creating the input data stream is shaped so as to change the rate information to the highest allowed value. 13. Urządzenie według zastrz. 11 albo 12, przy czym układ do tworzenia wejściowego strumienia danych jest ukształtowany tak, aby w przypadku zmiany informacji o przepływności ustawiać ją na najwyższą dozwoloną wartość. 14. A computer program with a program code for performing the method according to one of the claims 1 or 8 when the computer program runs on the computer. 14. Program komputerowy z kodem programu do wykonywania sposobu według jednego z zastrz. 1 albo 8, gdy program komputerowy zostanie uruchomiony w komputerze. Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e.V., Niemcy Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany Pełnomocnik: Proxy: EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 100 100 102 102 FIG. 6 FIG. 6 EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 FIG.3 FIG.3 EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 ELEMENT KANAŁU MP3 MP3 CHANNEL ELEMENT EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 FROM THE ENCODER (FRONT) 7QA Z KODERA (PRZÓD) 7Qa 70b 70b FROM ENCODER 2 (CENTRAL) \ 72a Z KODERA 2 (CENTRALNY) \ 72a 72b 72b FIG. 5 FIG. 5 DANE1C S\\XWW DATA1C S \\XWW WITH THE CODER _Ą \ 74a Z KODERA _Ą \ 74a DANE 2 Cl DATA 2 Cl 132a / 7 /// 7/77 ///) /// 7 ///////////// 4, 'ιν.νιϊΐ GIVEN FROM III ^DATA 0 Ł-> lowing nnf 132a /7///7/77///)///7/////////////4, 'ιν.νιϊΐ PODANE WYPEłnp III ^DANE 0 ł-> niające n nf 124 syncword = OxFFF ID layer protectionbit bitrate_index = OxE sampling_frequency padding bit = 1 124 syncword=OxFFF ID layer protectionbit bitrate_index=OxE sampling_frequency padding bit=1 132b 7777777777777Υ77777777777777777Λ ;dane 0 & 132b 7777777777777Υ77777777777777777Λ;data 0 & DANE WYPE: WYPE DATA: L134a Ł134a INSTANCE INSTANCJA DEKODERA RECEIVER MP3 MP3 EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 FIG. 7 FIG. 7 EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 FIG. 8 FIG. 8 EP 1 647 010 B1 EP 1 647 010 B1 Z-16441/17 Z-16441/17 FIG. 9 FIG. 9
126 paragraphs in 1 section, as filed
[0001] The present invention relates to the conversion of audio data streams encoding audio signals, and more specifically to a better handling of audio data streams in file format, in which audio data belonging to one time stamp are separated into different data blocks, as is the case e.g. in the MP3 format.
[0002] MPEG audio data compression is a particularly effective form of recording audio signals, such as music or film soundtrack, in digital form, while it allows on one hand to take up as little memory space as possible, on the other hand, maintaining good audio quality whenever possible. MPEG audio compression has proved to be one of the most effective solutions in this area in recent years.
[0003] There are currently various versions of the MPEG audio compression method. Generally speaking, the audio signal is sampled at a certain sampling frequency, with the resulting string of audio sampling values being assigned to overlapping time intervals or timestamps. The time stamps are then transmitted individually, e.g. to the hybrid filter bank, consisting of multi-phase filters and a modified, discrete cosine transformation (MDCT) that eliminates the effects of aliasing. Correct data compression then occurs when quantizing the MDCT coefficients. Quantized in this way, the MDCT coefficients are then converted into a Hufmann code, consisting of Hufmann code words, which creates a further level of compression in such a way that the more common coefficients are assigned shorter code words. All in all, MPEG compressions are therefore lossy, however, however, "audible" losses are within certain limits, because the method of quantizing DCT coefficients is influenced by knowledge of psychoacoustics.
[0004] Another widespread MPEG standard is the so-called MP3 standard, described in ISO / IEC 11172-3 and 13818-3. This standard makes it possible to adapt the loss of information due to compression to the bit rate at which audio information should be transmitted in real time. Also for MPEG standards, the transmission of the compressed data signal should be via a fixed rate channel. To ensure that also at low bit rates the listening quality on the receiving decoder will remain sufficient, the MP3 standard provides that the MP3 encoder has so-called bit reservoir. The significance of this fact is described below. Usually, due to the constant bit rate encoder
MP3 should encode each time stamp into a block of codewords of the same size so that the block can then be sent at a specific bit rate over a period of time corresponding to the refresh period. However, it would not be possible to take account of the fact that some parts of the audio signal, e.g. the sounds that follow a very loud sound in a music track require less accurate quantization at a consistent quality level than other parts of the audio signal, e.g. places covering many different instruments. Therefore, the MP3 encoder does not generate a simple bit stream format when each time stamp is encoded in a frame with the same frame length for all frames. Such a closed frame would consist of a frame header, side information and basic data belonging to the time stamp assigned to the frame, namely encoded MDCT coefficients, with the side information being information for the decoder on how to decode DCT coefficients, such as how many consecutive DCT coefficients is 0 to determine which DCT coefficients are sequentially contained in the master data. In the MP3 format, the backside or header is rather included in the side information or in the header. "Backpointer", which indicates the position within the master data in one of the preceding frames. This position contains the beginning of the master data that belongs to the time stamp assigned to the frame in which the corresponding back pointer is contained. The back indicator indicates e.g. number of bytes by which the beginning of the master data in the data stream is shifted. The end of these master data can be in any frame, depending on how high the compression ratio for this time stamp is. The length of the basic data of individual time stamps is therefore no longer constant. Therefore, the number of bits with which one block is coded can be adapted to the signal properties. At the same time, however, a constant bit rate can be obtained. This technique is called "bit reservoir". In general, the bit reservoir is a buffer or a bit store that can be used to provide more bits to encode a block of time sampling values than is actually allowed by the fixed output data rate. The bit reservoir technique takes into account the fact that some blocks of audio sampling values can be encoded using fewer bits, than it is determined by the constant baud rate, therefore these blocks fill the reservoir of bits, while in turn other blocks of audio sampling values have psychoacoustic properties, which do not allow such high compression, therefore, for these blocks for coding with low interference or interference-free, the available bits would not be sufficient. The required excess bits are taken from the bit reservoir, therefore the bit reservoir is emptied for such blocks. The bit reservoir technique is also described in the MPEG Layer 3 standard mentioned above.
[0005] While the MP3 format may also have advantages on the encoder side by providing back indicators, the disadvantages appear clearly on the decoder side. If the decoder e.g. it does not receive the MP3 bit stream from the beginning, but from a specific frame in the middle, then the encoded audio signal with the time stamp assigned to this frame can only be played back immediately if the reverse pointer accidentally has a value of 0, which would indicate that the beginning of the basic data for this frame is located immediately after the header or side data. Usually, however, this is not the case. The reproduction of the audio signal in place of such a time stamp is therefore impossible if the reverse indicator of the first received frame points to the preceding frame which, however, has not (yet) been received. In this case, you can (first) play the next frame.
[0006] Additional problems arise on the receiver side also in the case of general support for frames that are interrelated via a reverse pointer and thus are not introverted. Another problem with the bit stream with return addresses to the bit reservoir is that that if different audio channels are encoded to MP3 individually, belonging to each other in both data streams, as they belong to the same time stamp, master data can possibly be shifted relative to each other, namely with a variable offset within a frame, therefore, it is difficult to reconnect these individual MP3 streams into one multi-channel audio data stream.
[0007] Furthermore, there is a need for an easy possibility of simple to use and MP3-compliant creation of multi-channel audio data streams. Multi-channel MP3 audio data streams in accordance with ISO / IEC 13818-3 require matrixing to recover output channels from the channels transmitted on the decoder side and the use of multiple back indicators, and thus are complicated to operate.
[0008] MPEG 1/2 Layer 2 audio data streams are compatible with MP3 audio data streams in their successive frame composition and frame structure and arrangement, namely a structure composed of header, side information and part of the basic data, as well as the system with a quasi-static frame spacing, which depends on the sampling frequency and the changing bit rate for each frame, they differ, however, in the absence of a backward indicator or bit reservoir during encoding. More and less intensive coding intervals of the audio signal are encoded using the same frame length. The basic data belonging to one time stamp are in the appropriate frame with the appropriate header.
[0009] US-2003/009246 A1 describes a device for playing and / or editing a trick by which it is possible to easily edit MP3 data streams. To this end, it was proposed that, after loading the MP3 file into the MP3 preparation system, initially transform the file in the converter in such a way that a temporary MP3 stream is created, in which the frame data related to the frame are always immediately after the given identification block, in what are the back indicators or "Backpointers" are always 0. During conversion, the appropriate identification block is read first for a specific frame from the original MP3 data stream and the bit rate is set to the maximum possible value or the minimum possible value taking into account the resulting frame length in the intermediate MP3 stream. In addition, the fill bit is set or not set, depending on what is necessary in the resulting intermediate MP3 stream with closed frames. Other fields in another frame header are not changed. Of course, the backward index value is still changed to zero. In addition, the original MP3 data stream reads the frame data for the current frame, and is attached to the newly created identification block, and then the information about the fill is added to the frame usage data to set the length of the resulting and enclosed frame to such , which is determined by the changed bit rate. The resulting intermediate MP3 data stream is then fed to the device for playing and / or editing trick that can perform simple manipulations on it, because now the frames are enclosed within. The intermediate MP3 data stream thus changed is transferred to a typical MP3 decoder.
[0010] In Finlayson R. "A more loss tolerant RTP payload format for MP3 audio", June 2001, URL: http.//<a href="http://www.faqs.org/rfcs/rfc3119.html">www.faqs.org/rfcs/rfc3119.html</a> describes the conversion of an MP3 data stream to a real-time user data format, in short RTP, which is better in the event of packet loss. As part of this conversion, MP3 frames are changed to MP3 Application Data Units, in short ADU frames. Each ADU frame is preceded by an ADU descriptor. The ADU frame differs from the original MP3 frame in that the complete string of encoded audio data and any other arbitrary data for the ADU, i.e. those that start in the original MP3 data stream at the position indicated by the reverse pointer contained in the respective original MP3 frame header and end at the next position as indicated by the reverse indicator in the next MP3 frame are contained in this same ADU frame. In addition, these kinds of encapsulated ADU frames differ from the original MP3 frames only by the optional replacement of the first 11 bits of synchronization in the header of the MP3 frame with the help of the sequence number of the interrelationship, which was provided to allow the selection of a change in the sort order of the ADU frames for transmission in a way that deviates from the correct time order. ADU descriptors added to ADU frames created in this way contain three fields, namely a continuation mark, descriptor type character and information about the size of the ADU, which indicates the size of the ADU frame following the given ADU descriptor. Pairs composed of an ADU frame and an ADU descriptor are packaged in RTP packets, which in turn have an RTP header. If a pair of ADU frame and ADU descriptor does not fit into this packet, it is split into two more RTP packets. In this case, a continuation character is set in the ADU descriptor of the next ADU frame. The descriptor type character only specifies how many bits contain the ADU size information in the ADU descriptor. The RTP header fields include, but are not limited to, timestamp information that indicates the playback time point of the first ADU frame that is packaged in a given packet. This RTP packet data stream together with any the associated ADU frames can then be easily converted to a typical MP3 data stream, namely the original MP3 data stream.
[0011] The object of the present invention is to create a scheme for converting an audio data stream into a subsequent audio data stream or vice versa, thus handling audio data is facilitated, e.g. in terms of combining individual audio data streams into multi-channel audio data streams or also supporting data stream audio in general.
[0012] This task is solved by the method according to claim 1 or 8 and the device according to claim 9 or 11.
[0013] Operation with audio data can be facilitated, e.g. in terms of combining individual audio data streams into multi-channel audio data streams or also handling an audio data stream in general, by modifying a data block in an audio data stream that is divided into data blocks with an identification block and a block with audio data, e.g. with the help of add or replace parts of it in such a way that it contains information about the length, which determines the amount of data or The length of the audio data block data or the amount of data or data block length to get a second audio data stream with modified data blocks. Or, the audio data stream with indicators in the identification blocks that point to the identification blocks assigned to the identification blocks but are separated into different data blocks will be converted to an audio data stream in which the audio data of the identification blocks are grouped into associated forms with each other the audio data of the identification block. The associated audio data of the identification block may then be included together with their identification block in a channel element contained therein.
[0014] The observation within the present invention is that the indicator-based audio data stream, in which the indicator points to the beginning of the audio data of the identification block of the respective data block, is easier to operate if the audio data stream is manipulated, that all audio data of the identification block are combined in it, i.e. audio data that has one and the same time stamp or encode audio values belonging to one and the same audio tag into one associated block of related identification audio data, and if associated the corresponding identification block to which associated identification data audio is associated. The channel elements obtained in this way form after sorting a new audio data stream where all audio data that belongs to the same time stamp or encode audio values or the sampling values for this time stamp are also grouped in one channel element, so that it is easier to handle the new audio data stream.
[0015] According to an embodiment of the present invention, in the case of a new audio data stream, each identification block or each channel element is modified, e.g. by attaching or replacing one part to obtain information on the length, which determines the length or the amount of data of the channel element or related audio data contained therein to facilitate decoding of a new audio data stream using variable length channel elements. Preferably, the modification is performed in such a way that the excess portion of these identifying blocks identical for all identification blocks of the input audio data stream is replaced by the corresponding length information. Thanks to this action you can get that that the data rate of the resulting audio data stream is equal to the rate of the original audio data stream, despite the additional compared to the original, based on the audio stream data stream length information, and additionally, in the new audio data stream, properly redundant back indicators can be included, which can be obtained so that the original audio data stream can be reconstructed from the new audio data stream.
[0016] An identical, redundant portion of these identification blocks may be in the combined identification block placed at the beginning of the emerging new audio data stream. On the receiver side, the resulting second audio data stream can thus be converted back to the original audio data stream in order to use already existing decoders that are able to decode only the audio data streams of the original file format to decode the resulting audio data stream without indicators.
[0017] According to another embodiment of the present invention, the conversion of the first audio data stream to the second audio data stream of a different file format is used to create a multi-channel audio data stream from a plurality of audio data streams of the first file format. On the receiver side, the serviceability is improved compared to just purely combining the original audio data streams with the indicator, because in a multi-channel audio data stream all channel elements that belong to one time stamp or contain related audio data of identification blocks, obtained by coding the simultaneous channel time interval of a multi-channel audio signal, i.e. by coding the time intervals of different channels that belong to the same time stamp, they can be grouped into access units or "Access Units". In the case of indicator-based audio data formats this is not possible because there the audio data belonging to one time stamp can be divided into different data blocks. The placement of data in length data blocks in multiple data streams for different channels allows, when combining audio data streams to form a multi-channel data stream with access units, better syntax analysis by access units.
[0018] The present invention further arises from the observation that it is very easy to convert the audio data streams described above again to the original file format, which can then be decoded by existing decoders into an audio signal. Although the resulting channel elements have different lengths and thus are once longer and once shorter than the length available in the data block of the original audio data stream, it is not necessary to move or re-create the audio data stream in the new file format. merging the master data according to possibly retrieved backward indicators yet, but it is enough to increase the information about the bit rate in the identification blocks of the generated audio data stream of the original file format. The effect of this is that, according to this bit rate information, also the longest of the channel elements in the audio data stream to be decoded is less than or equal to the length of the data block that the data blocks have in the audio data stream of the first file format. Backward indicators are set to zero, and channel elements are extended to the length corresponding to the increased bit rate information by adding "don't care" bits. This creates data blocks of the audio data stream in the original file format, in which the assigned basic data is contained only in the data block itself and not in another. The backwardly converted audio data stream of the first file format can then be fed into the already existing decoder of the audio data stream of the first file format using an increased rate according to the increased bit information. As a result, laborious relocation operations for back conversion become unnecessary, as does the need to replace existing decoders with new ones.
[0019] On the other hand, according to a further embodiment, it is possible to recover the original audio data stream from the resulting audio data stream by using the information contained in the combined identification block of the resulting audio data stream of identical, redundant portion of the identification blocks to restore the portion overwritten by the information about the length.
[0020] Preferred embodiments of the present invention are further explained in more detail with reference to the accompanying drawings. The drawings show:
Fig. 1 a schematic drawing to explain the MP3 file format with a back pointer;
Fig. 2 a block diagram to explain the structure for converting an MP3 audio data stream to an MPEG-4 audio data stream;
Fig. 3 a logic diagram to explain a method of converting an MP3 audio data stream to an MPEG-4 audio data stream in accordance with an embodiment of the present invention;
Fig. 4 a schematic drawing to explain the step of grouping related audio data together with the attachment of identification blocks and the step of modifying the identification blocks as part of the method according to Figure 3;
Fig. 5 a schematic drawing to explain a method of converting multiple MP3 audio data streams into a multi-channel MPEG-4 audio data stream in accordance with another embodiment of the present invention;
Fig. 6 a block diagram of a system for converting the MPEG-4 audio data stream obtained in accordance with Fig. 3 into an MP3 audio data stream again to enable decoding with an existing MP3 decoder;
Fig. 7 a logic diagram of a method for converting an MPEG-4 audio data stream obtained in accordance with Fig. 3 into one or multiple audio data streams in MP3 format;
Fig. 8 is a logic diagram of the method of converting the MPEG-4 audio data stream obtained in accordance with Fig. 3 into one or multiple audio data streams in MP3 format in accordance with another embodiment of the present invention; and
Fig. 9 a schematic diagram of a method of converting an MP3 audio data stream to an MPEG-4 audio data stream in accordance with another embodiment of the present invention.
[0021] The present invention is described below with reference to drawings based on exemplary embodiments, where for the original audio data stream in file format, in which backward pointers are used in the data block identification blocks to indicate to the beginning of the basic data belonging to a given identification block, it's just an example of an MP3 audio data stream, while in the case of the resulting audio data stream, consisting of closed elements of channels, in which audio data belonging to a given time stamp are properly combined, it is also only an example of an MPEG-4 audio data stream. The MP3 format is described in the ISO / IEC 11172-3 and 13818-3 standards cited in the introduction, while the MPEG-4 file format is described in ISO / IEC 14496-3.
[0022] First, referring to Fig. 1, the MP3 format will be briefly explained. In fig. 1 a section of MP3 audio data stream 10 is shown. The audio data stream 10 consists of a series of frames or data blocks, of which in fig. 1 only three were presented in full, namely 10a, 10b and 10c. The MP3 audio data stream 10 is generated by the MP3 encoder from the audio signal or sound signal. The audio signal encoded by the data stream 10 is for example music, language or a mixture of these and the like. The data blocks 10a, 10b and 10c are respectively assigned to one of the successive or possibly overlapping time intervals for which the audio signal is divided by the MP3 encoder. Each time interval corresponds to an audio signal time stamp, and the term time stamp is also often used to refer to time intervals in the description. Each time period was encoded by an MP3 encoder individually, for example thanks to a hybrid filter bank consisting of a multi-phase filter bank and modified discrete cosine transformation with subsequent entropy coding, e.g. Huffman coding, in the main data ("main_data"). The master data, which belong to three consecutive time stamps to which the data blocks 10a-10c are assigned, are shown in Fig. 1 with designations 12a, 12b and 12c as related blocks outside the range of the actual audio data stream.
[0023] Data blocks 10a-10c of the audio data stream 10 are placed in the audio data stream 10 with an equal spacing. This in turn means that each data block 10a-10c has the same data block length or frame length. The frame length, in turn, depends on the bit rate at which the audio data stream should at least be real-time to create, and the sampling frequency that the MP3 encoder used to sample the audio signal before proper encoding. The relationship is that the sampling frequency in conjunction with a constant amount of sampling values per time stamp determines how long the time stamp is, and based on the bit rate and time stamp period, it is possible to calculate, for example, how many bits can be transmitted in that time period.
[0024] Both parameters, i.e. sampling rate and frequency, are indicated in the frame headers 14 in data blocks 10a-10c. Each data block 10a-10c thus has a frame header 14. In general, all the information necessary to decode the audio data stream is itself stored in each frame 10a-10c, so that the decoder has the ability to start decoding in the middle of the MP3 audio data stream 10.
[0025] In addition to the frame header 14 which is at the beginning, each data block 10a-10c still has a portion of 16 side information and a portion of 18 basic data that contains audio data of the data block. Part 16 of side information is placed immediately after heading 14. It contains information that is necessary for the decoder of the audio data stream 10 to find the basic data assigned to a given data block, or audio data of the identification block, which are only Hufmann code words arranged in a series and to decode them correctly in the form of DCT coefficients or MDCT. Part 18 of the master data forms the end of each block of data.
[0026] As already mentioned in the introduction to the description, the MP3 standard supports the reservoir function. Its use is made possible by the backward indicators contained in the side information inside the side 16 information, which are marked in Fig. 1 the number 20 If the backward indicator is 0, then the master data for this side information begins immediately after part 16 of the side information. Otherwise, indicator 20 ("main_data_begin") specifies the start of the master data that encodes the time stamp to which the data block is assigned, in which backward indicator information containing 20 backward information 20 is contained, in the preceding data block. In fig. 1 for example, data block 10a is assigned to a time stamp which is encoded by basic data 12a. The backward indicator in the 16 side information of this data block 10a refers e.g. by means of the bit or byte offset information, measured from the beginning of the header 14 of the data block 16a to the beginning of the basic data 12a, which is in the data block in the direction of stream 22 before the data block 10a. That is, at that time, when encoding the audio signal, the MP3 bit encoder reservoir generating the MP3 audio data stream 10 was not full, but could still be loaded by the height of the reverse indicator. Beginning with the position indicated by the back pointer 20 of the data block 10a, the basic data 12a in the audio data stream 10 are inserted with equally spaced header pairs and side information 14, 16. In the example shown, the master data 12a extends slightly beyond half of the master data portion 18 of the data block 10a. The back indicator 20 in the side information section 16 of the next data block 10b indicates the position immediately after the basic data 12a in data block 10a. Correspondingly, this refers to the back pointer in part 16 of the side information of data block 10c.
[0027] As can be seen, for MP3 audio data stream 10 it is rather an exception if the basic data belonging to one time stamp are actually only in the data block assigned to that time stamp. Data blocks are rather usually divided into one or more data blocks in which, depending on the size of the bit reservoir, not even the respective data block itself must be present. The value of the back indicator is limited by the size of the bit reservoir.
[0028] After describing the structure of the MP3 audio data stream based on Fig. 1, an arrangement will be described with reference to Fig. 2 which is suitable for converting the MP3 audio data stream to the MPEG-4 audio data stream or also from a signal audio receive an MPEG-4 audio data stream that can be easily converted to MP3.
[0029] Fig. 2 shows an MP3 encoder 30 and an MP3-MPEG4 converter 32. The MP3 30 encoder includes an input at which it receives the audio signal to be encoded and an output at which it provides an MP3 audio data stream that encodes the audio signal received at the input. The 30 MP3 encoder works according to the MP3 standard mentioned above.
[0030] An MP3 audio data stream whose structure has been explained with reference to Fig. 1, consists, as mentioned, of frames of fixed frame length, the latter depending on the set bit rate and the sampling frequency being its basis, as well as on setting or not setting the fill byte. The MP3-MPEG4 32 converter receives an MP3 audio data stream at the input and provides an MPEG-4 audio data stream at the output, which structure results from the following description of how the MP3-MPEG4 32 converter works. The meaning and purpose of converter 32 is to convert MP3 audio data stream from MP3 to MPEG-4. The MPEG-4 data format has the advantage that in it all the basic data belonging to a specific time stamp are contained in one related access unit ("Access Unit") or in a channel element, and therefore the handling of these is significantly simplified.
[0031] In Fig. 3 shows the individual steps of the method during the conversion of an MP3 audio data stream to an MPEG-4 audio data stream performed by converter 32. Initially, in step 40, an MP3 audio data stream is received. Reception may include saving the complete audio data stream or only the current portion thereof in temporary memory. Accordingly, subsequent steps can be performed during the conversion process during the real-time receiving process or only after it has been completed.
[0032] In step 42, all audio data or master data that belong to one time stamp are then grouped into an associated block, namely for all time stamps. Step 42 is in Fig. 4 shown in a schematic way closer, in this figure elements similar to the elements of the MP3 audio data stream shown in Fig. 1, have the same or similar designation and no further description of these elements has been given up.
[0033] As can be seen from the direction of the data stream 22, shown in Fig. 4 further on the left side of the MP3 audio data stream portion 10 reaches converter 32 earlier than the right side portions thereof. The two data blocks 10a and 10b are shown in full in Fig. 4. The time stamp belonging to data block 10a is encoded by MD1 basic data, which in Fig. 4 they are contained, for example, in one part in the data block before data block 10a and in the other part in data block 10a, namely in particular in part 18 of its basic data. These master data that encode the time stamp assigned to the next data block 10b are only contained in part 18 of the master data of data block 10a and designated as MD2. The MD3 master data belonging to the data block following data block 10b is divided into parts 18 of the basic data blocks 10a and 10b.
[0034] In step 42, the converter 32 combines all related to each other, i.e., coding one and the same time stamp, master data to form related blocks. In this way, the upstream segment 10a of data 44 and the downstream data block 18a of data section 10a, the MD1 basic data section 46 after completing step 42 together form one by one the associated data block 48. Correspondingly, this is done for the other master data MD2, MD3 ...
[0035] To perform step 42, the converter 32 reads the indicator in the side information 16 of the data block 10a, and then based on this indicator, respectively, the first portion 44 of the audio data block 12a of the identification block relating to that data block 10a, which is contained in field 18 preceding the data block, namely, starting from the location specified by the pointer to the header of the current data block 10a. Second part 46 of the identification block audio data, which is contained in Part 18 of the current data block 10a and includes the end of the audio data of the identification block, data block 10a, it then reads from the end of the 16 side information of the current block of audio data 10a to the beginning of the next audio data, here marked MD2, referring to the next block of data 10b, which is indicated by the indicator in the 16 side information of the next data block 10b, which converter 32 also loads. Joining together both parts 44 and 46 results in block 48 as described.
[0036] At step 50, the converter 32 then attaches the associated headers 14 together with the associated side information 16 to the blocks formed, to ultimately create MP3 channel elements 52a, 52b and 52c. Each element of the MP3 52a-c channel thus consists of the header 14 of the corresponding MP3 data block, the information part 16 of the same side data block of the same MP3 data block and the associated basic data block 48 that encode the time stamp assigned to that data block , from which the headline and side information come from.
[0037] The MP3 channel elements formed in steps 42 and 50 have different lengths of channel elements relative to each other which is indicated by double arrows 54a-54c. It should be remembered that data blocks 10a, 10b did have a fixed frame length of 56 in MP3 data stream 10, however, due to the function of the bit reservoir, the amount of basic data in relation to individual time stamps fluctuates around the average value.
[0038] To facilitate decoding now, and especially parsing or syntax analysis of individual MP3 channel elements 52a-52c on the decoder side, headers 14 H1-H3 are modified to include the length of a given channel element 52a-52c, i.e. 54a-54c. This is done in step 56. The length information is saved in the form of an identical or redundant for all headers 14 of the audio data stream 10. For the MP3 format, for example, each header receives, for example, a synchronization word ("syncword") at the beginning, consisting of 12 bits. At step 56, the synchronization word is occupied by the length of a given channel element. These 12 bits of the sync word are sufficient to reflect the length of a given channel element in binary form, therefore the length of the resulting elements 58a-58c MP3 channels with the modified header h1-h3 remains constant despite the step 56, i.e. equal to 54a-54c. In this way, the audio information after the items 58a-58c of the MP3 channels are ordered in the order of the time stamps they encode can, despite the addition of length information, be transmitted and played back at a constant bit rate in real time, just like the original MP3 audio stream 10 will be overwritten with additional headers.
[0039] In step 58, a file header is then created, or in case the data stream to be generated was not a file but a streaming, data stream header for the expected MPEG-4 audio data stream (step 60). Since the present embodiment generates an MPEG-4 compliant audio data stream, the file header is generated in accordance with the MPEG-4 standard, the file header being structured in this case by the "AudioSpecificConfig" function, which is defined in mentioned MPEG-4 standard. Compatibility with the MPEG-4 system is provided by the "ObjectTypeIndication" element, which is filled with the value 0x40, as well as the information "audioObjectType" number 29. The MPEG-4-specific "AudioSpecificConfig" header is expanded as per the original definition in ISO / IEC 14496-3 in the manner described below, with the following example including not all but only the content of the "AudioSpecificConfig" header relevant to this description:
AudioSpecificConfig () {audioObjectType;
samplingFrequencyIndex;
if (sairLplingFrequencyIndex<sup>= z</sup>= Oxf) samplingFrequency;
channelConfiguration;
if (audioObjectType == 29) {
MPEG_l_2_SpecificConfig ();
}}
[0040] The above list of the "AudiospecificConfig" header is only a presentation in a typical notation for the "AudioSpecificConfig" function, which is used in the decoder to parse or loading call parameters in the file header, namely "samplingFrequencyIndex" (index of sampling frequency), "channelConfiguration" (channel configuration) and "audioObjectType" (type of audio object) or provides instructions on how to decode or
parse the file header.
[0041] As can be seen, the file header that is generated in step 60 starts with the "audioObjectType" information, which as mentioned before is set to 29 (line 2). The "audioObjectType" parameter indicates the decoder and how the data was encoded, and especially how you can get further information about the encoding from the file header, which will be described later.
[0042] Next is the "samplingFrequencyIndex" invocation parameter that indicates a specific place in the normalized sampling rate table (row 3). If the index is 0 (row 4), the sampling frequency is indicated without reference to the standardized table (row 5).
[0043] Next, there is information about the channel configuration (line 6), which in the further described way determines how many channels are contained in the generated MPEG-4 audio data stream, but otherwise than in the present embodiment it is also possible combining more than one MP3 audio data stream to form an MPEG-4 audio data stream, as will be described later with reference to Fig. 5.
[0044] Next, if the parameter "audioObjectType" is 29, which is obviously the case here, the portion in the header of the "AudioSpecificConfig" file that contains the excess portion of the header of the MP3 frame in the audio data stream 10, that is, the portion that among frame headers 14 remain constant (line 8). In this case, this part is designated as "MPEG_1_2_SpecificConfig ()", which is another function that defines the structure of this part.
[0045] Although the "MPEG_1_2_SpecificConfig" structure can also be downloaded from the MP3 standard as it corresponds to the fixed part of the MP3 frame header, which does not change in the following frame, an example of its structure is given below:
MPEG_l_2_SpecificConfig (channelConfiguration) {syncword
ID layer 5 reserved sampling_frequency reserved reserved reserved if (channelConfiguration - 0) {
Kanalkonfigurationsbeschreibung;
}}
[0046] In the "MPEG_1_2_SpecificConfig" portion, all bits that differ from frame header to frame header in the MP3 audio data stream differ from each other are set to 0. For each frame header, the first parameter in "MPEG_1_2_SpecificConfig" is the same in every case, namely the 12-bit sync word "syncword", which is used to synchronize the MP3 encoder when receiving an MP3 audio data stream (line 2). The next ID parameter (line 3) specifies the MPEG version, i.e. 1 or 2, together with the corresponding ISO / IEC 13818-3 standard for version 2 and ISO / IEC 11172-3 for version 1.
The "layer" parameter (line 4) indicates "layer 3", which corresponds to the MP3 standard. The next bit is reserved ("reserved" in line 5) because its value can change from frame to frame and is transmitted by MP3 channel elements. This bit indicates, if necessary, that the header is followed by a CRC variable. The next variable "sampling_frequency" (line 6) refers to the table with sampling frequencies as defined in the MP3 standard, and thus determines the sampling frequency that is the basis of the MP3 DCT factor. Then in line 7 there is again the bit indication for specific applications ("reserved"), as in lines 8 and 9. Next (in line 11, 12) there is an exact definition of the channel configuration, if the parameter indicated in line 6 "AudioSpecificConfig" does not indicate a predefined channel configuration, but has a value of 0. Otherwise, the channel configuration from standard 14496-3 section 1 of table 1.11 applies.
[0047] Due to step 60, and especially because of the "MPEG_1_2_SpecificConfig" element in the file header, which contains all the redundant information in the headers of the 14 frames of the original MP3 audio stream 10, it is guaranteed that this excess part in the frame headers when inserting data to facilitate coding, as for example in step 56 by inserting the length of the channel element, does not lead to irreversible loss of this information in the generated MPEG-4 file, but this modified part can be reconstructed based on the MPEG-4 file header.
[0048] In step 62, then the MPEG-4 audio data stream is in the order generated in step 60 the MPEG-4 data header and channel elements sent in the order of the assigned time stamps, whereby the complete MPEG-4 audio data stream results in an MPEG-4 file. 4 or it is sent by MPEG-4 systems.
[0049] The previous description referred to the conversion of an MP3 audio data stream to an MPEG-4 audio data stream. As indicated, however, by the dotted lines in Fig. 2, it is also possible to convert two or more MP3 audio data streams from two MP3 encoders, namely 30 and 30 ', into an MPEG-4 multi-channel audio data stream. In this case, the MP3-MPEG4 converter 32 receives MP3 audio data streams from all 30 and 30 'encoders and transmits the multi-channel audio data stream in MPEG-4 format.
[0050] In Fig. 5 shown in the upper part based on the view of Fig. 4, how an MPEG-4 compliant MPEG-4 audio data stream can be obtained, wherein the conversion is again performed by the converter 32. Three strings 70, 72 and 74 of the channel elements are shown, which according to steps 40-56 were generated from another audio signal by the 30 encoder respectively. 30 'MP3 (fig. 2). From each sequence of 70, 72 and 74 channel elements, respectively, two channel elements are shown, namely 70a, 70b, 72a, 72b or 74a, 74b. In fig. 5 channel elements arranged one above the other, here 70a-74a or 70b-74b, are respectively assigned to the same time stamp. The channel elements of the string 70 encode, for example, an audio signal that, according to the relevant standard, was recorded from the front left, right ("front"), while strings 72 and 84 encode audio channels that reflect the recording of the same audio source from other directions, or using a different frequency spectrum, such as the center front speaker ("center") and rear right and left ("surround").
[0051] As indicated by arrows 76, the channel elements are during sending (compare step 62 in Fig. 3) in the MPEG-4 audio data stream, attached to the units hereinafter referred to as the "access unit" or access units 78. Therefore, data within one access unit 78 always refers to one time stamp in the MPEG-4 audio data stream. The arrangement of the elements 70a, 72a and 74a of the MP3 channels of the access unit 78, here in the order the front, center and rear channels, is in the header of the file that is created for the MPEG-4 audio data stream to be generated (see step 60 in fig. 3), included in the "AudioSpecificConfig" header by the appropriate parameter of the channel configuration call, referring to section 1 of ISO / IEC 144963. Access units 78 are then sequentially arranged in the MPEG-4 stream in order of timestamps one after the other, followed by the MPEG-4 file header. In the header of the MPEG4 file, the "channelConfiguration" parameter is set accordingly to indicate the order of the channel elements in the access units or their meaning on the decoder side.
[0052] As indicated by the previous description of Fig. 5, it is very easy to group MP3 audio data streams into one multi-channel audio data stream, if as proposed in accordance with the present invention, MP3 audio streams will be manipulated to get closed channel elements composed of data blocks, for which all data of one time stamp is contained in one element of the channel, wherein these channel elements of individual channels can then be easily grouped into access units.
[0053] The preceding description referred to the conversion of one or more MP3 audio data streams into one MPEG-4 audio data stream. An important observation of the present invention, however, also lies in the fact that all the advantages of the resulting MPEG-4 audio data stream, such as better ability to support single, encapsulated MP3 channel elements at the same bit rate and the possibility of multi-channel transmission, can be used without the need for complete replacement existing MP3 decoders via new decoders, but that you can also easily convert or convert backward transformation, so these decoders can be used when decoding the MPEG-4 audio data stream described above.
[0054] In Fig. 6 this is presented in the layout of the 100 MP3 reconstructor, whose operation will be explained in more detail later, and the decoders 102, 102 '... MP3. The MP3 reconstructor 100 receives at the input an MPEG-4 audio data stream that has been generated according to one of the preceding embodiments, and forwards it or, in the case of a multi-channel audio data stream, multiple MP3 audio data streams to one or more. many 102, 102 'decoders ... MP3s, which in turn decode the received MP3 audio data stream into an appropriate audio signal, e.g. they pass it to the appropriate speakers, which are arranged according to the channel configuration.
[0055] A particularly easy method of reconstructing the original MP3 audio data streams generated according to Fig. 5 of the MPEG-4 multi-channel audio data stream is described with reference to Fig. 5 below and Fig. 7, these steps being carried out by the reconstructor system 100 MP3 from Fig. 6.
[0056] At first, the MP3 reconstructor 100 in step 110 verifies that the MPEG-4 audio data stream received at the input is a reformatted MP3 audio data stream, thus checking according to the "AudioSpecificConfig" header the "audioObjectType" call parameter in whether the file header has a value of 29. If this is the case (line 7 in the "AudioSpecificConfig" header), the 100 MP3 reconstructor in its analysis of the syntax of the MPEG-4 audio data stream file header goes further and reads from the "MPEG_1_2_SpecificConfig" part the excess portion of all frame headers of the original MP3 audio data stream from which the MPEG-4 audio data stream (step 112).
[0057] After evaluating the "MPEG_1_2_SpecificConfig" parameter, the MP3 reconstructor 100 then replaces at step 114 in each element 74a-74c of the channel in the hF header therein, hc hs one or more parts of channel elements by means of the "MPEG_1_2_SpecificConfig" parameter components, and especially information about the length of the channel element, per sync word from the "MPEG_1_2_SpecificConfig" parameter, to re-receive the original MP3 HF audio stream headers, HC und HS, as marked by arrows 116. At step 118, the MP3 reconstructor 100 then modifies the Sf, Sc, and Ss side information in the MPEG4 audio data stream in each channel element. Especially the reverse indicator or The "backpointer" is set to 0 to get new S'F, S'C and S'S side information. The manipulation according to step 118 is indicated in Fig. 5 using the 120 arrows. At step 122, the MP3 reconstructor 100 then sets in each channel element 74a-74c a bit rate to the highest allowed value in the header HF, HC, HS, of the frame bearing the word synchronization according to step 114 instead of the channel element length information. Finally, the resulting headers deviate from the original headers, which is indicated in Fig. 5 using an apostrophe, i.e. H'F, H'C and H'S. The manipulation of the channel elements according to step 122 is also indicated by arrows 116.
[0058] To show once again the changes of steps 114-122 in Fig. 5 with respect to the H'F header and the part of the S'F minor index, individual parameters are listed below them. For step 124, the individual parameters of the H'F header are presented. The H'F frame header starts with the "syncword" parameter. The "syncword" parameter is set to its original value (step 114), as is the case with the MP3 audio data stream, namely to the value 0xFFF. Generally, the H'F frame header that is generated following steps 114-122 differs only from the original MP3 frame header that was contained in the original MP3 audio data stream 10 in that the bit rate is set to the highest allowed value, which according to the standard MP3 is 0xE.
[0059] The purpose and purpose of this change in the bit rate is to obtain a new frame length or frame for the newly created MP3 audio data stream. a data block length that is greater than the length of the original MP3 audio data stream from which the MPEG-4 audio data stream with the access unit 78 was generated. The trick here is that the length of the frame in bytes in MP3 format is always dependent on the bit rate, namely according to the formula:
For MPEG 1 layer 3:
Frame length [bit] = 1152 * bit rate [bit / s] / sampling frequency [bit / s] + + 8 * fill bit [bit]
For MPEG 2 layer 3:
Frame length [bit] = 576 * bit rate [bit / s] / sample rate [bit / s] + + 8 * fill bit [bit] [0060] In other words, the frame length of the MP3 audio data stream according to the standard is directly proportional to bit rate and inversely proportional to the sampling frequency. As an additional value, there is also the value of the fill bit, which is given in the headers hF, hC, hS of the MP3 frame and can be used to accurately set the bit rate. The sample rate is constant because it determines how fast the decoded audio signal is played. By changing the bit rate compared to the original setting, it is now also possible to include MP3 channel elements 74a-74c in the data block length of the newly created MP3 audio data stream that are longer than the original elements, because in order to generate the original audio data stream the basic data was generated by means of download bits from the bit reservoir.
[0061] Although in the present embodiment the bit rate is always set to the highest allowed value, it would also be possible to increase the bit rate only to such a value that is sufficient to produce the data block length in accordance with the MP3 standard, and therefore also the longest elements 74a-74c MP3 channels would fit in in length.
[0062] For step 126, it is shown that the back indicator "main_data_begin" in the resulting side information is set to 0. This does not mean anything other than that it was produced according to the method of Fig. 7 MP3 audio data stream data blocks are always closed in themselves, so the master data for a specific frame header and side information always starts immediately after the side information and ends within the same data block.
[0063] Steps 114, 118, 122 are performed on each channel element by extracting it from a function unit, respectively, with information about the length of the channel elements being useful during extraction.
[0064] In step 128, then as many padding data or data are attached to each channel element 74a-74c. "don't care" bits to increase the length of all MP3 channel elements uniformly to the length of the MP3 data block as defined by the new bit rate 0xE. Filling data are shown in Fig. 5 under the designation 128. The amount of padding data can be calculated for each channel element, for example, by analyzing information about the length of the channel element and the fill bit.
[0065] In step 130, then the channel elements modified in accordance with the preceding steps, shown in Fig. 5 as 74a'-74c ', they are transmitted as data blocks of the MP3 audio data stream in the order of coded time stamps to the respective MP3 decoder or Instances 134a-134c MP3 decoder. The MPEG-4 file header is skipped. The resulting MP3 audio data streams are shown in Fig. 5 in general under reference numbers 132a, 132b and 132c. Instances 134a-134c of MP3 decoders are, for example, already initiated earlier, namely in such a number as the number of channels contained in individual access units.
[0066] The MP3 rebuilder 100 knows which channel elements 74a-74c belong in the access unit 78 of the MPEG-4 audio data stream to which of the generated MP3 audio data streams 132a-132c based on the analysis of the "channelConfiguration" call parameter in the "AudioSpecificConfig" header of the stream MPEG-4 audio data. The MP3 decoder instance 134a, which is connected to the front speaker, therefore receives the audio data stream 132a that corresponds to the front channel, and in accordance with the above instances, the MP3 decoders 134b and 134c receive the audio streams 132B and 132c that correspond to the central channel, and rear, and transmit the resulting audio signals to appropriately arranged speakers, for example to the subwoofer or speakers at the back on the left and back on the right.
[0067] However, to encode the MPEG-4 audio data stream in real time by an instance system 102, 102 'of decoders or 134a-134c from Fig. 6 it is necessary for the newly generated MP3 audio data streams 132a-132c to be transmitted with an increased bitrate in step 122 that is higher than that of the original MP3 audio data stream 10, which is not a problem, however, since the arrangement between the MP3 reconstructor 100 and 102, 102 'MP3 or decoders 134a-134c is fixed, therefore the transmission sections can be designed as short and with sufficiently high data transmission speed while maintaining low costs and workload.
[0068] As described with reference to fig. 7 an embodiment obtained in accordance with Fig. 5 from the original MP3 audio data streams, the MPEG4 multi-channel audio data stream was not back-converted into the exact original MP3 audio data streams, but other MP3 audio data streams were generated from it, which, unlike the original audio data streams, all the reverse indicators were set to 0, and the bit rate to the highest value. The data blocks of these newly created MP3 audio data streams are therefore also closed in themselves insofar as all data that is assigned to a specific time stamp are contained in one and the same block 74'a-74'c , and if padding data has been used to extend the length of the data blocks to a uniform value.
[0069] Fig. 8 shows a method of performing a method in which it is possible to backconvert MPEG-4 audio data streams resulting from the embodiments of Figs. 1-5 into the original MP3 or audio streams again. MP3 primary audio data stream.
[0070] In this case, the MP3 reconstructor 100 checks again in step 150, exactly as in step 110, whether the MPEG-4 audio data stream is a reformatted MP3 audio data stream. Also, the subsequent steps 152 and 154 correspond to steps 112 and 114 of the procedure of Figure 7.
[0071] Instead of changing backward indicators in side information and bit rate in frame headers, however, the MP3 reconstructor 100 reconstructs, according to the method of Fig. 8, in step 156 the original length of the data block in the primary audio data streams
MP3s that have been converted to an MPEG-4 audio data stream based on sample rate, bit rate, and fill bits. Sampling frequency and information about the filling are given in the parameter "MPEG_1_2_SpecificConfig" and the bit rate in the channel element, if the latter differs for each frame.
[0072] The formula for calculating the original length of the original frame and the MP3 audio data stream to be reconstructed is again in the form as already indicated in the previous case:
For MPEG 1 layer 3:
Frame length [bit] = 1152 * bit rate [bit / s] / sampling frequency [bit / s] + + 8 * fill bit [bit]
For MPEG 2 layer 3:
Frame length [bit] = 576 * bit rate [bit / s] / sample rate [bit / s] + + 8 * fill bit [bit] [0073] Then the MP3 audio stream data. MP3 audio data streams are generated due to the fact that the given frame headers from a given channel are placed at an interval of the calculated length of the data block, and the spaces between them are filled by inserting audio data or master data on items indicated by indicators in side information. Unlike the embodiments of Fig. 7 or. 5, here, therefore, the master data associated with a given header or given side information is inserted at the beginning of the position determined by means of the reverse pointer in the MP3 audio data stream. Or, to put it differently, the start of dynamic master data is shifted according to the value from the "main_data_begin" parameter. The MPEG-4 file header is skipped. The resulting MP3 or audio audio data stream the resulting MP3 audio data streams correspond to the original MP3 audio data streams that formed the basis of the MPEG-4 audio data stream. Also, these MP3 audio data streams could therefore be decoded, as could the audio data streams of Fig. 7, through typical MP3 decoders for audio signals.
[0074] With reference to the preceding description, it should be pointed out that in some places described as single-channel MP3 audio data streams, MP3 audio data streams were in fact two-channel MP3 audio data streams that were defined in accordance with ISO / IEC 13818-3, however, this specification has not been analyzed in more detail since it does not change anything regarding the understanding of the present invention. Matrix operations from transmitted channels to recover input channels on the decoder side and the use of multiple backward indicators for these multi-channel signals are therefore not explained, but are referred to the appropriate standard.
[0075] Previous embodiments have enabled the saving of MP3 data blocks in an altered form in the form of the MPEG-4 file format. MPEG-1/2 audio layer 3, MP3 for short, and proprietary formats derived from them, such as MPEG2.5 or mp3PRO can be packed based on these methods into an MPEG-4 file, therefore it represents a new form of multi-channel presentation of any number of channels in an easy way. It is not necessary to use a complicated and not very common procedure in accordance with ISO / IEC 13818-3. In particular, MP3 data blocks are packed in such a way that each block - a channel element or access unit - belongs to one defined time stamp.
[0076] In the preceding embodiments, in order to change the digital signal presentation format, the presentation parts were overwritten with other data. In other words, the constant information necessary or useful for the decoder for different blocks of a portion of an MP3 data block is stored within the data stream.
[0077] Following the packaging of multiple blocks of mono or stereo data in one access unit of the MPEG-4 file format, it was also possible to obtain a multi-channel presentation that is much simpler to use than a presentation compliant with the ISO / IEC-13818-3 standard.
[0078] In the foregoing embodiments, the MP3 data block presentation has been formatted in a changed manner so that all data that belongs to a specific time stamp is also included within one access unit. This is not the case for general MP3 data blocks because the "main_data_begin" or the backward indicator in the original MP3 data block can refer to earlier in time data blocks.
[0079] A reconstruction of the original data stream was also possible (Fig. 8). This meant, as shown, that the reconstructed data streams could be processed by any compatible decoder.
[0080] Furthermore, the above-described embodiments have enabled the encoding and decoding of more than two channels. Subsequently, in the above embodiments, the ready-made, encoded MP3 data required only simple formatting operations to obtain a multi-channel format. On the other side, on the encoder side, all that was required was to undo this operation or these operations.
[0081] While the MP3 data stream usually contains data blocks of unequal length, because dynamic data that belongs to one block can be compressed into preceding blocks, the above examples included dynamic data immediately after the side information. The resulting MPEG-4 data stream then had a constant average bit rate but blocks of data of varying length. Element "main_data_begin" or the backward indicator is transmitted with them unchanged to guarantee the reproduction of the original data stream.
[0082] Furthermore, referring to Fig. 5 describes the extension of the MPEG-4 syntax that allows you to pack multiple blocks of MP3 data as MP3 channel elements into a multi-channel format within an MPEG-4 file. All entries of MP3 channel elements that belong to a given time point have been packed into one access unit. According to the MPEG-4 standard, information on the encoder side can be downloaded for configuration from the so-called "AudioSpecificConfig" header. It contains, in addition to the "audioObjectType" parameter, sample rate and channel configuration, etc., also a descriptor that is necessary for the given "audioObjectType" parameter. This descriptor has been described above with reference to "MPEG_1_2_SpecificConfig".
[0083] In accordance with the preceding embodiments, the sync word "12-Bit-MPEG1 / 2-syncword" in the header has been replaced by the length of a given MP3 channel element. According to ISO / IEC-13818-3, 12 bits are sufficient for this. The rest of the header was not further modified, which, however, can of course be done to, for example, shorten the header of the frame and the remaining redundant part, except for the word synchronization, and thus reduce the amount of information transmitted.
[0084] Various changes can be made without problems with respect to the above embodiments. Accordingly, the order can be changed in stages in Figs. 3, 7, 8, and in particular stages 42, 50, 56, 60 in Figs. 3, 11, 114, 118, 122 and 128 in Figs. 7 and 152, 154 , 156 in Fig. 8.
[0085] With reference to Figs. 3, 7, 8, it should further be pointed out that the stages represented therein are carried out by means of the appropriate properties in the converter or in the reconstructor in Fig. 2 or 6, which may for example be made in the form of a computer or a permanently wired connection system.
[0086] For the embodiments of Fig. 7 manipulation of headers or secondary information (steps (118, 122) in the form of an MP3 data stream slightly changed relative to the original MP3 data stream for the MP3 decoder takes place on the receiver side or decoder. In many applications, it may be advantageous to perform these steps on the encoder side or transmitter, because the receiving devices are most often mass production items, therefore the savings in electronics on the receiver side allow a much higher profit. According to an alternative embodiment, it can therefore be provided that these steps are already carried out during the conversion of the data format from MP3 to MPEG4. Stages compatible with this alternative method of converting the formats are shown in Fig. 9, wherein the steps that are identical to the steps of Fig. 3 they have the same designation and will not be described again to avoid repetition.
[0087] First, the MP3 audio data stream to be converted is received in step 40, and in step 42 audio data that belong to one time stamp or represent the coding of a time interval encoded by an MP3 audio data stream audio signal belonging to a given time stamp are combined into an associated block, namely with respect to all time stamps. Headers are reattached to related blocks to obtain channel elements (step 50). However, headers are modified not only, as in step 56, by replacing the sync word by the length of a given channel element. In steps 180 and 182 corresponding to steps 118 and 122 in Fig. 7 rather, further modifications follow. Namely, in step 180, the indicators in the side information of each channel element are set to zero, and in step 182, the bit rate in the header of each channel element is changed in such a way that, as described above, the bit-length dependent MP3 data block is sufficient, to include all audio data for this channel element or the assigned time stamp along with the header size and side information. Step 182 further includes, if desired, the displacement of the filling bits in the headers of the successive channel elements so that, later, in the case of the feed produced as an effect of the method of Fig. 9 an MPEG-4 audio data stream to the one operating in accordance with the method of FIG. 7, but without steps 118 and 122, the decoder get an accurate bit rate. Of course, filling can also be performed on the decoder side as part of step 128.
[0088] For step 182, it may be cost-effective to set the bit rate not to the highest possible value as described in relation to step 122. This value can also be set to the minimum value, which is sufficient to fit all audio data, header and side information of the channel element in the calculated length of the MP3 frame, which may also mean that in this case a short bit rate containing a small number of factors will be reduced encodable fragments of the encoded audio track.
[0089] After performing these modifications in steps 60 and 62, only the file header ("AudioSpecificConfig") is generated and it is transmitted along with the MP3 channel elements as an MPEG-4 audio data stream. This stream can, as already mentioned, be reproduced according to the method of Fig. 7, however, steps 118 and 122 can be omitted, which facilitates implementation on the decoder side. Steps 42, 50, 56, 180, 182 and 60 can of course be carried out in any order.
[0090] The previous description only referred to, for example, MP3 data streams with a fixed bit length of the data block. Of course, according to the preceding embodiments, MP3 data streams with variable data block length can also be processed, in which the bit rate, and thus also the length of the data block, changes for each frame.
[0091] The previous description referred to MP3 audio data streams. In the case of other non-indicator-based audio data streams, the embodiment of the present invention provides for modification of the header in data blocks of, for example, one MPEG 1/2 layer 2 audio data stream, which in addition to the headers also includes associated side information and assigned audio data, and therefore already closed in themselves to generate an MPEG-4 audio data stream. The modification gives each header information about the length, which indicates the amount of data of either a given data block or audio data in a given data block, so that it is easier to decode the MPEG4 data stream, especially if this stream is combined from multiple MPEG 1 / 2 layers 2 to form a multi-channel audio data stream similar to the preceding description with reference to Fig. 5. The modification is preferably obtained in a manner similar to the method previously described by replacing the synchronization words or other redundant portions of the stream in the headers of the MPEG 1/2 layer 2 data stream with length information. The preceding fig. 5 change of indicator format or its solution by connecting audio data belonging to one time stamp is unnecessary for layer 2 data streams, because there are no back indicators. The decoding of two MPEG 1/2 layer audio streams that represent two channels of a multi-channel audio data stream, the MPEG-4 audio data stream is simple because the length information is read and on its basis you can quickly access individual channel elements in access units. These can then be forwarded to typical decoders, compliant with MPEG 1/2 layer 2.
[0092] Furthermore, with respect to the present invention, it is not important exactly where the back pointer is located in data blocks based on the indicators of the audio data stream.
It could also be directly in the frame headers to define the associated identification block.
[0093] In particular, it should be pointed out that, depending on the circumstances, the inventive scheme may be implemented to convert the file format also to software. The implementation can be carried out on a digital storage medium, especially on a floppy disk or CD with electronically readable control signals that can work with a programmable computer system to perform the appropriate method. In general, the invention therefore also consists in a computer program product with a program code stored on the machine-readable medium for carrying out the method according to the invention if the computer program product is run on a computer. In other words, the invention can thus be implemented as a computer program with program code for performing the method if the computer program is run on the computer.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany
Proxy:
EP 1 647 010 B1 Z-16441/17
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
30 members in 17 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 10333071 | Germany | A | |
| 10339498 | Germany | A | |
| 04763200 | European Patent Office (EPO) | A | |
| 2004007744 | European Patent Office (EPO) | W | |
| 047632005 | – | – | – |
| 10333071 | – | – | – |
| 10339498 | – | – | – |
| DE2003133071 | – | – | – |
| DE2003139498 | – | – | – |
| EP20040763200 | – | – | – |
| WO2004EP07744 | – | – | – |
Members30
| Document | Office | Kind | |
|---|---|---|---|
| AU2004301746A1 | Australia | A1 | |
| CA2533056A1 | Canada | A1 | |
| WO2005013491A2 | World Intellectual Property Organization (WIPO) | A2 | |
| DE10339498A1 | Germany | A1 | |
| WO2005013491A3 | World Intellectual Property Organization (WIPO) | A3 | |
| MXPA06000750A | Mexico | A | |
| DE10339498B4 | Germany | B4 | |
| EP1647010A2 | European Patent Office (EPO) | A2 | |
| NO20060814L | Norway | L | |
| KR20060052854A | Republic of Korea | A | |
| IL173223A0 | Israel | A0 | |
| RU2006105203A | Russian Federation | A | |
| CN1826635A | China | A | |
| BRPI0412889A | Brazil | A | |
| US2006259168A1 | United States of America | A1 | |
| JP2006528368A | Japan | A | |
| KR100717600B1 | Republic of Korea | B1 | |
| AU2004301746B2 | Australia | B2 | |
| RU2335022C2 | Russian Federation | C2 | |
| JP4405510B2 | Japan | B2 | |
| US7769477B2 | United States of America | B2 | |
| CN1826635B | China | B | |
| IL173223A | Israel | A | |
| CA2533056C | Canada | C | |
| NO334901B1 | Norway | B1 | |
| EP1647010B1 | European Patent Office (EPO) | B1 | |
| PT1647010T | Portugal | T | |
| ES2649728T3 | Spain | T3 | |
| PL1647010T3This record | Poland | T3 | |
| BRPI0412889B1 | Brazil | B1 |
Numbers
- Publication
- 1647010
- Publication, DOCDB
- 1647010
- Publication, EPODOC
- PL1647010T
- Application
- 4763200
- Application, DOCDB
- 04763200
- Application, EPODOC
- PL20040763200T
Titles2
- English
- AUDIO FILE FORMAT CONVERSION
- Polish
- Sposób konwersji formatu pliku audio
Classification
- CPC, 1
- G10L19/173
- IPC, 1
- G10L19 16