Process for transmitting and/or storing digital signals of multiple channels
Abstract
Disclosed is a process for transmitting and/or storingdigital signals of multiple channels. This process issuited, in particular, for transmitting the five channelsof 3/2 stereophony as well as for transmitting two stereochannels and three additional commentary channels. In thismanner, by way of illustration, television programms withmulti-language audio signals can be transmitted. This process is distinguished that by reduction of the to-be-transmitted data, only a bit rate of 384 kbit/s isrequired for transmission. The reduction of the data is achieved by the K inputchannels being imaged in segments onto the N<=K virtualspectral data channels, by the spectral data channelsbeing quantized, coded and transmitted taking intoconsideration the principles of psychoacoustics and by Koutput channels being reproduced from the transmitted bitstream with the aid of a also transmitted list from the N<=K spectral data channels.

Term
Term ended
Expired 2 November 2013, 12.9 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
11 claims: 1 independent, 10 dependent
- 1CA 02148447 2003-12-03 - 12 What is claimed is:1. A process for transmitting or storing digital signals from K input channels, in which sampling values of signals from a time domain are transmitted in blocks into a frequency domain as spectral values, and said spectral values are coded and combined into a bit stream which is transmitted or stored and subsequently decoded and transmitted back into K output channels in the time domain, wherein a variable number of transmitted spectral data channels (TSC) are formed in segments during coding in dependence of the spectral values of the blocks of individual K input channels, with the number of transmitted spectral channels (NTSC) of the spectral data channels per segment being less than or the same as the number K of input and output channels, said NTSC and the structure of said spectral data channels being included in the bit stream as information (SEGMENT_DATA) and following transmission of said bit stream, said K output channels are combined in a decoder in segments from the transmitted spectral data channels (TSC) with the aid of the information ( SEGMENT_DATA).
- 8A process according to any one of claims 1 to 7, wherein, in said coding of spectral values, a coding is CA 02148447 2003-12-03 - 14 utilized in which groups of spectral values scaled by a common scale factor are combined (scale factor band) and spectral values belonging to scale factor bands not used in any information block (SEGMENT_ INFO) are not transmitted.
- 10A process according to any one of claims 1 to 9, wherein said coding of said input channels and imaging of said input channels onto said spectral data channels (TSC) occurs in such a manner that for reconstructing said output channels, linear combinations of reconstructed spectral values are formed from different spectral data channels.
Independent claims8
92 paragraphs in 8 sections, as filed
CA 02148447 2003-12-03
Process for Transmitting and/or Storing Digital Signals of Multiple Channels
Technical Field
The present invention relates to a process for transmitting and/or storing digital signals of multiple channels.
State of the Art
Processes in which digital signals, in particular audio signals, are transmitted frequency-coded are known, by way of illustration, from the PCT publications WO 88/01811 and WO 89/08357.
For one-channel and two-channel transmission, the standardization committee, Moving Pictures Experts Group (MPEG), of the International Standardization Organization (ISO) set the standard ISO-11172-3 for coding and the tobe-transmitted bit stream of audio signals.
Psycho-acoustic models which permit reducing the amount of to-be-transmitted data while exploiting the properties of human hearing with minimal quality loss are used in the mentioned coding process.
For explanation of all terms not made more apparent herein reference is explicitly made to the publications and the standard.
-2In further developing international standards, work is presently being done, among other things, in reducing the data in multi-channel transmission. The scientific publication MUSICAM-Surround: A Universal Multi-Channel Coding System Compatible with ISO 11172-3, 93rd AES convention, 1992, San Francisco, proposes a process for transmitting of up to 5 channels. By way of illustration, two stereo channels and one center channel as well as two side channels (3/2 stereophony) or two stereo channels and three commentary channels can be transmitted.
Further data reduction is achieved in that parts of the stereo signals, which are insignificant with regard to spatial perception, are transmitted in only one channel.
In addition transmitted are scale factors, which represent a measure of the intensity of the signals conducted from the mono-channel to the corresponding loud speakers. With this process, artefacts with lower audio pleasure are generated especially in the lower frequency range.
In addition to this, it has been proposed to reduce the to-be-transmitted amount of data by not determining a socalled intra-channel masking threshold for each channel for the coding, but rather by providing a common threshold for all the channels, taking into consideration the intra-channel masking effect. However, the use of a common masking threshold results in that interfering coding noises may be perceptable in the vicinity of a loud-speaker.
CA 02148447 2003-12-03
- 3 A verbal signal coding device which generates a number of values corresponding to the parts of the frequency spectrum of these verbal signals and subsequently codes them is known from EP 0 176 243 A2. The coded values are then converted by means of a bit converter according to .their energy equivalents, with the number of values generated and converted into bits remaining constant, however, the selection of values with varying bit allocations being variable.
Description of the Invention
The object of the present invention is to provide a process for transmitting and/or storing digital signals of multiple channels permitting further reduction of the amount of to-be-transmitted data and resulting in no subjectively perceptible disturbance of transmitted signals .
This object is solved by a process for transmitting or storing digital signals from K input channels, in which sampling values of signals from a time domain are transmitted in blocks into a frequency domain as spectral values, and said spectral values are coded and combined into a bit stream which is transmitted or stored and subsequently decoded and transmitted back into K output channels in the time domain, wherein a variable number of transmitted spectral data channels (TSC) are formed in segments during coding in dependence of the spectral values of the blocks of individual K input channels, with the number of transmitted spectral channels (NTSC) of the spectral data channels per segment being less than or the same as the number K of input and output channels, said
CA 02148447 2003-12-03
- 3a NTSC and the structure of said spectral data channels being included in the bit stream as information (SEGMENT_DATA) and following transmission of said bit stream, said K output channels are combined in a decoder in segments from the transmitted spectral data channels (TSC) with the aid of the information (SEGMENT__DATA) .
According to the present invention, the signals of the different channels are first converted into spectral values. Subsequently, using the spectral values of the corresponding segments of the different channels it is determined in which channels similar spectral portions occur.
Investigated is whether in combining the segments from the different channels, the interferences caused by the joint coding are below the audibility threshold or if artefacts are generated. If no artefacts are expected, combination is permitted. In this manner, the K input channels are imaged in segments on the spectral data channels (transmitted spectral channels).
The term spectral data channel is understood to be all spectral coded audio dara including the respective supplementary information required in order to transform a coded audio signal back into its entire or parts of its signal spectrum in the time range.
The more different the original input channels are, the more different spectral data channels have to be used in order to transmit the signal information. If, in an extreme case, all the channels in a segment are practically the same, a single spectral data channel may
CA 02148447 2003-12-03
- 3b suffice, in particular for the upper part of the spectrum. The amplitudes of the spectral segments can be controlled by means of the respective scale factors.
The number of NTSC (number of transmitted spectral channels) of the required spectral data channels is variable and may be less than or at most the same number K of the input respectively output channels.
CA 02148447 2003-12-03
-4In order to be able to control the combination of the Koutput channels from the NTSC spectral data channels in the decoder, a list of information data (SEGMENTJDATA) is transmitted in addition to the (reduced) signal data. This list describes how the spectral values of the output channels are combined from the sides of the spectral data channels.
Advantageous embodiments and further developments of the present invention are set forth below.
The control commands required to reconstruct a segment of an output channel are combined into an information block (SEGMENT^INFO) . This block contains fields for the length of the segment (SEG_LENGTH), for the selection of the spectral data channels (TSC_SELECT) and for the scale factors (scf). The coded spectral data (TSC_DATA) of a specific spectral data channel (TSC__NUM) are decoded in the decoder with the respective scale factors (sfc) determining the reconstruction matrix.
The information for reconstructing the segments ôf an output signal are lined up to form a list (SEGMENT_LIST).
The lists of segments of the individual output channels form the global list (SEGMENT__DATA) . Thus, the lists for a left channel (LEFT_CHANNEL), a right channel (RIGHT__CHANNEL) , a center channel (CENTER_CHANNEL) and further channels are listed in the global list.
According to an advantageous embodiment of the present invention, the channels form a
CA 02148447 2003-12-03
-5multi-channel tone. By way of illustration, the 5 channels of the 3/2 stereophony are transmitted. In addition to the two stereo channels, a center channel, a left side channel (LS_CHANNEL) and a right side channel (RS_CHANNEL) are transmitted. The transmitted channels are given following reverse transformation into the time range to the corresponding five loudspeakers of 3/2 stereophony.
In another embodiment, in addition to the conventional two stereo channels, several additional commentary channels are transmitted. In these channels, by way of illustration, in HDTV (high definition television) the audio signals can be transmitted in different languages. The viewer can then select the desired language for the given television picture. The language can be dynamically added to the two stereo channels.
The transmission of the spectral data channels occurs with a process that is compatible with the standard ISO 11172-3. All three layers of this standard can form the basis of the transmission of the spectral data channels.
During coding, groups of spectral values which are scaled by means of a common scale factor, thus belong to a scale factor band, are not transmitted if this scale factor band is not required in any information block (SEGMENT_INFO). In this event, the respective spectral values are not needed for reconstructing the output channels. Advantageous is if the fixed scale-band division of the ISO standard is utilized.
The signalization of the unused scale factor bands occurs implicitly, that means that no addiCA 02148447 2003-12-03
-6tional information has to be transmitted to this signalization.
Linear combinations of the to-betransmitted channels can also be formed with the invented process. This leads to a further reduction of the amount of data if the spectral values strongly resemble each other in the individual channels. The signals of the output channels are formed in this event by linear combinations of the reconstructed spectral values.
The bit rate required for transmitting the coded data from ail the channels does not exceed 384 kbit/s. In this way, the demands made on the maximum bit rate of layer III of the ISO standard are met.
The essential advantages of the present invention lie in that a distinct reduction of the to-be-transmitted amount of data is achieved by the imaging of the input channels on a small number of virtual spectral channels without any audible loss in quality. In this way, the audio signals are transmitted with especially high spatial resolution. This is of particular advantage in large rooms containing a large audience.
In addition, several channels can be made available to the individual users from which they can select the desired information. By way of illustration, the audio channels can be transmitted simultaneously for a television program in several languages of which the viewer can select the desired one.
The technically more complicated process steps for realizing the process are undertaken in the encoder providing ο
the to-be-transmitted bit stream. The decoder only processes the information of the arriving bit stream successively and is constructed substantially simpler than the encoder. The invented process, therefore, requires a higher degree of complexity in the few encoders whereas the decoders required in greater numbers for the user hardly increase in complexity.
Description of a Preferred Embodiment
The present invention is made more apparent in the following using a preferred embodiment.
The preferred embodiment is based on the standard ISO 11172-3 of the Moving Picture Expert Group (MPEG) of the International Organization for Standardization. This standard is referred to hereinafter as MPEG-1. The concept is restricted to layer III of the standard without the intention of limiting the overall inventive idea. Like in standard MPEG-1 / layer III, mathematical operations are used similar to program language C.
With the aid of filter banks, the signals of the 5 input channels from the time range are imaged in the frequency range, in subbands (Sb), the input signals being decomposed into undersampled spectral values.
Exploiting the regularities known in psychoacoustics, a computation is made as to which segments of the different channels can be combined without generating artefacts that lie above the audibility threshold.
The frequency value of the input channels are quantized and coded individually or as linear combinations. This occurs on the condition that the errors resulting from
-8the quantization lie below the audibility threshold.
Subsequently, the to-be-transmitted bit stream is combined. Xt contains the quantized and coded frequency values of the spectral data channels as well as supplementary information. These consist of scale factors, bit allocation, information of the tables and other parameters used in the current block and the lists indicating how the decoded frequency values of the spectral data channels are combined for reconstructing the output channels.
These lists are attached to the bit stream according to MPEG 1 as supplementary information.
This extension (MPEG2_extension_data) is displayed in the following:
MPEG2_extension_data() signalling_byte () :
NTSC;
SEGMENT-DATA ();
for (i=3? i<=NTSC; i++)
TSCJ3ATA (i);
bit uimsbf ? first two channelsThe term signalling_byte indicates how many and which input respectively output channels are employed. It determines whether a mono channel, stereo channel, a center channel or auxiliary channels, etc., are transmitted.
NTSC indicates the number of required spectral data
-92148447 channels .
SEGMENT_DATA describes the list for the reconstructing of the output channels and contains the lists for the reconstructions of the individual channels (SEGEMENT_LIST).
SEGMENTJ3ATA ( ) for (i-0; i<NTSC; 1++ , reset used_sb-map for (sb=0; sb<21; sb++) used_sb-map i sb - 0;
SEGMENT_LIST ( LEET_CHANNEL ) ;
SEGMENT_LIST(RIGHT<sub>—</sub>CHANNEL) ;
if (center_on)
SEGMENT_LIST(CENTER_CHANNEL) ; if ( stereo_surround) ; stereo surround
SEGMENT_LIST(LS_CHANNEL);
SEGMENT_LIST(RS_CHANNEL) ;
if (mono<sub>—</sub>surround) ; mono surround
SEGMENT_LIST ( MSJZHANNEL ) ;
for (i=0; i<no_of<sub>—</sub>coramentary_chan; i++)
SEGMENT-LIST(COM_CHANNEL i ) ;
The function used_sb_map indicates whether a scale factor band which contains the spectral values of a spectral data channel which is scaled by a common scale factor is used.
If such a scale factor band is not employed by any information block ( SEGMENT_INFO ), the corresponding spectral values are not transmitted.
-10The list for reconstruction of a specific channel (SEGMENT_LXST) contains information about the segment length(SEG__LENGTH), the size of the scale factors scalefac_size), the scale factors (scf) and the selection of the spectral data channel (TSC_SELECT).
SEGMEHT__LIST () sb - 0;
for (i—0 ; 1; i++)
SEGJCENGTH i ; if (SEG_LENGTH i <sup>: </sup>TSC_SELECT i ; if (SEG_LENGTH i == sign = +1? len else if (SEG_LENGTH sign = -1? len else sign = 0? len if (TSC_SELECT i scalefac_size; for (1=0; Klen?
scf i sb+1 ; used_sb_map TSi sb+1 = 1;
<sup>:</sup>= 0) break;
15) = SEG_LENGTH i == 14) = SEG_LENGTH = SEGJLENGTH ;»0)
1++)
SELECT bit uimsbf bit uimsbf i-1 ;
i-i ;
i ?
; 4 bit, bslbf ; 0..4 bits ; mark uses SBs if (Isign) sb += len;
In the preferred embodiment, the values 14 and 15 are re* served for forming the linear combinations of reconstructed spectral values. In the case of SEG_LENGTH == 15, is
-112148447 added and in the case of SEG_LENGTH == 14, is substracted.
The bit stream for the data from the spectral data channel corresponds to the bit stream of the main data in the case MPEG-l/layer III and reads:
TSC_DATA( TSCJIUM) part2_3_length? scalefac_compress; global_gain? block_type; big_values? table_select 3 ;
countltable_select ? region_count 2 ;
see MPEG-1/audio for (sb=0; sb<21>; sb++) if (used_sb_map TSC_NUM sb
Huffmancodesection TSC NUM sb
The used symbols stand for:
++ increase == same = allocation operator logical not bslbf bit string, left bit first ch channel sb subband uimsbf unsigned integer, most significant bit first
Contents8
1 sheet
Sheet 1
21 members in 11 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 4236989 | Germany | A | |
| 4236989 | Germany | A | |
| P42369894 | Germany | – | |
| 9301047 | Germany | W | |
| 9301047 | Germany | W | |
| DE19924236989 | – | – | – |
| P42369894 | – | – | – |
| PCTDE93001047 | – | – | – |
| WO1993DE01047 | – | – | – |
Members21
| Document | Office | Kind | |
|---|---|---|---|
| DE4236989A1 | Germany | A1 | |
| CA2148447A1 | Canada | A1 | |
| WO9410758A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU5333294A | Australia | A | |
| DE4236989C2 | Germany | C2 | |
| NO951549D0 | Norway | D0 | |
| NO951549L | Norway | L | |
| EP0667063A1 | European Patent Office (EPO) | A1 | |
| KR950704862A | Republic of Korea | A | |
| JPH08505739A | Japan | A | |
| EP0667063B1 | European Patent Office (EPO) | B1 | |
| AT143544T | Austria | T | |
| DE59304003D1 | Germany | D1 | |
| AU678407B2 | Australia | B2 | |
| US5706309A | United States of America | A | |
| RU2129336C1 | Russian Federation | C1 | |
| NO309629B1 | Norway | B1 | |
| KR100311604B1 | Republic of Korea | B1 | |
| CA2148447CThis record | Canada | C | |
| EP0667063B2 | European Patent Office (EPO) | B2 | |
| JP3792250B2 | Japan | B2 |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| ExpiryMKEX | MKEX | |
| ExpiryMKEX | MKEX | |
| Examination requestEEER | EEER |
Numbers
- Publication
- 2148447
- Publication, DOCDB
- 2148447
- Publication, EPODOC
- CA2148447
- Application
- 2148447
- Application, DOCDB
- 2148447
- Application, EPODOC
- CA19932148447
Titles2
- English
- PROCESS FOR TRANSMITTING AND/OR STORING DIGITAL SIGNALS OF MULTIPLE CHANNELS
- French
- METHODE DE TRANSMISSION SIGNAUX NUMERIQUES VIA PLUSIEURS CANAUX ET/OU DE STOCKAGE DE CES SIGNAUX
Classification
- CPC, 8
- H04H20/88
- H04B1/66
- G11B20/00007
- G11B20/00992
- H04N7/06
- H04N21/2368
- H04N21/4341
- H04N21/8106
- IPC, 12
- G11B20 10
- H04H20 88
- G11B20 00
- G11B20 18
- H04B1 66
- H04B14 00
- H04N7 06
- H04H1 00
- H04H5 00
- H04H20 18
- H04N21 2365
- H04N21 434