Images i.e. video, displaying method for e.g. videoconferencing application, involves controlling display of image set comprising delay image, where image of set is displayed during display duration lower than predefined duration
Abstract
The method involves forecasting availability delay of data, among data packets, defining a delay image, where the packets are transmitted via a communication network (100). Display of a set of images carried out with respect to the forecasting of the delay is controlled, where an image of the set of images is displayed during display duration higher than predefined display duration. Display of another set of images comprising the delay image is controlled, where an image of the latter set of images is displayed during another display duration lower than the predefined display duration. Independent claims are also included for the following: (1) a video display device comprising reception units (2) a computer program comprising instructions to perform a method for displaying images on a video display device.

Term
Projected expiry 7 August 2028.
- Priority and filed
- Published
- Today
- Projected expiry
23 claims: 2 independent, 21 dependent
- 1CLAIMS REVENDICATIONS 1. A method of displaying a plurality of images on a video display device receiving via a communication network data representing images to be displayed on a video screen for a predefined display time, characterized in that the method successively comprises a step of forecasting a data availability delay defining an image called a late image, a step of controlling the display of a first plurality of images performed as a function of said prediction, at least one image of the first plurality being displayed for a display duration greater than the predefined display duration, a step for controlling the display of a second plurality of images comprising the late image, at least one image of the second plurality being displayed for a display duration less than the predefined duration. 1. Procédé d’affichage d’une pluralité d’images sur un dispositif d’affichage vidéo recevant via un réseau de communication des données représentant des images destinées à être affichées sur un écran vidéo pendant une durée d’affichage prédéfinie, caractérisé en ce que le procédé comprend successivement une étape de prévision d’un retard de disponibilité de données définissant une image dite image en retard, une étape de commande de l’affichage d’une première pluralité d’images effectuée en fonction de ladite prévision, au moins une image de la première pluralité étant affichée pendant une durée d’affichage supérieure à la durée d’affichage prédéfinie, une étape de commande de l’affichage d’une deuxième pluralité d’images comprenant l’image en retard, au moins une image de la deuxième pluralité étant affichée pendant une durée d’affichage inférieure à la durée prédéfinie.
- 22Video display device comprising means for receiving via a communication network data representing images intended to be displayed on a video screen for a predefined display duration, characterized in that the device comprises means for forecasting a data delay defining an image called a late image, means for controlling the display of a first plurality of images, said control being performed as a function of said forecast, at least one image of the first plurality being displayed for a display duration greater than the predefined display duration, means for controlling the display of a second plurality of images comprising the late image, at least an image of the second plurality being displayed for a display duration less than the predefined duration. 22. Dispositif d’affichage vidéo comprenant des moyens de réception via un réseau de communication de données représentant des images destinées à être affichées sur un écran vidéo pendant une durée d’affichage prédéfinie, caractérisé en ce que le dispositif comprend des moyens de prévision d’un retard de données définissant une image dite image en retard, des moyens de commande de l’affichage d’une première pluralité d’images, ladite commande étant effectuée en fonction de ladite prévision, au moins une image de la première pluralité étant affichée pendant une durée d’affichage supérieure à la durée d’affichage prédéfinie, des moyens de commande de l’affichage d’une deuxième pluralité d’images comprenant l’image en retard, au moins une image de la deuxième pluralité étant affichée pendant une durée d’affichage inférieure à la durée prédéfinie.
Independent claims2
245 paragraphs, as filed
The present invention relates to a video display method which can be used, for example, in a context of live digital multimedia transmission (in real time - “live streaming”) such as videoconferencing or video surveillance, using a communication network such as the Internet. or a local network, for example of the IP (“Internet Protocol”) type.
The invention also relates to an image display device which can be used, for example, for video conferencing or video surveillance applications.
In a live multimedia transmission context, multimedia data is captured by a server using a camera and a microphone. For example, the camera captures an image every 20 ms, in the case of video at 50 frames per second, and the microphone records audio continuously.
The data is then encoded in a digital compression format before being transmitted over a data communications network to one or more receivers, also called clients, having a video display and audio reproduction function.
These customers receive and consume this data as it is received. In particular, the client reproduces the content encoded by the video data and the audio data as it receives them.
Multimedia data (images, portions of images - “slices” - or portions of sounds - “samples” -) have a limited useful life in such a context of broadcasting in real time. They must imperatively be received and processed by the customer before the moment when, depending on the application and the standards of use applicable to it, it is considered essential that the data be disseminated by the customer for the users.
Beyond this deadline, the data becomes useless and is generally ignored by the customer. The rendering of the information is then degraded, since part of the data has not been broadcast.
It is added that current network technologies are such that data loss occurs quite frequently between the server and the client. As long as these losses are limited, they can be in a certain way and under certain conditions corrected or masked. As these losses increase, the quality of rendering, especially video rendering, weakens to levels that may be considered unsatisfactory in some applications.
To ensure good quality audio-video rendering, it is therefore considered at present that a real-time broadcasting system must simultaneously meet the following four conditions.
First of all, the time between the capture of the audio or video data and its rendering must be less than 150 ms.
Second, the rendering of audio data and rendering of video data must be synchronous. It is considered that an acceptable level of synchronization is achieved if the audio data is broadcast less than 40 msec ahead or behind the video data.
Third, the time between the display of two successive images must be regular. On this subject, refer to the article IMPACT OF JITTER AND JERKINESS ON PERCEIVED VIDEO QUALITY by Quan Huynh-Thu and Mohammed Ghanbari (Proceedings of the Second International Workshop on Video Processing and Quality Metrics for Consumer Electronics VPQM-06).
Finally, fourth, the number of packet losses in the network must be limited, and the impact of packet losses must be controlled.
In this context, the Applicant is interested in situations where an image reaches the client with a significant delay in relation to the display rate. Such situations arise due to a delay in encoding the image and its transmission over the network, or due to a delay in the transmission of information over the network.
Conventionally, in such a case, the image preceding the late image is displayed for a duration greater than the average duration, then, when the late image is finally available, the latter is displayed. Finally, the display of the images following the late image is precipitated so as to make up for the delay taken by the display. The viewer therefore has the impression of freezing the image and then of abrupt acceleration. The quality of the visualization is therefore greatly degraded. The audio and video can be severely out of sync, and the video is no longer rendered live.
From the document US Pat. No. 6262776 “System and method for maintaining synchronization between audio and video” of Microsoft is known a system which adjusts the display instant of the images in order to try to maintain good synchronization between the audio and the video.
This system can in particular decide to skip an image and go directly to the next image. It also takes into account changes in the customer's calculation speed.
We also know from the article “Adaptive delay and synchronization control for Wi-fi based AV conferencing”, Proceedings of the First International Conférence on Quality of Service in Heterogeneous Wired / Wireless Networks (QSHINE'04) - Haining Liu and Magda El Zarki , a receiver which receives an audio / video stream transmitted over a Wifi wireless network and displays it synchronously.
The network considered in this document introduces variable delays depending on the conditions, for example because of radio interference or other communications.
The system adapts the display rate by increasing the display delay to allow better reception of the packets. We understand that this is done to the detriment of interactivity.
Finally, document US 7170545 “Method and apparatus for inserting variable audio delay to minimize latency in video conferencing” from Polycom Inc. is known to provide a receiver which receives an audio and video stream and displays it after having analyzed the content of the stream for determine if this is a monologue or dialogue situation.
In the first case, good synchronization between audio and video is desired, while in the second case, a low delay in rendering is desired so as to guarantee the interactivity of the application. Variable speeds are then used for audio rendering.
It can be seen that the known solutions are not completely satisfactory, in particular for applications where the requirements of the users in terms of quality of rendering and of transmission in real time are important.
The applicant's invention therefore seeks to limit the visual impact of the variable processing times in these contexts.
To this end, there is provided a method of displaying a plurality of images on a video display device receiving via a communication network data representing images intended to be displayed on a video screen for a period of time. predefined display, the method comprising successively a step of forecasting a delay in the availability of data, for example encoded data or else data having undergone decoding, and defining an image called a late image, a step of controlling the display of a first plurality of images performed as a function of said prediction, at least one image of the first plurality being displayed for a display duration greater than the predefined display duration, and a step of controlling the display of a second plurality of images comprising the late image, at least one image of the second plurality being displayed for a display duration less than the predefined duration.
This method makes it possible to improve the visual impact of the delay of the late image, thanks to the prediction of the delay and the dependence of the control of the display on this prediction.
More specifically, it is provided that the data can be received encoded via the communication network, and that the prediction of a delay in the availability of data defining an image comprises at least a prediction of a delay of reception of encoded data or a delay of. decoding of encoded data.
In other words, according to one characteristic, the data received via the communication network includes encoded data, and the prediction of a delay of decoded data may include only a prediction of a delay of reception of encoded data or at the time. both predicting a delay in receiving encoded data and predicting a delay in decoding encoded data.
It is therefore specified that alternatively the forecast of a delay of decoded data can include a forecast of a delay in the decoding of coded data received in time.
According to a first advantageous aspect, the display of the images is accompanied by a broadcast of a soundtrack intended to be broadcast by a loudspeaker associated with the video display device in a manner synchronized with the display of the images, and the The step of controlling the display of a first plurality of images is performed as a function of an acceptability criterion of desynchronization of the display of the images and of the broadcasting of the soundtrack for a given duration.
In a first case, the acceptability criterion is a criterion with two states, which are the state “criterion verified” and the state “criterion not verified”. In this case, the display of a first plurality of images is controlled if the acceptability criterion is verified. In a second case, the acceptability criterion can take a plurality of states, indicating a level of acceptability in a discrete or continuous scale at several levels. In this case, at least one parameter of the display of a first plurality of images may depend on the state of the acceptability criterion. This parameter can be the number of images of the plurality or the display time of at least one image of the plurality.
According to a general embodiment, the data comprises an independent image defined autonomously and a complementary image defined according to a previous image and the prediction of a delay in data availability includes a prediction of a delay of at least minus one datum defining an independent image.
More precisely, provision is made for an independent image to be able to include at least one Intra (Intra coded) block, within the framework of a mode of compressing the video data.
Furthermore, the method can include a step of sending via the network to the attention of a server a request to send an independent image. The forecasting step can moreover comprise a step of sending via the network to the attention of a server such a request made as a function of a request transmission relevance information (or simply information of relevance of request) determined at least on the basis of data received via the communication network.
According to an advantageous characteristic, the forecast comprises a step of determining information on the relevance of said transmission of a request.
It is specified that the determination of a piece of information on the relevance of the transmission can include at least one calculation or one reception of such information.
In particular, provision is made for the relevance information to be able to indicate at least that an error rate greater than a predefined value has been measured for an image received by the video display device via the communication network.
Said error rate may include at least an instantaneous error rate or a cumulative error rate.
In a particular embodiment, the display of the images is accompanied by a broadcast of a soundtrack intended to be broadcast by a loudspeaker associated with the video display device in a manner synchronized with the display of the images and the display. forecast includes an evaluation of an acceptability criterion of desynchronization of the display of the images and the broadcasting of the soundtrack for a given duration, the step of controlling the display of a first plurality of images being performed as a function of said desynchronization acceptability criterion.
In this context, provision is made for the relevance information to indicate at least that the criterion of acceptability of desynchronization of the display of the images and of the broadcasting of the soundtrack for a given duration is verified.
It is specified that the verification of the criterion can be determined as a function of the application made of the video display device on the basis of the content of the images and of the soundtrack.
Advantageously, the first plurality of images comprises a number of images calculated at least as a function of the criterion of acceptability of desynchronization.
Then, according to an interesting embodiment, provision has been made for the forecast to include a step of reception via the communication network of information announcing an upcoming transmission of an independent image. In particular, a second video display device receives the encoded data via the communication network and the forecast comprises a step of receiving information indicating that the second video display device has requested sending of an independent image. It is thus possible for the information announcing an upcoming transmission to be transmitted in response to a request to send an independent image by the second video display device.
It will be noted that this last aspect is combined with the use of the desynchronization acceptability criterion, and that, alternatively, it can be implemented independently.
According to also an embodiment, a plurality of independent images are sent with a given period, and said forecast comprises a step of determining a completion of an occurrence of the period.
It will be noted that this last aspect is combined with the use of the desynchronization acceptability criterion, and that, alternatively, it can be implemented independently.
Also according to a general and optional characteristic, the first plurality of images comprises a number of images calculated at least as a function of an average time between the sending of an image by a server and its reception by the display device. video.
It is also possible that during the step of displaying the first plurality of images, at least two images of the first plurality are displayed for a display duration greater than the predefined display duration.
According to one characteristic, the first plurality of images comprising at least a first, a second and a third image, the first image being displayed before the second image and the second image being displayed before the third image, the first, second and third image being displayed during display durations called the first, second and third display durations respectively, the first display duration is less than the second display duration, and the third display duration is also less than the second display duration.
There is therefore a phenomenon of smoothing the delay taken, on at least three images, or even on all the images of the plurality.
For example, during the step of displaying the first plurality of images, the first, second and third display durations are proportional to a sinusoidal function applied to an arrival number of the image and of period the duration of the stage.
In a parallel manner, but without the two aspects being necessarily implemented in the invention, it is possible that during the step of displaying the second plurality of images, at least two images of the second plurality are displayed for a display duration shorter than the preset display duration.
According to one characteristic, the second plurality of images comprising at least a first, a second and a third image, the first image being displayed before the second image and the second image being displayed before the third image, the first, second and third image being displayed during display durations called the first, second and third display durations respectively, the first display duration is greater than the second display duration, and the third display duration is also greater than the second display duration.
There is therefore a smoothing phenomenon of the acceleration imposed on the images, on at least three images, or even on all the images of the plurality.
For example, during the step of displaying the second plurality of images, the first, second and third display durations are proportional to a sinusoidal function applied to an arrival number of the image and of period the duration of the stage.
Finally, it is advantageous that a delay accumulated during the step of displaying the first plurality of images is compensated for by an advance taken during the step of displaying the second plurality of images. This can in particular be obtained by using symmetrical sinusoidal functions for the first plurality of images and for the second plurality of images.
Optionally, according to an advantageous characteristic, the step of predicting a delay comprises a step of estimating an estimated duration of the delay, and said display duration greater than the predefined display duration depends on said estimated duration of the delay. delay.
The invention also provides a. video display device comprising means for receiving via a data communication network representing images intended to be displayed on a video screen for a predefined display duration, the device comprising means for predicting a delay of decoded data defining an image called a late image, means for controlling the display of a first plurality of images, said command being performed as a function of said forecast, at least one image of the first plurality being displayed for a display duration greater than the predefined display duration, and means for controlling the display of a second plurality of images comprising the late image, at the at least one image of the second plurality being displayed for a display duration less than the predefined duration.
The invention also proposes a computer program comprising a series of instructions able, when they are executed by a microprocessor, to implement a method as presented above.
The invention will now be described in detail, with reference to the appended figures, relating to embodiments given by way of example.
FIG. 1 represents an audiovisual system comprising a client and a server to which the invention can be applied.
FIG. 2 represents an architecture of a receiver device to which the invention can be applied.
FIG. 3 and FIG. 4 represent the order and the size of the images of a transmission sequence of images encoded according to a type of video encoding to which the invention can be applied.
FIG. 5 represents different elements of a receiving device according to one embodiment of the invention.
FIG. 6 represents an aspect of the implementation of a display method according to the prior art, as a function of time.
FIG. 7 represents another aspect of the implementation of a display method according to the prior art similar to that presented in FIG. 6, also as a function of time.
FIG. 8 represents an algorithm used at the start of the implementation of a method according to an embodiment of the invention.
FIG. 9 represents an algorithm used during the implementation of a method according to an embodiment of the invention, before the display of each image.
FIG. 10 presents an aspect of the implementation of a method according to an embodiment of the invention, represented as a function of the images successively displayed.
FIG. 11 represents another aspect of the implementation of a method according to an embodiment of the invention, represented as a function of the images successively displayed, and corresponding to the implementation of FIG. 10.
FIG. 12 represents yet another aspect of the implementation of a method according to an embodiment of the invention, corresponding to the implementation of FIGS. 10 and 11, and represented as a function of time.
FIG. 13 also represents an aspect of the implementation of a method according to an embodiment of the invention, corresponding to the implementation of FIGS. 10 to 12, and represented as a function of time.
Referring to Figure 1, a client-server system to which the invention can be applied is shown. A sender or server device 101 transmits data packets of a data stream to a receiver or client device 102 via a data communication network 100.
The data communication network 100 (WAN - Wide Area Network or extended network; LAN - Local Area Network or local network) can be for example a wireless network (Wifi / 802.11a or b or g), an Ethernet network, or the Internet network.
As is often the case, the bit rate available on the network 100 is limited, for example by the presence of competing streams.
The data stream 104 provided by the server 101 includes multimedia information representing video and audio. The audio and video streams can be captured by the server 101 using a camera 105 and a microphone 106. The video and audio streams are encoded (in particular with a view to their compression) by the server 101.
For a better quality ratio for the quantity of data sent, the compression of the video can be of the motion compensation type, for example according to the H264 format or the MPEG2 format.
For each video image, the server 101 includes a capture date (timestamp) in the data packet of the image. It also includes a capture date in each audio data packet. This date is read on the same clock as for the image data.
The compressed data is cut into packets and transmitted to the client 102 by the network 100 using a communication protocol, for example the RTP protocol (real-time transport protocol or Real-time Transport Protocol), UDP (User Datagram Protocol or protocol). datagram) or DCCP (Datagram Congestion Control Protocol) or any other type of communication protocol.
In the case of the RTP protocol, for example, the capture dates are indicated in the header of each packet.
The client 102 decodes the data stream 104 received by the network 100 and reproduces the video images on a display device 107 and the audio data on a loudspeaker 108.
It is emphasized that the invention also applies to a client-server system where the same server provides information to several clients.
The dates included in the video images and the audio stream can be used by the client 102 to synchronize the video rendering and the audio rendering by matching the dates of the video display and the audio broadcast.
Feedback 109 is sometimes communicated from the client 102 to the server 101. It can make it possible to provide the server with feedback on the quality of the transmission of the video stream or the circumstances of its rendering, and then allow the server 101 to react according to this feedback in order to improve the quality of the application.
The server and the client are used here in the context of a real-time transmission application. The time between the shooting and the rendering date must be low, for example in a videoconferencing application, typically less than 150 milliseconds.
It is known that certain types of network introduce errors during transmission. Some data packets are lost during network congestion due to memory overflows of internal network elements such as routers, or due to transmission errors, such as interference on a wireless network.
Errors can have a very large effect on the quality of motion-compensated encoded video. Indeed, an error on one image is then propagated on the following images.
A conventional way to combat errors is for the client 102 to transmit refresh requests to the server 101 via the feedback 109 broadcast through the network 100.
Upon receipt of a refresh request, the server encodes the next image without using motion compensation to stop error propagation.
Figure 2 illustrates a block diagram of a receiver device 102 adapted to incorporate the invention.
Preferably, the receiver device 102 comprises a central processing unit (CPU, central unit) 201 capable of executing instructions from a program ROM 203 when the receiver device is powered on, and instructions relating to a software application. from main memory 202 after power on.
The main memory 202 is for example of the random access memory (RAM) type which functions as a work area of the central unit 201, and the memory capacity of the latter can be increased by an optional RAM connected to an expansion port. (not illustrated).
The instructions relating to the software application can be loaded into the main memory 202 from the hard disk 206 or else from the program ROM 203 for example.
In general, an information storage means, readable by a computer or by a microprocessor, integrated or not in the device, possibly removable, is suitable for storing one or more programs, the execution of which allows the implementation of the method according to the invention. Such a software application, when it is executed by the central unit 201, results in the execution of the steps of the flowcharts shown in FIGS. 8 and 9.
A network interface 204 allows the connection of the receiver device 102 to the communication network 100. The software application when it is executed by the central unit is adapted to react to messages from the server received through the network interface 204. and providing the server 101 with information via the network 100.
A user interface 205 enables rendering of multimedia data 104 to a user. This user interface 205 can also optionally receive an information input by the latter, for example via a keyboard (not shown). The user interface 205 can also be deported to another device connected by a high speed interface simultaneously transporting the decoded audio and video (for example an HDMI type interface).
A receiver device 102 implementing the invention is for example a microcomputer, a workstation, a digital assistant, a mobile telephone, a video projector, a television set or a box connected to a television set by an HDMI socket.
With reference to FIG. 3, we will now present a possible context of application of the invention, relating to the use of certain digital video formats according to which the video images are encoded in at least two types of images.
According to these formats, a type I image (Intra image) is an image compressed independently of the other images, otherwise called an independent image. The information flow (bitstream) defining the image I therefore contains all the information useful for decoding the image on the client. This is the case for images 11 or I2 in the example of FIG. 3.
Conversely, P images (Predicted images) are images compressed according to an earlier reference image, which may be an I image or a P image. P images are therefore compressed much more efficiently, and the information flow defining the image P is generally significantly smaller than an information flow defining an image I. This is the case for the images P1 or P2 in FIG. 3. The size factor between these two types of image may for example be of the order of 10 for the same level of quantization and therefore of equivalent visual quality. This is shown in Figure 3, where the y-axis shows the size of the data streams encoding the images.
In general, the more similar a P image is to the reference image to which it refers, the smaller it is.
It should be noted that it is possible to have parts (called blocks) of a P image encoded in Intra form. In this case, the size of the image is intermediate between the size of an I image and the size of a P image and increases as a function of the proportion of blocks of the image which are coded in Intra form.
In the following, when we will speak of an I image, it can also be a P type image with a high proportion of Intra coded blocks.
For example, the video compression format can be H.264, MPEG-2, or MPEG-4. These are known formats, based on the application of Discrete Cosine Transformation (DCT) by sequence blocks and on motion compensation, followed by quantization.
When using such a format, I-pictures may be essential for several reasons. First of all, in some formats they must be sent regularly. This is for example the case with the MPEG2 format, in which I images are sent every 25 images.
I-images can also in some systems be generated only when needed.
I images are particularly useful when there is a significant change in the content of the video images, for example a change of scene, which in the context of video surveillance can be a change of camera (in a system with several cameras) or a significant change in camera position. The new image has no resemblance to the previous images and therefore motion compensation compression is not useful or effective. In this case, sending an I image allows good image quality.
The I images are also used to reset the decoding when it is noted that too high a number of packet losses on the network have been reached. Indeed, an image P being coded on the basis of an earlier image, if the latter contains errors, these are liable to propagate, or even to grow to an unacceptable level. In this case, the receiving device can perform a refresh request so that the server sends an I image (or a P image with a high proportion of Intra blocks) and therefore stops the propagation of errors.
Another use case of an I image is the start of a sequence or a reinitialization of the transfer of information when a new client is connected to a server transmitting the same stream to several clients. In the latter case, an I image must be generated to allow the new client to start decoding the video stream.
It is specified that certain formats also provide for type B images, referring to two reference images, but type B images are generally not used for real-time video applications.
Finally, a GOP (Group Of Picture or group of images) is a series of images starting with an I image which makes no reference to an image in a previous GOP.
With reference to FIG. 4, the progression of the information flows encoding compressed images on the network has been shown schematically.
After compression, the information streams encoding images of the video are split into packets to be sent over the network. The size of a packet is suitable for transmission over the network. It is conventionally less than 1500 bytes in size. Normally this constraint is taken into account by the encoder which creates information units which can be decoded independently of one another and each representing a part of the image.
The bit rate provided by the network being limited, the sending-receiving time of an image is proportional to its number of packets and therefore to the amount of information in the image. For optimum quality, the server uses the maximum bit rate available to send P images and from time to time I images (for example at regular time or during a refresh request).
It appears that in the case of transmission over a network with a limited bit rate, the large size of the I image will induce a longer transmission time than that of the P images.
In some cases, the transmission time is added to the image processing time, which may include the image decoding time by the client or the image encoding time by the server.
This transmission and processing time may be of the same order or even greater than the delay accepted for the application. In this case, if no action is taken, the client will display the P-frame preceding the I-frame for longer than normal, then display the I-frame and P-frames following the I-frame for several seconds. display times shorter than normal, so as to catch up on delays and re-synchronize video and sound.
For example, in the case of video at 25 frames per second, one frame should be captured and displayed every 40 ms. On a server capable of providing a speed of 512 kbit.s'<sup>1</sup> the maximum size of P pictures is 512 divided by 25 = 20 kbit. When sending an I image whose size is for example 5 times greater than that of a P image, the number of packets to be sent is multiplied by 5, ie a quantity of information of up to 100 kbit.
If the network speed is used to the maximum, which is often the case, the sending time of this image is also up to five times higher, ie up to 200ms. This exceeds the point-to-point exchange constraint desired in the context of real-time application, which, as we saw in the introduction, must be less than 150 ms.
To compensate for the size of the I image and therefore the time taken to send it, it is necessary to encode the following P images more strongly. This is represented in Figures 3 and 4 by a size of the images P1, P2, P3 smaller than the images P4, P5.
The transmission time of an I picture can be evaluated by using the average size of the last received I pictures and the available throughput of the network.
FIG. 5 shows the main modules of a video display device 102 which can use the invention. A first buffer memory 605 receives the information packets coming from the network 100. These packets are then used in parallel by the audio (lower part of the figure) and video (upper part of the figure) processing chains.
The audio decoder 610 decodes the encoded audio packets and stores the decoded audio data in another buffer memory 615, the contents of which are coated progressively consumed by the audio reproduction module 620 controlling a loudspeaker.
The video decoder 630 in turn decodes the gathered information packets to recreate the information flow of the encoded video images to generate decoded images which are stored in a buffer memory 640, the content of which is consumed by the video reproduction module 645. The video decoder 630 and the video reproduction module 645 are clocked by a synchronizer 650.
Audio reproduction module 620 provides synchronizer 650 with timing information relating to the soundtrack. This information is on the one hand the date recorded in the soundtrack during its encoding by the server and on the other hand, the moment when the sound is actually reproduced by the audio reproduction module. This information will be used for the synchronization of the images with the soundtrack.
In the context of the embodiment shown in FIG. 5, the network 100 is such that data is lost between the server 101 and the client 102. The video display device 102 for its part comprises an error detection module 635.
Each image is encoded by several information packets transmitted over the network. The loss of packets from the video part therefore creates errors in the decoded images.
When an image comprising errors serves as a reference to another image encoded with respect to it (whether this reference image is an I image or a P image), the errors are propagated in the P images decoded according to the reference image.
The video decoder 630 incorporates here error correction algorithms which attempt to correct errors linked, for example, to transmission problems. Known algorithms can for example use spatial or temporal interpolations. These error corrections help to mitigate the effect of errors but are not always fully effective. For example in the case of a scene with large and non-continuous movements, the temporal interpolations are not very effective.
In order to best determine the actions to be taken in the face of the presence of an error, the error detection module 635 evaluates the error rate of the decoded image as a function of the number of packets lost in the current image, of the rate error of the previous image in the case of a P image and the efficiency of the error correction.
Since a packet represents a part of the image, we define the error rate of an image as the ratio of the number of packets lost to the number of packets sent for the image:
... Nb _ Packages _ Lost, to<sub>}</sub> -—
Nb _ Packages _ Sent,
Since the error of the reference image propagates to the P image encoded according to the reference image, the error rate of a P image is composed of its own error rate added to the rate d 'error of its reference image, following the formula:
Rate'<sub>pref</sub> = Rate<sub>ref</sub> + Rate?
If the decoder uses an error correction algorithm, the inherent error rate of the picture P will be multiplied by an improvement factor A between [0.1]. If the error correction corrected all the errors, then A = 0. If, on the contrary, the error correction algorithm could not make a correction, then A = 1. We therefore obtain the error rate corrected:
Rate'<sub>p ref</sub> = Rate<sub>ref</sub> + A<sub>p</sub> *Rate<sub>p</sub>
In the embodiment presented, the client 102 is able to make requests to the server 101 (as shown in FIG. 1, reference 109). It is in particular able to send to the server a request to send an I image, in order to use this I image to reset the display and get rid of the errors that may have been accumulated due to the succession of images P. Such a request is called a “refresh request”.
The error rate is used by a module for evaluating the relevance of a refresh request 636, deciding on the basis of the information communicated to it whether or not it is appropriate to ask the server to perform a refresh of the data. data. This module 636 uses in particular the error rate which is communicated to it for each image by the error detection module 635, as well as information on the audio and video contents and the general or instantaneous characteristics of the operation of the network 100, in particular. the available speed, and the time taken for the information to arrive from the server to the client.
The synchronizer 650 calculates the desired display instant of the available images taking into account the reproduction instant of the closest audio sample. The desired image display time is calculated to match the audio and video render times.
As indicated in the introduction, for the quality of the audiovisual rendering to be satisfactory, it is necessary to have a regular display of the images and a lag between the sound and the images not exceeding 40 ms. This is not always possible, however. Indeed, if an image has a long transmission time, the decoder must wait to receive the entire image before decoding it. The image is then displayed when decoding is complete.
It is specified that the end of reception of the packets of an image is characterized either by the reception of the last packet (in the case of transport of data encoded in MPEG4 format on RTP this last packet has a specific bit positioned in its header) or by the reception of a packet with a different date (available for example in the header of information packets in RTP format).
As mentioned previously with reference to FIG. 4, an I image causes a significant transmission delay which can cause an irregularity in the display of the images and a desynchronization of the audio and the video.
This is represented in FIG. 6 which shows the reproduction times of the images (bold line) and of the sound (thin line) in a conventional situation, where the invention is not implemented.
In this example relating to a rhythmic video at 25 images per second, the average display time of an image is 1/25 second or 40 ms.
During phase 1, before the reception of an I image, the images are displayed regularly and synchronously with the soundtrack.
During phase 2, during the reception of the image I, the processing thereof takes longer and the video display is frozen during this processing, while the audio reproduction continues.
During phase 3, a series of small size P images arrive quickly and the synchronizer resynchronizes. This produces an acceleration of the display on the video screen.
Finally, during phase 4, the images are again displayed regularly and synchronously with the soundtrack.
During this scenario, the change of stage produced a sudden change in the display rhythm, unpleasant for the viewer. Also, the sound and the display were out of sync during phases 2 and 3.
FIG. 7 represents for its part the evolution of the quantity of video information present in the buffer memory 605, during the same sequence as represented in the upper part of the figure.
Before the reception of the image I, during phase 1, the buffer memory is used at a regular rate: it fills up during the reception of each image with a speed depending on the instantaneous speed of the network, and empties some time after, when the received image is to be decoded and then displayed. From one image to another, the size of all the information stored may vary, since all the P images are not all the same size, but this size remains limited.
Usually, the buffer memory becomes completely empty when decoding an image begins, the next image starting at that time or immediately after being loaded into the buffer memory. It is also possible that the next image has already started to load into the buffer at this time, so that the buffer is not completely emptied at the time of display.
During the reception of the I image, during phase 2, the size of the total information stored gradually increases to a high value corresponding to the amount of information needed to encode the I image. remaining limited, the reception of this information takes time before the decoding of this image and its display are possible. The video channel is stopped until the end of reception of the image.
During phase 3, P images of small sizes sent by the server following the transmission of the I image are received by the client. These images are received and decoded quickly and allow acceleration of the video channel, and resynchronization of the image and sound.
Finally, during phase 4, the buffer memory is used again at a regular rate, corresponding to phase 1.
The implementation of the invention, one embodiment of which is described in detail with reference to the following figures, makes it possible to anticipate the reception of the image I by adapting in advance the display rate of the images and by thus avoiding freezing of the video caused by the processing time of a too large image
FIG. 8 shows the algorithm of the evaluation module 636 which decides whether an I image should be requested or not.
The evaluation module 636 receives as input the error rate calculated by the error detection module 635, which is calculated during an error evaluation calculation step 800.
A first test 801 compares this error rate with a first value called a low value, here 5%. If the error rate is lower than this value, it is considered that a refresh request is not necessary, and no particular action is taken by the module 636.
If the error rate is on the other hand greater than this first value, a second test 805 is carried out, so as to compare this rate with a second value, called a high value, and which is here equal to 30%.
If the error rate is also greater than this value, a refresh request is sent by the client to the server via the network. This refresh request aims to obtain the sending by the server of an I image, which will make it possible to delete the errors accumulated in the P images.
Step 810 represents the sending of a refresh request, as well as the prediction of the arrival of an I image.
If during the test 805, it is found that the error rate is less than the high value, a test 815 is carried out.
This test 815 is implemented on the basis of the content of the video images and of the associated soundtrack.
The impact of a possible long-term desynchronization between the reproduction of the audio data and the video data on the quality for the viewer of the audio-visual reproduction in the application context is evaluated.
For this, the audio and video contents are analyzed.
The audio intensity is first measured. If the sound level is less than 2dB, it is considered that the sound is not very audible. It is then estimated that a long desynchronization is possible.
The faces of people audible in the audio stream are searched for in the images. Known face and lip contour detection algorithms are used for this. On this subject, reference can be made to the article “Face Detection: A Survey”, by Erik Hjelmas in Computer Vision and Image Understanding 83, 236-274 (2001).
If audible people are outside the camera's range, long-term desynchronization is considered possible.
A refresh request is then sent from the client to the server, according to the implementation of step 810. The arrival of an I image is therefore also predicted.
If the analysis of the audio and video contents carried out in step 815 shows that a long-term desynchronization is not desirable, a step 820 is carried out, to determine whether a short-term desynchronization between the audio and the video is desirable.
If, for example, the video images show large movements (which can be detected for example by the presence of large movement vectors), it is considered in step 820 that a short-term desynchronization is not desirable either.
If the images show little movement, it is considered on the contrary in step 820 that a change in display rate over a short period of time can possibly be carried out.
In this case, a test 830 is carried out to find out whether the error rate is greater than a value called the mean value, here a value of 20%.
If so, the decision to issue a refresh request to the server is taken according to step 810. Otherwise, no particular action is taken by the client to correct the display.
It should be noted that in this case the duration of the desynchronization between the audio and the video must be short. This information is stored in order to set up the image delay curve calculations which will be presented later.
The step 810 of predicting an I image comprises an evaluation of the number of images, denoted N, remaining to be received via the network before the reception of the I image.
The value of N can be estimated as a function of the average time between the sending of the request and the reception of the image I, calculated on the previous implementations of the method.
The value of N can also be imposed by the client on the server as a function of the evaluations carried out in steps 815 and 820, provided that the throughput of the network is sufficient.
By default, a long synchronization duration with a value of N equal for example to 10 images is requested by the client from the server.
In the case where short synchronization is preferable a low value of N, for example 4, is chosen.
The information according to which an I image is expected is then transmitted to the synchronizer 650 with the value N representative of the number of images to be received before the I image.
With reference to FIG. 9, the operation of synchronizer 650 has been shown, which executes a synchronization calculation algorithm each time a new image is available or an image has just been displayed.
This algorithm makes it possible to calculate the delay At which must be added to the display date of the next image to be displayed. Its operation is broken down into three modes represented by the three branches separating at node 901.
The first mode is the “Normal” mode (left part of the figure). This mode is the default mode. For each image, a test 905 is performed to see if an I image is predicted. This test uses one of the methods described with reference to Figure 8.
If no I image is predicted, no delay At is added (reference 920), the cumulative delay is zero and the video is displayed at the normal rate. The synchronizer returns to test 901.
If an I image is predicted, the algorithm switches to “anticipation” mode and a delay curve is calculated during a step 925 from the number N defined previously.
This delay curve takes into account the minimum cumulative delay to be added to anticipate the arrival of the image I, as well as the current cumulative delay.
This curve represents the delay to be added, over a number N of images, to the display time of each image in order to have the minimum cumulative delay R allowing the anticipation of the reception of a type I image.
A first solution consists in adding a constant delay to the N images preceding the type I image to have a sufficient delay to anticipate the arrival of the I image:
Zv (0 = ^ a, i being the image number in the interval [0, N-1].
This solution is implemented in cases where the number of images N before the arrival of the image I is low (for example less than 4).
The total delay is evenly distributed over all the frames. This minimizes the effect of the change of pace.
A second solution is used when the number of images N before the arrival of the image I is greater than 4. It uses the following formula which makes it possible to smooth the accelerations:
/^,(/) <sub>=</sub> p <sup>+ sin</sup>fe (O] y θ ^ ant sampled over the interval
π. 3
ÏÏ'2 - ·, -π with the formula
<img file="FR2934918A1_D0001.tif" />
Finally, whatever the solution adopted for the distribution of the delays, the curve is normalized so that the sum of the samples is equal to a coefficient A:
<img file="FR2934918A1_D0002.tif" />
the number of the type I image being l<sub>NOT</sub>, lo being the number of the start of prediction image and I the number of each image in the sequence going from image l<sub>0</sub> in the image l<sub>NOT</sub>.
To minimize the accumulated delay and have the fastest possible type I image processing, the coefficient A is chosen equal to R, the minimum cumulative delay to anticipate the type I image, which is chosen to be equal to average processing time of an image of type I (Ci) from which we subtract the average processing time of an image P (C<sub>p</sub>) and the cumulative delay at time l<sub>0</sub> (R<sub>o</sub>).
R = C<sub>I</sub>-VS<sub>p</sub>-R<sub>0</sub>
During the delay calculation step 945, the last calculated delay curve is used.
Synchronizer 650 uses the g function<sub>AT</sub>,<sub>NJo</sub>(I) to calculate the delay increment to apply to the images: <sup>tol (i) gR</sup>'<sup>NJo</sup>
The images of the video then being displayed on the screen for a display time equal to D, =
Rhythm + At<sub>{</sub>, the display date of an image i + 1 then being equal to ti + 1 = ti + Di, with ti the display date of the image i.
The curve of the delays thus calculated is illustrated in figure 10 reference
1020. We see that this formula allows a smoothing of the added delay and avoids a sudden change of rhythm.
The delay curve and the number of the current image are stored.
The cumulative delay is also updated, during a step 950 to know at each image the cumulative delay that has been introduced into the video.
The delays are added over several images to obtain an accumulated delay R. R = © z = 0
After the update of the cumulative delay, the synchronizer passes to test 901, and if the algorithm is still in the “anticipation” mode, it then passes to test 915 to find out whether an I image has been received.
If this is not the case, the synchronizer remains in “anticipation” mode and the current delay curve is applied during a new occurrence of step 945.
In the contrary case where an I image has been received, we go to step 940 to calculate a new delay curve to allow resynchronization as a function of the current cumulative delay.
The same method is used for this resynchronization step as for the anticipation step, except that the normalization coefficient A is calculated to compensate for the cumulative delay R:
A = - R
This value removes all the accumulated delay to return to synchronization with the soundtrack.
In the case of using a sinusoidal curve with the same period as that used during the anticipation step, the delay curve 1025 shown in FIG. 10 is obtained. It should be noted that a sinusoidal curve with period different could be used, or even a non-sine curve, so that the curve of the delay increments is not symmetrical about the I frame.
By default, the number N 'of images used to resynchronize can be calculated as a function of R so as not to speed up the video too much during resynchronization.
It is also possible to use a duration N 'to take into account the impact in quality of the audio video desynchronization by using the results of the audio and video content analysis steps 815 and 820 or by redoing these calculations if necessary (in the case of where module 636 is not active)
The resynchronization curve 1025 is therefore as follows
The synchronizer mode is then set to “resynchronization” (item 940). After an occurrence of steps 945 and 950, where a delay value At is applied and where the cumulative delay is updated, the synchronizer resumes step 901, and directs itself towards the central branch of the algorithm represented in FIG. 8.
The synchronizer performs test 910 to verify if a new image is predicted while the resynchronization is not finished. If this is the case, a new delay curve is calculated at step 925 according to the algorithm presented previously.
If the test 910 is negative, the test 935 which checks whether the cumulative delay is zero makes it possible to know whether the resynchronization is finished.
If so, in step 930 the synchronizer is switched to "normal" mode, and it resumes step 901.
If test 935 is negative, the algorithm remains in “resynchronization” mode and the last memorized delay curve is applied at step 945.
FIGS. 11, 12 and 13 represent an implementation of the method according to the invention, using the delay increment calculation as presented in relation to FIG. 10.
FIG. 11 shows the evolution of the cumulative delay as stored at each occurrence of step 950. This starts at the value 0, then gradually increases (reference 1010) during the anticipation phase starting with the display d 'an image numbered 0 and ending with the display of an image I numbered N, up to a cumulative delay value sufficient so that the delay of the image I does not jerk into the 'display.
The cumulative delay then gradually decreases (reference 1015) to an cumulative delay value equal to the value 0 during the resynchronization phase which begins with the display of the image I numbered N, and which ends with the display of an image numbered N + N '.
FIG. 12 shows the reproduction times of the images (bold line) and of the sound (thin line) in a situation where the invention is implemented with the cumulative delay shown in FIG. 11.
During phase 1, before the reception of the I-image, the images are displayed synchronously at a constant rate.
During phase 2, an I-frame is provided, a delay is gradually added in a flexible manner until the I-frame is processed.
The display gets out of sync with the soundtrack and the rhythm slowly changes. The cumulative delay added must be the same as the delay that frame I will cause.
During phase 3, the synchronizer sets up a resynchronization: the added delay is gradually removed in a flexible manner over a certain number of frames. The display resynchronizes with the soundtrack and the rhythm slowly evolves.
Finally, during phase 4, once the resynchronization is complete, the images are displayed synchronously. The rhythm is again constant.
It is observed that thanks to the invention, the rhythm evolves in a flexible manner over time and does not interfere with the viewing of the video.
Referring to Fig. 13, the buffer memory 605 is used as follows.
Before the reception of the I image, during phase 1, the buffer memory is used at a regular rate. It is filled by the arriving images, and it empties regularly, the arrival time of an image being of the same order of magnitude as the display time of the image preceding it.
An I image is then provided, following step 810 shown in FIG. 8. The use of the buffer memory becomes desynchronized and remains continuous, the arrival time of an image then being shorter than the display time of. the image preceding it. The average level of occupation of the buffer memory then increases, during phase 2, to a value equivalent to the size of an I image. The device then proceeds to display the image I, at which point the buffer memory is completely emptied.
During a resynchronization phase 3, the average value of the occupation of the buffer memory increases during a first part 3a of phase 3, then decreases during a second part 3b of phase 3. During part 3a, Small sized P pictures are received quickly one behind the other, while during part 3b medium sized P pictures are received.
In phase 4, after resynchronization, the buffer is used again at a regular rate.
The buffer is used continuously over time. The synchronizer 650 transmits to the video decoder 630 the desired decoding instant, which has the consequence of modulating the waiting time of the packets in the buffer memory 605.
This algorithm also has the advantage of thus taking into account any information packets reaching the buffer memory 605 with a slight delay.
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| WO2007031918A2 | Cites | World Intellectual Property Organization (WIPO) | A | Search report | 3,9 |
| US6665751B1 | Cites | United States of America | A | Search report | 1-23 |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 0855475 | France | A | |
| FR20080055475 | – | – | – |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Notification of lapseLapsedST | ST |
Numbers
- Publication
- 2934918
- Publication, DOCDB
- 2934918
- Publication, EPODOC
- FR2934918
- Application
- 855475
- Application, DOCDB
- 0855475
- Application, EPODOC
- FR20080055475
Titles
- French
- PROCEDE D'AFFICHAGE D'UNE PLURALITE D'IMAGES SUR UN DISPOSITIF D'AFFICHAGE VIDEO ET DISPOSITIF ASSOCIE.
Classification
- CPC, 7
- H04N21/4341
- H04N21/43072
- H04N21/2368
- H04N21/4305
- H04N21/44209
- H04N21/6373
- H04N21/6379
- IPC, 4
- G09G5 12
- H04N5 04
- H04N7 52
- H04N7 56