Method and system for producing a video synopsis
14 claims: 8 independent, 6 dependent
- 1ビデオカメラによって取得された第1の動的シーンの映像フレームの ソースシーケンス を、第2の動的シーンを表示する映像フレームのより短い 概要シーケンス に変換することによって、映像概要(video synopsis)を生成するための方法であって、当該方法が、 少なくとも一つのオブジェクトの動作を示す前記 ソースシーケンス の映像フレームのサブセットを取得するステップであって、各オブジェクトが、前記 ソースシーケンス の少なくとも3の異なるフレームからの 、各フレーム内で互いに連結された ピクセル の サブセットであるステップを具え、当該方法が、 前記 ソースシーケンス から少なくとも3の ソース オブジェクトを選択し、 時間サンプリング(temporal sampling) によって、各々の選択されたソースオブジェクトから1又はそれ以上の概要オブジェクトを サンプリング するステップ であって、前記時間サンプリングが、ソースシーケンスのN個のフレームを映像シーケンスのM個(NはMとは異なる)のフレームにマッピングするものであり、前記選択が、オブジェクトの現れる時間、オブジェクトの種類、およびオブジェクトの動作の種類のうちの少なくとも1つを含む選択基準に基づくものであるステップ と、 各概要オブジェクトについて前記映像概要における表示を開始する各表示時間を決定するステップ であって、映像概要の全体の再生時間が前記ソースシーケンスの再生時間よりも短くなるように、概要オブジェクトの各々について表示時間を決定するステップ と、 前記 ソースシーケンス のそれぞれ異なる時間で得られた 少なくとも2の概要オブジェクト が、前記映像概要中に同時に表示されるように、前記第1の動的シーンの前記オブジェクトの空間的な位置を変更することなく 、概 要オブジェクトをそれぞれの所定の表示時間で表示することによって前記映像概要を生成するステップと、によって特徴付けられる方法。
- 2請求項1に記載の方法において、前記オブジェクトのうちの1つが背景オブジェクトであることを特徴とする方法。
- 3請求項2に記載の方法において、前記オブジェクトと前記背景とを継ぎ目のない映像につなぎ合わせることを特徴とする方法。
- 4請求項1乃至3のいずれか1項に記載の方法において、前記ソースオブジェクトが選択され、各概要オブジェクトの表示を開始する各時間が費用関数を最適化するように決定され、 前記費用関数が、前記映像概要の長さとその質的な映像基準(qualitative video metrics)との間の妥協点を決定するものである ことを特徴とする方法。
- 5請求項1乃至4のいずれか1項に記載の方法において、前記 ソースシーケンス が固定位置で軸に対し回転するカメラによって取得され、前記 概要シーケンス にコピーする前に少なくとも前記ソースオブジェクトの少なくとも2つを空間的に移動させることを特徴とする方法。
- 6請求項1乃至4のいずれか1項に記載の方法において、前記 ソースシーケンス が固定位置で静止カメラによって取得されることを特徴とする方法。
- 7請求項1乃至6のいずれか1項に記載の方法において、少なくとも3の異なるソースフレームが時間的に連続していることを特徴とする方法。
- 8請求項1乃至7 のいずれか1項に記載の方法において、前記ソース映像シーケンスにおいて同時に発生する2のイベントが、前記 概要シーケンス において異なる時間に表示されることを特徴とする方法。
- 9請求項1乃至8 のいずれか1項に記載の方法が、監視用映像概要、動画の活動密度(activity density)の増加、映像インデックスのいずれかに用いられることを特徴とする方法。
- 10請求項9 に記載の方法が、前記 概要シーケンス 内の各ピクセルについて、前記 ソースシーケンス 内の対応するピクセルへのポインタを維持するステップを含むことを特徴とする方法。
- 11請求項1乃至10 のいずれか1項に記載の方法であって、 (a)前記 ソースシーケンス のフレーム間の画像動作パラメータを算出し、 (b)前記第1の動的シーン内の静止したオブジェクトが整列された ソースシーケンス 内で静止するように、前記 ソースシーケンス 内の映像フレームを移動させることにより、前記整列された ソースシーケンス を与えるように、当該 ソースシーケンス を事前に整列させるステップを含むことを特徴とする方法。
- 12第1の動的シーンの映像フレームの ソースシーケンス を、第2の動的シーンを表示する少なくとも2つの映像フレームの 概要シーケンス に変換するシステム(10)であって、当該システムが、 少なくとも一つのオブジェクトの動作を示す前記 ソースシーケンス 内の映像フレームのサブセットを格納し、各々のオブジェクトが、少なくとも3の異なるソースフレームからの 、各フレーム内で互いに連結された ピクセル の サブセットであり、 前記時間サンプリングが、ソースシーケンスのN個のフレームを映像シーケンスのM個のフレーム(NはMとは異なる)にマッピングするものであり、前記選択が、オブジェクトの現れる時間、オブジェクトの種類、およびオブジェクトの動作の種類のうちの少なくとも1つを含む選択基準に基づくものである 、第1のメモリ(12)を具え、前記システムが、 前記 ソースシーケンス から少なくとも3のオブジェクトを選択し、 時間サンプリング によって、各々の選択されたソースオブジェクトから1又はそれ以上の概要オブジェクトを サンプリング する、前記第1のメモリ(12)に連結された選択ユニット(13)と、 各概要オブジェクトについて、映像概要における表示を開始するために各表示時間を決定し、前記 ソースシーケンス のそれぞれ異なる時間で得た 少なくとも2の概要オブジェクト が、前記映像概要中に同時に表示されるように、前記オブジェクトの空間的位置又は前記第1の動的シーンにおいてそこから得られる前記それぞれのオブジェクトを変更することなく 、概 要オブジェクト、又は、各所定の表示時間でそこから得られたオブジェクトとを表示することによって前記映像概要を生成する、フレーム生成部 であって、映像概要の全体の再生時間が前記ソースシーケンスの再生時間よりも短くなるように、概要オブジェクトの各々について表示時間を決定するフレーム生成部 (14)と、 前記フレーム生成部に連結されて、前記 概要シーケンス のフレームを蓄積する、第2のメモリ(15)と、 前記第2の動的シーンを表示すべく、前記第2のメモリ(15)にディスプレイ装置(16)を連結する手段とによって特徴付けられるシステム。
- 13請求項12 に記載のシステムにおいて、前記フレーム生成部(14)が、前記 概要シーケンス にコピーする前に、前記ソースオブジェクトの少なくとも2つを空間的に移動させる移動ユニット(17)を具えることを特徴とするシステム。
- 14コンピュータプログラ ムで あって、当該プログラムがコンピュータ上で実行されるときに、 請求項1乃至11 のいずれか1項に記載の方法を実行するコンピュータプログラムコードを具えることを特徴とするコンピュータプログラ ム。
Independent claims14
91 paragraphs, as filed
Related application This application is a partial continuation application of the US 10 / 556,601 (Peleg et al.) "Method and sytem for spatio-temporal video warping" filed on November 2, 2006, and was published on May 11, 2006. Corresponding to WO 2006/048875, and further claiming the priority of provisional application Nos. 60 / 736,313 filed on November 15, 2005 and 60 / 759,044 filed on January 17, 2006, these The entire contents of the application are incorporated herein by reference.
The present invention generally relates to rendering-based images and videos in which new images and videos are produced by integrating parts from multiple original videos of the scene. In particular, the present invention relates to such techniques for the purpose of video abstraction or synopsis.
Conventional technology References to the prior art which are considered to be relevant as the background of the present invention are described below, the contents of which are incorporated herein by reference. Additional references are set forth in US Provisional Applications 60 / 736,313 and 60 / 759,044 described above, the contents of which are incorporated herein by reference. The confirmation of the citations described herein does not imply that it is a method related to the patentability of the present invention disclosed herein. Each citation is identified by a number enclosed in square brackets, and therefore the prior art is cited by a number enclosed in square brackets throughout the specification. [1] A. Agarwala, M. Dontcheva, M. Agrawala, S. Drucker, A. Colburn, B. Curless, D. Salesin, and M. Cohen. Interactive digital photomontage. From page 302. [2] A. Agarwala, KC Zheng, C. Pal, M. Agrawala, M. Cohen, B. Curless, D. Salesin, and R. Szeliski. Panorama video textures. pp. 821-827 of the 2005 SIGGRAPH. [3] J. Assa, Y. Caspi, and D. Cohen-Or. Action synopsis: Pose selection and illustration. pp. 667-676 of the 2005 SIGGRAPH. [4] O. Boiman and M. Irani. Detecting irregularities in images and in video. 2005 Beijing ICCV pages I: 462-469. [5] AM Ferman and AM Tekalp. Multiscale content extraction and representation for video indexing. Proc. Of SPIE, 3229: 23-31, 1997. [6] M. Irani, P. Anandan, J. Bergen, R. Kumar, and S. Hsu. Efficient representations of video sequences and their applications. Signal Processing: Image Communication, 8 (4): 327-351, 1996. [7] C. Kim and J. Hwang. An integrated scheme for object-based video abstraction. ACM Multimedia 303-311 in New York 200 years. [8] S. Kirkpatrick, CD Gelatt, and MP Vecchi. Optimization by simulated annealing. Science, 4598 (13): 671-680, 1983. [9] V. Kolmogorov and R. Zabih. What energy functions can be minimized via graph cuts? 2002 ECCV pages 65-81. [10] Y. Li, T. Zhang, and D. Tretter. An overview of video abstraction techniques. Technical Report HPL-2001-191, HP Laboratory, 2001. [11] J. Oh, Q. Wen, J. lee, and S. Hwang. Video abstraction. In S. Deb, editor, Video Data Mangement and Information Retrieval, pages 321-346. Idea Group Inc. and IRM Press, 2004. [12] C. Pal and N. Jojic. Interactive montages of sprites for indexing and summarizing security video. In Video Proceedings of CVPR05, page II: 1192, 2005. [13] A. Pope, R. Kumar, H. Sawhney, and C.Wan. Video abstraction: Summarizing video content for retrieval and visualization. In Signals, Systems and Computers, pages 915-919, 1998. [14] WO2006 / 048875 Method and system for spatio-temporal video warping, pub. May 11, 2006 by S. Peleg, A. Rav-Acha and D. Lischinski. Corresponds to US 10 / 556,601. [15] AM Smith and T. Kanade. Video skimming and characterization through the combination of image and language understanding. Pages 61-70 of CATVD in 1998. [16] A. Stefanidis, P. Partsinevelos, P. Agouris, and P. Doucette. Summarizing video datasets in the spatiotemporal domain. 906-912 at the 2000 DEXA Workshop. [17] H. Zhong, J. Shi, and M. Visontai. Detecting unusual activity in video. 2004 CVPR pages 819-826. [18] X. Zhu, X. Wu, J. Fan, AK Elmagarmid, and WG Aref. Exploring video content structure for hierarchical summarization. Multimedia Syst, 10 (2): 98-115, 2004. [19] J. Barron, D. Fleet, S. Beauchemin and T. Burkitt .. Performance of optical flow techniques, volume 92, pages 236-242. [20] V. Kwatra, A. Schδdl, I. Essa, G. Turk and A. Bobick. Graphcut textures: image and video synthesis using graph cuts. pp. 227-286 of the 2003 SIGGRAPH. [21] C. Kim and J. Hwang, Fast and Automatic Video Object Segmentation and Tracking for Content-Based Applications, IEEE Transactions on Circuits and Systems for Video Technology, Vol. 12, No. 2, February 2002, pp 122-129. [22] US Pat. No. 6,665,003
Video synopsis is a temporary compact representation that allows you to browse and search for video.
There are two main approaches to video overview. In one approach, a set of distinctive images (keyframes) is selected from the original video sequence. The selected keyframe is the frame that best represents the image [7,18]. The other approach chooses to collect short video sequences [15]. The second approach is not compact, but it has a positive effect on the dynamics of the scene. These approaches (and other approaches) have been described in an extensive study of video abstraction [10,11].
In both approaches described above, all frames are used as the basic basic elements. Another method uses mosaic images with metadata for video indexing [6,13,12]. In this way, static synopsis images contain objects at different times.
An object-based approach is known, in which objects are extracted from the input video [7,5,16]. However, these methods utilize the detection of objects that identify important keyframes and do not combine activities at different time intervals.
A method of generating a single panoramic image using the iterated min-cut [1] and generating a panoramic image using the repeated minimum cut [2] is a method in this field. Is known for. These two methods approximate the problem of exponential complexity (in the number of input frames), so they are more appropriate for a small number of frames. Related research in this area involves combining the two images using the smallest cuts [20].
WO2006 / 048875 [14] discloses methods and systems for manipulating the temporary flow of video. The first sequence of video frames in the first dynamic scene is transformed into a second sequence of video frames displaying the second dynamic scene, and in one aspect, at least in the first dynamic scene. For one feature, each part of the first sequence of video frames is sampled at a different rate than the peripheral parts of the first sequence of video frames, and the sampled parts correspond to the second sequence. It is copied to the frame to be used. This allows you to change the temporary synchronization of features in a dynamic scene.
<p> According to the first aspect of the present invention, a computer that converts a first sequence of video frames of a first dynamic scene into a second sequence of at least two video frames displaying a second dynamic scene. A method of implementation is provided and the method is (a) A step of obtaining a subset of the video frames of the first sequence showing the behavior of at least one object having a plurality of pixels arranged at each of the x and y coordinates. (b)<u style="single">In each video frame</u>A step of selecting from said subsets showing a non-spatially overlapping appearance of at least one object in the first dynamic scene. (c) The second sequence is copied from at least three different input frames into at least two consecutive frames of the second sequence, without changing the x, y coordinates of the pixels of the object, respectively. It comprises a step of making at least one frame of the sequence include at least two parts appearing in different frames of the first sequence.</p><p> According to a second aspect of the present invention, a system that transforms a first sequence of video frames in a first dynamic scene into a second sequence of at least two video frames displaying a second dynamic scene. Provided and the system concerned A first memory that stores a subset of the video frames in the first sequence that show the behavior of at least one object with multiple pixels arranged in x, y coordinates.<u style="single">In each video frame</u>A selection unit attached to the first memory for selecting from the subset portion showing a spatially non-overlapping appearance of at least one object in the first dynamic scene. The part is copied from at least three different input frames into at least two consecutive frames of the second sequence without changing the x, y coordinates of the object, respectively, and the frames of the second sequence. A frame generator containing at least two parts in which at least one of them appears in different frames in the first sequence. It includes a second memory for accumulating frames of the second sequence.</p><p> A third aspect of the present invention further comprises a data carrier that reliably implements a series of output video frames displaying a dynamic scene, wherein at least two consecutive frames of the output video have x, y coordinates, respectively. Of the output video frame, the plurality of pixels having a plurality of pixels having a plurality of pixels originating from the object parts of at least three different input frames without changing the x, y coordinates of the pixels in the object. At least one contains at least two parts that appear in different input frames.</p><p> The dynamic video synopsis disclosed by the present invention differs from the conventional video abstraction means described above in the following two identifications. (i) The video outline itself is a video that displays the dynamics of the scene. (ii) Change the relative timing of activities to reduce spatio-temporal redundancy as much as possible.</p><p> For example, consider a schematic video clip represented by the space-time volume in FIG. The footage begins with a person walking on the ground, with birds flying in the sky after a period of inactivity. Inactive frames are omitted in many video abstraction methods. The video overview is substantially more compact by displaying people and birds at the same time. This allows optimal utilization of the image area by shifting the event from the original time interval to another time interval if no other activity occurs at this spatial location. Such an operation relaxes the chronological consistency of the chronology, as shown at the beginning [14].</p><p> The present invention also provides a low-level method, utilizing the optimization of Markov random fields.<u style="single">Video overview</u>To generate.</p><p> One of the options provided by the present invention is the ability to display multiple dynamic appearances of a single object. This effect is a generalization of the "strobo-scopic" image used in traditional video overviews of moving objects [6,1]. Two different schemes are provided to do this. In the first scheme, snapshots of objects in different time cases are provided in the output video to provide an indicator of the progress of the object throughout the video from the start position to the end position. In the second scheme, the object does not define a start or end position, but moves randomly and unpredictably. In this case, snapshots of the objects in different time cases are again provided as output footage, which gives the impression of more objects than they really are. Both schemes have in common that multiple snapshots taken from the input video at different times avoid spatial duplication and do not copy the dynamic course of the object of interest. Is to be copied to the output video from.</p><p> In the content of the present invention and the accompanying claims, the term "video" is synonymous with the most common term "movie", which is accessible as a computer image file subject to post-processing. And includes any kind of video file, eg digital, analog. The camera is preferably in a predetermined position where it can rotate and zoom, but it is not a sensitive translational motion as is done with previously proposed techniques. Scenes related to the present invention are dynamic as opposed to, for example, US Pat. No. 6,665,003 [22] and static scenes processed by other references aimed at displaying stereoscopic images that do not display dynamic scenes. And the continuous frames have spatial and temporal continuity. According to one variety of the invention, we formulate this problem as a single smallest cut problem that can be explained in polynomial time by finding the largest flow on the graph [5].</p><p> In order to explain the present invention, a configuration called "spatial time amount" is used to generate a dynamic panoramic image. The amount of spatial time consists of an image input sequence by sequentially stacking all frames along the time axis. However, as far as the actual implementation is concerned, it should be understood that it is not necessary to actually configure the amount of spatial time, for example, by stacking 2D frames of a dynamic source scene in time. Moreover, the source frame is usually processed individually to form the target frame, but it is helpful to refer to the amount of space and time as if it were a physical configuration rather than a conceptual configuration. There will be.</p><p> In order to understand the present invention and to understand how the present invention is actually practiced, preferred examples will be described by reference to the accompanying drawings without limitation.</p>
1. Activity detection The present invention assumes that all input pixels have a "importance" level. From now on, the activity level will be used as the "importance" level, but it will be clear that other measurements can be used as the "importance" based on the desired application. Evaluation of importance (or activity) level is envisioned and is not a feature of the invention in itself. This can be done using one of a variety of means of detecting irregularities [4,17], moving object detection, and performing object tracking. Alternatively, this can be done on the basis of recognition algorithms such as face recognition.
For example, if a single and commonly used activity indicator is selected and the input pixel I (x, y, t) has a color difference from the temporal median at position (x, y) greater than a given threshold. , "Active". The active pixel is defined by a unique function.<img file="JP4972095B2_D0001.tif" />
A median filter is applied to the χ before continuing the synopsis process to remove noise from the activity indicator.
Although continuous activity measurements are available, the present invention focuses on a dual case. Measurements of continuous activity can be made with only minor changes to almost all equations described in detail below [4,17,1].
We describe two different examples for the calculation of video summaries. One approach (Section 2) utilizes graph representation and cost function optimization with graph cuts. Another approach (Section 3) utilizes object segmentation and tracking.
2. Image overview by minimizing energy N frames of the input video sequence are displayed in 3D spatial time I (x, y, z), in which case (x, y) is the spatial coordinate of this pixel, 1 t N. Is the number of frames.
We have the following properties<u style="single">Video overview</u>Generate S (x, y, z). -Video overview S should be substantially shorter than the original video I. -The greatest "activity" of the original video should appear in the summary video. -The behavior of the objects in the video overview should be similar to the behavior in the original video. The video overview should look good and visible seams or fragmented objects should be blocked.
Has the above-mentioned characteristics<u style="single">Video overview</u>S is generated using the mapping M and assigns the coordinates of the source pixel from I to all the coordinates (x, y, z) in the overview S. We focus on pixel time-shifting while fixing the spatial position. Therefore, the summary pixel S (x, y, z) arises from the input pixel I (x, y, M (x, y, t)). The time shift M is obtained by solving the problem of energy minimization, and the cost function is obtained by the following equation.<img file="JP4972095B2_D0002.tif" />Where E<sub>a</sub>(M) indicates loss of activity, E<sub>d</sub>(M) indicates the discontinuity at the seam. Loss of activity<u style="single">Video overview</u>It will be the number of active pixels of the input video I that does not appear in S.<img file="JP4972095B2_D0003.tif" />
Discontinuity cost E<sub>d</sub>Is<u style="single">Video overview</u>It is defined as the sum of the color differences at the seams between the spatial and temporal adjacencies within and the corresponding adjacencies in the input video (a similar equation exists in [1]).<img file="JP4972095B2_D0004.tif" />Where e<sub>i</sub>Is six units that represent six spatial and temporal adjacencies.
Figures 2a and 2b are schematic representations of the spatial and temporal processing that produces a short video overview by minimizing the cost function, and the behavior of moving objects is represented by the "activity strips" in the diagram. .. The upper figure shows the original video, and the lower figure shows the outline of the video. In particular, in FIG. 2a, a short video overview S is generated from the input video l by including many active pixels. For smoothness, if pixel A in S corresponds to pixel B in l, the "cross-boundary" adjacencies should be similar. Finding the optimal M minimization (3) is a very big optimization problem. An approximate solution is shown in Figure 2b,<u style="single">Video overview</u>Consecutive pixels within are limited to be generated from contiguous input pixels.
The cost function E (M) (equation 1) corresponds to a 3D Markov random field, and each node corresponds to a pixel in the 3D volume of the output video, and the time value corresponding to the input frame is assigned. Keep in mind that you can. The weight of the nodes is determined by the activity cost, and the edges between the nodes are determined by the discontinuity cost. Therefore, the cost function can be minimized by algorithms such as iterative graph-cuts [9].
2.1.2 Limited solution using D graph The optimization of equation (1), which allows each pixel in the video overview to occur at any time, is a major problem. For example, a 3-minute input video summarized in a 5-second video overview has about 2 each with a 5400 label.<sup>25</sup>Produces a graph with nodes.
In the case of dynamic textures or objects moving in a horizontal path, 3D MRF can be effectively solved by reducing this problem to a 1D problem [2]. In this study, we work on moving objects in a more general way, and therefore we use different constraints.<u style="single">Video overview</u>Consecutive pixels in S are restricted to originate from contiguous pixels in input video I. Under this limitation, the 3D graph is reduced to a 2D graph, and each node corresponds to a spatial position in the overview video. The label of each node M (x, y) determines the number of frames t in I shown in the first frame of S, as shown in FIG. 2b. M (x<sub>1</sub>, y<sub>1</sub>) M (x)<sub>2</sub>, y<sub>2</sub>), The seam is at two adjacent positions (x)<sub>1</sub>, y<sub>1</sub>) And (x<sub>2</sub>, y<sub>2</sub>) And the discontinuity cost along the seam E<sub>d</sub>(M) is the sum of the color differences at this spatial position across all frames in S.<img file="JP4972095B2_D0005.tif" />Where e<sub>i</sub>Is four unit vectors that represent four spatially adjacent parts.
The number of labels on each node is NK, where N and K are the number of frames for each of the input and output video. The loss of activity for each pixel is<img file="JP4972095B2_D0006.tif" />Is.
3. Object-based overview The low-level approach to dynamic video overview described above is limited to local properties such as preventing visible seams. High-level object-based characteristics can be incorporated where the object can be detected. For example, avoiding the stroboscopic effect requires detection and tracking of each object in the volume. This section describes the implementation of an object-based approach for dynamic video overview. Some object-based video summarization methods exist in the literature (eg, [7,5,16]), all of which utilize the detected objects to select important frames. Unlike these methods, the present invention moves objects in time, creates new overview frames that do not appear in the input sequence, and makes good use of space and time.
In one embodiment, moving objects are detected by comparing each pixel to a temporal median and thresholding this difference, as described above. This is followed by noise cleaning using a spatial median filter to combine spatially and temporally relevant components. It will be appreciated that there are many other methods in the literature for object detection and tracking used for this task (eg [7,17,21]). Each object detection and tracking process creates a set of objects, and each object b is represented by a unique function.<img file="JP4972095B2_D0007.tif" />
3a, 3b, and 3c are diagrams showing an example of temporal rearrangement according to the present invention. The upper figure of each figure shows the original image, the lower figure shows the outline of the image, and the movement of the moving object is represented by the "activity strip" in the figure. Figure 3a shows two objects recorded at different times and moving at the same time interval in the video overview. Figure 3b shows a single object that moves over a long period of time and is split into segments with short time intervals, which are displayed simultaneously and produce a dynamic stroboscopic effect. Figure 3c shows that the intersection of objects does not interfere with the overview when the volume of the object is divided into segments.
From each object, a segment is created by selecting a subset of the frames in which the object appears. Such segments can represent different time intervals and are optionally acquired at different sampling rates.
The video outline S is composed of the input video I by using the following processing. (1) Object b<sub>1</sub>... b<sub>r</sub>Is extracted from the input video I. (2) A set of non-overlapping segments B is selected from the original objects. (3) Temporal shift M is applied to each selected segment to generate a short video overview, prevent overlapping between objects, and allow seamless stitching. This is illustrated in FIGS. 1 and 3a-3c. FIG. 4 is a diagram showing an example in which a single frame of the video outline utilizes the dynamic stroboscopic effect shown in FIG. 3b.
Processes (2) and (3) are related because we select and shift processes (2) and (3) in time to get a short, seamless video overview. It can be understood that the above processes (2) and (3) do not have to be complete. When we say "non-spatially overlapping segment", we allow small overlaps, and when we say "prevent overlap", we allow small overlaps between objects that move in time. However, this should be minimized to obtain a visually appealing image.
In an object-based view, the pixels in the generated overview have multiple sources (generated from different objects), so we add a processing step after all the objects have been merged. The background image is generated by getting the median values of the pixels in every frame of the sequence. The "selected objects" can then be fused using weighting proportional to the distance (in RGB space) between the pixels in each frame and the median image. This stitching mechanism is similar to that used in [6].
We define a set of all pixels (x, y, t) S that map to a single overview pixel as src (x, y, t) and we are an object (or segment). # B = Σ for the number of (active) pixels in b<sub>x, y, t I</sub>χ<sub>b</sub>Let it be (x, y, t).
We define an energy function that measures the subset selection of segment B and the cost of the temporal shift M. Cost is loss of activity E<sub>α</sub>, Object E<sub>O</sub>Overlapping penalties, and long<u style="single">Video overview</u>Penalty for item E<sub>l</sub>including.<img file="JP4972095B2_D0008.tif" />here,<img file="JP4972095B2_D0009.tif" />Is.
3.1. Video with a given length-Overview Here, a short predetermined length K<u style="single">Video overview</u>However, I will explain an example consisting of a long video. For this scheme, each object is divided into overlapping, contiguous segments with length K. All segments change over time and start at time t = 1, which segment<u style="single">Video overview</u>It is necessary to decide whether it is included in. Obviously, with this scheme, some objects do not appear in the summary data.
First, we determine the cost of all paired overlaps in the segment. b<sub>i</sub>And b<sub>j</sub>Appearance time t<sub>i</sub>And t<sub>j</sub>Two segments with, and the support of each segment is represented by a unique function χ (such as equation 5).
The cost of these two segments is determined as the sum of the color differences of the two segments after time t = 1.<img file="JP4972095B2_D0010.tif" />
<u style="single">Video overview</u>In the case of, we choose a partial set of segment B that minimizes the cost of equation 6, where E<sub>l</sub>Is a constant K, and the overlap cost is<img file="JP4972095B2_D0011.tif" />Demanded by.
To prevent the same spatial and temporal pixels from appearing twice (which is acceptable but useless), we have a segment b across the original video.<sub>i</sub>And b<sub>j</sub>For v (b<sub>i</sub>, b<sub>j</sub>) = . In addition, if the stroboscopic effect is not needed, the stroboscopic effect is all b sampled from the same object.<sub>i</sub>And b<sub>j</sub>For v (b<sub>i</sub>, b<sub>j</sub>It can be avoided by setting) = .
Simulated annealing [8] is used to minimize the energy function. Each state shows a subset of the segments contained within the overview, and adjacent states are retrieved for configuration, the segment is removed, added, or replaced with another segment.
After segment selection, an overview video of length K is composed by combining all the altered segments. An example of one frame of the video outline using this means. It is shown in Figure 4.
<u style="single">3.2.</u>Lossless video overview For some applications such as video surveillance, we are long<u style="single">Video overview</u>However, all activities are guaranteed to appear. In this case, the object is not intended to select a set of object segments made in the preceding section, but to discover a compact temporal rearrangement of the object's segments.
In addition, we use pseudo-annealing to minimize energy. In this case, the states correspond to a set of time-shifts in all segments, and the two states are defined as adjacent parts if the time-shifts change only for a single segment. In this case, there are two issues to be aware of. -The segment of the object that appears in the first or last frame is<u style="single">Video overview</u>Should remain within (otherwise these will suddenly appear or disappear). We note that each state satisfies this constraint by eventually fixing the temporal shift of all these objects. The temporal arrangement of the input video is a local minimization of the common energy function and is therefore not a good option to initialize the annealing process. We initialize a pseudo-anneal with a short image and all the objects overlap.
Figures 5a, 5b and 5c show this means when a short overview can display a long sequence with no stroboscopic effect and no loss of activity. The three objects can be time-shifted so that they are displayed at the same time. In particular, FIG. 5a is a spatial and temporal schematic diagram of the original image (upper figure) and the image outline (lower figure). Figure 5b shows three frames of the original video, and as shown in Figure 5a, each person appears individually in the original video,<u style="single">Video overview</u>Now all three objects appear together. Figure 5c shows three people at the same time<u style="single">Video overview</u>Shows one frame of.
4. Panorama image overview When a video camera scans a scene, a panoramic mosaic is used to remove many redundancy. However, conventional methods construct a single panoramic image that impairs the dynamics of the scene. Limited dynamics can be represented by strobe images [6,1,3], and moving objects are displayed at several positions along these paths.
Panorama<u style="single">Video overview</u>Can be generated by simultaneously displaying activities that occur at different positions in the scene at different times. It is displayed by the camera because it provides substantial compression and the duration of activity of each object is time-limited. A special case is when the camera tracks an object, such as a running lion, as shown in Figure 6. If the camera tracks a running lion<u style="single">Video overview</u>Is a panoramic mosaic of the background, and the foreground contains some dynamic copies of the running lion. In this case, a short video overview can be obtained simply by allowing the stroboscopic effect.
Constructing the panoramic video overview is a preliminary step in aligning all the frames with some reference frames, and is done in the same way as a normal video overview. After alignment, the coordinates of the object's image are obtained from the global coordinates, which is the coordinate system of one of the input images.
To allow the video to be processed, we impose a penalty on the moving object instead of completely blocking it, even if the segmentation of the moving object is not perfect. This overlap penalty allows for adaptability of the object's temporal placement, and even if the segmentation is not perfect, the object's pixels will contain some background.
An additional period has been added,<u style="single">Video overview</u>The temporal order of is biased to the order of the input video.
Minimizing the aforementioned energies of all possible segment sections B and temporal shift M is very likely and therefore very exhausting. However, the problem can be reduced by limiting the solutions. The two restricted schemes are described in the sections below.
5. Monitoring example An interesting application of video overview is access to recorded surveillance video. If a particular event in the video needs to be investigated, the investigation can be done very quickly by using the video overview.
As mentioned earlier, Figure 5 shows an example of the power of a video overview that condenses all activities in a short period of time without losing any activity. This was done using footage collected by cameras monitoring the coffee station. Two examples are taken from real surveillance cameras. 8a, 8b and 8c are diagrams showing the details of the video outline of street surveillance. Figure 8a shows a normal frame (22 seconds) of the original video. Figure 8b shows a frame of a video overview video (2 seconds) showing condensed activity. Figure 8c shows a short video overview (0.7 seconds) frame, showing more condensed activity. The images shown in these figures are generated from footage captured by cameras monitoring the streets, with pedestrians occasionally crossing the field of view. Many of these are summarized in a very condensed overview.
8a and 8b are diagrams showing the details of the video outline of the fence surveillance. There is very little activity near the fence, and occasionally soldiers can be seen patrolling. The video overview shows all cases of patrolling and walking soldiers at the same time, optionally by displaying with stroboscopic effect.<u style="single">Video overview</u>Can even be shortened.
6. Video index based on video overview The video overview can be used as a video index, providing users with efficient and intuitive links to access actions in the video. This can be done by associating all summary pixels with pointers to the appearance of the corresponding objects in the original video. In the video overview, the information in the video is projected into the "space of activity", and the activity is important regardless of the temporal content (although we can protect the spatial content). Since the activities are concentrated within a short period of time, certain activities in the video are easily accessible.
From the above explanation, when the video camera scans a dynamic scene, the absolute "chronological time" at which the area becomes visible in the input video is not part of the scene's dynamics. Will be understandable. The "local time" within the visible period of each region is more related to the depiction of dynamics in the scene and should be protected when constructing a dynamic mosaic. The above-described embodiment is the first aspect of the present invention. In the second aspect, we show how to generate a seamless panoramic mosaic, in which in this example stitches between images are as intra-scene as possible when objects in the scene move. Prevents omitting the object part of.
7. Generate panoramic images with dimensional minimum cuts I<sub>l</sub>, ...., I<sub>N</sub>Is the frame of the input sequence. We assume that the sequences are aligned in a single reference frame using one of the traditional methods. For simplicity, we assume that all frames after alignment are the same size (pixels outside the camera's field of view are recorded as invalid). We also assume that the camera pans clockwise (other movements can be handled in a similar way).
P (x, y) is a configured panoramic image. For each pixel (x, y) in P, we need to select the frame M (x, y) from which this pixel is retrieved (ie, if M (x, y) = k, then P (x, y) = I<sub>k</sub>(x, y)). Assuming the camera pans clockwise, it is clear that the left column is taken from the first frame, while the right column is taken from the last frame (producing a panoramic image with a small field of view). Other boundary conditions can be selected for this).
Our aim is to generate seamless panoramic images. For this reason, we try to prevent stitching within an object, especially if the object is moving. We utilize a seam score similar to the score used in [1], but instead of solving the difficult problem of NP (by approximation), we find the optimal solution to a more limited problem.
8. Formulation of problems such as energy minimization problems The main difference from the above equation is<img file="JP4972095B2_D0012.tif" />It is the stitch cost required by. here, min = min (M (x, y), M (x', y')), max = max (M (x, y), M (x', y')).
This cost is reasonable assuming that the frame allocations are continuous, that is, if (x, y) and (x', y') are adjacent pixels, these source frames M (x, y) ) And M (x', y') are close. The main advantage of this cost is that this problem can be solved as a minimum cut problem on the diagram.
The energy function we minimize is<img file="JP4972095B2_D0013.tif" />And here, N (x, y) is the adjacent pixel of (x, y). E (x, y, x', y') is the stitch cost for each adjacent pixel, as explained in Equation 1. Valid (x, y, k) is 1 and I<sub>k</sub>(x, y) is the effective pixel (ie, within the field of view of the camera). D is a very large number (meaning infinity).
9. Generate a single panorama Next, we transform a two-dimensional multi-label problem (with complex exponents) into a three-dimensional two-label problem (having complex polynomials that can actually be solved immediately). Is shown. For each pixel x, y and input frame k, we determine a binary variable b (x, y, k) equal to 1 if X (x, y) <= k (M (x, y) ) Is the source frame of the pixel (x, y)). It is clear that b (x, y, k) = 1.
For each 1 k N and b (x, y, k), we define M (x, y) as the minimum k of b (x, y, k) = 1. We describe the energy term where minimization gives a seamless panorama. For each adjacent pixel (x, y) and (x', y') and for each k, we assign b (x, y, k) b (x', y', k) Add an error term to (this error term is symmetric).<img file="JP4972095B2_D0014.tif" />
We also add an infinite penalty for allocations where b (x, y, k) = 1 but b (x, y, k + 1) = 0 (M (x, y) <= k). But M (x, y)> k is not possible).
Finally, I<sub>k</sub>If (x, y) is an invalid pixel, we give an infinite penalty to the allocation so that if k> 1, b (x, y, k) = 1 b (x, y, k +) 1) If = 0 or k = 1, b (x, y, k) = 1, which can prevent the selection of this pixel (these assignments are M (x, y) = k. Means).
All the terms mentioned above are pairs of variables in the 3D grid, so we can explain the energy function of the 3D binary MRF as a minimization and minimize it in polynomial time with a minimum cut. Can [9].
10. Generate panoramic videos with 4D minimum cuts To generate a panoramic video (of length L), we should generate a series of panoramic images. It is not good to generate each panoramic image individually because it does not enhance temporal consistency. Alternatively, if a continuous mosaic image gets each pixel from the consecutive frames used in the preceding mosaic, it starts with the first mosaic image as the first frame (M).<sub>l</sub>(x, y) = M (x, y) +1). This possibility is similar to that described above with reference to FIG. 2b.
In a second aspect of the invention, we instead use another formula that gives the stitches the opportunity to change from one panorama frame to another, which is used to successfully stitch moving objects. It's very important.
We generate a four-dimensional diagram consisting of Case L of the three-dimensional diagram described above.<img file="JP4972095B2_D0015.tif" />
To enhance temporal consistency, we impose an infinite penalty on the allocation, i.e. b (x, y, N, l) = 1 for each case of l <L, b for each case of l> 1. (x, y, 1, l) = 0.
Furthermore, for each (x, y, k, l) (1 l L-1, 1 k N-1), we Set the cost function for the allocation b (x, y, k, l) = 1 b (x, y, k + 1, l + 1) (if k = N-1, we are to the left of the cost Use only terms).<img file="JP4972095B2_D0016.tif" />This cost enhances the display of consecutive pixels (in time) in the generated video (for example, these pixels are in the background).
An example of a modification of this method is that instead of connecting each pixel (x, y) to the same pixel in a contiguous frame, the optical flow at the pixel (u, v) allows the corresponding pixel (x + u, y +). v) Connect to. A suitable method for calculating optical flow is, for example, [19]. By using optical flow, we handle the case of moving objects better.
In addition, we can minimize the energy function using the smallest cut in the 4D diagram, and the binary solution determines a panoramic video that alleviates the stitching problem.
11. Practical improvements Very large memory is required to store a 4D diagram. Therefore, we take advantage of some improvements that reduce both memory requirements and algorithm operating time. As mentioned earlier, energy can be minimized without explicitly preserving vertices for invalid pixels. Therefore, the number of vertices decreases to the number of pixels in the input video and increases with the number of frames in the output video. Instead of resolving each frame in the output video, we resolve only the sampled set of output frames and insert a stitch function between them. This improvement assumes that the movement in the scene is not very large. We can only generate each pixel from a partial set of input frames. This is especially useful for a series of frames obtained from video, where the movement between each pair of consecutive frames is very small. In this case, we don't lose much by sampling a set of source frames for each pixel. However, it is preferable to sample the source frame in a consistent manner. For example, if frame k is a possible source for pixels (x, y) in the l-th output frame, then k + 1 frames are pixels (x, y) in the l + 1-th output frame. Should be a possible source frame for. We utilize a multi-resolution framework (eg, as used in [2]) and a coarse solution is found for low resolution images (buffer). After ringing and sampling), the solution is refined only at the boundaries.
Here we describe how to combine videos by score of interest. There are several applications, such as generating video with dense or sparser activity, or controlling the scene in a user-specific way.
The dynamic panorama described in [14] can be considered as a special case, where different parts of the same video are combined to obtain a video with a larger field of view, in which case we are at each time. The "visibility" of each pixel determines the score of interest. More generally, combining different parts of the same video (shifting in time or space) can be used in other cases as well. For example, to increase the density of activity within a video, we combine different parts of the video in which the action occurs into a new video with many actions. The embodiments described above with reference to FIGS. 1-8 illustrate special cases of activity and utilize different methods.
Two issues to tackle are 1. How to combine videos into "good looking" videos. For example, we want to avoid stitching problems. 2. Maximize the score of interest.
We start by explaining the different scores available and explain the scheme used to combine the videos.
One of the main features available as an interest function is the level of pixel "importance". In our experiments, we consider "activity" within a pixel to indicate its importance, but other measurements of importance are equally suitable. The activity level assessment is not the feature of the invention itself, but can be done using one of the various methods referred to in Section 1 above (activity detection).
13. Other scores Other scores available to combine videos: Visibility score: There are pixels that cannot be seen when the camera moves or when trying to fill a hole in the image. We can penalize invalid pixels (not necessarily by infinite scores). This way we can help fill the holes (or widen the field of view), but if bad stitches occur, we may not want to fill the holes or have a small field of view. You may use it. Direction: The measurement of activity can be replaced with the measurement of direction. For example, we would prefer a horizontally moving area to a vertically moving area. -User specifications: The user specifies suitable features of interest such as color and texture. In addition, users manually specify regions (and time slots) with different scores. For example, the user can control the dynamics in the scene by drawing a mask where 1 means that maximum activity is desired and 0 means no activity is desired. That is, it can be generated at a specific position.
14. Algorithm We utilize a method similar to the method used by [20] with the following modifications. We add an area of interest for each pixel selected from one video or another. This score is added to the vertices (source and sink) of the terminal utilizing the edges of each pixel in each video, and the weights of these edges are the scores of interest. We calculate (optionally) the optical flow between each successive pair of frames. Next, for added consistency, we set the edges of the temporally adjacent parts ((x, y, t) to (x, y, t + 1)) to the edges of the adjacent parts by optical flow ( It can be exchanged from (x, y, t) to (x + u (x, y), y + v (x, y), t + 1)). This enhances the transition between stitch videos so that the stitches are encouraged to follow a less prominent flow. When deciding which part of the video (or moving part) will be combined, not only the stitch cost but also the score of interest should be considered. For example, when generating a video with a high density of activity levels, we choose a set S of videos that maximizes the score.<img file="JP4972095B2_D0017.tif" />
Figure 9b shows the effect, such as increased activity density of the video, with the original frame shown in Figure 9a. When two or more videos are combined, we use a repetitive means, and this iteration combines the new video with the generated video. For accuracy, old seams and scores generated by previous iterations should be considered. This scheme is described in [20], even though there are no scores of interest. A sample frame of the generated video is shown in Figure 9b.
FIG. 10 is a schematic diagram of the process. In this example, the video is combined with a time-shifted video. Coupling is done using the minimum cut, based on the criteria described above, that is, maximizing the score of interest while minimizing stitch costs.
Referring here to FIG. 11, the first sequence of video frames of the first dynamic scene acquired by the camera 11 is converted into the second sequence of at least two video frames showing the second dynamic scene. A block diagram of the system 10 according to the present invention is shown. The system includes a first memory 12 that stores a subset of video frames without a first sequence showing the behavior of at least one object with multiple pixels located at each x, y coordinate. A selection unit 13 that selects from a subset of the spatially non-overlapping appearance of at least one object in the first dynamic scene is concatenated to the first memory 12. The frame generator 14 copies at least three different parts of the input frame into at least two consecutive frames in the second sequence without changing the x, y coordinates of each pixel in the object, and the second At least one frame of the sequence contains at least two parts that appear in different frames within the first sequence. The frames of the second sequence are stored in the second memory 15 for subsequent processing or display by the display unit 16. The frame generator 14 may include a warping unit 17 that spatially moves at least two parts before copying the second sequence.
In practice, system 10 is implemented by an optimally programmed computer with a graphics card or workstation and suitable peripherals, as is well known in the art.
In System 10, at least three different input frames are close in time. The system 10 may further include any alignment unit 18 coupled to a first memory for pre-aligning the first sequence of video frames. In this case, the camera 11 is connected to the alignment unit 18 to store the pre-arranged video frames in the first memory 12. Alignment unit 18 Calculate the image behavior parameters between frames in the first sequence and It works by moving the video frame in the first sequence so that the stationary object in the first dynamic scene is still in the video.
Similarly, system 10 includes any time slice generator 19 coupled to selection unit 13 to clear the aligned space and amount of time with a "time front" surface and generate a series of time slices. You may.
These optional features are not described in detail as the terms "timefront " and " timeslice" are similar to those fully described in the aforementioned WO2006 / 048875 referenced. ..
For completeness, FIG. 12 is a flow diagram showing the main processes performed by the system 10 according to the present invention.
15. Consideration The video outline has been proposed as a means of condensing the activities in the video in a very short period of time. This condensed display allows efficient access to activities within the video sequence. Two methods have been proposed, one using low-level graph optimization.<u style="single">Video overview</u>Each pixel in is a node in this graph. This means is directly from the input video<u style="single">Video overview</u>It has the advantage of obtaining a solution, but the complexity of the solution is very high. An alternative is to detect the first moving object and optimize the detected object. The second method requires a preliminary step in segmenting the behavior, which is very fast and allows object-based constraints. The activity of the generated video summary is much more condensed than the original video, and such a summary may be awkward for those who have never used it. However, if the goal is to observe a lot of information in a short amount of time, the video overview achieves this goal. Special attention should be paid to the possibility of acquiring a dynamic stroboscope. Along with further reducing the length of the video overview, the dynamic stroboscope may require user adjustment. Multiple spatial appearances of a single object require some training to be found to exhibit long activity times. We describe a particular example of a dynamic video overview, but many extensions are possible. For example, instead of having a binary "activity" indicator, the activity indicators may be contiguous. Continuous activity can be done, for example, by controlling the speed of the displayed object based on the activity level.<u style="single">Video overview</u>You can expand the options that can be generated. The video overview may also be used for long moving images consisting of many shots. Theoretically, our algorithm does not combine parts of different scenes due to the disadvantage of overlapping (or discontinuity). In this case, the single background model utilized for a single shot is replaced with an adjustable background estimator. Another method used for long videos is to use the traditional method of binary detection of shots to generate a video summary for each shot individually.
It will be understood that the system according to the present invention is optimal for a programmed computer. Similarly, the present invention contemplates a computer program readable by a computer in order to carry out the methods of the present invention. Furthermore, the present invention contemplates a machine-readable memory that explicitly executes program instructions that can be executed by a machine that implements the methods of the invention.
<figref num="1">FIG. 1 is a diagram showing an approach of the present invention that generates a compact video outline by simultaneously displaying temporally arranged features.</figref><figref num="2a">FIG. 2a is a schematic view showing an outline of an image generated by the present invention.</figref><figref num="2b">FIG. 2b is a schematic view showing an outline of an image generated by the present invention.</figref><figref num="3a">FIG. 3a is a diagram showing an example of temporary rearrangement according to the present invention.</figref><figref num="3b">FIG. 3b is a diagram showing an example of temporary rearrangement according to the present invention.</figref><figref num="3c">FIG. 3c is a diagram showing an example of temporary rearrangement according to the present invention.</figref><figref num="4">FIG. 4 is a diagram showing a single frame of the video outline using the dynamic stroboscopic effect shown in FIG. 3b.</figref><figref num="5a">FIG. 5a is a diagram showing an example in which a short overview can display a long sequence without compromising activity and without stroboscopic effect.</figref><figref num="5b">FIG. 5b is a diagram showing an example where a short overview can display a long sequence without compromising activity and without stroboscopic effect.</figref><figref num="5c">FIG. 5c is a diagram showing an example where a short overview can display a long sequence without compromising activity and without stroboscopic effect.</figref><figref num="6">FIG. 6 is a diagram showing a further example of the panoramic image outline according to the present invention.</figref><figref num="7a">FIG. 7a is a diagram showing details of a video outline of street surveillance.</figref><figref num="7b">FIG. 7b is a diagram showing details of a video outline of street surveillance.</figref><figref num="7c">FIG. 7c is a diagram showing details of a video outline of street surveillance.</figref><figref num="8a">FIG. 8a is a diagram showing the details of the video outline of the fence monitoring.</figref><figref num="8b">FIG. 8b is a diagram showing the details of the video outline of the fence monitoring.</figref><figref num="9a">FIG. 9a is a diagram showing the increasing activity density of moving images according to a further embodiment of the present invention.</figref><figref num="9b">FIG. 9b is a diagram showing the increasing activity density of moving images according to a further embodiment of the present invention.</figref><figref num="10">Figure 10 shows the figure<u style="single">9</u>It is a schematic diagram of the process used to generate the moving image shown in.</figref><figref num="11">FIG. 11 is a block diagram showing the main functions of the system according to the present invention.</figref><figref num="12">FIG. 12 is a flow chart showing the main operations performed by the present invention.</figref>
38 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR102493856B1 | Cited by | Republic of Korea | Search report |
| US11501534B2 | Cited by | United States of America | Applicant |
| JP2005210573A | Cites | Japan | – |
| WO2004040480A1 | Cites | World Intellectual Property Organization (WIPO) | – |
| WO01078050A1 | Cites | World Intellectual Property Organization (WIPO) | – |
51 members in 11 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 60736313 | United States of America | – | |
| 73631305 | United States of America | P | |
| 60759044 | United States of America | – | |
| 75904406 | United States of America | P | |
| 2006001320 | Israel | W |
Members51
| Document | Office | Kind | |
|---|---|---|---|
| AU2006314066A1 | Australia | A1 | |
| CA2640834A1 | Canada | A1 | |
| WO2007057893A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2007057893A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AU2007345938A1 | Australia | A1 | |
| CA2676632A1 | Canada | A1 | |
| WO2008093321A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1955205A2 | European Patent Office (EPO) | A2 | |
| KR20080082963A | Republic of Korea | A | |
| CN101366027A | China | A | |
| JP2009516257A | Japan | A | |
| IL191232A0 | Israel | A0 | |
| US2009219300A1 | United States of America | A1 | |
| KR20090117771A | Republic of Korea | A | |
| EP2119224A1 | European Patent Office (EPO) | A1 | |
| CN101689394A | China | A | |
| IL199678A0 | Israel | A0 | |
| US2010092037A1 | United States of America | A1 | |
| US2010125581A1 | United States of America | A1 | |
| JP2010518673A | Japan | A | |
| JP2010134923A | Japan | A | |
| AU2007345938B2 | Australia | B2 | |
| BRPI0620497A2 | Brazil | A2 | |
| US8102406B2 | United States of America | B2 | |
| US2012092446A1 | United States of America | A1 | |
| JP4972095B2This record | Japan | B2 | |
| EP1955205B1 | European Patent Office (EPO) | B1 | |
| DK1955205T3 | Denmark | T3 | |
| AU2006314066B2 | Australia | B2 | |
| US8311277B2 | United States of America | B2 | |
| IL199678A | Israel | A | |
| US2013027551A1 | United States of America | A1 | |
| CN101366027B | China | B | |
| IL191232A | Israel | A | |
| US8514248B2 | United States of America | B2 | |
| JP5355422B2 | Japan | B2 | |
| JP5432677B2 | Japan | B2 | |
| BRPI0720802A2 | Brazil | A2 | |
| CN101689394B | China | B | |
| KR101420885B1 | Republic of Korea | B1 | |
| CA2640834C | Canada | C | |
| US8818038B2 | United States of America | B2 | |
| KR101456652B1 | Republic of Korea | B1 | |
| US8949235B2 | United States of America | B2 | |
| CA2976801A1 | Canada | A1 | |
| WO2016131129A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CA2676632C | Canada | C | |
| US2018042388A1 | United States of America | A1 | |
| EP3297272A1 | European Patent Office (EPO) | A1 | |
| BRPI0620497B1 | Brazil | B1 | |
| BRPI0720802B1 | Brazil | B1 |
33 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Written request for registration of change of nameJAPANESE INTERMEDIATE CODE: R313533S533 | S533 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4972095
- Application
- 2008539616
Titles2
- Japanese
- 映像概要を生成する方法およびシステム
- English
- How and system to generate video overview
Classification
- CPC, 7
- H04N5/2625
- H04N21/8549
- G11B27/034
- G11B27/28
- G06F16/739
- G06V20/40
- G06T3/16
- IPC, 2
- G06T3 00
- G06T7 20
