Method, apparatus and computer program product for providing adaptive gesture analysis
Abstract
A method for providing adaptive gesture analysis may include the following steps: dividing a distance range into multiple depth ranges, generating multiple intensity images of at least two image frames, and each intensity image of the intensity images provides an indication Image data of an object existing in a corresponding depth range of a respective image frame, determine the motion change between the two image frames in each corresponding depth range, and based at least in part on the motion change, Determine the depth of a target. A device and computer program product corresponding to the method are also provided.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
26 claims: 26 independent, 0 dependent
- 1A method includes the following steps:dividing a distance range into a plurality of depth ranges;generating a plurality of intensity images of at least two image frames, each intensity image of the intensity images providing the indication of a respective image frame Wait for the image data of the object in the corresponding depth range of one of the multiple depth ranges;determine the motion change between the two image frames in each corresponding depth range;and based at least in part on the motion change, Determine the depth of a target. 一種方法,包含以下步驟:將一距離範圍劃分為多個深度範圍;產生至少二個影像訊框的多個強度影像,該等強度影像之每一強度影像提供指示在一各自影像訊框的該等多個深度範圍之一相對應深度範圍中存在物體的影像資料;判定在每一相對應深度範圍之該等二個影像訊框之間的動作變化;及至少部分地基於該動作變化,來判定一目標的深度。
- 2According to the method described in item 1 of the scope of the patent application, the division of the distance range includes the following steps:selecting a distance interval between the short distance and the long distance that is wider than the distance between the short distance and the long distance. 如申請專利範圍第1項所述之方法,其中劃分該距離範圍包含以下步驟:在近距離與遠距離處選擇比該等近距離與遠距離之間的距離間隔還寬的距離間隔。
- 3In the method described in claim 1, wherein generating the multiple intensity images includes the following steps:generating multiple intensity images of adjacent frames. 如申請專利範圍第1項所述之方法,其中產生該等多個強度影像包含以下步驟:產生相鄰訊框之多個強度影像。
- 4For example, the method described in item 1 of the scope of patent application, wherein generating the multiple intensity images includes the following steps:For each corresponding depth range of a specific frame, it will correspond to an object that is not in a current depth range The data is set to a predetermined value, and the data corresponding to the current depth range is retained. 如申請專利範圍第1項所述之方法,其中產生該等多個強度影像包含以下步驟:對於一特定訊框的每一相對應深度範圍,將與不在一目前深度範圍中之物體相對應的資料設定為一預定值,且保留與該目前深度範圍相對應的資料。
- 5For the method described in item 4 of the scope of patent application, the determination of the action change includes the following steps:compare the object data in a current frame with a previous The object data in the previous frame is compared to determine a change in the intensity from the previous frame to the current frame, and the data of the change in the unindicated intensity is set to the predetermined value. 如申請專利範圍第4項所述之方法,其中判定動作變化包含以下步驟:將在一目前訊框中的物體資料與在一先 前訊框中的物體資料相比較,以判定在從該先前訊框至該目前訊框之強度中的一改變,且將未指示強度中該變化的資料設定為該預定值。
- 6According to the method described in claim 1, wherein determining the depth of the target includes the following steps:determining the target depth based on an integration of the movement change and the target position clue information. 如申請專利範圍第1項所述之方法,其中判定該目標的深度包含以下步驟:基於該動作變化與目標位置線索資訊的一整合,來判定目標深度。
- 7The method described in item 1 of the scope of patent application further includes the following steps:tracking the target's action based on the determination of the depth of the target. 如申請專利範圍第1項所述之方法,更包含以下步驟:基於該目標之該深度的該判定,追蹤該目標的動作。
- 8The method described in item 7 of the scope of patent application further includes the following steps:recognizing the gesture feature of the target, and initiating a user interface command based on a recognized gesture. 如申請專利範圍第7項所述之方法,更包含以下步驟:辨識該目標的手勢特徵,以基於一所辨識的手勢,啟始一使用者介面命令。
- 9A device includes:a processor;and a memory, which includes computer program code. The memory and the computer program code are combined with the processor to cause the device to perform at least the following steps: dividing a distance range into Multiple depth ranges;generating multiple intensity images of at least two image frames, each intensity image of the intensity images provides an indication that one of the multiple depth ranges of a respective image frame exists in a corresponding depth range The image data of the object;determine the motion change between the two image frames in each corresponding depth range;and determine the depth of a target based at least in part on the motion change Spend. 一種裝置,包含:一處理器;及一記憶體,其包括電腦程式碼,該記憶體和該電腦程式碼與該處理器組配來致使該裝置執行至少下面的步驟:將一距離範圍劃分為多個深度範圍;產生至少二個影像訊框的多個強度影像,該等強度影像之每一強度影像提供指示在一各自影像訊框的該等多個深度範圍之一相對應深度範圍中存在物體的影像資料;判定在每一相對應深度範圍之該等二個影像訊框之間的動作變化;及至少部分地基於該動作變化,來判定一目標的深 度。
- 10For the device described in item 9 of the scope of patent application, the memory including the computer code and the processor are further assembled to cause the device to select between the short and long distances than the short and long distances. The distance between the gaps is also a wide distance gap to divide the distance range. 如申請專利範圍第9項所述之裝置,其中包括該電腦程式碼之該記憶體與該處理器進一步組配來致使該裝置透過在近距離及遠距離處選擇比該等近距離與遠距離之間的距離間隔還寬的距離間隔,來劃分該距離範圍。
- 11For the device described in item 9 of the scope of patent application, the memory including the computer code and the processor are further combined to cause the device to generate multiple intensity images of adjacent frames. Intensity images. 如申請專利範圍第9項所述之裝置,其中包括該電腦程式碼之該記憶體與該處理器進一步組配來致使該裝置透過產生相鄰訊框的多個強度影像,來產生該等多個強度影像。
- 12For the device described in item 9 of the scope of the patent application, the memory including the computer code and the processor are further assembled to cause the device to pass through each corresponding depth range for a specific frame. The data corresponding to an object in the current depth range is set to a predetermined value, and the data corresponding to the current depth range is retained to generate the multiple intensity images. 如申請專利範圍第9項所述之裝置,其中包括該電腦程式碼之該記憶體與該處理器進一步組配來致使該裝置透過對於一特定訊框的每一相對應深度範圍,將與不在一目前深度範圍中的物體相對應的資料設定為一預定值,且保留與該目前深度範圍相對應的資料,來產生該等多個強度影像。
- 13For the device described in item 12 of the scope of patent application, the memory including the computer code and the processor are further assembled to cause the device to combine object data in a current frame with a previous frame The object data is compared to determine a change in an intensity from the previous frame to the current frame, and the data of the change in the unindicated intensity is set to the predetermined value to determine the action change. 如申請專利範圍第12項所述之裝置,其中包括該電腦程式碼之該記憶體與該處理器進一步組配來致使該裝置透過將一目前訊框中的物體資料與在一先前訊框中的物體資料相比較,來判定在從該先前訊框至該目前訊框的一強度中的一改變,且將未指示強度中之該改變的資料設定為該預定值,從而判定動作變化。
- 14For the device described in item 9 of the scope of patent application, the memory including the computer code and the processor are further combined to cause the device to integrate the target location clue information based on the movement change Determine the depth of the target, thereby determining the depth of the target. 如申請專利範圍第9項所述之裝置,其中包括該電腦程式碼之該記憶體與該處理器進一步組配來致使該裝置透過基於該動作變化與目標位置線索資訊的一整合來 判定目標深度,從而判定該目標的深度。
- 15For the device described in item 9 of the scope of patent application, the memory including the computer code and the processor are further combined to cause the device to track the target's action based on the determination of the depth of the target. 如申請專利範圍第9項所述之裝置,其中包括該電腦程式碼之該記憶體與該處理器進一步組配來致使該裝置基於該目標之該深度的該判定,來追蹤該目標的動作。
- 16The device described in item 15 of the scope of patent application, in which the memory including the computer code and the processor are further combined to cause the device to recognize the gesture feature of the target, and start based on a recognized gesture A user interface command. 如申請專利範圍第15項所述之裝置,其中包括該電腦程式碼之該記憶體與該處理器進一步組配來致使該裝置辨識該目標的手勢特徵,以基於一所辨識的手勢,啟始一使用者介面命令。
- 17A computer program product comprising at least one computer-readable non-transitory storage medium having computer-executable program code instructions stored thereon. The computer-executable program code instructions include:a plurality of first program code instructions, which It is used to divide a distance range into multiple depth ranges;multiple second code commands are used to generate multiple intensity images of at least two image frames, and each intensity image of the intensity images provides instructions Image data of an object exists in one of the multiple depth ranges of a respective image frame;multiple third code commands are used to determine the two of each corresponding depth range The motion change between the image frames;and a plurality of fourth code commands, which are used to determine the depth of a target based at least in part on the motion change. 一種電腦程式產品,其包含具有儲存於其上之電腦可執行程式碼指令的至少一個電腦可讀非暫時性儲存媒體,該等電腦可執行程式碼指令包含:多個第一程式碼指令,其等用以將一距離範圍劃分為多個深度範圍;多個第二程式碼指令,其等用以產生至少二個影像訊框的多個強度影像,該等強度影像之每一強度影像提供指示在一各自影像訊框的該等多個深度範圍之一相對應深度範圍中存在物體的影像資料;多個第三程式碼指令,其等用以判定每一相對應深度範圍之該等二個影像訊框之間的動作變化;及多個第四程式碼指令,其等用以至少部分地基於該動作變化,來判定一目標的深度。
- 18For example, the computer program product described in item 17 of the scope of patent application, wherein the first program code instructions include options for selecting a distance between the short distance and the long distance that is wider than the distance between the short distance and the long distance between Every instruction. 如申請專利範圍第17項所述之電腦程式產品,其中該等第一程式碼指令包括用以在近距離與遠距離處選擇比該等近距離與遠距離之間的距離間隔還寬的距離間 隔的指令。
- 19For the computer program product described in item 17 of the scope of patent application, the second program code instructions include instructions for generating the multiple intensity images of adjacent frames. 如申請專利範圍第17項所述之電腦程式產品,其中該等第二程式碼指令包括用以產生相鄰訊框之該等多個強度影像的指令。
- 20For example, the computer program product described in item 17 of the scope of patent application, wherein the second program code instructions include each corresponding depth range for a specific frame, which will correspond to an object that is not in a current depth range The data of is set to a predetermined value, and the command to retain the data corresponding to the current depth range. 如申請專利範圍第17項所述之電腦程式產品,其中該等第二程式碼指令包括用以對於一特定訊框的每一相對應深度範圍,將與不在一目前深度範圍中之物體相對應的資料設定為一預定值,且保留與該目前深度範圍相對應的資料的指令。
- 21For example, the computer program product described in item 20 of the scope of patent application, wherein the third program code instructions include the object data used to compare the object data in a current frame with the object data in a previous frame to determine from the An instruction to change the intensity of the previous frame to the current frame, and to set the changed data in the unindicated intensity to the predetermined value. 如申請專利範圍第20項所述之電腦程式產品,其中該等第三程式碼指令包括用以將一目前訊框中的物體資料與一先前訊框中的物體資料相比較,來判定從該先前訊框至該目前訊框的強度中的一改變,且將未指示強度中之該改變的資料設定為該預定值的指令。
- 22For example, in the computer program product described in item 17 of the scope of patent application, the fourth program code instructions include instructions for determining the target depth based on an integration of the movement change and the target position clue information. 如申請專利範圍第17項所述之電腦程式產品,其中該等第四程式碼指令包括用以基於該動作變化與目標位置線索資訊的一整合,來判定目標深度的指令。
- 23For example, the computer program product described in item 17 of the scope of patent application further includes a plurality of fifth program code commands, which are used to track the action of the target based on the determination of the depth of the target. 如申請專利範圍第17項所述之電腦程式產品,更包含多個第五程式碼指令,其等用以基於該目標之該深度的該判定,來追蹤該目標的動作。
- 24For example, the computer program product described in item 23 of the scope of patent application further includes a plurality of sixth program code commands, which are used to identify the gesture feature of the target, and initiate a user interface command based on a recognized gesture . 如申請專利範圍第23項所述之電腦程式產品,更包含多個第六程式碼指令,其等用以辨識該目標的手勢特徵,以基於一所辨識的手勢,啟始一使用者介面命令。
- 25A device that includes:A device for dividing a distance range into multiple depth ranges;a device for generating multiple intensity images of at least two image frames, each of the intensity images provides an indication of a respective image frame Image data of an object in one of the multiple depth ranges corresponding to the depth range;a device for determining the motion change between the two image frames in each corresponding depth range;and at least partially Based on the movement change, a device for determining a target depth is determined. 一種裝置,包含: 用以將一距離範圍劃分為多個深度範圍的裝置;用以產生至少二個影像訊框的多個強度影像的裝置,該等強度影像的每一強度影像提供指示在一各自影像訊框的該等多個深度範圍之一相對應深度範圍中存在物體的影像資料;用以判定每一相對應深度範圍之該等二個影像訊框之間的動作變化的裝置;及用以至少部分地基於該動作變化,來判定一目標深度的裝置。
- 26For example, the device described in item 25 of the scope of patent application, wherein the device for generating the multiple intensity images includes, for each corresponding depth range of a specific frame, it will be different from the one that is not in the current depth range. The data corresponding to the object is set to a predetermined value, and the device retains the data corresponding to the current depth range. 如申請專利範圍第25項所述之裝置,其中用以產生該等多個強度影像的裝置包含,用以對於一特定訊框的每一相對應深度範圍,將與不在一目前深度範圍中之物體相對應的資料設定為一預定值,且保留與該目前深度範圍相對應之資料的裝置。
Independent claims26
74 paragraphs, as filed
Method, device and computer program product for providing adaptive gesture analysis
METHOD, APPARATUS AND COMPUTER PROGRAM PRODUCT FOR PROVIDING ADAPTIVE GESTURE ANALYSIS
Technical field
The embodiments of the present invention are generally related to user interface technology, and more particularly to a method, device, and computer program product for providing gesture analysis of a visual interactive system.
background
The modern communication era has caused a huge expansion of wired and wireless networks. Computer networks, television networks, and telephone networks are experiencing an unprecedented technological expansion due to consumer demand. Wireless and mobile network technologies have met relevant consumer needs and provide more flexible and real-time information transmission.
Current and future network technologies continue to promote the simplicity of information transmission and the convenience of users. One area where there is a need to improve the simplicity of information transmission and the convenience of users is related to the human-computer interface for HCI (Human-Computer Interaction). With recent developments in computing equipment and handheld or mobile devices that have improved the performance of these devices, many people are thinking about the development of the next generation of HCI. Furthermore, considering that these devices will tend to improve their ability to generate content, store content, and/or receive content fairly quickly upon request, and also consider that mobile electronic devices such as a mobile phone usually face display size, text The input speed and the physical implementation of the user interface (UI) are limited, so the challenge usually arises in the context of HCI.
Furthermore, the improvements in HCI can also enhance user enjoyment, and for the user interface with computing devices in the environment, open up the possibility that effective HCI may have changed in other ways. One such improvement is related to gesture recognition. Compared with other interaction mechanisms such as keyboards and mice currently used in HCI, some people may consider gesture recognition to improve the naturalness and ease of communication. In this way, some applications have been developed to enable gesture recognition to be used as a command controller in digital home devices, in file/web page navigation, or as an alternative to the commonly used remote control. However, the current gesture analysis mechanism is usually slow or troublesome to use. Moreover, many of the currently appropriate gesture analysis mechanisms may experience difficulties in detecting or tracking gestures in an unconstrained environment. For example, embodiments where it is difficult to distinguish the background under varying or certain lighting configurations and environments may present challenges in gesture tracking. Therefore, considering the general usability of next-generation HCI, improvements in gesture analysis may be expected.
A brief summary of the invention with some examples of the invention
Therefore, a method, device and computer program product are provided to realize the use of gesture analysis in, for example, a visual interactive system. In some exemplary embodiments, an adaptive gesture tracking solution can use depth data to analyze image data. For example, for various depth ranges, multiple intensity images can be considered. In an exemplary embodiment, the intensity images under the various depth ranges can be analyzed to determine the depth at which a target (for example, a hand or other gesture appendages) is located, so that the target can be tracked . As such, in some cases, changes in motion in each depth range can be used with other cues to provide relatively fast and accurate target tracking. In some embodiments, three-dimensional (3D) depth data can be used to provide depth data of intensity images to enable adaptive gesture analysis in an unrestricted environment. In this manner, some exemplary embodiments of the present invention can provide relatively robust and fast gesture analysis.
In an exemplary embodiment, a method for providing adaptive gesture analysis is provided. The method may include the following steps: dividing a distance range into a plurality of depth ranges, generating a plurality of intensity images of at least two image frames, wherein each intensity image of the intensity images can provide an indication of a respective image frame The image data of an object existing in a corresponding depth range is determined, the motion change between the two image frames in each corresponding depth range is determined, and the depth of a target is determined based at least in part on the motion change.
In another exemplary embodiment, a computer program product for providing adaptive gesture analysis is provided. The computer program product includes a computer readable storage medium with a computer executable program code portion stored thereon. The computer executable code portions may include first, second, third, and fourth code portions. The first code part is used to divide a distance range into multiple depth ranges. The second code part is used to generate a plurality of intensity images of at least two image frames. Each of the intensity images can provide image data indicating the presence of objects in a corresponding depth range of a respective image frame. The third code part is used to determine the motion change between the two image frames in each corresponding depth range. The fourth code portion is used to determine the depth of a target based at least in part on the motion change.
In another exemplary embodiment, a device for providing adaptive gesture analysis is provided. The device may include a processor. The processor can be configured to divide a distance range into multiple depth ranges to generate multiple intensity images of at least two image frames, wherein each intensity image of the intensity images can provide an indication of a respective image signal The image data of the object existing in a corresponding depth range of the frame is determined, the motion change between the two image frames in each corresponding range is determined, and the depth of a target is determined based at least in part on the motion change.
In yet another exemplary embodiment, a device for providing adaptive gesture analysis is provided. The device may include a device for dividing a distance range into a plurality of depth ranges, and a device for generating a plurality of intensity images of at least two image frames, wherein each intensity image of the intensity images provides an indication The image data of an object existing in a corresponding depth range of each image frame is used to determine the motion change between the two image frames in each corresponding depth range, and is used at least partly based on The movement changes to determine a target depth device.
The embodiments of the present invention can provide a method, device, and computer program product for use in, for example, a mobile or fixed environment. Therefore, for example, computing device users can enjoy improved performance when interacting with their respective computing devices.
A brief description of the multiple views of the schema
Thus, some embodiments of the present invention have been described in general, and reference will now be made to additional drawings that are not necessarily drawn to scale, in which: Figure 1 illustrates an exemplary embodiment of the present invention, one of a UI controller is adaptable An example of a gesture analysis process; Figure 2 shows a diagram of different depth intervals dividing the entire depth range of a 3D image according to an exemplary embodiment of the present invention; Figure 3 shows an exemplary implementation according to the present invention For example, a schematic block diagram of a device that can realize gesture analysis; Figure 4 (including Figures 4A to 4I) shows an exemplary embodiment of the present invention, corresponding to the intensity generated by each of various different depths Image; Figure 5 (including Figures 5A to 5N) shows an indication of the intensity images of adjacent frames and their relative actions according to an exemplary embodiment of the present invention; Figure 6 (including Figures 6A to 6F ) Shows the stages of a process of determining the target depth according to an exemplary embodiment of the present invention; Figure 7 shows a block diagram of a mobile terminal that may benefit from an exemplary embodiment of the present invention; and Figure 8 is According to an exemplary embodiment of the present invention, a flowchart of an exemplary method for gesture analysis is provided.
Detailed description of some embodiments of the present invention
Now, some embodiments of the present invention will be more completely described below with reference to the additional drawings, in which some but not all of the embodiments of the present invention are shown. Indeed, various embodiments of the present invention can be embodied in many different forms, and should not be construed as being limited to the embodiments presented herein; in addition, these embodiments are provided so that this disclosure will meet appropriate legal requirements. The same reference numbers refer to the same elements. As used herein, the terms "data", "content", "information" and similar terms can be used interchangeably to indicate data that can be transmitted, received, and/or stored according to embodiments of the present invention. Moreover, the term "demonstration" as used herein is not provided to convey any qualitative assessment, but instead only conveys an example. In addition, the terms near and far are used in this text in a relative sense to indicate that an object is closer and farther away from a certain point with respect to another object, but does not otherwise indicate any specific or measurable location. Therefore, the use of any of these terms should not be construed as limiting the spirit and scope of the embodiments of the present invention.
Some embodiments of the present invention can provide a mechanism that can experience improvements related to gesture analysis. In this regard, for example, some embodiments may provide a real-time gesture analysis solution, which can be applied to interactive activities on handheld or other computing devices. Therefore, a user can control the device by gestures instead of manually operating a device (for example, the user's handheld or computing device, or even a remote device). Some exemplary embodiments may provide automatic gesture analysis through a solution that integrates various components such as a 3D camera, a depth analyzer, a motion analyzer, a target tracker, and a gesture recognizer. Target tracking according to some embodiments of the present invention can provide a relatively accurate target (for example, hand) position with relatively low sensitivity to background, lighting, hand size changes, and movement.
Target tracking can be implemented in an exemplary embodiment by a detection-based strategy. In this regard, for example, the target position in each frame can be determined based on skin detection and multiple useful clues such as size and position information (e.g., a previous frame). Detection-based tracking according to some embodiments can provide relatively accurate and fast tracking, which can be used in connection with real-time applications.
FIG. 1 shows an example of an adaptive gesture analysis process of a UI controller according to an exemplary embodiment of the present invention. It should be understood that although an exemplary embodiment will be described below in the context of hand gesture analysis based on target detection, for gesture analysis, other parts of the body may also be included. For example, for gesture analysis, arm positioning, foot positioning, etc. can also be considered, assuming that the arms, feet, etc. are exposed to achieve target detection. Furthermore, the process of FIG. 1 is only exemplary, and thus other embodiments may include, for example, the deletion of the same or different operations including additional or different operations, different orders, and/or detection of some operations.
As shown in FIG. 1, the first FIG. 1 is a flowchart showing various operations that can be performed in association with an exemplary embodiment. Image data (such as video data) can be received at operation 10 initially. When a camera may be part of or communicating with a device, the image data may be received from the camera associated with the device performing gesture recognition according to an exemplary embodiment. In some embodiments, the communication between the camera and other components used in gesture analysis may be instant or at least relatively delayed. In an exemplary embodiment, the camera may be a 3D camera 40 capable of providing 3D depth data and an intensity image at the same time, as shown in FIG. 2. It can be seen from Figure 2 that the 3D camera can provide data that can be separated into various depth intervals. The depth intervals may be equidistant or may have varying distances between them (for example, the near depth 42 and the far depth 44 may have a larger interval 46, and the intermediate distance 48 may have a smaller interval 50).
At operation 12, depth data and intensity data can be extracted from the image data collected by the 3D camera. A frame of the image data received at operation 10 can then be segmented into different depth ranges to provide intensity images at varying depth ranges at operation 14. The analysis of the image data can be carried out frame by frame, so that at operation 16, an intensity image of a previous (or subsequent) frame can be compared with the segmented images of each respective different depth range. At operation 18, the motion difference can be analyzed at each depth range. Because in many cases, it is expected that the target of the gesture may have the greatest change from one frame to the next, so motion analysis for adjacent frames can be used to identify a target area (for example, a gesture capable of Appendages, such as hands). Therefore, although it may not necessarily be able to predict the depth of the target, it can be predicted that compared to other depths, it is expected that more significant actions can be seen at the depth of the target.
Based on the action analysis for various depth ranges, candidate depth ranges can be identified at operation 20. In this regard, for example, an action displayed above a certain threshold, or as a maximum value or at least a candidate depth range higher than other depth ranges or a given depth range can be identified as a candidate depth range. In some exemplary embodiments, at operation 22, one or more cues (eg, location, size, etc.) may be considered together with the motion analysis of the candidate depth range. Based on the actions of the candidate depth ranges (and in some cases also based on the clues), a rough target depth range can be determined at operation 24. At operation 26, an updated average target depth may be determined (for example, by averaging all pixels in the target depth range). At operation 28, based on the finally determined target depth range, a target area can be determined and tracked.
Thereafter, the target can be tracked continuously, and the actions or changes in the features can be used for gesture analysis, and these features can be captured from the tracked target area (for example, the area of a hand). In an exemplary embodiment, gesture analysis can be performed by comparing features from the tracked target area with features in a feature storage database corresponding to a specific gesture. By judging that the features in the database (such as a matching database) match (or are substantially similar within a critical value) between the features extracted from the tracked target area, the corresponding can be identified In a gesture of the specific gesture, the specific gesture is associated with the matching features from the database.
If a specific gesture is recognized, then a corresponding command can be executed. In this way, for example, a database can store information that associates gestures with respective commands or UI functions. Thus, for example, if a clenched fist is recognized when playing music or video content, and the clenched fist is associated with a stop command, the demonstrated music or video content stops.
FIG. 3 is a schematic block diagram of an apparatus capable of implementing adaptive gesture analysis according to an exemplary embodiment of the present invention. An exemplary embodiment of the present invention will now be described with reference to FIG. 3, in which certain elements of a device capable of adaptive gesture analysis are displayed. The device in FIG. 3 can be used in, for example, a mobile terminal (or the mobile terminal 110 in FIG. 7) or various other mobile and fixed devices (such as a network device, a personal computer, a laptop, etc.). Alternatively, the embodiment can be used on a bonding device. Therefore, some embodiments of the present invention may be fully embodied at a single device (such as the mobile terminal 110), or through a device in a client/server relationship. Furthermore, it should be understood that the devices or elements described below may not be mandatory, and thus some devices or elements may be omitted in some embodiments.
Referring now to Figure 3, a device capable of adaptive gesture analysis is provided. The device may include a processor 70, a user interface 72, a communication interface 74 and a memory device 76 or otherwise communicate with the processor 70, the user interface 72, the communication interface 74 and the memory device 76 . The memory device 76 may include, for example, electrical and/or non-dependent memory. The memory device 76 can be configured to store information, data, application programs, commands, etc., used to enable the device to perform various functions according to the exemplary embodiment of the present invention. For example, the memory device 76 may be configured to buffer input data for processing by the processor 70. Additionally or alternatively, the memory device 76 may be configured to store instructions executed by the processor 70. As another alternative, the memory device 76 may be one of multiple databases storing information and/or media content.
The processor 70 can be embodied in a number of different ways. For example, the processor 70 may be embodied as various processing devices such as a processing element, an auxiliary arithmetic unit, and a controller, or include such as an ASIC (application-specific integrated circuit), an FPGA (field programmable gate array) , A variety of other processing equipment such as a hardware accelerator and other integrated circuits. In an exemplary embodiment, the processor 70 may be configured to execute instructions stored in the memory device 76 or accessible by the processor 70 in other ways.
At the same time, the communication interface 74 can be any device such as a device or circuit embodied in hardware, software, or a combination of hardware and software, and any device is configured to communicate with the device from a network And/or receive data from any other device or module, and/or transmit data to the network and/or any other device or module communicating with the device. In this regard, the communication interface 74 may include, for example, an antenna (or multiple antennas), and supporting hardware and/or software capable of communicating with a wireless communication network. In a fixed environment, the communication interface 74 may alternatively or similarly support wired communication. In this manner, the communication interface 74 may include a communication modem and/or other hardware/software for supporting communication via cable, digital subscriber line (DSL), universal serial bus (USB), or other mechanisms.
The user interface 72 can communicate with the processor 70 to receive a user input instruction at the user interface 72 and/or provide an audible, visible, mechanical or other output to the user. In this manner, the user interface 72 may include, for example, a keyboard, a mouse, a joystick, a display, a touch screen, a microphone, a speaker, or other input/output mechanisms. In an exemplary embodiment where the device is embodied as a server or some other network device, the user interface 72 can be restricted or excluded. However, in an embodiment where the device is embodied as a mobile terminal (such as the mobile terminal 110), the user interface 72 may include a speaker, a microphone, a display, and a keyboard in other devices or components. Any one or all of etc.
In an exemplary embodiment, the processor 70 may be embodied to include or otherwise control a depth analyzer 78, a motion analyzer 80, a target tracker 82, and a gesture recognizer 84. The depth analyzer 78, the motion analyzer 80, the target tracker 82, and the gesture recognizer 84 can be embodied in hardware, software, or a combination of hardware and software (for example, processing operations under software control). In the device 70), any device such as a device or circuit is configured to perform the corresponding functions of the depth analyzer 78, the motion analyzer 80, the target tracker 82, and the gesture recognizer 84, respectively , As described below. In an exemplary embodiment, the depth analyzer 78, the motion analyzer 80, the target tracker 82, and/or the gesture recognizer 84 may be respectively associated with a media capture module (such as the camera module in FIG. 7). 137) Communication to receive image data for analysis as described below.
The depth analyzer 78 can be configured to segment the input image data of each frame into data corresponding to each of various depth ranges. The depth analyzer 78 can then generate intensity images corresponding to each of the various depth ranges. In an exemplary embodiment, the depth analyzer 78 may be configured to separate the entire distance range into many small intervals, such as<i>D</i> ={<i>D</i><sub>1</sub><i>D</i><sub>2</sub> …<i>D</i><sub><i>N</i></sub>}. The equal intervals can be unequal, as shown in Figure 3. The depth analyzer 78 can then for each intensity frame<i>I</i><sub><i>1</i></sub>, Produce each depth range<img file="TWI489397B_D0001.tif" he="55" id="i0001" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="214" />A corresponding intensity image of<img file="TWI489397B_D0002.tif" he="55" id="i0002" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="24" />. For the generation of the intensity image of each depth range, the intensity of any point not under the depth range of the corresponding intensity image is set to zero. Therefore, only for each intensity image in the various depth ranges, the image intensity in the respective depth range is provided.
Figure 4 (including Figures 4A to 4I) shows an example of intensity images that can be generated by the depth analyzer 78 in different depth ranges of a frame. In this regard, Figure 4A shows a captured intensity image that has not been modified by the depth analyzer 78. Figures 4B to 4I show the intensity images produced in each of the various ranges that extend from the closest (at Figure 4B) the camera that produced the captured intensity image to the distance from the camera The farthest (at Figure 4I). In Figure 4B, the intensity image of the data closest to the camera detects a hand of the person focused in the intensity image. It means that all the image data of objects that are not at the depth of the hand have been set to zero. In Figure 4C, the intensity image represents the data at the approximate depth of the person's face and torso in the intensity image, and the data at all other depths are set to zero. In Figure 4D, the intensity image represents the data at the approximate depth of the person's chair, and all other data are set to zero. This process continues so that, for example, in Figure 4F, the intensity image represents data at the approximate depth of another persons chair , and all other data are set to zero, and in Figure 4H, the intensity image represents behind the other person For the data at the approximate depth of the workstation, set all other data to zero. In an exemplary embodiment, the intensity images generated by the depth analyzer 78 can be communicated with the motion analyzer 80 for continued processing.
The motion analyzer 80 can be configured to analyze the data in each depth range relative to adjacent data frames in the same corresponding depth range. Thus, for example, the motion analyzer 80 can compare an intensity image of a first frame (such as the frame in Figure 4B) and a second frame (such as an adjacent frame) in the same depth range according to the intensity image. Sequence phase comparison to detect the movement from one frame to the next. As shown above, compared with objects in these various depth ranges, a gesturing hand may cause more motions to be detected in subsequent frames. In this way, for example, if the person focused in the intensity image gestures to the camera, it is possible that an adjacent frame (previous or subsequent) of the frame in Figure 4B will be displayed compared to the Actions (or at least more actions) of other frames (such as those in Figures 4C to 4I), because these other frames display data corresponding to objects that are not moving, or at least less likely to display Move as much as the person's hand.
In an exemplary embodiment, the motion analyzer 80 can be configured to calculate in the current frame according to the following formula<sup><i>I</i></sup><sub><i>t</i></sub>With the previous frame<sup><i>I</i></sup><sub><i>t-1</i></sub>An image difference between each depth range<img file="TWI489397B_D0003.tif" he="109" id="i0003" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="734" />,in<i>W</i>and<i>H</i>Are the image width and height, and<img file="TWI489397B_D0004.tif" he="141" id="i0004" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="1381" />. The value<img file="TWI489397B_D0005.tif" he="51" id="i0005" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="101" />It can be considered as a relative action, and the action is given in each depth range<img file="TWI489397B_D0006.tif" he="53" id="i0006" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="100" />Example.
Figure 5 (which includes Figures 5A to 5N) illustrates an example of the operation of the motion analyzer 80 according to an exemplary embodiment. In this regard, Figure 5A shows an intensity image of a previous frame, and Figure 5B shows an intensity image of the current frame. Figures 5C to 5J show the intensity images in each of the four different depth ranges relative to the current and previous frames generated by the depth analyzer 78. In particular, Figure 5C shows an intensity image corresponding to the previous frame in a first depth range. Figure 5D shows an intensity image corresponding to the previous frame in a second depth range. Figure 5E shows an intensity image corresponding to the previous frame in a third depth range. Figure 5F shows an intensity image corresponding to the previous frame in a fourth depth range. At the same time, the 5G image represents an intensity image corresponding to the current frame in the first depth range. Figure 5H shows an intensity image corresponding to the current frame in the second depth range. Figure 51 shows an intensity image corresponding to the current frame in the third depth range. Figure 5J shows an intensity image corresponding to the current frame in the fourth depth range. Figures 5K to 5N show an output of the motion analyzer 80, the output indicating the relative motion from the previous frame to the current frame. In particular, the 5K image shows that under the first depth range, the relative movement from the previous frame to the current frame is not noticed. Fig. 5L shows that the relative movement from the previous frame to the current frame is noticed because the hand moves from the third depth range (see Fig. 5E) to the second depth range (see Fig. 5H). Figure 5M shows that a small amount of movement into the third depth range has been detected from the previous frame to the current frame because the hand has left the third depth range and entered the second depth range. Figure 5N shows that in the fourth depth range, there is substantially no relative motion between the current and previous frames.
From the calculation steps performed by the motion analyzer 80, it can be understood that the relative motion<img file="TWI489397B_D0007.tif" he="51" id="i0007" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="100" />Only the non-zero pixels in the specific depth range of the current frame are involved. Because the non-zero pixels correspond to the objects located in the specific depth range (such as the target or other objects), the relative action<img file="TWI489397B_D0008.tif" he="53" id="i0008" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="99" />It can be used as a measure to find the depth of the candidate target. In this regard, the depth at which the maximum relative motion (for example, the second depth range) is displayed is indicated in the 5K image, and the target (for example, the hand) is in the second depth range.
In some embodiments, it is also possible, for example, via<img file="TWI489397B_D0009.tif" he="92" id="i0009" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="669" />To calculate absolute actions. The absolute motion operation can be used to evaluate whether motion occurs in a certain depth range. These two measurements (relative action and absolute action) that provide an indication of an action change can be further used in the target depth determination, which will be described in more detail below.
The movement changes obtained using each depth range may distinguish the target from other objects. In some cases, one or more additional clues such as location, size, etc. can be used to help distinguish the target from other objects. The target tracker 82 can be configured to automatically determine (for example, based on the motion change between adjacent frames in various depth ranges) the depth of the target. In this regard, once the action change has been determined as described above, the target tracker 82 can be configured to be based on the relative action determined in the corresponding depth range, and may also be based on the additional clue(s) , To select one of the possible depth ranges (for example, the candidate depth range) as the target depth range. The target tracker 82 can then capture the target from the selected intensity image of the corresponding depth range, and track the target.
In an exemplary embodiment, the target tracker 82 can be configured to perform the following operations:
1) The target depth range determined in the previous frame<i>D</i><sub><i>k</i></sub>In, the image difference<img file="TWI489397B_D0010.tif" he="73" id="i0010" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="131" />With a predetermined threshold<i>T</i>Compared. if<img file="TWI489397B_D0011.tif" he="72" id="i0011" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="202" />, Then the depth of the target is considered unchanged.
2) if<img file="TWI489397B_D0012.tif" he="72" id="i0012" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="201" />, Then the depth of the target is deemed to have changed. Taking into account the continuity of the target action, the target tracker 82 can start from multiple adjacent depth ranges<i>D</i><sub><i>k</i></sub>Choose among<i>m</i>Candidate depth ranges. The target tracker 82 can then follow the relative action<img file="TWI489397B_D0013.tif" he="54" id="i0013" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="97" />To sort the depth ranges. Then, according to the highest<i>m</i>Piece<img file="TWI489397B_D0014.tif" he="56" id="i0014" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="97" />, You can select the corresponding depth ranges as candidates for the target depth.
3) Non-target objects can then be excluded to reveal the depth range of the target. In this regard, for example, the target tracker 82 can be configured to further analyze the multiple intensity images (corresponding to each candidate depth) by integrating multiple constraints (such as location information, size factors, etc.) At<img file="TWI489397B_D0015.tif" he="53" id="i0015" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="100" />) To get the target position. In an exemplary embodiment, the position may be expressed as the center of gravity of a certain object area. For the two adjacent frames, the size and position of the target area may not change significantly. Therefore, the position change can be used as an effective limiting condition for determining which object is the target to be tracked. In the example shown in Figure 6, the position change in a certain depth range can be defined as the distance between the center of gravity of the object in the current frame and the target position in the previous frame. Therefore, the object corresponding to the minimum position change and a similar size is determined as the target. Therefore, the object depth is considered to be the rough depth of the target<i>D</i><sub><i>k'</i></sub>, As shown in the bottom column of Figure 6 (Figure 6F).
4) Once the depth range of the target is determined<i>D</i><sub><i>k'</i></sub>, You can use a subsequent target segmentation formula to provide a more accurate target depth:<img file="TWI489397B_D0016.tif" he="101" id="i0016" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="324" />, Where n is the depth range<i>D</i><sub><i>k'</i></sub>The number of this pixel in it. Then the target pixel can be restricted by a depth condition:<img file="TWI489397B_D0017.tif" he="50" id="i0017" img-content="character" img-format="tif" inline="no" orientation="portrait" wi="475" />To get, where<i>d</i><sub><i>T</i></sub>Is the empirical threshold.
The target tracker 82 can be configured to track the target (for example, the user's hand) through the depth determination process and the segmentation of the intensity image containing the target. In order to achieve gesture recognition for one hand, the accurate position of the hand can improve the quality of the analysis and the output produced. Based on the aforementioned mechanism for determining the position of the hand, the hand tracking can be completed on a continuous frame to enable gesture detection.
Some embodiments may also use the gesture recognizer 84, which may be configured to perform gesture matching between features associated with the target and features associated with a particular known gesture. For example, a database of known gestures and their respective characteristics can be provided to compare with the characteristics of a current gesture. If the similarity between the compared gestures is sufficient, the gesture recognizer 84 can associate a current gesture with the specific known gesture, thereby identifying or recognizing the current gesture.
In an exemplary embodiment, the database of known gestures can be generated by the user (or by another user) in the offline phase. Therefore, multiple samples of each gesture can be collected to form a gesture gallery. In an exemplary embodiment, size normalization can be performed initially, and each sample can be converted into a feature vector according to the above scheme and recorded as a template for matching purposes. A recognized gesture can be used to trigger or cause the execution of a specific command associated with the recognized gesture. In this regard, for example, the gesture recognizer 84 can communicate the identity of a recognized gesture with the processor 70, and the processor 70 can execute (for example, via the user interface 72) a corresponding UI command. The command can be used to guide a UI system to perform a corresponding operation.
Fig. 6 (which includes Figs. 6A to 6F) shows an example of the entire procedure for determining the target depth. Figure 6A shows an intensity image of a previous frame, and Figure 6B shows an intensity image of the current frame. As shown in the series of images shown in FIG. 6C, the depth analyzer 78 determines the intensity image of each respective depth range for various different depths related to the previous frame. The depth analyzer 78 also determines the intensity images of each respective depth range for the respective different depths related to the current frame, as shown in a series of images in the 6D figure. The motion analyzer 80 then determines the motion changes between the current and previous frames in each depth range, as shown in a series of images in Figure 6E. The target tracker 82 then performs target depth range determination by integrating the motion and position information of each depth range, as shown in a series of images in Fig. 6F. As shown in Figure 6, the target (such as the human hand shown in the current and previous frames) is located at depth 3 in the previous frame. However, through the analysis of the adjacent depth range, in the current frame, it is determined that the new position of the target is at depth 4. The action in depth 2 is generated by another object (in this case, another person's hand), and it is excluded in order to take into account the constraints of the position.
Based on the above description, the embodiments of the present invention can provide image segmentation to locate a target (for example, a hand), so as to achieve robust tracking in an effective manner. Therefore, the relatively accurate target tracking result and hand gesture recognition rate can be improved. The use of the 3D camera can provide real-time 3D depth data, which can be used by embodiments of the present invention to eliminate or substantially reduce the influence of background and lighting on the accuracy of gesture recognition. The division of the depth range can also help the analysis of the content under the different depths of each frame. The motion calculation described in this article may realize the object motion capture in each depth range, and the object motion includes the motion of the target and other objects. By comparing the actions in different ranges and integrating multiple useful clues, the depth of the target can be automatically determined, so that the hand can be captured. Therefore, based on the accurate hand segmentation and tracking results, the gesture recognition accuracy can be improved. Thus, for example, tracking and identification performance can be improved, and the usability of interaction can also be improved.
An exemplary embodiment of the present invention will now be described with reference to FIG. 7, in which certain elements of a device capable of adaptive gesture analysis are shown. In this manner, FIG. 7 shows a block diagram of a mobile terminal 110 that can benefit from an exemplary embodiment of the present invention. However, it should be understood that the mobile terminal shown and described later is only an illustrative type of mobile terminal that can benefit from some embodiments of the present invention, and therefore should not be construed as limiting the present invention The scope of the embodiment. Such as portable digital assistants (PDA), pagers, mobile TVs, gaming devices, all types of computers (such as laptops or mobile computers), cameras, audio/video players, radios, global positioning system (GPS) devices Or any combination of multiple types of mobile terminals and other types of communication systems can easily use the embodiments of the present invention.
In addition, although the various embodiments of the method of the present invention can be executed or used in conjunction with a mobile terminal 110, the method can be executed or used in addition to a mobile terminal (such as a personal computer (PC), server, etc.) Equipment to use. Moreover, the system and method of the embodiment of the present invention may have been initially described in conjunction with mobile communication applications. However, it should be understood that the system and method of the embodiments of the present invention can be used in combination with various other applications in the mobile communication industry and industries outside the mobile communication industry.
The mobile terminal 110 may include an antenna 112 (or multiple antennas) in operative communication with a transmitter 114 and a receiver 116. The mobile terminal 110 may further include a device such as a controller 120 (or processor 70) or other processing elements. The device provides signals to the transmitter 114 and receives signals from the receiver 116, respectively. The signals may include information sent according to the air interface standard of the applicable cellular system, and/or may include data corresponding to voice, received data, and/or user generated/transmitted data. In this regard, the mobile terminal 110 can operate under one or more air interface standards, communication protocols, modulation types, and access types. By way of illustration, the mobile terminal 110 can operate according to any one of a plurality of first, second, third, and/or fourth generation communication protocols. For example, the mobile terminal 110 can be based on the second generation (2G) wireless communication protocol IS-136 (Time Division Multiple Access (TDMA), GSM (Global System for Mobile Communications)) and IS-95 (Code Division Multiple Access ( CDMA)), or according to third-generation (3G) wireless communication protocols such as Universal Mobile Telecommunications System (UMTS), CDMA2000, Wideband CDMA (WCDMA), and Time-sharing Synchronous CDMA (TD-SCDMA), according to protocols such as E-UTRAN (Evolution It operates according to the fourth-generation (4G) wireless communication protocol and so on. Alternatively (or additionally), the mobile terminal 110 can operate according to a non-cellular communication mechanism. For example, the mobile terminal 110 can communicate in a wireless local area network (WLAN) or other communication networks.
It should be understood that the device such as the controller 120 may include circuits for implementing audio/video and logic functions of the mobile terminal 110 in particular. For example, the controller 120 may include a digital signal processor device, a microprocessor device, and various analog-to-digital converters, digital-to-analog converters, and/or other support circuits. The control and signal processing functions of the mobile terminal 110 can be allocated among these devices according to their respective capabilities. The controller 120 may thus also include functions for encoding and interleaving messages and data before modulation and transmission. The controller 120 may additionally include an internal voice encoder, and may include an internal data modem. Moreover, the controller 120 may include functions for operating one or more software programs stored in the memory. For example, the controller 120 can operate a connection program, such as a conventional web browser. The connection program can then allow the mobile terminal 110 to transmit and receive web content, such as location-based content and/or other web content, according to, for example, a wireless application protocol (WAP), hypertext transfer protocol (HTTP), etc.
The mobile terminal 110 may also include a user interface that includes an output device such as a headset or speaker 124, a microphone 126, a display 128, and a user interface operatively coupled to the controller 120 Input interface. The user input interface that allows the mobile terminal 110 to receive data may include any one of multiple devices that allow the mobile terminal 110 to receive data, such as a keyboard 130, a touch display (not shown), or other input devices. In the embodiment including the keyboard 130, the keyboard 130 may include numbers (0-9) and related keys (#, *), and other hard and soft keys for operating the mobile terminal 110. Alternatively, the keyboard 130 may include a QWERTY keyboard arrangement. The keyboard 130 may also include various soft keys with associated functions. Additionally or alternatively, the mobile terminal 110 may include an interface device such as a joystick or other user input interface. The mobile terminal 110 further includes a battery 134 such as a vibrating battery pack for supplying power to various circuits for operating the mobile terminal 110 and optionally providing mechanical vibration as a detectable output.
The mobile terminal 110 may further include a user identity module (UIM) 138. The UIM 138 is typically a memory device with a built-in processor. The UIM 138 may include, for example, a user identity module (SIM), a universal integrated circuit card (UICC), a universal user identity module (USIM), a removable user identity module (R-UIM), etc. . The UIM 138 typically stores information elements related to a mobile user. The mobile terminal 110 can be equipped in addition to the UIM Memory other than 138. The mobile terminal 10 may include an electrical memory 140 and/or a non-electric memory 142. For example, the dependent memory 140 may include random access memory (RAM), which includes dynamic and/or static memory RAM, on-chip or off-chip cache memory, and the like. The embeddable and/or removable non-electrical memory 142 may include, for example, read-only memory, flash memory, magnetic storage devices (such as hard disks, floppy drives, tapes, etc.), optical disk drives, and/ Or media, non-dependent random access memory (NVRAM), etc. Like the electrical memory 140, the non-electric memory 142 may include a cache area for temporarily storing data. The memories can store any one of multiple pieces of information and data used by the mobile terminal 110 to implement the functions of the mobile terminal 110. For example, the memories may include an identifier such as an International Mobile Equipment Identification (IMEI) code, which can uniquely identify the mobile terminal 110. Furthermore, the memories can store commands used to determine cell id information. In particular, the memories can store an application program executed by the controller 120, and the application program determines an identity of the current cell that the mobile terminal 110 communicates with, that is, cell id identity or cell id information .
In an exemplary embodiment, the mobile terminal 110 may include a media capture module that communicates with the controller 120, such as a camera, video, and/or audio module. The media capture module can be any device used to capture an image, video, and/or audio for storage, display, or transmission. For example, in an exemplary embodiment where the media capture module is a camera module 137, the camera module 137 may include a digital camera capable of forming a digital image file from a captured image. In this way, the camera module 137 may include all hardware and software such as a lens or other optical devices required to create a digital image file from a captured image. In an exemplary embodiment, the camera module 137 may be a 3D camera capable of capturing 3D image information representing depth and intensity.
FIG. 8 is a flowchart of a system, method, and program product according to some exemplary embodiments of the present invention. It should be understood that each block or step of the flowchart and the combination of blocks in the flowchart can be implemented by various devices including one or more computer program instructions, such as hardware, firmware, and/or software. For example, one or more of the above-mentioned procedures can be embodied by computer program instructions. In this regard, the computer program instructions embodying the above-mentioned procedures can be stored in a mobile terminal or a memory device using other devices of the embodiments of the present invention, and stored in the mobile terminal or other device. One processor to execute. It should be understood that any of these computer program instructions can be loaded into a computer or other programmable device (ie, hardware) to generate a machine for execution on the computer (for example, via a processor) or other programmable device These instructions create a device for implementing the functions specified in the flowchart block or step(s). These computer program instructions can also be stored in a computer-readable memory, which can guide a computer (such as the processor or another computing device) or other programmable devices to function in a specific way , So that the instructions stored in the computer-readable memory produce a product, and the product includes an instruction device that implements the function specified in the block or step of the flowchart (etc.). The computer program instructions can also be loaded on a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device, and a computer-executed process is generated, so that it can be executed on the computer Or the instructions on other programmable devices provide steps for implementing the functions specified in the flowchart block or step(s).
Therefore, the blocks or steps of the flowchart support the combination of devices for performing the designated functions, the combination of steps for performing the designated functions, and the program instruction devices for performing the designated functions. It should also be understood that one or more blocks or steps of the flowchart, and the combination of the blocks or steps in the flowchart, can be the basic computer system by the specific-purpose hardware that performs the specified functions or steps, or Special purpose hardware and computer instructions are combined to implement.
In this regard, for example, an embodiment of a method for providing adaptive gesture analysis shown in FIG. 8 may include the following steps: at operation 200, a distance range is divided into multiple depth ranges, and in operation At 210, multiple intensity images of at least two image frames are generated. Each of the intensity images can provide image data representing the existence of objects in a corresponding depth range of a respective image frame. The method further includes the following steps: determining the motion change between the two image frames in each corresponding depth range at operation 220, and determining the depth of a target based at least in part on the motion change at operation 230.
In an exemplary embodiment, dividing the distance range includes selecting a wider distance interval between the short distance and the long distance than the distance interval between the short distance and the long distance. In some cases, generating the multiple intensity images may include the following steps: generating multiple intensity images of adjacent frames, or for each corresponding depth range of a specific frame, will be different from the current depth range The data corresponding to the object is set to a predetermined value, and the data corresponding to the current depth range is retained. In an exemplary embodiment, determining the action change may include the following steps: comparing the object data in a current frame with the object data in a previous frame to determine the change from the previous frame to the current frame The intensity of is changed, and data that does not indicate the intensity change is set to the predetermined value. In some embodiments, determining the depth of the target may include determining the depth of the target based on an integration of the movement change and cue information of the target position.
In an exemplary embodiment, the method may also include other optional operations, and some examples of these optional operations are shown in dotted lines in FIG. 8. In this regard, exemplary additional operations may include an operation 240 that includes tracking the target motion based on the determination of the target depth. In some embodiments, the method may further include recognizing the gesture feature of the target at operation 250 to initiate a user interface command based on a recognized gesture.
In an exemplary embodiment, a device for performing the method in Figure 8 above may include a processor (such as the processor 70) that is configured to perform the aforementioned operations (200-250) Some or every operation. The processor can, for example, be configured to perform the operations (200-250) by executing the logic functions implemented by the hardware, executing the stored instructions or executing the algorithms used to perform each operation of the operations (200-250) . Optionally, the device may include a device for performing each of the aforementioned operations. In this regard, according to an exemplary embodiment, an example of a device for performing operations 200-250 may include, for example, the processor 70, the depth analyzer 78, the motion analyzer 80, the target tracker 82, and the gesture recognition Each of the devices 84, or an algorithm executed by the processor to control the aforementioned gesture recognition, hand tracking, and depth determination.
Those with general knowledge in the art to which these inventions benefit from the teachings set forth in the foregoing description and associated drawings will think of many modifications and other embodiments of the present invention presented here. Therefore, it should be understood that the embodiments of the present invention are not limited to the specific embodiments disclosed, and the modifications and other embodiments are intended to be included in the scope of the appended patents. Moreover, although the foregoing descriptions and the associated drawings describe exemplary embodiments in the context of certain exemplary combinations of elements and/or functions, it should be understood that different combinations of elements and/or functions may be other combinations of elements and/or functions. Examples are provided without departing from the scope of these additional patent applications. In this regard, for example, in comparison with those explicitly described above, different combinations of elements and/or functions can also be considered as being proposed in some of the scopes of these patent applications. Although specific terms are used here, they are only in a general and descriptive sense, and not for restrictive purposes.
<p>10~28. . . operate</p><p>40. . . 3D camera</p><p>42. . . Near depth</p><p>44. . . Far depth</p><p>46. . . interval</p><p>48. . . Middle distance</p><p>50. . . interval</p><p>70. . . processor</p><p>72. . . user interface</p><p>74. . . Communication interface</p><p>76. . . Memory device</p><p>78. . . In-depth analyzer</p><p>80. . . Motion analyzer</p><p>82. . . Target tracker</p><p>84. . . Gesture recognizer</p><p>110. . . Mobile terminal</p><p>112. . . antenna</p><p>114. . . launcher</p><p>116. . . receiver</p><p>120. . . Controller</p><p>124. . . Headphones or speakers</p><p>126. . . microphone</p><p>128. . . monitor</p><p>130. . . keyboard</p><p>134. . . Battery</p><p>137. . . Camera module</p><p>138. . . User Identity Module (UIM)</p><p>140. . . Dependent memory</p><p>142. . . Non-electrical memory</p><p>200~250. . . operate</p>
Figure 1 shows an example of an adaptive gesture analysis process of a UI controller according to an exemplary embodiment of the present invention;
Figure 2 shows a diagram of different depth intervals dividing the entire depth range of a three-dimensional image according to an exemplary embodiment of the present invention;
Figure 3 shows a schematic block diagram of a device that can implement gesture analysis according to an exemplary embodiment of the present invention;
Figure 4 (including Figures 4A to 4I) shows intensity images corresponding to each of various different depths according to an exemplary embodiment of the present invention;
Figure 5 (which includes Figures 5A to 5N) shows the intensity images of adjacent frames and an indication of the relative motion between them according to an exemplary embodiment of the present invention;
Figure 6 (including Figures 6A to 6F) illustrates a stage of a process of determining the target depth according to an exemplary embodiment of the present invention;
Figure 7 shows a block diagram of a mobile terminal that may benefit from an exemplary embodiment of the present invention; and
FIG. 8 is a flowchart of an exemplary method for providing gesture analysis according to an exemplary embodiment of the present invention.
44 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| TWI811336B | Cited by | Taiwan Province of China | Examiner |
| TWI799700B | Cited by | Taiwan Province of China | Examiner |
| US11662253B2 | Cited by | United States of America | Applicant |
| US11543296B2 | Cited by | United States of America | Applicant |
| CN101166237A | Cites | China | Examiner |
| CN1433236A | Cites | China | Examiner |
| US2002175921A1 | Cites | United States of America | Examiner |
| WO2005114556A2 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| US20020175921A1 | Cites | United States of America | – |
| WO2005114556A2 | Cites | World Intellectual Property Organization (WIPO) | – |
12 members in 6 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 12261108 | United States of America | – | |
| 26110808 | United States of America | A | |
| 26110808 | United States of America | A | |
| 12261108 | – | – | – |
| US20080261108 | – | – | – |
Members12
| Document | Office | Kind | |
|---|---|---|---|
| US2010111358A1 | United States of America | A1 | |
| WO2010049790A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201019239A | Taiwan Province of China | A | |
| EP2344983A1 | European Patent Office (EPO) | A1 | |
| KR20110090973A | Republic of Korea | A | |
| CN102257511A | China | A | |
| US8325978B2 | United States of America | B2 | |
| KR101300400B1 | Republic of Korea | B1 | |
| CN102257511B | China | B | |
| TWI489397BThis record | Taiwan Province of China | B | |
| EP2344983A4 | European Patent Office (EPO) | A4 | |
| EP2344983B1 | European Patent Office (EPO) | B1 |
Numbers
- Publication
- I489397
- Publication, DOCDB
- I489397
- Publication, EPODOC
- TWI489397B
- Application
- 98136283
- Application, DOCDB
- 98136283
- Application, EPODOC
- TW200998136283
Titles2
- English
- METHOD, APPARATUS AND COMPUTER PROGRAM PRODUCT FOR PROVIDING ADAPTIVE GESTURE ANALYSIS
- Chinese
- 用於提供適應性手勢分析之方法、裝置及電腦程式產品
Classification
- CPC, 3
- G06T7/20
- G06V40/20
- G06V20/64
- IPC, 1
- G06K9 62