An adaptable framework for cloud assisted augmented reality
40 claims: 29 independent, 11 dependent
- 1方法であって、 モバイルプラットフォームを用いて画像データを獲得することと、ここで、前記画像データは、オブジェクトの少なくとも1つのキャプチャ画像からのものである、 前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングすることと、 以前に獲得された画像データと比べて前記画像データにおける変化を備えるトリガイベントが存在するか否かを判定することと、ここにおいて、前記トリガイベントは、以前のキャプチャ画像に対して前記少なくとも1つのキャプチャ画像中に異なるオブジェクトが現れるシーン変化を備える、 前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングし続けている間に、前記トリガイベントが存在する場合、前記画像データをサーバに送信することと、 前記画像データに関連する情報を前記サーバから受信することと、ここにおいて、前記画像データに関連する前記情報は、前記オブジェクトの二次元(2D)モデル、前記オブジェクトの三次元(3D)モデル、前記オブジェクト上の点の三次元座標推定、拡張情報、前記オブジェクトについての顕著性情報、および、オブジェクトマッチングに関連する情報のうちの少なくとも1つを備える、 を備えており、 前記オブジェクトに対する前記モバイルプラットフォームのポーズを取得することと、 前記ポーズ、および前記画像データに関連する前記情報を用いて、前記オブジェクトをトラッキングすることと を さらに 備える、方法。
- 2前記オブジェクトをトラッキングすることは、前記サーバから受信された前記オブジェクトの基準画像を用いることをさらに備える、請求項1に記載の方法。
- 3前記画像データを前記サーバに送信する前に前記少なくとも1つのキャプチャ画像の品質を決定することをさらに備え、前記画像データは、前記少なくとも1つのキャプチャ画像の前記品質が閾値よりも良い場合にのみ、前記サーバに送信される、請求項1に記載の方法。
- 4前記少なくとも1つのキャプチャ画像の前記品質を決定することは、前記少なくとも1つのキャプチャ画像の鮮明度を分析することと、前記少なくとも1つのキャプチャ画像内の検出されたコーナの数を分析することと、学習分類子とともに前記少なくとも1つのキャプチャ画像から得られた統計値を使用することと、のうちの少なくとも1つを備える、請求項3に記載の方法。
- 5前記サーバから受信された前記画像データに関連する前記情報に基づいて、前記オブジェクトに対して拡張をレンダリングすることをさらに備える、請求項1に記載の方法。
- 6前記画像データに関連する前記情報は、前記オブジェクトの識別を備える、請求項1に記載の方法。
- 7前記少なくとも1つのキャプチャ画像は、複数のオブジェクトを備え、前記画像データに関連する前記情報は、前記複数のオブジェクトの識別を備える、請求項1に記載の方法。
- 8前記モバイルプラットフォームに対する前記複数のオブジェクトの各々についてのポーズを取得することと、 前記ポーズ、および前記画像データに関連する前記情報を用いて、前記複数のオブジェクトの各々をトラッキングすることと をさらに備える、請求項7に記載の方法。
- 9前記画像データに関連する前記情報は、前記オブジェクトの基準画像を備え、前記ポーズを取得することは、前記少なくとも1つのキャプチャ画像および前記基準画像に基づいて第1のポーズを前記サーバから受信することを備える、請求項 1 に記載の方法。
- 10視覚ベースのトラッキングにより前記オブジェクトをトラッキングし続けることが、前記第1のポーズが前記サーバから受信されるまで前記オブジェクトの基準フリートラッキングを実行することを備える、請求項 9 に記載の方法。
- 11前記第1のポーズが前記サーバから受信されたとき、前記オブジェクトの第2のキャプチャ画像を獲得することと、 漸次的な変化を決定するために、前記少なくとも1つのキャプチャ画像と前記第2のキャプチャ画像との間で前記オブジェクトをトラッキングすることと、 前記オブジェクトに対する前記モバイルプラットフォームの前記ポーズを取得するために、前記漸次的な変化および前記第1のポーズを使用することと をさらに備える、請求項 9 に記載の方法。
- 12前記オブジェクトの第2のキャプチャ画像を獲得することと、 前記基準画像を用いて、前記第2のキャプチャ画像内の前記オブジェクトを検出することと、 前記オブジェクトに対する前記モバイルプラットフォームの前記ポーズを取得するために、前記基準画像および前記第2のキャプチャ画像において検出された前記オブジェクトを使用することと、 前記オブジェクトの基準ベースのトラッキングを初期化するために前記ポーズを使用することと をさらに備える、請求項 9 に記載の方法。
- 13前記シーン変化が存在するか否かを判定することは、 前記少なくとも1つのキャプチャ画像および前記以前のキャプチャ画像を用いて第1の変化メトリックを決定することと、 以前のトリガイベントからの第2の以前のキャプチャ画像と前記少なくとも1つのキャプチャ画像とを用いて第2の変化メトリックを決定することと、 前記少なくとも1つのキャプチャ画像についてのヒストグラム変化メトリックを生成することと、 前記シーン変化を判定するために、前記第1の変化メトリック、前記第2の変化メトリック、および前記ヒストグラム変化メトリックを使用することと を備える、請求項1に記載の方法。
- 14前記画像データに関連する前記情報は、オブジェクト識別を備え、前記方法は、さらに、 前記オブジェクトの追加のキャプチャ画像を獲得することと、 前記オブジェクト識別を用いて、前記追加のキャプチャ画像内の前記オブジェクトを識別することと、 前記オブジェクト識別に基づいて前記追加のキャプチャ画像のためのトラッキングマスクを生成することと、ここで、前記トラッキングマスクは、前記オブジェクトが識別される前記追加のキャプチャ画像内の領域を示す、 前記追加のキャプチャ画像の残りの領域を識別するために、前記トラッキングマスクを前記オブジェクトの前記追加のキャプチャ画像と共に使用することと、 前記追加のキャプチャ画像の前記残りの領域におけるシーン変化を備えるトリガイベントを検出することと を備える、請求項1に記載の方法。
- 15動きセンサデータ、位置データ、バーコード認識、テキスト検出結果、またはコンテキスト情報のうちの少なくとも1つを備えるセンサデータを獲得することと、前記画像データとともに前記センサデータを前記サーバに送信することと、をさらに備える、請求項1に記載の方法。
- 16前記コンテキスト情報は、ユーザ挙動、ユーザ選好、ロケーション、ユーザについての情報、時刻、および照明品質のうちの1つまたは複数を含む、請求項 15 に記載の方法。
- 17前記画像データは、異なる位置にあるカメラを用いてキャプチャされた前記オブジェクトの複数の画像からのものであり、前記方法は、前記オブジェクトに対する前記カメラのポーズの粗な推定を決定することと、前記画像データとともに前記ポーズの前記粗な推定を送信することとをさらに備え、前記サーバから受信された前記情報は、前記ポーズのリファインメントおよび前記オブジェクトの三次元モデルのうちの少なくとも1つを備える、請求項1に記載の方法。
- 18前記画像データは、異なる位置にあるカメラを用いてキャプチャされた前記オブジェクトの複数の画像からのものであり、前記サーバから受信された前記情報は、前記カメラに対する前記オブジェクトのポーズをさらに備える、請求項1に記載の方法。
- 19モバイルプラットフォームであって、 画像データを獲得するように適合されたセンサと、ここで、前記センサは、カメラであり、前記画像データは、オブジェクトの少なくとも1つのキャプチャ画像からのものである、 無線トランシーバと、 前記センサと前記無線トランシーバとに結合されたプロセッサとを備え、前記プロセッサは、前記センサを介して前記画像データを獲得し、前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングし、以前に獲得された画像データと比べて前記画像データにおける変化を備えるトリガイベントが存在するか否かを判定し、ここにおいて、前記トリガイベントは、以前のキャプチャ画像に対して前記少なくとも1つのキャプチャ画像中に異なるオブジェクトが現れるシーン変化を備え、前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングし続けている間に、前記トリガイベントが存在する場合、前記無線トランシーバを介して前記画像データを外部プロセッサに送信し、前記無線トランシーバを介して前記外部プロセッサから前記画像データに関連する情報を受信するように適合され、 ここにおいて、前記画像データに関連する前記情報は、前記オブジェクトの二次元(2D)モデル、前記オブジェクトの三次元(3D)モデル、前記オブジェクト上の点の三次元座標推定、拡張情報、前記オブジェクトについての顕著性情報、および、オブジェクトマッチングに関連する情報のうちの少なくとも1つを備 え、 前記プロセッサは、さらに、前記オブジェクトに対する前記モバイルプラットフォームのポーズを取得し、前記ポーズおよび前記画像データに関連する前記情報を用いて前記オブジェクトをトラッキングするように適合される、 モバイルプラットフォーム。
- 20前記プロセッサは、前記外部プロセッサから受信された前記オブジェクトの基準画像を用いて前記オブジェクトをトラッキングするようにさらに適合される、請求項 19 に記載のモバイルプラットフォーム。
- 21前記プロセッサは、前記画像データが前記外部プロセッサに送信される前に前記少なくとも1つのキャプチャ画像の品質を決定するようにさらに適合され、前記画像データは、前記少なくとも1つのキャプチャ画像の前記品質が閾値よりも良い場合にのみ、前記外部プロセッサに送信される、請求項 19 に記載のモバイルプラットフォーム。
- 22前記プロセッサは、前記少なくとも1つのキャプチャ画像の鮮明度の分析、前記少なくとも1つのキャプチャ画像内の検出されたコーナの数の分析、および前記少なくとも1つのキャプチャ画像から得られた統計値での学習分類子の処理のうちの少なくとも1つを実行するように適合されることによって、前記少なくとも1つのキャプチャ画像の前記品質を決定するように適合される、請求項 21 に記載のモバイルプラットフォーム。
- 23前記プロセッサは、さらに、前記無線トランシーバを介して受信された前記画像データに関連する前記情報に基づいて前記オブジェクトに対して拡張をレンダリングするように適合される、請求項 19 に記載のモバイルプラットフォーム。
- 24前記画像データに関連する前記情報は、前記オブジェクトの識別を備える、請求項 19 に記載のモバイルプラットフォーム。
- 25前記少なくとも1つのキャプチャ画像は、複数のオブジェクトを備え、前記画像データに関連する前記情報は、前記複数のオブジェクトの識別を備える、請求項 19 に記載のモバイルプラットフォーム。
- 26前記プロセッサは、さらに、前記モバイルプラットフォームに対する前記複数のオブジェクトの各々についてのポーズを取得し、前記ポーズおよび前記画像データに関連する前記情報を用いて前記複数のオブジェクトの各々をトラッキングするように適合される、請求項 25 に記載のモバイルプラットフォーム。
- 27前記画像データに関連する前記情報は、前記オブジェクトの基準画像を備え、前記プロセッサは、前記少なくとも1つのキャプチャ画像および前記基準画像に基づいて第1のポーズを前記外部プロセッサから受信するように適合される、請求項 19 に記載のモバイルプラットフォーム。
- 28前記プロセッサは、前記第1のポーズが前記外部プロセッサから受信されるまで、前記オブジェクトの基準フリートラッキングを実行するように適合されることによって、視覚ベースのトラッキングにより前記オブジェクトをトラッキングし続けるように構成される、請求項 27 に記載のモバイルプラットフォーム。
- 29前記プロセッサは、さらに、前記第1のポーズが前記外部プロセッサから受信されるとき、前記オブジェクトの第2のキャプチャ画像を獲得し、漸次的な変化を決定するために、前記少なくとも1つのキャプチャ画像と前記第2のキャプチャ画像との間で前記オブジェクトをトラッキングし、前記オブジェクトに対する前記モバイルプラットフォームの前記ポーズを取得するために前記漸次的な変化と前記第1のポーズとを使用するように適合される、請求項 27 に記載のモバイルプラットフォーム。
- 30前記プロセッサは、さらに、前記オブジェクトの第2のキャプチャ画像を獲得し、前記基準画像を用いて前記第2のキャプチャ画像内の前記オブジェクトを検出し、前記オブジェクトに対する前記モバイルプラットフォームの前記ポーズを取得するために前記基準画像および前記第2のキャプチャ画像において検出された前記オブジェクトを使用し、前記オブジェクトの基準ベースのトラッキングを初期化するために前記ポーズを使用するように適合される、請求項 27 に記載のモバイルプラットフォーム。
- 31前記プロセッサは、前記少なくとも1つのキャプチャ画像および前記以前のキャプチャ画像を用いて第1の変化メトリックを決定し、以前のトリガイベントからの第2の以前のキャプチャ画像と前記少なくとも1つのキャプチャ画像とを用いて第2の変化メトリックを決定し、前記少なくとも1つのキャプチャ画像についてのヒストグラム変化メトリックを生成し、前記シーン変化を判定するために、前記第1の変化メトリック、前記第2の変化メトリック、および前記ヒストグラム変化メトリックを使用するように適合されることによって、前記シーン変化が存在するか否かを判定するように適合される、請求項 19 に記載のモバイルプラットフォーム。
- 32前記画像データに関連する前記情報は、オブジェクト識別を備えており、前記プロセッサは、さらに、前記オブジェクトの追加のキャプチャ画像を獲得し、前記オブジェクト識別を用いて前記追加のキャプチャ画像内の前記オブジェクトを識別し、前記オブジェクト識別に基づいて前記追加のキャプチャ画像のためのトラッキングマスクを生成し、ここで、前記トラッキングマスクは、前記オブジェクトが識別される前記追加のキャプチャ画像内の領域を示しており、前記追加のキャプチャ画像の残りの領域を識別するために前記トラッキングマスクを前記オブジェクトの前記追加のキャプチャ画像と共に使用し、前記追加のキャプチャ画像の前記残りの領域におけるシーン変化を備えるトリガイベントを検出するように適合される、請求項 19 に記載のモバイルプラットフォーム。
- 33動きセンサデータ、位置データ、バーコード認識、テキスト検出結果、またはコンテキスト情報のうちの少なくとも1つを備えるセンサデータを獲得するように適合された少なくとも1つの追加のセンサをさらに備え、前記センサデータは、前記画像データとともに前記外部プロセッサに送信される、請求項 19 に記載のモバイルプラットフォーム。
- 34前記コンテキスト情報は、ユーザ挙動、ユーザ選好、ロケーション、ユーザについての情報、時刻、および照明品質のうちの1つまたは複数を含む、請求項 33 に記載のモバイルプラットフォーム。
- 35前記画像データは、異なる位置にある前記カメラを用いてキャプチャされた前記オブジェクトの複数の画像からのものであり、前記プロセッサは、前記オブジェクトに対する前記カメラのポーズの粗な推定を決定し、前記画像データとともに前記ポーズの前記粗な推定を送信するようにさらに構成され、前記外部プロセッサから受信される前記情報は、前記ポーズのリファインメントおよび前記オブジェクトの三次元モデルのうちの少なくとも1つをさらに備える、請求項 19 に記載のモバイルプラットフォーム。
- 36前記画像データは、異なる位置にあるカメラを用いてキャプチャされた前記オブジェクトの複数の画像からのものであり、前記外部プロセッサから受信された前記情報は、前記カメラに対する前記オブジェクトのポーズをさらに備える、請求項 19 に記載のモバイルプラットフォーム。
- 37モバイルプラットフォームであって、 画像データを獲得する手段と、ここで、画像データを獲得する前記手段は、カメラであり、前記画像データは、オブジェクトの少なくとも1つのキャプチャ画像からのものである、 前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングする手段と、 以前に獲得された画像データと比べて前記画像データにおける変化を備えるトリガイベントが存在するか否かを判定する手段と、ここにおいて、前記トリガイベントは、以前のキャプチャ画像に対して前記少なくとも1つのキャプチャ画像中に異なるオブジェクトが現れるシーン変化を備え、 前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングし続けている間に、前記トリガイベントが存在する場合、前記画像データをサーバに送信する手段と、 前記画像データに関連する情報を前記サーバから受信する手段と、ここにおいて、前記画像データに関連する前記情報は、前記オブジェクトの二次元(2D)モデル、前記オブジェクトの三次元(3D)モデル、前記オブジェクト上の点の三次元座標推定、拡張情報、前記オブジェクトについての顕著性情報、および、オブジェクトマッチングに関連する情報のうちの少なくとも1つを備える、 を備えており、 前記オブジェクトに対する前記モバイルプラットフォームのポーズを取得する手段と、 前記ポーズ、および前記画像データに関連する前記情報を用いて、前記オブジェクトをトラッキングする手段と を さらに 備える、モバイルプラットフォーム。
- 38前記オブジェクトをトラッキングする前記手段は、前記サーバから受信された前記オブジェクトの基準画像をさらに用いる、請求項 37 に記載のモバイルプラットフォーム。
- 39プログラムコードを記憶した非一時的なコンピュータ読取可能な媒体であって、 画像データを獲得するためのプログラムコードと、ここで、前記画像データは、オブジェクトの少なくとも1つのキャプチャ画像からのものである、 前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングするためのプログラムコードと、 以前に獲得された画像データと比べて前記画像データにおける変化を備えるトリガイベントが存在するか否かを判定するためのプログラムコードと、ここにおいて、前記トリガイベントは、以前のキャプチャ画像に対して前記少なくとも1つのキャプチャ画像中に異なるオブジェクトが現れるシーン変化を備える、 前記オブジェクトの前記少なくとも1つのキャプチャ画像を用いる視覚ベースのトラッキングにより前記オブジェクトをトラッキングし続けている間に、前記トリガイベントが存在する場合、前記画像データを外部プロセッサに送信するためのプログラムコードと、 前記外部プロセッサから前記画像データに関連する情報を受信するためのプログラムコードと、ここにおいて、前記画像データに関連する前記情報は、前記オブジェクトの二次元(2D)モデル、前記オブジェクトの三次元(3D)モデル、前記オブジェクト上の点の三次元座標推定、拡張情報、前記オブジェクトについての顕著性情報、および、オブジェクトマッチングに関連する情報のうちの少なくとも1つを備える、 を備えており、 前記オブジェクトに対するモバイルプラットフォームのポーズを取得するためのプログラムコードと、 前記ポーズ、および前記画像データに関連する前記情報を用いて、前記オブジェクトをトラッキングするためのプログラムコードと を さらに 備える、非一時的なコンピュータ読取可能な媒体。
- 40前記オブジェクトをトラッキングするための前記プログラムコードは、さらに、前記外部プロセッサから受信された前記オブジェクトの基準画像を用いる、請求項 39 に記載の非一時的なコンピュータ読取可能な媒体。
Independent claims40
70 paragraphs, as filed
0001U.S. Patent Application No. 61/384667, entitled "An Adaptable Framework For Cloud Assisted Augmented Reality," filed September 20, 2010, both of which are transferred to the assignee of the present application and incorporated herein by reference. Claims the priority of issue and US patent application 13/235847 entitled "An Adaptable Framework For Cloud Assisted Augmented Reality" filed on September 19, 2011.
0002Augmented reality systems can insert virtual objects into the user's field of view in the real world. There can be many components in a traditional AR system. These include data acquisition, data processing, object detection, object tracking, registration, refinement, and rendering components. These components can interact with each other to provide the user with a rich AR experience. However, some components in detection and tracking in traditional AR systems use computationally intensive operations, which can interfere with the AR experience for the user.
0003The mobile platform uses distributed processing, where latency-sensitive operations are performed on the mobile platform and non-latency-sensitive but computationally intensive operations are performed on remote servers. Efficiently process sensor data including image data. The mobile platform acquires sensor data such as image data and determines if there is a trigger event that sends the sensor data to the server. A trigger event is a change in sensor data relative to previously acquired sensor data, such as a scene change in a captured image. When changes are present, sensor data is sent to the server for processing. The server processes the sensor data and returns information related to the sensor data, such as identification of objects in the image. The mobile platform can then use the identified objects to perform reference based tracking.
0004In one embodiment, the method is to acquire the sensor data using the mobile platform and determine if there is a trigger event with a change in the sensor data compared to the previously acquired sensor data. Includes sending sensor data to the server in the presence of a trigger event and receiving information related to the sensor data from the server. The sensor data can be a captured image of the object, eg, a photo or video frame.
0005In another implementation, the mobile platform includes a sensor adapted to acquire sensor data and a wireless transceiver. The sensor can be, for example, a camera for capturing an image of an object. The processor is coupled to the sensor and the wireless transceiver to acquire the sensor data through the sensor and determine if there is a trigger event with a change in the sensor data compared to the previously acquired sensor data. If a trigger event is present, it is adapted to send sensor data to an external processor via a wireless transceiver and receive information related to the sensor data from the external processor via a wireless transceiver.
0006In another embodiment, the mobile platform has a means of acquiring sensor data, a means of determining if there is a trigger event with a change in the sensor data compared to previously acquired sensor data, and a trigger event. Includes means for transmitting sensor data to the server when is present and means for receiving information related to the sensor data from the server. The means for acquiring the sensor data is a camera, and the sensor data is a captured image of an object.
0007In yet another embodiment, the non-temporary computer-readable medium containing the program code comprises the program code for acquiring the sensor data and changes in the sensor data compared to previously acquired sensor data. The program code for determining whether or not the trigger event exists, the program code for transmitting the sensor data to the external processor when the trigger event exists, and the information related to the sensor data are received from the external processor. Includes program code for.
0008<figref num="1">Figure 1 shows a block diagram showing a system for distributed processing, including mobile platforms and remote servers.</figref><figref num="2">FIG. 2 is a flow chart showing a distributed processing process in which latency-sensitive operations are performed by a mobile platform and latency-insensitive but computationally intensive operations are performed by an external processor.</figref><figref num="3">Figure 3 shows a block diagram of system operations for server-assisted AR.</figref><figref num="4">Figure 4 shows a call flow diagram for server-assisted AR, where pauses are provided by the remote server.</figref><figref num="5">Figure 5 shows another call flow diagram for server-assisted AR where pauses are not provided by the remote server.</figref><figref num="6">FIG. 6 shows a flow chart of the method performed by the scene change detector.</figref><figref num="7">FIG. 7 is a chart showing the performance of a distributed processing system showing the required network transmission as a function of the minimum trigger gap.</figref><figref num="8">Figure 8 shows an approach to face recognition using a server-assisted AR process.</figref><figref num="9">Figure 9 shows an approach to face recognition using a server-assisted AR process.</figref><figref num="10">Figure 10 shows an approach to visual search using a server-assisted AR process.</figref><figref num="11">Figure 11 shows an approach to visual search using a server-assisted AR process.</figref><figref num="12">Figure 12 shows an approach to criteria-based tracking using server-assisted processes.</figref><figref num="13">Figure 13 shows an approach to criteria-based tracking using server-assisted processes.</figref><figref num="14">Figure 14 shows an approach to 3D model creation using server-assisted processes.</figref><figref num="15">FIG. 15 is a block diagram of a mobile platform capable of distributed processing using server based detection.</figref>
Detailed explanation
0009The distributed processing system disclosed herein may determine when to deliver the data to be processed to a server over a wireless network or to another device over a network in a cloud computing environment. Including devices. The device can also process data on its own. For example, for more efficient processing, latency-sensitive operations may be selected to be performed on the device, and latency-insensitive operations may be selected to be performed remotely. .. Factors that determine when to send the data to be processed to the server are, among other things, whether the operations being performed on the data are latency sensitive or insensitive, and the amount of computation required. , Processor speed / availability on either device or server, network conditions, or quality of service.
0010In one embodiment, a system that includes a mobile platform and an external server is provided for Augmented Reality (AR) applications, in which latency-sensitive operations are performed for efficient processing. Latency-insensitive but computationally intensive operations that are performed on mobile platforms are performed remotely, eg, on servers. The result can then be sent by the server to the mobile platform. By using distributed processing for AR applications, end users can seamlessly enjoy the AR experience.
0011As used herein, the mobile platform is a cellular or other wireless communication device, personal communication system (PCS) device, personal navigation device (PND), personal information manager (PIM), personal digital assistant device (PDA), Or refers to any portable electronic device such as any other suitable mobile device. The mobile platform can receive wireless communications and / or navigation signals such as navigation positioning signals. The term "mobile platform" also refers to short distances, regardless of whether satellite signal reception, assistive data reception, and / or location-related processing is performed on the device or personal navigation device (PND). It is intended to include devices that communicate with a PND, such as by wireless, infrared, wired, or other connection. The "mobile platform" is also intended to include all electronic devices, including wireless communication devices, computers, laptops, tablet computers, etc. with AR capabilities.
0012FIG. 1 shows a block diagram showing a system 100 for distributed processing using server-based object detection and identification. System 100 includes a mobile platform 110 that performs latency-sensitive operations such as tracking, and remote server 130 performs latency-insensitive but computationally intensive operations such as object identification. The mobile platform may include a camera 112 and a display 114, and / or a motion sensor 164. The mobile platform 110 can acquire an image 104 of the object 102, which can be shown on the display 114. The image 104 captured by the mobile platform 110 can be a single frame from a still image, eg, a photo, or a video stream, both of which are captured herein. It is called image). The mobile platform 110 additionally or alternatives includes, for example, a satellite positioning system (SPS) receiver 166, or, for example, an accelerometer, a gyroscope, an electronic compass, or other similar motion sensing element. Multiple motion sensors 164 can also be used to obtain other sensor data, including position and / or orientation data, from sensors other than camera 112. SPS can be used for Global Positioning Systems (GPS), Global Navigation Satellite Systems (GNSS) such as Galileo, Gronas or Compass, or, for example, the Quasi-Vertical Satellite System (QZSS) over Japan, over India. Can it be associated with one Indian Regional Navigation Satellite System (IRNSS), various other regional systems such as Beidou over China, and / or one or more global and / or regional navigation satellite systems? Various expansion systems that can otherwise be used with them (eg, Satellite Based Augmentation (SBAS)) It can be a system)) constellation.
0013The mobile platform 110 transmits acquired data information such as captured image 104 and / or sensor data such as SPS information or location information from the onboard motion sensor 164 to the server 130 via the network 120. The acquired data information is also, or alternative, contextual, such as the identification of any object currently tracked by the mobile platform 110. data) can be included. The network 120 can be any wireless communication network such as a wireless wide area network (WWAN), a wireless local area network (WLAN), a wireless personal area network (WPAN), and the like. The server 130 processes the data information provided by the mobile platform 110 and generates information related to the data information. For example, server 130 may use object database 140 to perform object detection and identification based on the image data provided. Server 130 returns information related to the acquired data to mobile platform 110. For example, when the server 130 identifies an object from the image data provided by the mobile platform 110, the server 130 identifies the object, including, for example, an identifier such as a title or identification number, or a reference image 106 of the object 102. It can be returned with any desired side information such as saliency indicators, information links, etc. that can be used by mobile platforms for augmented reality applications.
0014If desired, the server 130 poses (positions) of the mobile platform 110 at the time the image 104 is captured, as compared to, for example, the object 102 in the reference image 106, which is an image of the object 102 from a known position and orientation. And orientation) can be determined and provided to mobile platform 110. The returned pose can be used to bootstrap the tracking system within the mobile platform 110. In other words, the mobile platform 110, from the time it captures the image 104 to the time it receives the reference image 106 and the pose from the server 130, for example, visually or using the motion sensor 164, everything in that pose. Can track gradual changes in. The mobile platform 110 can then use the received pose with a gradual change in the tracked pose to quickly determine the current pose for object 102.
0015In another embodiment, the server 130 returns a reference image 106, but does not provide pose information, and the mobile platform 110 uses an object detection algorithm to capture the object 102's current capture of the object 102's reference image 106. The current pose for object 102 is determined by comparing the images. Pauses can be used as inputs to the tracking system so that relative movements can be estimated.
0016In yet another embodiment, the server 130 returns only pose information and does not provide a reference image. In this case, the mobile platform 110 can use the captured image 104 with pose information to create a reference image that can later be used by the tracking system. Alternatively, the mobile platform 110 tracks the gradual change in position between the captured image 104 and the subsequent captured image (also known as the current image) and gradually tracks the poses taken from the server 130. It can also be used with the results to calculate the pose of the current image with respect to the reference image generated by the mobile platform. In the absence of the reference image 102, the current image can be warped (or modified) with the estimated pose to obtain an estimate of the reference image that could be used to bootstrap the tracking system. sell).
0017In addition, to minimize the frequency of discovery requests sent by the mobile platform 110 to the server 130, the mobile platform 110 can initiate discovery requests only in the presence of a trigger event. The trigger event can be based on a change in sensor data from the image data or motion sensor 164 with respect to previously acquired image data or sensor data. For example, the mobile platform 110 can use the scene change detector 304 to determine whether or not a change in image data has occurred. Thus, in some embodiments, the mobile platform 110 can communicate with the server 130 over the network for detection requests only when triggered by the scene change detector 304. The scene change detector 304 triggers communication with the server for object detection only, for example, when new information is present in the current image.
0018FIG. 2 is a flowchart showing a distributed processing process in which latency-sensitive operations are performed by the mobile platform 110 and non-latency-sensitive but computationally intensive operations are performed by an external processor such as server 130. Is. As shown, sensor data is acquired by mobile platform 110 (202). The sensor data can be information that includes captured images, such as captured photo or video frames, or character recognition or extracted key points derived from them. Sensor data can also, or alternatively, include, for example, SPS information, motion sensor information, barcode recognition, text detection results, or other results obtained from partial processing of images, as well as user behavior, users. It can include contextual information such as preferences, location, user information or data (eg, social network information about the user), time, lighting quality (natural vs. artificial), and people standing nearby (in the image). ..
0019The mobile platform 110 determines that a trigger event, such as a change in sensor data relative to previously acquired sensor data, is present (204). For example, a trigger event can be a scene change in which a new or different object appears in the image. The acquired sensor data is sent to server 130 after a trigger event such as a scene change is detected (206). Of course, if no scene change is detected, the sensor data does not need to be sent to the server 130, thereby reducing communication and detection requests.
0020The server 130 processes the acquired information to perform, for example, object recognition, which is well known in the art. After the server 130 processes the information, the mobile platform 110 receives information related to the sensor data from the server 130 (208). For example, the mobile platform 110 may receive the result of object identification, including, for example, a reference image. Information related to sensor data is additionally or alternative to items placed near the mobile platform 110 (eg, buildings, restaurants, available products in stores, etc.), as well as two-dimensional (2D) from the server. ) Or information such as a three-dimensional (3D) model, or information that can be used in other processes such as gaming. If desired, additional information may be provided, including the pose of the mobile platform 110 with respect to the objects in the reference image at the time the image 104 was captured, as discussed above. If the mobile platform 110 includes a local cache, the mobile platform 110 can store a plurality of reference images sent by the server 130. These stored reference images can be used, for example, for subsequent rediscovery that can be performed on the mobile platform 110 if tracking is lost. In some embodiments, the server identifies multiple objects in the image from the sensor. In such an embodiment, a reference image or other object identifier can be sent to the mobile platform 110 for only one of the identified objects, or multiple object identifiers corresponding to each object. Can be transmitted to and received by the mobile platform 110.
0021In this way, the information that can be provided by the server 130 is the recognition result, information about the identified object, a reference image (s) about the object that can be used for various functions such as tracking, and the recognized object. Can contain 2D / 3D models of, the absolute pose of the recognized object, extended information used for display, and / or saliency information about the object. In addition, server 130 can send information related to object matching that can enhance classifiers on mobile platform 110. One possible example is when the mobile platform 110 uses a decision tree for matching. In this case, server 130 can send values for individual nodes in the tree to facilitate more precise tree construction and subsequent better matching. Examples of decision trees are, for example, k-means, kd tree, vocabulary tree. tree), and other trees. In the case of k-means tree, server 130 also sends seeds to initialize the hierarchical k-means tree structure on mobile platform 110, which causes mobile platform 110 to load the appropriate tree. Allows you to perform lookups to do so.
0022Optionally, the mobile platform 110 may acquire a pose for the mobile platform with respect to object 102 (210). For example, the mobile platform 110 captures another image of object 102 and compares the newly captured image with the reference image to pose for the object in the reference image without receiving pose information from the server 130. Can be obtained. If the server 130 provides pose information, the mobile platform takes the pose provided by the server 130, which is the pose of the mobile platform 110 with respect to the object in the reference image at the time the initial image 104 was captured, the initial image 104. Combined with the tracked changes in the poses of the mobile platform 110 since it was captured, the current pose can be determined quickly. Note that whether poses are taken with or without the help of server 130 can depend on the capabilities of network 120 and / or mobile platform 110. For example, if server 130 supports pose estimation and the mobile platform 110 and server 130 agree to an application programming interface (API) for sending poses, pause information is sent to mobile platform 110 for tracking purposes. Can be used. The pose (210) of object 102 sent by the server can be in the form of a relative rotation and transformation matrix, a homography matrix, an affine transformation matrix, or any other form.
0023Optionally, the mobile platform 110 then uses the data received from the server 130 to track the AR, eg, the target, to estimate the object pose at each frame, and to virtual objects. You can insert or extend the user view or image through the rendering engine with the estimated pose (212).
0024Figure 3 shows a block diagram of the operation of System 100 for Server 130 assisted AR. Reference-free, as shown in Figure 3. A new captured image 300 is used to launch tracker) 302. The reference free tracker 302 performs tracking based on optical flow, normalized cross-correlation (NCC), or any similar method known in the art. The reference free tracker 302 identifies features such as points, lines, and regions within the new captured image 300 and tracks these features from frame to frame, for example using a flow vector. The flow vector obtained from the tracking result is useful for estimating the relative movement between the previous captured image and the current captured image, and then for identifying the speed of movement. The information provided by the reference free tracker 302 is received by the scene change detector 304. The scene change detector 304, for example, features tracked features from the reference free tracker 302, other types of image statistics (such as histogram statistics), as well as other available from sensors within the mobile platform. Use with information to estimate changes in the scene. If no trigger is sent by the scene change detector 304, the process continues with the reference free tracker 302. If the scene change detector 304 identifies a substantial change in the scene, the scene change detector 304 is a server based detector. detector) 308 sends a trigger signal that can start the detection process. If desired, image quality estimator 306 can be used to analyze image quality and further control the transmission of requests to server-based detector 308. The image quality estimator 306 examines the quality of the image and triggers a detection request if the quality is good, i.e. above the threshold. If the image quality is poor, the detection will not be triggered and the image will not be sent to the server-based detector 308. In one embodiment of the invention, the mobile platform 110 waits for a good quality image for a finite period of time after a scene change is detected before sending a good quality image to the server 130 for object recognition. Can be done.
0025Image quality can be based on known image statistics, image quality measurements, and other similar approaches. For example, the sharpness of a captured image can be quantified by high-pass filtering as well as generating a set of statistics representing, for example, edge intensity and spatial distribution. An image is classified as a good quality image if the sharpness value exceeds, or is comparable to, the "prevailing sharpness" of the scene, for example averaged over several previous frames. Can be done. In another implementation, FAST (Features from Accelerated Segment) Fast corner detection algorithms, such as Test corners or Harris corners, can be used to analyze the image. If there are enough corners, for example, the number of detected corners exceeds the threshold, or is more than the "general number of corners" of the scene, for example averaged over some previous frames. If many or comparable, the image can be classified as a good quality image. In another implementation, statistics from the image, such as the mean or standard deviation of the magnitude of the edge gradient, can be used to distinguish between good quality and poor quality images. Learning classifier ) Can be used to inform.
0026Image quality can also be measured using sensor inputs. For example, an image captured by the mobile platform 110 while moving fast can be blurry, so it is of poorer quality than if the mobile platform 110 is stationary or moving slowly. Sometimes. Therefore, motion from sensor data, for example from motion sensor 164 or from visual-based tracking, to determine if the resulting camera image is of sufficient quality to be sent for object detection. Estimates can be compared to thresholds. Similarly, image quality can be measured based on a determined amount of image blur.
0027In addition, a trigger time manager 305 may be provided to further control the number of requests sent to the server-based detector 308. Trigger time manager 305 maintains the state of the system and can be based on heuristics and rules. For example, if the number of images from the last trigger image is greater than a threshold, eg 1000 images, the trigger time manager 305 will time out and automatically start the detection process on the server-based detector 308. Yes, you can generate a trigger. Therefore, if there is no trigger for the expanded number of images, the trigger time manager 305 can force a trigger, which is useful for determining if an additional object is in the camera's field of view. In addition, the trigger time manager 305 can be programmed to maintain the minimum spacing between two triggers at the selected value η. That is, the trigger time manager 305 suppresses the trigger if the trigger is within the image of η from the last triggered image. Separating the triggered image can be useful, for example, when the scene is changing at high speed. Therefore, if the scene change detector 304 produces more than one trigger within the image of η, only one trigger image will be sent to the server-based detector 308, thus communicating from the mobile platform 110 to the server 130. Reduce the amount. The trigger time manager 305 can also manage the trigger schedule. For example, if the scene change detector 304 generates a new trigger less than η and more than μ before the last trigger, the new trigger is remembered and the image gap between consecutive triggers is It can be postponed by trigger time manager 305 until at least η. As an example, μ can be two images, η μ, and, for example, η can vary as 2, 4, 8, 16, 32, 64.
0028Trigger time manager 305 can also manage detection failures on server 130. For example, if a previous server-based detection attempt was unsuccessful, the trigger time manager 305 may periodically generate a trigger to retransmit the request to the server-based detector 308. Each of these attempts may use a different query image based on the most recent captured image. For example, after a detection failure, a periodic trigger can be generated by the trigger time manager 305 with a time gap η, for example, if the last failed detection attempt precedes the image of η. Is sent, where the value of η can be a variable.
0029When the server-based detector 308 is started, the data associated with the new captured image 300 is provided to the server 130. It may include the new captured image 300 itself, information about the new captured image 300, as well as sensor data associated with the new captured image 300. When an object is identified by server-based detector 308, the found object, such as a reference image, a 3D model of the object, or other relevant information is provided to mobile platform 110, which updates its local cache 310. .. If the server-based detector 308 does not find the object, the process can return to a periodic trigger, for example, using the trigger time manager 305. If no object is found after Г attempts, for example four attempts, the object is considered non-existent in the database and the system resets to a scene change detector based trigger.
0030With the found objects stored in the local cache 310, the object detector 312 running on the mobile platform 110 detects the objects in the current camera's field of view and the poses for them. Run the process and reference based on object identities and poses tracker) Send to 314. The poses and object identities sent by the object detector 312 can be used to initialize and start the reference base tracker 314. In each of the images subsequently captured (eg, a frame of video), the reference-based tracker 314 provides a pose for the object to the rendering engine within the mobile platform 110, which can be on top of or on the displayed object. Make the desired extension in the image. In one implementation, the server-based detector 308 can send a 3D model of the object instead of the reference image. In such cases, the 3D model is stored in the local cache 310 and later used as input to the reference base tracker 314. After the reference base tracker 314 is initialized, the reference base tracker 314 receives each of the new captured images 300 and identifies the location of the tracked object in each of the new captured images 300, thereby extending. Allows data to be displayed for tracked objects. Criteria-based tracker 314 can be used for many applications such as pose estimation, face recognition, building recognition, or other applications.
0031In addition, after the reference base tracker 314 is initialized, the reference base tracker 314 identifies each area of the new captured image 300 where the identified object resides, and this information is a means for the tracking mask. Remembered by. Therefore, an area within the new camera image 300 where the system has complete information about it is identified and provided as input to the reference free tracker 302 and the scene change detector 304. The reference free tracker 302 and the scene change detector 304 receive each of the new captured images 300 and use the tracking mask to operate in the remaining area of the new captured image 300, that is, in the area where complete information does not exist. Continue to do that. Using the tracking mask as feedback not only helps reduce false triggers from the scene change detector 304 by the tracked object, but also in the calculation of the reference free tracker 302 and the scene change detector 304. It also helps reduce complexity.
0032In one embodiment, the server-based detector 308 can additionally provide pose information for the object in the new captured image 300 with respect to the object in the reference image, as shown by the dotted line in FIG. .. The pose information provided by the server-based detector 308 can be used by the pose updater 316 to generate updated poses, along with the pose changes determined by the reference free tracker 302. The updated pose can then be provided to the reference base tracker 314.
0033In addition, if tracking is temporarily lost, subsequent rediscovery can be performed using the local detector 318, which searches the local cache 310. Figure 3 shows the local detector 318 and the object detector 312 separately for clarity, but if desired, the local detector 318 implements the object detector 312, i.e. the object detector. 312 can perform rediscovery. If the object is found in the local cache, the object identity is used to reinitialize and start the baseline base tracker 314.
0034FIG. 4 shows a call flow diagram for server-assisted AR in which pauses are provided by server 130, as shown by the dashed line and pause updater 316 in FIG. When the scene change detector 304 indicates that the field of view has changed (step A), the system manager 320 asks the server-based detector 308 for new images, which may be in, for example, jpeg or other formats, and object detection. (Step B) initiates the server-based discovery process. In addition, sensor data including information related to images and information from sensors such as SPS, orientation sensor display values, gyros, compasses, pressure sensors, altitude meters, and user data, such as application usage data, user profiles. Additional or alternative information such as, social network information, past searches, location / sensor information, etc. may also be sent to detector 308. System manager 320 also sends a command to reference free tracker 302 to track the object (step C). Detector 308 processes the data and poses back to the reference image for the object (s), features such as SIFT features, lines with descriptors, metadata (such as for extension), and AR applications. A list of objects (s), such as (step D), can be returned to system manager 320. A reference image for the object is added to the local cache 310 (step E), and the local cache 310 acknowledges the addition of the object (step F). The reference free tracker 302 provides the detector 312 with a change in pose between the initial image and the current image (step G). The detector 312 uses the reference image to find the object in the current captured image and provides the object ID to system manager 320 (step H). In addition, the pose provided by the server-based detector 308 is from the reference free tracker 302. Used by detector 312 with changes in pose to generate the current pose, which is also provided to system manager 320 (step H). System manager 320 instructs the reference free tracker 302 to stop object tracking (step I) and the reference base tracker 314 to start object tracking (step J). Tracking continues with reference-based tracker 314 until tracking is lost (step K).
0035FIG. 5 shows another call flow diagram for server-assisted AR where pauses are not provided by server 130. The call flow is similar to the call flow shown in FIG. 4, except that the detector 308 does not provide pause information to the system manager 320 in step D. Thus, the detector 312 determines the pose based on the current image and the reference image provided by the detector 308 and provides the pose to the system manager 320 (step G).
0036As mentioned above, the scene change detector 304 controls the frequency of detection requests sent to the server 130 based on the changes in the current captured image relative to the previous captured image. The scene change detector 304 is used when it is desirable to communicate with the external server 130 to initiate object detection only when important new information is present in the image.
0037FIG. 6 shows a flow chart of the method performed by the scene change detector 304. The process for scene change detection is based on a combination of metrics from the reference free tracker 302 (Figure 3) and an image pixel histogram. As mentioned earlier, the reference free tracker 302 tracks relative movement between successive images as an optical flow, an approach such as normalized cross-correlation, and / or, for example, a point, line, or region correspondence. Use any approach that you like. Histogram-based methods work well for certain use cases, such as book flipping, where there are significant changes in the information content of the scene in a short period of time, and are therefore suitable for use in the scene detection process. Informative, the criteria-free tracking process can efficiently detect changes in other use cases, such as panning, where there is a gradual change in the information content in the scene.
0038Thus, the input image 402 is provided, as shown in FIG. The input image is the current captured image, which can be the current video frame or photo. If the last image did not trigger a scene change detection (404), the scene change detector initialization (406) is performed (406). Initialization divides the image into blocks, for example, 8x8 blocks in the case of QVGA images (408), and FAST (Features from Accelerated Segment Test), which holds, for example, M strongest corners. ) Using a corner detector to extract key points from each block (410), where M can be 2. Of course, other methods can be used to extract key points, such as Harris corners, SIFT (Scale Invariant Feature Transform) feature points, SURF (Speeded-up Robust Features), or any other desired method. Can be used for No trigger signal is returned (412).
0039If the last image triggered scene change detection (404), metrics are obtained from the reference free tracker 302 (Figure 3) shown as optical flow process 420 and from the image pixel histogram shown as histogram process 430. .. If desired, the reference free tracker 302 can generate metrics using a process other than optical flow, such as normalized cross-correlation. Optical flow process 420 tracks corners from previous images (422), using, for example, normalized cross-correlation, and identifies their location in the current image. The corners are divided into blocks, for example, each using a FAST corner detector that holds M strongest corners based on the FAST corner threshold, as described in Initialization 406 above. By selecting a key point from the block, it may have been previously extracted, or in the case of Harris corners, it retains the M strongest corners based on the Hessian threshold. Criteria free tracking is performed on selected corners across successive images to determine the location of the corners in the current image and the corners lost in tracking. In the current iteration, ie, the total intensity of the corners lost between the current image and the preceding image (d at 424) is calculated as the first change metric and since the previous trigger, ie, the current The total intensity of the corners lost between the image and the previous trigger image (D at 426) is calculated as a second change metric, which is provided in the video statistic calculation 440. Histogram process 430 divides the current input image (called C) into B × B blocks, and for each block the color histogram H<sup>C</sup><sub>i, j</sub>(432), where i and j are block indexes in the image. The block-wise comparison of the histogram is the Nth past image H, for example using the Chi-Square method.<sup>N</sup><sub>i, j</sub>Performed with a histogram of the corresponding blocks from (434). Histogram comparisons help determine the similarity between the current image and the Nth past image to identify whether the scene has changed significantly. Using one example, B can be selected to be 10. To compare the histogram of the current image with the Nth past image using the chi-square method, the following calculation is performed: <maths num="1"><img id="000002" he="21" wi="158" file="JP6290331B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>
0040Block-by-block comparison is an array of different values f<sub>ij</sub>To generate. Array f<sub>ij</sub>Is sorted and the histogram change metric h is, for example, the sorted array f<sub>ij</sub>Determined as the average of half the elements in the center of (436). The histogram change metric h is also provided for video statistics calculation.
0041As mentioned above, if desired, the tracking mask provided by the reference base tracker 314 (Figure 3) is used during scene change detection to reduce the area of the input image that should be monitored for scene changes. be able to. The tracking mask identifies the area where the object is identified and therefore scene change monitoring can be omitted. Thus, for example, when the input image is divided into a plurality of blocks in, for example, 422,432, the tracking mask can be used to identify the blocks that are within the area with the identified objects. As a result, those blocks can be ignored.
0042The video statistic calculation 440 receives the optical flow metrics d, D, histogram change metric h and generates an image quality decision provided with the metrics d, D, h to see if the detection should be triggered. To determine. The change metric Δ is calculated, compared to the threshold (458), and returns the trigger signal (460). Of course, if the change metric Δ is less than the threshold, no trigger signal is returned. The change metric Δ can be calculated based on the optical flow metric d, D, and the histogram change metric h, for example: (456): <maths num="2"><img id="000003" he="14" wi="158" file="JP6290331B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>
0043Here, α, β, and γ are weights that are appropriately selected to provide relative importance to the three statistics d, D, and h (452). In one embodiment, the values of α, β, and γ can be set constant throughout the execution time. In an alternative embodiment, the values of α, β, and γ can be adapted according to the received, possible feedback on the performance of the system, or according to the targeted use case. For example, in an application involving panning-type scene change detection, the values of α and β can be set relatively high compared to γ, as the statistics d and D may have higher reliability in this case. Alternatively, in applications that primarily include book flipping type use cases where the histogram statistic h can be more informative, the values of α and β can be set relatively low compared to γ. The threshold can be adapted, if desired, based on the output of the video statistic calculation 440 (454).
0044In one case, if desired, the scene detection process can be based on the metric from the reference free tracker 302, without the metric from the histogram. For example, the change metric Δ obtained from Equation 2 can be used with γ = 0. In another implementation, the input image is from multiple blocks and, for example, from each block using a FAST (Features from Accelerated Segment Test) corner detector that holds the M strongest corners, as described above. It can be divided into multiple extracted key points. If a sufficient number of blocks change between the current image and the previous image, for example compared to the threshold, the scene is determined to have changed and a trigger signal is returned. A block can be considered changed, for example, if the number of tracked corners is less than another threshold.
0045In addition, if desired, the scene detection process can simply be based on the total intensity of the corners lost since the previous trigger (426 D) compared to the intensity of the total number of corners in the image. For example, the change metric Δ obtained from Equation 2 can be used with α = 0 and γ = 0. The total strength of the corners lost since the previous trigger can be determined as follows: <maths num="3"><img id="000004" he="22" wi="158" file="JP6290331B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>
0046In Equation 3, s<sub>j</sub>Is the intensity of the corner j, t is the last triggered image number, c is the current image number, and Li is the set containing the lost corner identifier in frame i. If desired, different change metrics Δ can be used: <maths num="4"><img id="000005" he="32" wi="158" file="JP6290331B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>
0047Where N<sub>T</sub>Is the total number of corners in the triggered image. The change metric Δ can be compared to the threshold (458).
0048Additionally, as mentioned above, a tracking mask can be used by the scene change detector 304 to limit the area of each image searched for for scene changes. In other words, the loss of strength at the corners outside the area of the trigger mask is a related metric. The reduction in the size of the area explored by the scene change detector 304 results in a corresponding reduction in the number of corners that can be expected to be detected. Thus, additional parameters may be used to compensate for corner loss due to the tracking mask, for example: <maths num="5"><img id="000006" he="18" wi="158" file="JP6290331B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>
0049The compensation parameter λ can be used to adjust the change metric Δ. For example, if the scene detection process is simply based on the total intensity (D) of corners lost in the unmasked area since the previous trigger, the change metric Δ obtained from Equation 4 is modified as follows: Can be: <maths num="6"><img id="000007" he="32" wi="158" file="JP6290331B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>
0050Where D<sub>c</sub>Is given by Equation 3 (where Li is defined as a set containing the identifiers of the lost corners in the unmasked area at frame i), A<sub>c</sub>Is the mask area for image c and A is A<sub>t + 1</sub>Is initially set to.
0051Figure 7 is a chart showing the performance of the system for a typical book flipping use case where 5 pages are turned in 50 seconds. Figure 7 shows the number of network transmissions required to request object detection as a function of the minimum trigger gap in seconds. The fewer network transmissions required for the same minimum trigger gap, the better the performance. Curve 480 for periodic triggers and curve 482 for optical flow-based scene change detectors (SCDs) with no histogram statistics (γ = 0) and no reference-based tracker 314 (Figure 3). No histogram statistics (γ = 0), but with curve 484 for optical flow based scene change detector (SCD) using reference base tracker 314, reference base tracker 314 and timing manager 305 (Figure 3) Several curves are shown, including curve 486 for a scene change detector (SCD) (shown in Figure 6) based on a combination of optical flow and histogram. As can be seen from Figure 7, the combination system outperforms other systems in flipping use cases.
0052Figure 8 shows an approach to face recognition using a server-assisted AR process. As shown in FIG. 8, the mobile platform 110 includes acquiring facial images as well as any other useful sensor information such as SPS or position / motion sensor data. To execute. Mobile platform 110 performs face detection 504, such as face data (which can be a face image) for one or more faces, as well as SPS or position / motion sensor data, as indicated by arrow 506. Providing any other useful data to server 130. Mobile platform 110 tracks 2D movement of the face (508). The server 130 executes face recognition 510 based on the provided face data, for example, using the data read from the global database 512 and stored in the local cache 514. Server 130 provides data about the face, such as identity or other desired information, to the mobile platform 110, which uses the received data to annotate the face displayed on the display 114 by name, etc. Or provide rendered extended data (516).
0053Figure 9 shows another approach to face recognition using a server-assisted AR process. Figure 9 is similar to the approach shown in Figure 8, and the similarly specified elements are identical. However, as shown in FIG. 9, the image is provided to server 130 (508') and face detection (504') is performed by server 130.
0054Figure 10 shows an approach to visual search using a server-assisted AR process. As shown in FIG. 10, the mobile platform 110 includes acquiring an image of a desired object, as well as any other useful sensor information such as SPS or position / motion sensor data. Perform acquisition (520). The mobile platform 110 performs feature detection (522) and provides the server 130 with the detected features as well as any other useful data such as SPS or position / motion sensor data, as indicated by arrow 526. To do. Mobile platform 110 tracks feature 2D movements (524). The server 130 performs object recognition 528 based on the provided features, for example, using the data read from the global database 530 and stored in the local cache 532. Server 130 may also perform global registration (534) to obtain, for example, reference images, poses, and so on. Server 130 provides data about objects such as reference images and poses to mobile platform 110, which uses the received data to perform local registration (536). The mobile platform 110 can then render the desired extended data for the objects displayed on the display 114 (538).
0055Figure 11 shows another approach to visual search using a server-assisted AR process. FIG. 11 is similar to the approach shown in FIG. 10, and the similarly specified elements are identical. However, as shown in FIG. 11, the entire image is provided to server 130 (526') and feature recognition (522') is performed by server 130.
0056Figure 12 shows an approach to criteria-based tracking using server-assisted processes. As shown in FIG. 12, the mobile platform 110 includes acquiring an image of a desired object, as well as any other useful sensor information such as SPS or position / motion sensor data. Perform acquisition (540). In some embodiments, the mobile platform 110 may generate side information, such as text recognition or barcode reading (541). The mobile platform 110 performs feature detection (542) and, as indicated by arrow 546, is generated along with the detected features, as well as any other useful data such as SPS or position / motion sensor data. If so, the side information is provided to the server 130. The mobile platform 110 uses, for example, point, line, or region tracking, or dense optical flow to track 2D movement of features (544). In some embodiments, the server 130 uses the provided features to multiple planes. recognition) (548) can be executed. Once a plane is identified, object recognition (550) can be performed on individual planes or groups of planes, for example, using the data read from global database 552 and stored in the local cache 554. .. If desired, any other recognition method is used. In some embodiments, server 130 also performs pose estimation (555), if desired, using a homography matrix, an affine matrix, a rotation matrix, and a transformation matrix, with 6 degrees of freedom (six-). degrees of degrees of It can be provided by freedom). Server 130 provides data about the object, such as a reference image, to mobile platform 110, which can be a local homography registration or a local essential matrix registration using the received data. Perform ration (556). As mentioned above, the mobile platform 110 can include a local cache 557 to store the received data, which is a subsequent rediscovery that can be performed on the mobile platform 110 if tracking is lost. Can be beneficial. The mobile platform 110 can then render the desired extended data for the objects displayed on the display 114 (558).
0057Figure 13 shows another approach to criteria-based tracking using server-assisted processes. FIG. 13 is similar to the approach shown in FIG. 12, with similarly designated elements being identical. However, as shown in FIG. 13, the entire image is provided to server 130 (546') and feature recognition (542') is performed by server 130.
0058Figure 14 shows an approach to 3D model creation using server-assisted processes. As shown in FIG. 14, the mobile platform 110 includes acquiring an image of a desired object as well as any other useful sensor information such as SPS or position / motion sensor data. Perform acquisition (560). Mobile platform 110 performs 2D image processing (562) and tracks motion using criteria-free tracking, such as an optical flow or normalized cross-correlation-based approach (564). The mobile platform 110 performs a local 6-DOF registration (568) to obtain a rough estimate of the pose. In certain embodiments, this data may be provided to server 130 along with the images. Server 130 then bundles to refine the registration. Adjustment) can be performed (570). Given the set of images and the 3D point correspondence from different perspectives, the bundle adjustment algorithm helps estimate the 3D coordinates of the points in known reference coordinate systems and identifies the relative movement of the camera between different perspectives. Help. Bundle adjustment algorithms are generally computationally intensive operations that can be performed efficiently on the server side by passing side information from the mobile platform 110 and additional information from the local cache 572 if available. Can be done. After the location of the 3D points and the relative poses have been estimated, they can be provided directly to the mobile platform 110. Alternatively, a 3D model of the object can be built on the server based on the data and such data can be sent to the mobile platform 110. The mobile platform 110 can then use the information obtained from the server 130 to render the desired extended data on the objects displayed on the display 114 (576).
0059Note that the entire system configuration can be adapted depending on the capabilities of the mobile platform 110, server 130, and communication interfaces such as network 120. If the mobile platform 110 is a low-end device that does not have a dedicated processor, most of the operations can be offloaded to server 130. On the other hand, if the mobile platform 110 is a high-end device with good computing power, the mobile platform 110 may choose to perform some of the tasks and offload fewer tasks to the server 130. In addition, the system may be adapted to handle different types of communication interfaces, for example, depending on the bandwidth available on the interface.
0060In one implementation, the server 130 can provide feedback to the mobile platform 110 regarding the task and what parts of the task can be offloaded to the server 130. Such feedback may be based on the capabilities of the server 130, the type of operation to be performed, the available bandwidth within the communication channel, the power level of the mobile platform 110 and / or the server 130, and the like. For example, server 130 can recommend that mobile platform 110 send a lower quality version of the image if the network connection is poor and the data rate is low. Server 130 may also suggest that the mobile platform perform more processing on the data and send the processed data to server 130 if the data rate is low. For example, the mobile platform 110 can calculate features for object detection and send features instead of sending the entire image if the communication link has a low data rate. If server 130 has a good network connection, or if past attempts to recognize objects in the image have failed, the mobile platform 110 will instead send a higher quality version of the image, or , It may be recommended to send the image more frequently, thereby reducing the minimum frame gap η.
0061In addition, the mobile-server architecture described herein can be extended to scenarios where more than one mobile platform 110 is used. For example, two mobile platforms 110 are looking at the same 3D object from different angles, and server 130 performs a joint bundle adjustment from the data obtained from both mobile platforms 110 to make a good 3D object. You can create model objects. Such applications can be useful for applications such as multiplayer gaming.
0062FIG. 15 is a block diagram of the mobile platform 110 capable of distributed processing using server-based detection. The mobile platform 110 includes a camera 112 and a user interface 150 including a display 114 capable of displaying images captured by the camera 112. The user interface 150 may also include a keypad 152, or other input device through which the user can enter information into the mobile platform 110. If desired, the keypad 152 is removed by integrating the virtual keypad with a display 114 equipped with a touch sensor. The user interface 150 may also include a microphone 154 and a speaker 156, for example if the mobile platform is a cellular phone.
0063As mentioned above, the mobile platform 110 can include a wireless transceiver 162 that can be used to communicate with the external server 130 (FIG. 3). The mobile platform 110 can optionally receive positioning signals from motion sensors 164, including, for example, accelerometers, gyroscopes, electronic compasses, or other similar motion sensing elements, and SPS systems. It may include additional features that may be useful in AR applications, such as the Positioning System (SPS) receiver 166. Of course, mobile platform 110 may include other elements unrelated to this disclosure.
0064The mobile platform 110 is also connected to and communicates with the camera 112 and wireless transceiver 162, along with other features such as user interface 150, motion sensor 164, and SPS receiver 166 when used. Can include unit 170. As described above, the control unit 170 receives and processes data from the camera 112, and in response, controls communication with the external server through the wireless transceiver 162. The control unit 170 may be provided by a processor 171 and associated memory 172 which may include software 173 executed by the processor 171 to perform the methods or parts of the methods described herein. Control unit 170 may additionally or optionally include hardware 174 and / or firmware 175.
0065The control unit 170 includes a scene change detector 304 that triggers communication with an external server as described above. Additional components such as the trigger time manager 305 and image quality estimator 306 shown in Figure 3 may also be included. The control unit 170 is further used to discover objects in the current image based on reference free tracker 302, reference base tracker 314, and objects stored in the local cache, for example in memory 172. Includes unit 312 and. The control unit 170 further includes an augmented reality (AR) unit 178 for generating AR information and displaying it on the display 114. The scene change detector 304, reference free tracker 302, reference base tracker 314, detection unit 312, and AR unit 178 are shown separately and away from processor 171 for clarity, but they are simply It can be a unit and / or can be implemented in processor 171 based on instructions in software 173 that are read by processor 171 and executed in processor 171. As used herein, one or more of the processor 171 and the scene change detector 304, reference free tracker 302, reference base tracker 314, detection unit 312, and AR unit 178 may be one or more. It will be appreciated that microprocessors, embedded processors, controllers, application specific integrated circuits (ASICs), digital signal processors (DSPs), etc. can be included, but not necessarily. The term processor is intended to describe the functionality provided by a system rather than specific hardware. In addition, as used herein, the term "memory" refers to any type of computer storage medium, including long-term memory, short-term memory, or other memory associated with mobile platforms. Any particular type of memory or
0066The methods described herein can be implemented by various means depending on the application. For example, these methods may be implemented in hardware 174, firmware 175, software 173, or a combination thereof. For hardware implementation, the processing unit is one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays ( It may be implemented within an FPGA), processor, controller, microcontroller, microprocessor, electronic device, other electronic unit designed to perform the functions described herein, or a combination thereof. Thus, the device for acquiring sensor data is based on the camera 112, the SPS receiver 166, the motion sensor 164, and the image generated by the camera 112 or other means for acquiring sensor data. It can be equipped with a processor capable of generating side information such as text recognition or barcode reading. The device for determining if there is a trigger event with a change in sensor data compared to previously acquired sensor data is by processor 171 executing instructions embedded in software 173 or by hardware 174 or It comprises a detection unit 312 that can be implemented in firmware 175, or other means for determining if there is a trigger event with a change in the sensor data compared to previously acquired sensor data. A device that sends sensor data to a server in the presence of a trigger event comprises a wireless transceiver 162, or other means for sending sensor data to a server in the presence of a trigger event. Devices that receive information related to sensor data from the server include a wireless transceiver 162, or other means for receiving information related to sensor data from the server. To. The device for obtaining a mobile platform pose for an object comprises a reference free tracker 302, a radio transceiver 162, or other means for obtaining a mobile platform pose for an object. A device that tracks an object using the object's reference image and pose comprises a reference base tracker 314 or other means of tracking the object using the object's reference image and pose. The device that determines if there is a scene change in the captured image compared to the previous captured image is a scene change that can be performed by processor 171 executing instructions embedded in software 173 or in hardware 174 or firmware 175. The detector 304, or other means of determining whether or not there is a scene change in the captured image as compared to the previous captured image.
0067In firmware and / or software implementations, the method may be implemented in modules (eg, procedures, functions, etc.) that perform the functions described herein. Any machine-readable medium that tangibly incorporates the instructions can be used in carrying out the methods described herein. For example, software 173 may contain program code stored in memory 172 and executed by processor 171. Memory can be implemented within processor 171 or outside processor 171.
0068For firmware and / or software implementations, the function may be stored as one or more instructions or codes on a computer-readable medium. Examples include non-transitory computer-readable media encoded with data structures and computer-readable media encoded with computer programs. Computer-readable media include physical computer storage media. The storage medium can be any available medium that can be accessed by a computer. As a non-limiting example, such computer-readable media include RAM, ROM, flash memory, EEPROM®, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic disk storage. It may include a device, or other medium that can be accessed by a computer and used to store the desired program code in the form of instructions or data structures. Disks and discs, as used herein, are compact discs (CDs), laserdiscs (registered trademarks), optical discs, digital versatile discs (DVDs), floppy (registered trademarks) discs, Includes Blu-ray® discs. Here, the disk usually reproduces data by magnetic action, and the disk (disc) optically reproduces data by a laser. The above combinations should also be included in the range of computer readable media.
0069Although the present invention has been described with respect to specific embodiments for teaching purposes, the invention is not limited thereto. Various indications and modifications can be made without departing from the scope of the invention. Therefore, the purpose and scope of the attached claims should not be limited to the above description.<u style="single">The invention described in the initial claims of the present application will be added below.</u><u style="single">[C1]</u><u style="single">The way</u><u style="single">Acquiring sensor data using a mobile platform</u><u style="single">Determining if there is a trigger event with a change in the sensor data compared to previously acquired sensor data.</u><u style="single">When the trigger event is present, sending the sensor data to the server and</u><u style="single">Receiving information related to the sensor data from the server</u><u style="single">How to prepare.</u><u style="single">[C2]</u><u style="single">The method according to C1 above, wherein the sensor data comprises a captured image of an object.</u><u style="single">[C3]</u><u style="single">Further comprising determining the quality of the captured image before transmitting the sensor data to the server, the sensor data is transmitted to the server only if the quality of the captured image is better than the threshold. , The method described in C2 above.</u><u style="single">[C4]</u><u style="single">Determining the quality of the captured image is obtained from the image together with analyzing the sharpness of the captured image, analyzing the number of detected corners in the captured image, and a learning classifier. The method according to C3 above, comprising using the statistical values and at least one of them.</u><u style="single">[C5]</u><u style="single">The method according to C2, further comprising rendering an extension for the object based on the information associated with the sensor data received from the server.</u><u style="single">[C6]</u><u style="single">The method according to C2 above, wherein the information related to the sensor data comprises identifying the object.</u><u style="single">[C7]</u><u style="single">The method according to C2, wherein the captured image comprises a plurality of objects, and the information related to the sensor data comprises identifying the plurality of objects.</u><u style="single">[C8]</u><u style="single">Acquiring a pose for each of the plurality of objects for the mobile platform,</u><u style="single">Using the pose and the information related to the sensor data to track each of the plurality of objects.</u><u style="single">The method according to C7 above, further comprising.</u><u style="single">[C9]</u><u style="single">Acquiring the pose of the mobile platform with respect to the object,</u><u style="single">Tracking the object using the pose and the information related to the sensor data.</u><u style="single">The method according to C2 above, further comprising.</u><u style="single">[C10]</u><u style="single">The information related to the sensor data includes a reference image of the object, and acquiring the pose comprises receiving a first pose from the server based on the captured image and the reference image. The method described in C9 above.</u><u style="single">[C11]</u><u style="single">The method according to C10, further comprising performing reference free tracking of the object until the first pose is received from the server.</u><u style="single">[C12]</u><u style="single">Acquiring a second captured image of the object when the first pose is received from the server,</u><u style="single">Tracking the object between the captured image and the second captured image to determine the gradual change.</u><u style="single">Using the gradual change and the first pose to obtain the pose of the mobile platform with respect to the object.</u><u style="single">The method according to C10 above, further comprising.</u><u style="single">[C13]</u><u style="single">Acquiring a second captured image of the object,</u><u style="single">Using the reference image to detect the object in the second captured image,</u><u style="single">Using the object detected in the reference image and the second captured image to obtain the pose of the mobile platform with respect to the object.</u><u style="single">Using the pose to initialize the reference-based tracking of the object</u><u style="single">The method according to C10 above, further comprising.</u><u style="single">[C14]</u><u style="single">The information related to the sensor data includes a two-dimensional (2D) model of the object, a three-dimensional (3D) model of the object, extended information, saliency information about the object, and information related to object matching. The method according to C2 above, comprising at least one of them.</u><u style="single">[C15]</u><u style="single">The method according to C2, wherein determining whether or not the trigger event is present comprises determining whether or not there is a scene change in the captured image as compared to a previous captured image.</u><u style="single">[C16]</u><u style="single">Determining whether or not the scene change exists is</u><u style="single">Determining the first change metric using the captured image and the previous captured image,</u><u style="single">Determining the second change metric using the second previous captured image and the captured image from the previous trigger event,</u><u style="single">To generate a histogram change metric for the captured image,</u><u style="single">Using the first change metric, the second change metric, and the histogram change metric to determine the scene change.</u><u style="single">The method according to C15 above.</u><u style="single">[C17]</u><u style="single">The information associated with the sensor data comprises object identification, and the method further comprises.</u><u style="single">Acquiring additional captured images of the object</u><u style="single">Using the object identification to identify the object in the additional captured image,</u><u style="single">Generating a tracking mask for the additional captured image based on the object identification, where the tracking mask indicates an area within the additional captured image from which the object is identified.</u><u style="single">Using the tracking mask with the additional captured image of the object to identify the remaining area of the additional captured image, and</u><u style="single">To detect a trigger event with a scene change in the remaining area of the additional captured image.</u><u style="single">The method according to C2 above.</u><u style="single">[C18]</u><u style="single">The method according to C1 above, wherein the sensor data includes one or more of image data, motion sensor data, position data, barcode recognition, text detection results, and contextual information.</u><u style="single">[C19]</u><u style="single">The method according to C18 above, wherein the context information includes one or more of user behavior, user preference, location, information about the user, time of day, and lighting quality.</u><u style="single">[C20]</u><u style="single">The method according to C1 above, wherein the sensor data comprises an image of a face, and the information received from the server comprises an identity associated with the face.</u><u style="single">[C21]</u><u style="single">The sensor data includes a plurality of images of an object captured by cameras at different positions and a rough estimate of the pose of the camera with respect to the object, and the information received from the server is the pose. The method according to C1 above, comprising refinement of the above and at least one of the three-dimensional models of the object.</u><u style="single">[C22]</u><u style="single">It s a mobile platform</u><u style="single">With sensors adapted to acquire sensor data,</u><u style="single">Wireless transceiver and</u><u style="single">Is there a processor coupled to the sensor and the wireless transceiver that has a trigger event that acquires sensor data through the sensor and has a change in the sensor data compared to previously acquired sensor data? If the trigger event is present, the sensor data is transmitted to the external processor via the wireless transceiver, and information related to the sensor data is received from the external processor via the wireless transceiver. With a processor adapted to</u><u style="single">Mobile platform with.</u><u style="single">[C23]</u><u style="single">The mobile platform according to C22, wherein the sensor is a camera and the sensor data comprises a captured image of an object.</u><u style="single">[C24]</u><u style="single">The processor is further adapted to determine the quality of the captured image before the sensor data is transmitted to the external processor, the sensor data only if the quality of the captured image is better than a threshold. , The mobile platform according to C23 above, transmitted to the external processor.</u><u style="single">[C25]</u><u style="single">The processor performs at least one of the analysis of the sharpness of the captured image, the analysis of the number of corners detected in the captured image, and the processing of the learning classifier with the statistical values obtained from the image. The mobile platform according to C24 above, adapted to determine the quality of the captured image by being adapted to perform.</u><u style="single">[C26]</u><u style="single">The mobile platform according to C23, wherein the processor is further adapted to render an extension to the object based on the information associated with the sensor data received via the radio transceiver.</u><u style="single">[C27]</u><u style="single">The mobile platform according to C23, wherein the information associated with the sensor data comprises identifying the object.</u><u style="single">[C28]</u><u style="single">The mobile platform according to C23, wherein the captured image comprises a plurality of objects, and the information related to the sensor data comprises identifying the plurality of objects.</u><u style="single">[C29]</u><u style="single">The processor is further adapted to take poses for each of the plurality of objects with respect to the mobile platform and track each of the plurality of objects using the pose and the information associated with the sensor data. The mobile platform described in C28 above.</u><u style="single">[C30]</u><u style="single">The mobile platform according to C23, wherein the processor further acquires a pose of the mobile platform with respect to the object and is adapted to track the object using the pose and the information associated with the sensor data. ..</u><u style="single">[C31]</u><u style="single">The information associated with the sensor data comprises a reference image of the object, the processor being adapted to receive a first pose from the external processor based on the captured image and the reference image. Mobile platform described in C30.</u><u style="single">[C32]</u><u style="single">The mobile platform according to C31, wherein the processor is further adapted to perform reference free tracking of the object until the first pause is received from the server.</u><u style="single">[C33]</u><u style="single">The processor further acquires a second captured image of the object when the first pose is received from the external processor, and the captured image and the second to determine a gradual change. C31 adapted to track the object to and from the captured image and use the gradual change and the first pose to obtain the pose of the mobile platform with respect to the object. Mobile platform described in.</u><u style="single">[C34]</u><u style="single">The processor further acquires a second captured image of the object, detects the object in the second captured image using the reference image, and acquires the pose of the mobile platform with respect to the object. C31 above, wherein the second captured image and the object detected in the reference image are used in order to use the pose to initialize reference-based tracking of the object. Mobile platform.</u><u style="single">[C35]</u><u style="single">The processor further comprises at least one of a two-dimensional (2D) model of the object, a three-dimensional (3D) model of the object, extended information, saliency information about the object, and information related to object matching. The mobile platform according to C23 above, adapted to receive from the external processor via the radio transmitter.</u><u style="single">[C36]</u><u style="single">The processor is adapted to determine if the trigger event is present by being adapted to determine if there is a scene change in the captured image relative to the previous captured image. The mobile platform described in C23 above.</u><u style="single">[C37]</u><u style="single">The processor uses the captured image and the previous captured image to determine a first change metric, and a second change metric using the second previous captured image and the captured image from a previous trigger event. To generate the histogram change metric for the captured image and to use the first change metric, the second change metric, and the histogram change metric to determine the scene change. The mobile platform according to C36 above, adapted to determine if the scene change is present or not.</u><u style="single">[C38]</u><u style="single">The information associated with the sensor data comprises object identification, and the processor further acquires an additional captured image of the object and uses the object identification to capture the object in the additional captured image. Identify and generate a tracking mask for the additional captured image based on the object identification, where the tracking mask indicates an area within the additional captured image from which the object is identified. The tracking mask is used with the additional captured image of the object to identify the remaining area of the additional captured image to detect a trigger event with a scene change in the remaining area of the additional captured image. The mobile platform described in C23 above, adapted to.</u><u style="single">[C39]</u><u style="single">The mobile platform according to C22, wherein the sensor data includes one or more of image data, motion sensor data, position data, barcode recognition, text detection results, and contextual information.</u><u style="single">[C40]</u><u style="single">The mobile platform according to C39 above, wherein the context information includes one or more of user behavior, user preference, location, information about the user, time of day, and lighting quality.</u><u style="single">[C41]</u><u style="single">The mobile platform according to C22, wherein the sensor comprises a camera, the sensor data comprises an image of a face, and the information received via the wireless transceiver has an identity associated with the face.</u><u style="single">[C42]</u><u style="single">The sensor comprises a camera, and the sensor data is received from a server with a plurality of images of an object captured by the cameras at different positions and a rough estimate of the camera's pose with respect to the object. The mobile platform according to C22, wherein the information comprises at least one of a refinement of the pose and a three-dimensional model of the object.</u><u style="single">[C43]</u><u style="single">It s a mobile platform</u><u style="single">Means to acquire sensor data and</u><u style="single">A means of determining whether or not there is a trigger event with a change in the sensor data compared to previously acquired sensor data.</u><u style="single">When the trigger event is present, the means for transmitting the sensor data to the server and</u><u style="single">As a means for receiving information related to the sensor data from the server</u><u style="single">Mobile platform with.</u><u style="single">[C44]</u><u style="single">The means for acquiring sensor data is a camera, the sensor data is a captured image of an object, the information associated with the sensor data comprises a reference image of the object, and the mobile platform further comprises.</u><u style="single">A means of obtaining the pose of the mobile platform with respect to the object, and</u><u style="single">The mobile platform according to C43, comprising means for tracking the object using the pose and a reference image of the object.</u><u style="single">[C45]</u><u style="single">The means for acquiring sensor data is a camera, the sensor data is a captured image of an object, and the means for determining whether or not the trigger event is present is the captured image as compared to a previous captured image. The mobile platform according to C43 above, which comprises means for determining whether or not there is a scene change in.</u><u style="single">[C46]</u><u style="single">A non-temporary computer-readable medium containing program code</u><u style="single">Program code for acquiring sensor data and</u><u style="single">Program code for determining whether or not there is a trigger event with a change in the sensor data compared to previously acquired sensor data.</u><u style="single">When the trigger event is present, the program code for transmitting the sensor data to the external processor and</u><u style="single">A non-transitory computer-readable medium comprising program code for receiving information related to the sensor data from the external processor.</u><u style="single">[C47]</u><u style="single">The sensor data is a captured image of the object, the information related to the sensor data includes a reference image of the object, and the non-temporary computer-readable medium is:</u><u style="single">The program code for acquiring the pose for the object and</u><u style="single">The non-transitory computer-readable medium according to C46, further comprising a program code for tracking the object using the pose and the reference image of the object.</u><u style="single">[C48]</u><u style="single">The sensor data is a captured image of an object, and the program code for determining whether or not the trigger event exists is whether or not there is a scene change in the captured image as compared with the previous captured image. The non-temporary computer-readable medium according to C46 above, comprising the program code for determining.</u>
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2006293912A | Cites | Japan |
| JP2005182350A | Cites | Japan |
| JP2004145448A | Cites | Japan |
| Stephan Gammeter ET AL.,"Server-side object recognition and client-side object tracking for mobile augmented reality",Computer Vision and Pattern Recognition Workshops(CVPRW), 2010 IEEE Computer Society Conference on,2010年 6月13日,p.1-8 | Non-patent | – |
19 members in 8 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 61384667 | United States of America | – | |
| 38466710 | United States of America | P | |
| 13235847 | United States of America | – | |
| 201113235847 | United States of America | A |
Members19
| Document | Office | Kind | |
|---|---|---|---|
| WO2012040099A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2012243732A1 | United States of America | A1 | |
| CN103119627A | China | A | |
| KR20130060339A | Republic of Korea | A | |
| EP2619728A1 | European Patent Office (EPO) | A1 | |
| JP2013541096A | Japan | A | |
| JP2015144474A | Japan | A | |
| KR101548834B1 | Republic of Korea | B1 | |
| JP5989832B2 | Japan | B2 | |
| US2016284099A1 | United States of America | A1 | |
| JP6000954B2 | Japan | B2 | |
| US9495760B2 | United States of America | B2 | |
| JP2017011718A | Japan | A | |
| CN103119627B | China | B | |
| US9633447B2 | United States of America | B2 | |
| JP6290331B2This record | Japan | B2 | |
| EP2619728B1 | European Patent Office (EPO) | B1 | |
| ES2745739T3 | Spain | T3 | |
| HUE047021T2 | Hungary | T2 |
13 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 6290331
- Application
- 157488
Titles2
- Japanese
- クラウド支援型拡張現実のための適応可能なフレームワーク
- English
- Adaptable framework for cloud-assisted augmented reality
Classification
- CPC, 7
- G06T7/246
- G06T19/006
- G06T7/73
- G06V20/20
- G06T7/292
- G06T7/269
- G06T2207/10004
- IPC, 3
- H04N5 232
- G06T19 00
- H04N5 222
