Device and method with sensor-specific image recognition
23 claims: 8 independent, 15 dependent
- 1イメージセンサによって受信された入力イメージから特徴抽出レイヤを用いて特徴データを抽出するステップと、前記抽出された特徴データに固定マスク及び可変マスクを適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するステップと、を含み、前記可変マスクは、前記抽出された特徴データに応答して調整され 、 前記認識結果を出力するステップは、 前記抽出された特徴データに前記固定マスクを適用することで、第1認識データを算出するステップと、 前記抽出された特徴データに前記可変マスクを適用することで、第2認識データを算出するステップと、 前記第1認識データ及び前記第2認識データに基づいて前記認識結果を決定するステップと、 を含む、 イメージ認識方法。
- 2前記第1認識データを算出するステップは、前記抽出された特徴データに前記固定マスクを適用することで、オブジェクト関心領域に関する汎用特徴マップを生成するステップと、前記汎用特徴マップから前記第1認識データを算出するステップと、を含む、請求項 1 に記載のイメージ認識方法。
- 3前記第2認識データを算出するステップは、前記抽出された特徴データに対応する対象特徴マップに対して前記可変マスクを適用することで、前記イメージセンサの関心領域に関するセンサ特化特徴マップを生成するステップと、前記センサ特化特徴マップから前記第2認識データを算出するステップと、を含む、請求項 1 に記載のイメージ認識方法。
- 4前記センサ特化特徴マップを生成するステップは、前記対象特徴マップの個別値に対して前記可変マスクにおいて対応する値を適用するステップを含む、請求項 3 に記載のイメージ認識方法。
- 5前記抽出された特徴データから完全接続レイヤ及びソフトマックス関数を用いて第3認識データを算出するステップをさらに含み、前記認識結果を決定するステップは、前記第1認識データ及び前記第2認識データと共に、前記第3認識データにさらに基づいて前記認識結果を決定するステップを含む、請求項 1 に記載のイメージ認識方法。
- 6イメージセンサによって受信された入力イメージから特徴抽出レイヤを用いて特徴データを抽出するステップと、 前記抽出された特徴データに固定マスク及び可変マスクを適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するステップと、 を含み、 前記可変マスクは、前記抽出された特徴データに応答して調整され、 前記認識結果を出力するステップは、前記可変マスクを含むセンサ特化レイヤの少なくとも一部のレイヤを用いて、前記特徴データにより前記可変マスクの1つ以上の値を調整するステップを含む、 イ メージ認識方法。
- 7前記可変マスクの1つ以上の値を調整するステップは、前記特徴データに対して畳み込みフィルタリングが適用された結果であるキー特徴マップ及び転置されたクエリ特徴マップ間の積結果から、ソフトマックス関数を用いて前記可変マスクの値を決定するステップを含む、請求項 6 に記載のイメージ認識方法。
- 8前記認識結果を出力するステップは、前記固定 マ スクに基づいた第1認識データ及び前記可変マスクに基づいた第2認識データの加重和を前記認識結果として決定するステップを含む、請求項1に記載のイメージ認識方法。
- 9前記加重和を前記認識結果として決定するステップは、前記第1認識データに適用される加重値よりも大きい加重値を前記第2認識データに適用するステップを含む、請求項 8 に記載のイメージ認識方法。
- 10イメージセンサによって受信された入力イメージから特徴抽出レイヤを用いて特徴データを抽出するステップと、 前記抽出された特徴データに固定マスク及び可変マスクを適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するステップと、 を含み、 前記可変マスクは、前記抽出された特徴データに応答して調整され、 アップデート命令に応答して、外部サーバから前記可変マスクを含むセンサ特化レイヤのパラメータを受信するステップと、受信された前記パラメータをセンサ特化レイヤにアップデートするステップと、をさらに含む、 イ メージ認識方法。
- 11前記外部サーバに対して、前記イメージセンサの光学特性と同一又は類似の光学特性に対応するセンサ特化パラメータを要求するステップをさらに含む、請求項 10 に記載のイメージ認識方法。
- 12前記センサ特化レイヤのパラメータをアップデートする間に、前記固定マスクの値を保持するステップをさらに含む、請求項 10 に記載のイメージ認識方法。
- 13前記認識結果を出力するステップは、前記固定マスク及び複数の可変マスクに基づいて前記認識結果を算出するステップを含む、請求項1に記載のイメージ認識方法。
- 14イメージセンサによって受信された入力イメージから特徴抽出レイヤを用いて特徴データを抽出するステップと、 前記抽出された特徴データに固定マスク及び可変マスクを適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するステップと、 を含み、 前記可変マスクは、前記抽出された特徴データに応答して調整され、 前記認識結果を出力するステップは、前記固定マスク及び複数の可変マスクに基づいて前記認識結果を算出するステップを含み、 前記複数の可変マスクのうち、1つの可変マスクを含むセンサ特化レイヤのパラメータ及び他方の可変マスクを含む他のセンサ特化レイヤのパラメータは互いに異なる、 イ メージ認識方法。
- 15イメージセンサによって受信された入力イメージから特徴抽出レイヤを用いて特徴データを抽出するステップと、 前記抽出された特徴データに固定マスク及び可変マスクを適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するステップと、 を含み、 前記可変マスクは、前記抽出された特徴データに応答して調整され、 前記認識結果を出力するステップは、前記オブジェクトがリアルオブジェクトであるか、又は、偽造オブジェクトであるかを指示する真偽情報を前記認識結果として生成するステップを含む、 イ メージ認識方法。
- 16イメージセンサによって受信された入力イメージから特徴抽出レイヤを用いて特徴データを抽出するステップと、 前記抽出された特徴データに固定マスク及び可変マスクを適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するステップと、 を含み、 前記可変マスクは、前記抽出された特徴データに応答して調整され、 前記認識結果に基づいて権限を付与するステップと、前記権限により電子端末の動作及び前記電子端末のデータのうち少なくとも1つに対するアクセスを許容するステップと、をさらに含む、 イ メージ認識方法。
- 17前記認識結果を出力するステップは、前記認識結果が生成された後、前記認識結果をディスプレイを介して可視化するステップを含む、請求項1 ~請求項16のいずれか一項 に記載のイメージ認識方法。
- 18請求項1~請求項 17 のいずれか一項に記載の方法を実行するための命令語を含む1つ以上のコンピュータプログラムを格納したコンピュータで読み出し可能な記録媒体。
- 19入力イメージを受信するイメージセンサと、前記入力イメージから特徴抽出レイヤを用いて特徴データを抽出し、前記抽出された特徴データに固定マスク及び可変マスクを適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するプロセッサと、を含み、前記可変マスクは、前記抽出された特徴データに応答して調整され、 前記プロセッサは、 前記抽出された特徴データに前記固定マスクを適用することで、前記抽出された特徴データから第1認識データを算出し、 前記抽出された特徴データに前記可変マスクを適用することで、前記抽出された特徴データから第2認識データを算出し、 前記第1認識データ及び前記第2認識データの和に基づいて前記認識結果を決定する、 イメージ認識装置。
- 20前記和は、前記第1認識データに適用される加重値よりも大きい加重値を前記第2認識データに適用することで決定される、請求項 19 に記載のイメージ認識装置。
- 21前記プロセッサは、前記抽出された特徴データに前記固定マスクを適用することで、オブジェクト関心領域に関する汎用特徴マップを生成し、前記汎用特徴マップから前記第1認識データを算出し、前記抽出された特徴データに対応する対象特徴マップに対して前記可変マスクを適用することで、前記イメージセンサの関心領域に関するセンサ特化特徴マップを生成し、前記センサ特化特徴マップから前記第2認識データを算出する、請求項 19 に記載のイメージ認識装置。
- 22受信された入力イメージから特徴抽出レイヤを用いて特徴データを抽出し、可変マスク及び固定されたマスクを前記抽出された特徴データに適用することで、前記入力イメージに示されるオブジェクトに関する認識結果を出力するイメージ認識装置と、認識モデルのセンサ特化レイヤに対する追加トレーニング完了及びアップデート要求のうち少なくとも1つに応答して、前記イメージ認識装置に追加的にトレーニングされたセンサ特化レイヤのパラメータを配布するサーバを含み、前記可変マスクは、前記イメージ認識装置の前記センサ特化レイヤに含まれて前記抽出された特徴データに応答して調整され、前記イメージ認識装置は、配布された前記パラメータに基づいて前記イメージ認識装置の前記センサ特化レイヤをアップデートする、イメージ認識システム。
- 23前記サーバは、前記イメージ認識装置のイメージセンサに類似していると判断されたイメージセンサを含む他のイメージ認識装置に前記追加的にトレーニングされたセンサ特化レイヤの前記パラメータを配布する、請求項 22 に記載のイメージ認識システム。
Independent claims23
125 paragraphs, as filed
Below, a technique for recognizing an image is provided.
In recent years, in order to solve the problem of classifying input patterns into specific groups, there has been much research being conducted on applying efficient pattern recognition methods possessed by humans to actual computers. One such research is on artificial neural networks that model the characteristics of human biological nerve cells using mathematical expressions. To solve the problem of classifying input patterns into specific groups, artificial neural networks use algorithms that mimic the learning ability possessed by humans. Using this algorithm, an artificial neural network can generate a mapping between input patterns and output patterns, and the ability to generate such a mapping is expressed as the learning ability of an artificial neural network. In addition, an artificial neural network has a generalization ability that can generate a relatively correct output for input patterns that have not been used in learning based on the learned results.
<p>An image recognition apparatus according to an embodiment outputs a recognition result for an object by using a mask that is variable according to feature data together with a fixed mask.</p>
<p>An image recognition method according to one embodiment includes the steps of extracting feature data from an input image received by an image sensor using a feature extraction layer, and outputting a recognition result for an object shown in the input image by applying a fixed mask and a variable mask to the extracted feature data, the variable mask being adjusted in response to the extracted feature data.</p><p>The step of outputting the recognition result may include the steps of calculating first recognition data by applying the fixed mask to the extracted feature data, calculating second recognition data by applying the variable mask to the extracted feature data, and determining the recognition result based on the first recognition data and the second recognition data.</p><p>The step of calculating the first recognition data may include a step of generating a generic feature map for an object region of interest by applying the fixed mask to the extracted feature data, and a step of calculating the first recognition data from the generic feature map.</p><p>The step of calculating the second recognition data may include a step of generating a sensor specific feature map for a region of interest of the image sensor by applying the variable mask to a target feature map corresponding to the extracted feature data, and a step of calculating the second recognition data from the sensor specific feature map.</p><p>Generating the sensor specific feature map may include applying corresponding values in the deformable mask to distinct values of the object feature map.</p><p>The image recognition method may further include a step of calculating third recognition data from the extracted feature data using a fully connected layer and a softmax function, and the step of determining the recognition result may include a step of determining the recognition result further based on the third recognition data together with the first recognition data and the second recognition data.</p><p>The step of outputting the recognition result may include a step of adjusting one or more values of the variable mask according to the feature data, using at least a portion of a sensor-specific layer that includes the variable mask.</p><p>The step of adjusting one or more values of the variable mask may include determining the values of the variable mask using a softmax function from a product result between a key feature map, which is a result of applying convolution filtering to the feature data, and a transposed query feature map.</p><p>The step of outputting the recognition result may include the step of determining a weighted sum of first recognition data based on the fixed mask and second recognition data based on the variable mask as the recognition result.</p><p>The step of determining the weighted sum as the recognition result may include applying a weighting to the second recognition data that is greater than a weighting applied to the first recognition data.</p><p>The image recognition method may further include the steps of receiving parameters of a sensor specific layer including the variable mask from an external server in response to an update command, and updating the received parameters to the sensor specific layer.</p><p>The image recognition method may further include a step of requesting sensor specific parameters corresponding to optical characteristics identical or similar to optical characteristics of the image sensor from the external server.</p><p>The image recognition method may further comprise the step of maintaining values of the fixed mask while updating parameters of the sensor specific layer.</p><p>The step of outputting the recognition result may include the step of calculating the recognition result based on the fixed mask and a plurality of variable masks.</p><p>Among the plurality of variable masks, a parameter of a sensor-specific layer including one variable mask and a parameter of another sensor-specific layer including the other variable mask may be different from each other.</p><p>The step of outputting the recognition result may include the step of generating, as the recognition result, authenticity information indicating whether the object is a real object or a counterfeit object.</p><p>The image recognition method may further include the steps of granting authority based on the recognition result, and allowing access to at least one of operation of the electronic terminal and data of the electronic terminal based on the authority.</p><p>The step of outputting the recognition result may include the step of visualizing the recognition result via a display after the recognition result is generated.</p><p>An image recognition device according to one embodiment includes an image sensor that receives an input image, and a processor that extracts feature data from the input image using a feature extraction layer, and outputs a recognition result relating to an object shown in the input image by applying a fixed mask and a variable mask to the extracted feature data, the variable mask being adjusted in response to the extracted feature data.</p><p>The processor can calculate first recognition data from the extracted feature data by applying the fixed mask to the extracted feature data, calculate second recognition data from the extracted feature data by applying the variable mask to the extracted feature data, and determine the recognition result based on the sum of the first recognition data and the second recognition data.</p><p>The sum may be determined by applying a weighting to the second recognition data that is greater than a weighting applied to the first recognition data.</p><p>The processor can apply the fixed mask to the extracted feature data to generate a generic feature map for an object region of interest and calculate the first recognition data from the generic feature map, and apply the deformable mask to a target feature map corresponding to the extracted feature data to generate a sensor-specific feature map for the image sensor region of interest and calculate the second recognition data from the sensor-specific feature map.</p><p>An image recognition system according to one embodiment includes an image recognition device that extracts feature data from a received input image using a feature extraction layer and outputs a recognition result for an object shown in the input image by applying a variable mask and a fixed mask to the extracted feature data, and a server that distributes parameters of an additionally trained sensor-specific layer to the image recognition device in response to at least one of completion of additional training and an update request for a sensor-specific layer of a recognition model, wherein the variable mask is included in the sensor-specific layer of the image recognition device and is adjusted in response to the extracted feature data, and the image recognition device can update the sensor-specific layer of the image recognition device based on the distributed parameters.</p><p>The server may distribute the parameters of the additionally trained sensor-specific layer to other image recognition devices that contain image sensors determined to be similar to an image sensor of the image recognition device.</p>
<p>An image recognition apparatus according to an embodiment can minimize a recognition error rate by generating a recognition result optimized for optical characteristics of a sensor through a deformable mask.</p>
<figref num="1">FIG. 2 is a diagram illustrating a recognition model according to an embodiment.</figref><figref num="2">1 is a flowchart illustrating an image recognition method according to an embodiment.</figref><figref num="3">FIG. 2 illustrates an exemplary structure of a recognition model according to one embodiment.</figref><figref num="4">FIG. 2 illustrates an exemplary structure of a recognition model according to one embodiment.</figref><figref num="5">FIG. 13 is a diagram illustrating an exemplary structure of a recognition model according to another embodiment.</figref><figref num="6">FIG. 13 is a diagram illustrating an exemplary structure of a recognition model according to another embodiment.</figref><figref num="7">FIG. 1 is a diagram illustrating an attention layer according to an embodiment.</figref><figref num="8">FIG. 10 illustrates an exemplary structure of a recognition model according to a further embodiment.</figref><figref num="9">FIG. 2 illustrates training of a recognition model according to an embodiment.</figref><figref num="10">FIG. 13 is a diagram illustrating parameter updates of a sensor specialization layer in a recognition model according to an embodiment.</figref><figref num="11">1 is a block diagram showing a configuration of an image recognition device according to an embodiment;</figref><figref num="12">1 is a block diagram showing a configuration of an image recognition device according to an embodiment;</figref>
The embodiments described below may be modified in various ways. The scope of the patent application is not limited or restricted by such embodiments. The same reference numerals in each drawing indicate the same elements.
Specific structural or functional descriptions disclosed herein are merely exemplary for purposes of describing embodiments, and the embodiments may be embodied in many different forms and the present invention is not limited to the embodiments described herein.
The terms used in this specification are merely used to describe certain embodiments and are not intended to limit the present invention. A singular expression includes a plural expression unless the context clearly indicates otherwise. In this specification, the terms "include" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood as not precluding the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention belongs. Commonly used predefined terms should be interpreted as having a meaning consistent with the meaning they have in the context of the relevant art, and should not be interpreted as having an ideal or overly formal meaning unless expressly defined herein.
In addition, when describing the present invention with reference to the drawings, the same components are denoted by the same reference numerals regardless of the reference numerals, and duplicated descriptions thereof will be omitted. In the description of the embodiments, if a detailed description of related known technology is determined to unnecessarily obscure the gist of the present invention, the detailed description thereof will be omitted.
FIG. 1 is a diagram illustrating a recognition model according to an embodiment.
An image recognition device according to an embodiment may recognize a user using feature data extracted from an input image. For example, the image recognition device extracts feature data from an input image based on at least some layers (e.g., a feature extraction layer) of a recognition model. The feature data is data in which an image is abstracted, and may be represented, for example, in the form of a vector. Feature data having a vector form of two or more dimensions may also be referred to as a feature map. In this specification, a feature map may mainly represent feature data in the form of a two-dimensional vector or a two-dimensional matrix.
The recognition model is a model designed to extract feature data from an image and output a result of recognizing an object shown in the image from the extracted feature data, and may be, for example, a machine learning structure or may include a neural network 100.
The neural network 100 corresponds to an example of a deep neural network (DNN). DNNs include a fully connected network, a deep convolutional network, and a recurrent neural network. The neural network 100 can perform object classification, object recognition, voice recognition, image recognition, and the like by mapping input data and output data having a nonlinear relationship to each other based on deep learning. Deep learning is a machine learning method for solving problems such as image or voice recognition from a big data set, and maps input data and output data to each other through supervised or unsupervised learning.
In this specification, recognition includes data verification and/or data identification. Verification refers to an operation of determining whether input data is true or false. For example, verification refers to a discrimination operation of determining whether an object (e.g., a human face) indicated by an arbitrary input image is identical to an object indicated by a reference image. As a different example, liveness verification refers to a discrimination operation of determining whether an object indicated by an arbitrary input image is a real object or a fake object.
The image recognition device may also verify whether the data extracted from the input image and acquired is identical to registered data previously registered in the device, and determine that the verification of the user corresponding to the input image has been successful in response to the verification that the two pieces of data are identical. In addition, when a plurality of registered data are stored in the image recognition device, the image recognition device may sequentially verify the data extracted from the input image and acquired against each of the plurality of registered data.
Identification refers to a classification operation that determines which of a number of labels the input data indicates, where each label may indicate a class (e.g., the identity of a registered user). For example, an identification operation may indicate whether the user included in the input data is male or female.
1, a neural network 100 includes an input layer 110, a hidden layer 120, and an output layer 130. The input layer 110, the hidden layer 120, and the output layer 130 each include a number of artificial nodes.
For convenience of explanation, the hidden layer 120 is illustrated in FIG. 1 as including three layers, but the hidden layer 120 may include any number of layers. Also, although the neural network 100 is illustrated in FIG. 1 as including a separate input layer for receiving input data, the input data may be directly input to the hidden layer 120. The artificial nodes of the layers of the neural network 100 except the output layer 130 may be connected to the artificial nodes of the next layer via links for transmitting output signals. The number of links corresponds to the number of artificial nodes included in the next layer.
Each artificial node included in the hidden layer 120 receives the output of an activation function related to weighted inputs of an artificial node included in a previous layer. The weighted inputs are inputs of an artificial node included in a previous layer multiplied by weights. The weights may be referred to as parameters of the neural network 100. The activation functions may include sigmoid, hyperbolic tangent (tanh) and ReLU (rectified linear unit), and the activation functions form nonlinearity in the neural network 100. Each artificial node included in the output layer 130 may receive the weighted inputs of an artificial node included in a previous layer.
According to one embodiment, when input data is given, the neural network 100 calculates a function value according to the number of classes to be identified in the output layer 130 via the hidden layer 120, and can identify the input data as the class having the largest value among them. The neural network 100 can identify the input data, but is not limited thereto, and the neural network 100 can also verify the input data against reference data (e.g., enrollment data). The following description of the recognition process will be mainly described in terms of the verification process, but may also be applied to the identification process unless otherwise specified.
If the width and depth of the neural network 100 are large enough, it can have the capacity to implement any function. If the neural network 100 learns a sufficient amount of training data through an appropriate training process, it can achieve optimal recognition performance.
Although the neural network 100 has been described above as an example of the recognition model, the recognition model is not limited to the neural network 100. Next, a verification operation using feature data extracted using the feature extraction layer of the recognition model will be mainly described.
FIG. 2 is a flowchart illustrating an image recognition method according to an embodiment.
First, the image recognition device receives an input image through an image sensor. The input image may be an image of an object, in which at least a part of the object is captured. The part of the object may be a body part related to a unique biometric feature of the object. For example, if the object is a person, the part of the object may be a face, a fingerprint, an iris, and veins of the person. In this specification, the input image is mainly described as including a human face, but is not limited thereto. The input image may be a color image and may include a plurality of channel images for each channel constituting a color space. For example, in an RGB color space, the input image may include a red channel image, a green channel image, and a blue channel image. The color space is not limited thereto, and may be configured as YCbCr, etc. However, the input image is not limited thereto, and may include a depth image, an infrared image, an ultrasound image, a radar scan image, etc.
Then, in step S210, the image recognition device extracts feature data from the input image received by the image sensor using the feature extraction layer. For example, the feature extraction layer may be the hidden layer 120 described with reference to FIG. 1 and include one or more convolution layers. The output of each convolution layer is a result of applying a convolution operation by sweeping a kernel filter to data input to the corresponding convolution layer. If the input image is composed of multiple channel images, the image recognition device can extract feature data for each channel image using the feature extraction layer of the recognition model, and propagate the feature data for each channel to the next layer of the recognition model.
Then, in step S220, the image recognition device outputs a recognition result for an object shown in the input image based on a fixed mask and a variable mask that is adjusted in response to the extracted feature data from the feature data extracted in step S210. The fixed mask may be a mask that has the same value for different input images. The variable mask may be a mask that has different values for different input images.
A mask includes a mask weight for excluding, storing, and modifying a value included in any data. A mask is applied to data including multiple values through an element-wise operation. For example, any value in the data may be multiplied by a mask weight corresponding to the value in the mask. As described below, a mask includes a mask weight for emphasizing and/or storing values corresponding to a region of interest in the data and weakening and/or excluding values corresponding to the remaining regions. For example, the mask weight has a real value between 0 and 1, but the value range of the mask weight is not limited thereto. Data to which a mask is applied may be referred to as masked data.
For reference, the following description will be focused on a case where the size and dimensions of the mask are the same as the size and dimensions of the data to which the mask is applied. For example, if the data to which the mask is applied is a two-dimensional vector having a size of 32×32, the mask may also be a two-dimensional vector having a size of 32×32. However, this is merely an example and is not limiting, and the size and dimensions of the mask may be different from the size and dimensions of the data.
According to an embodiment, the image recognition device applies a mask to the extracted feature data and the target data re-extracted from the feature data to calculate a plurality of masked data. The image recognition device can calculate a recognition result using the plurality of masked data.
3 and 4 are diagrams illustrating an exemplary structure of a recognition model according to an embodiment.
3 shows a schematic structure of an exemplary recognition model 310. According to one embodiment, an image recognition device uses the recognition model 310 to output a recognition result 309 from an input image 301. For example, the image recognition device can use the recognition model 310 to output a recognition result 309 from a single image without an image pair.
The recognition model 310 includes a feature extraction layer 311, a fixed layer 312, and a sensor-specific layer 313. The feature extraction layer 311 may represent a layer designed to extract feature data from the input image 301. The fixed layer 312 represents a layer designed to apply a fixed mask 321 to data (e.g., feature data) propagated from the feature extraction layer 311, and output a first recognition data from the data to which the fixed mask 321 is applied. The sensor-specific layer 313 represents a layer designed to apply a deformable mask 322 to data (e.g., an object feature map extracted from the feature data through one or more convolutional layers) propagated from the feature extraction layer 311, and output a second recognition data from the data to which the deformable mask 322 is applied.
Furthermore, the recognition model 310 may be customized according to the type of image sensor of the electronic terminal to which the recognition model 310 is attached. For example, the parameters of the fixed layer 312 of the recognition model 310 are invariant regardless of the type of image sensor, and the parameters of the sensor specialization layer 313 (e.g., connection weights between artificial nodes, etc.) may vary according to the type of image sensor. The type of image sensor may be classified according to, for example, the optical characteristics of the image sensor. Even if the model numbers of various image sensors are different, if the optical characteristics are the same or similar, the image sensors are classified into the same type.
An image recognition apparatus according to an embodiment extracts feature data from an input image 301 through a feature extraction layer 311. As described above, the feature data may be data in vector form (e.g., feature vector) that is data in which image features are abstracted, but is not limited thereto.
The image recognition device calculates a plurality of recognition data from the same feature data by individually using masks. For example, the image recognition device may calculate first recognition data from extracted feature data based on a fixed mask. The first recognition data indicates a result calculated from data to which the fixed mask is applied, and may be referred to as generic recognition data. As another example, the image recognition device may calculate second recognition data from extracted feature data based on a variable mask 322. The second recognition data indicates a result calculated from data to which the variable mask 322 is applied, and may be referred to as sensor-specific data.
The image recognition device determines a recognition result 309 based on the first recognition data and the second recognition data. The first recognition data and the second recognition data each indicate at least one of a probability that an object shown in the input image 301 is a real object and a probability that the object is a counterfeit object. As will be described later, the probability of being a real object has a real value between 0 and 1, and the closer the probability is to 0, the more likely the object shown in the input image is a counterfeit object, and the closer the probability is to 1, the more likely the object shown in the input image is a real object. The image recognition device determines the recognition result 309 by integrating the first recognition data and the second recognition data. For example, the image recognition device may calculate the recognition result 309 as a weighted sum of the first recognition data and the second recognition data.
FIG. 4 shows a more detailed structure of the recognition model shown in FIG.
3, the image recognition device can extract feature data 492 from an input image 401 using the feature extraction layer 405 of the recognition model 400. In the following, an example of calculating first recognition data 494 from the feature data 492 using the fixed layer 410 and an example of calculating second recognition data 498 using the sensor specialization layer 420 will be described.
First, the image recognition device generates a generic feature map 493 for the object region of interest by applying a fixed mask 411 to the feature data 492. For example, the image recognition device may apply a mask weight value corresponding to the corresponding value in the fixed mask 411 to the element-wise calculation for each value of the feature data 492. The object region of interest may be a region of interest for a part of an object in the data, for example, a region including components related to a human face. The mask weight value within the object region of interest in the fixed mask 411 may be higher than the mask weight value of the remaining region. Thus, the generic feature map 493 may be a feature map in which components related to a human face in the feature data 492 are emphasized, and the remaining components are less emphasized (e.g., weakened) or eliminated.
The image recognition device calculates the first recognition data 494 from the generic feature map 493. For example, the image recognition device calculates the first recognition data 494 using a recognizer 412 in a fixed layer 410. The recognizer 412 in the fixed layer 410 is designed to output the recognition data from the generic feature map 493. For example, the recognizer outputs a first verification score vector (e.g., first verification score vector=[probability of real object, probability of fake object]) indicating the probability that an object shown in the input image 401 is a real object and a probability that it is a fake object as a classifier. The classifier includes a fully connected layer (FC layer) and a softmax operation.
For reference, in this specification, a verification score is mainly described as an example of recognition data, but is not limited thereto. The recognition data may include information indicating the probability that an object shown in an input image belongs to each of k classes, where k is an integer equal to or greater than 2. In addition, a softmax operation is mainly described as a representative operation for calculating the recognition data, but is not limited thereto, and other non-linear mapping functions may be used.
The image recognition apparatus may then adjust the deformable mask 495 in response to the propagation of the feature data 492 before applying the deformable mask 495 to the target feature map 496. For example, the image recognition apparatus may use at least a portion of the sensor specialization layer 420 (e.g., a mask adjustment layer 421) including the deformable mask 495 to adjust one or more values of the deformable mask 495 according to the feature data 492. Thus, the mask weights of the deformable mask 495 may be updated for each input of the input image 401. The mask adjustment layer 421 may be embodied as a portion of the attention layer, for example, and will be described with reference to FIG. 7 below.
The image recognition device generates a sensor-specific feature map for the region of interest of the image sensor by applying the deformable mask 495 adjusted as described above to the object feature map 496 corresponding to the feature data 492. For example, the image recognition device may extract the object feature map 496 from the feature data 492 using the object extraction layer 422. The object extraction layer 422 may include one or more convolution layers, and the object feature map 496 may be a feature map obtained by applying one or more convolution operations to the feature data 492. The image recognition device may generate the sensor-specific feature map 497 by applying corresponding values in the deformable mask 495 to individual values of the object feature map 496. For example, the image recognition device may apply a mask weight value corresponding to the corresponding value in the deformable mask 495 to each value of the object feature map 496 in the element-wise operation.
In this specification, the region of interest of the image sensor refers to a region of interest related to a part of an object and optical characteristics of the image sensor in the data. For example, the region of interest of the image sensor is a region including a main component in object recognition in consideration of optical characteristics of the image sensor in the data (e.g., lens shading and sensitivity of the image sensor, etc.). As described above, since the mask weight value of the variable mask 495 is adjusted for each input, the region of interest of the image sensor may also change for each input. The sensor-specific feature map may be a feature map in which the region of interest related to the object and optical characteristics of the image sensor is emphasized in the target feature map. For reference, the optical characteristics of the image sensor are reflected in parameters of the sensor-specific layer 420 determined through training, which will be described later with reference to FIG. 9 and FIG. 10.
The image recognition device calculates the second recognition data 498 from the sensor-specific feature map 497. For example, the image recognition device calculates the second recognition data 498 through the recognizer 423 of the sensor specialization layer 420. The recognizer 423 of the sensor specialization layer 420 is designed to output the recognition data from the sensor specialization feature map 497. For example, the recognizer 423 outputs a second verification score vector (e.g., second verification score vector=[probability of real object, probability of fake object]) indicating the probability that the object shown in the input image 401 is a real object and the probability that it is a fake object as a classifier. For reference, even if the recognizer 412 of the fixed layer 410 and the recognizer 423 of the sensor specialization layer 420 have the same structure (e.g., a structure configured of a fully connected layer and a softmax operation), the parameters may be different from each other.
The image recognition device applies an integration operation 430 to the first recognition data 494 and the second recognition data 498 to generate a recognition result 409. For example, the image recognition device may determine the weighted sum of the first recognition data 494 based on a fixed mask and the second recognition data 498 based on a variable mask 495 as the recognition result 409. For example, the image recognition device determines the recognition result as shown in the following Equation (1).
<math num="1"><img file="JP7635495B2_D0001.tif" /></math>
In the above formula (1), the recognition result 409 may be a liveness verification score.<sub>1</sub>is the validation score of the first recognition data 494, score<sub>2</sub>denotes a verification score of the second recognition data 498. α denotes a weighting value for the first recognition data 494, and β denotes a weighting value for the second recognition data 498. According to an embodiment, the image recognition apparatus may apply a weighting value to the second recognition data 498 that is greater than the weighting value for the first recognition data 494. For example, in the above-mentioned Equation (1), β may be greater than α. For reference, Equation (1) is merely an example, and the image recognition apparatus may calculate n pieces of recognition data according to the structure of a recognition model, and apply n weighting values to each of the n pieces of recognition data to calculate a weighted sum. Here, among the n weighting values, the weighting value applied to the recognition data based on the variable mask may be higher than the weighting value applied to the remaining recognition data. Here, n is an integer equal to or greater than 2.
5 and 6 are diagrams illustrating an exemplary structure of a recognition model according to another embodiment.
As shown in FIG. 5, the image recognition device can calculate recognition data based on a verification layer 530 in addition to the recognition data based on a fixed mask 511 and a variable mask (Attention mask) 521 described above with reference to FIG. 3 and FIG. 4. The verification layer 530 includes a recognition device. The first recognition data 581 based on the fixed layer 510 including the fixed mask 511 is indicated as a hard mask score, the second recognition data 582 based on the sensor specific layer 520 including the variable mask 521 is indicated as a soft mask score, and the third recognition data 583 based on the basic liveness verification model is indicated as a 2D liveness score. The image recognition device can calculate the first recognition data 581, the second recognition data 582, and the third recognition data 583 individually from the feature data x commonly extracted from one input image 501 through the feature extraction layer 505.
The image recognition device may determine a recognition result 590 based on the first recognition data 581, the second recognition data 582, and further based on the third recognition data 583. For example, the image recognition device may generate authenticity information indicating whether the object is a real object or a counterfeit object as the recognition result 590. The recognition result 590 may include a value indicating the probability of being a real object as a liveness score.
FIG. 6 illustrates in more detail the structure shown in FIG.
The recognition model includes a fixed layer 610, a sensor specialization layer 620, and a liveness verification model 630. When the image recognition device executes the recognition model using an input image 601, feature data x extracted by the feature extraction layer 605 of the liveness verification model 630 is propagated to the fixed layer 610 and the sensor specialization layer 620.
The fixed layer 610 illustratively includes a fixed mask 611, a fully connected layer 613, and a softmax operation 614. For example, the image recognition device applies the fixed mask 611 to the feature data x, and calculates a generalized feature map 612 as shown in the following Equation (1).
<math num="2"><img file="JP7635495B2_D0002.tif" /></math>
In the above formula (2), Feat<sub>generic</sub>denotes the generic feature map 612, and M<sub>hard</sub>denotes the fixed mask 611, x denotes the feature data,<img file="JP7635495B2_D0003.tif" />denotes an element-wise operation (e.g., element-wise multiplication). The image recognition device uses the generic feature map 612 Feat<sub>generic</sub>is propagated to the fully connected layer 613, and the softmax operation 614 is applied to the output value to calculate the first recognition data 681.<sub>generic</sub>, and the size of the data output from the fully connected layer 613 (e.g., 32×32) may be the same as each other.
The sensor-specific layer 620 illustratively includes an attention layer 621, a fully connected layer 623, and a softmax operation 624. The attention layer 621 will be described in detail with reference to FIG. 7 below. For example, the image recognition device uses the attention layer 621 from the feature data x to generate a sensor-specific feature map 622Feat<sub>specific</sub>The attention feature map can be calculated as
<math num="3"><img file="JP7635495B2_D0004.tif" /></math>
In the above formula (3), Feat<sub>specific</sub>denotes the sensor specific feature map 622, and M<sub>Soft</sub>is a variable mask, and h(x) is an object feature map corresponding to feature data x. The calculation of the object feature map h(x) will be described with reference to FIG. 7 below. The image recognition device calculates the sensor specific feature map 622Feat<sub>specific</sub>is propagated to the fully connected layer 623, and the softmax operation 624 is applied to the output value to calculate the second recognition data 682.<sub>specific</sub>, and the size of the data output from the fully connected layer 623 (e.g., 32×32) may be the same as each other.
The liveness verification model 630 includes a feature extraction layer 605 and a recognition device. According to an embodiment, the image recognition device calculates third recognition data 683 from the extracted feature data x using a fully connected layer 631 and a softmax operation 632. For example, the sizes (e.g., 32×32) of data output from the fully connected layers 613, 623, and 631 may be the same.
The image recognition device calculates a liveness score 690 through a weighted sum operation 689 on the first recognition data 681, the second recognition data 682, and the third recognition data 683.
According to an embodiment, the image recognition device performs the liveness verification model 630, the fixed layer 610, and the sensor specialization layer 620 in parallel. For example, the image recognition device may propagate the feature data x extracted by the feature extraction layer 605 to the fixed layer 610, the sensor specialization layer 620, and the verification model 630 simultaneously or in adjacent times. However, without being limited thereto, the image recognition device may propagate the feature data x sequentially to the liveness verification model 630, the fixed layer 610, and the sensor specialization layer 620. The first recognition data 681, the second recognition data 682, and the third recognition data 683 may be calculated simultaneously, but without being limited thereto, may be calculated at different times depending on the computation time required for each of the fixed layer 610, the sensor specialization layer 620, and the liveness verification model 630.
FIG. 7 is a diagram illustrating an attention layer according to an embodiment.
According to one embodiment, the image recognition device can adjust one or more values of the variable mask 706 using an attention layer 700. For example, the attention layer 700 includes, for example, a mask adjustment layer 710, an object extraction layer 720, and a masking operation. The mask adjustment layer 710 includes a query extraction layer 711 and a key extraction layer 712. The query extraction layer 711, the key extraction layer 712, and the object extraction layer 720 may each include, but are not limited to, one or more convolution layers.
The image recognition apparatus extracts a query feature map f(x) from the feature data 705 using a query extraction layer 711. The image recognition apparatus extracts a key feature map g(x) from the feature data 705 using a key extraction layer 712. The image recognition apparatus extracts an object feature map h(x) using an object extraction layer 720. As described above with reference to FIG. 2, when the input image includes a multi-channel image (e.g., a three-channel image) as a color image, the feature data 705 may be extracted for each channel. The query extraction layer 711, the key extraction layer 712, and the object extraction layer 720 are configured to extract features for each channel.
For example, the image recognition device can determine the value of the variable mask 706 using a softmax function from the product result between the key feature map g(x), which is the result of applying convolution filtering to the feature data 705, and the transposed query feature map f(x). The product result between the key feature map g(x) and the transposed query feature map f(x) indicates the similarity level of all keys for a given query. The variable mask 706 is determined as shown in the following formula (4).
<math num="4"><img file="JP7635495B2_D0005.tif" /></math>
In the above formula (4), M<sub>Soft</sub>is a variable mask 706, f(x) is a query feature map, and g(x) is a key feature map. The image recognition device uses the variable mask 706M determined by the above-mentioned formula (4).<sub>Soft</sub>is applied to the object feature map h(x) using the above-mentioned formula (3). The sensor-specific feature map 709 is obtained by applying the object feature map h(x) to the deformable mask 706M<sub>Soft</sub>The masked result is shown by: The sensor-specific feature maps 709 are generated for each channel as many times as the number of channels.
The attention layer 700 described with reference to FIG. 7 can prevent the vanishing gradient problem by referring to the entire image of the encoder once again at each time point in the decoder. The attention layer 700 can focus on and refer to parts of the entire image that are not the same value and have high relevance to recognition. For reference, in FIG. 7, the attention layer is shown as a self-attention structure in which the same feature data is input as a query, key, and value, but is not limited thereto.
FIG. 8 illustrates an exemplary structure of a recognition model according to a further embodiment.
According to an embodiment, the recognition model 800 includes a feature extraction layer 810, a fixed layer 820, and a first sensor specialization layer 831 to an n-th sensor specialization layer 832, where n may be an integer equal to or greater than 2. The first sensor specialization layer 831 to the n-th sensor specialization layer 832 may each include a variable mask, and the value of each variable mask may be adjusted in response to feature data extracted by the feature extraction layer 810 from the input image 801. The image recognition device calculates a recognition result 809 based on the fixed mask of the fixed layer 820 and the multiple variable masks of the multiple sensor specialization layers. The image recognition device integrates the recognition data calculated from the fixed layer 820 and each of the first sensor specialization layer 831 to the n-th sensor specialization layer 832 to determine the recognition result 809. For example, the image recognition device determines the recognition result 809 as a weighted sum of the multiple recognition data.
Among the above-mentioned multiple variable masks, the parameters of the sensor specializing layer including one variable mask and the parameters of the other sensor specializing layer including the other variable mask may be different from each other. Also, the first sensor specializing layer 831 to the n-th sensor specializing layer 832 may be layers with different structures. For example, among the first sensor specializing layer 831 to the n-th sensor specializing layer 832, one sensor specializing layer may be embodied as an attention layer, and the remaining layers among the first sensor specializing layer 831 to the n-th sensor specializing layer 832 may be embodied as a structure other than the attention layer.
FIG. 9 is a diagram illustrating training of a recognition model according to an embodiment.
According to an embodiment, the training device can train the recognition model using training data. The training data includes a pair of training input and training output. The training input can be an image, and the training output can be a ground truth value for the recognition of an object shown in the image. For example, the training output has a value indicating that the object shown in the training input image is a real object (e.g., 1) or a value indicating that the object is a fake object (e.g., 0). After the training is completed, the recognition model outputs a real value between 0 and 1 as recognition data, which indicates the probability that the object shown in the input image is a real object, but is not limited thereto.
The training device propagates the training input to the temporary recognition model to calculate the temporary output. The recognition model before the training is completed can be referred to as the temporary recognition model. The training device calculates feature data using the feature extraction layer 910 of the temporary recognition model, and propagates the feature data to the fixation layer 920, the sensor specialization layer 930, and the validation layer 940, respectively. In the propagation process, a temporary general feature map 922 and a temporary attention feature map 932 are calculated. The training device calculates a first temporary output from the fixation layer 920, a second temporary output from the sensor specialization layer 930, and a third temporary output from the validation layer 940. The training device calculates a loss based on a loss function from each temporary output and the training output. For example, the training device calculates a first loss based on the first temporary output and the training output, a second loss based on the second temporary output and the training output, and a third loss based on the third temporary output and the training output.
<math num="5"><img file="JP7635495B2_D0006.tif" /></math>
The training device calculates the weighted loss of the losses calculated as shown in the above formula (5). In the above formula (5), the Liveness loss is the total loss 909, and the Loss<sub>1</sub>is the first loss, Loss<sub>2</sub>is the second loss, Loss<sub>3</sub>denotes the third loss. α denotes a weighting value for the first loss, β denotes a weighting value for the second loss, and γ denotes a weighting value for the third loss. The training device updates the parameters of the temporary recognition model until the overall loss 909 reaches the threshold loss. Depending on the design of the loss function, the training device can increase or decrease the overall loss 909. For example, the training device can update the parameters of the temporary recognition model via back propagation.
According to one embodiment, for an initial recognition model that is not trained, the training device updates all parameters of the feature extraction layer 910, the fixation layer 920, the sensor specialization layer 930, and the validation layer 940 during training. Here, the training device can train the initial recognition model using generic training data 901. The generic training data 901 can include images acquired by any image sensor as training input. The training images of the generic training data 901 can be acquired by any type of image sensor, but are not limited to this, and can be acquired by various types of image sensors. The recognition model trained using the generic training data 901 can be referred to as a generic recognition model. The generic recognition model can be used for, for example, high-end performance (high-end The image sensor of a flagship-level electronic device has excellent optical performance. The general-purpose recognition model may output FR (False Rejection) and FA (False Acceptance) results for a specific type of image sensor. This is because the optical characteristics of the image sensor of that type are not reflected in the general-purpose recognition model. The FR result indicates a result of mistaking a true for a false, and the FA result indicates a result of mistaking a false for a true.
The training device generates a recognition model for a specific type of image sensor from the generic recognition model. For example, the training device fixes the values of the fixed mask 921 included in the fixed layer 920 and the parameters of the validation layer 940 in the generic recognition model during training. The training device updates the parameters of the sensor-specific layer 930 during training with the provisional recognition model. The training device calculates the overall loss 909 as described above, and iteratively adjusts the parameters of the sensor-specific layer 930 until the overall loss 909 reaches a threshold loss. For example, the training device can update the parameters of the attention layer 931 (e.g., connection weights) and the parameters of the fully connected layer from the sensor-specific layer 930.
Here, the training device can use both the generic training data 901 and the sensor-specific training data 902 to train the sensor-specific layer 930 of the recognition model. The sensor-specific training data 902 may be data consisting only of training images acquired by a specific type of image sensor. The type of image sensor may be classified according to the optical characteristics of the image sensor as described above. The training device can use the sensor-specific training data 902 to update parameters of the sensor-specific layer 930 based on the calculated loss as described above.
In the early stage of the release of a new product, the amount of sensor-specific training data 902 may not be sufficient, but in order to prevent over-fitting due to a lack of training data, the training device also uses the generic training data 901 for training. The amount of generic training data 901 is greater than the amount of sensor-specific training data 902. In other words, the training device can generate a recognition model having a sensor-specific layer 930 specialized for individual optical characteristics through the conventional generic training data 901 (e.g., a database of millions of existing images) together with a small amount (e.g., tens of thousands) of sensor-specific training data 902. Thus, the training device can generate a recognition model specialized for a specific type of image sensor from a generic recognition model in a relatively short time. A previously undiscovered spoofing attack Even if a new spoofing attack occurs, the training device learns the parameters of the sensor-specific layer so as to more quickly defend against new spoofing attacks, and the trained parameters of the sensor-specific layer are urgently distributed to each image recognition device (e.g., an electronic terminal shown in the following FIG. 10). The sensor-specific training data 902 includes images corresponding to newly reported FR and FA results.
FIG. 10 is a diagram illustrating parameter updates of a sensor specific layer in a recognition model according to an embodiment.
The image recognition system includes a training device 1010, a server 1050, and electronic terminals 1060, 1070, and 1080.
The processor 1011 of the training device 1010 may train the recognition model as described above with reference to Fig. 9. The training device 1010 may also perform additional training on the sensor-specific layer 1043 of the recognition model 1040 after the initial training on the initial recognition model 1040 is completed. For example, in response to the occurrence of a new spoofing attack, the training device 1010 may re-train the sensor-specific layer 1043 of the recognition model 1040 based on training data related to the new spoofing attack.
The memory 1012 of the training device 1010 stores the recognition model 1040 before and after the training is completed. The memory 1012 also stores the generic training data 1020, the sensor-specific training data 1030, and parameters of the feature extraction layer 1041, the sensor-specific layer 1043, and the fixed layer 1042 in the recognition model 1040. Upon completing the training described above with reference to Fig. 9, the training device 1010 can distribute the trained recognition model 1040 through communication with the server 1050 (e.g., wired communication or wireless communication).
Also, the server 1050 may distribute only some of the parameters of the recognition model 1040 to each electronic terminal instead of distributing all of the parameters. For example, the training device 1010 may upload the parameters of the retrained sensor specialization layer 1043 to the server 1050 in response to the completion of additional training for the sensor specialization layer 1043 of the recognition model 1040. The server 1050 may provide only the parameters of the sensor specialization layer 1043 to the electronic terminals 1060, 1070, and 1080 of the electronic terminal group 1091 having a specific type of image sensor. The electronic terminals 1060, 1070, and 1080 belonging to the electronic terminal group 1091 are equipped with image sensors having the same or similar optical characteristics. The server 1050 may distribute the additionally trained sensor specialization layer 1043 to the corresponding electronic terminal in response to at least one of the completion of additional training for the sensor specialization layer 1043 of the recognition model 1050 and an update request received from the electronic terminal. The update request may be a signal from any terminal requesting an update of the recognition model to the server.
10, the training device 1010 is illustrated as storing only one type of recognition model 1040, but is not limited to this. The training device may store other types of recognition models and provide updated parameters to other terminal groups 1092 as well.
Each of the electronic terminals 1060, 1070, and 1080 belonging to the electronic terminal group 1091 described above receives parameters of the sensor specialization layer 1043 including the variable mask from the external server 1050 in response to an update command. The update command may be a command input by a user, but is not limited thereto, and may be a command received by the electronic terminal from the server. Each of the electronic terminals can update the received parameters to the sensor specialization layers 1062, 1072, and 1082. Here, each of the electronic terminals 1060, 1070, and 1080 fixes the parameters of the remaining feature extraction layers 1061, 1071, and 1081 and the fixed layers 1063, 1073, and 1083. For example, each of the electronic terminals 1060, 1070, and 1080 can hold the value of the fixed mask before, during, and after updating the parameters of the sensor specialization layers 1062, 1072, and 1082. For reference, when FR results and FA results that depend on the unique optical characteristics of an individual image sensor are reported, parameters can be distributed as a result of a training device training the above-mentioned FR results and FA results to the sensor-specific layer 1043.
As another example, the electronic terminal may request sensor specific parameters 1043 corresponding to optical characteristics identical or similar to those of the currently attached image sensor from the external server 1050. The server 1050 may search for sensor specific parameters 1043 corresponding to the optical characteristics requested by the electronic terminal, and provide the searched sensor specific parameters 1043 to the corresponding electronic terminal.
In FIG. 10, an example in which the server 1050 distributes the parameters of the sensor specific layer 1043 has been described, but the present invention is not limited thereto. When a change occurs in the fixed mask value of the fixed layer 1042, the server 1050 distributes it to the electronic terminals 1060, 1070, and 1080. The electronic terminals 1060, 1070, and 1080 may update the fixed layers 1063, 1073, and 1083 as necessary. For example, when a general FR result and FA result that are unrelated to the specific optical characteristics of an individual image sensor are reported, the training device may adjust the fixed mask value of the fixed layer 1042. For reference, the update of the fixed mask may improve the recognition performance in various electronic terminals having various types of image sensors in a general purpose manner. The update of the variable mask corresponding to the individual optical characteristics may improve the recognition performance in an electronic terminal having an image sensor of the corresponding optical characteristics.
If a neural network learns data acquired using only a specific device, the recognition rate of the device is high. However, when the same neural network is installed in another device, the recognition rate decreases. As described above with reference to FIG. 9 and FIG. 10, the recognition model according to an embodiment may have a sensor-specific layer specialized for each image sensor through a small amount of additional training instead of retraining the entire existing network. Therefore, since an emergency patch for the recognition model is possible, the privacy and security of the electronic terminals 1060, 1070, and 1080 may be more safely protected.
11 and 12 are block diagrams showing a configuration of an image recognition device according to an embodiment.
The image recognition device 1100 shown in FIG. 11 includes an image sensor 1110, a processor 1120, and a memory 1130.
The image sensor 1110 receives an input image. For example, the image sensor 1110 may be a camera sensor that captures a color image. The image sensor 1110 may also be a dual phase detection sensor (2PD sensor) that can obtain a disparity image for any pixel using a phase difference between the left and right. Since the disparity image is generated immediately by the dual phase detection sensor, a depth image can be calculated from the corresponding disparity image without using a stereo sensor or a conventional depth extraction method.
Unlike time-of-flight (MToF) and structured light depth sensors, the 2PD sensor is mounted on the device 1100 without additional form factor and sensor cost. For example, the 2PD sensor is a Contact Image Sensor (CIS). Unlike a CIS sensor, the 2PD sensor includes detection elements each composed of two photodiodes (e.g., a first photodiode and a second photodiode). Thus, two images are generated through imaging by the 2PD sensor. The two images may include an image detected by a first photodiode (e.g., a left photodiode) and an image detected by a second photodiode (e.g., a right photodiode). The two images are slightly different from each other due to the difference in physical distance between the photodiodes. The image recognition device 1100 uses the two images to calculate disparity according to the difference in distance using a triangulation method or the like, and estimates the depth of each pixel from the calculated disparity. Unlike a CIS sensor that outputs three channels, the output of the 2PD sensor outputs one channel image for each of the two photodiodes, thereby reducing the memory and amount of calculation used. This is because to estimate disparity from an image acquired by a CIS sensor, a three channel image pair (e.g., six channels in total) is required, whereas to estimate disparity from an image acquired by a 2PD sensor, only a one channel image pair (e.g., two channels in total) is required.
However, without being limited thereto, the image sensor 1110 may include an infrared sensor, a radar sensor, an ultrasonic sensor, a depth sensor, and the like.
The processor 1120 extracts feature data from the input image using the feature extraction layer. The processor 1120 outputs a recognition result for an object shown in the input image based on a fixed mask from the extracted feature data and a deformable mask that is adjusted in response to the extracted feature data. The processor 1120 updates the parameters of the sensor specialization layer stored in the memory 1130 when it receives the parameters of the sensor specialization layer from the server via communication.
The memory 1130 temporarily or permanently stores the recognition model and data generated in the process of implementing the recognition model. When new parameters of the sensor specific layer are received from the server, the memory 1130 can replace existing parameters with the newly received parameters.
Referring to Figure 12, a computing device 1200 is a device for recognizing images using the image recognition method described above. In one embodiment, the computing device 1200 corresponds to the electronic terminal described with reference to Figure 10 and/or the device 1100 described with reference to Figure 11. The computing device 1200 may be, for example, an image processing device, a smartphone, a wearable device, a tablet computer, a netbook, a laptop, a desktop, a personal digital assistant (PDA), or a head mounted display (HMD).
12, a computing device 1200 includes a processor 1210, a storage device 1220, a camera 1230, an input device 1240, an output device 1250, and a network interface 1260. The processor 1210, the storage device 1220, the camera 1230, the input device 1240, the output device 1250, and the network interface 1260 communicate via a communication bus 1270.
The processor 1210 executes functions and instructions for execution within the computing device 1200. For example, the processor 1210 processes instructions stored in the storage device 1220. The processor 1210 may perform one or more of the operations described above with reference to Figures 1-11.
The storage device 1220 stores information or data necessary for the execution of the processor 1210. The storage device 1220 may include a computer-readable storage medium or a computer-readable storage device. The storage device 1220 stores instructions for execution by the processor 1210 and stores relevant information during the execution of software or applications by the computing device 1200.
The camera 1230 captures an input image for image recognition. The camera 1230 captures multiple images (e.g., multiple frame images). The processor 1210 outputs a recognition result for a single image using the above-mentioned recognition model.
Input device(s) 1240 receive input from a user via haptic, video, audio, or touch input. Input device(s) 1240 may include a keyboard, mouse, touch screen, microphone, or any other device capable of detecting input from a user and communicating the detected input.
The output device 1250 provides the output of the computing device 1200 to the user through a visual, auditory, or tactile channel. The output device 1250 may include, for example, a display, a touch screen, a speaker, a vibration generator, or any other device capable of providing output to a user. The network interface 1260 communicates with external devices through a wired or wireless network. The output device 1250 may provide the result of recognizing the input data (e.g., access allowed and/or access denied) to the user using at least one of visual information, auditory information, and tactile information.
According to an embodiment, the computing device 1200 grants authority based on the recognition result. The computing device 1200 may allow access to at least one of the operations and data of the computing device 1200 based on the authority. For example, the computing device 1200 may grant authority in response to a case where the recognition result verifies that the user is a user registered in the computing device 1200 and is a real object. If the computing device 1200 is in a locked state, the computing device 1200 may unlock the locked state based on the authority. As a different example, the computing device 1200 may allow access to a financial settlement function in response to a case where the recognition result verifies that the user is a user registered in the computing device 1200 and is a real object. As a further example, the computing device 1200 may visualize the recognition result via the output device 1250 (e.g., a display) after the recognition result is generated.
The above-described devices may be realized by hardware components, software components, or a combination of hardware and software components. For example, the devices and components described in the present embodiment may be realized by a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a programmable logic controller ... The processor may be embodied using one or more general-purpose or special-purpose computers, such as a processor unit, a microprocessor, or a different device that executes and responds to instructions. The processor executes an operating system (OS) and one or more software applications that run on the operating system. The processor also accesses, stores, manipulates, processes, and generates data in response to the execution of the software. For ease of understanding, a single processor may be described, but one skilled in the art will appreciate that a processor may include multiple processing elements and/or multiple types of processing elements. For example, a processor may include multiple processors or a processor and a controller. Other processing configurations are also possible, such as parallel processors.
The software may include computer programs, codes, instructions, or any combination of one or more thereof, capable of configuring or instructing a processing device to operate as desired, either independently or in combination. The software and/or data may be embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or transmitted signal wave, either permanently or temporarily, to be interpreted by the processing device or to provide instructions or data to the processing device. The software may be distributed across computer systems coupled to a network, and may be stored and executed in a distributed manner. The software and data may be stored on one or more computer readable recording media.
The method according to the present invention may be embodied in the form of program instructions to be executed by various computer means and recorded on a computer-readable recording medium. The recording medium may include program instructions, data files, data structures, and the like, alone or in combination. The recording medium and the program instructions may be specially designed and constructed for the purposes of the present invention, or may be well known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, flash memory, and the like. Examples of program instructions include not only machine code, such as generated by a compiler, but also high-level language code executed by a computer using an interpreter, and the like. The hardware devices may be configured to operate as one or more software modules to perform the operations shown in the present invention, and vice versa.
Although the embodiments have been described above with reference to limited drawings, those skilled in the art may apply various technical modifications and variations based on the above description. For example, the described techniques may be performed in a different order than described, and/or the components of the described systems, structures, devices, circuits, etc. may be combined or combined in a different manner than described, or may be replaced or substituted with other components or equivalents to achieve appropriate results.
Therefore, the scope of the present invention should not be limited to the disclosed embodiment, but should be determined by the appended claims and their equivalents.
100 Neural Networks
301 Input Image
309 Recognition results
310 Recognition Model
311 Feature Extraction Layer
312 Fixed Layer
313 Sensor Specialized Layer
321 Fixed Mask
322 Variable Mask
1010 Training Equipment
1100 Image Recognition Device
1200 Computing Device
18 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20110050883A1 | Cites | United States of America |
| JP2011237932A | Cites | Japan |
| CN110210543A | Cites | China |
| 積田 貴幸 外3名,競泳プール映像における泳者位置とストローク数の推定,電子情報通信学会技術研究報告,Vol.117 No.485,日本,電子情報通信学会,2018年03月01日,p.11-p.16 | Non-patent | – |
7 members in 5 offices
Members7
| Document | Office | Kind | |
|---|---|---|---|
| CN112926574A | China | A | |
| EP3832542A1 | European Patent Office (EPO) | A1 | |
| US2021174138A1 | United States of America | A1 | |
| KR20210071410A | Republic of Korea | A | |
| JP2021093144A | Japan | A | |
| US11354535B2 | United States of America | B2 | |
| JP7635495B2This record | Japan | B2 |
8 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 7635495
- Application
- 184118
Titles2
- Japanese
- センサ特化イメージ認識装置及び方法
- English
- Sensor specific image recognition device and method
Classification
- CPC, 16
- G06V10/25
- G06V10/7715
- G06V10/26
- G06V10/10
- G06N3/08
- G06V10/22
- G06V10/147
- G06V10/40
- G06N3/045
- G06F18/214
- G06V10/454
- G06V10/82
- G06N3/09
- G06N3/0464
- G06V20/80
- G06F18/217
- IPC, 1
- G06T7 00
