Textual and image based search
12 claims: 7 independent, 5 dependent
- 1計算システムであって、1つまたはそれ以上のプロセッサと、前記1つまたはそれ以上のプロセッサによって実行されたときに、前記1つまたはそれ以上のプロセッサに少なくとも、ユーザ装置からテキストクエリを受信することと、前記テキストクエリに対応する複数の結果を判定することと、前記テキストクエリが定義されたカテゴリに対応することを判定することと、前記テキストクエリが前記定義されたカテゴリに対応するという判定に応じて、画像の絞り込みオプションを前記ユーザ装置に提供させることと、前記ユーザ装置から前記画像の絞り込みオプションの一部としてのオブジェクトの画像を受信することと、前記オブジェクトを表すオブジェクト特徴ベクトルを生成することと、前記オブジェクト特徴ベクトルを、前記複数の結果として返される画像のセグメントに対応する複数の保存された特徴ベクトルと比較し、前記オブジェクト特徴ベクトルと前記複数の保存された特徴ベクトルとのそれぞれの間の類似度を示す複数の類似度スコアを生成することと、前記複数の類似度スコアに基づいて、前記複数の結果のランク付けされたリストを生成することと、前記オブジェクトの画像の前記受信に応じて前記ランク付けされたリストを前記ユーザ装置に提示することと、を引き起こすプログラム命令を保存するメモリと、を備えた計算システム。
- 2前記複数の保存された画像セグメントのそれぞれは、画像全体より少ない部分に対応する、請求項1に記載の計算システム。
- 3コンピュータ実装方法であって、ユーザ装置からクエリを受信することと、前記クエリに基づいて第1の複数の画像を判定することと、定義されたカテゴリに前記クエリが対応することを判定することと、前記定義されたカテゴリに前記クエリが対応する判定に応じて、視覚的な改良オプションを前記ユーザ装置に提示させることと、前記ユーザ装置から前記視覚的な改良オプションの一部としてのオブジェクトの画像を受信することと、前記オブジェクトの前記画像を、前記第1の複数の画像のそれぞれの少なくとも1つの画像セグメントと比較し、 前記オブジェクトの前記画像と前記第1の複数の画像のそれぞれの1つとの間の類似度を示す複数のそれぞれの類似度スコアを生成することと、前記複数のそれぞれの類似度スコアに基づいて、前記第1の複数の画像の少なくとも一部のランク付けされたリストを判定することと、前記ランク付けされたリストに従って、前記第1の複数の画像の前記少なくとも一部を前記ユーザ装置に提示することと、を含む、コンピュータ実装方法。
- 4前記オブジェクトの画像を処理して、前記オブジェクトの画像に表される前記オブジェクトのオブジェクトタイプを判定することをさらに含み、前記オブジェクトの画像を比較することは、前記オブジェクトを表すオブジェクト特徴ベクトルを生成することと、前記オブジェクト特徴ベクトルを、同じオブジェクトタイプを有する前記第1の複数の画像で表されるオブジェクトに対応する複数の保存された特徴ベクトルと比較することと、を含む、請求項3に記載のコンピュータ実装方法。
- 5前記ユーザ装置に、前記ユーザ装置のカメラの視野内の前記オブジェクトを検出させることと、前記ユーザ装置に、前記オブジェクトに対応するオブジェクトタイプを判定させることと、前記ユーザ装置に、前記ユーザ装置のディスプレイにオブジェクトタイプ識別子を提示させることと、をさらに含む、請求項3または4に記載のコンピュータ実装方法。
- 6オブジェクトタイプに対応するキーワードを生成することと、クエリの一部としてキーワードを含めることと、をさらに含む、請求項5に記載のコンピュータ実装方法。
- 7前記第1の複数の画像の前記一部、前記第1の複数の画像、または前記クエリの少なくとも1つに対応する複数のキーワードを判定することと、提示および前記ユーザ装置のユーザによる選択のために前記複数のキーワードのそれぞれを提供することと、をさらに含む、請求項3-6のいずれか一項に記載のコンピュータ実装方法。
- 8前記オブジェクトの前記画像を比較することは、前記オブジェクトを表すオブジェクト特徴ベクトルを生成することと、前記オブジェクトが定義されたカテゴリーの1つのオブジェクトタイプのどれに対応するか判定することと、前記オブジェクトタイプの種類に対応するラベルを生成することと、前記ラベルを使用して、 可能性が最も高い位置に基づいて、 前記第1の複数の画像内の 画像セグメントに関連付けられた保存された特徴ベクトルを選択する ことと、前記オブジェクト特徴ベクトルと、 選択された 前記画像セグメントに関連付けられた保存された特徴ベクトルとを比較することと、をさらに含む、請求項3-7のいずれか一項に記載のコンピュータ実装方法。
- 9計算システムの少なくとも1つのプロセッサによって実行されたときに、前記計算システムに少なくとも、ユーザ装置からクエリを受信することと、定義されたカテゴリに前記クエリが対応することを判定することと、前記カテゴリに前記クエリが対応することに応じて、前記ユーザ装置において視覚的な改良オプションを有効にさせることと、前記視覚的改良オプションの一部として、前記ユーザ装置からストリーミングビデオを受信することと、前記ストリーミングビデオの少なくとも一部を処理して、前記ストリーミングビデオで表される1つまたはそれ以上のオブジェクトのオブジェクトタイプを識別することと、前記ユーザ装置のディスプレイ上に、前記ストリーミングビデオのプレゼンテーションと同時に、前記1つまたはそれ以上のオブジェクトの前記オブジェクトタイプを提示させることと、前記ユーザ装置から前記オブジェクトタイプの選択を受け取ることと、前記クエリと前記選択されたオブジェクトタイプとの両方に対応する複数の保存画像を判定することと、前記ユーザ装置の前記ディスプレイ上に前記複数の保存された画像を提示させることと、を実行させる命令を保存する非一時的コンピュータ可読記憶媒体。
- 10前記定義されたカテゴリは食品であり、前記ストリーミングビデオは、前記ユーザ装置のカメラの視野内に現在ある食べ物の表現を含む、請求項9に記載の非一時的コンピュータ可読記憶媒体。
- 11前記命令は、前記計算システムに少なくとも、前記定義されたカテゴリに基づいて、前記複数の保存された画像を判定する際に考慮される候補画像を判定させる、請求項9または10に記載の非一時的コンピュータ可読記憶媒体。
- 12前記少なくとも1つのプロセッサに前記複数の記憶された画像を判定させる前記命令は、前記計算システムに少なくとも、前記クエリ、前記少なくとも1つのオブジェクトタイプ、または前記定義されたカテゴリに対応する少なくとも1つのキーワードを判定することと、前記少なくとも1つのキーワードに基づいて、前記複数の保存された画像を判定することと、を生じさせる、請求項9-11のいずれか一項に記載の非一時的コンピュータ可読記憶媒体。
Independent claims12
144 paragraphs, as filed
This application claims the benefit of U.S. Application No. 15/713,567, filed September 22, 2017, entitled "Text and Image-Based Search," which is incorporated herein by reference in its entirety. .
As more and more accessible digital content becomes available to users and customers, it becomes increasingly difficult for users to find the content they are searching for. Although several different search techniques exist, such as keyword searches, there are many inefficiencies in such systems.
<figref num="1A">3 shows an input image obtained by a user device according to the described implementation;</figref><figref num="1B">1A shows a visual search result for a selected object in the input image of FIG. 1A, where, according to the described implementation, the result is an image containing objects that are visually similar to the selected object; FIG.</figref><figref num="2">2 is an exemplary image processing process according to the described implementation.</figref><figref num="3">A representation of a segmented image according to the implementation.</figref><figref num="4">1 is an exemplary object matching process according to the described implementation;</figref><figref num="5">2 is another example of an object matching process in accordance with the described implementation.</figref><figref num="6A">3 illustrates an input image obtained by a user device according to the described implementation;</figref><figref num="6B">FIG. 6A shows the visual search results for objects of interest in the input image, and according to the described implementation, the results include images related to the objects of interest.</figref><figref num="7">3 is an example object category matching process in accordance with the described implementation;</figref><figref num="8A">In accordance with the described implementation, a query is shown with options that provide visual enhancements.</figref><figref num="8B">Figure 3 illustrates visually improved input according to the described implementation.</figref><figref num="8C">8B shows search results for the query of FIG. 8A refined based on the visual refinement of FIG. 8B, according to the described implementation; FIG.</figref><figref num="9">1 is an exemplary text and image matching process according to the described implementation;</figref><figref num="10A">FIG. 4 illustrates an example visual refinement input of a query, according to the described implementation; FIG.</figref><figref num="10B">FIG. 10A shows search results and visual improvements to the query of FIG. 10A according to the described implementation.</figref><figref num="11">1 illustrates an example computing device according to one implementation.</figref><figref num="12">12 shows an exemplary configuration of components of a computing device such as that shown in FIG. 11;</figref><figref num="13">1 is a pictorial diagram of an example implementation of a server system that may be used in a variety of implementations.</figref>
Systems and methods are described herein that facilitate retrieval of information based on selection of one or more objects of interest from larger images and/or videos. In some implementations, images of objects may be supplemented with other forms of search input, such as text or keywords to narrow the results. In other implementations, images of objects may be used to supplement or improve existing searches, such as text or keyword searches.
In many image-based queries (e.g., fashion design, interior design, etc.), the user is not interested in the specific objects represented in the image (e.g., a dress, a couch, a lamp, etc.); the entire image, including the objects and how those objects are placed (e.g., the choice of style between a shirt and a skirt, the placement of a couch relative to the television). For example, a user may provide an image containing a pair of shoes, indicate the shoe as an object of interest, and present other images containing shoes that are visually similar to the selected shoe, along with those other shoes and pants, You may want to display style combinations with other objects such as shirts, hats, wallets, etc.
In one embodiment, a user may initiate a search by providing or selecting an image that includes an object of interest. The described implementations may then process the image to detect objects of interest and/or receive selections from the user indicating objects of interest as represented in the images. The portion of the image that includes the object of interest may be segmented from the rest of the image, the determined object of interest, and/or the generated object feature vector representing the object of interest. Based on the determined object of interest and/or object feature vector, the stored feature vectors of other stored image segments are compared with the object feature vector of the object of interest to identify the object of interest and the object feature vector of the object of interest. Other images containing visually similar objects can be determined. The stored image may be a particular image of another visually similar object, or an image of multiple objects, often including one or more objects that are visually similar to the object of interest. , thereby providing an image showing how objects, such as the object of interest, are combined with other objects. The user can select one of the presented images, select additional or other objects, or perform other actions.
In some implementations, a stored image may be segmented into various regions, objects represented by those segments may be determined, and feature vectors representing those objects may be generated and associated with the segments of the image. . Once an object feature vector is generated for an object of interest, the object feature vector may be compared with stored feature vectors of different segments of the stored image to detect images containing visually similar objects. . Comparing object feature vectors with stored feature vectors corresponding to segments of the image to ensure visual similarity to the object of interest, even if the image of interest contains representations of many other objects. Images containing objects can be identified.
In yet another embodiment, the user may select multiple objects of interest, and/or the selected object of interest may be an object of positive interest or an object of negative interest. You can specify whether Objects of positive interest are objects selected by a user who is interested in seeing images of other visually similar objects. Objects of negative interest are user-selected objects that the user does not want included in other images. For example, if a user selects a chair and lamp positive object and a rug negative object from an image, the implementation described here selects a chair and lamp that are visually similar to the selected chair and lamp. identifying other images containing, which may include representations of other objects, but do not include lags that are visually similar to the selected lag;
In some implementations, images of objects of interest may be processed to detect types of objects of interest. Based on the determined type of object of interest, it can be determined whether the type of object of interest corresponds to a defined category (eg, food, fashion, home decor). If the type of object of interest corresponds to a defined category, multiple query types can be selected, from which the results of different queries are returned and mixed as a result of the input image. For example, some query types can be configured to receive query keywords and provide image results based on the keywords. Other query types can be configured to receive image-based queries, such as feature vectors, compare the image queries to stored image information, and return results corresponding to the queries.
Different query types can be used to provide results corresponding to defined categories. For example, if the object of interest is determined to be a type of food, one query type might include content (text, images, videos, etc.) that is related to, but does not include, a visual representation of the object of interest. , audio, etc.). Another query type may return images or videos containing objects that are visually similar to the object of interest. In such examples, results from various query types can be determined and mixed to provide a single response to the query that includes results from each query type.
In yet another example, a user may initiate a text-based query and then refine the text-based query using images of objects of interest. For example, a user can enter a text-based query such as "summer clothes," and the described implementation can process the text-based query to determine that the query corresponds to a predefined category (such as fashion). The user can then provide an image containing an object of interest and use that object of interest to improve or change the results of the text-based query. For example, if the object of interest is red tops, the search results that match the text query are processed to return results that include representations of other tops that are visually similar to the object of interest (in this example, red tops). Can be detected. The results may then be ranked such that results that match the text-based search and include objects that are visually similar to the object of interest are ranked highest and presented to the user first.
FIG. 1A shows an input image obtained by a user device 100 in accordance with the described implementation. In this example, the user wants to search for images that include objects that are visually similar to the object of interest 102, in this example a high heeled shoe. As will be appreciated, any object that can be represented in an image can be an object of interest. To provide an object of interest, the user generates images using one or more cameras of user device 100, provides images from memory of user device 100, and provides images external to user device 100. providing images stored in memory, selecting images provided by the systems and methods described herein (e.g., resulting images), and/or obtaining images from another source or location. can be provided or selected.
In this example, the user generated image 101 using the camera of user device 100. The image includes multiple objects, such as a high heel shoe 102, a lamp 104-2, a bottle 104-1, and a table 104-3. Once an image is received, the image can be segmented and processed to detect objects within the image and determine objects of interest on which to perform a search. As described further below, the image may be processed using any one or more of a variety of image processing techniques, such as object recognition, edge detection, etc., to identify objects within the image.
Objects of interest may be determined based on the object's relative size, whether the object is in focus in the image, the object's location, and so on. In the illustrated example, high-heeled shoes 102 are determined to be an object. This is because it is positioned towards the center of the image 101 and is physically in front and in focus of other objects 104 represented in the image. Other implementations allow the user to select or specify objects of interest.
Once the object of interest is determined, the input image is segmented to generate a feature vector representing the object of interest. Generation of feature vectors will be described in detail below. In contrast to typical image processing, objects of interest are extracted or segmented from other parts of the image 101, and the object of interest feature vector is generated as shown. Generating object feature vectors that represent only the object of interest rather than the entire image improves the quality of the matching described herein. Specifically, as described further below, the stored images are segmented, objects are detected in different segments of the stored images, and respective feature vectors representing objects represented in those images are generated. As such, each saved image may include multiple segments and multiple different feature vectors, each feature vector representing an object represented in the image.
Once object feature vectors representing objects of interest are generated, they may be compared to stored feature vectors representing individual objects included in segments of the stored image. As a result, even if the entire stored image is significantly different from the input image 100, based on the comparison between the object of interest and the stored feature vector representing a segment of the stored image that is smaller than the entire image, , it may be determined that the stored image includes a representation of the object that is visually similar to the object feature vector.
In some implementations, the type of object of interest may be determined and used to limit or reduce the number of stored feature vectors that are compared to the object feature vector. For example, if the object of interest is determined to be a shoe (such as a high-heeled shoe), the object feature vector may only be compared to stored feature vectors known to represent other shoes. In another example, saved feature vectors may be selected for comparison based on locations within images where objects of a certain type are commonly located. For example, if again it is determined that the object of interest is a type of shoe, it may be further determined that shoes are typically represented in the bottom third of the image. In such an example, only the stored feature vectors corresponding to segments of the image that are in the lower third of the stored image may be compared to the object feature vector.
Once the saved feature vector is compared to the object feature vector, a similarity score is determined representing the similarity between the object feature vector and the saved feature vector, and the one determined to have the highest similarity score. Stored images associated with the stored feature vectors are returned as a result of the search. For example, FIG. 1B shows a visual search result for an object of interest, which is a high heel shoe 102 in the input image 100 of FIG. 1A, according to the described implementation.
In this example, an object feature vector representing a high heel shoe 102 is compared to stored feature vectors representing objects represented in different segments of the stored image returned as result image 110. As described below, the saved image is segmented, objects are detected, object feature vectors are generated, and the saved image, segments, locations of those segments within the saved image, and datastore An association can be made between feature vectors held in .
In this example, the object feature vector representing object of interest 102 is compared to the stored feature vectors representing objects 113-1, 113-2A, 113-2B, 113-2C, 113-3, 113-4, etc. The similarity between the object feature vector and the stored feature vector is determined. As shown, the images 110 returned in response to the search include objects in addition to those determined to be visually similar to the object of interest. For example, a first image 110-1 includes a segment 112-1 that includes an object 113-1 that has been determined to be visually similar to the object of interest 102, as well as other objects such as a person 105, clothing, etc. include. As discussed further below, the saved image returned may include a number of segments and/or objects. Alternatively, the saved images returned may include only visually similar objects. For example, the fourth image 110-4 includes a single segment 112-4 that includes an object 113-4 that is visually similar to the object of interest 102, but no other objects are represented in the image.
The second image 110-2 includes a plurality of segments 112-2A, 112-2B, 112-2C and a plurality of objects 113-2A, 113-2B, 113-2C of the same type as the object of interest 102 . In such an example, the object feature vector may be associated with the second image 110-2 and compared to one or more feature vectors representing different objects. In some implementations, the similarities between the object feature vectors and the saved feature vectors associated with the second image 110-2 are averaged, and the average is the similarity of the second image 110-2. used as a gender. In other implementations, the highest similarity score, lowest similarity score, median similarity score, or other similarity score is selected as representative of the visual similarity between the object of interest and the image. obtain.
Upon receiving the results from the comparison of the generated object feature vector and the stored feature vector, the user may view and/or interact with the image 110 provided in response. Images are ranked and ranked such that images associated with saved feature vectors with higher similarity scores are ranked higher and displayed before images associated with feature vectors with lower similarity scores. can be presented.
FIG. 2 is an example image processing process that may be performed to generate stored feature vectors and labels representing segments and objects of stored images maintained in a data store, in accordance with the described implementation. The example process 200 begins, as at 202, by selecting an image to process. Any image can be processed according to the implementation described with respect to Figure 2. For example, an image stored in an image data store, an image generated by a camera of the user device, an image held in the memory of the user device, or other image may be selected for processing according to the example process 200. can. In some cases, the image processing process 200 is used to identify segments, labels, and/or corresponding feature vectors for all objects in the saved image such that the segments, labels, and/or feature vectors are associated with the saved image. can be generated and used in determining whether an object of interest is visually similar to one or more objects represented in the stored image. In another example, image processing process 200 may be performed on an input image to generate labels and/or object feature vectors for determined objects of interest.
When you select an image, the image will be divided as shown in 204. Various segmentation techniques can be used, such as circle packing algorithms, superpixels, etc. The segment of the image may then be processed to remove background regions of the image from consideration, such as 206. Determining the background region is, for example, a combination of careful constraints (e.g., a salient object is likely to be in the center of an image segment) and unique constraints (e.g., a salient object is likely to be different from the background). This can be done using In one embodiment, each segment (S<sub>i</sub>), unique constraints can be computed using a combination of color, texture, shape, and/or other feature detection. Pairwise Euclidean distance of all pairs of segments: L2(S<sub>i</sub>,S<sub>j</sub>)teeth,<img file="JP7472016B2_D0001.tif" />is also calculated. segment S<sub>i</sub>Unique constraint U or U<sub>i</sub>teeth,<img file="JP7472016B2_D0002.tif" />It can be calculated as Each segment S<sub>i</sub>A careful constraint on<img file="JP7472016B2_D0003.tif" />It can be calculated as Here, X' and Y' are the center coordinates of the image.
Then one or more segments S', a subset of S are defined as U(s)-A(<u style="Single">s</u>)>t. t is a threshold set manually or learned from data. The threshold t can be any defined number or amount utilized to distinguish segments as background information or potential objects. or<img file="JP7472016B2_D0004.tif" />,and<img file="JP7472016B2_D0005.tif" /> 、<img file="JP7472016B2_D0006.tif" />is an element of S' and r<sub>i</sub>is the element R-, where R- is the set of non-salient regions (background) of the image, which can be computed and used as the similarity between each segment against a labeled database of labeled salient and non-salient segments. The final scores are as follows.
<img file="JP7472016B2_D0007.tif" />
In another embodiment, selection of portions of interest for past interactions of the same user may be determined. The final segment S' is then clustered to form one or more segments. Each segment is a distinctive part of the image.
Returning to FIG. 2, once the background segment is removed, the remaining objects in the image are determined, as at 208. The remaining objects in the image can be determined by calculating a score for each possible hypothesis of the object's location, for example using a sliding window approach. Each segment can be processed to determine potential matching objects using approaches such as Harr-like wavelet boosted selection or multi-part-based models. For example, a feature vector can be determined for the segment and compared to information stored for the object. Based on the feature vector and the stored information, a determination may be made about how similar the feature vector is to the stored feature vector for a particular object and/or a particular type of object.
The sliding window approach can be run N times, each using a different trained object classifier or label (e.g., person, bag, shoe, face, arm, hat, pants, top, etc.). After determining the hypotheses for each object classifier, the output is a set of optimal hypotheses for each object type. Position-dependent constraints can also be considered, since objects are usually not displayed randomly in an image (for example, eyes and noses are usually displayed together). For example, the location of the root object (e.g., a person) is defined as W(root), and each geometric constraint for each object k is a six-element vector<img file="JP7472016B2_D0008.tif" />are shown relative to each other as Root object W<sub>root</sub>for each object W<sub>oi</sub>The geometric "fit" of<img file="JP7472016B2_D0009.tif" />defined by
<img file="JP7472016B2_D0010.tif" />Here, dx, dy are object box W<sub>oi</sub>is the average geometric distance between each pixel of the root object box and each pixel of the root object box. optimal value<img file="JP7472016B2_D0011.tif" />The problem of finding arg minλ<sub>i</sub><img file="JP7472016B2_D0012.tif" />It can be formulated as Here, D<sub>train</sub>(Θ<sub>i</sub>) is Θ in training or other saved images.<sub>i</sub>is the observed value of
To optimize this functionality, the position of objects within the image can be determined. For example, the center of the root object (eg, a person) in the image is marked as (0,0), and the positions of other objects in the processed image are shifted with respect to the root object. Next, a linear support vector machine (SVM) is<sub>i</sub>is applied as a parameter. The input to the SVM is D<sub>train</sub>(Θ<sub>i</sub>). Other optimization techniques such as linear programming, dynamic programming, convex optimization, etc. may also be used alone or in combination with the optimizations described herein. Training data D<sub>train</sub>(Θ<sub>k</sub>) can be collected by having the user place a bounding box over both the entire object and the landmark. Alternatively, semi-automated approaches such as face detection algorithms, edge detection algorithms, etc. may be used to identify objects. In some implementations, other shapes may be used to represent objects, such as ellipses, ellipses, and/or irregular shapes.
Returning to FIG. 2, feature vectors and labels are generated and associated with each identified object, such as 210 and 212. Specifically, bounding boxes containing objects are associated with feature vectors and associations generated for labels and segments maintained in data store 1303 (FIG. 13). Additionally, the location and/or size of bounding boxes forming segments of the image may be associated and saved with the image. The size and/or position of the segment can be stored, for example, as pixel coordinates (x,y) corresponding to the edges or corners of the bounding box. As another example, the size and/or position of a segment may be stored as a column and/or row position and size.
A label may be a unique identifier (such as a keyword) that represents an object. Alternatively, the label may include classification information or object type. For example, a label associated with a representation of clothing may include an apparel classifier (such as a prefix classifier) in addition to the object's unique identifier. In yet other implementations, the label may indicate attributes of the object represented in the image. Attributes include, but are not limited to, the object's size, shape, color, texture, pattern, etc. Other implementations may determine a set of object attributes (e.g., color, shape, texture) for each object in the image and concatenate that set to form a single feature vector representing the object. . The feature vectors can then be transformed into visual labels using a visual vocabulary. A visual vocabulary can be generated by running a clustering algorithm (such as K-means) on features generated from a large dataset of images, with cluster centers resulting in a vocabulary set. Each single feature vector may be stored and/or translated into one or more lexical terms that are most similar in the feature space (eg, n).
After associating a label and a feature vector with each object represented in the image, the image segment corresponding to the object is indexed, as at 214. Each object can be indexed using standard text-based search techniques. However, unlike standard text or visual search, multiple indexes can be maintained in data store 1303 (Figure 13) and each object can be associated with one or more of the multiple indexes.
FIG. 3 is a representation of a segmented image that may be maintained in a data store, according to one embodiment. Images, such as image 300, can be segmented using the segmentation techniques described above. Using the example routine 200, the background segment was removed and six objects in the image were segmented and identified. Specifically, a body object 302, a head object 304, an upper object 306, a pants object 308, a bag object 310, and a shoe object 312. As part of the segmentation, the root object (body object 302 in this example) is determined and the positions of other objects 304-312 are considered when identifying those other objects. Once the object type is determined, a label or other identifier is generated and associated with the image segment and image.
In addition to indexing the segments, determining the objects, generating labels, and associating the segments and labels with the image 300, feature vectors representing each object in the image 300 are generated and stored in a data store, and the image 300, segment , and associated with the label. For example, a feature vector representing the size, shape, color, etc. of the wallet object can be generated and associated with image 300 and segment 310. Feature vectors representing other objects detected within the image may be similarly generated and associated with those objects, segments, and image 300.
In other implementations, the image may be segmented using other segmentation and identification techniques. For example, images can be segmented using crowdsourcing techniques. For example, when viewing an image, a user can select areas of the image that contain objects and label those objects. The more users identify objects in an image, the more reliable the identification of those objects becomes. Based on the segmentation and identification provided by the user, objects in the image can be indexed and associated with other visually similar objects in other images.
FIG. 4 is an example object matching process 400 in accordance with the described implementation. The example process 400 begins, such as 402, by receiving an image that includes a representation of one or more objects. Similar to other examples described herein, images can be received from any of a variety of sources.
Upon receiving the image, the image is processed using all or a portion of the image processing process 200 described above to determine objects of interest represented in the image, as indicated at 404. In some implementations, the entire image processing process 200 may be performed and then objects of interest may be determined from the detected objects as part of the example process 200. Other implementations run one or more object detection algorithms to determine latent objects in an image, then select one of the latent objects as an object of interest, and perform the example process 200 on that Can be performed on latent objects.
For example, run an edge detection or object detection algorithm to detect potential objects in an image and utilize the location of the potential object, the clarity or focus of the potential object, and/or other information. objects of interest can be detected. For example, in some implementations, an object of interest may be determined to be in focus and in the foreground of the image, toward the center of the image. In other implementations, the user may provide an indication or selection of a segment of the image that includes the object of interest.
Once an object of interest is determined, an image processing process 200 is performed on the object and/or a segment of the image containing the object to identify the object and generate an object feature vector representative of the object, , generate a label corresponding to the type of object.
The generated object feature vector and/or label is then compared, as at 408, with the stored feature vector corresponding to the object represented by the segment of the stored image, and the object feature vector and each stored Generate a similarity score between the calculated feature vectors. In some implementations, rather than comparing the object feature vector with all saved feature vectors, a label representing the type of object is used to reduce the saved feature vectors to have the same or similar labels. It can only contain things. For example, if the object of interest is determined to be a shoe, the object feature vector is only compared to the saved feature vector with the label shoe, thereby restricting comparisons to objects of the same type.
In other implementations, in addition to, or as an alternative to, comparing the object feature vector to saved feature vectors with the same or similar labels, the location of the saved image segment is determined by the location of the object of interest. expected to be placed in a particular segment of the image. For example, if the object of interest is determined to be a shoe, it is determined that the lower third segment of the saved image most likely contains the shoe object, and the feature vector comparison is saved. may be limited to a segment in the bottom third of the image. Alternatively, the position of the object of interest relative to the root object (e.g., a person) can be determined and utilized, and features corresponding to segments of the saved image based on their position relative to the root object, as described above. Vectors can be selected.
Comparing the object feature vector and the saved feature vector generates a similarity score indicating the similarity of the object feature vector and the saved feature vector being compared. Images associated with stored feature vectors with higher similarity scores are determined to be more sensitive to search and image matching than stored images associated with feature vectors with lower similarity scores. Ru. Because a saved image is associated with multiple saved feature vectors that can be compared to the object feature vector, some implementations base the similarity score determined for each associated saved feature vector , the average similarity score of the images is determined. In other implementations, the similarity score for an image with multiple saved feature vectors that is compared to the object feature vector is the median similarity score, the lowest similarity score, or the feature vectors associated with the saved images. may be other variations of the similarity score.
Based on the similarity score determined for each image, a ranked list of saved images is generated, such as 410. In some implementations, the ranked list may be based solely on similarity scores. Other implementations may include the popularity of the saved image, whether the user has previously viewed and/or interacted with the saved image, some saved feature vector associated with the saved image, object features, etc. Due to other factors, such as many feature vectors associated with the saved image compared to the vector, many saved feature vectors associated with the saved image and with labels identical or similar to the object of interest. One or more of the stored images can be weighted higher or lower based on the image size.
Finally, as at 412, a plurality of results for the stored images is returned, eg, to the user device, based on the ranked results list. In some implementations, the example process 400 is performed, in whole or in part, by a remote computing resource remote from the user device, such that the plurality of results of the images corresponding to the ranked results list is displayed on the user device. may be transmitted to the user device for presentation to the user device in response to transmitting the image of the object of interest. In other implementations, portions of the example process 400 may be executed on a user device, and portions of the example process 400 may be executed on a remote computing resource. For example, program instructions stored in memory of the user device may be executed to cause one or more processors on the user device to receive images of objects, determine objects of interest, and/or label or label objects of interest. Generation of object feature vectors representing objects can be performed. Object feature vectors and/or labels are sent from the user device to a remote computing resource, and code running on the remote computing resource sends one or more of the received object feature vectors to one or more processors of the remote computing resource. or more, generate a ranked result list, and respond to an input image containing the desired object by generating a ranked result list and sending the images corresponding to the ranked result list to the user device. to the user. In other implementations, different aspects of example process 400 may be performed by different computing systems at the same or different locations.
FIG. 5 is another example object matching process 500 according to the described implementation. The example process 500 begins, such as at 502, by receiving an image that includes a representation of one or more objects. Similar to other examples described herein, images can be received from any of a variety of sources.
Once the image is received, the image is processed using all or a portion of the image processing process 200 described above such that one or more objects of interest represented in the image are It will be judged. In some implementations, the entire image processing process 200 can be performed to determine candidate objects of interest from the detected objects as part of the example process 200. In other implementations, one or more object detection algorithms may be performed to determine candidate objects within the image.
For example, run an edge detection or object detection algorithm to detect objects in an image and use the location of potential objects, the clarity or focus of potential objects, and/or other information to identify objects of interest. Certain candidate objects can be detected. For example, in some implementations, candidate objects of interest are determined to be in focus, located in the foreground of the image, and/or located close to each other toward the center of the image. can be done.
A determination is then made, as at 506, as to whether there are multiple candidate objects of interest represented within the image. If it is determined that there are no candidate objects of interest, a single detected object is utilized as the object of interest, as at 507. If it is determined that there are multiple candidate objects of interest, the image is displayed along with an identifier indicating each of the candidate objects of interest so that the user can select one or more candidate objects as objects, such as 508. presented to the user. For example, an image may be presented on a touch-based display of a user device with a visual identifier placed adjacent to each candidate object. The user can then provide input by selecting one or more candidate objects as objects of interest. Next, as at 510, user input is received and utilized by the example process to determine objects of interest.
In some implementations, a user may be able to specify both objects of interest and objects of no interest, or objects that are given negative weights in determining images that match a search. For example, if multiple objects are detected in an image and presented to the user for selection, the user can choose between a positive selection indicating the object as an object of interest, a negative selection indicating the object as an object of no interest, or a search No selections are taken into account when determining which stored images match the .
Upon determining objects of interest, or if there is only one object of interest, the image processing process 200 performs operations on those objects and/or segments of the image containing the objects that identify the objects, such as at 512. It is executed to generate feature vectors that identify objects and create labels corresponding to each object type. In an example involving both objects of interest and objects of no interest, the example process 200 (Figure 2) includes the creation of object and feature vectors/labels for both objects of interest and objects of no interest. It can be performed for both types.
Each generated object feature vector and/or label is compared with the stored feature vector corresponding to the object represented by the segment of the stored image, such as 514, and each object feature vector and each stored Generate similarity scores between feature vectors. In some implementations, rather than comparing an object feature vector with all saved feature vectors, a label representing the object type is used to determine whether only the saved feature vectors are object features of the same or similar type. As the vectors are compared, the stored feature vectors that are compared with different object feature vectors can be reduced. For example, if one of the objects of interest is determined to be a shoe, the object feature vector for that object can only be compared to the stored feature vector with the label shoe. Similarly, if the second object of interest is determined to be a tops, then the object feature vector for that object can only be compared to the stored feature vector with the tops label.
In other implementations, in addition to, or as an alternative to, comparing the object feature vector to saved feature vectors with the same or similar labels, the location of the saved image segment is determined by the location of the object of interest. expected to be placed in a particular segment of the image. For example, if the object of interest is determined to be a shoe, it is determined that the lower third segment of the saved image most likely contains the shoe object, and the feature vector comparison is saved. may be limited to a segment in the bottom third of the image. Alternatively, the position of the object of interest relative to the root object (e.g., a person) can be determined and utilized, and features corresponding to segments of the saved image based on their position relative to the root object, as described above. Vectors can be selected.
Comparison of the object feature vectors and the saved feature vectors generates a similarity score indicating the similarity of each object feature vector to the saved feature vectors being compared. Images associated with stored feature vectors with higher similarity scores are determined to be more sensitive to search and image matching than stored images associated with feature vectors with lower similarity scores. Ru. Because a saved image may be associated with multiple saved feature vectors that can be compared to one or more object feature vectors, some implementations An average similarity score of the images is determined based on the similarity scores determined by the above steps. In other implementations, a similarity score for an image with multiple stored feature vectors that is compared to multiple object feature vectors may produce two similarity scores, one for each object feature vector. In examples involving similarity scores for objects of no interest, the similarity scores may similarly be determined by comparing the object feature vectors of no interest to the stored feature vectors.
Based on the similarity score determined for each image, a ranked list of saved images is generated, such as 516. In some implementations, the ranked list may be based solely on similarity scores. In implementations where multiple similarity scores are determined for different objects of interest, an image associated with a high similarity score of both objects of interest may be associated with a high similarity score of only one object of interest. The ranked list can be determined to be ranked higher than the images of . Similarly, if the user specifies an object of no interest, an image containing an object that is visually similar to the object of no interest will have a rank of interest and one or more saved feature vectors associated with the image. can be lowered. Depending on the implementation, other factors may be considered in ranking the saved images. For example, the popularity of the saved image, whether the user has previously viewed and/or interacted with the saved image, the number of feature vectors associated with the saved image, the object feature vectors associated with the saved image compared to one or more of the stored images based on a large number of stored feature vectors associated with the stored images, many stored feature vectors with the same or similar labels as one of the objects of interest, etc. Can be weighted higher or lower.
Finally, as at 518, a plurality of results for the stored images is returned, eg, to the user device, based on the ranked results list. In some implementations, the example process 500 is performed, in whole or in part, by a remote computing resource remote from the user device, such that the plurality of results of the images corresponding to the ranked results list is displayed on the user device. may be transmitted to the user device for presentation to the user device in response to the transmitting the image of the object of interest. In other implementations, portions of the example process 500 may be executed on a user device, and portions of the example process 500 may be executed on a remote computing resource. For example, program instructions stored in memory of the user device may be executed to cause one or more processors on the user device to receive images of objects, determine objects of interest, and/or label or label objects of interest. Generation of object feature vectors representing objects can be performed. Object feature vectors and/or labels are sent from the user device to a remote computing resource, and code running on the remote computing resource sends one or more of the received object feature vectors to one or more processors of the remote computing resource. or more, generate a ranked result list, and respond to an input image containing the desired object by generating a ranked result list and sending the images corresponding to the ranked result list to the user device. to the user. In other implementations, different aspects of example process 500 may be performed by different computing systems at the same or different locations.
FIG. 6A shows an input image 601 obtained by a user device 600 used to generate search results, according to the described implementation. Similar to the example above, the input image can be received or obtained from any source. In this example, the input image is captured by the camera of user device 600 and includes a representation of pineapple 602, a bottle of water 604-1, and a sheet of paper 604-2. In other implementations, the user can select image control 608 to select an image that is stored in the user device's memory or otherwise accessible to the user device. Alternatively, the user may select remote image control 606 to display/select an image from a plurality of images stored in memory remote from the user device.
In this example, in addition to processing an image to detect one or more objects of interest in the image, it is also possible to determine whether the object of interest corresponds to a defined category. . Defined categories include, but are not limited to, food, home decor, fashion, etc. A category may contain multiple different types of objects. For example, food may include thousands of food objects, such as pineapples.
If an object of interest is determined to correspond to a predefined category, multiple query types can be selected and utilized to generate results that are mixed to respond to the input image query. Different query types may include different types or styles of queries. For example, one query type may be a vision-based search that includes images that are visually similar to objects of interest, or image segments that are visually similar to objects of interest, as described above. Another query type is a text-based query that searches for and determines content that indicates how to use the object of interest or how to combine it with other objects of interest. For example, if the defined category is food, the first query type may return results that include images of food that are visually similar to the object of interest. The second query type may return results containing images of various food combinations, or recipes containing foods determined to be objects of interest.
In an example of multiple query types, the inputs used for each query type may be different. For example, a first query type that utilizes visual or image-based search can be configured to receive an object feature vector representing an object of interest, and that object feature vector may be a stored feature vector, as described above. Saved images containing objects that are visually similar to the object of interest can be detected by comparison with the vector. In contrast, the query type receives text/keyword input and stores stores that are not visually similar to the object of interest, but that contain labels that match the keyword or that are related to the object of interest. Can be configured to determine images.
In an example where one of the query types is configured to receive text/keyword input and search a data store of saved images, keywords or labels corresponding to the desired objects and/or categories are generated, respectively used to query stored images.
In some implementations, each query type can search content held in the same data source, but differences in the query type and how the stored content is queried can return different results. In other implementations, one or more of the query types may search different content held in the same data store or different data stores.
FIG. 6B shows the visual search results for objects of interest selected from FIG. 6A. Here, according to the described implementation, the results include images obtained from multiple query types related to the object of interest 602.
In this example, the object of interest, a pineapple, is food and is therefore determined to correspond to the defined category of food. Additionally, there are two different query types associated with food categories, one that is determined to be a visual or image-based search and one that is a text or keyword-based search.
In this example, the first query type generates an object feature vector representing a pineapple and compares the object feature vector to the stored feature vectors to find images containing objects that are visually similar to the object of interest 602. judge. The second query type generates a text query containing the keyword "pineapple + recipe" to search for images related to recipes that use pineapple. In some implementations, keywords may be determined based on objects and/or categories of interest. For example, based on image processing, it may be determined that the object of interest is a pineapple, and therefore one of the labels may be the object type of interest (eg, pineapple). Similarly, food categories may include or have labels associated with them, such as "recipe," which are used in creating text-based queries.
In other implementations, the keywords utilized by the text-based query may be based on labels associated with images determined from the image-based query. For example, if the first query type is an image-based search and returns images that are similar to, or contain similar image segments to, the object of interest, then the labels associated with those returned images are compared and the most frequently The labels used for are used as keywords in the second query type.
The results for each query type may be mixed and presented as a ranked list of images on user device 600. In this example, a first image 610-1 related to a recipe for making a pina colada is returned for the second query type, and a second image 610-2 is an object visually similar to the desired object. 602 is returned for the first query type containing (pineapple), and the two are displayed as a mixed result in response to the image input by the user.
In some implementations, the determined keywords or labels, such as keywords 611-1 through 611-N, may be presented on the user device and selectable by the user to further refine the query. Users can also add their own keywords by selecting add control 613 and entering additional keywords. Similarly, as explained below, in this example multiple objects are detected in the input image and indicators 604-1, 604-2 are also included to allow the user to specify different or additional objects of interest. Visible to other objects. As the user selects different or additional objects of interest, the search results are updated accordingly.
The user may interact with the displayed results returned to the user device to refine the search, provide additional or different keywords, select additional or different objects of interest, and/or perform other actions.
FIG. 7 is an example object category matching process 700 in accordance with the described implementation. The example process 700 begins, such as 702, by receiving an image that includes a representation of one or more objects. Similar to other examples described herein, images can be received from any of a variety of sources.
Upon receiving the image, the image is processed using all or a portion of the image processing process 200 (FIG. 2) described above to determine the one or more interests represented in the image, as shown at 704. An object with is determined. In some implementations, the entire image processing process 200 can be performed to determine candidate objects of interest from the detected objects as part of the example process 200. In other implementations, one or more object detection algorithms may be performed to determine candidate objects within the image.
For example, run an edge detection or object detection algorithm to detect objects in an image and use the location of potential objects, the clarity or focus of potential objects, and/or other information to identify objects of interest. Certain candidate objects can be detected. For example, in some implementations, candidate objects of interest are determined to be in focus, located in the foreground of the image, and/or located close to each other toward the center of the image. can be done. In some implementations, object detection scans only images for certain types of objects that correspond to one or more predefined categories. Defined categories include, but are not limited to, food, home decor, fashion, etc. In such an implementation, image processing simply processes the image to determine whether an object type associated with one of the defined categories is potentially represented in the image. As mentioned above, multiple types of objects can be associated with each category, and in some implementations, object types can be associated with multiple categories.
A determination is then made, as at 706, as to whether the object of interest corresponds to the defined category, or whether an object corresponding to the defined category has been identified in the image.
The objects of interest correspond to defined categories based on the type of object of interest that is determined when the objects of interest are identified (e.g., identified as part of example process 200). It can be determined as follows. In implementations where more than one object is determined to be of interest, some implementations may require that both objects of interest correspond to the same predefined category. Other implementations require that objects of interest be associated with only one predefined category.
If it is determined that the object of interest does not correspond to the defined category, the received image is compared with stored image information, as at 707. For example, a feature vector representing a received image rather than an object of interest may be generated and compared to a stored feature vector corresponding to a stored image. In other embodiments, segment feature vectors representing one or more objects identified in the received image may be generated and compared to the stored segment feature vectors, as discussed above with respect to FIG. . Next, stored images determined to be visually similar to the received image and/or segments of the received image are returned, such as 709.
If the object of interest is determined to correspond to a predefined category, a query type associated with the predefined category is determined, such as 708. As mentioned above, multiple query types can be associated with predefined categories and used to retrieve different types or styles of content depending on the search.
A determination is then made, as at 710, as to whether the one or more query types are text-based queries for retrieving content. If one of the query types is determined to be a text-based query, query keywords are determined based on objects of interest, categories, users, or other factors, such as 712. For example, as explained above, some implementations allow a visual or image-based query to be followed by a text-based query, which can be associated with content items/images that match the visual or image-based query. Keywords can be determined from labels. For example, the frequency of words in labels associated with images returned to an image-based query may be determined, and keywords may be selected as those words in the most frequent labels.
The keywords are then used to query labels and/or annotations associated with the saved content, and a ranked list of results is returned based on keyword matches, such as 714.
If none of the query types is determined to be a text-based query, or in addition to generating and sending a text query, the received images are also compared to the stored images, as in 715. Similar to block 709, the comparison may include a comparison of a feature vector representing the received image with a stored feature vector representing the stored image, and/or one or more corresponding to an object in the received image (e.g., an object of interest). The above segment feature vector may be compared with a feature vector of a saved segment. Comparison of segment feature vectors can be performed in a manner similar to that described above with respect to FIG. 4 to determine images that contain objects that are visually similar to the object of interest.
Next, a result ratio is determined, such as 716, that indicates the proportion or percentage of content returned by each query type that is included in the ranked results returned to the user. The ratio or percentage of results can be determined based on various factors, such as category, user preferences, objects of interest, quantity or quality of results returned from each query type, user location, and so on.
Based on the ratio or percentage of results, the ranked results for each query type are mixed to produce 718 mixed results. Finally, as at 720, the blended results are returned to the user device and presented to the user as responsive to the input image containing the object of interest.
In some implementations, the example process 700 is performed, in whole or in part, by a remote computing resource remote from the user device, such that the plurality of results of the images corresponding to the ranked list of results is displayed on the user device. may be transmitted to the user device for presentation to the user device in response to the transmitting the image of the object of interest. In other implementations, portions of the example process 700 may be executed on a user device, and portions of the example process 700 may be executed on a remote computing resource. For example, program instructions stored in memory of the user device may be executed to cause one or more processors on the user device to receive images of objects, determine objects of interest, and/or label or label objects of interest. Generation of object feature vectors representing objects can be performed. Object feature vectors and/or labels are sent from the user device to a remote computing resource, and code running on the remote computing resource sends one or more of the received object feature vectors to one or more processors of the remote computing resource. or more, generate a ranked result list, and respond to an input image containing the desired object by generating a ranked result list and sending the images corresponding to the ranked result list to the user device. to the user. In other implementations, different aspects of example process 700 may be performed by different computing systems at the same or different locations.
By providing mixed results, users can choose between images containing objects that are visually similar to the provided object of interest and images that are related to, but not necessarily visually similar to, the object of interest. It is possible to display both images that do not contain representations of objects. Users can search for information about the object of interest, combinations of the object of interest with other objects, recipes related to the object of interest, rather than other images of the object of interest, in defined categories. Such a mixture is beneficial because there are many
FIG. 8A shows a query on a user device with options to provide visual enhancements, according to the described implementation. In the illustrated example, the user has entered a text-based query 807 that includes the keyword "summer clothes." In this example, the search input begins with a text-based input, and it is determined whether the text-based input corresponds to a defined category, such as food, fashion, home decor, etc. If the text input is related to a defined category, the user is presented with a visual refinement option and the user provides an image containing the object of interest that is used to narrow down results matching the text-based query. can.
For example, because the text-based query 807 returns images 810-1, 810-2, 810-3~810-N that are determined to contain annotations, keywords, or labels that correspond to the text-based query "summer clothes" may be used for. In some implementations, other keywords or labels 811 may also be presented to the user to allow the user to further refine the query. In some implementations, visual refinement options 804 are presented if the input keyword is determined to correspond to a defined category.
In Figure 8B, selecting the visual refinement option activates the user device's camera and processes the images captured by the camera and/or the camera's field of view to determine the shape of the object represented in the captured image/field of view. Detected. For example, if the camera is pointed at sweater 802, the shape of the sweater may be detected and suggested object types 805 may be presented to the user to confirm the object type of interest to the user. Similarly, a shape overlay 803 may also be presented on the display 801 of the user device 800 to indicate the shape of the currently selected object type.
In this example, the determined object category is fashion, and the currently detected object type of object 802 in the field of view corresponds to object type "tops" 805-3. The user can select different object types by selecting different indicators, such as "skirt" 805-1, "dress" 805-2, "jacket" 805-N, etc. As will be appreciated, fewer, additional, and/or different object types or indicators may be displayed. For example, the user may be presented with options to choose from based on color, fabric, style, size, texture, pattern, etc.
Similarly, in some implementations, rather than utilizing an image from the user device's camera, the user selects the image control 808 and selects an image from the user device's memory or from images accessible to the user device. You can. Alternatively, the user may select remote image control 806 and select an image from a remote data store as input data.
As with other examples, when an image is input, the image is processed to determine objects of interest, generate labels corresponding to the objects of interest, and generate feature vectors representing the objects of interest. Ru. The labels and/or feature vectors can then be utilized to refine or re-rank the images determined to be responsive to the keyword search. For example, FIG. 8C shows search results for the query of FIG. 8A, with "summer clothes" 807 refined based on the visual input of FIG. 8B, indicated by top icon 821, according to the described implementation. As with other examples, the labels and/or object feature vectors generated for the object of interest are combined with the stored feature vectors corresponding to objects contained in the stored images determined to match the original query. The comparison is used to generate a similarity score. In this example, the object feature vector representing sweater 802 (FIG. 8B) is compared to the saved feature vectors corresponding to segments of the image determined to correspond to the text query. Next, as described above, the rank of the image is changed based on the similarity score determined from the comparison of the feature vectors. The re-ranked images are then sent to the user device and displayed on the user device's display in response to the input image. For example, saved images 820-1, 820-2, 820-3, and 820-4 are visually similar to objects of interest and are ranked at the top of the re-ranked list and and may be displayed on a display of the user device.
FIG. 9 is an example text and image matching process 900, according to the described implementation. The example process 900 begins, such as 902, upon receipt of a text-based query, such as the entry of one or more keywords into a search input box presented on a user device. The stored content is then queried, as at 904, to determine which images have associated labels or keywords that correspond to or match the text input of the query. Further, as at 906, it is determined whether the text query corresponds to a predefined category. For example, if you define a category and can include one or more keywords or labels, and your text-based input includes keywords or labels such as "outfits," then if your query input corresponds to the defined category, It will be judged. If it is determined that the query does not correspond to the defined category, the example process is complete, as at 908, and the user can interact with the results presented in response to the text-based query.
If the query is determined to correspond to the defined category, the user is presented with options to visually narrow down the search results, such as 910. The visual enhancement may be, for example, a graphical button or icon presented with the search results selected by the user to generate an image and/or launch a camera to select an existing image. . In some implementations, determining whether a query corresponds to a defined category may be omitted, and each instance of process 900 may present the user with options for visual refinement of the search results, such as at 910.
It is also determined, as at 912, whether images used to narrow down the results of the query have been received. If no images are received, the example process 900 completes as at 908. However, if an image is received, the image may be processed using all or a portion of the image processing process 200 (FIG. 2) described above to determine the objects of interest represented in the image, such as at 914. be done. In some implementations, the entire image processing process 200 may be performed and then objects of interest may be determined from the detected objects as part of the example process 200. Other implementations run one or more object detection algorithms to determine latent objects in an image, then select one of the latent objects as the object of interest, and perform the example process 200 on that Can be performed on latent objects.
For example, run an edge detection or object detection algorithm to detect potential objects in an image and utilize the location of the potential object, the clarity or focus of the potential object, and/or other information. objects of interest can be detected. For example, in some implementations, an object of interest may be determined to be in focus and in the foreground of the image, toward the center of the image. In other implementations, the user may provide an indication or selection of a segment of the image that includes the object of interest.
Once an object of interest is determined, an image processing process 200 is performed on the object and/or a segment of the image containing the object to identify the object, generate a feature vector representative of the object, and generate a feature vector such as 916. , generate a label corresponding to the type of object.
The generated object feature vector and/or label is then compared to the stored feature vector corresponding to the object in the stored image determined to match the text-based query, such as 918, and the object feature vector and/or label are Generate a similarity score between each saved feature vector.
As described above, the comparison of the object feature vector and the stored feature vector generates a similarity score that indicates the similarity between the object feature vector and the stored feature vector to which it is compared. Images associated with stored feature vectors having higher similarity scores are determined to be more responsive to visually sophisticated searches than stored images associated with feature vectors having lower similarity scores. Ru. Because a saved image is associated with multiple saved feature vectors that can be compared to the object feature vector, some implementations base the similarity score determined for each associated saved feature vector , the average similarity score of the images is determined. In other implementations, the similarity score for an image with multiple saved feature vectors that is compared to the object feature vector is the median similarity score, the lowest similarity score, or the feature vectors associated with the saved images. may be other variations of the similarity score.
Based on the similarity score determined for each image, the results of the text-based query are re-ranked into an updated ranked list, such as 920. In some implementations, the ranked list may be based solely on similarity scores. Other implementations may include the popularity of the saved image, whether the user has previously viewed and/or interacted with the saved image, some saved feature vector associated with the saved image, object features, etc. Due to other factors, such as many feature vectors associated with the saved image compared to the vector, many saved feature vectors associated with the saved image and with labels identical or similar to the object of interest. One or more of the stored images can be weighted higher or lower based on the image size.
Finally, the image with the highest rank in the ranked list is returned to the user device for presentation, such as 922. In some implementations, the example process 900 is performed, in whole or in part, by a remote computing resource remote from the user device, such that the plurality of results of the images corresponding to the ranked results list is displayed on the user device. may be transmitted to the user device for presentation to the user device in response to the transmitting the image of the object of interest. In other implementations, a portion of the example process 900 may be executed on a user device and a portion of the example process 900 may be executed on a remote computing resource. For example, program instructions stored in memory of the user device may be executed to cause one or more processors on the user device to receive images of objects, determine objects of interest, and/or label or label objects of interest. Generation of object feature vectors representing objects can be performed. Object feature vectors and/or labels are sent from the user device to a remote computing resource, and code running on the remote computing resource sends one or more of the received object feature vectors to one or more processors of the remote computing resource. or more, generate a ranked result list, and respond to an input image containing the desired object by generating a ranked result list and sending the images corresponding to the ranked result list to the user device. to the user. In other implementations, different aspects of the example process 900 may be performed by different computing systems at the same or different locations.
FIG. 10A shows yet another exemplary visual refinement input of a query in accordance with the described implementation. In this example, the user has entered the text-based query "salmon recipe" 1007. It is determined that the query corresponds to a predefined category (such as a recipe) and that the user provides a visual enhancement. In this example, streaming video of a camera's field of view on user device 1000 is processed in real time or near real time to detect objects within the camera's field of view. In this example, the field of view in the streaming video is inside the refrigerator. In other examples, streaming video may include other regions. Processing may be performed on user device 1000 by computing resources remote from the user device, or a combination thereof.
When an object in a streaming video is detected, for example using an edge detection algorithm and/or some or all of the example process 200 (Figure 2), a keyword or label indicating the type of detected object is added to the streaming video. It is presented on the display 1001 of the device at the same time as the presentation.
In this example, strawberries, avocados, and eggs have been detected as candidate objects of interest within the field of view of the user device's camera. When an object is detected, a label 1002 is visually displayed adjacent to the object to indicate that the object has been detected.
In some implementations, a corpus of potential objects can be text-queried to detect candidate objects of interest and speed the process to improve the user experience by identifying only candidate objects of interest that correspond to keyword queries. Objects that match the corpus are identified as candidate objects. For example, a text query can be processed to determine that a user is looking for a recipe that includes salmon. Based on that information, a corpus of potential objects included or referenced in images associated with recipes that also include salmon is determined, and only objects matching that corpus are identified as candidate objects of interest.
In this example, the candidate objects detected in the field of view of the user device's camera are identified by the identifiers "strawberry" 1002-2, "egg" 1002-1, and "avocado" 1002-3. As the user moves the camera's field of view, the position of the identifier 1002 is updated to correspond to the relative position of the detected object, and if additional candidate objects enter the field of view and are included in the streaming video, the identifiers of those objects are updated. will be presented as well.
The user can select one of the identifiers to indicate that the object is of interest. In FIG. 10B, the user has selected the object egg as the object of interest. Accordingly, the keyword egg is added to the query "salmon recipe" 1001-1, as indicated by the egg icon 1001-2, and images such as images 1010-1, 1010-2, 1010-3, and 1010-N is determined to contain or be associated with the labels/keywords "salmon," "recipe," and "egg" and is returned to the user for display in response to the query. In some implementations, other keywords 1011 may be presented to the user device 1000 as well for further refinement of the query results.
Provides the ability for users to utilize visual search and/or a combination of visual search and text-based search to generate results based on predefined categories determined from the input and/or objects detected in the input This allows users to input the type of content they want to explore, improving the quality of results through better inference. The increased flexibility provided by the described implementation provides images that are visually similar to the input image by focusing visual search (e.g., feature vectors) on segments or parts of the stored image rather than the entire image. Technical improvements are provided for visual search only and/or automatically add different forms of search (such as keywords) to visual search. Additionally, by adding visual refinements to text-based queries that utilize either feature vectors or visual matching via keyword matching, or both, users can express input parameters in different contexts (keyword, visual). This allows you to more appropriately determine and search for the desired information.
FIG. 11 illustrates an example user device 1100 that may be used in accordance with various implementations described herein. In this example, user device 1100 includes a display 1102 and optionally at least one input component 1104, such as a camera, on the same and/or opposite side of the device as display 1102. User equipment 1100 may also include an audio transducer, such as a speaker 1106, and optionally a microphone 1108. In general, user device 1100 may have any type of input/output component that allows a user to interact with user device 1100. For example, various input components to enable user interaction with the device may include a touch-based display 1102 (resistive, capacitive, etc.), a camera, a microphone, a Global Positioning System (GPS), a compass, or the like. Includes any combination of One or more of these input components may be included in or in communication with the device. Various other input components and combinations of input components may also be used within various implementations, as should be apparent in light of the teachings and suggestions contained herein.
To provide the various functionality described herein, FIG. 12 illustrates an example set of basic components 1200 of a user device 1100, such as user device 1100 described with respect to FIG. 11. In this example, the device includes at least one central processing unit 1202 for executing instructions that may be stored in at least one memory device or element 1204. As will be apparent to those skilled in the art, the apparatus may include many types of memory, data storage, or computer-readable storage media, such as first data storage for program instructions for execution by processor 1202. can. Removable storage memory can be used to share information with other devices, etc. Typically, the device includes some type of display 1206, such as a touch-based display, electronic ink (e-ink), organic light emitting diode (OLED), or liquid crystal display (LCD).
As described, the device in many implementations includes at least one imaging device 1208, such as one or more cameras that can image objects in the vicinity of the device. The imager may include, or be based at least in part, on any suitable technology, such as a CCD or CMOS imager having a determined resolution, focal range, visible area, and capture rate. The apparatus may include at least one search component 1210 for performing the process of generating search terms, labels, and/or identifying and presenting results matching selected search terms. For example, a user device may constantly or intermittently communicate with a remote computing resource and exchange information such as selected search terms, images, labels, etc. with the remote computing system as part of a search process.
The device may also include at least one location component 1212, such as GPS, NFC location tracking, Wi-Fi location monitoring, etc. The location information obtained by location component 1212 may be used with various implementations described herein as a factor in selecting images that match objects of interest. For example, if a user is in San Francisco and actively selects a bridge (object) shown in an image, the user's location will be a factor in identifying visually similar objects such as the Golden Gate Bridge. considered to be.
The example user device may also include at least one additional input device that can receive conventional input from a user. This traditional input includes, for example, push buttons, touch pads, touch-based displays, wheels, joysticks, keyboards, mice, trackballs, keypads, or other devices or elements that allow a user to enter commands into the device. It will be done. In some implementations, these I/O devices may also be connected wirelessly, infrared, Bluetooth, or other links.
FIG. 13 is a pictorial diagram of an example implementation of a server system 1300, such as a remote computing resource, that can be used with one or more of the implementations described herein. Server system 1300 may include a processor 1301, such as one or more redundant processors, a video display adapter 1302, a disk drive 1304, an input/output interface 1306, a network interface 1308, and memory 1312. Processor 1301, video display adapter 1302, disk drive 1304, input/output interface 1306, network interface 1308, and memory 1312 may be communicatively coupled to each other by communication bus 1310.
Video display adapter 1302 provides display signals to a local display that allow an operator of server system 1300 to monitor and configure the operation of server system 1300. Input/output interface 1306 similarly communicates with external input/output devices such as a mouse, keyboard, scanner, or other input/output devices that can be operated by an operator of server system 1300. Network interface 1308 includes hardware, software, or any combination thereof for communicating with other computing devices. For example, network interface 1308 may be configured to provide communication between server system 1300 and other computing devices, such as user device 1100.
Memory 1312 typically includes random access memory (RAM), read only memory (ROM), flash memory, and/or other volatile or permanent memory. Memory 1312 is shown storing an operating system 1314 for controlling the operation of server system 1300. Also stored in memory 1312 is a binary input/output system (BIOS) 1316 for controlling low-level operations of server system 1300.
Memory 1312 further stores program codes and data for providing network services that allow user device 1100 and external sources to exchange information and data files with server system 1300. Accordingly, memory 1312 may store browser application 1318. Browser application 1318 includes computer-executable instructions that, when executed by processor 1301, generate or obtain a configurable markup document, such as a web page. Browser application 1318 communicates with data store manager application 1320 to facilitate data exchange and mapping between data store 1303, user devices such as user device 1100, external sources, and the like.
As used herein, the term "data store" refers to any device or combination of devices capable of storing, accessing, or retrieving data, including any combination and number of data servers, databases, data storage devices, etc. , and data storage media can be included in standard, distributed, or clustered environments. Server system 1300 includes appropriate hardware and software to integrate with data store 1303 as necessary to execute aspects of one or more applications of user equipment 1100, external sources and/or search services 1305. can include. Server system 1300 works with data store 1303 to provide access control services and generate content such as matching search results, images containing visually similar objects, and indexes of images containing visually similar objects. can.
Data store 1303 may include a number of separate data tables, databases, or other data storage mechanisms and media for storing data related to particular aspects. For example, the illustrated data store 1303 includes digital items (eg, images) and corresponding metadata about those items (eg, labels, indexes). Search history, user settings, profiles, and other information can be stored in the data store as well.
It should be appreciated that there may be many other aspects that may be stored in the data store 1303, which may be stored in any of the mechanisms listed above, or additional mechanisms of the data store, as appropriate. Data store 1303, through logic associated therewith, may be operable to receive instructions from server system 1300 and retrieve, update, or otherwise process data in response.
Memory 1312 may also include search service 1305. Search service 1305 may be executable by processor 1301 to implement one or more of the functions of server system 1300. In one implementation, search service 1305 may represent instructions embedded in one or more software programs stored in memory 1312. In another implementation, search service 1305 may represent hardware, software instructions, or a combination thereof. Search service 1305 may perform some or all of the implementations described herein alone or in combination with other devices, such as user device 1100.
Server system 1300, in one implementation, is a distributed environment that utilizes multiple computer systems and components interconnected via communication links using one or more computer networks or direct connections. However, those skilled in the art will appreciate that such a system can operate equally well in systems having fewer or more components than shown in FIG. 13. Accordingly, the depiction of FIG. 13 should be construed as illustrative in nature and not limiting the scope of the present disclosure.
Implementations disclosed herein may include a computing system having one or more processors and memory for storing program instructions. The program instructions, when executed by the one or more processors, cause the one or more processors to receive a text query from at least a user device, and to determine and return a plurality of results corresponding to the text query. can. Following receipt of the text query, program instructions, when executed on the one or more processors, cause the one or more processors to receive an image of an object from a user device and generate an object feature vector representing the object. generating and comparing a plurality of stored feature vectors corresponding to segments of a plurality of resultant images, an object feature vector based at least in part on a comparison of the object feature vector and the plurality of stored images; Generating a ranked list of results and presenting the ranked list in response to receipt of images.
Each of the plurality of saved image segments may correspond to less than an entire image. The image is at least one of an image generated by a camera of the user device, an image obtained from a memory of the user device, an image obtained from a plurality of results, or an image obtained from a storage medium remote from the user device. There may be. The program instructions further cause the one or more processors to at least determine that the text query corresponds to a predefined category and provide image enhancement options in response to the determination that the text query corresponds to a predefined category. Good too. The defined categories may be at least one of fashion, clothing, home decor, personal, or food.
Implementations disclosed herein may include computer-implemented methods. A computer-implemented method includes: receiving a query from a user device; determining a first plurality of images based at least in part on the query; receiving an image of an object after receiving the query; at least one image segment of each of the first plurality of images, determining a ranked list of at least some of the first plurality of images based at least in part on the comparison; providing a presentation of at least a portion of the first plurality of images in the attached list.
Optionally, the computer-implemented method also includes processing the images to determine an object type of an object represented in the images, and comparing the images includes generating an object feature vector representative of the object, and generating an object feature vector representing the object. and comparing with stored feature vectors corresponding to objects represented in the first plurality of images having the same object type. The computer-implemented method may also include determining that the query corresponds to a defined category and presenting visual refinement options. The computer-implemented method may also include detecting an object within a field of view of a camera of the user device, determining an object type corresponding to the object at the user device, and displaying the object type identifier on a display of the user device. The object type identifier may include at least one of a graphical representation corresponding to the shape of the object type, or a name of the object type. The computer-implemented method may also include enabling selection of a second object type identifier, thereby indicating a second object type corresponding to the object. The computer-implemented method may also include generating keywords corresponding to the object type and including the keywords as part of the query. A computer-implemented method determines a plurality of keywords corresponding to at least one of a portion of a first plurality of images, a first plurality of images, or a query, and presents each of the plurality of keywords to a user device by a user. and providing a selection. Comparing the images of the object includes: generating an object feature vector representative of the object; determining an expected position of the object in the first plurality of images; determining an image segment based at least in part; The method may further include one or more of comparing the object feature vector to a stored feature vector associated with the image segment. The computer-implemented method may further include determining an object type of the object, and determining the expected position is based at least in part on the object type.
Implementations disclosed herein may include a non-transitory computer-readable storage medium that stores instructions that may be executed by at least one processor of a computing system. The instructions, when executed by the at least one processor, cause the computing system to receive a query at at least a user device, determine that the query corresponds to a defined category, and enable visual enhancement options for the user device; Streaming video from a user device as part of the receiving visual enhancement option includes processing at least a portion of the streaming video to create one or more images represented by the streaming video present on the display of the user device. Identifies the object type of an object, receives an object type of one or more objects, an object type selection, and determines multiple saved images corresponding to both the query and the selected object type, simultaneously with streaming video presentation and display the plurality of stored images on a display of the user device.
The defined category may be food, and the streaming video may include a representation of the food that is currently within the field of view of the user device's camera. The query may be a text-based query. The instructions may further cause the computing system to at least determine candidate images to be considered in determining the plurality of stored images based at least in part on the defined categories. The instructions causing the at least one processor to determine a plurality of stored images cause the computing system to at least determine at least one keyword corresponding to a query, at least one object type, or a defined category, and the at least one keyword The plurality of stored images may be determined based at least in part on the stored images.
Implementations disclosed herein may include a computing system. The computing system can include an image data storage device capable of storing one or more of a plurality of image segments and/or a first plurality of stored images having image information corresponding to each image. The image information may indicate, for each image, one or more of the respective plurality of image segments, each image segment being smaller than the entire stored image and the plurality of stored feature vectors, respectively. , and each stored feature vector corresponds to an object represented by an image segment of the plurality of image segments. A computing system may also include one or more processors and memory for storing program instructions. When program instructions are executed by one or more processors, the one or more processors receive an image from a user device, process the image, and render it into an image as part of a visual-based search. of interest, generate an object feature vector representing the object of interest, and compare the object feature vector with a plurality of saved feature vectors to generate a second plurality of saved feature vectors containing visually similar representations of the object. Determining a feature vector of the object of interest indicative of a rank of the second plurality of images of the first plurality of stored images based at least in part on a comparison of the object feature vector and the plurality of stored feature vectors. Determine the attached list. The image includes at least one image segment that includes a representation of an object that is determined to be visually similar to the object of interest. each image of the second plurality of images is sent to a display of the user device such that the entire image of each image of the plurality of images of the second plurality of images is included.
Program instructions for processing an image to determine an object of interest cause one or more processors, at runtime, to process the image to determine a first candidate object of interest and a second candidate object of interest. An input may be received indicating a first candidate object of interest as the object of interest. An image may contain representations of multiple objects, where the object of interest includes the location of the object of interest within the image, the part of the image that is in focus, and the size of the object of interest represented within the image. , or based at least in part on the color of the object of interest compared to the background color. Optionally, the program instructions further cause the one or more processors to process at least the image to determine a second object represented in the image and to generate a second object feature vector representing the second object. generate a second object feature vector, among the stored feature vectors corresponding to each of the plurality of image segments, a third plurality of stored feature vectors that include representations of objects that are visually similar to the second object; and the ranked list further includes an object feature vector having the plurality of saved feature vectors determined based at least in part on the comparison and a second of the first plurality of saved images. comparing the second object feature vector with the plurality of stored feature vectors to identify the plurality of images of the object, wherein at least one image segment containing a representation of the visual object is similar to the object of interest; and further includes at least one second image segment that includes a representation of an object that is visually similar to the second object. The program instructions, when executed, may cause the one or more processors to receive at least a selection of an object of interest and a second object from at least a user device.
Implementations disclosed herein may include computer-implemented methods. A computer-implemented method includes: receiving instructions of an image from a user device; processing the image to determine a first object represented in the image; generating an object feature vector representing the first object; may include one or more of comparing the feature vector to a plurality of stored feature vectors, each of the plurality of stored feature vectors representing a respective image segment of the first plurality of images; The image segment generates a ranked list of a second plurality of images from the first plurality of images, each image having less than all of the images, and each image of the second plurality of images has at least a portion of the comparison. A plurality of images are presented by the user device, including at least one respective image segment determined to include a representation of the object based on the first object and visually similar to the first object.
Optionally, the computer-implemented method includes, prior to receiving the image instructions, segmenting the second image into a plurality of segments; for each of the plurality of segments, each feature vector corresponds to a segment; and storing the second image and each respective feature vector in a data store, the respective feature vector being included in the plurality of stored feature vectors. The computer-implemented method may further include one or more of storing, for each of the plurality of segments, position information indicative of the position of the respective segment within the second image. The computer-implemented method performs, for each of a plurality of image segments, one or more of the following: determining an object represented in the image segment, generating a label corresponding to the object, and that a feature vector represents the object. May include. Optionally, the label may indicate at least one of the type of object or category of object. Optionally, the computer-implemented method includes determining a label of a first object represented in the image, and determining a plurality of stored feature vectors based at least in part on the label of the first object. It may further include one or more. Optionally, processing the image may include one or more of processing the image to determine a plurality of candidate objects, receiving a selection of a first candidate object, and the first candidate object being a first object. Optionally, the computer-implemented method includes receiving a selection of a second candidate object, generating a second object feature vector representing the second candidate object, and comparing the second object feature vector with at least some of the plurality. may further include one or more of the following: Generating the ranked list of the second plurality of images from the first plurality of images includes, at least in part, comparing the second object feature vector with at least a portion of the plurality of stored feature vectors. It is based on Optionally, the computer-implemented method further comprises one or more of: determining an object type of the first object; and determining a plurality of stored feature vectors based at least in part on the object type. may be included. Optionally, the plurality of saved feature vectors may have the same object type as the object type of the first object.
Implementations disclosed herein may include a non-transitory computer-readable storage medium that stores instructions. The instructions, when executed by at least one processor of the computing system, may cause the computing system to maintain image information corresponding to a plurality of images in a data store. The image information may indicate, for each image, one or more of the respective plurality of image segments, where each image segment represents a portion of the respective image, the respective plurality of feature vectors, and each The feature vectors represent objects within each image segment and multiple labels corresponding to each image segment. The instructions further cause the computing system to determine an object represented in the image, generate an object feature vector representing the object, determine a label for the object, and generate a plurality of feature vectors, an object feature vector based at least in part on the label. comparing the feature vector with each of the plurality of feature vectors to determine a similarity score, each similarity score representing the similarity of the object feature vector to each of the plurality of feature vectors; generating a ranked list of saved images based at least in part;
Optionally, the label may indicate at least one of the object or the object's object type. Optionally, each of the plurality of feature vectors may represent an image segment of the respective stored image that is smaller than the entire image, and the instructions further present to the computing system at least the images indicated in the stored ranked list. Each presented image includes a respective image segment that is smaller than the entire image. Optionally, each saved feature vector may represent an object of an image segment that is smaller than the entire image. Optionally, the similarity score may represent the Euclidean distance between the feature vector and the saved feature vector.
Implementations disclosed herein may include a computing system. A computing system may include one or more processors and memory for storing program instructions. Program instructions, when executed by the one or more processors, cause the one or more processors to receive an image of an object from at least a user device and process the image to determine an object represented in the image. , the object corresponds to a defined category, determining a first query type and a second query type based at least in part on the object or the defined category, and generating keywords corresponding to the object for use in the first query type. , a first query type based at least in part on the keyword, generating a feature vector representing objects for use in the second query type, determining a second result for the second query type based at least in part on the feature vector. , a mixture of the first result, the second result produces a mixed result that contains the first proportion of the first result and the second proportion of the second result, and/or image.
Optionally, the program instructions may further cause the one or more processors to determine a result ratio of the mixing result indicating the first ratio and the second ratio. Optionally, the result ratio may be based at least in part on the object, the defined category, the first query type, the second query type, the user who submitted the image of the object, the user device, or user settings. Optionally, keywords may be generated based at least in part on labels associated with objects, defined categories, or images included in the second result. Optionally, the first result may include content corresponding to an item that utilizes or includes the object, and the second result may include content that includes a representation of the object.
Implementations disclosed herein may include computer-implemented methods. A computer-implemented method includes receiving an image of an object from a user device, determining that the object corresponds to a defined category, and determining a first query type and a second query type based at least in part. may include one or more of the following: a defined category or object, obtaining a first query result of a first query type corresponding to the object, obtaining a second query result of a second query type corresponding to the object, at least a first portion of the first query result and at least A second portion of the second query results generates a blended result and transmits the blended result for presentation by a user device in accordance with the image of the object.
Optionally, the defined category may be at least one of food, home decor, or fashion. Optionally, the computer-implemented method can further include one or more processing the image to determine an object type of an object represented in the image, the object being at least partially conformed to the object type. It is determined that the category corresponds to the category defined based on the category. Optionally, the first query type is a text-based query and the second query type is an image-based query. Optionally, the computer-implemented method includes determining a third query type associated with the defined category, obtaining results of the third query corresponding to the object, and at least one object identifier matching the third query type. further including. Optionally, the first query type returns content related to the object and the second query type returns content containing representations of objects of the same object type as the object. Optionally, the computer-implemented method can include one or more generating keywords based at least in part on the object or the object type of the object, and retrieving the first query result includes generating a keyword based at least in part on the object or the object type of the object. determining the first query result based on the keyword. Optionally, the computer-implemented method further includes one or more determining a result ratio indicating a first proportion of the first query results and a second proportion of the second query results to include in the mixed result. and the blending is based at least in part on the resulting ratios. Optionally, the computer-implemented method may further include one or more of determining an object type of the object and determining a plurality of stored feature vectors based at least in part on the object type. Optionally, obtaining the first query result may be based at least in part on keywords determined for the object, the obtaining of the second query result being based at least in part on keywords that are textual representations of the object; The feature vector may be based at least in part on a feature vector generated from the object, the feature vector visually representing the object.
Implementations disclosed herein may include a non-transitory computer-readable storage medium that stores instructions. The instructions, when executed by at least one processor of the computing system, cause the computing system to determine at least an object represented in the image, to determine that the object corresponds to a defined category, and to determine a first query type and a first query type. a type associated with a defined category upon which a second query may be determined; generating keywords corresponding to objects for use in the first query type; generating a feature vector representing the objects; based at least in part on the keywords; obtaining a first result of a first query type, obtaining a result of a second query type based at least in part on a second obtained feature vector, and obtaining at least one result from the first result and a second result. Send an instruction to present at least one result from the results.
Optionally, the instructions may further cause the computing system to segment the image into at least a plurality of image segments and process each of the plurality of image segments to determine an object represented in the image. Optionally, the first result may be obtained based at least in part on a comparison of the keyword and a label associated with the saved image. Optionally, the second result may be obtained based at least in part on a comparison of the feature vector and a stored feature vector representative of an object represented in the stored image. Optionally, the first result includes content that describes the object and the second result includes content that is visually similar to the object.
The concepts disclosed herein may be applied within a number of different devices and computer systems, including, for example, general purpose computing systems and distributed computing environments.
The above aspects of the disclosure are intended to be exemplary. They have been selected to illustrate the principles and application of disclosure and are not intended to be exhaustive or to limit disclosure. Many modifications and variations of the disclosed embodiments may be apparent to those skilled in the art. Those skilled in the art will recognize that the components and process steps described herein can be interchanged with other components or steps, or combinations of components or steps, and still achieve the benefits and advantages of this disclosure. It should be. Furthermore, it will be apparent to those skilled in the art that the present disclosure may be practiced without some or all of the specific details and steps disclosed herein.
Aspects of the disclosed system may be implemented as a computer method or as an article of manufacture such as a memory device or non-transitory computer-readable storage medium. A computer-readable storage medium may be readable by a computer and may include instructions for causing a computer or other device to perform the processes described in this disclosure. Computer-readable storage media may be implemented by volatile computer memory, non-volatile computer memory, hard drives, solid state memory, flash drives, removable disks, and/or other media. Additionally, one or more modules and engine components can be implemented in firmware or hardware.
Unless otherwise specified, letters such as "a" and "an" should generally be construed to include one or more of the described items. Accordingly, phrases such as "device configured for" are intended to include one or more of the listed devices. One or more such enumerated devices may also be collectively configured to perform the described enumeration. For example, "a processor configured to run enumerations A, B, and C" includes running enumeration A working in conjunction with a second processor configured to run enumerations B and C. A first processor may be included.
As used herein, language such as "about," "approximately," "generally," "approximately," "similar," or "substantially" refers to Represents a value, quantity, or characteristic that performs a desired function or achieves a desired result. For example, the terms "about," "approximately," "generally," "approximately," "similar," or "substantially" mean less than 10%, less than 5%, less than 1%, less than 1% of the stated amount; May represent an amount of less than 0.1%, less than 0.01%.
Although the subject matter has been described in language specific to structural features and/or methodological acts, it is understood that the subject matter as defined in the appended claims is not necessarily limited to the particular features or acts described. I want to be understood. Rather, the specific features and acts are disclosed as example forms of implementing the claims.
30 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20100205202A1 | Cites | United States of America |
| JP2015143951A | Cites | Japan |
| JP2017504861A | Cites | Japan |
12 members in 4 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 15713567 | United States of America | – | |
| 201715713567 | United States of America | A | |
| 2018051823 | United States of America | W |
Members12
| Document | Office | Kind | |
|---|---|---|---|
| US2019095467A1 | United States of America | A1 | |
| WO2019060464A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP3685278A1 | European Patent Office (EPO) | A1 | |
| JP2020534597A | Japan | A | |
| US10942966B2 | United States of America | B2 | |
| US2021256054A1 | United States of America | A1 | |
| US11620331B2 | United States of America | B2 | |
| US2023252072A1 | United States of America | A1 | |
| JP2023134806A | Japan | A | |
| JP7472016B2This record | Japan | B2 | |
| US12174884B2 | United States of America | B2 | |
| JP7670764B2 | Japan | B2 |
12 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Re-examination (zenchi) completed and case transferred to appeal boardAppealJAPANESE INTERMEDIATE CODE: A912A912 | A912 | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 7472016
- Application
- 2020514552
Titles2
- Japanese
- テキストおよび画像ベースの検索
- English
- Text and image-based search
Classification
- CPC, 7
- G06F16/5866
- G06F16/5838
- G06F16/583
- G06F16/738
- G06F16/24578
- G06F16/248
- G06F16/56
- IPC, 4
- G06F16 532
- G06F16 538
- G06F16 58
- G06T7 00
