System for displaying information in vision field i.e. public internet site, to user, has determination unit determining information based on acquired sound sequence, and superimposition unit superimposing representation of information
Abstract
The system has an acquisition unit for acquiring a sound sequence, and a determination unit for determining information according to acquired sound sequence. A superimposition unit superimposes representation of the determined information on an image corresponding to a vision field, where the information is related to a person. The determination unit comprises a speech recognition unit to associate the person with the acquired sound sequence, where the information is identity of the person. Independent claims are also included for the following: (1) a method for displaying information in a vision field (2) a computer program comprising a set of instructions for implementing a method for displaying information in a vision field.

Term
3.8 yearsto projected expiry
Projected expiry 30 June 2030, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
14 claims: 3 independent, 11 dependent
- 1CLAIMS REVENDICATIONS 1. System for displaying information in a field of vision, characterized in that it comprises:1. Système d'affichage d'informations dans un champ de vision, caractérisé en ce qu'il comprend : - Acquisition means (MIC, PROC) of a sound sequence;- des moyens d'acquisition (MIC, PROC) d'une séquence sonore ;- Means for determining (PROC, SERV) information as a function of the acquired sound sequence;- des moyens de détermination (PROC, SERV) d'une information en fonction de la séquence sonore acquise ;- Superimposition means (PROC, VIS) of a representation of the determined information on an image corresponding to the field of vision. - des moyens de superposition (PROC, VIS) d'une représentation de l'information déterminée sur une image correspondant au champ de vision.
- 9Device for displaying information in a field of vision, characterized in that it comprises:9. Dispositif d'affichage d'informations dans un champ de vision, caractérisé en ce qu'il comprend : - Acquisition means (MIC, PROC) of a sound sequence;- des moyens d'acquisition (MIC, PROC) d'une séquence sonore ;- Means for determining (PROC) information as a function of the acquired sound sequence;- des moyens de détermination (PROC) d'une information en fonction de la séquence sonore acquise ;- Superimposition means (PROC, VIS) of a representation of the determined information on an image corresponding to the field of vision. - des moyens de superposition (PROC, VIS) d'une représentation de l'information déterminée sur une image correspondant au champ de vision.
- 13Method for displaying information in a field of view, characterized in that it comprises the following steps:13. Procédé d'affichage d'informations dans un champ de vision, caractérisé en ce qu'il comprend les étapes suivantes : - acquisition (E2) d'une séquence sonore ;- acquisition (E2) of a sound sequence;- determination (E4, E6, E8, E10;E4, E14, E16) of information as a function of the acquired sound sequence;- détermination (E4, E6, E8, E10;E4, E14, E16) d'une information en fonction de la séquence sonore acquise ;- superposition (E12, E18) d'une représentation de l'information déterminée sur une image correspondant au champ de vision. - superposition (E12, E18) of a representation of the determined information on an image corresponding to the field of vision.
Independent claims3
55 paragraphs, as filed
The invention relates to a device and a method for displaying information in a field of view.
We know under the name of augmented reality the idea of superimposing on a real environment, generally corresponding to the field of vision of a user, additional information (for example images, symbols or characters), generally qualified as virtual because they are produced by a computer system, precisely in order to enrich what the user sees.
Patent application FR 2 876 820 describes for example such a system in which a correlation is sought between images captured in the real environment and images from a database in order to provide a virtual information element to About the captured images.
In this context, the invention proposes a system for displaying information in a field of view, characterized in that it comprises means for acquiring a sound sequence, means for determining information as a function of the acquired sound sequence and means for superimposing a representation of the determined information on an image corresponding to the field of vision.
The field of vision is thus enriched by means of information determined on the basis of the sound environment of the user.
The information relates, for example, to a person and the determination means may comprise voice recognition means capable of associating said person with the acquired sound sequence.
This application of the system proposed above is particularly interesting and can tend to the emergence of a community as explained below.
The information is for example in this case the identity of said person.
According to another conceivable embodiment, the determination means may comprise voice recognition means capable of identifying at least one word of the acquired sound sequence. The determining means can then determine the information by reading from a database on the basis of the identified word.
The display of contextual information relating to the speech of the interlocutor is thus obtained, which makes it possible in particular to enrich the understanding of the user.
The superposition means can for example in practice superimpose said representation on a device for displaying said image. Said image is for example generated by means of an image acquisition device directed towards the field of vision.
As a variant, the superposition means comprise a head-up vision device capable of displaying the representation in the field of vision.
These different devices are suitable for putting the invention into practice.
The invention also proposes a device for displaying information in a field of vision, characterized in that it comprises means for acquiring a sound sequence, means for determining information as a function of the sequence. acquired sound and means of superimposing a representation of the determined information on an image corresponding to the field of vision.
In such a device like the one described below, the determination means comprise, for example, means for transmitting data relating to the acquired sound sequence to a remote server and means for receiving information from the server. distant.
A method of displaying information in a field of vision is thus proposed, characterized in that it comprises the following steps:
- acquisition of a sound sequence;
- determination of information as a function of the acquired sound sequence;
- superposition of a representation of the determined information on an image corresponding to the field of vision.
Finally, a computer program is envisaged comprising instructions for implementing this method when this program is executed by a processor.
This device, method and program may further include the optional features presented above in terms of system, with associated advantages.
Other characteristics and advantages of the invention will appear better in the light of the description which follows, given with reference to the appended drawings in which:
FIG. 1 represents an example of a system produced in accordance with the teachings of the invention;
- Figure 2 shows a method in accordance with the teachings of the invention.
The system shown in FIG. 1 comprises a module for acquiring a sound sequence which notably includes a MIC microphone.
The microphone MIC is for example (but not necessarily) worn on VIS glasses capable of superimposing, on the field of vision of the user who wears these glasses, graphic elements (such as symbols or characters), on the command of a PROC processor, typically based on a microprocessor. The steps of the method described below which are implemented by the processor PROC thus result, for example, from the execution of a computer program whose instructions are stored in the processor PROC and which are executed by the microprocessor.
The processor PROC is also in communication with a remote server SERV, for example by means of a wireless link of a cellular network which is itself connected to the server via the Internet network. Other types of connection (wired or wireless) between the processor and the server are naturally conceivable.
As explained in more detail below with reference to FIG. 2, the sound sequences acquired by the micro MIC are transmitted to the remote server SERV (in general after a pre-processing which notably includes the digitization of the sound sequence). On the basis of the data received from the processor PROC (whether these are the digitized sound sequences or data resulting from the processing of the sound sequences as explained below), the remote server SERV performs an analysis and determines by means of this analysis the associated information to the data received.
As will be described with reference to FIG. 2, this information is for example the identity of a person whose voice corresponds to the sound sequence, or associated information (in a database stored for example on the remote server SERV ) to words identified in the sound sequence.
The remote server SERV can thus transmit to the processor PROC this information associated with the previously acquired sound sequence and the processor PROC can thus control the display of graphic elements representing this information in the VIS glasses superimposed on the user's visual field. .
Note that, as shown in Figure 1, the system may optionally also include a CAM camera, for example in order to identify images in the field of vision of the user in order for example to locate the interlocutor of the person. user and consequently display the graphic elements relating to this interlocutor superimposed at the level of the latter in the user's field of vision.
Note that the example described here provides for the use of head-up vision glasses for the superposition of graphic elements in the user's field of vision. As a variant, provision could be made for the camera CAM to acquire the user's field of view and for the superposition to be carried out on a display device (for example a screen) which simultaneously displays the image captured by the camera. camera and graphic elements in overlay.
With reference to FIG. 2, an example of a method implemented in the system which has just been described is now described.
The method begins with the acquisition of a voice sequence by means of the acquisition module comprising the microphone MIC. The sound environment of the microphone is thus converted in particular by digitization into data representing the captured sound sequence.
When an interlocutor addresses the user who wears the VIS glasses, step E2 is thus carried out with the acquisition of a voice sequence (that is to say of a sound sequence which includes the voice of the interlocutor).
The acquired voice sequence (that is to say the data representative of the captured voice) is then transmitted to the remote server SERV in step E4.
The remote server SERV then proceeds to step E6 to analyze the voice sequence received, here with the aim of recognizing the identity of the interlocutor (that is to say of the speaker whose voice is present in the voice sequence received).
This analysis comprises for example the determination of a fingerprint of the received voice sequence and the comparison of this determined fingerprint with voice prints stored in a database of voice prints hosted by the remote server SERV: one thus seeks determining whether the voiceprint of a person stored in the database matches the voiceprint determined on the basis of the received voice sequence.
It should be noted that, as already indicated, provision could be made, as a variant, to carry out a preprocessing of the voice sequence at the level of the processor PROC, for example to determine within the processor PROC the voice print corresponding to the voice sequence acquired at the level of the processor. step E2 and consequently to transmit only this fingerprint from the processor PROC to the remote server SERV. The remote server SERV can then proceed to search for the stored voice print corresponding to the received voice print.
In any event, if the analysis implemented in step E6 makes it possible to recognize (step E8) the person whose voice print corresponds to that of the voice sequence acquired in step E2, the remote server SERV transmits to the processor PROC the identity of the speaker thus identified (step E10).
If, on the other hand, the speaker is not recognized at step E8, one proceeds to step E20 described below.
Following step E10, the processor PROC receives the identity of the speaker as determined and transmitted by the remote server SERV.
The processor PROC then controls in step E12 the display in the VIS glasses of information relating to the identified speaker superimposed on the field of vision of the user who wears the VIS glasses. The information displayed is typically the name of the interlocutor as well as possibly other information associated with it. Strictly speaking, graphic elements such as characters which, taken together, represent the name of the interlocutor are displayed in the VIS glasses.
In the embodiment described here, the method continues at step E14 by the analysis by the remote server SERV of the previously received voice sequence according to a semantic recognition algorithm which makes it possible to identify within the voice sequence the words spoken by the interlocutor, for example during a predefined time.
After possible filtering of certain words in order to keep only the words of interest (for example, after deletion of the articles), the remote server SERV searches a database for contextual information relating to the words identified in step E14.
The research, which can be carried out on dedicated content or on the contrary on public websites, possibly predefined, can be oriented according to parameters defined by the user, in particular according to the context (professional, leisure, etc.) or other data, such as for example resulting from the analysis of the images captured by the camera CAM.
The remote server SERV then transmits in step E16 this information to the processor PROC. The processor PROC controls in step E18 the display of graphic elements (typically characters) representing the information received in step E16 superimposed in the user's field of vision.
The user's field of vision is thus enriched by information relating to the speech delivered by the interlocutor and which therefore completes his understanding of the latter.
The process implemented if the speaker is not recognized in step E8 is now described.
In this case, a message signaling the failure of the recognition is transmitted from the remote server SERV to the processor PROC so that the processor PROC asks the user wearing the VIS glasses to associate (step E20) the footprint determined on the basis of the voice sequence acquired in step E2 to a person (that is to say to the interlocutor) by giving (for example on a user interface, not shown, provided for this purpose) the identity of the interlocutor.
The identity thus entered by the user can then optionally be transmitted to the remote server SERV in order to be stored there in step E22.
The fingerprint-speaker association thus stored can naturally be used during a future implementation of the method illustrated in FIG. 2 (in which case the speaker will naturally be recognized in step E8).
It is also possible to provide that the information concerning a given interlocutor will be classified within the database hosted by the remote server SERV and consider that the user can share this information with other people who will thus be able to recognize the interlocutor. by means of a method of the type described in Figure 2.
We could thus provide for the emergence of a community or social network for sharing information applicable to virtual reality or augmented reality.
The embodiments which have just been presented are only possible examples of the invention, which is not limited thereto.
2 sheets
Sheet 1 Sheet 2
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN108777145A | Cited by | China | – | Search report | – |
| WO2013155154A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| US9857451B2 | Cited by | United States of America | – | Applicant | – |
| US10107887B2 | Cited by | United States of America | – | Applicant | – |
| US9360546B2 | Cited by | United States of America | – | Applicant | – |
| US9354295B2 | Cited by | United States of America | – | Applicant | – |
| WO2013155251A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| US9291697B2 | Cited by | United States of America | – | Applicant | – |
| WO2013155154A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| WO2013155251A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| US10909988B2 | Cited by | United States of America | – | Applicant | – |
| US2008186255A1 | Cites | United States of America | XY | Search report | 1,9,13,14 |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 1002745 | France | A | |
| FR20100002745 | – | – | – |
Numbers
- Publication
- 2962235
- Publication, DOCDB
- 2962235
- Publication, EPODOC
- FR2962235
- Application
- 1002745
- Application, DOCDB
- 1002745
- Application, EPODOC
- FR20100002745
Titles
- French
- DISPOSITIF ET PROCEDE D'AFFICHAGE D'INFORMATIONS DANS UN CHAMP DE VISION
Classification
- IPC, 2
- G06F3 01
- G02B27 01