Method and system for multimodal and asynchronous conference with intervention of computer existing in virtual space
Abstract
This record has no abstract on file.
Term
Term ended
Expired 12 July 2019, 7.2 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
17 claims: 2 independent, 15 dependent
- 1A system that records asynchronous multimodal events that occur in different conference sessions with a computer in a virtual environment. A synchronous multimodal data structure stored in memory that can record a description of the virtual environment and multiple aspects of each of the plurality of events that have occurred in connection with the virtual environment as a single conference session. The one that stores the document and It has a plurality of computer devices interconnected to a network, one of which stores the above data structure. Each of the rest of the plurality of computer devices allows communication between the user and the data structure. Further, the system is performed at a time after recording the one conference session in the synchronous multimodal document recorded as one conference session in which each user interacts with the data structure. Allows you to insert at least one of the events that occurred in one conference session and in a different conference session, and the synchronous multimodal document in a session different from the one conference session by at least one user. A synchronous multi that is subordinate to at least one comment, the at least one comment is merged with the synchronous multimodal document, and the at least one comment is updated in a manner that appears to have been entered in the one conference session. A system that records multimodal events that occur in different conference sessions, characterized by generating modal documents. 仮想環境においてコンピュータを介在させてなされ異なる会議セッションで起こった非同期なマルチモーダルなイベントを記録するシステムであって、 メモリに記憶されたデータ構造であって、上記仮想環境と上記仮想環境に関連して起こった複数のイベントの各々の複数の側面との説明を1つの会議セッションとして記録することができる同期マルチモーダルドキュメントをストアするものと、 ネットワークに相互接続された複数のコンピュータ装置であって、そのうちの1のコンピュータ装置に上記データ構造が記憶されるものとを有し、 上記複数のコンピュータ装置の残りの各々はユーザおよび上記データ構造との間の通信を可能にし、 さらに、上記システムは、各ユーザが上記データ構造と相互作用して1つの会議セッションとして記録された上記同期マルチモーダルドキュメントに、当該1つの会議セッションを記録した時点より後の時点で行なわれた当該1つの会議セッションと異なる会議セッションで発生したイベントのうちの少なくとも1つのイベントを、挿入させることができるようにし、上記同期マルチモーダルドキュメントが少なくとも1人のユーザによる上記1つの会議セッションと異なるセッションにおける少なくとも1つの注釈に従属し、当該少なくとも1つの注釈が上記同期マルチモーダルドキュメントとマージして、上記少なくとも1つの注釈が、上記1つの会議セッションにおいて入力されたように見える態様で更新された同期マルチモーダルドキュメントを生成することを特徴とする、異なる会議セッションで起こったマルチモーダルなイベントを記録するシステム。
- 14A computer-mediated method of recording asynchronous multimodal events that occur in different conference sessions in a virtual environment. Create a data structure to store events in the virtual environment, and use the data structure to send multiple events related to the virtual environment.As one conference session in relation to the virtual environmentRecord in a synchronous multimodal document andAt least one annotation in a conference session that is different from that one conference session at a time later than when the one conference session was recorded in association with the virtual environment.And at least saidOne annotation and already recorded in the data structureSort events,In a manner in which at least one annotation appears to have been entered in the one conference session above.Including generating updated synchronous multimodal documents,SaidMethod. 仮想環境においてコンピュータを介在させてなされ異なる会議セッションで起こった非同期なマルチモーダルなイベントを記録する方法であって、 前記仮想環境におけるイベントを保存するデータ構造を作成し、 前記仮想環境に関連して行われる複数のイベントを前記データ構造を使用して前記仮想環境と関連させて1つの会議セッションとして同期マルチモーダルドキュメントに記録し、前記1つの会議セッションを前記仮想環境と関連させて記録した時点より後の時点で当該1つの会議セッションと異なる会議セッションで少なくとも1つの注釈をを入力し、 前記少なくとも1つの注釈と、すでに前記データ構造中に記録されているイベントを並べ替え、上記少なくとも1つの注釈が、上記1つの会議セッションにおいて入力されたように見える態様で更新された同期マルチモーダルドキュメントを生成することを含む、前記方法。
Independent claims2
131 paragraphs, as filed
[Technical Fields to which the Invention Belongs] The present invention relates to a computer-supported collaborative work environment. In particular, it relates to computer-mediated methods and systems for conducting, reviewing, and supplementing multimodal and asynchronous meetings in virtual space.
[0002] [Conventional Technology] As more companies are dispersed worldwide and it becomes possible to work from home due to the development of network infrastructure, the need for meetings in an alternative form to the meeting format is increasing. There is. Existing teleconferencing technologies such as video conferencing and video conferencing meet this need to some extent, but each has its drawbacks. Traditional teleconferencing systems are expensive and unsuitable for unprepared, rushed discussions. Although the newer desktop teleconferencing system solves this problem to some extent, it tends to hinder its use due to privacy issues due to the real-time transmission of participants' photos and videos. Furthermore, in the conventional technique, all participants must be able to participate at the same time. It is especially difficult to meet these conditions, for example, when the members of one working group are scattered all over the world and there is a time difference between them.
[0003] [Problems to be Solved by the Invention] A conference session in a video conference or a video conference system can be recorded for later playback, but the viewer of the conference played in this way is not a participant and is not a participant. It's just a bystander. Annotation tools are the first step in supporting participation in replayed meetings. However, in existing systems, annotations are added at the end of sessions that have already taken place, making them distinct from the conference itself. In other words, it cannot be incorporated as one session with simultaneity.
[0004] The present invention provides a system and a method for conferencing in a virtual environment.
[0005] Further, the system and method in the present invention provide a virtual environment in which a plurality of participants meet and interact synchronously or asynchronously, which are not necessarily in the same place.
[0006] Further, it provides a method and a system for conducting, reviewing, and supplementing a meeting in a virtual environment.
[0007] Conference participants interact through avatars in a multidimensional, place-based graphical environment in a virtual environment. An avatar is a graphic representation of each participant. The meeting is saved as a multimodal document for later playback and addition. The system uses multiple "tracks" of discussion text data, voice commands, graphics, and documents in a multi-modal document to store meeting records for later participants. .. Later participants can play the multimodal document of the meeting and supplement it indefinitely, creating a single meeting flow. Furthermore, by using an avatar, it is possible to easily generate a composite screen incorporating the contents of a plurality of conference sessions, and it is possible to reduce the concern about the privacy of the user.
[0008] Participants do not have to be in the same location because the dialogue takes place over a computer network. There is no need to "meet" at the same time. In addition to supporting real-time conversations between participants who are not in the same location, dialogues in virtual environments are recorded electronically. These recorded dialogues are preserved and can be reviewed and supplemented by those who later "participate" in the meeting.
Dialogues can range from concise, with one person asking a simple question to another, to a more neat structure that takes place at a fixed time and agenda. These sessions are recorded by the systems and methods of the invention, allowing others to view and participate asynchronously. Later participants will be able to see themselves interacting with the original participants when they replay the meeting and add their own remarks. Therefore, the sense of participation is higher than when simply watching the dialogue of others later, and it is possible to feel more strongly that they are a member of the group.
Later participants' dialogue will be inserted into the recorded conference session, and later Shin participants will see new remarks mixed with earlier discussions. What is integrated as one conference can be regenerated and supplemented by later participants if necessary.
[0011] Since the use of the system and method in the present invention is not limited to irregular asynchronous conferences, the user does not need to remember special and occasionally used commands or operations. In addition, it can be part of a regular desktop environment, so there is no additional cost to set up an environment for asynchronous conferencing. In addition, frequent and synchronous interactions with co-workers' avatars will eliminate discomfort with the avatars, and asynchronous interactions with avatars that exist in the same virtual space will feel more natural.
[0012] A virtual environment that allows synchronous and asynchronous participation is a multidimensional, place-based system in the virtual environment where participants discuss through avatars. The content of each remark is emitted from the avatar in the form of a cartoon balloon. In such a place-based system in a virtual environment, there are multiple "rooms" where participants can gather, and a discussion held in a room is that particular "room." Only the participants who are "in" can experience it. Participants can also use different avatars or use the "Paint" tool to draw question marks, etc. to enhance the discussion and express emotions such as confusion and dissatisfaction.
Materials that supplement the meeting include meeting agendas and idealists that can be generated during the brainstorming phase, and suggestions and design materials that are shared through commands that update each user's interface. These work-related materials play an important role in harmonizing the activities of the participants. Furthermore, in special rooms, it is possible to vote using "tokens" and write on a shared whiteboard.
[0014] All events in a conference conducted according to the systems and methods of the invention are recorded for later reproduction and further supplementation if necessary. Later participants can play a recording of the actual meeting and observe the decision-making process, rather than just reading a summary of what happened at the meeting. This playback process involves observing each participant's avatar "speaking" and changing facial expressions, as well as changing the content of the material and the description on the whiteboard as the meeting progresses. .. In addition, users can "fit" themselves into the meeting by adding their own opinions and become more involved than simply observing. These additional opinions will be incorporated into the conference document and will be included in subsequent playbacks.
[0015] The use of playback functions is emerging in computer-backed collaboration. However, according to the systems and methods of the present invention, a once passive observer can enter a recorded conference and become an active participant in an already held conference. Further, the present invention allows a plurality of users to watch a conference and supplement it, unlike the conventional technique in which a user can only see a conference session held earlier. Since the backgrounds such as avatars and rooms use simple graphics, it is easy to create a composite screen that inserts the avatars of people who later participated in the previous scene. This compositing process is much easier than using video. When using video, it takes a considerable amount of effort to synthesize a part selected from one source among a plurality of video sources so that the part selected from one source appears in another source. In addition, since the avatar, which is a graphic, is used, there is no privacy problem associated with video conferencing.
[0016] Recording and live participants to prevent confusion between participants participating in "live" on the composite screen and participants in the recorded session as later conference sessions take place. Can give a different look. In addition, a number label can be placed near the avatar to clearly identify which session a participant was attending.
Distributed with the text / graphics-based system of the present invention by integrating shared collaborative tools such as whiteboards and editors with features for synchronous and asynchronous meetings and avatars in virtual environments. It makes the methods for awareness and collaboration in the working group more applied, user-friendly, comprehensive and productive.
[0018] The first invention is data stored in memory that stores a virtual environment and a synchronous multimodal document capable of recording a description of a plurality of aspects of each of a plurality of events taking place in relation to the virtual space. It has a structure and multiple computer devices that are connected to a network, store the data structure in one of them, and each otherwise communicates with the user and the data structure, at least by one user. For each user to enter at least one of the plurality of events in the synchronous multimodal document with an asynchronous addition that produces one, merged and updated synchronous multimodal document with the synchronous multimodal document. It is a system that records multimodal asynchronous events mediated by a computer in a virtual environment, which is characterized by being able to interact with the data structure.
[0019] The second invention is a system characterized in that the updated synchronous multimodal document in the first invention can be further supplemented and merged.
[0020] The third invention is a system characterized in that, in the first invention, each user can further examine the updated synchronous multimodal document.
[0021] In the first invention, the fourth invention includes at least one multimode recording control means for recording each event in the multimodal document, and supplementary control for making additions to the synchronous multimodal document. The system comprises means and at least one data control means having a replay control means for at least one user to replay at least one event of the synchronous multimodal document.
[0022] A fifth invention is a system according to the first invention, wherein the virtual environment has at least one prop that is located within the virtual environment and can be moved.
[0023] The sixth invention is a system characterized in that, in the first invention, the virtual environment has a link to another virtual environment.
[0024] The seventh invention is a system according to the first invention, wherein the virtual environment has a virtual space in which an event is performed.
[0025] The eighth invention is a system characterized in that the virtual space is a virtual conference room in the seventh invention.
[0026] The ninth invention is a system characterized in that a character string input by a user in the first invention is displayed as a balloon in the virtual environment.
[0027] The tenth invention is a system according to the first aspect, wherein the virtual environment has at least one area for displaying information related to the plurality of events.
[0028] The eleventh invention is a system according to the first invention, wherein the virtual environment has a means for each user to input a command to the system.
[0029] The twelfth invention is a system according to the eleventh invention, wherein the command includes recording start, recording stop, reproduction, information loading, fast forward, and rewind.
[0030] A thirteenth invention is a system characterized in that, in the eleventh invention, the command further includes a new conference, goto, information loading, audio reproduction, and video reproduction.
[0031] The fourteenth invention creates a data structure for storing an event in the virtual environment in a method of recording and adding a multimodal asynchronous conference via a computer in the virtual environment, and the virtual environment has the data structure. Record multiple related events in a synchronous multimodal document using the data structure, enter at least one asynchronous addition, and sort the recorded events with at least one asynchronous addition. A computer-mediated, multimodal, asynchronous conference recording and addition method in a virtual environment, including generating updated synchronous multimodal documents.
[0032] The fifteenth invention inputs a further asynchronous addition in the fourteenth invention, rearranges the events of the further asynchronous addition and the updated synchronous multimodal document, and the synchronous multimodal document. Is a method that involves updating.
[0033] A sixteenth invention is a method comprising monitoring at least one asynchronous addition and updating at least one user interface based on each asynchronous addition in the fourteenth invention.
[0034] The seventeenth invention is a method according to the fourteenth invention, which comprises playing at least one event for at least one user through the user interface.
[Embodiment of the Invention] FIG. 1 shows an embodiment of a multimodal asynchronous conferencing system 100 via a computer in the present invention. The system 100 includes a multimodal document 105 and several user interface displays 120. The user interface and display 120 are each connected to the multimodal document 105. Users who are spatially and / or temporally separated from other users should use each user interface and display 120 to interact asynchronously with the multimodal document 105 associated with one or more users. Can be done. As users interact with each other through the user interface and display 120, asynchronous conferences in virtual space are established and recorded. One of the users chairs the meeting.
[0036] The multimodal document 105 is a data structure stored in memory capable of recording all aspects of an asynchronous conference. Data structures that are usually stored include discussion text data, voice commands, graphics, materials, and conference agendas. When these and other data structures are combined as a multimodal document 105, it can provide the complete content of the conference for later participants.
FIG. 2 shows a general purpose computer 200, such as a server for running system 100 in FIG. Generally, various user interface displays 120 (clients) and general purpose computer 200 can be run in a traditional client-server architecture. In particular, the general-purpose computer 200 includes a multimodal document 105, an input / output interface 210, a supplementary control means 220, a multimodal recording control means 230, a playback control means 240, and a system clock 250.
[0038] The multimodal recording control means 230, supplementary control means 220, and playback control means 240 input data to the multimodal document 105 and / or receive data output from the multimodal document 105. The user interface display 120 is connected to the general purpose computer 200 implementing the asynchronous conferencing system 100 through the link 130 and the input / output interface 210. Conference participants interact with each other through the input / output interface 210. Participants' inputs and displays are routineized from their respective user interface displays 120 through input / output interfaces 210 and one or more appropriate 220-240 controls. Which control means is used depends on whether the asynchronous conferencing system 100 is recording the conference, supplementing the already recorded conference, and playing the conference.
[0039] The multimodal recording control means 230, supplementary control means 220, and playback control means 240 are configured to allow conference participants to record, play, and supplement an ongoing conference or a conference that has already been recorded.
[0040] The system clock 250 is connected to the multimodal document 105 and stamps the dialogue between the participants stored in the multimodal document 105. This ensures that each event is played back in the correct order when the meeting is played back later. The system clock 250 has two imprints on each of the user's interactions. The first is a "real" stamp that records the time when the interaction actually took place. The second is to record and store the sequence of participants' interactions in the supplemented conference session. These two markings will be described in detail in the following discussion of conference playback options.
[0041] In order to record, play back, or supplement the conference recorded in the multimodal document 105, the asynchronous conference system 100 in this embodiment uses a client-server architecture in which the conference actually ". It is "held" and the server is programmed to record and store each event in the meeting. Each event at the meeting is recorded in the multimodal document 105. This multimodal document 105 is text input data, voice commands such as pre-recorded sound files, graphical representations of participants such as avatars and props, conference scenes such as rooms and backgrounds, and created by participants. It contains information on multiple "tracks" such as supplied materials and whiteboard notes. The meeting recording includes a stamp and a record of the meeting event. Each event in a conference brings changes to one or more of these "tracks."
[0042] The engraving has two functions. One is to record the date and time when the event, or user interaction, actually took place. The second is that it is added to the event to function as a number to indicate the position in the order of the events that make up the conference.
[0043] FIG. 3 shows a virtual environment in operation. The computer-mediated, multimodal asynchronous conference according to the system and method of the present invention begins in virtual space. The meeting chair usually synthesizes or selects the virtual environment he wants to use from the default scene templates. As shown in Figure 3, the virtual environment usually includes an environment where participants can place their avatars, such as a conference room. However, any virtual environment can be created to meet the needs of the participants.
[0044] During the first session of the conference, each user launches a client software program of the present invention to join another user in virtual space. Events from the client are accepted by the server and queued. The server dequeues and processes each event. For each event, the server can identify the user and the time the event took place, the type of event, and some of the content of the event. These events include spoken text, voice commands, room appearance changes, avatar changes, avatar position changes, prop appearances and changes, prop position changes, browser loading and whiteboard writing. Etc., includes input from the user.
[0045] A written statement is an event that includes a character string. During the meeting, each statement 1005 appears to be emitted from the speaker's avatar in a cartoon balloon-like style, as shown in Figure 3. The text data of the content of the remark is stored in the multimodal document 105 together with the personal information and stamps attached to the user.
[0046] A voice command is a pre-recorded sound. When a user presents a voice command, the server saves the user, time, identification of the event, and a periodic redundancy check of the voice file in which the voice command is stored. The server then keeps a copy of each audio file in the multimodal document 105. When an audio file is needed, the client downloads the required audio file from the server, caches the audio file, and plays it. Similar on-demand uploads by the server occur when new sound files, prop descriptions, avatar descriptions and virtual environment background files are referenced in the event.
The virtual environment 1000 in FIG. 3, that is, the conference room environment, is the room in which the chair is present at the start of the conference. The room description event is sent from the client (in this case the chair's client) to the server. The room description event contains the room name of the virtual environment 1000, the graphics file corresponding to the actual background of the room, and the time. The server stores the name and identity of the graphics file and the client ID, stamping and periodic redundancy checks. Users entering the room during the meeting will see the graphic background shown in Figure 3 and described at the room description event.
By capturing and managing files, audio and / or graphics each time they are used, the user or client can download the stored conference and revisit it on the client machine where the network connection has been lost. It is also possible.
[0049] Avatar 1010 is a graphic representation of the conference participants. The user may create an avatar 1010 or choose from pre-prepared characters. One or more of the parts that make up Avatar 1010 may be a video. When the user first appears, or when the user changes the appearance of the avatar, the client sends a user appearance event to the server. User appearance events are graphics files that store the client's identity, time, new avatar appearance, periodic redundancy checks of the Avatar 1010 parts list, and several graphics files to create one avatar. Includes an array size that usually indicates whether is overlaid, the coordinate position of the avatar in the virtual environment 1000. You can create more complex avatars by overlaying multiple graphics files. In addition, overlaying multiple graphics files allows the user to partially change the "base" of the avatar, for example when expressing emotions. For audio and room graphics, if name, identity, or periodic redundancy checks indicate that the server does not have the latest Avatar 1010 graphics files, then the server Separately requests the client to upload the required files and generate identification information for the components of that graphics file. Therefore, when a user changes the appearance of Avatar 1010 during a meeting, all users will see the updated appearance. The server stores an array of component identification information, client identification information, time, and avatar position after each event for each participant.
[0050] The user position indicates the position of the user's avatar 1010 in the virtual space 1000. When a user moves their avatar in virtual space, the client sends a user position event to the server, including client identification information, time, and the new position of the avatar in virtual space. This information is stored in the multimodal document 105.
Props 1020 are other objects that can be moved in virtual space. The prop 1020 can be incorporated into the virtual space 1000 and moved in the virtual space like a component of an avatar. When the props 1020, such as the coffee cup 1021 and memo 1022 shown in Figure 3, first appeared, or when the appearance of the props 1020 changed, the client was asked to identify the client, the name of the graphics file, and the graphics file. Sends a prop appearance event that includes identification information, periodic redundancy checks, and the position of the prop in virtual space. For audio and room graphics, if name, identity, or periodic name redundancy checks indicate that the server does not have the latest prop 1020 graphics files, then the server Separately requests the client to upload the required files and generate identification information for the components of that graphics file. The server generates prop identification information for the prop 1020. The system then stores prop identification information, client identification information, time and location.
When the prop 1020 is moved within the virtual space 1000, the client sends a prop position event containing the client identification information, the time and the new position of the prop in the virtual space. During the meeting, each user can see the movement of prop 1020 in virtual space 1000. The server stores the client identity, time, and the new location of prop 1020 in virtual space in the multimodal document 105.
FIG. 4 shows a browser. The browser load command is used with a browser or similar display interface. When the client sends a browser load command, this event contains a string such as the URL (uniform resource locator) or address of the data to be loaded, the client identity, and the time. The page is loaded into each user's graphical user interface, as in the case of World Wide Web Document 1030 during the meeting and shown in Figure 4.
FIG. 5 shows an example of a paint stroke 1040. Pen or paint strokes, and supplementary eraser strokes, are "pictures" that the user can draw anywhere in Virtual Environment 1000. All aspects such as the position, color, line width and direction of stroke 1040 are determined by the user. When a user makes a paint stroke or eraser stroke on a shared whiteboard or anywhere in the virtual environment 1000, the client gives the server stroke information, including the color, line width, and the location of the points that determine stroke 1040. With time and identification information. The server stores time, client identification information, and stroke information along with information about the start and end points of the stroke. This information allows the server to later reproduce stroke 1040 in the correct direction.
[0055] There are at least two ways to revisit a conference recorded by the asynchronous conference system 100 of the present invention. In sequential mode, the user follows either the actual time as described above from the beginning of the meeting or the time to indicate the logical event order when playing the meeting. Just go. In random access mode, the user can jump in the middle of an already recorded meeting.
[0056] In sequential mode, the recorded meeting is played from the time the chair begins the meeting. The real-time array allows the user to play the meeting in the actual order in which the events took place. In the case of a logical array of times, the server reorganizes the events as if the meetings were held synchronously.
[0057] In random access mode, a user who wants to go directly to a particular event or part of a recorded conference can use one of the two conference summaries as a navigation tool. The user specifies whether to display the summary in chronological order based on the actual time stamp or in a logical order based on the event stamp. The logical order is the order of event markings, including events inserted later in an asynchronous conference.
[0058] In either mode, the user is presented with a plurality of summary screens. The summary screen links a particular action, event, remark, etc. to where the event took place during the meeting.
FIG. 12 shows a text log of discussions, a summary 1045 and an audio log. Each entry in log 1045 is a link that starts playing the recorded conference from the point corresponding to the selected entry.
[0060] This view supports two important processes. One is the outline of a meeting that can be scanned quickly by supporting browsing and search goals, and the other is a part of the meeting based on a particular interest by providing direct access to the area of interest. It is designed to be reproducible. So if one user is particularly interested in the agenda of a conference that is trying to restructure a meeting with other members of the group, the user can go directly to the point where the agenda was being discussed. You can. Because the summary tells you, for example, when the material for the conference was presented. Similarly, if a member of the group was browsing and was interested in someone's opinion on a topic, the summary would allow participants to say when and what material they could see. It becomes clear. In this simple example, the virtual environment 1000 changes little during the meeting. However, in an actual meeting, people may move in and out of the room, change positions as their roles change, and the meeting itself may change the venue.
[0061] Therefore, there are a number of points of access for participants playing the meeting for the first time to add comments, and a summary review should facilitate an understanding of the structure of the meeting and the discussions that took place during the meeting. Can be done. The summary provides a quick and meaningful restructuring of the meeting for participants who replay the meeting to see what has changed from the previous dialogue.
[0062] Playing back a recorded conference from the beginning can be easily done using any of the event sequences. The server simply sorts and plays the events in the order the user wants. The state of the virtual environment 1000, props 1020 and avatar 1010 are restored to the state when the selected event first took place.
[0063] However, if one wants to join at some point in the middle of a meeting, the system must be able to reconstruct the graphical state of the meeting. More specifically, the presence and appearance of Avatar 1010, the presence and appearance of Props 1020, the appearance of Virtual Environment 1000, and the content of the shared browser need to be reintegrated. Since the text remarks 1005 and voice occupy only a very short time and are events rather than contributing to state changes, the asynchronous conferencing system 100 in the present invention is responsible for regenerating the state of the conference. Do not rebuild the newest text or audio events. Asynchronous conferencing system 100 simply plays text remarks 1005 and voice made after the selected time.
[0064] Two different classes of events must be reconstructed when regenerating the state of the conference. That is, events that update by replacing one state with another, such as changing the background of a room in virtual environment 1000, and events that contribute to the accumulated state, such as writing to a whiteboard. There are two.
Events that are updated by replacing one state with another include the presence or absence of props and avatars, appearance, movement, room background changes in virtual environments, browser content changes, and the like. For such events, the state is reconstructed by going back in the event history and the asynchronous conferencing system 100 finding the newest, relevant overwrite event. It is possible that such an event only takes place at the beginning of the meeting. In such cases, the asynchronous conferencing system 100 needs to scan all the way to the start of the conference to reproduce the state at the desired restart point.
[0066] With respect to the accumulation event, the state is reconstructed by tracing back the event history and accumulating all the related events from the start of the conference until the target restart point is reached.
[0067] During a conference replay, most users will see the same thing as when attending a "live" sync session, with one exception. The exception is that some avatars 1010 show a "raw" look and some show a "recorded" look. Participants represented by the "recorded" appearance of Avatar 1010 indicate that they are not in the online environment at that time, indicating that their dialogue was previously recorded. If a user's avatar's "live" and "recorded" appearances may both appear when the meeting is playing, it means that the user is currently attending the meeting. It also indicates that he was also involved in a previously recorded dialogue.
[0068] During the playback of the conference, the user can pause the playback at any time and add his / her own comments, figures, and the like. Events taken to supplement the meeting during playback are recorded in the same way as events taken in the original session, except for the engraving. The actual date and time of the event is stored as a true stamp. However, the events in the original session are stamped with integer events, but the events added during playback are in the order of the original event stamps by adding the period and supplement number Y to the original event stamp X. Inserted inside. That is, the additional number of the event added immediately after the playback event X is X.1. The event and addition number XY of each event that is continuously added before the next playback event may be incremented by 1 from the event and addition number XY of the event that was added before that. For example, if you add three new events between replay events 3 and 4, those events will be given event markings of 3.1, 3.2, and 3.3. Then, all events including those added at the end of playback are newly numbered based on the integer event stamp. A similar process is performed for later regeneration and supplementation.
[0069] Figures 6 and 7 show how user events are recorded and displayed on the displays of other participants in the conference. Figure 6 roughly illustrates how the user's request for the next scene is handled. In (1), the user interface display 120 requests the asynchronous conference system 200 for the next scene. In (2), the asynchronous conferencing system 200 advances the logical time, and in (3) sends the request for the next event to the multimodal document 105. In (4), the multimodal document 105 sends an event to the asynchronous conferencing system 200, and in (5) the asynchronous conferencing system 200 distributes the event to each user. In (6), each user interface display 120 plays a "knock" file.
[0070] Figure 7 illustrates a record of a new event. In (1), the input character string "Hi" is received from the user 120. In (2), the user ID and the stamp are added to the character string by the asynchronous conference system 200. In (3), the new event is recorded in the multimodal document 105. In (4), the user ID and associated string are distributed to each user interface display 120 and then displayed.
FIGS. 8-11 show a set of graphical user interfaces displaying a virtual environment 1000 in which a conference in which two conference participants 1100 and 1120 are attending is taking place. The virtual environment 1000 contains a number of props, including a clock 1130 that can be used to arrange comments in chronological order during playback. Conference participants are represented by static avatars 1100 and 1120, respectively, made of photographs. These avatars are static in the sense of gestures, but can be placed anywhere in the virtual environment 1000. In this embodiment, the user needs to select the desired prop and drag it to a new location in order to move the prop. Conversely, to move an avatar 1100 or 1120, the user simply selects the location in the room where he or she wants to place his or her avatar. In response, the user's avatar is automatically placed in the selected position. In addition, there are several so-called "hot spots" within the virtual environment 1000. When the avatar is placed in one of the "hot spots", the user is automatically moved to another virtual environment. Here, another whiteboard virtual environment is created as an example. You can enter the whiteboard virtual environment by placing the avatar 1100 or 1120 on the whiteboard 1145 at the top left of the virtual environment 1000 shown in Figures 8-11. In the whiteboard room, all group members can draw explanatory diagrams and leave messages to each other. Since the figures and messages will remain in this whiteboard virtual environment, members who participate later can also see the entries on the whiteboard and comment on them on the whiteboard.
Users can interact in a conference and control navigation by typing text into the input area 1140 of each user interface display 120. The user selects the input area 1140 and types characters into it. If it is a command as described below, it will be sent to the server. If it is a conversation, it is represented as a "speech" in the form of a speech bubble 1005 adjacent to the user's avatar 1100 or 1120. In this way, users can have discussions in real time. FIG. 12 shows an example of log 1045. Log 1045 can be viewed in a web browser along with the associated link 1160. It should be noted that the dialogue in Figure 12 is different from the part of the conference illustrated in Figure 8-11. Log 1045 will be described in detail later along with navigation.
[0073] The commands that the user presents to the server for starting, ending, playing, visiting a specific scene, loading a web browser, and loading a conference summary include "NewMeeting", "EndMeeting", and ". Includes "ReplayMeeting", "LoadBrowser", "GotoScene", "NextScene" and more. The system responds to the "NewMeeting" command by returning the name of the new conference document. The user issuing this command is the chair of the conference. For the "EndMeeting" command, the system ends the current meeting. That is, it stops recording and closes the conference document. For "logical" or "chronological" "ReplayMeeting <meeting document name>" commands, the system displays the virtual environment at the start of the requested meeting. For "LoadBrowser, <URL>", "LoadSoundClip, <filename>" or "LoadVideoClip, <filename>", the system will use the specified URL, audio file or video file for all participants' browsers. Load to. For "GotoScene <number>", the system displays the indicated scene. That scene becomes the current scene. This command has no effect when no playback is taking place. For the "NextScene" command, the system advances to the next scene after the current scene. If it is just after the playback starts, the "NextScene" command proceeds to the playback of scene 1. This command has no effect when no playback is taking place.
[0074] Further, it is possible to retrieve data from this virtual space. In Figure 13-15, web pages related to various topics are illustrated. Figure 13 shows the conference agenda. In many cases, one user is the chairman, but we will follow the items on the agenda one by one. This agenda is responsible for the chair organizing the conference and for embodying the replay and supplementary sessions of the conference for later participants. The asynchronous conferencing system 100 in the present invention records and logs everything that has been done. In the example conference illustrated in Figure 8-11, a number of resources are presented on topics of interest to the participants in the conference. These materials are shown in Figures 14 and 15. As shown in Figure 16, all participants will be able to see these materials on the graphical user interface 1170, which is displayed adjacent to the virtual environment 1000 when presented. By launching the page in this way, participants can share the information they have provided or external materials related to the conference.
[0075] Figure 17-20 shows the following dialogue conducted by participants who later attended the conference. The two new participants, represented by avatars 1190 and 1200, visit the first part of the meeting to understand the content of the discussions that have already taken place. The avatars 1100 and 1120, which represent the previous participants, have a recorded appearance here, as the previous participants are not actually participating online now.
[0076] In Figure 21, the names of the new participants have been added to the Participant Names column on the agenda displayed in the window of the graphical user interface 1170. Each line of dialogue previously typed in the meeting is replayed. The avatars 1100 and 1120, which represent the original speakers, are displayed where they made the text statement 1005, respectively. However, these avatars now have a recorded appearance that distinguishes them from the "live" participants. Two participants currently in the meeting can insert comments at any time during playback, and the inserted comments will be added to and stamped in log 1045. Materials were presented during the previous meeting session and, if updated, will also be displayed on display 120 of the new participants. For example, in this scenario, Agenda 1175 shows the current progress items, removes old items (such as training time and place), and shows new information (later participants), as shown in Figure 22. Has been updated in various ways for (name). The updated information is an annotation, such as being displayed in a different color, and which participant added each item.
[0077] During the replay of the conference, Agenda 1175 is being updated to show that an agreement has been reached on who will attend the CSCW conference, especially as shown in Figure 32. At the end of both sessions of the conference, Agenda 1175 has partially turned into a record of what happened at the conference. And some of the action items in Agenda 1175 are assigned to specific people.
By using version control and keeping a log of everything that happened during the meeting, we can unravel what the dissenting opinions were and how they were resolved before reaching a decision. Can be done. Further, the asynchronous conferencing system 100 in the present invention provides a specific method for accessing a specific part of a recorded conference to facilitate later playback and supplementation. The conference itself, including all dialogues and supplementary material referenced during the conference, defines the multimodal document 105. The asynchronous conferencing system 100 in the present invention provides participants with a number of ways to access important parts of the recorded conferencing document 105. These are playback view, text view and multi-view.
The playback view allows you to replay and annotate the entire conference by following each scene. In the replay view, later participants view the dialogue in its original order, insert comments, review and update supplementary material presented during the meeting, and are new at any point in the replayed time flow. You can create materials. By using the playback view, group members in different locations can share the same graphics state, voice commands, and materials in virtual environment 1000 and actually participate in a meeting with a colleague as if the meeting was taking place. You can participate as you did.
[0080] The text view is the content of the conference and the basic presentation content of the log 1045 of all strings entered in the virtual space. As shown in Figure 12, for each statement in the log, the text view shows the speaker 1220, the stamp 1230, and the time series number 1240, which indicates in which thread the statement was made. Figure 12 shows a meeting taking place online. The asynchronous conferencing system 100 in the present invention assigns the first dialogue to the time series number 1 and gives the next smaller time series number to any subsequent visits by the user.
[0081] The text view allows a quick quick scan of the text log 1045, and facilitates keyword look ahead and backup search to support fast scanning. Using the text view, the user can browse through the text, navigate using keyword search, and / or scroll through the meeting document 105. From any column of the text view, the reader can follow links to other materials that form the complete meeting document. Those other materials are displayed as they were when the text string was entered. From the text view, the user can navigate from the text of the meeting content, such as the remarks of the participants, to the context of the meeting. The context of the meeting includes the statements already made, who was attending the meeting, and the state of the shared tools. That is, the text view can represent a meeting in a much simpler way, and allows quick navigation to media with more information.
[0082] Figure 23-27 shows an example of a simple meeting flow. Figures 23-27 show, in particular, the first recording of this conference, the first time it took place. Figure 28-32 shows this meeting taking place a second time, that is, when the original meeting ended and later participants were reviewing and / or supplementing the already recorded meeting. ..
[0083] FIG. 23 shows what kind of display can be displayed on each user's display. The conference agenda 1175 is displayed on the left side of the graphical user interface screen 1150 within the graphical user interface 1170, and the conference virtual environment 1000 is displayed on the graphical user interface window 1180 located in the upper right part of the graphical user interface screen 1150. In addition, the input area 1140 where participants enter characters is displayed in the lower part of the virtual environment 1000. From Figure 23, it is clear that there are two participants in the first session of this conference. In Figure 24, it can be seen that when the meeting started, one of the participants made a written statement 1005. Figure 25 is a snapshot of the midpoint of the meeting. As can be seen from Agenda 1175, participants are currently working on "Item 3", and the corresponding part 1255 of the "Agenda" in Agenda 1175 is highlighted and is a topic currently under discussion. It is shown. In addition, a question mark 1040 can be seen above one user's avatar 1010 on the graphical user interface window 1180. As mentioned earlier, these question marks are the result of paint input added to the graphical user interface window 1180, or indicate that Avatar 1010 has been modified to show the user's emotions.
[0084] FIG. 26 is a snapshot of the same meeting as it progresses further. In Figure 26, the left side of the graphical user interface 1170 has changed to a screen showing the conference program discussed earlier. In Figure 27, one of the participants marks the end of the meeting.
[0085] Figures 28-32 show that this conference is for the second session, the avatars 1100 and 1120 of the previous participants are shown in the recorded appearance, and the avatars 1190 and 1200 of the new participants. Is represented by its raw appearance. Figure 29 shows the start of the second session of the conference. Looking at Figure 29, the number of participants in Agenda 1175 has increased from 2 to 4. Figure 30 is a snapshot of the discussion of "Item 2" in Agenda 1175. As shown in FIG. 30, it can be seen that blowout 1260s showing simultaneous remarks from different participants can be displayed at the same time. Figure 31 shows that previously linked programs are also displayed for new participants. In addition, Figure 32 shows how the agenda in Agenda 1175 progresses as decisions are made and the meeting ends.
[0086] FIG. 33 is a flowchart showing how a multimodal conference is generated and is complemented asynchronously. In the S200, the system determines if the meeting is a new meeting. If so, proceed to S300, otherwise fly to S900.
[0087] In S300, the chair is decided and the asynchronous conference system 100 in the present invention is initialized. Initialization involves defining the virtual environment in which the conference will take place and interconnecting the participants who will attend the conference from the beginning. Recording of multimodal documents starts on S400. At S500, a meeting will start among the participants and proceed to S600.
[0088] In the S600, participants' interactions are recorded and stored in a multimodal document. As mentioned earlier, these interactions can be recorded as necessary or desirable to provide text input, voice commands, graphics, whiteboard writing, and other complete content of the meeting to later participants. Interactions can be included. Recording is stopped on the S700 at the end of the meeting. And on the S800, participants are asked if they want to replay the entire or part of the meeting after the meeting is over. If no participant chooses to play the conference, proceed to S1600. If more than one participant chooses to replay the meeting for review or further supplementation of the meeting content, proceed to S900.
[0089] In S900, the multimodal document 105 is played for the original or new participants, the current participants. Next, in the S1000, current participants have the option of supplementing the already recorded multimodal document. If more than one participant chooses to supplement the already recorded multimodal document, proceed to S1200. If no participant has selected to add to the already recorded multimodal document, proceed to S1100. In S1100, the playback of the conference ends and proceeds to S1600.
[0090] At S1200, supplementation to already recorded multimodal documents by current participants begins. In the S1300, participants' dialogue is inserted in a multimodal document in a time-series manner in the meeting to create a synchronous meeting. The S1400 stops recording when one or more participants finish supplementing an already recorded multimodal document. In the S1500, one or more participants choose to continue playing the already recorded conference, or end playback if they do not make further remarks. If the current participants have finished considering the meeting at S1500, proceed to S1600. If you choose to continue playing the conference for further review or supplementation, you can return to the S900 to continue the playback and supplementation process. The process ends in S1600.
That is, the method for asynchronous multimodal document generation and asynchronous supplementation is divided into three basic processes. The first process incorporates the conference into a multimodal document. The second process plays the conference captured in the multimodal document, and the third process supplements the conference captured in the multimodal document to create an integrated recording for later review.
[0092] Although the exemplary system in this embodiment operates by interposing an application such as a web browser, other applications can also be intervened. For example, in an asynchronous meeting, when a participant in a later session provides an answer to a question presented by the previous participant, the system can email the previous participant. In addition, if the user is familiar with the events that took place at the meeting, they can annotate without playing the meeting.
As shown in FIGS. 1 and 2, the asynchronous multimodal conferencing system is preferably implemented on a general purpose computer with a single program or a distributed program. However, asynchronous multimodal conferencing systems also include electronic wires such as one or more dedicated computers, programmed microprocessors or microcontrollers, peripheral integrated circuits, ASICs or other integrated circuits, digital signal processors, digital circuits. It can also be implemented on circuits or logic wire circuits, PLDs, PLAs, FPGAs, PALs and other programmable logic devices. In general, any device capable of implementing the flowchart shown in FIG. 33 can be used to implement one or more asynchronous multimodal conferencing devices.
[0094] Further, the asynchronous multimodal conferencing system can be realized by a server or other node on a local area network, a wide area network, an intranet, the Internet, or other distributed networks. Also, virtual environments and avatars are not limited to 2D graphics. A three-dimensional avatar can provide a more realistic image. Also, the link 130 may be either a wired or wireless link to the asynchronous conferencing system 200.
In addition, conference control is shown in relation to the character commands entered in text box 1140, but all selectable elements in the asynchronous conference system, such as avatars and props, are connected to at least user interface 120. You can select and control the mouse and / or keyboard, or any other controller used for the graphical user interface. In addition to the mouse and keyboard, other controllable selection devices such as joysticks, trackballs, touchpads, optical pens, and touch screen displays can also be used.
Although the present invention has been described in line with the particular embodiments described above, it will be apparent to those skilled in the art that many alternatives, modifications, and variations will be readily apparent. Therefore, the preferred embodiments of the invention described above are not for the purpose of limitation, but for the purpose of explanation, and deviate from the purpose and scope of the invention defined in the claims. It can be changed in various ways without any need.
[Effectiveness of the Invention] As described above, according to the present invention, a conference session is recorded in a multimodal document so that it can be played back and supplemented later, and played back and supplemented endlessly by later participants. And can collaborate without being bound by simultaneity. In addition, avatars that represent participants can be used to easily generate composite screens for multiple sessions and reduce user privacy concerns.
BRIEF DESCRIPTION OF THE DRAWINGS [Fig. 1] Fig. 1 is a block diagram showing a function of an asynchronous conference system in the present invention.
FIG. 2 is a block diagram showing FIG. 1 in more detail.
FIG. 3 is a diagram showing a typical virtual environment of the asynchronous conference system in the present invention.
FIG. 4 is a halftone image showing an example of a first graphical user interface screen displayed during a conference.
FIG. 5 is a halftone image showing an example of a second graphical user interface screen displayed during a conference.
FIG. 6 is a block diagram showing processing of a request to the next scene.
FIG. 7 is a block diagram showing event recording and user display updates.
FIG. 8 is a halftone image showing a graphical user interface screen of a virtual environment at the start of the first session of a conference.
FIG. 9 is a halftone image showing a graphical user interface screen of a virtual environment showing a second event in the first session of a conference.
FIG. 10 is a halftone image showing a graphical user interface screen of a virtual environment showing a third event in the first session of a conference.
FIG. 11 is a halftone image showing a graphical user interface screen of a virtual environment showing a fourth event in the first session of a conference.
FIG. 12 is an example of a text log of an asynchronous conference system according to the present invention.
FIG. 13 is a halftone image showing a graphical user interface screen displaying a conference agenda.
FIG. 14 is a halftone image showing a graphical user interface screen displaying data accessed using a browser.
FIG. 15 is a halftone image showing a graphical user interface screen displaying data accessed using a browser.
FIG. 16 is a halftone image showing a graphical user interface screen displaying an example of user and representative screens during the first session of a conference.
FIG. 17 is a halftone image showing a graphical user interface screen of a virtual environment at the start of the second session of a conference.
FIG. 18 is a halftone image showing a graphical user interface screen of a virtual environment showing a second event in the second session of a conference.
FIG. 19 is a halftone image showing a graphical user interface screen of a virtual environment showing a third event in the second session of a conference.
FIG. 20 is a halftone image showing a graphical user interface screen of a virtual environment showing a fourth event in the second session of a conference.
FIG. 21 is a halftone image showing a graphical user interface screen displaying an example of user and representative screens during the second session of a conference.
FIG. 22 is a halftone image showing a graphical user interface screen showing that the agenda has been updated as the second session of the conference progresses.
FIG. 23 is a halftone image showing an example of a further graphical user interface screen during the first session of a conference.
FIG. 24 is a halftone image showing an example of a further graphical user interface screen during the first session of a conference.
FIG. 25 is a halftone image showing an example of a further graphical user interface screen during the first session of a conference.
FIG. 26 is a halftone image showing an example of a further graphical user interface screen during the first session of a conference.
FIG. 27 is a halftone image showing an example of a further graphical user interface screen during the first session of a conference.
FIG. 28 is a halftone image showing an example of a further graphical user interface screen during the second session of the conference.
FIG. 29 is a halftone image showing an example of a further graphical user interface screen during the second session of the conference.
FIG. 30 is a halftone image showing an example of a further graphical user interface screen during the second session of the conference.
FIG. 31 is a halftone image showing an example of a further graphical user interface screen during the second session of the conference.
FIG. 32 is a halftone image showing an example of a further graphical user interface screen during the second session of the conference.
FIG. 33 is a flowchart showing a method of an asynchronous conference system in the present invention.
[Code description] 100 Multimodal asynchronous conference system 105 Multimodal document 120 User interface display 200 General-purpose computer 210 Input / output interface 220 Supplementary control means 230 Multimodal recording control means 240 Playback control means 1000 Virtual environment 1005 Text remarks 1010 Avatar 1020 Props 1021 Coffee Cup 1022 Memo 1030 WW Document 1040 Paint Stroke 1045 Text Log Summary 1100, 1120 Static Avatar 1130 Clock used to arrange comments in chronological order during playback 1140 Input Area 1145 White board
3 members in 2 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 09123518 | United States of America | – | |
| 12351898 | United States of America | A | |
| 1998123518 | – | – | – |
| US19980123518 | – | – | – |
Members3
| Document | Office | Kind | |
|---|---|---|---|
| JP2000050226A | Japan | A | |
| US6119147A | United States of America | A | |
| JP3675239B2This record | Japan | B2 |
15 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentA521 | A521 | |
| Notification of reasons for refusalA131 | A131 |
Numbers
- Publication
- 3675239
- Publication, DOCDB
- 3675239
- Publication, EPODOC
- JP3675239B
- Application
- 19748699
- Application, DOCDB
- 19748699
- Application, EPODOC
- JP19990197486
Titles2
- Japanese
- 仮想空間におけるコンピュータを介在したマルチモーダルおよび非同期な会議を行なう方法およびシステム
- English
- Computer-mediated multimodal and asynchronous conferencing methods and systems in virtual space
Classification
- CPC, 3
- H04L12/1831
- G06Q10/109
- H04L12/1822
- IPC, 3
- H04N7 15
- G06Q10 10
- H04L12 18