Voice communication system
7 claims: 3 independent, 4 dependent
- 1仮想空間を用いた複数のユーザの会話を実現する音声コミュニケーション・システムであって、 前記複数のユーザ各々の実空間上の位置を管理するサーバ装置と、前記複数のユーザが使用する複数のクライアント端末とを有し、 前記複数のクライアント端末各々は、 自クライアント端末の自ユーザの実空間上の位置に関する位置情報を検知する位置検知手段と、 前記検知手段が検知した自ユーザの実空間上の位置情報を前記サーバ装置に送信するクライアント送信手段と、 前記サーバ装置から自ユーザ以外のユーザである他ユーザ各々の実空間上の位置に関する位置情報を受信するクライアント受信手段と、 前記自ユーザおよび前記他ユーザ各々の実空間における前記位置情報に基づいて前記複数のユーザ各々の前記仮想空間における位置を算出する空間モデル化手段と、 前記空間モデル化手段が算出した位置に基づいて前記他ユーザの各々の音声に適用する音響効果を制御する音響制御手段と、を有し、 前記サーバ装置は、 前記複数のクライアント端末各々から、前記クライアント端末の自ユーザの実空間上の前記位置情報を受信するサーバ受信手段と、 前記サーバ受信手段が受信した前記複数のユーザ各々の実空間上の前記位置情報を記憶する記憶手段と、 前記複数のクライアント各々に、前記記憶手段が記憶している前記クライアント端末の他ユーザ各々の前記位置情報を送信するサーバ送信手段と、を有し、 前記位置検知手段は、実空間において前記自ユーザが向いている方向を、さらに検知し、 前記位置情報には、実空間における前記自ユーザまたは前記他ユーザの向きを示す方位情報が含まれ、 前記複数のクライアント端末各々の前記空間モデル化手段は、前記自ユーザを前記仮想空間の中心に配置し、前記自ユーザおよび前記他ユーザの前記位置情報から算出される前記自ユーザと前記他ユーザ各々との実空間上での距離および方向に応じて、前記他ユーザ各々の前記仮想空間における位置を算出し、 前記音響制御手段は、実空間における前記自ユーザまたは前記他ユーザの前記方位情報に基づいて前記他ユーザ各々の音声に適用する音響効果を制御すること を特徴とする音声コミュニケーション・システム。
- 2請求項1記載の音声コミュニケーション・システムにおいて、 前記音響再生手段は、前記複数のユーザ各々の前記仮想空間における位置および前記仮想空間の属性情報に基づいて、前記他ユーザ各々の音声に適用する音響効果を制御すること を特徴とする音声コミュニケーション・システム。
- 3請求項1記載の音声コミュニケーション・システムにおいて、 前記複数のクライアント端末各々は、前記空間モデル化手段が算出した位置に基づいて表示画面に出力するイメージデータを作成するイメージ作成手段を有すること を特徴とする音声コミュニケーション・システム。
- 4請求項3記載の音声コミュニケーション・システムにおいて、 前記イメージ作成手段は、前記仮想空間における自ユーザの位置と向きを常に固定し、自ユーザを中心として前記仮想空間および前記他ユーザを相対的に移動または回転させたイメージデータを作成すること を特徴とする音声コミュニケーション・システム。
- 5請求項1記載の音声コミュニケーション・システムにおいて、 前記サーバ装置の前記記憶手段には、前記仮想空間の属性が記憶され、 前記サーバ送信手段は、前記複数のクライアント各々に、前記仮想空間の属性を送信し、 前記クライアント受信手段は、前記サーバ装置から前記仮想空間の属性を受信し、 前記空間モデル化手段は、前記仮想空間の属性に基づいて、前記複数のユーザ各々の前記仮想空間における位置を算出し、 前記音響制御手段は、前記空間モデル化手段が算出した位置に基づいて前記他ユーザの各々の音声に適用する音響効果を制御すること を特徴とする音声コミュニケーション・システム。
- 6仮想空間を用いた複数のユーザの会話を実現する音声コミュニケーション・システムにおける、前記ユーザが使用するクライアント端末であって、 自クライアント端末の自ユーザの実空間上の位置に関する位置情報を検知する位置検知手段と、 前記検知手段が検知した自ユーザの実空間上の位置情報を、前記複数のユーザ各々の実空間上の位置を管理するサーバ装置に送信する送信手段と、 前記サーバ装置から自ユーザ以外のユーザである他ユーザ各々の実空間上の位置に関する位置情報を受信する受信手段と、 前記自ユーザおよび前記他ユーザ各々の実空間における前記位置情報に基づいて前記複数のユーザ各々の前記仮想空間における位置を算出する空間モデル化手段と、 前記空間モデル化手段が算出した位置に基づいて前記他ユーザの各々の音声に適用する音響効果を制御する音響制御手段と、を有し、 前記位置検知手段は、実空間において前記自ユーザが向いている方向を、さらに検知し、 前記位置情報には、実空間における前記自ユーザまたは前記他ユーザの向きを示す方位情報が含まれ、 前記空間モデル化手段は、前記自ユーザを前記仮想空間の中心に配置し、前記自ユーザおよび前記他ユーザの前記位置情報から算出される前記自ユーザと前記他ユーザ各々との実空間上での距離および方向に応じて、前記他ユーザ各々の前記仮想空間における位置を算出し、 前記音響制御手段は、実空間における前記自ユーザまたは前記他ユーザの前記方位情報に基づいて前記他ユーザ各々の音声に適用する音響効果を制御すること を特徴とするクライアント端末。
- 7仮想空間を用いて複数のユーザが複数のクライアント端末を用いて会話を実現する音声コミュニケーション・システムにおける音響サーバ装置であって、 前記複数のクライアント端末各々から、前記クライアント端末のユーザの音声を受信する音声受信手段と、 外部システムから、前記複数のクライアント端末の複数のユーザの実空間上の前記位置情報を受信し、当該複数のユーザ各々の位置情報に基づいて、前記複数のユーザ各々の前記仮想空間における位置を算出する空間モデル化手段と、 前記空間モデル化手段が算出した位置に基づいて、前記複数のクライアント毎に、前記複数のユーザの各々の音声に適用する音響効果を制御する音響制御手段と、 前記複数のクライアント毎に、前記音響制御手段が制御した複数のユーザの音声を送信する音声送信手段と、を有し、 前記位置情報には、実空間における前記複数のクライアント端末の各自ユーザまたは 前記 自ユーザ以外のユーザである他ユーザの向きを示す方位情報が含まれ、 前記空間モデル化手段は、前記複数のクライアント端末の各自ユーザを前記仮想空間の中心に配置し、前記自ユーザおよび前記他ユーザの前記位置情報から算出される前記自ユーザと前記他ユーザ各々との実空間上での距離および方向に応じて、前記他ユーザ各々の前記仮想空間における位置を算出し、 前記音響制御手段は、実空間における前記自ユーザまたは前記他ユーザの前記方位情報に基づいて前記他ユーザ各々の音声に適用する音響効果を制御すること を特徴とする音響サーバ装置。
Independent claims7
96 paragraphs, as filed
The present invention relates to a technique for humans to talk with each other mainly by using voice through media.
Patent Document 1 discloses a navigation system that uses GPS technology to display the relative location information of a mobile phone calling user based on the location information of the calling user and the location information of the communication partner. ..
In addition, as a conference system using virtual space, there is a conference system FreeWalk developed at Kyoto University (see, for example, Non-Patent Document 1 and Non-Patent Document 2). Freewalk is a system in which users of a conference system share a virtual space and users in the same space can have a conversation. Each user can see the image of this virtual space from his / her own viewpoint, or a viewpoint close to it but also within his / her own field of view, by 3D graphics. 3D graphics technology is a technology that simulates 3D space by computer graphics, and API (Application Programming) that realizes it. Interface) includes the industry standard OpenGL (http://www.opengl.org/) and Microsoft's Direct3D. The image of the conversation partner is shot by a video camera and projected in real time on a virtual screen placed in the image that can be seen from one's own point of view. In addition, each user can move freely in this virtual space. That is, it is possible to change its position in this virtual space using a pointing device or keyboard keys. In Non-Patent Documents 1 and 2, the sound is attenuated according to the distance, but the three-dimensional audio technology described later is not used.
There is also a conferencing system Somewire developed by Interval Research Corporation (see, for example, Patent Document 2, Patent Document 3 and Non-Patent Document 3). Somewire is a system in which users of a conference system share a virtual space and users in the same space can talk to each other. At Somewire, audio is played with high quality stereo audio. It also has an intuitive physical (tangible) interface rather than a GUI (graphical user interface) that makes it possible to control the position of the conversation partner in virtual space by moving something like a doll. In Somewire, the sound is not attenuated according to the distance, and 3D audio technology is not used.
There is also a conference system using 3D distributed audio technology developed by Hewlett-Packard Company (see, for example, Non-Patent Document 4). 3D distributed audio technology is a technology that applies 3D audio technology in a system connected by a network (so-called distributed environment). And 3D audio technology is a technology that simulates 3D acoustic space, and the API for realizing this is Open AL (http: //), which is an industry standard defined by Loki Entertainment Software and others. www.opengl.org/), Microsoft DirectSound 3D, Creative Technology EAX 2.0 (http://www.sei.com/algorithms/eax20.pdf) and so on. By using this 3D audio technology, it is possible to simulate the direction and distance of the sound source as seen from the listener and localize the sound source in the acoustic space in the sound reproduction by speakers such as headphones, 2 channels or 4 channels. it can. In addition, by simulating acoustic attributes such as reverberation, reflection by objects such as walls, sound absorption depending on the distance by air, and sound interruption by obstacles, the presence of the room and the presence of objects in the space You can express a feeling.
<patcit num="1"><text>JP 2002-236031</text></patcit><patcit num="2"><text>US 5,889,843</text></patcit><patcit num="3"><text>US 6,262,711 B1</text></patcit><nplcit num="1"><text>Hideyuki Nakanishi, Riki Yoshida, Toshikazu Nishimura, Toru Ishida, "Free Walk: Support for Informal Communication Using 3D Virtual Space", IPSJ Journal, Vol.39, No.5, pp.1356-1364 , 1998.</text></nplcit>
<nplcit num="2"><text>Nakanishi, H., Yoshida, C., Nishimura, T., and Ishida, T., "Free Walk: A 3D Virtual Space for Casual Meetings", IEEE MultiMedia, April-June 1999, pp.2028.</text></nplcit><nplcit num="3"><text>Singer, A., Hindus, D., Stifelman, L., and White, S., "Tangible Progress: Less Is More In Somewire Audio Spaces", ACM CHI '99 (Conference on Human Factors in Computing Systems), pp. 104-112, May 1999.</text></nplcit><nplcit num="4"><text>Low, C., and Babarit, L., "Distributed 3D Audio Rendering", 7th International World Wide Web Conference (WWW7), 1998, http://www7.scu.edu.au/programme/fullpapers/1912/com1912. htm.</text></nplcit>
<p> By the way, even when the communication partner of the mobile phone exists in a place close to oneself (a position where the figure can be seen), it may be difficult to find it. For example, in a crowded amusement park or a train station in the center of the city, it is difficult to find the communication partner in the crowd and approach it even if you are talking via a mobile phone at a distance where you can see the communication partner. .. In addition, at construction sites, it may be necessary to grasp the work place (arrangement) of invisible collaborators.</p><p> In addition, when the communication partner in the virtual space (the other party communicating through the media) exists nearby in the real space, the media sound by the 3D audio technology in the virtual space of the communication partner and the media sound in the real space The direct sound may be heard from a different direction or distance. This causes inconveniences such as responding to a call from a communication partner that exists nearby in the real space in a different direction.</p><p> In Patent Document 1, the position of the other party is displayed on the map, but the recognition of the position of the other party through voice is not considered. Further, in the conference systems described in Patent Documents 2 and 3 and Non-Patent Documents 1 to 4, the position of the communication partner in the real space is not considered.</p><p> The present invention has been made in consideration of the above circumstances, and an object of the present invention is a voice capable of associating a real space with a virtual space and grasping the relative position and direction of a communication partner in the real space as a physical sensation. To provide a communication system.</p>
<p> In order to solve the above problems, in the present invention, the positions of each of the plurality of users in the virtual space are calculated based on the position information of each of the plurality of users in the real space.</p><p> For example, a voice communication system that realizes conversations between a plurality of users using a virtual space, a server device that manages the positions of the plurality of users in the real space, and a plurality of client terminals used by the plurality of users. And have. Each of the plurality of client terminals transmits the position detection means for detecting the position information regarding the position of the own user of the own client terminal in the real space and the position information of the own user detected by the detection means in the real space to the server device. Based on the client transmitting means, the client receiving means for receiving the position information about the position in the real space of each other user who is a user other than the own user from the server device, and the position information in the real space of the own user and each other user. It has a spatial modeling means for calculating a position in a virtual space of each of a plurality of users, and an acoustic control means for controlling an acoustic effect applied to each voice of another user based on the position calculated by the spatial modeling means. .. The server device stores the server receiving means for receiving the position information in the real space of the client terminal's own user from each of the plurality of client terminals and the position information in the real space for each of the plurality of users received by the server receiving means. The storage means is provided, and the server transmission means for transmitting the position information of each of the other users of the client terminal stored in the storage means to each of the plurality of clients.</p>
<p> According to the present invention, the relative position and direction of the communication partner in the real space can be easily grasped as a physical sensation by the voice (media sound) of the communication partner. Therefore, the user can have a natural conversation in the virtual space and the real space.</p>
Embodiments of the present invention will be described below.
FIG. 1 shows a system configuration diagram of a voice communication system to which one embodiment of the present invention is applied. As shown in the figure, this system consists of multiple clients 201, 202, 203, a presence server 110 that manages presence, a SIP proxy server 120 that controls sessions, and a registration server 130 that registers and authenticates users. , Is connected via a network 101 such as the Internet. Presence is the virtual space itself and the position information (presence) of each user in the virtual space.
Although the present embodiment has three clients, the number of clients is not limited to three, and may be two or four or more. Further, in the present embodiment, the network 101 is composed of a single domain, but the network is composed of a plurality of domains, and each domain can be combined to perform communication across a plurality of domains. In that case, there are a plurality of presence servers 110, SIP proxy servers 120, and registration servers 130.
Next, the hardware configuration of the voice communication system will be described.
FIG. 2 shows the hardware configurations of the clients 201, 202, 203, the presence server 110, the SIP proxy server 120, and the registration server 130.
The clients 201, 202, and 203 include a CPU 301 that processes and calculates data according to a program, a memory 302 that the CPU 301 can directly read and write, an external storage device 303 such as a hard disk, and a communication device for data communication with an external system. A general computer system having a 304, an input device 304, and an output device 306 can be used. For example, it is a portable computer system such as a PDA (Personal Digital Assistant), a wearable computer, and a PC (Personal Computer). The input device 305 and the output device 306 will be described later in FIG.
The presence server 110, SIP proxy server 120, and registration server 130 include a CPU 301 that processes and calculates data according to at least a program, a memory 302 that the CPU 301 can directly read and write, an external storage device 303 such as a hard disk, and an external system and data. A general computer system having a communication device 304 for communication can be used. Specifically, it is a server, a host computer, and the like.
Each function described later in each of the above devices is a predetermined program loaded or stored in the memory 302 (a program for a client in the case of clients 201, 202, 203, a program for a presence server in the case of a presence server 110). This is realized by the CPU 301 executing the program for the SIP proxy server in the case of the SIP proxy server 120 and the program for the registration server in the case of the registration server 130.
Next, the input device 305 and the output device 306 of the client 201 and the functional configuration will be described with reference to FIG. The same configuration is used for the clients 202 and 203.
The client 201 has a microphone 211, a camera 213, a GPS receiving device 231, a geomagnetic sensor 232, and an operation unit (not shown) as the input device 305. As the output device 306, it has a headphone 217 compatible with 3D audio technology and a display 220. GPS receiver 231 receives GPS signals from at least three GPS satellites. Then, the GPS receiver 231 measures the distance between the client 201 and the GPS satellites and the rate of change of the distance for at least three GPS satellites, and the current position of the user carrying the client 201 in the real space. Is calculated. The geomagnetic sensor 232 detects the magnetic field held by the earth, and calculates the orientation (direction) of the user carrying the client 201 in real space from the detection result. The geomagnetic sensor 232 may be a gyro that detects the angle at which the moving body rotates.
The functional configuration includes an audio encoder 212, an audio renderer 216, a video encoder 214, a graphics renderer 219, a spatial modeler 221, a presence provider 222, an audio communication unit 215, a video communication unit 218, and a session control unit 223. And have.
The audio encoder 212 converts audio into a digital signal. The audio renderer 216 uses 3D audio technology to perform processing that results from the attributes of the virtual space, such as reverberation and filtering. The video encoder 214 converts the image into a digital signal. The graphics renderer 219 performs processing resulting from the attributes of the virtual space. The space modeler 221 receives the position information and the direction information in the real space from the GPS receiver 231 and the geomagnetic sensor 232, and calculates the presence such as the position and orientation of the user in the virtual space. The presence provider 222 transmits and receives the user's position information and orientation information in the real space to and from the presence server 110. The audio communication unit 215 transmits and receives audio signals to and from other clients in real time. The video communication unit 218 transmits and receives video signals to and from other clients in real time. The session control unit 223 controls the communication session with other clients and the presence server 110 via the SIP proxy server 120.
Here, the virtual space is a space virtually created for a plurality of users to hold a meeting or a conversation, and is managed by the presence server 110. When a user enters a virtual space, the presence server 110 transmits the attributes of the virtual space and the position information and orientation information of other users existing in the virtual space in the real space. Then, the spatial modeler 221 transfers the transmitted information and the position information and the orientation information of the own user in the real space input from the GPS receiving device 231 and the geomagnetic sensor 232 to the memory 302 or the external storage device 303. Store. The attributes of the virtual space include, for example, the size of the space, the height of the ceiling, the reflectance / color / texture of the walls and the ceiling, the reverberation characteristics, and the absorption rate of sound by the air in the space. Of these, the reflectance of walls and ceilings, reverberation characteristics, and the absorption rate of sound by air in the space are auditory attributes, and the color and texture of walls and ceilings are visual attributes, and the size of the space. , Ceiling height is an attribute related to both hearing and vision.
Next, the operation of each function will be described in the order of presence, audio, and video.
Regarding the presence, the GPS receiver 231 and the geomagnetic sensor 232 calculate the position and orientation of the own user in the real space, and input the position information and the orientation information of the own user to the space modeler 221. The space modeler 221 stores the attributes of the virtual space (space size, reverberation characteristics, etc.) previously transmitted from the presence server 110 and the position information and orientation information of each other user in the virtual space in the real space 302. Alternatively, it is held in the external storage device 303. The space modeler 221 maps the real space and the virtual space from the attributes of the virtual space and the position information of other users and the own user in the real space. When the own user and a plurality of other users exist in the virtual space, the space modeler 221 arranges another user who is relatively close to the own user in the real space at a position relatively close to the own user in the virtual space. To do. Note that the mapping from the real space to the virtual space is a non-linear mapping (non-linear mapping) even if it is a linear mapping (linear mapping) in which the position information in the real space is scaled down to a position in the virtual space. You may. Non-linear mapping will be described below.
Figure 4 schematically shows an example of nonlinear mapping between real space and virtual space using arctan (x). In the non-linear mapping shown, real-space coordinates (position information) are used as a common coordinate system. In FIG. 4, a plane p perpendicular to the paper surface indicating the real space, a position u in the real space of the own user, and a position c in the real space of the third other user are shown. That is, the cutting line including u and c of the plane p is described on the paper surface (Fig. 4). Further, FIG. 4 shows a cross section of the sphere s tangent to the plane p indicating the virtual space of the own user and a cross section of the sphere q tangent to the plane p indicating the virtual space of the third other user. .. Then, it is assumed that the first other user exists at the point a of the plane p in the real space, and the second other user exists at the point b.
In this case, the spatial modeler 221 converts the distance d to another user into arctan (d / r) (r is a constant), that is, the length of the arc on the sphere s (a constant multiple). Specifically, the first other user who is at point a in real space (the distance from the own user in real space is the length of the line segment from u to a) is the point a'in virtual space. Map (place) to (so that the distance from your user is the length of the arc from u to a'). Similarly, in the space modeler 221, the second other user at the b point in the real space is set to the b'point in the virtual space, and the third other user at the c point in the real space is set to the c in the virtual space. 'Map (place) at a point. That is, the space modeler 221 transforms each point on the plane p, which is the real space, into s on the sphere, which is the virtual space. In the above description, it is assumed that all other users are on the above-mentioned cutting line due to space limitations (drawings). However, even if two or more other users do not exist on the same straight line including the own user, the mapping can be performed in the same manner in the three-dimensional space.
If another user exists at an infinity point (not shown) in the real space, the user is mapped (placed) at the d ́ point in the virtual space. By mapping infinity to a finite distance in this way, it is possible to have a conversation no matter how far away other users existing in the same virtual space are. The space modeler 221 maps each point a', b', c', and d'in a state where the upper hemisphere of the sphere s, which is a virtual space, is extended flat.
Further, the space modeler 221 holds the radius r (or a constant multiple of the radius r) of the sphere s, which is a virtual space, as a virtual space attribute in the memory 302 or the external storage device 303. Then, the space modeler 221 sets the sphere s which is a virtual space by using the radius r of the sphere s held in the memories 302 and 303. The radius r of the sphere s, which is a virtual space attribute, is managed by the presence server 110 and notified to the space modeler 221 of each client. That is, the radii r of the sphere s, which is the virtual space of all users existing in the same virtual space, are the same. As a result, each user's sense of distance can be matched.
Further, the sphere q is a virtual space of a third other user existing at point c in the real space. Like the space modeler 221 of the own user, the space modeler 221 of the third other user uses arctan (x) to move the own user who is at the u point in the real space to the u'' point in the virtual space. Map (place).
Then, the space modeler 221 sets the direction of each user by using the orientation information of each user mapped on the virtual space. If the direction of the geomagnetic sensor 232 and the direction of the user do not match (for example, if the mounting position of the geomagnetic sensor 232 is not fixed), or because of magnetic disturbance, the geomagnetic sensor 232 will accurately orient. If not instructed, the following operations may be performed. For example, a user can use a specific direction (for example, north) to accurately indicate the direction. And press the reset button on the operation unit 226 (see Fig. 8). The spatial modeler 221 receives the signal from the reset button and corrects the output from the geomagnetic sensor 232 so that the direction at that time is regarded as the specific direction described above. Further, instead of the correction based on the absolute orientation (specific orientation) as described above, a method of matching the direction in the real space of another user with the direction in the virtual space can be considered. For example, the user can correct the direction in the real space and the relative direction in the virtual space by pressing the reset button while facing the direction of another user in the vicinity. If multiple such correction methods are implemented in the client, the user first selects the method and then presses the reset button.
The spatial modeler 221 transmits the position information and the orientation information of the own user in the real space to the presence server 110 via the presence provider 222. Further, the spatial modeler 221 receives the position information and the orientation information of another user in the real space from the presence server 110 via the presence provider 222. That is, since the spatial modeler 221 receives the position information and the orientation information of another user in the real space via the network 101, delay and jitter occur with respect to the position and orientation of the other user in the virtual space. Is inevitable. On the other hand, since the position and orientation of the own user are directly input from the GPS receiver 231 and the geomagnetic sensor 232 to the space modeler 221, almost no delay occurs.
For voice, the microphone 211 collects the voice of the user using the client 201 and sends it to the audio encoder 212. Then, the audio encoder 212 converts the above-mentioned voice into a digital signal and outputs it to the audio renderer 216. Further, the audio communication unit 215 transmits and receives an audio signal in real time to and from another one or a plurality of clients, and outputs the audio signal to the audio renderer 216.
The digital output signal output from the audio encoder 212 and the audio communication unit 215 is input to the audio renderer 216. Then, the audio renderer 216 uses 3D audio technology to virtualize based on the auditory virtual space attributes held by the space modeler 221 and the positions of the own user and other users mapped on the virtual space. Calculate how the voice of another user (communication partner) can be heard in space. Hereinafter, the audio renderer 216 will be specifically described with reference to FIGS. 5 and 6.
FIG. 5 is a diagram schematically showing the direction and distance of a sound source that is a communication partner (other user). In FIG. 5, a human head 1 showing a person from directly above and a sound source 2 as a communication partner are shown. Human head 1 has a nose 11 to indicate orientation. That is, the human head 1 faces the direction 3 in which the nose 11 is added. In 3D audio technology, HRIR (Head Related Impulse Response), which expresses how the sound changes around the head 1 (impulse response), and pseudo-reverberation generated by a virtual environment such as a room are used. Represents the direction and distance of sound. And HRIR is the distance 4 between the sound source 2 and the head 1 and the angle (horizontal angle and vertical angle) 5 between the head 1 and the sound source. Determined by. It is assumed that the memory 302 or the external storage device 303 stores the HRIR value measured in advance for each distance and each angle using a dummy head (human head 1). In addition, by using different values for the left channel (measured with the left ear of the dummy head) and for the right channel (measured with the right ear of the dummy head) for the HRIR values, left and right, Express the sense of direction in the front-back or up-down direction.
FIG. 6 is a diagram showing the processing of the audio renderer 216. The audio renderer 216 performs the following calculations for each sound source (other user) for each packet (usually every 20 ms) received by RTP (Real-time Transport Protocol). As shown, the audio renderer 216 has a signal sequence s for each sound source.<sub>i</sub>[t] (t = 1, ...) and the coordinates of the sound source in virtual space (x)<sub>i</sub>, y<sub>i</sub>) Is accepted (S61). The coordinates of each sound source in the virtual space are input from the space modeler 221. The space modeler 221 maps (arranges) each sound source (other user) in the virtual space, and then inputs the coordinates of each sound source (position information in the virtual space) to the audio renderer 216. Further, the signal sequence of each sound source is input from the audio communication unit 215.
Then, the audio renderer 216 calculates the distance and azimuth between the own user and the sound source for each sound source using the input coordinates (S62). It is assumed that the own user exists at the center of the virtual space (coordinates (0,0)). Then, the audio renderer 216 identifies the HRIR corresponding to the distance and the angle (azimuth) with the own user from the HRIR numerical values stored in the memory 302 or the external storage device 303 (S63). The audio renderer 216 may use the HRIR value calculated by interpolating the HRIR value stored in the memory 302 or the like.
Then, the audio renderer 216 performs a convolution calculation using the signal sequence input in S61 and the HRIR for the left channel of the HRIR specified in S63, and generates a left channel signal (S64). Then, the audio renderer 216 adds all the left channel signals from each sound source (S65). In addition, the audio renderer 216 performs a convolution calculation using the signal sequence input in S61 and the HRIR for the right channel of the HRIR specified in S63, and generates a right channel signal (S66). Then, the audio renderer 216 adds all the right channel signals from each sound source (S67).
Next, the audio renderer 216 adds reverberation to the signal on the left channel after addition (S68). That is, the audio renderer 216 calculates the reverberation based on how the sound changes depending on the attributes of the virtual space (impulse response). There are two types of reverberation calculation: FIR (finite impulse response) and IIR (infinite impulse response). Since these calculation methods are basic methods related to digital filters, description thereof will be omitted here. In addition, the audio renderer 216 adds reverberation to the signal of the right channel after addition in the same manner as the signal of the left channel (S69). The HRIR identification (S63) and the reverberation calculation (S68, S69) are performed for each packet as described above, but in the convolution calculation (S64, S66), there is a part to be carried over to the next packet. Therefore, it is necessary to retain the specified HRIR or input signal sequence until the processing of the next packet.
In this way, the audio renderer 216 performs processing such as adjusting the volume by the above calculation, superimposing reverberation and reverberation sound, filtering, etc. on the voice of the communication partner user output from the audio communication unit 215, and the own user. Control the sound effect to the sound that should be heard at the position in the virtual space of. That is, the voice is localized and reproduced by the process resulting from the relative position between the attribute of the virtual space and the communication partner. As a result, it is possible to easily grasp the direction in which the communication partner who cannot directly hear the voice is present.
The audio renderer 216 uses the client 201 after performing processing resulting from the attributes of the virtual space such as reverberation and filtering on the own user voice output from the audio encoder 212, if necessary. It may be rendered at the position of the user's head. The voice of the own user generated by the audio renderer 216 is output to the headphone 217, and the own user listens to the voice. That is, if the direct sound of the own user's voice is heard by the own user, a strange impression may be given, and especially if the delay is large, the own voice is hindered. Therefore, the own user's own voice is usually given to the own user. Don't let me hear. However, it is also possible not to hear the direct sound, but to hear only the reverberation with the delay within the range of several tens of ms. This makes it possible to grasp the physical sensation regarding the position of the own user in the virtual space and the size of the virtual space.
As for the image, the camera 213 captures the user's head and continuously sends the captured image to the video encoder 214. Then, the video encoder 214 converts the image into a digital signal and outputs it to the graphics renderer 219. In addition, the video communication unit 218 transmits and receives a video signal in real time to and from another one or a plurality of clients, and outputs the video signal to the graphics renderer 219. Next, the graphics renderer 219 inputs digital output signals from the video encoder 214 and the video communication unit 218.
Then, the graphics renderer 219 calculates how the communication partner can be seen in the virtual space based on the visual virtual space attribute held by the space modeler 221, the position of the communication partner in the virtual space, and one's own position (. Coordinate conversion). Next, the graphics renderer 219 performs a process on the image of the user of the communication partner output from the video communication unit 218, which results from the attributes of the virtual space from the viewpoint viewed from its own position by the above calculation, and is displayed on the screen. Create image data to be output to. The video generated by the graphics renderer 219 is output to the display 220 and reproduced as a video from the viewpoint of the user who uses the client 201, and the user refers to the output of the display 220 as necessary.
FIG. 7 is an example of the virtual space displayed on the display 220. The display content shown in FIG. 4 is an example in which the local user who uses the client 201 shares a virtual space with the first and second other users who use the client 202 and the client 203. .. In the illustrated example, the virtual space is displayed in a plan view. A client 201 placed in the virtual space from directly above based on the attributes of the virtual space stored in the memory 302 or the external storage device 303 by the space modeler 221, its position in the virtual space, and information of other users. A two-dimensional image obtained by looking at the own avatar 411 representing the own user of the user, the first other avatar 412 representing the user of the communication partner, and the second other avatar 413 is displayed. The graphics renderer 219 fixes the position and orientation of the own user of the client 201, and displays the client 201 so that the virtual space and other users in the virtual space move and rotate relative to the local user. When the user moves or changes direction in the real space, the space modeler 221 maps the virtual space in response to the input from the GPS receiver 231 or the geomagnetic sensor 232, so that the virtual space or the virtual space is changed. The screen that other users have moved / rotated relatively is displayed in real time. Further, in the illustrated example, the directional information 420 indicating the north is displayed.
This makes it possible to express the positional relationship between the own user and another user (clients 202, 203) who is the communication partner in the virtual space. Further, by fixing the direction of the own user to the front, the consistency between the voice and the graphics display can be ensured, and the position and direction of the other user can be grasped as a physical sensation. Further, since other users existing behind the own user can also be displayed, there is an advantage that there is little risk of overlooking other users approaching from behind.
Although not shown, by displaying the scale on the display 220, the distance in the virtual space from other users can be accurately expressed. For example, it is conceivable to make it possible to select a scale from a plurality of scale candidates by using a radio button or the like, or to make it possible to continuously change the scale by using a scroll bar slider. By operating these buttons and the scroll bar slider, the scale of the displayed plan view is changed immediately so that you can check the distant situation and the position of your user in the room (in the virtual space). Can be confirmed, or the vicinity can be observed in more detail.
Although not shown, the image of the own user taken by the camera 213 of the client 201 is on the avatar 411, the image of the first other user taken by the camera 213 of the client 202 is on the avatar 412, and the image of the first other user is on the camera 213 of the client 203. The second user's video taken by the camera is pasted on the avatar 413 by a texture map. When the user with whom you are communicating rotates, the texture also rotates, so you can see which orientation the first and second users are facing in the virtual space.
For real-time communication of voice or images, RTP (Real-time Transport Protocol), which is a protocol described in the document RFC 3550 issued by the IETF (Internet Engineering Task Force), is used. If a slight increase in delay is allowed in voice or image communication, a communication proxy server that performs voice or image communication is provided separately for communication between the audio communication unit 215 or video communication unit 218 and other clients. , It is also possible to perform audio or image communication with other clients via this communication proxy server.
This is the end of the description of the client 201 in FIG. Among the clients 201, the microphone 211, the camera 213, the GPS receiver 231 and the geomagnetic sensor 232, the headphones 217 and the display 220 are realized by hardware. In addition, the audio encoder 212 and the video encoder 214 are realized by software, hardware, or a combination thereof. In addition, the audio communication unit 215, the video communication unit 218, the spatial modeler 221 and the session control unit 223 are usually realized by software.
Next, with reference to FIG. 8, the types of clients 201, 202, and 203 are illustrated.
The client shown in Figure 8 (a) has a size and functionality similar to that of a PDA or handheld computer. The client body 230 includes a camera 213, a display 220, an operation unit 226, an antenna 237, and a GPS receiver 231. Further, the headset connected to the main body 230 has a headphone 217, a microphone 211 and a geomagnetic sensor 232. By installing the geomagnetic sensor 232 inside the headphone 217 (upper part of the headband, etc.), the geomagnetic sensor 232 can be worn at a substantially constant angle (facing forward) to the user at all times. The operation unit 226 has instruction buttons 241 to 245 for inputting various instructions to the client 201. The instruction buttons 241 to 245 include a reset button for aligning the direction of the geomagnetic sensor 232 installed on the headphone 217 when the headset is worn. Further, although the illustrated headset is connected to the main body 230 by wire, it can also be wirelessly connected by Bluetooth, IrDA (infrared data), or the like. Further, the client connects to the network 101 by wireless LAN using the antenna 237.
The client shown in Figure 8 (b) shows an example of a wearable computer. The client body 241 like the vine of glasses has a microphone 211, a camera 213, headphones 217, a display 220, a GPS receiver 231 and a geomagnetic sensor 232. The display 220 is a head-mounted display, and forms a virtual image in front of the user who wears the client body 241 by several tens of centimeters, or forms a three-dimensional image in front of the user. The client of FIG. 8B has an operation unit 226 (not shown) connected by wire or wirelessly.
Next, the processing procedure in the client 201 will be described with reference to FIGS. 9 to 12.
Figure 9 shows the processing procedure when connecting the client 201 to the network 101. The illustrated connection procedure is performed when the client 201 is powered on. First, the session control unit 223 sends a login message including the user's identification information and the authentication information to the SIP proxy server 120 (S901). The SIP proxy server 120 accepts the login message and sends the user's authentication request message to the registration server 130. Then, the registration server 130 authenticates the user's identification information and the authentication information, and sends the user's identification information to the presence server 110. For communication between the client and the registration server 130, it is conceivable to use the REGISTER message of the protocol SIP (Session Initiation Protocol) specified in the IETF document RFC 3261. The client periodically sends a REGISTER message to the registration server 130 via the SIP proxy server 120.
In addition, the SIP SUBSCRIBE message described in the IETF document RFC 3265 can be used to communicate between the presence provider 222 of client 201 and the presence server 110. The SUBSCRIBE message is an event request message that requests to be notified in advance when an event occurs. The presence provider 222 requests the presence server 110 to notify the event that occurred regarding the room list and the attendee list of the virtual space managed by the presence server 110. When using the SUBSCRIBE message, the presence provider 222 communicates with the presence server 110 via the session control unit 223 and the SIP proxy server 120.
The presence provider 222 then receives a room list from the presence server 110 (S902). When the SUBSCRIBE message is used in S901, the room list is sent using the NOTIFY message as the event notification message. Presence provider 222 then displays the received room list on display 220 (S903).
FIG. 10 shows the processing procedure of the client 201 when the user selects the room he / she wants to enter from the room list displayed on the display 220. The presence provider 222 of the client 201 accepts the room selection instruction input using the operation unit 226 (S1001). The presence provider 222 then sends an admission message (enter) to the presence server 110 (S1002). The admission message includes the identification information of the own user and the position information and the orientation information of the own user in the real space. The position information and orientation information of the own user are calculated by the GPS receiving device 321 and the geomagnetic sensor 322 and input to the space modeler 221. Then, the spatial modeler 221 stores the input position information and orientation information in the memory 302 or the external storage device 303. The presence provider 222 reads the position information and the orientation information stored in the memory 302 or the external storage device 303, includes them in the admission message, and transmits the information.
You can also use the SIP SUBSCRIBE message to send the admission message. That is, the SUBSCRIBE message with the selected room as the recipient is used as the admission message. The SUBSCRIBE message requests notification of events that occur in the virtual space of the selected room (eg, user entry / exit or movement, change of virtual space attributes, etc.).
Next, the presence provider 222 receives the attendee list of other users who are admitted to the room selected from the presence server 110 (S1003). If you use the SUBSCRIBE message as the admission message, the attendee list is sent to the presence provider 222 in the corresponding NOTIFY message format. It should be noted that the attendee list includes at least the identification information of other users who have entered the room, the position information and the orientation information in the real space, and the virtual space attribute of the designated room. To do. The virtual space attribute includes the radius r of the sphere s, which is the virtual space shown in FIG. 4, or a constant multiple of the radius r (hereinafter, virtual space radius, etc.).
Although not shown, the processing procedure when the user leaves the room is not shown, but the presence provider 222 sends the exit message including the user identification information to the presence server 110 in response to the user's exit instruction.
FIG. 11 shows a processing procedure when the user changes the presence, that is, when the user moves in the real space. First, the spatial modeler 221 receives input of position information and direction information (hereinafter, position information, etc.) from the GPS receiver 231 and the geomagnetic sensor 232 (S1101). Then, the space modeler 221 compares the position information or the like stored in the memory 302 or the external storage device 303 (hereinafter, memory or the like) with the position information or the like received by the S711, and whether or not they are different. Is determined (S1102). The memory or the like stores the position information or the like previously input from the GPS receiver 231 and the geomagnetic sensor 232.
When the received position information, etc. is the same as the position information, etc. stored in the memory, that is, when the own user does not move in the real space and the orientation does not change (S1102: NO), the space modeler 221 It returns to S1101 without performing the subsequent processing.
When the received position information or the like is different from the position information or the like stored in the memory or the like, that is, when the own user moves or changes the direction in the real space (S1102: YES), the space modeler 221 is set to the received position. Store information etc. in memory etc. Then, the space modeler 221 changes the mapping of the virtual space or the direction of the own user by using the position information after the movement (S1103). The virtual space mapping is a non-linear mapping between the real space and the virtual space described in FIG. The space modeler 221 arranges its own user in the center of the virtual space, and non-linearly rearranges the positions of other users existing in the same virtual space.
Next, the space modeler 221 provides the position information of the virtual space after movement to the audio renderer 216, the graphics renderer 219, and the presence provider 222. Notify (S1104). As explained in FIG. 6, the audio renderer 216 can hear the voice of another user who is the communication partner at the position and orientation of the own user in the virtual space mapped based on the position information in the real space. To calculate. Then, the audio renderer 216 performs processing such as volume adjustment, reverberation, and filtering based on the above calculation on the voice of another user of the communication partner output from the audio communication unit 215, and virtualizes the own user who uses the client 201. It controls the sound effect to the sound that should be heard at the position in the space, and updates the 3D sound. In addition, the graphics renderer 219 changes the viewpoint based on the position of the own user in the virtual space mapped based on the position information in the real space and the orientation of the own user, and how the communication partner can be seen in the virtual space. Is calculated (coordinate conversion) (see Fig. 7). Then, the graphics renderer 219 creates image data to be output on the screen in view from the position and orientation, and updates the display screen.
Next, the presence provider 222 notifies the presence server 110 of the position information in the real space after the movement (S1105). If you use the SIP protocol, use the NOTIFY message. Note that the NOTIFY message is usually sent as a result of receiving the SUBSCRIBE message. Therefore, when the presence server 110 receives the admission message from the client 201, it is conceivable to return the attendee list and send the SUBSCRIBE message corresponding to the NOTIFY message. The presence server 110 receives the location information and the like in the real space notified by the presence provider 222, and updates the location information and the like of the user in the attendee list.
FIG. 12 shows a change input of presence, that is, a processing procedure when the presence server 110 notifies the client 201 of the position information in the real space of another user.
The spatial modeler 221 receives the location information and the like in the real space of another user of another client from the presence server 110 via the presence provider 222 (S1201). The presence server 110 notifies (transmits) the location information and the like transmitted from the client 201 in S1105 of FIG. 11 to clients other than the client of the transmission source. Then, the space modeler 221 stores the notified real space position information and the like in a memory or the like in a storage unit. Then, the space modeler 221 maps another user on the virtual space or changes the direction of the other user by using the notified position information or the like (see FIG. 4). Then, the space modeler 221 notifies the audio renderer 216 and the graphics renderer 219 of the position information of the virtual space after the movement (S1203). The audio renderer 216 and the graphics renderer 219 update the 3D sound and display screen of the other user based on the notified position and orientation of the other user, as described in S1104 of FIG.
Next, the functional configuration and processing procedure of the presence server 110 will be described. Since the registration server 130 and the SIP proxy server 120 are the same as the conventional communication using SIP, the description thereof will be omitted.
FIG. 13 shows the functional configuration of the presence server 110. The presence server 110 includes an interface unit 111 for sending and receiving various information to and from the client, a determination unit 112 for determining the message type from the client, a processing unit 113 for performing processing according to the determination result, and attributes of the virtual space. It has a storage unit 114 that manages and stores events (user entry / exit, movement, etc.), room list, visitor list, etc. that occur in the virtual space. The storage unit 114 stores in advance some attributes of the virtual space managed by the presence server 110. As mentioned above, the user selects the virtual space he / she wants to enter from these virtual spaces (see FIGS. 9 and 10). After that, the client sends various events of the user who entered the virtual space to the presence server 110. As a result, various events occur in each virtual space. The storage unit 114 stores these information in the memory 302 or the external storage device 303.
FIG. 14 shows the processing procedure of the presence server 110. The presence server 110 receives a request from the client and performs processing for the request until the presence server 110 stops. First, the interface unit 111 waits for a message from the client (S1411). Upon receiving the message, the determination unit 112 determines the type of message received by the interface unit 111 (S1412).
If the message is a login message, the processing unit 113 instructs the interface unit 111 to send the room list to the client that sent the message (S1421). The interface unit 111 sends the room list to the client that sent the message, then returns to S1411, and waits for the next message.
If the message is an admission message, processing 113 adds the user of the message source client to the attendee list for the specified room (S1431). That is, the processing unit 113 adds the identification information of the user and the position information and the orientation information of the user in the real space, which are included in the admission message, to the attendee list. Next, the processing unit 113 causes the interface unit 111 to transmit the identification information of all the visitors (however, other than the user) of the designated room, the position information in the real space, and the orientation information to the message transmission source client. Instruct. Further, the processing unit 113 instructs the interface unit 111 to transmit the virtual space attribute of the designated room to the message source client. The virtual space attribute includes the radius r of the sphere s, which is the virtual space shown in FIG. 4, or a constant multiple of the radius r (hereinafter, virtual space radius, etc.). The interface unit 111 transmits to the source client according to the above instruction (S1432). Then proceed to S1436, which will be described later.
In the case of a moving message, the processing unit 113 updates the position information and the orientation information in the real space of the message source client (user) in the attendee list (S1435). The position information and the direction information in the real space are included in the movement message. Then, the processing unit 113 provides the client of all the visitors in the target room (excluding the message source client) with the identification information of the user of the message source client, and the position information and the orientation information in the real space. , Is instructed to the interface unit 111 (S1436). The interface unit 111 transmits to the client according to the above instruction, and returns to S1411. The same applies to the admission message (S1431).
In the case of an exit message, the processing unit 113 deletes the user of the message source client from the attendee list (S1441). Then, the processing unit 113 instructs the interface unit 111 to notify the clients of all the visitors of the target room (excluding the message sender client) that the user has left the room ( S1442). The interface unit 111 transmits to the client according to the above instruction, and returns to S1411.
Although not shown, the presence server 110 may change the virtual space attribute by receiving a request (input) from the administrator of the presence server 110. For example, the determination unit 112 receives a change instruction such as a virtual space radius input from the input means 305 of the presence server 110. The change instruction of the virtual space radius and the like includes identification information for identifying the room to be changed and the changed virtual space radius and the like. Then, the processing unit 113 changes the virtual space radius and the like of the room to be changed stored in the storage unit 114. Then, the processing unit 113 reads out the visitor list stored in the storage unit 114, and notifies the clients of all the users who have entered the room to be changed of the changed virtual space radius and the like. The spatial modeler 221 of the client receiving the notification maps each user in the real space on the sphere s such as the changed virtual space radius shown in FIG.
An embodiment of the present invention has been described above.
In the voice communication system of the present embodiment, each user is mapped on the virtual space based on the position and direction of each user in the real space. As a result, even if the communication partner is in a remote place where the voice (direct sound) cannot be heard in the real space, the relative position and direction of the communication partner can be physically sensed by the voice (media sound) of the communication partner. Can be easily grasped as. Therefore, even in the crowd, it is possible to easily find the communication partner and approach the communication partner.
Further, in the present embodiment, the directions in which the communication partners exist are the same in the real space and the virtual space. Therefore, even if there is a communication partner at a close distance where the sound in the real space (direct sound) can be heard, the sound in the real space (direct sound) and the sound in the virtual space (media sound) can be heard. You can't hear it from different directions. Therefore, there is no inconvenience such as answering a voice (media sound) call in the virtual space in a different direction.
The present invention is not limited to the above embodiment, and many modifications can be made within the scope of the gist thereof.
For example, the client 201 of the present embodiment has a camera 213, a video encoder 214, and the like, and outputs image data of a virtual space to a display 220. However, since the present invention is a voice communication system mainly for voice communication, the client 201 does not have to output the image data of the virtual space to the display 220. In this case, the client 201 does not have a camera 213, a video encoder 214, a display 220, or the like.
Further, the graphics renderer 219 of the present embodiment expresses a virtual space by using a plan view (two-dimensional data) (see FIG. 7). However, the graphics renderer 219 may use 3D graphics technology to provide a clearer virtual space display. That is, the size of the space stored in the memory 302 or the external storage device 303 by the space modeler 221, the attributes of the virtual space such as the material of the wall and the ceiling, the position and direction of the own user and other users in the virtual space, and the like. A two-dimensional image may be created from the three-dimensional data of the above and displayed on the display 220.
Further, the audio renderer 216 may perform the following processing on the voice (media sound) of the communication partner user output from the audio communication unit 215. For example, the audio renderer 216 filters media sounds with an impulse response that cannot be real voice (direct sound). Alternatively, the audio renderer 216 adds reverberation to the voice (media sound) of the user of the communication partner to recognize a sense of distance from the sound source, which is different from the actual voice (direct sound). Alternatively, the audio renderer 216 adds noise to the voice (media sound) of the user of the communication partner. This makes it easy to determine whether the voice of the communication partner is a direct sound or a media sound even when the user of the communication partner exists at a close distance where the actual voice (direct sound) can be heard in the real space. Can be determined.
In addition, when the communication partner exists at a distance where the actual voice (direct sound) can be heard in the real space, the actual voice (direct sound) of the communication partner and the voice (media sound) output from the audio communication unit 215 I can hear both. In this case, if the delay of the media sound is small, it is localized by the media sound, and conversely, if the delay of the media sound is too large, it is heard as an independent sound source unrelated to the direct sound, causing confusion. Therefore, when the communication partner exists at a predetermined close distance, the audio renderer 216 may control the delay time of the voice (media sound) of the communication partner within a certain range. If the delay of the media sound is larger than that of the direct sound and within a certain range, the media sound is heard as the reverberation (echo) of the direct sound, so that it can be localized by the direct sound and the occurrence of confusion can be prevented. Further, the audio renderer 216 may reduce the volume of the voice (media sound) of the communication partner existing at a predetermined close distance by a certain amount or a certain percentage. This makes it possible to balance the volume of a distant communication partner who can only hear the media sound.
In addition, it is conceivable to use Bluetooth, which is a wireless communication technology, to determine whether or not a communication partner exists at a predetermined close distance where direct sound can be heard in a real space. That is, when data can be transmitted / received by Bluetooth, it is determined that a communication partner exists at a predetermined close distance.
Further, the client of the present embodiment detects the position and orientation of the user (client) by using the GPS receiving device 231 and the geomagnetic sensor 232. However, a sensor net may be used to detect the position and orientation of the user (client). By using the sensor net, the position and orientation of the user can be detected even when the user uses the client indoors.
Further, in the present embodiment, each client directly performs voice communication to make the voice input from the other client three-dimensional (see FIG. 6). However, if the processing power and communication power of the client are low, the server may perform these processes. That is, it is conceivable to add a new acoustic server to the network configuration shown in FIG. An embodiment having an acoustic server will be described below.
FIG. 15 is a network configuration diagram of an embodiment having an acoustic server. The illustrated network configuration differs from the network configuration of FIG. 1 in that it has an acoustic server 140. Further, the clients 201, 202, and 203 are different from the client configuration shown in FIG. 3 in the following points. That is, the audio renderer 216 is a simple audio decoder that does not perform audio three-dimensional processing (see Fig. 6). Further, the audio communication unit 215 communicates with the sound server 140 instead of directly communicating with other clients.
FIG. 16 is a configuration diagram of the acoustic server 140. As shown in the figure, the acoustic server 140 has at least one audio receiving unit 141, an audio renderer 142, a mixer 143, and an audio transmitting unit 144, respectively. That is, it is assumed that the acoustic server 140 has as many processing units 141 to 144 as the number of clients (that is, for each client). The audio server 140 is realized by using one program or device in a time-division manner without having as many audio receivers 141, audio renderers 142, mixers 143 and audio transmitters 144 as there are clients. May be.
The acoustic server 140 also has a spatial modeler 145. The space modeler 145 receives the position of each user in the real space and the attributes of the virtual space (virtual space radius, etc.) from the presence server 110, and performs the same processing as the client space modeler 221 shown in FIG. 3, in the virtual space. Map (place) the position of each user above.
The audio receiving unit 141 receives the voice input from the audio communication unit 215 of each client. The audio renderer 142 performs three-dimensional audio, and outputs signal data (signal strings) of two channels (left channel and right channel) to each mixer 143 associated with each client in response to each client. To do. That is, the audio renderer 142 calculates the sound source input (FIG. 6: S61), distance and angle of the client audio renderer 216 shown in FIG. 3 based on the position of each user in the virtual space arranged by the spatial modeler 145. (S62), HRIR identification (S63), and convolution calculation (S64, S66) and the same processing are performed. The mixer 143 receives signal data of two channels from each audio renderer 142, and performs the same processing as the mixing process (S65, S67) and the reverberation calculation (S68, S69) of the client audio renderer 216 shown in FIG. Then, the mixer 143 outputs the signal data of two channels to the audio transmission unit 144. The audio transmitter 144 transmits this signal data to the client.
Next, the processing of the presence server 110 and the client will be described. When the presence server 110 notifies each client of the user name, the position of the user, the virtual space radius, etc. in S1432, S1436, and S1442 of FIG. Notify the position, virtual space radius, etc. As a result, each client performs voice communication with the default communication port of the acoustic server 140 (or with the port notified by the presence server 110 at the time of entry) when entering the room. That is, the audio communication unit 215 of each client transmits a 1-channel audio stream to the acoustic server 140, and receives a 2-channel audio stream from the acoustic server 140.
Next, the processing of the acoustic server 140 will be described. Each of the audio receivers 141 associated with each client receives the audio stream from each client and buffers the synchronized (corresponding) signal data between the audio streams from all input clients to the client. It is sent to the audio renderer 142 associated with each. The method of this buffering (playout buffering) is described in, for example, the following document. By Colin Perkins: RTP: Audio and Video for the Internet, Addison-Wesley Pub Co; 1st edition (June 11, 2003). Then, the audio renderer 142 calculates the distance / angle, identifies the HRIR, and performs the convolution calculation (Fig. 6: S62 to S64, S66) based on the position of each user in the virtual space arranged by the space modeler 145. Perform processing. Then, the mixer 143 performs mixing processing (FIGS. 6: S65 and S67) and reverberation calculation (FIG. 6: S68 and S69), and outputs signal data of two channels to each client. Then, the audio transmission unit 144 transmits this signal data to the corresponding client. As a result, even when the processing capacity of the client is low, it is possible to realize three-dimensional voice.
Further, the presence server 110 may have the function of the acoustic server 140 described above. That is, the presence server 110 may not only manage the user's position, virtual space attributes, etc., but also process the sound server 140 without separately providing the sound server 140.
<figref num="1">It is a network block diagram in this embodiment.</figref><figref num="2">It is a hardware block diagram of each apparatus in this Embodiment.</figref><figref num="3">It is a block diagram of the client in this embodiment.</figref><figref num="4">It is a figure which showed typically the mapping of the real space and the virtual space in this embodiment.</figref><figref num="5">It is a figure which showed typically the direction and the distance of the sound source in this embodiment.</figref><figref num="6">It is a figure which showed typically the process of the audio renderer in this embodiment.</figref><figref num="7">It is an example of a display display screen of a virtual space in this embodiment.</figref><figref num="8">This is an example of the types of clients in this embodiment.</figref><figref num="9">It is a connection processing flow diagram of a client to a network in this embodiment.</figref><figref num="10">It is an admission processing flow diagram of a client in this embodiment.</figref><figref num="11">It is a movement processing flow diagram of the own user of the client in this embodiment.</figref><figref num="12">It is a movement processing flow diagram of another user of a client in this embodiment.</figref><figref num="13">It is a functional block diagram of the presence server in this embodiment.</figref><figref num="14">It is a processing flow diagram which shows the processing procedure of the presence server in this embodiment.</figref><figref num="15">It is a network block diagram in an embodiment which has an acoustic server.</figref><figref num="16">It is a functional block diagram of the acoustic server in embodiment which has an acoustic server.</figref>
Code description
101 ... Network, 110 ... Presence Server, 120 ... SIP Proxy Server, 130 ... Registration Server, 201, 202, 203 ... Client, 211 ... Microphone, 212 ... Audio Encoder , 213 ... Camera, 214 ... Video Encoder, 215 ... Audio Communication Unit, 216 ... Audio Renderer, 217 ... Headphones, 218 ... Video Communication Unit, 219 ... Graphics Renderer, 220 ... display, 221 ... spatial modeler, 222 ... presence provider, 223 ... session control, 231 ... GPS receiver, 232 ... geomagnetic sensor
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2002281468A | Cites | Japan |
| JP10056626A | Cites | Japan |
| JP2003287426A | Cites | Japan |
| JP2001251698A | Cites | Japan |
| JP2003069968A | Cites | Japan |
5 members in 3 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004155733 | Japan | A | |
| JP20040155733 | – | – | – |
Members5
| Document | Office | Kind | |
|---|---|---|---|
| CN1703065A | China | A | |
| US2005265535A1 | United States of America | A1 | |
| JP2005341092A | Japan | A | |
| US7634073B2 | United States of America | B2 | |
| JP4546151B2This record | Japan | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4546151
- Publication, DOCDB
- 4546151
- Publication, EPODOC
- JP4546151B
- Application
- 155733
- Application, DOCDB
- 2004155733
- Application, EPODOC
- JP20040155733
Titles2
- Japanese
- 音声コミュニケーション・システム
- English
- Voice communication system
Classification
- CPC, 5
- H04M3/567
- H04M3/42093
- H04M3/42365
- H04M2242/30
- H04M3/568
- IPC, 9
- H04W4 06
- H04W4 02
- H04W88 02
- H04W4 16
- G06F15 16
- H04B7 26
- H04M3 42
- H04M3 56
- H04W64 00
