Method and apparatus for teleconference
11 claims: 3 independent, 8 dependent
- 1テレカンファレンスの方法であって、第1デバイスのプロセッシング回路によって、第2デバイスから、第1オーディオを運ぶ第1メディアストリームと、第2オーディオを運ぶ第2メディアストリームとを受信するステップであり、前記第2デバイスは、メディアストリームを夫々供給する複数のデバイスからユーザによる選択又はアクティブスピーカの検出に応答して選択されたいずれか1つのデバイスである、ステップと、前記第2デバイスから、前記第1オーディオを重み付けする第1オーディオ重みと、前記第2オーディオを重み付けする第2オーディオ重みとを受信するステップと、前記第1デバイスの前記プロセッシング回路によって、前記第1オーディオ重みに基づいた重み付き第1オーディオと、前記第2オーディオ重みに基づいた重み付き第2オーディオとを結合することで、混合オーディオを生成するステップとを有 し、 前記第1メディアストリームは、没入型メディアコンテンツを含み、該没入型メディアコンテンツは、360度ビデオであり、 前記第2メディアストリームは、前記没入型メディアコンテンツに対するオーバーレイメディアコンテンツを含み、該オーバーレイメディアコンテンツは、前記360度ビデオにオーバーレイされる2次元画像であり、 前記第1オーディオ重みは、前記第2オーディオ重みとは異なる、 方法。
- 2前記第1デバイスに関連したスピーカを通じて前記混合オーディオを再生するステップを更に有する、請求項1に記載の方法。
- 3カスタマイズパラメータに基づき前記第1オーディオ重み及び前記第2オーディオ重みをカスタマイズするために前記第2デバイスへ前記カスタマイズパラメータを送信するステップを更に有する、請求項1又は2に記載の方法。
- 4前記第1オーディオ及び前記第2オーディオの音の強さに基づき前記第2デバイスによって決定される前記第1オーディオ重み及び前記第2オーディオ重みを受信するステップを更に有する、請求項1乃至3のうちいずれか一項に記載の方法。
- 5前記第1オーディオ及び前記第2オーディオは、オーバーレイオーディオであり、当該方法は、前記第1オーディオ及び前記第2オーディオのオーバーレイ優先度に基づき前記第2デバイスによって決定される前記第1オーディオ重み及び前記第2オーディオ重みを受信するステップを更に有する、請求項1乃至3のうちいずれか一項に記載の方法。
- 6アクティブスピーカの検出に基づき前記第2デバイスによって調整される前記第1オーディオ重み及び前記第2オーディオ重みを受信するステップを更に有する、請求項1乃至3のうちいずれか一項に記載の方法。
- 7前記プロセッシング回路によって前記混合オーディオを第3メディアストリームにエンコードするステップと、前記第1デバイスのインターフェース回路を介して、前記第3メディアストリームを第3デバイスへ送信するステップとを更に有する、請求項1乃至 6 のうちいずれか一項に記載の方法。
- 8前記第1デバイスの前記インターフェース回路を介して、前記第3メディアストリームと、没入型メディアコンテンツを含む第4メディアストリームとを送信するステップを更に有し、前記第3メディアストリームは、前記第4メディアストリームに対するオーバーレイメディアストリームである、請求項 7 に記載の方法。
- 9テレカンファレンスの方法であって、第1デバイスのプロセッシング回路によって、メディアストリームを夫々供給する複数のデバイスからユーザによる選択又はアクティブスピーカの検出に応答して選択されたいずれか1つのデバイスから、テレカンファレンスセッションの第1メディアコンテンツを運ぶ第1メディアストリームと、前記テレカンファレンスセッションの第2メディアコンテンツを運ぶ第2メディアストリームとを受信するステップと、前記第1デバイスの前記プロセッシング回路によって、前記第1メディアコンテンツと前記第2メディアコンテンツとを混合する第3メディアコンテンツを生成するステップと、前記第1デバイスの伝送回路を介して、前記第3メディアコンテンツを運ぶ第3メディアストリームを第2デバイスへ送信するステップとを有 し、 前記第1メディアストリームは、没入型メディアコンテンツを含み、該没入型メディアコンテンツは、360度ビデオであり、 前記第2メディアストリームは、前記没入型メディアコンテンツに対するオーバーレイメディアコンテンツを含み、該オーバーレイメディアコンテンツは、前記360度ビデオにオーバーレイされる2次元画像であり、 前記第1メディアコンテンツ内の第1オーディオ及び前記第2メディアコンテンツ内の第2オーディオは、前記第1オーディオに割り当てられた第1オーディオ重みと、前記第2オーディオに割り当てられた、前記第1オーディオ重みとは値が異なる第2オーディオ重みとに基づいて混合される、 方法。
- 10テレカンファレンスの方法であって、第1デバイスによって、第2デバイスへ、第1オーディオを運ぶ第1メディアストリームと、第2オーディオを運ぶ第2メディアストリームとを送信するステップであり、前記第1メディアストリーム及び前記第2メディアストリームは、複数のデバイスからユーザによる選択又はアクティブスピーカの検出に応答して選択された前記第1デバイス又はいずれか1つの他のデバイスから供給されたメディアストリームに基づく、ステップと、前記第1デバイスによって、前記第1オーディオを重み付けする第1オーディオ重みと、前記第2オーディオを重み付けする第2オーディオ重みとを決定するステップと、前記第1デバイスによって、前記第2デバイスへ、前記第1オーディオと前記第2オーディオとを混合するために前記第1オーディオ重み及び前記第2オーディオ重みを送信するステップとを有 し、 前記第1メディアストリームは、没入型メディアコンテンツを含み、該没入型メディアコンテンツは、360度ビデオであり、 前記第2メディアストリームは、前記没入型メディアコンテンツに対するオーバーレイメディアコンテンツを含み、該オーバーレイメディアコンテンツは、前記360度ビデオにオーバーレイされる2次元画像であり、 前記第1オーディオ重み及び前記第2オーディオ重みは、互いに異なる値を有するよう決定される、 方法。
- 11セッション記述プロトコルに基づきカスタマイズパラメータを受け取るステップと、前記カスタマイズパラメータに基づき前記第1オーディオ重み及び前記第2オーディオ重みを決定するステップとを更に有する、請求項 10 に記載の方法。
Independent claims11
162 paragraphs, as filed
[Incorporated by reference]
This patent application claims benefit of priority to U.S. Provisional Patent Application No. 63/088,300, filed October 6, 2020, entitled "NETWORK BASED MEDIA PROCESSING FOR AUDIO AND VIDEO MIXING FOR TELECONFERENCING AND TELEPRESENCE FOR REMOTE TERMINALS," and U.S. Provisional Patent Application No. 63/124,261, filed December 11, 2020, entitled "AUDIO MIXING METHODS FOR TELECONFERENCING AND TELEPRESENCE FOR REMOTE TERMINALS," and claims benefit of priority to U.S. Provisional Patent Application No. 17/327,400, filed May 21, 2021, entitled "METHOD AND APPARATUS FOR TELECONFERENCE." The disclosures of all of these prior applications are incorporated herein by reference in their entirety.
[Technical field]
This disclosure describes embodiments generally relating to teleconferencing.
The background statement provided herein is intended to generally present the context of the present disclosure. To the extent that the work of the currently named inventors is described in this background section, and any aspects of the description that may not otherwise qualify as prior art at the time of filing, are not admitted, either explicitly or implicitly, as prior art to the present disclosure.
A teleconferencing system allows users at two or more remote locations to interact with each other via media streams, such as video streams, audio streams, or both. Some teleconferencing systems also allow users to exchange digital documents, such as images, text, videos, applications, etc.
Aspects of the present disclosure provide a method and apparatus for teleconferencing. In some examples, the teleconferencing apparatus includes a processing circuit. The processing circuit of a first device (e.g., a user device or a server for network-based media processing) receives a first media stream carrying a first audio and a second media stream carrying a second audio from a second device. The processing circuit receives a first audio weight for weighting the first audio and a second audio weight for weighting the second audio from the second device, and generates a mixed audio by combining the weighted first audio based on the first audio weight and the weighted second audio based on the second audio weight.
In some examples, the first device is a user device. The first device can play the mixed audio through a speaker associated with the first device.
In an example, the first device sends the customization parameters to the second device to customize the first audio weighting and the second audio weighting based on the customization parameters.
In some examples, the first audio weighting and the second audio weighting are determined by the second device based on the intensity of the first audio and the second audio.
In some examples, the first audio and the second audio are overlay audio, and the processing circuitry receives first audio weights and second audio weights determined by the second device based on overlay priorities of the first audio and the second audio.
In some examples, the first audio weight and the second audio weight are adjusted by the second device based on detection of an active speaker.
In some examples, the first media stream includes immersive media content and the second media stream includes overlay media content, and the first audio weighting is different from the second audio weighting.
In some embodiments, the first device is a network-based media processing device. The processing circuitry encodes the mixed audio into a third media stream and transmits the third media stream to the user device via an interface circuitry of the first device. In some examples, the processing circuitry transmits the third media stream and a fourth media stream including the immersive media content via the interface circuitry. The third media stream is an overlay to the fourth media stream.
In accordance with some aspects of the disclosure, a processing circuit of a first device (e.g., a server for network-based media processing) receives a first media stream carrying first media content of a teleconference session and a second media stream carrying second media content of the teleconference session, the processing circuit generates third media content that mixes the first media content and the second media content, and transmits the third media stream carrying the third media content to the second device via the transmission circuit.
In some embodiments, the processing circuitry of the first device mixes the first audio in the first media content with the second audio in the second media content to generate the third audio based on the first audio weight assigned to the first audio and the second audio weight assigned to the second audio. In some examples, the first audio weight and the second audio weight are received from a host device transmitting the first media stream and the second media stream. In some examples, the first device may determine the first audio weight and the second audio weight.
In some examples, the first media stream is an immersive media stream and the second media stream is an overlay media stream, and the processing circuitry of the first device mixes the first audio with the second audio based on first and second audio weights that have different values.
In some examples, the first media stream and the second media stream are overlay media streams, and the processing circuitry of the first device mixes the first audio with the second audio based on equal values of the first audio weight and the second audio weight.
In some examples, the first media stream and the second media stream are overlay media streams, and the processing circuitry of the first device mixes the first audio with the second audio based on first audio weights and second audio weights associated with overlay priorities of the first media stream and the second media stream.
In accordance with some aspects of the disclosure, a first device (e.g., a host device generating immersive media content) can transmit a first media stream carrying a first audio and a second media stream carrying a second audio to a second device. The first device can determine a first audio weight for weighting the first audio and a second audio weight for weighting the second audio, and transmit the first audio weight and the second audio weight to the second device for mixing the first audio and the second audio.
In some examples, the first device receives the customization parameters based on a session description protocol and determines the first audio weighting and the second audio weighting based on the customization parameters.
In some examples, the first device determines the first audio weight and the second audio weight based on the intensity of the first audio and the second audio.
In some examples, the first audio and the second audio are overlay audio, and the first device determines the first audio weight and the second audio weight based on an overlay priority of the first audio and the second audio.
In some examples, the first device determines the first audio weight and the second audio weight based on detection of an active speaker in one of the first audio and the second audio.
In some examples, the first media stream includes immersive media content and the second media stream is an overlay media stream. The first device determines different values for the first audio weight and the second audio weight.
Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for teleconferencing, cause the computer to perform a teleconferencing method.
Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
<figref num="1">1 illustrates a teleconferencing system according to some examples of the present disclosure.</figref><figref num="2">1 illustrates another teleconferencing system in accordance with some examples of the present disclosure.</figref><figref num="3">1 illustrates another teleconferencing system in accordance with some examples of the present disclosure.</figref><figref num="4">1 shows a flowchart illustrating a process according to some examples of the present disclosure.</figref><figref num="5">1 shows a flowchart illustrating a process according to some examples of the present disclosure.</figref><figref num="6">1 shows a flowchart illustrating a process according to some examples of the present disclosure.</figref><figref num="7">FIG. 1 is a schematic diagram of a computer system according to the present disclosure.</figref>
Aspects of the present disclosure provide techniques for media mixing, such as audio mixing, video mixing, and the like, for teleconferencing. In some examples, a teleconferencing can be an audioconference, where teleconference participants interact with audio streams. In some examples, a teleconferencing can be a videoconference, where teleconference participants interact with media streams that may include video and/or audio. In some examples, media mixing is performed by a network-based media processing element, such as a server device. In some examples, media mixing is performed by an end-user device (also referred to as a user device).
In accordance with some aspects of the present disclosure, the media mixing techniques can be implemented in a variety of teleconferencing systems. Figures 1-3 show several teleconferencing systems.
FIG. 1 illustrates a teleconferencing system (100) according to some examples of the present disclosure. The teleconferencing system (100) includes a subsystem (110) and a number of user devices, such as user devices (120) and (130). The subsystem (110) is installed in a location, such as a conference room A. In general, the subsystem (110) is configured to have a relatively higher bandwidth than the user devices (120) and (130) and can provide host services for a teleconferencing session (also called a teleconferencing call). The subsystem (110) can enable users or participants in the conference room A to participate in the teleconferencing session, and can also enable several remote users, such as user B of the user device (120) and user C of the user device (130), to participate in the teleconferencing session from remote locations. In some examples, the subsystem (110) and the user devices (120) and (130) are referred to as terminals in the teleconferencing session.
In some embodiments, the subsystem (110) includes various audio, video, and control components suitable for a conference room. The various audio, video, and control components can be integrated into the device or can be distributed components linked by appropriate communication technology. In some examples, the subsystem (110) includes a wide-angle camera (111), such as a fisheye camera, an omnidirectional camera, or the like, having a relatively wide field of view. For example, an omnidirectional camera can be configured to have a field of view that covers approximately the entire sphere, and video captured by an omnidirectional camera can be referred to as omnidirectional video or 360-degree video.
Additionally, in some examples, the subsystem (110) includes a microphone (112), such as an omnidirectional (also called non-directional) microphone that can capture sound waves from approximately any direction. The subsystem (110) may include a display screen (114), a speaker device, etc., to allow users in conference room A to play multimedia corresponding to video and audio of users located outside of conference room A. In examples, the speaker device may be integrated with the microphone (112) or may be a separate component (not shown).
In the example, subsystem (110) includes a controller (113). While a laptop computing device is shown as controller (113) in Figure 1, other suitable devices may be used as controller (113), such as a desktop computer, a tablet computer, etc. It is also noted that in the example, controller (113) may be integrated with other components of subsystem (110).
The controller 113 may be configured to perform various control functions for the subsystems 110. For example, the controller 113 may be used to initiate a teleconference session and manage communication of values between the subsystems 110 and the user devices 120 and 130. In an example, the controller 113 may encode video and/or audio captured in conference room A (e.g., captured by the camera 111 and microphone 112) to generate a media stream carrying the video and/or audio, and cause the media stream to be transmitted to the user devices 120 and 130.
Additionally, in some examples, the controller 113 can receive media streams from each of the user devices (user devices 120 and 130) of the teleconferencing system 100, carrying audio and/or video captured at each of the user devices. The controller 113 can address and transmit the received media streams to other user devices of the teleconferencing system. For example, the controller 113 can receive a media stream from the user device 120 and address and transmit the media stream to the user device 130, and can receive other media streams from the user device 130 and address and transmit the other media stream to the user device 120.
Additionally, in some examples, the controller (113) can determine appropriate teleconference parameters, such as audio, video mixing parameters, etc., and transmit the teleconference parameters to the user devices (120) and (130).
In some examples, the controller (113) can facilitate user input in conference room A to the display of a user interface on a screen, such as a display screen (114), the screen of a laptop computing device, or the like.
Each of user devices 120 and 130 can be any suitable teleconferencing-enabled equipment, such as a desktop computer, a laptop computer, a tablet computer, a wearable device, a handheld device, a smartphone, a mobile device, an embedded device, a game console, a gaming device, a personal digital assistant (PDA), a telecommunications device, a global positioning system ("GPS") device, a virtual reality ("VR") device, an augmented reality (AR) device, an implantable computing device, an automobile computer, a network-enabled television, an Internet of Things ("IoT") device, a workstation, a media player, a personal video recorder (PVR), a set-top box, a camera, an internal component (e.g., a peripheral) included in a computing device, an appliance, or any other type of computing device.
In the example of FIG. 1, the user device (120) includes a wearable multimedia component that enables a user, such as user B, to participate in a teleconferencing system. For example, the user device (120) includes a head mounted display (HMD) that can be worn on the head of user B. The HMD can include display optics in front of one or both eyes of user B to play video. In another example, the user device (120) includes a headset (not shown) that can be worn by user B. The headset can include a microphone that captures the user's voice and one or two earpieces that output audio sound. The user device (120) also includes suitable communication components (not shown) that can transmit and/or receive media streams.
In the example of FIG. 1, the user device (130) may be a mobile device such as a smartphone that incorporates together communications components, imaging components, audio components, etc. to enable a user, such as User C, to participate in a teleconference session.
1, the subsystems (110), user devices (120), and user devices (130) include suitable communication components (not shown) capable of interfacing with a network (101). The communication components may include one or more network interface controllers (NICs) or other types of transceiver circuitry to transmit and receive communications and/or data over a network, such as network (101).
The network (101) may include, for example, a public network such as the Internet, a private network such as a public and/or personal intranet, or some combination of private and public networks. The network (101) may also include any type of wired and/or wireless network, including, but not limited to, a local area network ("LAN"), a wide area network ("WAN"), a satellite network, a cable network, a Wi-Fi network, a Wi-Max network, a mobile communication network (e.g., 3G, 4G, 5G, etc.), or any combination thereof. The network (101) may utilize communication protocols including packet-based and/or datagram-based protocols such as the Internet Protocol ("IP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), or other types of protocols. Additionally, the network (101) may also include a number of devices that facilitate network communication and/or form the hardware infrastructure for the network, such as switches, routers, gateways, access points, firewalls, base stations, relay stations, backbone devices, etc. In some examples, the network (101) may further include devices that allow for connection to a wireless network, such as a wireless access point ("WAP").
In the example of Figure 1, the subsystem (110) can host a teleconference session using peer-to-peer technology. For example, after a user device (120) joins the teleconference session, the user device (120) can properly address packets (e.g., using the IP address of the subsystem (110)) and send the packets to the subsystem (110), and the subsystem (110) can properly address packets (e.g., using the IP address of the user device (120)) and send the packets to the user device (120). The packets can carry various information and data, such as media streams, acknowledgments, control parameters, etc.
In some examples, the teleconferencing system (100) may be an immersive teleconferencing system. In one embodiment, the subsystem (110) can provide a teleconferencing (teleconferencing) session. For example, during the teleconferencing session, the subsystem (110) is configured to generate immersive media, such as omnidirectional video/audio, using an omnidirectional camera and/or an omnidirectional microphone. In an example, the HMD of the user device (120) can detect a head movement of the user B and determine an orientation of the user B's viewport based on the head movement. The user device (120) can transmit the orientation of the user B's viewport to the subsystem (110), which can then transmit viewport-dependent streams, such as a video stream adjusted based on the orientation of the user B's viewport (a media stream carrying video adjusted based on the orientation of the user B's viewport), an audio stream adjusted based on the orientation of the user B's viewport (a media stream carrying audio adjusted based on the orientation of the user B's viewport), etc., to the user device (120) for playback on the user device (120).
In another example, user C can use the user device (130) to input the orientation of user C's viewport (e.g., using a touch screen on a smartphone). The user device (130) can transmit the orientation of user C's viewport to the subsystem (110), which can then transmit viewport-dependent streams, such as a video stream (a media stream carrying video that is adjusted based on the orientation of user C's viewport) adjusted based on the orientation of user C's viewport, an audio stream (a media stream carrying audio that is adjusted based on the orientation of user C's viewport), etc., to the user device (130) for playback on the user device (130).
It is noted that during the teleconference session, the viewport orientation of user B and/or user C may change. The change in viewport orientation may be notified to subsystem 110, which may adjust the viewport orientation in each viewport-dependent stream sent to user device 120 and user device 130, respectively.
For ease of explanation, immersive media is used to refer to wide-angle media, such as omnidirectional video and omnidirectional audio, and to refer to viewport-dependent streams that are generated based on the wide-angle media. In this disclosure, 360-degree media, such as 360-degree video and 360-degree audio, are used to describe techniques for teleconferencing, although it is noted that teleconferencing techniques can be used with immersive media that is less than 360 degrees.
FIG. 2 illustrates another teleconference system (200) according to some examples of the present disclosure. The teleconference system (200) includes multiple subsystems, such as subsystems (210A)-(210Z), and multiple user devices, such as user devices (220) and (230), that are installed in conference rooms A through Z, respectively. One of the subsystems (210A)-(210Z) can initiate a teleconference session and enable other subsystems and user devices, such as user device (220) and user device (230), to participate in the teleconference session. Thus, users, such as users in conference rooms A through Z, user B of user device (220), and user C of user device (230), can participate in the teleconference session. In some examples, the subsystems (210A)-(210Z) and user devices (220) and (230) are referred to as terminals in the teleconference session.
In some embodiments, each of the subsystems 210A-Z operates in a similar manner to the subsystem 110 described above. Additionally, each of the subsystems 210A-Z utilizes certain components that are the same or equivalent to those used in the subsystem 110. A description of these components has been provided above and will be omitted here for clarity. It is noted that the subsystems 210A-Z may be configured differently from one another.
User devices 220 and 230 are configured similarly to user devices 120 and 130 described above, and network 201 is configured similarly to network 101. Descriptions of these components have been provided above and are omitted here for clarity.
In some embodiments, one of the subsystems (210A)-(210Z) can initiate a teleconference session, and the remainder of the subsystems (210A)-(210Z) and the user devices (220) and (230) can participate in the teleconference session.
According to aspects of the present disclosure, during a teleconference session of an immersive teleconference, multiple subsystems among subsystems 210A-210Z can generate respective immersive media, and user devices 220 and 230 can select one of subsystems 210A-210Z to provide the immersive media. Generally, subsystems 210A-210Z are configured with relatively high bandwidth and can each act as a host providing the immersive media.
In an example, after the user device 220 joins a teleconference session, the user device 220 can select one of the subsystems 210A-Z, e.g., the subsystem 210A, as a host for the immersive media. The user device 220 can address and send packets to the subsystem 210A, and the subsystem 210A can address and send packets to the user device 220. The packets can include any suitable information/data, such as media streams, control parameters, and the like. In some examples, the subsystem 210A can send adjusted media information to the user device 220. It is noted that the user device 220 can change the selection from the subsystems 210A-Z during the teleconference session.
In an example, the HMD of the user device (220) can detect head movement of user B and determine an orientation of the viewport for user B based on the head movement. The user device (220) can send the orientation of the viewport for user B to the subsystem (210A), which can then send viewport-dependent streams, such as a video stream adjusted based on the orientation of the viewport for user B, an audio stream adjusted based on the orientation of the viewport for user B, etc., to the user device (220) for playback on the user device (220).
In another example, after the user device 230 joins the teleconference session, the user device 230 can select one of the subsystems 210A-210Z, e.g., the subsystem 210Z, as a host for the immersive media. The user device 230 can address and send packets to the subsystem 210Z, and the subsystem 210Z can address and send packets to the user device 230. The packets can include any suitable information/data, such as media streams, control parameters, and the like. In some examples, the subsystem 210Z can send adjusted media information to the user device 230. It is noted that the user device 230 can change the selection from the subsystems 210A-210Z during the teleconference session.
In another example, user C can use the user device (230) to input an orientation of user C's viewport (e.g., using a touch screen on a smartphone). The user device (230) can transmit the orientation of user C's viewport to the subsystem (210Z), which can then transmit viewport-dependent streams, such as a video stream adjusted based on the orientation of user C's viewport, an audio stream adjusted based on the orientation of user C's viewport, etc., to the user device (230) for playback on the user device (230).
It is noted that during a teleconference session, the viewport orientation of a user (e.g., User B, User C) may change. For example, a change in the viewport orientation of User B may be notified to a subsystem selected by User B, which may then adjust the viewport orientation accordingly in the viewport-dependent streams sent to the user device (220).
For ease of explanation, immersive media is used to refer to wide-angle media, such as omnidirectional video and omnidirectional audio, and to refer to viewport-dependent streams that are generated based on the wide-angle media. In this disclosure, 360-degree media, such as 360-degree video and 360-degree audio, are used to describe techniques for teleconferencing, although it is noted that teleconferencing techniques can be used with immersive media that is less than 360 degrees.
3 shows another teleconference system (300) according to some examples of the present disclosure. The teleconference system (300) includes a network-based media processing server (340), multiple subsystems such as subsystems (310A)-(310Z) installed in conference rooms A to B, respectively, and user devices such as user devices (320) and (330). The network-based media processing server (340) can set up a teleconference session and enable the subsystems (310A)-(310Z) and user devices such as user devices (320) and (330) to participate in the teleconference session. Thus, users in conference rooms A to Z, user B of user device (320), and user C of user device (330) can participate in the teleconference session.
In some examples, the subsystems 310A-310Z and the user devices 320 and 330 are referred to as terminals in a teleconference session, and the network-based media processing server 340 can bridge the terminals in the teleconference session. In some examples, the network-based media processing server 340 is referred to as a media-aware networking element. The network-based media processing server 340 can perform a media resource function (MRF) and can perform a media control function as a media control unit (MCU).
In some embodiments, each of the subsystems 310A-Z operates in a similar manner to the subsystem 110 described above. Additionally, each of the subsystems 310A-Z utilizes certain components that are the same or equivalent to those used in the subsystem 110. A description of these components has been provided above and will be omitted here for clarity. It is noted that the subsystems 310A-Z may be configured differently from one another.
User devices 320 and 330 are configured similarly to user devices 120 and 130 described above, and network 301 is configured similarly to network 101. Descriptions of these components have been provided above and are omitted here for clarity.
In some examples, the network-based media processing server 340 can initiate a teleconference session. For example, one of the subsystems 310A-Z and the user devices 320 and 330 can access the network-based media processing server 340 to initiate a teleconference session. The subsystems 310A-Z and the user devices 320 and 330 can participate in the teleconference session. Furthermore, the network-based media processing server 340 is configured to provide media-related functions for bridging the terminals in the teleconference session. For example, the subsystems 310A-Z can each address packets carrying their respective media information, such as video and audio, and send the packets to the network-based media processing server 340. It is noted that the media information sent to the network-based media processing server 340 is viewport dependent. For example, the subsystems 310A-310Z can send their respective videos, such as full 360-degree videos, to a network-based media processing server 340. In addition, the network-based media processing server 340 can receive viewport orientations from the user devices 320 and 330, perform media processing to condition the media, and transmit the conditional media information to each of the user devices.
In an example, after the user device (320) joins the teleconference session, the user device (320) can address and send packets to the network-based media processing server (340), and the network-based media processing server (340) can address and send packets to the user device (320). The packets can include any suitable information/data, such as media streams, control parameters, and the like. In an example, user B can use the user device (320) to select a conference room to view video from a subsystem in the conference room. For example, user B can use the user device (320) to select conference room A to view captured video from a subsystem (310A) installed in conference room A. Furthermore, the HMD of the user device (320) can detect the head movement of user B and determine the orientation of the viewport of user B based on the head movement. The user device (320) can send the selection of conference room A and the orientation of user B's viewport to the network-based media processing server (340), which can process the media sent from the subsystem (310A) and send viewport-dependent streams, such as a video stream adjusted based on the orientation of user B's viewport, an audio stream adjusted based on the orientation of user B's viewport, etc., to the user device (320) for playback at the user device (320). In some examples, when the user device (320) selects conference room A, the user device (320), the subsystem (310A), and the network-based media processing server (340) can communicate with each other based on the Session Description Protocol (SDP).
In another example, after the user device (330) joins the teleconference session, the user device (330) can address and send packets to the network-based media processing server (340), which can address and send packets to the user device (330). The packets can include any suitable information/data, such as media streams, control parameters, and the like. In some examples, the network-based media processing server (340) can send the adjusted media information to the user device (330). For example, user C can use the user device (330) to input (e.g., using a touch screen on a smartphone) a selection of a conference room, e.g., conference room Z, and an orientation of user C's viewport. The user device (330) can send the selection of conference room Z and the orientation of user C's viewport to the network-based media processing server (340), which can process the media sent from the subsystem (310Z) and send viewport-dependent streams, such as a video stream adjusted based on the orientation of user C's viewport, an audio stream adjusted based on the orientation of user C's viewport, etc., to the user device (330) for playback at the user device (330). In some examples, when the user device (330) selects conference room Z, the user device (330), the subsystem (310Z), and the network-based media processing server (340) can communicate with each other based on the Session Description Protocol (SDP).
It is noted that during a teleconference session, the viewport orientation of the users (e.g., User B, User C) may change. For example, a change in the viewport orientation of User B may be notified by User B to the network-based media processing server (340), which may then adjust the viewport orientation accordingly in the viewport-dependent streams sent to the user device (320).
For ease of explanation, immersive media is used to refer to wide-angle media, such as omnidirectional video and omnidirectional audio, and to refer to viewport-dependent streams that are generated based on the wide-angle media. In this disclosure, 360-degree media, such as 360-degree video and 360-degree audio, are used to describe techniques for teleconferencing, although it is noted that teleconferencing techniques can be used with immersive media that is less than 360 degrees.
It is noted that the conference room selection may change during the teleconference session. In an example, a user device such as the user device (320), the user device (330), etc., may trigger a switch from one conference room to another conference room based on an active speaker. For example, in response to an active speaker being in conference room A, the user device (330) may determine to switch the conference room selection to conference room A and send the selection of conference room A to the network-based media processing server (340). The network-based media processing server (340) may then process the media sent from the subsystem (310A) and send viewport-dependent streams, such as a video stream adjusted based on the orientation of the viewport of user C, an audio stream adjusted based on the orientation of the viewport of user C, etc., to the user device (330) for playback at the user device (330).
In some examples, the network-based media processing server (340) can suspend receiving video streams from any conference rooms that do not have any active users. For example, the network-based media processing server (340) can determine that conference room Z does not have any active users, and then the network-based media processing server (340) can suspend receiving video streams from subsystem (310Z).
In some examples, the network-based media processing server (340) can include distributed computing resources and can communicate with the subsystems (310A)-(310Z) and user devices (320) and (330) via the network (301). In some examples, the network-based media processing server (340) can be a separate system tasked with managing aspects of one or more teleconference sessions.
In various examples, the network-based media processing server (340) may include one or more computing devices operating in a cluster or other grouped configuration to share resources, balance load, improve performance, provide failover support or redundancy, or for other purposes. For example, the network-based media processing server (340) may belong to various types of devices, such as traditional server-type devices, desktop computer-type devices, and/or mobile-type devices. Thus, even though represented as one type of device (a server-type device), the network-based media processing server (340) may include a wide variety of device types and is not limited to a particular type of device. The network-based media processing server (340) may represent, but is not limited to, a server computer, a desktop computer, a web server computer, a personal computer, a mobile computer, a laptop computer, a tablet computer, or any other type of computing device.
In accordance with aspects of the present disclosure, the network-based media processing server (340) can perform certain media functions to reduce the processing burden on terminals such as the user devices (320), (330), etc. For example, the user devices (320) and/or (330) may have limited media processing capacity or may have difficulty encoding and rendering multiple video streams, so the network-based media processing server (340) can perform media processing, such as decoding/encoding audio and video streams, to offload the media processing on the user devices (320) and (330). In some examples, the user devices (320) and (330) are battery-powered devices, and the battery life of the user devices (320) and (330) is extended when the media processing of the user devices (320) and (330) is offloaded to the network-based media processing server (340).
Media streams from different resources may be processed and mixed. In some examples, for example, in International Organization for Standardization (ISO) 23090-2, an overlay may be defined as a second media rendered on top of a first media. In accordance with aspects of the present disclosure, for a teleconference session of an immersive teleconference, additional media content (e.g., video and/or audio) may be overlaid on the immersive media content. The additional media (or media content) may be referred to as overlay media (or overlay media content) on the immersive media (or immersive media content). For example, the overlay content may be a piece of visual/audio media rendered on top of an omnidirectional video or image item or on top of a viewport.
Using FIG. 2 as an example, when a presentation is shared by participants in conference room A, in addition to being displayed by the subsystem (210A) in conference room A, the presentation is also broadcast as a stream (also called an overlay stream) to other participant parties, such as the subsystem (210Z), the user device (220), the user device (230), and the like. For example, the user device (220) selects conference room A, and the subsystem (210A) can send a first stream of immersive media, such as a 360-degree video captured by the subsystem (210A), and an overlay stream to the user device (220). At the user device (220), the presentation can be overlaid on the 360-degree video captured by the subsystem (210A). In another example, the user device (230) selects conference room Z, and the subsystem (210Z) can send a first stream carrying immersive media, such as a 360-degree video captured by the subsystem (210Z), and an overlay stream to the user device (230). At the user device (230), the presentation may be overlaid on top of the 360-degree video captured by the subsystem (210Z). It is noted that the presentation may in some instances be overlaid on top of the 2D video.
In another scenario, user C can be a remote speaker, and a media stream carrying audio corresponding to user C's speech (called an overlay stream) can be sent from the user device (230) to, for example, the subsystem (210Z) and broadcast to other participating parties, such as the subsystem (210A). For example, the user device (220) selects conference room A, and the subsystem (210A) can send a first stream of immersive media, such as a 360-degree video captured by the subsystem (210A), and an overlay stream to the user device (220). At the user device (220), the audio corresponding to user C's speech can be overlaid with the 360-degree video captured by the subsystem (210A). The media stream carrying audio corresponding to user C's speech can be called an overlay stream, in an example, and the audio can be called an overlay audio, in an example.
Some aspects of the present disclosure provide techniques for audio and video mixing, and more specifically, for combining audio and/or video of multiple media streams, such as an immersive stream and one or more overlay streams. In accordance with aspects of the present disclosure, audio and/or video mixing may be performed by a network-based media processing element, such as a network-based media processing server (340), or by an end-user device, such as a user device (120), a user device (130), a user device (220), a user device (230), a user device (320), a user device (330), or the like.
In the example of Figure 1, the subsystem (110) is referred to as a sender capable of transmitting multiple media streams each carrying media (audio and/or video), and the user devices (120) and (130) are referred to as receivers. In the example of Figure 2, the subsystems (210A)-(210Z) are referred to as senders capable of transmitting multiple media streams each carrying media (audio and/or video), and the user devices (220) and (230) are referred to as receivers. In the example of Figure 3, the network-based media processing server (340) is referred to as a sender capable of transmitting multiple media streams each carrying media (audio and/or video), and the user devices (320) and (330) are referred to as receivers.
According to some aspects of the present disclosure, mixing levels such as audio weights can be assigned to the overlay streams and the immersive streams in an immersive teleconference for audio mixing. Furthermore, in some embodiments, the audio weights can be appropriately adjusted, and the adjusted audio weights can be used for audio mixing. In some examples, the audio mixing is also referred to as audio downmix.
In some examples, such as immersive teleconferencing, when overlay media is overlaid on the immersive media, overlay information may need to be provided, such as an overlay source, an overlay rendering type, overlay rendering properties, user interaction properties, etc. In some examples, the overlay source specifies the media, such as an image, audio, or video, that is used as the overlay, the overlay rendering type describes whether the overlay is fixed relative to the viewport or sphere, the overlay rendering properties may include opacity, transparency, etc.
In the example of FIG. 2, multiple conference rooms, each equipped with an omnidirectional camera, can participate in a teleconference session. A user, for example, user B, can select a source of immersive media, for example, one of the multiple conference rooms, each equipped with an omnidirectional camera, via a user device (220). To add additional media, such as audio or video, to the immersive media, the additional media can be transmitted to the user device (220) separately from the immersive media as an overlay stream carrying the additional media. The immersive media can be transmitted as a stream carrying the immersive media (called an immersive stream). The user device (220) can receive the immersive stream and the overlay stream and can overlay the additional media with the immersive media.
According to aspects of the disclosure, user devices such as user device (220), user device (230), etc., can receive multiple media streams carrying respective audio in a teleconference session. The user devices can decode the media streams to extract the audio and mix the decoded audio from the media streams. In some examples, during a teleconference of an immersive teleconference, a subsystem of a selected conference room can transmit multiple media streams and can provide mixing parameters for the audio carried in the multiple media streams. In an example, user B can select conference room A via the user device (220) to receive an immersive stream carrying 360-degree immersive video captured by subsystem (210A). Subsystem (210A) can transmit the immersive stream along with one or more overlay streams to the user device (220). Subsystem (210A) can provide mixing levels for the audio carried in the immersive stream and one or more overlay streams based on, for example, a Session Description Protocol (SDP). It is noted that the subsystem (210A) may also update the audio mixing level during the teleconference session and send a signal to notify the user device (220) of the updated mixing level based on the SDP.
In an example, the mixing levels of the audio are defined using audio mixing weights. For example, the subsystem (210A) transmitting the immersive and overlay streams carrying each audio can determine an audio mixing weight for each audio. In an example, the subsystem (210A) determines the default audio mixing weights based on sound intensity. Sound intensity can be defined as the power carried by a sound wave per unit area orthogonal to the unit area. For example, a controller of the subsystem (210A) can receive an electrical signal indicating the sound intensity of each audio and determine the default audio mixing weights based on the electrical signal, e.g., based on the signal level, power level, etc. of the electrical signal.
In another example, the subsystem 210A determines the audio mixing weights based on the overlay priority. For example, the controller of the subsystem 210A can detect a particular media stream from the immersive stream and the overlay stream that carries audio for an active speaker. The controller of the subsystem 210A can determine a higher overlay priority for that particular media stream and can then determine a higher mixing weight for the audio carried by that particular media stream.
In another example, an end user can customize the overlay priority. For example, user B can use the user device (220) to send customization parameters to the subsystem (210A) based on the SDP. The customization parameters can indicate, for example, a particular media stream carrying audio on which user B wants to focus. The subsystem (210A) can then determine a higher overlay priority for that particular media stream and a higher mixing weight for the audio carried by that particular media stream.
In some embodiments, when overlay priorities are used, a sender, e.g., subsystem 210A, may be informed of all overlays of other senders, e.g., subsystem 210Z, and the priorities of those overlays in the teleconference session, and assign weights accordingly, so that when a user device switches to another subsystem, audio mixing weights can be appropriately determined.
In some embodiments, the audio mixing weights may be customized by an end user. In one scenario, the end user may want to hear or focus on one particular audio carried by a media stream. In another scenario, the quality of the downmixed audio with the default audio mixing weights is unacceptable due to reasons such as audio level fluctuations, audio quality, or poor signal-to-noise ratio (SNR) channel, in which case the audio mixing weights may be customized. In an example, user B wants to focus on audio from a particular media stream, in which case user B may use the user device (220) to indicate customization parameters for adjusting the audio mixing weights. For example, the customization parameters indicate an increase in the audio mixing weight for the audio of the particular media stream. The user device (220) can send the customization parameters to a sender of the media stream, e.g., the subsystem (210A), during a teleconference session based on the SDP. Based on the customization parameters, the controller of the subsystem (210A) can adjust the audio mixing weights to increase the audio mixing weights for the audio of the particular media stream, and the subsystem (210A) can send the adjusted audio mixing weights to the user device (220) so that the user device (220) can mix the audio based on the adjusted audio mixing weights.
It is also noted that in some examples, a user device, such as the user device (120), the user device (130), the user device (220), the user device (230), the user device (320), the user device (330), etc., may overwrite the received audio mixing weights with different values based on user preferences.
In the example of FIG. 3, multiple conference rooms, each equipped with an omnidirectional camera, can participate in a teleconference session. A user, for example, user B, can select a source of immersive media, for example, one of the multiple conference rooms, each equipped with an omnidirectional camera, via a user device (320). To add additional media, such as audio or video, to the immersive media, the additional media can be sent to the user device (320) separately from the immersive media as an overlay stream carrying the additional media. In some embodiments, a network-based media processing server (340) receives media streams from participant parties in the teleconference (e.g., subsystems (310A)-(310Z), user devices (320) and (330)), processes the media streams, and sends the appropriate processed media streams to the participant parties. For example, the network-based media processing server (340) can send an immersive stream carrying the immersive media captured by the subsystem (310A) and an overlay stream carrying the overlay media to the user device (320). The user device (320) can receive the immersive stream and the overlay stream, and in some embodiments can overlay the overlay media with the immersive media.
According to aspects of the disclosure, user devices such as user device (320), user device (330), etc., can receive multiple media streams carrying respective audio in a teleconference session. The user devices can decode the media streams to extract the audio and mix the decoded audio from the media streams. In some examples, during a teleconference of an immersive teleconference, the network-based media processing server (340) can transmit multiple media streams to end-user devices. In an example, user B can select conference room A via the user device (320) to receive an immersive stream carrying 360-degree immersive video captured by the subsystem (310A). According to aspects of the disclosure, audio mixing parameters such as loudness can be defined by the sender of the immersive media or customized by the end user. In some examples, the subsystem (310A) can provide a mixing level for the audio carried in one or more overlay streams to the network-based media processing server (340) via, for example, a Session Description Protocol (SDP)-based signal. It is noted that the subsystem (310A) may also update the audio mixing level during the teleconference session and send a signal to notify the network-based media processing server (340) of the updated mixing level based on the SDP.
In an example, audio mixing levels are defined using audio mixing weights. In an example, the subsystem (310A) can determine and send the audio mixing weights to the network-based media processing server (340) based on the SDP. In an example, the subsystem (310A) determines default audio mixing weights based on sound intensity.
In another example, the subsystem (310A) determines the audio mixing weights based on the overlay priority. For example, the subsystem (310A) can detect a particular media stream that carries audio for an active speaker. The subsystem (310A) can determine a higher overlay priority for that particular media stream and can determine a higher mixing weight for the audio carried by that particular media stream.
In another example, an end user can customize the overlay priority. For example, user B can use the user device (320) to send customization parameters to the subsystem (310A) based on the SDP. The customization parameters can indicate, for example, a particular media stream carrying audio on which user B wants to focus. The subsystem (310A) can then determine a higher overlay priority for that particular media stream and a higher mixing weight for the audio carried by that particular media stream.
In some embodiments, when overlay priorities are used, a sender, e.g., subsystem 310A, may be informed of all overlays of other senders, e.g., subsystem 310Z, and the priorities of those overlays in the teleconference session, and assign weights accordingly, so that when a user device switches to another subsystem, audio mixing weights can be appropriately determined.
In some embodiments, the audio mixing weights may be customized by an end user. In one scenario, the end user may want to hear or focus on one particular audio carried by a media stream. In another scenario, the quality of the downmixed audio with the default audio mixing weights is unacceptable due to reasons such as audio level fluctuations, audio quality, or poor signal-to-noise ratio (SNR) channel, in which case the audio mixing weights may be customized. In an example, user B wants to focus on audio from a particular media stream, in which case user B may use the user device (320) to indicate customization parameters for adjusting the audio mixing weights. For example, the customization parameters indicate an increase in the audio mixing weight for the audio of the particular media stream. The user device (320) can send the customization parameters to a sender of the media stream, e.g., the subsystem (310A), during the teleconference session based on the SDP. Based on the customization parameters, the subsystem (310A) can adjust the audio mixing weights to increase the audio mixing weights for the audio of the particular media stream and send the adjusted audio mixing weights to the network-based media processing server (340). In an example, the network-based media processing server (340) can send the adjusted audio mixing weights to the user device (320). Thus, the user device (320) can mix the audio based on the adjusted audio mixing weights. In another example, the network-based media processing server (340) can mix the audio according to the adjusted audio mixing weights.
In examples, the immersive stream and one or more overlay streams are provided from a sender, such as one of the subsystems 210A-210Z or one of the subsystems 31A-310Z, and N represents the number of overlays and is a positive integer. Furthermore, a0 represents audio carried in the immersive stream, a1-aN represent audio carried in the overlay streams, respectively, and r0-rN represent audio mixing weights of a0-aN, respectively. In some examples, the sum of the default audio mixing weights r0-rN is equal to 1. The mixed audio (also referred to as audio output) may be generated according to Equation 1: Audio Output = r0 x a0 + r1 x a1 + ... + rn x an Equation 1
In some embodiments, audio mixing may be performed by an end user device, such as user device (220), user device (230), user device (320), user device (330), etc., based on the audio mixing weights, for example, according to Equation 1. The end user device may decode the received media streams to extract the audio, and mix the audio according to Equation 1 to generate an audio output for playback.
In some embodiments, audio mixing or parts of audio mixing may be performed by the MRF or MCU, for example, by a network-based media processing server (340). Referring to FIG. 3, in some examples, the network-based media processing server (340) receives various media streams carrying audio. Furthermore, the network-based media processing server (340) may perform media mixing, such as audio mixing based on audio mixing weights. Using the subsystem (310A) and the user device (330) as an example (e.g., the user device (330) selects conference room A), when the user device (330) is in a low power state or has limited media processing capabilities, audio mixing or parts of audio mixing may be offloaded to the network-based media processing server (340). In an example, the network-based media processing server (340) may receive media streams to send to the user device (330) and audio mixing weights for mixing the audio in the media streams. The network-based media processing server (340) then decodes the media streams to extract the audio and mixes the audio according to Equation 1 to generate mixed audio. It is noted that the network-based media processing server (340) can appropriately mix the video portions of the media streams into a mixed video. The network-based media processing server (340) can encode the mixed audio and/or mixed video in another stream (called a mixed media stream) and send the mixed media stream to the user device (330). The user device (330) can receive the mixed media stream, decode the mixed media stream to extract the mixed audio and/or mixed video, and play the mixed audio/video.
In another example, the network-based media processing server (340) receives an immersive media stream and multiple overlay media streams for providing media content to the user device (330) and audio mixing weights for mixing audio in the immersive media stream and multiple overlay media streams. When multiple overlay media streams need to be transmitted, the network-based media processing server (340) decodes the multiple overlay media streams to extract the audio, and then mixes the audio to generate a mixed overlay audio, for example, according to Equation 2: Mixed Overlay Audio=r1×a1+...+rn×an Equation 2
It is noted that the network-based media processing server (340) can appropriately mix the video portion of the overlay media stream into a mixed overlay video. The network-based media processing server (340) can encode the mixed overlay audio and/or mixed overlay video in another stream (called a mixed overlay media stream) and send the mixed overlay media stream together with the immersive media stream to the user device (330). The user device (330) can receive the immersive media stream and the mixed media stream and decode the immersive media stream and the mixed media stream to extract the immersive media audio (a0), the mixed overlay audio and/or the mixed overlay video. Based on the immersive media audio (a0) and the mixed overlay audio, the user device (330) can generate the mixed audio (called audio output) for playback, for example, according to Equation 3: Audio Output = r0 x a0 + Mixed Overlay Audio Equation 3
In an example, when there is no background noise or disturbance from any audio from the overlay media stream or the immersive media stream (the audio from the immersive media stream is referred to as background in some examples), or when the audio intensity levels of all media streams are approximately the same or the variance is relatively small (e.g., less than a predefined threshold), audio mixing may be performed by adding together the audio taken from all streams, such as the overlay media stream and the immersive media stream (e.g., with equal mixing weights of 1, respectively), to generate an aggregate audio. The aggregate audio may be normalized (e.g., divided by the number of audios). The audio mixing in this example may be performed by end-user devices, such as the user device (120), the user device (130), the user device (220), the user device (230), the user device (320), the user device (330), and the network-based media processing server (340).
In some embodiments, the audio weights may be used to select a portion of the audio to mix. In an example, when a large amount of audio is collected and then normalized, it may be difficult to distinguish one audio stream from another. Using the audio weights, a selected number of audios may be collected and then normalized. For example, when the total number of audios is 10, the audio weights for the 5 selected audios may be 0.2, and the audio weights for the 5 unselected audios may be 0. The selection of audio may be based on the mixing weights defined by an algorithm, or may be based on the overlay priority.
In some embodiments, a user device may choose to vary the selection of audio from the media streams to be mixed by modifying the audio mixing weights of each, or by using a subset of the media streams to take the audio and mix the audio.
In some embodiments, when there is a large variation in audio loudness in the media stream, the audio mixing weights for the overlay audio and the immersive audio may be set to the same level.
In some embodiments, the number of audio streams to be downmixed may be limited because the user device has limited resource capacity or has difficulty distinguishing between audio from different conference rooms. When such limitations apply, a sender device, such as a subsystem 210A-Z or a network-based media processing server 340, may select the media streams to be audio downmixed based on sound intensity or overlay priority. It is noted that the user device can send customization parameters based on the SDP to change the selection during the teleconference session.
In some scenarios, during a teleconference session, the speaking/presenting person needs to be in focus, so the media stream containing the speaking person's audio may be assigned a relatively large audio mixing weight and the audio mixing weights for other audio in other media streams may be lowered.
In some scenarios, when a remote user is giving a presentation and the immersive audio of the immersive media stream has background noise, a sender, such as a subsystem 210A-Z or a network-based media processing server 340, can lower the audio mixing weighting of the immersive audio to be less than the overlay audio associated with the remote user. This can be customized by end users already in the session by lowering the audio weighting during the teleconference session, but changing the default audio mixing weighting provided by the sender can allow a new remote user who has just joined the conference to get the default audio mixing weighting for the audio stream from the sender to downmix the audio with good acoustic quality.
In an embodiment, audio mixing parameters such as audio mixing weights are defined by a sender device, such as a subsystem (310A)-(310Z). The sender device can determine the audio mixing weights to set the audio streams to the same loudness level. The audio mixing parameters (audio mixing weights) can be sent from the sender device to the network-based media processing server (340) via SDP signaling.
In another embodiment, a sender device, such as subsystems 310A-310Z, can set the audio mixing weights for the audio of the immersive media content to be higher than the audio mixing weights for the other overlay audio of the overlay media stream. In an example, the overlay audio may have the same audio mixing weights. The audio mixing parameters (audio mixing weights) can be sent from the sender device to the network-based media processing server 340 via SDP signaling.
In other embodiments, a sender device, such as subsystems 310A-310Z, may set the audio mixing weights for the audio of the immersive media content to be higher than the audio mixing weights for the overlay audio of the overlay media streams. The audio mixing parameters (audio mixing weights) may be transmitted from the sender device to the network-based media processing server 340 via SDP signaling.
In some instances, the network-based media processing server (340) may transmit the same audio stream to multiple end-user devices, for example, when the end-user devices may not have sufficient processing capacity.
In some examples, individual audio streams may be encoded per user device by the sender device or by the network-based media processing server (340), e.g., when audio mixing parameters are user-defined or customized by the user. In examples, audio mixing parameters may be based on the user's field of view (FoV), e.g., audio streams for overlays that are within the user's field of view (FoV) may be mixed at a greater loudness compared to other streams. Audio mixing parameters (audio mixing weights) may be negotiated by the sender device, the user device, and the network-based media processing server (340) via SDP signaling.
In an embodiment, for example, if an end user supports multimedia telephony services for the Internet Protocol Multimedia Subsystem (MTSI) but does not support MTSI Immersive Teleconferencing and Telepresence for Remote Terminals (ITT4RT), the network-based media processing server (340) may mix both the audio and video to generate mixed audio and video and provide a media stream carrying the mixed audio and video to the end user device, thereby providing backward compatibility for the MTSI terminal.
In other embodiments, for example when the end-user device has limited capabilities, the network-based media processing server (340) may mix both audio and video to generate mixed audio and video and provide a media stream carrying the mixed audio and video to the end-user device.
In other embodiments, when the network-based media processing server (340) has limited capabilities and some end user devices are limited capability MITI devices, the network-based media processing server (340) can mix both audio and video from the same sender device to generate mixed audio and video and provide a media stream carrying the mixed audio and video to the end user devices that are limited capability MSTI devices.
In another embodiment, the network-based media processing server (340) can use SDP signaling to negotiate a common set of settings for audio mixing with all or a subset of the end-user devices that are MSTI devices, the common set of settings being for immersive media and a single video composition of various overlay media. Based on the common set of settings, the network-based media processing server (340) can then perform audio mixing and/or video mixing to generate mixed audio and video, and provide media streams carrying the mixed audio and video to all or a subset of the end-user devices that are MSTI devices.
4 shows a flow chart illustrating a process (400) according to an embodiment of the present disclosure. In various embodiments, the process (400) may be performed by a processing circuit within a device, such as the processing circuit of the user device (120), the user device (130), the user device (220), the user device (230), the user device (320), the user device (330), the network-based media processing server (340), etc. In some embodiments, the process (400) is implemented with software instructions, such that the processing circuit performs the process (400) when the processing circuit executes the software instructions. The process starts at (S401) and proceeds to (S410).
At (S410), a first media stream carrying a first audio and a second media stream carrying a second audio are received.
At (S420), a first audio weight for weighting the first audio and a second audio weight for weighting the second audio are received.
At (S430), the weighted first audio based on the first audio weights and the weighted second audio based on the second audio weights are combined to generate a mixed audio.
In some examples, the device is a user device, and the processing circuitry of the user device can receive the first and second audio weights determined by, for example, a host device for immersive content (e.g., subsystems (110), (210A)-(210Z), (310A)-(310Z)), and the user device can play the mixed audio through a speaker associated with the user device. In examples, to customize the audio weights, the user device can send customization parameters to the host device such that the host device customizes the first and second audio weights based on the customization parameters.
In some examples, the host device may determine the first audio weight and the second audio weight based on the intensity of the first audio and the second audio.
In some examples, the first audio and the second audio are overlay audio, and the host device can determine the first audio weight and the second audio weight based on an overlay priority of the first audio and the second audio.
In some examples, the host device may determine the first audio weighting and the second audio weighting based on the detection of an active speaker.
In some examples, the first media stream includes immersive media content and the second media stream corresponds to overlay media content, and the host device may determine the first audio weighting to be different from the second audio weighting.
In some embodiments, the process (400) is performed by a network-based media processing server that performs media processing offloaded from the user device. The network-based media processing server can encode the mixed audio into a third media stream and transmit the third media stream to the user device. In some examples, the process (400) is performed by a network-based media processing server that performs overlay media processing offloaded from the user device. The network-based media processing server can transmit the third media stream and a fourth media stream that includes immersive media content. The third media stream includes overlay media content to the immersive media content.
The process proceeds to (S499) and ends.
5 shows a flow chart illustrating a process (500) according to an embodiment of the present disclosure. In various embodiments, the process (500) may be performed by a processing circuit in a device for network-based media processing, such as a network-based media processing server (340). In some embodiments, the process (500) is implemented with software instructions, such that the processing circuit performs the process (500) when the processing circuit executes the software instructions. The process starts at (S501) and proceeds to (S510).
At (S510), a first media stream carrying a first media content and a second media stream carrying a second media content are received.
At (S520), a third media content is generated that mixes the first media content and the second media content.
In some examples, a first audio in the first media content is mixed with a second audio in the second media content to generate a third audio, the first audio being weighted based on a first audio weight assigned to the first audio and the second audio being weighted based on a second audio weight assigned to the second audio. In examples, the first audio weight and the second audio weight are determined by a host device providing the immersive media content and transmitted from the host device to a network-based media processing server.
In an example, the first media stream is an immersive media stream and the second media stream is an overlay media stream, in which case the first audio weight and the second audio weight are different in value.
In an example, the first media stream and the second media stream are overlay media streams, and the first audio weight and the second audio weight are equal in value.
In another example, the first media stream and the second media stream are overlay media streams, and the first audio weight and the second audio weight depend on the overlay priority of the first media stream and the second media stream.
At (S530), a third media stream carrying a third media content is transmitted to the user device.
The process then proceeds to (S599) and ends.
6 shows a flow chart illustrating a process (600) according to an embodiment of the present disclosure. In various embodiments, the process (600) may be performed by a processing circuit in a host device for immersive media content, such as the processing circuit of subsystems (110), (210A)-(210Z), (310A)-(310Z), etc. In some embodiments, the process (600) is implemented with software instructions, such that the processing circuit performs the process (600) when the processing circuit executes the software instructions. The process starts at (S601) and proceeds to (S610).
At (S610), a first media stream carrying a first audio and a second media stream carrying a second audio are transmitted.
At (S620), a first audio weight for weighting the first audio and a second audio weight for weighting the second audio are determined.
In some examples, the host device receives the customization parameters based on a session description protocol and determines the first audio weighting and the second audio weighting based on the customization parameters.
In some examples, the host device determines the first audio weight and the second audio weight based on the intensity of the first audio and the second audio.
In some examples, the first audio and the second audio may be divided into overlay audio, and the first audio weight and the second audio weight may be determined based on the host device and the overlay priority of the first audio and the second audio.
In some examples, the host device determines the first audio weighting and the second audio weighting based on detection of an active speaker in one of the first audio and the second audio.
In some examples, the first media stream includes immersive media content and the second media stream includes an overlay media stream, and the host device determines different values for the first audio weight and the second audio weight.
At (S630), the first audio weights and the second audio weights are transmitted for mixing the first audio with the second audio.
The process then proceeds to (S699) and ends.
The techniques described above may be implemented using computer readable instructions and as computer software physically stored on one or more computer readable media. For example, Figure 7 illustrates a computer system (700) suitable for implementing certain embodiments of the disclosed subject matter.
The computer software may be coded using any suitable machine code or computer language that may be subjected to assembly, compilation, linking, or similar mechanisms to produce code including instructions, microcode, etc. that may be executed directly or through interpretation by one or more central processing units (CPUs), etc.
The instructions may be executed in a variety of computers and components thereof, including personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.
7 for computer system (700) are illustrative of course and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components depicted in the exemplary embodiment of computer system (700).
The computer system (700) may include certain human interface input devices. Such human interface input devices may respond to input by one or more users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
The human interface input devices may include one or more of a keyboard (701), a mouse (702), a trackpad (703), a touch screen (710), a dataglove (not shown), a joystick (705), a microphone (706), a scanner (707), and a camera (708) (only one of each is shown).
The computer system (700) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the user's senses, for example, through haptic output, sound, light, and smell/taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (710), data glove (not shown), or joystick (705), although there may also be haptic feedback devices that do not function as input devices), audio output devices (e.g., speakers (709), headphones (not shown)), visual output devices (e.g., screens including CRT screens, LCD screens, plasma screens, OLED screens, each of which may or may not have touch screen input capabilities and each of which may or may not have haptic feedback capabilities (some of which may be capable of outputting two-dimensional visual output or output in greater than three dimensions through means such as stereoscopic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
The computer system (700) may also include human accessible storage devices and associated media, such as optical media or similar media (721), including CD/DVDs as well as CD/DVD ROM/RW (720), thumb drives (722), removable hard disks or solid state drives (723), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM/ASIC/PLD based devices such as security dongles, and the like.
Those skilled in the art should also understand that the term "computer-readable medium" as used with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.
The computer system (700) may also include an interface (754) to one or more communication networks. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, and the like. Examples of networks include local area networks such as Ethernet, including SGSM, 3G, 4G, 5G, LTE, and the like, WLAN, cellular networks, TV wired or wireless wide area digital networks, including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks, including CANBus, and the like. Certain networks typically require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (e.g., a USB port on the computer system (700)), while others are typically built into the core of the computer system (700) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using any of these networks, computer system (700) can communicate with other entities. Such communications can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., CANbus to a specific CANbus device), or bidirectional, for example, to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks can be used with each of these networks and network interfaces described above.
The above-mentioned human interface devices, human accessible storage devices, and network interfaces may be attached to a core 740 of the computer system 700 .
The core (740) may include one or more central processing units (CPUs) (741), graphics processing units (GPUs) (742), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (743), hardware accelerators for specific tasks (744), graphics adapters (750), etc. These devices may be connected through a system bus (748), along with internal mass storage devices (747) such as read only memory (ROM) (745), internal non-user accessible hard drives, SSDs, etc. In some computer systems, the system bus (748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (748) or through a peripheral bus (749). In an example, a screen (710) may be connected to the graphics adapter (750). Architectures for peripheral buses include PCI, USB, etc.
The CPU (741), GPU (742), FPGA (743), and accelerator (744) may execute certain instructions that may combine to constitute the above-mentioned computer code. The computer code may be stored in ROM (745) or RAM (746). Persistent data may be stored, for example, in an internal mass storage device (747), while transient data may also be stored in RAM (746). Rapid storage and retrieval from any of the memory devices may be made possible through the use of cache memory, which may be closely associated with one or more of the CPU (741), GPU (742), mass storage device (747), ROM (745), RAM (746), etc.
The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.
By way of example, and not by way of limitation, a computer system (700) having the architecture shown in FIG. 7, and specifically the core (740), may provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with the mass storage accessible by the user described above, as well as specific storage of the core (740) that is non-transitory in nature, such as the core's internal mass storage (747) or ROM (745). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (740). The computer-readable media may include one or more memory devices or chips, depending on the particular needs. The software may cause the core (740) and specifically the processor therein (including a CPU, GPU, FPGA, etc.) to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM (746) and modifying such data structures according to the processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerator (744)) that may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software may encompass logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) that store software to be executed, circuitry that embodies logic to be executed, or both, where appropriate. The present disclosure includes any suitable combination of hardware and software.
While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of the disclosure. Thus, those skilled in the art will appreciate that numerous systems and methods will be able to devise that embody the principles of the disclosure and are thus within its spirit and scope, even if not explicitly shown and described herein.
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| WO2019229299A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP2004072354A | Cites | Japan |
| JP2009159367A | Cites | Japan |
| JP2004534457A | Cites | Japan |
| JP2005045737A | Cites | Japan |
31 members in 6 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 63088300 | United States of America | – | |
| 202063088300 | United States of America | P | |
| 63124261 | United States of America | – | |
| 202063124261 | United States of America | P | |
| 17327400 | United States of America | – | |
| 202117327400 | United States of America | A | |
| 2021038370 | United States of America | W |
Members31
| Document | Office | Kind | |
|---|---|---|---|
| US2022107779A1 | United States of America | A1 | |
| US2022109758A1 | United States of America | A1 | |
| WO2022076046A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2022076183A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20220080184A | Republic of Korea | A | |
| CN114667727A | China | A | |
| EP4042673A1 | European Patent Office (EPO) | A1 | |
| EP4042673A4 | European Patent Office (EPO) | A4 | |
| JP2023508130A | Japan | A | |
| KR20230048107A | Republic of Korea | A | |
| EP4165830A1 | European Patent Office (EPO) | A1 | |
| CN116018782A | China | A | |
| US11662975B2 | United States of America | B2 | |
| US2023229384A1 | United States of America | A1 | |
| JP2023538548A | Japan | A | |
| US11847377B2 | United States of America | B2 | |
| EP4165830A4 | European Patent Office (EPO) | A4 | |
| KR102626555B1 | Republic of Korea | B1 | |
| US11914922B2 | United States of America | B2 | |
| US2024069855A1 | United States of America | A1 | |
| JP7521112B2 | Japan | B2 | |
| CN116018782B | China | B | |
| EP4042673B1 | European Patent Office (EPO) | B1 | |
| JP7548488B2This record | Japan | B2 | |
| JP2024133169A | Japan | A | |
| KR102774376B1 | Republic of Korea | B1 | |
| KR20250035594A | Republic of Korea | A | |
| US12254242B2 | United States of America | B2 | |
| CN114667727B | China | B | |
| JP7765559B2 | Japan | B2 | |
| CN121056440A | China | A |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 7548488
- Application
- 2022535698
Titles2
- Japanese
- テレカンファレンスの方法
- English
- How to Teleconference
Classification
- CPC, 22
- H04L65/403
- G06F3/165
- H04N7/15
- G10L19/008
- H04L65/764
- H04L65/762
- H04L65/70
- H04N19/503
- H04N19/597
- H04N19/46
- H04N7/147
- H04N7/152
- H04M3/568
- H04M2201/50
- H04M3/567
- H04N21/233
- H04N21/439
- H04N21/4341
- H04N21/2368
- H04N21/816
- H04N19/51
- H04R5/04
- IPC, 2
- H04N7 15
- H04M3 56
