Systems and methods for scalable distributed global infrastructure for real-time multimedia communication
Abstract
The present invention provides a new solution, which relates to a system and method that supports the operation of a virtual media room or virtual meeting room (VMR), where each VMR can receive multiple videos from multiple participants in different geographical locations Multiple video conference sources of the conference terminal, the video conference terminal may be dedicated or standardized, and they implement a multi-party video conference session between multiple participants. A globally distributed infrastructure supports the operation of the virtual conference room through multiple multi-party control units (MCUs) as media processing nodes, where the multi-party control units are constructed based on non-customized components rather than customized hardware, and each multi-party control unit Used for real-time processing of multiple audio and video streams from the multiple video conference terminals. Each individual VMR can span the infrastructure of a series of globally distributed servers and media processing nodes co-located at Internet access points (POPS). This large-scale distributed architecture can support thousands of active VMRs at the same time and it is also transparent to the users of VMRs.

Term
4.6 yearsto projected expiry
Projected expiry 12 May 2031, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
24 claims: 2 independent, 22 dependent
- 1一个系统,包括: 一个虚拟会议室(VMR)引擎,在其运行时,接受来自多个视频会议终端的多个音频和 视频流,其中每个视频会议终端与一个视频会议的多个参与者中的一个关联; 支持所述虚拟会议室的运行的一个全球分布的基础设施,包括多个作为媒体处理节点 的多方控制单元(MCU),每个多方控制单元用于实时处理来自所述多个视频会议终端的多 个音频和视频流,其中所述多方控制单元基于非定制部件而不是定制硬件构建。
- 2根据权利要求1中所述的系统,其特征在于: 所述非定制部件包括Linux/x86的中央处理器和PC的图形处理单元中的一个或多个。
- 3根据权利要求1中所述的系统,其特征在于: 所述全球分布的基础设施以堆架式云计算风格部署所述多方控制单元,以获得较高的 可扩展性和性价比。
- 4根据权利要求1中所述的系统,其特征在于: 所述MCU支持所述多个音频和视频流的可伸缩混合和组合。
- 5根据权利要求1中所述的系统,其特征在于: 所述全球分布的基础设施为所述MCU将本地局域网上的以及跨地域的非定制部件集 群,以实现无限扩展。
- 6根据权利要求5中所述的系统,其特征在于: 所述集群的MCU利用网络或应用层多路传送和多比特率的数据流分布方案来实现无 限扩展。
- 7根据权利要求1中所述的系统,其特征在于: 所述全球分布的基础设施在第三方数据中心通过接入点(POP)方式在全球分配所述 MCU以处理有着不同通信协议的视频会议终端的音频和视频流,其中每个POP有着所要求 的能够处理来自所述接入点所在的地理区域的负载的处理能力。 根据权利要求7中所述的系统,其特征在于: 所述全球分布的基础设施将来自连接到所述视频会议的每个参与者的音频和视频流 引导至最近的POP以便最小化他们的延迟时间。
- 89. 根据权利要求7中所述的系统,其特征在于: 所述全球分布的基础设施分配来自每个接入点的本地局域网内的或在广域网上跨越 多个接入点的参与者的音频和视频流。
- 910. 根据权利要求7中所述的系统,其特征在于: 所述全球分布的基础设施允许一个或多个其它全球分布的专用网络去连接它,其包括 对视频会议服务的部署,所述视频会议服务要求在边缘节点处联合以及几个通讯和传输协 议的转换和解码。
- 1011. 根据权利要求7中所述的系统,其特征在于: 所述全球分布的基础设施限制来自所述视频会议的每个参与者的音频和视频通过接 入点的两个跃点。
- 1112. 根据权利要求1中所述的系统,其特征在于: 所述全球分布的基础设施通过将协议控制信息和音频与视频流的处理隔离在一个或 多个单独的、独立的、无特权的进程中以提供一个高级的容错协议处理机制以防止不当的 输入造成不稳定和安全漏洞。
- 1213. 根据权利要求1中所述的系统,其特征在于: 所述全球分布的基础设施支持基于互联网的客户端-服务器体系结构的分布式容错 消息传送。
- 1314. 根据权利要求1中所述的系统,其特征在于: 所述全球分布的基础设施使得传统视频会议终端能够无缝穿越防火墙,所述传统视 频会议终端执行标准化协议,并未采用可供视频会议终端与防火墙外通讯的公共互联网地 址。
- 1415. 一个方法,包括: 接受来自多个视频会议终端的多个音频和视频流,其中每个视频会议终端与一个虚拟 视频会议的多个参与者中的一个关联; 支持虚拟会议室的运行的一个全球分布的基础设施,包括多个作为媒体处理节点的多 方控制单元(MCU),每个多方控制单元用于实时处理来自所述多个视频会议终端的多个音 频和视频流,其中所述多方控制单元基于非定制部件而不是定制硬件构建。
- 1516. 根据权利要求15中所述的方法,其特征在于:还包括: 以堆架式云计算风格部署所述多方控制单元,以获得较高的可扩展性和性价比。
- 1617. 根据权利要求15中所述的方法,其特征在于:还包括: 支持所述多个音频和视频流的可伸缩混合和组合。 1 根据权利要求15中所述的方法,其特征在于:还包括: 为所述MCU将本地局域网上的以及跨地域的非定制部件集群,以实现无限扩展。
- 1719. 根据权利要求15中所述的方法,其特征在于:还包括: 在第三方数据中心通过接入点(POP)方式在全球分配所述MCU以处理有着不同通信协 议的视频会议终端的音频和视频流,其中每个POP有着所要求的能够处理来自所述接入点 所在的地理区域的负载的处理能力。
- 1820. 根据权利要求19中所述的方法,其特征在于:还包括: 将来自连接到所述视频会议的每个参与者的音频和视频流引导至最近的POP以便最 小化他们的延迟时间。
- 1921. 根据权利要求19中所述的方法,其特征在于:其还包括: 分配来自每个接入点的本地局域网内的或在广域网上跨越多个接入点的参与者的音 频和视频流。
- 2022. 根据权利要求19中所述的方法,其特征在于:还包括: 允许一个或多个其它全球分布的专用网络去连接它,其包括对视频会议服务的部署, 所述视频会议服务要求在边缘节点处联合以及几个通讯和传输协议的转换和解码。
- 2123. 根据权利要求19中所述的方法,其特征在于:其还包括: 限制来自所述视频会议的每个参与者的音频和视频通过接入点的两个跃点。
- 2224. 根据权利要求15中所述的方法,其特征在于:其还包括: 通过将协议控制信息和音频与视频流的处理隔离在一个或多个单独的、独立的、无特 权的进程中以提供一个高级的容错协议处理机制以防止不当的输入造成不稳定和安全漏 洞。
- 2325. 根据权利要求15中所述的方法,其特征在于:其还包括: 支持基于互联网的客户端-服务器体系结构的分布式容错消息传送。
- 2426. 根据权利要求15中所述的方法,其特征在于:其还包括: 使得传统视频会议终端能够无缝穿越防火墙,所述传统视频会议终端执行标准化协 议,并未采用可供视频会议终端与防火墙外通讯的公共互联网地址。
Independent claims24
170 paragraphs, as filed
System and method for scalable distributed global infrastructure in real-time multimedia communication Background technology
[0001] In recent years, video conferencing has been rapidly developed in enterprises. This is because business has become more and more globalized, and employees are connected to a large-scale system consisting of remote employees, partners, suppliers, and customers. To interact. At the same time, in the consumer sector, the availability of cheap software solutions and the widespread use of cameras in laptops and mobile devices have all contributed to the use of video chat to keep in touch with family and friends.
[0002] However, the option screens that can be used for video conferences are still divided into isolated states, and they cannot communicate well with each other. There are also hardware-based conference rooms within the enterprise, equipped with video conferencing systems provided by vendors such as Polycom, TANDBERG, and LTV, as well as high-end telepresence systems promoted by Cisco. At the low end of the price range are software based on enterprise video conferencing applications, such as Microsoft's Lyne and such as Skype, GoogleTalk and Apple's FaceTime.
[0003] When choosing to use any of the above-mentioned video phone systems, there are great trade-offs in terms of price, quality, and scope. In order to achieve low-latency, high-definition calls, large companies invest hundreds of thousands of dollars in their telepresence systems, but they can only cover a small number of people who access the similar systems. Small and medium enterprises invest tens of thousands of dollars in their hardware systems that can achieve high-definition resolutions of up to 720. They buy multi-party conference units with a fixed number of ports worth hundreds of thousands of dollars and communicate between their different branches through these multi-party conference units. However, it is easy to communicate with the external system of the company when it is done. Very confused at the time. The company cannot afford to use such a high-input and low-quality solution experience using the Skype client, but on the other hand it can easily connect with other people, whether they are inside or outside of their company. Ordinary users will have concerns when they find that using these video conferences is too complicated and incomprehensible compared to using mobile phones or fixed phones that do not need to consider these concerns but just "work". Therefore, although the technology of video conferencing is feasible and most people can afford the price, the proportion of video conferencing used in business is low.
[0004] Today, people need more than ever before to eliminate this concern and provide a high-quality video call at almost the same price as a voice call without the user's consideration of complicated concerns. Such a service will connect the hardware and software of different video conferences and chat systems from different vendors. When talking to each other, the video call involves different protocols (H·323, SIP, XMPP, proprietary ) And have different video and audio codecs. The video call will provide lower latency and better visual experience than existing solutions. It will be hosted on the Internet or the cloud, so the company no longer needs to spend a lot of money and operational investment in complex equipment. The ease of use will be as simple as installing audio conference calls, without complicated preparation arrangements from the company's IT.
[0005] With regard to other objects, features and advantages of the present invention, the following will be described in detail in the specific embodiments with reference to the accompanying drawings.
Description of the drawings
[0006] FIG. 1 depicts an example of a virtual meeting room (Virtual Meeting Room, VMR) operating system that spans multiple standards and a proprietary video conferencing system.
[0007] Figure 2 depicts a virtual meeting room operating process flow across multiple standards and proprietary video conferencing systems
Diagram example.
[0008] FIG. 3 depicts an example of various components of a media processing node.
[0009] FIGS. 4A-4B describe a diagram of an embodiment of media encoding in a simple dyadic scene.
[0010] Figure 5 depicts an example of a highly scalable audio mixing and sampling chart.
[0011] FIG. 6 depicts an example of an audio sound wave echo canceller based on Internet/cloud technology.
[0012] FIG. 7 depicts an example of a multi-phase media stream distribution process between each post office protocol in a local area network or between multiple post office protocols (POPs) across a wide area network.
[0013] FIG. 8 depicts an example of the software components of a global infrastructure engine that supports virtual meeting rooms (VMR).
[0014] FIG. 9 depicts an embodiment of a high-level mechanism for a fault-tolerant protocol, where the fault-tolerant protocol is used to prevent erroneous input to avoid instability or security breaches.
[0015] FIG. 10 depicts an example illustrating firewall traversal technology.
[0016] FIG. 11 depicts an example illustrating the management and control of a video conference.
[0017] FIG. 12 depicts an example illustrating high-quality event sharing using a multi-party conference unit (MCU) of a global infrastructure engine.
[0018] FIG. 13 depicts an example of a method of associating a laptop computer or a mobile phone with a conference room system.
[0019] FIG. 14 depicts an example illustrating the provision of welcome screen content to participants.
[0020] FIG. 15 depicts a graphical example of a personalized video conference room on a per-call basis.
[0021] FIG. 16 depicts an example of a single online "home page" used to personalize the sharing of someone's desktop, laptop, and/or mobile phone screen.
[0022] FIG. 17 depicts an example of one-click video conference login through a mailbox.
[0023] FIG. 18 depicts an example illustrating the delivery of a virtual reality experience to participants.
[0024] FIG. 19 depicts an example illustrating the provision of enhanced real user interaction services to participants.
Detailed ways
[0025] The method of the present invention is described by way of example and is not limited to the accompanying drawings, where the same reference numerals correspond to the same elements. It is worth noting that the usage of "a" or "an" or "some" in the embodiments of the present disclosure does not necessarily refer to the same embodiment, and such references mean at least one.
[0026] The present invention provides a new method, the new method has a mature system and method to support a virtual media room or virtual meeting room (Virtual Meeting Room, VMR) operation, wherein each (VMR) can receive from Multiple participants in different geographic locations receive multiple video conferencing sources of audio, video, presentation, and other media streams from video conferencing terminals and other multimedia function devices. The other multimedia function devices may also be proprietary or It is standards-based and enables multi-party communication or point-to-point video conferencing between most participants. For a non-limiting example, video sources from video conferencing terminals include but are not limited to Skype, while video sources from standard video conferencing terminals include but are not limited to H. 323 and SIPO. Each single VMR can be distributed as Global infrastructure support for a series of commodity servers, which serve as media processing nodes located at the access point of Internet access, where this large-scale distributed architecture can simultaneously support production in a way that does not require reservations. Thousands of active VMRs are also transparent to the users of VMRs. Each VMRs provides its users with a rich set of meetings and collaboration interactions that have never been experienced by the participants of a video conference so far. The interactions described include
Conversation control, configuration, visual layout of conference participants, VMR customization, and adaptation of conference rooms to adapt to different participants. For a non-limiting example, such a VMR is used for point-to-point calls between two different terminals such as a Skype client and a standard H·323 terminal, where the Skype user does not know the terminal technology of other users. In the case of initiating a call to another, after determining the necessity of switching between the two terminals, a VMR is automatically established between the two parties.
[0027] The method further utilizes virtual reality and augmented reality technologies to transform video and audio streams from participants in various customized ways to achieve a series of rich user experiences. The globally distributed infrastructure supports the sharing of events between participants in geographically dispersed locations through a multipoint control unit for real-time processing of multiple audio and video streams from multiple video conference terminals.
[0028] Compared with the traditional video conference system that requires each video conference participant to comply with the same communication standard or protocol, the VMR of the present invention allows the user or participant of the video conference to use independent style equipment and protocol to participate in multi-party or Point-to-point video conference sessions. By performing transparent operations on video and audio streams in the Internet or the cloud without the intervention of end users, the method combines video conferencing systems with different devices and different protocols and various video chats that exist in the world today to make it one The overall system.
[0029] Hosting VMR on the Internet or the cloud enables participants to transparently initiate calls to anyone, and call them through VMR in all registered terminal devices, and allow the callee to use any terminal device they wish to call Answer the call. A VMR hosted on the Internet or in the cloud enables any participant to upload media content to the cloud and rebroadcast it to other participants in the format of their choice, with or without modification.
[0030] FIG. 1 depicts an example of a virtual meeting room (VMR) operating system 100 that spans multiple standards and proprietary video conferencing systems. Although the components described in the figure are functionally independent, this description is only for ease of explanation. Obviously, the description of the components in the figure can be arbitrarily combined or divided into separate software, firmware, and/or hardware. In addition, obviously, any combination or separation of such components can be executed on the same one or more hosts, or several virtualized instances on one or more hosts, where one or Multiple hosts can be connected to one or more networks distributed around the world.
[0031] As shown in FIG. 1, the system 100 includes at least one VMR engine 102 for operating the VMRs, a global infrastructure engine 104 for supporting the operation of the VMRs, and a user experience engine 106 for enhancing the user experience of VMRs. .
[0032] The engine used here refers to software, firmware, hardware, or other components used to accomplish a purpose. The engine mainly includes software instructions stored in non-volatile memory (also referred to as auxiliary memory). When the software instructions are executed, at least a subset of the software instructions are input by the processor into the memory (also referred to as the main memory). The processor executes the software instructions in the memory. The processor may be a shared processor, a dedicated processor, or a combination of shared and dedicated processors. A typical program includes calls to hardware components (such as input/output devices), which usually require the execution of drivers. The driver may or may not be considered part of the engine, but the difference is not the main one.
[0033] As shown in FIG. 1, each engine runs in one or more hosting devices (hosts). Wherein, the host may be a computer device, a communication device, a storage device or any electronic device with the ability to run software components. For a non-limiting example, a computer device can be, but is not limited to, a laptop computer, a desktop computer, a tablet computer, an Apple tablet computer, an Apple multimedia player, an Apple mobile phone, an Apple MP4 player, Googles Android device, PDA, or a server machine, the server is
A physical or virtual server that is hosted by a service provider or a third party provided by the service provider in the Internet public or private data center, or located in the private information center or office of the enterprise. The storage device can be, but is not limited to, a hard disk drive, a flash memory drive, or any portable storage device. The communication device can be, but is not limited to, a mobile phone.
[0034] As shown in FIG. 1, each of the VMR engine 102, the global infrastructure engine 104, and the user experience engine 106 has one or more communication interfaces (not shown), and these communication interfaces are software components that enable the The engines communicate with each other through one or more communication networks according to a certain communication protocol such as TCP/IP. Wherein, the communication network may be, but is not limited to, the Internet, an intranet, a wide area network, a local area network, a wireless network, a Bluetooth, a wireless local area network, and a mobile communication network. The physical connection and communication protocol of the network are well known to those skilled in the art.
[0035] FIG. 2 depicts an example of a flow chart of a virtual meeting room operating process that spans multiple standards and proprietary video conferencing systems. For the sake of illustration, although the figure depicts the functional steps in a specific order, the process of the present invention is not limited to any specific or arranged steps. Those skilled in the art should know that the different steps described in this figure can be omitted, rearranged, combined and/or adapted in various ways.
[0036] As shown in FIG. 2, the flowchart 200 starts at block 202, which receives multiple video conferencing sources from multiple video conferencing terminals from multiple participants, wherein each video conferencing terminal communicates with all video conferencing sources. A certain participant among the multiple participants in the virtual conference room is associated. The flowchart 200 is followed by block 204, which describes that for each participant of the VMR, the multiple video conference sources are converted and mixed into a composite video and video conference terminal compatible with the video conference terminal associated with the participant. Audio streaming. The flowchart 200 is followed by a block 206, which describes the initiation of a multi-party video conference session that can be initiated in real time among multiple participants, wherein the types of the multiple video conference terminals are different. The flowchart 200 ends in block 208, which describes the presentation of composite audio and video streams to each VMR participant to enhance the user experience.
Virtual meeting/media room (VMR)
[0037] As shown in FIG. 1, the VMR engine 102 allows participants to enter and participate in a video conference room through all types of video conference terminals. The VMR engine 102 combines video equipment and/or video conferencing systems provided by a variety of different manufacturers with video conferencing sources implemented by the software of the terminal in real time, so as to effectively process the multi-party video conferencing. More specifically, the VMR engine 102 converts and combines multiple video conference sources from VMR participants in real time into a composite video and audio stream compatible with each video conference terminal, such as all the video and audio streams associated with each VMR participant. The video conference system. Wherein, the conversion of the video conference source includes at least one or more of the following aspects of the video conference source: Video coding format (such as H. 264, proprietary, etc.) Video coding profile and level (such as H. 264, proprietary, etc.) 264 main level, H. 264 constrained baseline contour, etc.) Audio coding format (such as SILK, G7xx, etc.) Communication protocol (such as H. 323, SIP, XMPP, proprietary, etc.) Video resolution (such as QC/SIF, C/SIF, QA/GA, High Definition-720p/1080p, etc.) Screen ratio (such as 4:3, 19: 9. Customization, etc.) The bit rate of the audio stream (narrow band, wide band, etc.) The bit rate of the video stream (1.5Mbps, 768kbps, etc.) Encryption standard (the new standard for symmetric encryption, AES, ownership, etc.) ) Acoustic design (such as echo cancellation, noise reduction, etc.) In order to transform the video conference source, the technology involved in the present invention includes but not limited to code conversion, frequency up, frequency down, conversion, mixing
Combine, add and remove video, audio and other multimedia streams, noise reduction, and automatic gain control (AGC) of video conferencing sources. [0038] In some embodiments, the VMR engine 102 runs on two different terminals such as Skype. The client and a standard H. A point-to-point call is established between the 323 terminals, where the Skype user initiates a call to another user without knowing the terminal technology of the other user, and after determining the necessity of the required conversion between the two terminals, the Skype user automatically communicates between the two terminals. Provide a virtual meeting room between parties. In this case, the VMR is used to allow communication between two terminals, There is no need for users to know or worry about the differences between the protocols, video encoding, audio encoding, or other technologies used by the terminal. [0039] In some embodiments, the VMR engine mixes and renders the composite video and audio streams to closely match the performance of the video conference terminal associated with the multiple participants, so that the participants can get a more efficient conference experience . When synthesizing the finally presented video and audio stream frames, the VMR engine 102 may consider the innovative video layout of the participants and the activity of various participants in the video conference. For a non-limiting example, the VMR may highlight active speakers relative to other participants. In some embodiments, the VMR engine 102 may also mediate the multimedia data stream/content accompanying the video conference source into a part of the composite audio/video stream for collaborative processing, wherein the multimedia data stream will include, but is not limited to, a slideshow Film sharing, electronic whiteboard, media streaming and desktop display screen. It also supports chat style information for real-time communication between participants. The status information of the participants includes, but is not limited to, the type of terminal used, the quality of the received signal, and the soft status of the audio/video that can be displayed to all participants.
[0040] In some embodiments, the media processing node 300 is used to convert and combine several video and audio streams of video conference sources in real time to create and present one or more composite media streams for each participant of the VMR. As shown in FIG. 3, the media processing node 300 may include one or more of the following components: a video synthesizer 302, a video transcoder 304, a distributed multi-channel video switch 306, an audio transcoder/preprocessor 308, a distributed Multi-channel multi-channel audio mixer 310, protocol connector 312 and a distributed conference controller 314. In the case of video, the following three (or more) forms of video streams from participants are available to the media processing node 300: Original compressed video; Uncompressed original video; Low-resolution compressed abbreviation video.
[0041] As shown in FIG. 3, the video synthesizer 302 in the media processing node 300 selects which video stream it needs based on the video that needs to be composed and presented to the participants. The two or more compressed video streams listed above are converted by the video transcoder 304 and sent through the distributed multi-channel video switch 306 using multiple addresses on the Internet, so that other remote users want these video streams Media processing nodes can subscribe to them as needed. This solution allows nodes (local and global) in the entire cluster to share and/or exchange audio and video streams in the most efficient manner. The data stream can be transmitted through a public network, a private network, or through a pre-allocated overlay network with service level guarantee. Using this method, the video synthesizer 302 can display various synthesis results, including but not limited to only the active speaker, two people in conversation displayed side by side, and any other custom formats required by the participants. It may include converting the video into other forms of presentation.
[0042] As shown in FIG. 3, the video transcoder 304 in the media processing node 300 effectively encodes and decodes the composite video stream, where the characteristics of each different data stream will be extracted during the decoding process. Here, the video transcoder 304 collects the information provided by the encoded bitstream of the composite video, where the collected information includes but is not limited to: Motion Vectors (MVs); CBP and skipping; Static macroblocks ( 0 motion vector and no CBP);
Quantization value (Qp);
Frame rate.
The features described are used to create a metadata field related to the uncompressed video stream and the compressed composite stream or other transformed data.
[0043] In some embodiments, the video synthesizer 302 not only combines the original video streams into a composite video stream, but also creates a composite metadata field to summarize the same operations in the metadata field ( (Including 2D and 3D operations) applied to the individual video streams of the composite video. For a non-limiting example, the motion vector needs to use the same transformation, so that the video synthesizer 302 can be applied to each original video stream. The motion vector includes, but is not limited to, scaling, rotation, translation, and cutting. operating. The metadata may be applied to other non-real-time multimedia services, where the multimedia services include, but are not limited to, recorded data streams and annotated data streams used for offline search and indexing.
[0044] FIGS. 4A-4B are diagrams of an embodiment in which the input media code is scaled down to one-half of the original in a dyadic vector scene, where FIG. 4A describes how to process macroblocks, and FIG. 4B describes How to deal with motion vectors. The video synthesizer 402 aligns the original video to the boundary of the macroblock as much as possible to achieve a good result. For this purpose, the more sensible choice of the video synthesizer 402 is to minimize the number of macroblocks covered by the video stream while maintaining the target size of the video stream. Optionally, the video synthesizer 402 can mark the border area of the video sources, because in the case of a video conference, these video sources usually have less information. The metadata field may include information such as the position of the speaker in the video stream, which information may be segmented and individually compressed as shown. The composite metadata field is processed to provide meaningful information on a macroblock basis and to best match the coding technology. For non-limiting embodiments, In the case of H.264, the processing considers the case where the macroblock is subdivided into 4X4 sub-blocks.
In the case of H.263, the macro block cannot be subdivided or can only be subdivided into 8X8 blocks according to the attachment used. In the case of H.261, the macro block is not subdivided.
[0045] In some embodiments, the video transcoder 304 is input into the composite original video and the composite metadata field, and then the video transcoder 304 uses the information provided by the metadata field to reduce calculations and focus on meaningful area. For non-limiting embodiments: Skip macro block detection: If the composite metadata points to a static megabyte MB, you can quickly skip macro block detection by selecting auto skip.
MV search range: The search range of the MV can be dynamically adjusted according to the synthetic metadata information, and the search range can be directly estimated according to the megabytes in the metadata field that matches the MV.
MV prediction: The MV displayed in the synthetic metadata is used for primary prediction during motion estimation.
Quantization: The quantization value used in the encoding process is constrained by the numerical value provided by the synthetic metadata field.
Frame rate adaptation: When the given frame rate is not updated, the area with lower frame rate composition will be marked and skipped.
The composite area without motion gets fewer bits.
The border area of each video is coded with fewer bits.
[0046] In the case of audio, the audio transcoder/pre-processor 308 combines the audio streams of each participant received at the media processing node 300 through the distributed multi-channel audio mixer 310 with those received by other participants. The audio streams are mixed together. The mixed output can also be sent via the network through the distributed multi-channel audio mixer 310, so that other nodes that want to receive this data stream can subscribe to it and mix it with the local data stream at their media processing node
Together. This method enables the global infrastructure engine 104 to provide mixed video output to participants in the VMR video conference in a decentralized manner.
[0047] In some embodiments, as shown in FIG. 5, the audio transcoder/preprocessor 308 can make it possible to mix audio signals from multiple codecs at an optimal sampling rate to obtain highly extended audio. . More specifically, the audio transcoder/preprocessor 308 first determines the best possible sampling rate to mix the audio based on the video terminal associated with the participant in the specific VMR. The audio transcoder/preprocessor 308 then estimates the noise entering each channel and determines the voice activity on each channel. Only the active channel is mixed to eliminate all noise of VMR, and the channel is balanced to enhance the signal and reduce noise. Finally, the audio transcoder/preprocessor 308 mixes all channels through channel standardization, and creates a unique data stream for each participant based on all other audio streams in the VMR to eliminate echo on the channel.
[0048] In some embodiments, the audio transcoder/preprocessor 308 can provide real-time language translation and other speech-to-text or speech-to-video conversion services. These services include, but are not limited to, English language translation and real-time subtitles. Translate, interact and modify content through voice commands during calls, and introduce data from the Internet in real time through voice-to-video services.
[0049] Due to the bad or non-existent echo cancellation on the side of one or more participants of the video conference often destroys the entire conference of all people in the VMR, in some embodiments, such as an embodiment shown in FIG. 6 The audio transcoder/preprocessor 308 can automatically determine the need for sound echo cancellation in a video conference terminal based on the Internet/cloud. First, the audio transcoder/preprocessor 308 determines the delay time of the round trip during which the audio stream comes out of the MCU to the terminal and back. Then the audio transcoder/preprocessor 308 estimates the long-term and short-term energy loss of the sound signal from the MCU on the speaker and microphone. The natural loss at the terminal can be calculated by the following formula:
ERL=logiO (energy (microphone signal)/energy (speaker signal)) If the natural loss of the terminal is greater than 24 decibels, there is no need to do echo cancellation.
Distributed infrastructure
[0050] For video conferencing, the traditional method is to establish an infrastructure. In order to meet these requirements, it is often necessary to use FPGA (Field Programmable Logic Array) and DSP (Digital Signal Processor) custom hardware to implement low-latency media processing. The hardware is linked together to handle large loads. Such a custom hardware system is not very flexible in the AA/format and communication protocol, because the hardware logic and DSP encoding are written and optimized for a specific set of AA/codecs and formats, so it can handle it. The establishment of this system is very expensive, it requires a large number of R & D teams and many years of design cycle and professional engineering technology.
[0051] Supporting the operation of the VMR engine 102 shown in FIG. 1 requires a multi-protocol video bridging solution, such as a multi-point control unit well-known in the industry, and the aforementioned media processing nodes to process and compose various types of The video conference source of the terminal. Traditionally, an MCU is built by custom hardware by combining a special FPGA (Field Programmable Logic Array) and a 100s DSP (Digital Signal Processor) together, and finally forms a tool in the expensive blade frame installation system. There are many multipoint control units for digital signal processor boards. Even with such an expensive system, when participants use HD video, it can only achieve a speed of 10s or 100s when connected to the MCU. In order to achieve greater scope, service providers have to purchase many of these blade boxes and put load balancers and custom scripts together. But this method is very expensive and difficult to implement. It is difficult to program the DSP software and FPGA code used, and it is difficult to seamlessly distribute globally. In addition, the system usually runs a dedicated operating system, which makes it difficult to add third-party software
Software and quickly provide new features and some functions, so that when the virtual meeting room spans multiple MCUs, the ability to provide flexible image synthesis for participants in the virtual meeting room is lost.
[0052] As shown in the embodiment shown in FIG. 1, the global infrastructure engine 104 can enable the use of off-the-shelf (off-the-shelf, or off-the-shelf) by establishing MCUs (Multipoint Control er Units, multi-party control units) as media processing nodes. Called off-the-shelf) components, such as Linux/x86 CPU (central processing unit) and PC GPU (image processing unit) instead of custom hardware to process media streams to efficiently and expandably process and synthesize media streams. These MCUs can be deployed in a rack-and-stack cloud computing style to achieve the most scalable and cost-effective method to support VMR services. In the past 5 years, the x86 architecture has been greatly improved in terms of digital signal processing (DSP) functions. In addition, the existing graphics processing unit (GPU) for rendering PC graphics can be used to enhance the processing power of the CPU.
[0053] As shown in FIG. 1, the global infrastructure engine 104 supports and enables the operation of the VMR to have at least one or more of the following attributes: The ability to support multiple audio formats and protocols; Scalable mixing and synthesis Audio and video streaming; Service delivery with minimal delay on a global scale; Establish capital efficiency and operating cost efficiency.
[0054] In some embodiments, the global infrastructure engine 104 can use x86 server clusters on the local area network and across regions as the media processing node 300 of the MCU to achieve near-infinite expansion. All media processing nodes 300 work together as a huge MCU. In some embodiments, for example, the design of these aggregated MCUs uses network layer multiplexing and a novel multi-bit rate data stream distribution scheme to achieve unlimited expansion. Under this design, the global infrastructure engine 104 can achieve great scalability according to the number of participants in each call, the geographic distribution of callers, and the distribution of calls across multiple POP protocols around the world.
[0055] In some embodiments, the global infrastructure engine 104 distributes the MCUs globally in a third-party data center through Points of Presence (POP) to process video conference sources of video terminals with different communication protocols. . Each POP has the required processing capacity (such as a server) that can handle the load from the geographic area where the access point is located. Users or participants are connected to the video conferencing system 100 and are directed by the global infrastructure engine 104 to the nearest POP (connector) so that they minimize their delay time. Once participants reach the POP of the global infrastructure engine 104, their audio and video streaming conference sources can continue on the high-performance network between the POPs. The global infrastructure engine 104 distributed in this way allows the largest media processing engine (VMR engine 102) ever built to operate as a single system 100. If you use the traditional DSP/FPGA-based custom hardware method to build, the system requires a lot of capital, research and development costs and huge matching operating scripts.
[0056] FIGS. 7A-7C describe an example of a multi-phase media stream distribution process all from a local area network of a POP or across multiple POPs on a wide area network. As shown in Figure 7A, it describes the media stream distribution stage 1, a single-node media distribution with a POP. In this stage, as a non-limiting example, a video conference source from a video conference participant passes through and runs H The 323 conference room system, the PC running H.323, and the PC running Skype are all connected to a node in the POP according to the principle of neighboring conference hosts. In this stage, the video conference source is load-balanced But it is not clustered between POP nodes. Figure 7B depicts the second stage of media stream distribution. The media distribution of an aggregation node with a POP is in which the video conference source from the participants is load-balanced among the nodes aggregated in the POP, and the audio/video stream is in the POP. The nodes are scattered/overflowing. Figure 7C describes the third stage of media stream distribution.
Media distribution is completed between the set node and different POPs. At this stage, some conference participants can connect to the nearest POP instead of a single POP.
[0057] In some embodiments, the global infrastructure engine 104 will allow multiple other globally distributed private networks to connect to it, including but not limited to the deployment of video conferencing services: for example, Microsofts Lycn chat tool requires Union at edge nodes (such as cooperation between multiple organizations), as well as conversion and decoding of several communication and transmission protocols.
[0058] In some embodiments, the global infrastructure engine 104 may limit the video conference source from each participant of a video conference to only pass two hops of media nodes and/or POPs in the system at most. However, it is possible to acquire other levels with intermediate media processing nodes that perform transcoding and transcoding. Using this solution, the global infrastructure engine 104 can provide participants with pseudo-scalable video coding associated with devices that do not support Scalable Video Coding (SVC). For example, each participant of a video conference supports an appropriate bit rate acceleration/ Reduced-capable audio and video coding (AVC). In a media distributed network, the global infrastructure engine 104 takes the AVC data streams and adapts them into multi-bitrate AVC data streams. According to the solution, SVC can still be used on devices that support SVC. With the increase in the use of SVC by participant client devices and the adoption and growth of such networks, it is also possible to use SVC in internal networks instead of multi-bitrates. AVC data stream.
[0059] FIG. 8 depicts an example of the software components of the global infrastructure engine 104 that supports virtual meeting rooms (VMR). Some of the components include, but are not limited to, a media gateway engine, a media processing engine, and a media gateway engine used to process code conversion, image synthesis, mixing, and echo cancellation between H. 26x, G. 7xx, and SILK. Multi-protocol connectors between SIP, XMPP, and intranet interconnection, such as conference control, screen and presentation sharing, chat, and other website applications, are distributed through these nodes and the POPs of the global infrastructure engine 104 to achieve real-time communication . Some of the components, including but not limited to user/account management, billing systems, and network operations center (NOC) systems for guidance, monitoring, and node management, are running on one or more centralized but redundant nodes. Other components include, but are not limited to, general application frameworks and platforms (such as Linux/x86 CPU, GPU, package management, aggregation), which can be run on distributed nodes and centralized management nodes
[0060] When the input to the server is accepted through an open network, especially untrusted resources, the input must be verified in order to prevent security vulnerabilities, denial of service attacks, and service instability. In the case of a video conference, it is necessary to verify an audio/video stream input from a conference terminal. The verification includes control protocol information and a compressed media stream, both of which must be verified. The encoding process of untrusted input is responsible for the verification and clear inspection of untrusted input before it enters the system for propagation. This is very important. History has shown that relying on this as the only verification countermeasure is not enough. For a non-limiting example, H. 323 makes extensive use of Abstract Syntax Notation One (ASN. 1) for encoding. Over the past few years, most public ASN. 1 implementations have some security issues. The complexity of ASN. 1 makes It is almost impossible to compile a completely safe parsing program by hand. For another non-limiting example, many implementations of H.264 video decoders do not include boundary checking for performance reasons, but instead include a system-specific codec restart when it performs an invalid storage read and triggers a failure. Code.
[0061] In some embodiments, FIG. 9 depicts an example illustration in which the global infrastructure engine 104 provides an advanced fault-tolerant protocol processing mechanism through the protocol connector 212 to prevent improper input from causing instability and possible security vulnerabilities. All processing protocol control information and the encoding of compressed audio and video streams are isolated in one or more separate, independent, unprivileged processes. More specifically, a separate process: each incoming connection should cause a new process to be created and processed by the protocol connector 212. The process will be responsible for decompressing incoming media streams, converting incoming control information into internal API calls and
Decompress the media into an internally uncompressed form. For a non-limiting example, inbound H.264 video can be converted to YUV420P frames before being passed to another process. The purpose of this is that if this process crashes, other parts of the system will not be affected.
Independent process: Each connection should be handled in its own process. A given process should only be responsible for one videophone terminal, so if this process crashes, only a single terminal will be affected, and the rest of the system will not notice anything.
Unprivileged processes: Each process should be as independent as possible from the rest of the system. In order to do this, ideally each process is best run with its own user certificate, and can use the change root directory system call to make most of the file system inaccessible.
Performance considerations: The protocol connector 212 will introduce some processes. Usually there is only one of these processes, which will bring the possibility of performance degradation, especially in the system that handles audio and video streams. It needs to move between processes. massive data. To achieve this, you can use shared memory facilities to reduce the amount of data that needs to be copied. [0062] In some embodiments, the global infrastructure engine 104 supports distributed fault-tolerant messaging based on an Internet/cloud client-server architecture, where the distributed fault-tolerant messaging provides one or more of the following features: Under the reliable and unreliable transmission mechanism, it has the ability to directly unicast, broadcast, multicast and anycast traffic.
The ability to balance the load between service requests to access media processing nodes and sensitive content or idle server classification.
A synchronous and asynchronous transmission mechanism that can transmit information regardless of whether the program crashes or not.
Including the use of effective fan-out technology for atomic broadcasting based on priority and time sequence transmission mechanism.
Use single write and atomic broadcasting to achieve an effective fan-out capability.
Selectively abandon non-real-time queued information to improve the ability of real-time response.
A priority-based queuing mechanism with the ability to discard non-transmitted non-real-time events.
A transaction-aware messaging system.
Integrate with a hierarchical entry naming system based on the content of the conference room, IP address, process name and process number.
[0063] Traditionally, the traditional video terminal of the video conference participant, such as the video conference terminal using the H.323 protocol, is mainly connected with other terminals in the local area network of the enterprise or organization. People have tried many times to make H. 323 terminals pass through firewalls to communicate seamlessly with terminals outside the corporate network. Some of these firewalls have been extended to H. 323 by standardized firewalls in the ITU protocol, called H. 460. 17 , 1& 19, 23, 24, and other attempts adopted by video conferencing equipment suppliers, including deploying gateway hardware or software in the DMZ of the companys network. However, the facts have proved that none of these attempts have been successful. Inter-organizational calls have always been cumbersome and can only be carried out through heavy IT support and participation.
[0064] In some embodiments, the global infrastructure engine 104 enables traditional video conference terminals of video conference participants to communicate with other terminals through seamless firewall traversal. Since traditional video conferences usually implement standardized protocols, and these standardized protocols do not assume available Internet/cloud-based services, the global infrastructure engine 104 uses at least one or more of the following technologies to implement the embodiment shown in FIG. 10 The shown firewall traversal: Restrict all video conferences from terminals located outside the firewall to a server of the global infrastructure setting engine 104, where the server can pass through a public IP on the network accessible to every user Address is obtained.
Maintain a series of UDP/IP ports used to reach the global infrastructure engine 104, and the global infrastructure engine
104 distributes media to a small designated subset of ports through the series of UDP/IP ports. This allows companies with strict firewall policies to only open the firewall in a relatively narrow range compared to fully opening the firewall.
Provides a simple web browser-based application that allows any user to easily run a series of checks to determine the performance and behavior of the company's firewall, and to determine whether the firewall is a problem when using H. 323 terminals Or need to change any rules.
Provide an enhanced browser-based application program that acts as a tunnel proxy server to enable any user to run software on the browser or local PC operating system to allow the terminal to tunnel the software to one or more A public server on the Internet. Optionally, the software can be run on any PC or a server on a local or virtual network in an independent manner to form a proxy server to achieve the same tunneling.
[0065] In the embodiment shown in FIG. 1. The user experience engine 106 presents multimedia content including but not limited to composite audio/video streams to participants of the VMR to enhance the participant user experience (UE). The user experience (UE) provided by the user experience engine 106 to the participants of the VMR conference hosted on the VMR engine 102 mainly includes one or more of the following aspects: Physical interaction with the video conference terminal. The user experience engine 106 can control the establishment and management of a multi-party communication video conference in a VMR in a device/manufacturer-independent manner. Most of the physical interaction with the remote control provided by the manufacturer can be classified as a web application, where the web application can be launched from any computer or communication device, including laptops, smart phones, and tablets. In some embodiments, given that Internet/cloud-based software can recognize and convert these interactions into actionable events, these interactions can also be driven by voice or visual commands.
The user interface (U1) related to the web application controls the interaction between the participants and the VMR engine 102. Here, the user experience engine 106 controls the interaction between the host and the conference participants. Through an intuitive user interface (U1) provided by the user experience engine 106, video conference participants can control features such as video layout, mute participants, send chat messages, share screens, and add third-party video content.
Video/multimedia content: User experience engine control, the content that is actually seen on the screen during the video conference and when participants log in to a VMR, and is presented in the form of screen output, synthesized conference sources, welcome banners, etc. In some embodiments, the user interface U1 and/or multimedia content may include information related to the performance indicators of the participant's call experience, including but not limited to video resolution, video and audio bit rate, connection quality, The connection packet loss rate, the carbon compensation obtained as a result of the call, the savings in transportation costs, and the cost savings compared to traditional microprocessor-based MCU calls. This is also an environmentally friendly solution, such as saving money spent on infrequent flight journeys or the same mileage, and the right to show off associated with various status levels similar to the flight status level. Motivation programs can be formulated according to the different levels of status obtained, and participants are encouraged to use video conferences as opposed to travel conferences. This gives them personal rewards for using video conferencing for business interests.
Customize video conferences with special applications (such as vertical industries). In order to make the video conference meet the needs of special industries, the user experience engine 106 allows users to customize the VMR, which enables conference participants to experience a new level of collaboration and conference effectiveness. These vertical industries or professions include, but are not limited to, employment and recruitment, distance education, telemedicine, security legal confession, shared viewing of instant events such as sports, concerts, and customer support, etc.
Personalize the VMR according to the preferences and privileges of each host and/or participant. When scheduling a video conference, the user experience engine 106 provides the moderator with the ability to personalize the conference. Examples of these customizations include, but are not limited to, the initial welcome banner, agenda upload, specifying the video layout to be used in the meeting, and giving privileges to meeting participants.
[0066] Although most traditional video conferencing systems cost hundreds of thousands of dollars, they are
The organizer may provide participants with very limited freedom and flexibility in controlling the user experience. The layout comes from several pre-configured options, and the settings that can be modified during a call are also limited.
[0067] In some embodiments, during the call of the participants, the user experience engine 106 provides management and control of security and privacy settings in the conference initiated by the host, wherein the management and control features include but are not limited to , Mute a specific speaker in a video conference, control and/or broadcast the layout related to a video conference terminal to all or some participants, and selectively share additional materials with some participants (for non-restrictive implementation For example, in a human resources vertical application, where multiple interviewers interview candidates in a common call).
[0068] Providing video conferencing services through the Internet/cloud, the user experience engine 106 reduces many limitations of traditional video conferencing systems. For a non-limiting example, in a video conference, the user experience engine 106 enables participants associated with different types of video conference terminals to talk to each other over the Internet. For a non-limiting example, a participant from an H·323 terminal and a participant from a desktop client such as Skype talk to each other. The host and the participant can choose from many options. In addition, compared with traditional passive bridge conferences, by providing termination services in the cloud, the user experience engine 106 can access richer functions of the video conference that participants can use. More specifically, each participant can control one or more of the following:
1. The video window of which active participant of the VMR conference is displayed on the screen of his/her conference terminal.
2. Options for how different participants should be arranged on the screen of his/her conference terminal.
3. Where and how to view the layout options of the second video channel (screen sharing, presentation sharing, co-viewing other content) on the screen of his/her conference terminal.
Using this kind of in-conference control, the host can control the security and privacy settings of a specific call in a way that is not allowed or provided by the prior art.
[0069] As shown in FIG. 11, in addition to the above options, the host of the call has a wealth of options to manage and control the video conference selected through a web interface, where these options include but are not limited to:
1. Mute some participants during the call.
2. Share content with some participants during the call.
3. Define a standard layout of his/her video conference terminal screen and some display caller that other participants can see.
4. Choose to display metadata of the caller description in the respective video windows of some participants, including user name, website name, and any other metadata.
5. Add and delete participants in a video conference call in an easy and seamless way through a real-time, dynamic web interface.
6. Easily customize the welcome screen displayed to video callers who join the call, which can display information about the call and any audio or video materials that the service provider or call host wants participants to see.
[0070] In some embodiments, the user experience engine 106 can make a private meeting in the VMR possible by creating sub-rooms in the main VMR, and any subset of participants can join the private meeting and have a private chat. For a non-limiting example, participants can invite others to have a quick audio/video or text conversation while remaining in the main VMR.
[0071] An experience sharing activity between participants of a video conference usually requires all participants in the same place to be physically present. Otherwise, when it happens on the Internet, the quality is often very poor.
The necessary steps are quite challenging for an average person to turn it into a practical technology.
[0072] In some embodiments, the user experience engine 106 provides co-browsing events through VMR, these events can be scheduled and shared between participants, so that they can experience the fun of multiple people participating in one thing at the same time, and Share the experience together via video conference. For a non-limiting embodiment, the shared event may be a Super Bowl game that people want to enjoy with friends, or a quick conversation where a group of friends watch some movie trailers together, so as to decide which one to watch in a movie theater.
[0073] In some embodiments, as described in the embodiment in FIG. 12, the user experience engine 106 utilizes the MCU in the global infrastructure engine 104 to provide a convenient, fast and high-quality solution for event sharing. More specifically, the user experience engine 106 makes the call initiator 1202 invite a group of other participants 1204 to share the video conference in the VMR through the web interface 1206. Once added to the web interface 1206 in VMR, everyone in the web interface 1206 shares online videos and content, Then the initiating participant 1202 can present a link to the website where the content to be shared is located, and the content flows directly from the content source into the same VMR1206, regardless of whether the content is with the initiating participant 1202 or a third party located on the network On a website or content server. Participant 1202 can continue to have conversations with other participants 1204 while watching content 1210. The characteristics of the content to watch include, but are not limited to, for example, where to see it, its audio level, whether it needs to be muted, whether to pause or temporarily remove it, etc. The layout of various items, which are controlled by 1210 people who share these contents, are similar to the control of the video conference host who manages and controls the video conference discussed above. This method provides an eye-catching and novel way for a group of people distributed around the world who want to experience an event together to watch the live broadcast. This makes possible a whole new set of applications around active participation in live events, such as social events such as meetings or weddings.
[0074] In some embodiments, the user experience engine 106 can implement multi-view display and device-independent control of participants in the video conference. Here, each video terminal has its own user interface and is available to the hardware video system in the conference room, and each video conference terminal has a remote control that is not easy to use. In order to make the user experience of connecting to the VMR simple, the user experience engine 106 minimizes the operations that need to be implemented using the local interface, and moves all these functions to the interface that runs in the devices that most users are familiar with. The devices that these users are familiar with include For example, desktop computers, notebook computers, mobile phones or mobile tablet computers, through which the user experience is as independent as possible from the terminal device user interface functions to control the VMR. Through this device-independent video conference control, the user experience engine 106 provides flexibility, ease of use, rich experience and extended functions, which makes this experience far more personal and interesting for the participants. .
[0075] In some embodiments, the user experience engine 106 may also allow participants to use multiple devices/video conference terminals to participate in and/or control the video conference. On a device, such as a video conference room system, participants can receive audio and video streams. On another device, such as a laptop or tablet, participants can send/receive presentation materials, chat messages, etc., and can also use it to control the meeting, such as mute one or more participants, and use it for presentation Picture-in-picture to change the layout of the video conference terminal screen, etc. The operations on the laptop are reflected to the video conference room system, because both are connected to the same hosted video conference VMR.
[0076] At present, connecting a video conference from an H.232 terminal often requires frequent steps to be performed through remote control of the device. In addition to logic issues, such as locating the remote control in the requested conference room, there are also learning curve related issues such as obtaining the correct number from the directory to call and entering the designated code for remote calling. When terminal participants and desktop participants turn on their video devices to join the conference, they are directly added to the conference.
[0077] In some embodiments, the user experience engine 106 provides a brand new
In order to improve and simplify the user experience, the welcome screen content presented to the participants includes, but is not limited to, the interactive welcome handshake used in the video conference in the embodiment of FIG. 14, the splash screen, the interaction of the information related to the room entry number, and the welcome Video etc. In order to join a call from a video conference terminal, all the host has to do is to call the personal VMR number he/she has subscribed to. Then the host can set the detailed information of the call, the information includes rich media content forming the part of the welcome handshake with other participants, which can be set as the default option for all calls hosted by the host. Other participants call into the VMR and enter the room with the designated number to make a conference call. When joining VMR, they first enjoy the rich media content set as the welcome page, including meeting content descriptions, such as agenda, names of callers, company-related statistics, etc. These contents can also be more general non-commercial applications, including any flash contents such as videos, music, animations, and advertisements. After joining the call, the display also displays a participant-specific code on his/her screen of the participant, which can be used to deliver content to the conference for shared content. The code can also be input through a web application for calling or run through voice or visual instructions, where the voice or visual instructions are recognized and processed by software in the Internet cloud, and then converted into executable events.
[0078] FIG. 13 depicts an example of a connection between a conference room system and a laptop or a mobile phone, where the participants use a home network (HAN) room conference system 1002 and use the entry directory on the remote control to send to a well-known VMR1004 After dialing out, once the connection is successful, the user experience engine 106 plays a welcome screen and a "conference ID" related to the video conference. The participant enters a web application 1006 or a mobile application 1008, and enters the conference ID together with the conference number of the VMR that the participant wishes to join into the application, so that the participant enters the VMR. Optionally, a participant can enter a video conference hosted in VMR according to one of the following methods: Use buttons in the conference room system to fail to mention One-time dialing access via voice recognition From a notebook to the conference room The system plays the recognized music or sound Once connected, the camera in the meeting room shows some gestures or patterns.
[0079] The experience described above also provides an opportunity for any participant to open the audio or video stream without using the default. When all participants enter and the call is ready to start, the moderator can start the call globally, and each participant can finely control whether to turn on/off their audio/video. In some embodiments, this also allows monetizable services to be provided while participants are waiting, such as streaming media advertisements to a small range of participants, time zone, population, and other characteristics determined by Internet cloud services. In other embodiments, when people wait for the call to start, a video about the new features introduced in the server can be displayed. In other embodiments, the detailed information of the participants in the conference call can be displayed in a rich multimedia format. There can be no such technology in any technology.
[0080] At present, consumers who wish to organize a video call have only two choices, either choose to use the commercial/professional option of the H323 terminal, such as Polycom or TED systems, or use a desktop application with limited functions/quality. This kind of application shows the user a video of the participants stamp size, and in particular, is displayed on a simple or boring background or interface.
[0081] In order to solve this problem, FIG. 15 depicts that the user experience engine 106 provides a personalized VMR that allows participants to customize (or customize) or personalize his/her meeting experience on a per-call basis. These significant changes, Revolutionized and popularized the video conferencing experience. For business users, the user experience engine 106 finds conference rooms for calls during the meeting, or similar professional settings and different types of backgrounds, welcome music, welcome banners, other status and chat information, and colorful paper tapes, etc. to provide layout and background .
[0082] The host can choose one from a series of options provided to him/her based on a predetermined plan. For retail consumers, the experience will be more informal and varied. The caller can decorate his/her room in any way he/she likes. Participants' personal websites for VMR can be decorated and personalized in the same way. During the call, the user experience engine 106 can extract or read in these customizable options specified by the participants, and put them into this customized VMR, so that this experience is richer than the traditional call experience.
[0083] Providing personalized conference services via the Internet/cloud has a clear advantage of eliminating processor computing power in any terminal. As long as the terminal can receive and process the encoded video stream, the user experience engine 106 can provide participants with any level of rich media content as part of their call experience, all of which can be controlled and set by the host in the VMR.
[0084] For a two-person dialogue, unlike the traditional flat layout where the two parties are displayed side by side, the user experience engine 106 can present a 3D layout, and the input videos from the two participants are set to make them appear to be looking at each other. In this way, other participants in the video conference see the conversation more naturally. Similarly, for a non-traditional application, such as telemedicine, or a conference call where a patient can talk to a doctor remotely, the conference itself can be made like a doctor's office. Patients can watch videos of health-related issues while waiting for the doctor. Once the doctor calls in, these experiences can simulate a virtual doctor's office visit. Other applications include but are not limited to some scenarios. For example, recruitment may have a custom layout and a video call that looks different, where in their video, the interviewees resume can be seen by the interviewer, and the resume Can be edited and annotated by the interviewer, but this may be hidden from the interviewee.
[0085] A "tandem" service, such as our service to keep callers anonymous, which allows them to call in from any software or hardware terminal without being discovered by the receiver of any personally identifiable information of the caller.
[0086] At present, one of the most important pain points for web users is the lack of convenient and complete solutions for remote collaboration. There are many scenarios in which users need to share with remote users what is currently on their screen, such as a drawing, a video, the current state of their device during online fault detection, and some specified content. Currently, the only way to do this is to sign a desktop client that supports screen sharing and request permission to start sharing. If one party does not have such a client, or the person he wants to share the screen with is a contact on the client's client, this method is invalid. In addition, these solutions are not applicable to mobile phones and other small screen devices.
[0087] In some embodiments, the user experience engine 106 creates a single online "homepage" for personalized sharing of someone's desktop, laptop, and/or mobile screen with other video conferencing terminals. In the discussion of this article, screen sharing refers to the action of displaying a persons screen on the screen of the other end/remote video conferencing terminal, so that the screen of a remote machine can be seen, or in a streaming media at the same time. In the form. Some minor changes to this sharing include allowing remote users to see only a part of someones screen, giving them extra ability to interact with someones screen, and so on. For a non-restrictive embodiment, for screen sharing, the user experience engine 106 provides one or more of the following features: Ability to access HTTP or HTTPS in a personalized and consistent manner. As shown in FIG. 16, for a non-restrictive In an embodiment, a user of a service will be assigned a website address (URL) in the form of http://myscre. en/jo eb low, which will be used whenever the user shares one or more of his/her screens. A permanent access link to access the users screen. Then the user can share this URL with his/her friends, colleagues, online social networks, etc. Here, the URL can be a so-called TinyURL (domain name is usually less than 10 characters), which can make a simple shorthand for the location on the network.
Access to someones screen sharing URL can be customized by default options, so that access can only be achieved when the user actively chooses to share his/her screen. In addition, users are provided with the option of combining participant pass codes, timed screen sharing sessions, and IP address filtering options to ensure maximum control over the group of users who share his/her screen.
When in screen sharing mode, a list of available screens can be displayed to participants, and participants can select one or more of these lists to watch. According to the host's permission settings, they can also issue remote access to interact with the shared screen.
[0088] For example, the Skype company has created a browser plug-in, which allows a participant to click to display a "Skype phone" icon next to any number displayed on his/her browser. She makes calls to any number on her browser and sends these calls through a Skype desktop client. On the other hand, the online contacts of today's users may be one of multiple storage methods, such as Google contacts, exchanges, Yahoo contacts, Skype, Facebook, etc. Although these contacts can interact in different ways in the native application (for non-limiting embodiments, for example, hovering over a Google contact, the user will be presented with a menu with options to email or instant call with the contact ), but there is no simple and universal way to provide a one-click video call function that spans different contact protocols, similar to Skype's one-click video call to a number.
[0089] In some embodiments, the user experience engine 106 supports a web browser and/or desktop plug-in, which enables intelligent one-click video conference calls to participants in VMR contacts from the authentication protocol of the video conference terminal. (Not the number). As described in this article, a plug-in refers to a small piece of software that extends the functionality of a larger program. The plug-in is often used in web browsers and desktop applications to extend their functionality in a specific area.
[0090] In some embodiments, the plug-in created by the user experience engine 106 provides such a function, no matter what the situation, the contacts from the verification agreement are displayed in the browser (such as Gma subscription, YI Ma subscription, etc.) and / Or desktop applications (such as MS Outlook, Thunderbird, etc.). As shown in Figure 17, the one-click video conference call plug-in provided by the user experience engine 106 provides at least the following features:
1. The user must agree to install these plug-ins and activate the application of said plug-ins.
2. In order to enable the application, there is a "video call" icon next to each contact from the authentication protocol (assuming it is a message collaboration and Google contact). For a non-limiting example, if the sender of the mail in a user's Exchange mailbox is an authentication protocol, the display interface of the mailbox is added with a video call icon and a small arrow showing more options.
3. Click this icon to establish a video conference call between the user and the contact through VMR, where both parties select the appropriate video terminal.
4. Clicking on the arrow provides the user with a complete list of the ways in which the user can interact with the contact through the VMR service. The VMR service includes audio calls, scheduling a future call, and so on.
[0091] In some embodiments, when some video conference call rooms are too bright and cause distraction, the user experience engine 106 automatically performs video gain control. Similar to the AGC (Auto Gain) in the audio system, the brightness of all rooms of the video conference terminal is part of the video conference, which can be adjusted to produce the illusion that the conference is taking place in the same place. Optionally, the video automatic gain control can be turned on by participants in the conference who feel that the brightness of one or more rooms makes them disturbed.
[0092] In some embodiments, the user experience engine 106 provides real-time information about the cost savings of the ongoing video conference, such as the miles saved for each conference, gas costs, and hotel costs. Among them, the distance between participants can be based on
The geographic location and mileage of the IP address are calculated, and the federal mile points can be calculated based on the miles saved per call to present a total amount on the screen. The claimed carbon offset can also be calculated based on the location of the participant and the duration of the call, and appropriately displayed to the participant.
[0093] Virtual reality (VR) represents the spectral span between a computer avatar and a static image, where the computer avatar is pre-configured for selection, and the static image is uploadable and animated to a limited extent. ). In a multiparty video call setting, there is no way to either migrate the participants to a virtual world while maintaining their real-world characters, or to transplant them into a virtual reality (VR) world and animate them at the same time Persona.
[0094] In some embodiments, in order to solve the two problems discussed above, the user experience engine 106 presents realistic virtual reality (VR) events to participants of a video conference through the MCU of the global infrastructure engine 104. Like a traditional video conference call, the VMR accepts the input of audio/video streams from the cameras of different participants and composites them into a whole, and then encodes and transmits the composite video to each participant separately. As shown in FIG. 18, when a participant wants a virtual reality (VR) version of an event, the user experience engine 106 accepts one or more of the following additional steps to deliver the VR experience to the participant.
1. The image detection and segmentation component 1802 accepts video input from each participant.
2. The segmentation component 1802 detects and extracts participants from the background of the video stream, and provides metadata about his/her position and other characteristics in the video stream.
3. Then the user experience engine 106 adds various features to the participant's face through the virtual reality rendering component 1804 or transforms the face by applying any image transformation algorithm, so that the participant becomes vivid. The user experience engine 106 can perform further analysis of the face and feature detection and fully make the participant's face vivid, and then create a semi-animated version of the face itself.
4. The video synthesizer 302 in the media processing node 300 replaces the selected background with the background covered by the (lively) participants of the virtual real scene, and provides the video stream to other participants.
In this way, the user experience engine 106 can obtain and convert video and audio streams input from different participants in different customized ways to achieve different user experiences.
[0095] In some embodiments, the user experience engine 106 may extract all participants from their video stream environment, and then add them to a common environment and send them as a video stream. For a non-limiting embodiment, calls made by different participants located in different geographical locations may all look like a conversation sitting at the conference table with each other.
[0096] Providing small-scale, real-time advertisements about services available to users in a specific geographic area has broad market application prospects and huge benefits. The few existing solutions rely heavily or completely on GPS-related information, or high-performance processors on mobile devices to execute the required programs to generate the information. In some embodiments, the user experience engine 106 allows the MCUs of the global infrastructure engine 104 to implement Internet/cloud computing-based augmented reality user interaction services. Further, the user experience engine 106 parses the video stream collected by the participant/user video conference terminal (for example, a mobile phone with a camera), and provides augmented reality video together with the services available in the geographical area of the participant. The comments are fed back to the user together, such as local events, entertainment, and dining options. The user only needs to place a video phone on the VMR and point his/her video camera at the place of interest to him or her. As shown in FIG. 19, the user experience engine 106 and the global infrastructure engine 104 process the received video in the cloud, so that the user equipment does not need to have processor capabilities. For a non-limiting embodiment, image detection and analysis are performed
The cutting component 1902 analyzes the billboards and the road signs that can be confirmed to confirm and verify the GPS information obtained from the location service database 1904 or the information obtained from the user's GPS to determine his/her whereabouts. The user experience engine 106 then changes the input video source with the collected user geographic information and overwrites the video stream through the metadata synthesizer 1906 to generate an augmented reality video stream to the user, where the metadata synthesizer has metadata including, for example, Names of restaurants within walking distance, names of local entertainment venues, etc.
[0097] As shown in FIG. 19, the image detection and segmentation component 1902 is the core of the logic, which is used to analyze the input video and extract regions of interest from the video. The location service database 1904 is filled with information about different compression codes, which can take GPS information and/or compression codes as input, and provide rich data about services in the area and other information that may be of interest. As mentioned earlier, the metadata synthesizer 1906 obtains input information from the image detection and segmentation component 1902 and the location service database 1904, and presents useful metadata in real time. The input video source is covered with useful metadata about the surrounding environment. .
[0100] In some embodiments, when the user is walking around the area, the user experience engine 106 may provide a guided tour of the area, and the user experience engine 106 may fill the screen with scenery and information about the area in advance. More information on the sound. In other embodiments, the user experience engine 106 can also fill in information in the video and display photos of friends who may be nearby, so as to embed this augmented reality service into an existing service to locate friends in the same area. .
[0101] In some embodiments, the augmented reality service provided by the user experience engine 106 is customizable, not only when installing any downloaded software on the user's mobile terminal, but also in every way of use. The user may use the augmented reality service for various purposes. For a non-limiting example, a 411 query is made at a certain moment. After that, the user can immediately make a call and get a virtual virtual tour of the local tour. During the journey, soon, the user may still need relevant information about the restaurant to have a big meal. As more information about each place can be obtained on third-party websites, the user experience engine 106 provides a seamless way to closely associate with each provider on the Internet/cloud to provide more information to each user. More current information. . Since this method is completely stackable and only depends on the plan selected by the user, the call can be run through a system with more powerful operating capabilities to obtain and provide users with more useful information, so that each user The required features provide a complete set of pricing options.
[0102] In some embodiments, the user experience engine 106 supports the free translation of real-time multimedia communications in a live broadcast video conference by real-time translation of different languages in the cloud, so that VMR participants can use different languages to translate Communicate with each other and conduct intelligent conversations. Furthermore, the real-time cloud-based translation may include, but is not limited to, one or more of the following options: In a common language, real-time voice and subtitles, for example, different speeches in a video conference Participants can speak in different languages, while translation and letters are seamlessly completed in the cloud; Voice-initiated services, such as search and location-based services, provide translation from language to visualization; With each participant The voice translated into the corresponding language selected for him/her; The same language as the speaker, but a different voice can be selected to replace the speakers voice; Multimedia speech is provided to different users in different languages through cloud technology / Simultaneous translation of conferences; Not only audio/video transmission is used, but also real-time transmission of files/inputs from the format of the sender to the format of the data selected by the receiver, where this conversion is realized in the cloud.
Taking into account the delay of construction, when two or more parties communicate in different languages, this function can be replaced by manual translation.
[0103] Another embodiment can be implemented by using a general or dedicated digital computer or microprocessor according to the guidance of the disclosure of the present invention, which is obvious to those skilled in the computer field. Under the teaching of the content disclosed in this application, a skilled programmer can easily write a suitable software code, which is also obvious to those skilled in the software field. Moreover, it is also obvious to those skilled in the art that the present invention forms a suitable network by using integrated circuits or interconnected by suitable conventional components and circuits.
[0104] An embodiment includes a computer program product. The computer program product is a machine-readable medium on which instructions are stored, which can instruct one or more hosts to perform any function proposed herein. The machine-readable storage medium may include, but is not limited to, one or more types of magnetic disks including floppy disks, optical disks, DVDs, CD-ROMs, DRAMs, VRAMs, flash memory, magnetic or optical cards, nanosystems (including molecular memory Ics), or any other type of media and equipment suitable for storing instructions and/or data. Stored on any computer readable medium, the present invention includes software. This software is used to control the hardware of a general/special computer or microprocessor, as well as the hardware that allows the computer or microprocessor to interact with the user or execute the present invention. Other mechanisms, such software may include, but are not limited to, device drivers, operating systems, operating environment platforms/containers, and applications.
[0105] The description of the various embodiments of the claimed subject matter above has been provided for purposes of illustration and description. Its purpose is not to exhaust or limit the exact form disclosed for the claimed subject matter. Many modifications and changes are obvious to those skilled in the art. The interface described in the embodiment and method of the system, these contents can obviously be replaced by equivalent software concepts, such as classes, methods, types, models, components, beans, modules, object models, Programs, threads and other appropriate content. For the components described in the above system embodiments and methods, it is obvious that these concepts can be replaced with equivalent concepts, such as classes, methods, types, interfaces, modules, object models, and other suitable content. The selected and described embodiments are used to best describe the principles and practical applications of the present invention, so that those skilled in the relevant fields can understand the subject matter claimed by the present invention. Therefore, various specific embodiments and specific applications are provided. Various modifications that can be expected are suitable.
18 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN108513088A | Cited by | China | – | Search report | – |
| CN105635632A | Cited by | China | – | Search report | – |
| CN109756694A | Cited by | China | – | Search report | – |
| CN115516836A | Cited by | China | – | Search report | – |
| US10609334B2 | Cited by | United States of America | – | Applicant | – |
| US9900552B2 | Cited by | United States of America | – | Applicant | – |
| CN105282477A | Cited by | China | – | Search report | – |
| WO2018153267A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| CN105830031A | Cited by | China | – | Search report | – |
| CN114125369A | Cited by | China | – | Search report | – |
| WO2016008457A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| CN107371069A | Cited by | China | – | Search report | – |
| CN105094957A | Cited by | China | – | Search report | – |
| US10261834B2 | Cited by | United States of America | – | Applicant | – |
| WO2015172437A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| CN108513089A | Cited by | China | – | Search report | – |
| CN112102501A | Cited by | China | – | Search report | – |
| CN104754282A | Cited by | China | – | Search report | – |
| CN115203655A | Cited by | China | – | Search report | – |
| CN101047531A | Cites | China | A | Search report | 1-24 |
| CN101335869A | Cites | China | A | Search report | 1-24 |
| CN101517963A | Cites | China | A | Search report | 1-24 |
| CN101635635A | Cites | China | A | Search report | 1-24 |
| CN1372416A | Cites | China | A | Search report | 1-24 |
| CN1820505A | Cites | China | A | Search report | 1-24 |
| EP1830568A2 | Cites | European Patent Office (EPO) | Y | Search report | 1-26 |
| US7035230B1 | Cites | United States of America | A | Search report | 1-24 |
| WO9424803A1 | Cites | World Intellectual Property Organization (WIPO) | Y | Search report | 1-26 |
50 members in 4 offices
Priority claims29
| Document | Office | Kind | Date |
|---|---|---|---|
| 33404310 | United States of America | P | |
| 33404310 | United States of America | P | |
| 33404510 | United States of America | P | |
| 33404510 | United States of America | P | |
| 33405010 | United States of America | P | |
| 33405010 | United States of America | P | |
| 33405410 | United States of America | P | |
| 33405410 | United States of America | P | |
| 61334043 | United States of America | – | |
| 61334045 | United States of America | – | |
| 61334050 | United States of America | – | |
| 61334054 | United States of America | – | |
| 13105699 | United States of America | – | |
| 201113105699 | United States of America | A | |
| 201113105699 | United States of America | A | |
| 2011036263 | United States of America | W | |
| 2011036263 | United States of America | W | |
| 13105699 | – | – | – |
| 61334043 | – | – | – |
| 61334045 | – | – | – |
| 61334050 | – | – | – |
| 61334054 | – | – | – |
| PCTUS2011036263 | – | – | – |
| US20100334043P | – | – | – |
| US20100334045P | – | – | – |
| US20100334050P | – | – | – |
| US20100334054P | – | – | – |
| US201113105699 | – | – | – |
| WO2011US36263 | – | – | – |
Members50
| Document | Office | Kind | |
|---|---|---|---|
| US2011279634A1 | United States of America | A1 | |
| US2011279635A1 | United States of America | A1 | |
| US2011279636A1 | United States of America | A1 | |
| US2011279637A1 | United States of America | A1 | |
| US2011279638A1 | United States of America | A1 | |
| US2011279639A1 | United States of America | A1 | |
| US2011283203A1 | United States of America | A1 | |
| WO2011143427A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2011143434A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2011143438A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2011143440A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2012082226A1 | United States of America | A1 | |
| WO2012047849A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2012047849A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2569937A1 | European Patent Office (EPO) | A1 | |
| EP2569938A1 | European Patent Office (EPO) | A1 | |
| EP2569939A1 | European Patent Office (EPO) | A1 | |
| EP2569940A1 | European Patent Office (EPO) | A1 | |
| US8482593B2 | United States of America | B2 | |
| CN103238317AThis record | China | A | |
| CN103250410A | China | A | |
| EP2625856A1 | European Patent Office (EPO) | A1 | |
| US8514263B2 | United States of America | B2 | |
| CN103262529A | China | A | |
| CN103270750A | China | A | |
| CN103493479A | China | A | |
| US2014092203A1 | United States of America | A1 | |
| US2014098180A1 | United States of America | A1 | |
| US2014313278A1 | United States of America | A1 | |
| US2014313282A1 | United States of America | A1 | |
| US8875031B2 | United States of America | B2 | |
| US8885013B2 | United States of America | B2 | |
| US9035997B2 | United States of America | B2 | |
| US9041765B2 | United States of America | B2 | |
| US9124757B2 | United States of America | B2 | |
| US9143729B2 | United States of America | B2 | |
| US9232191B2 | United States of America | B2 | |
| US9300705B2 | United States of America | B2 | |
| US9369673B2 | United States of America | B2 | |
| EP2569939B1 | European Patent Office (EPO) | B1 | |
| CN103238317B | China | B | |
| EP2569937B1 | European Patent Office (EPO) | B1 | |
| CN103493479B | China | B | |
| EP2569938B1 | European Patent Office (EPO) | B1 | |
| CN103262529B | China | B | |
| CN103270750B | China | B | |
| CN103250410B | China | B | |
| EP2625856B1 | European Patent Office (EPO) | B1 | |
| EP2569940B1 | European Patent Office (EPO) | B1 | |
| EP2569940B8 | European Patent Office (EPO) | B8 |
4 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Transfer of patent rightTR01 | TR01 | |
| Patent grantGrantedGR01 | GR01 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 103238317
- Publication, DOCDB
- 103238317
- Publication, EPODOC
- CN103238317
- Application
- 800344437
- Application, DOCDB
- 201180034443
- Application, EPODOC
- CN201180034443
Titles2
- Chinese
- 实时多媒体通讯中可伸缩分布式全球基础设施的系统和方法
- English
- System and method for scalable distributed global infrastructure in real-time multimedia communication
Classification
- CPC, 5
- H04N7/152
- H04N7/141
- H04L12/1827
- H04L51/10
- H04N5/265
- IPC, 2
- H04N7 14
- H04N7 15