Hosting and broadcasting virtual events using streaming interactive video
Abstract
The present invention is a method comprising broadcasting a live game tournament in the form of a multicast streaming interactive video stream to a large number of viewers over the Internet in a hosting service. Audio from the announcer is interactively overlaid onto the video stream multicast by the hosting service.

Term
Projected expiry 4 December 2028.
- Priority
- Filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1호스팅 서비스에서 인터넷을 통해 다수의 관찰자에게 멀티캐스트 실시간 스트리밍 인터랙티브 비디오 스트림의 형태로 라이브 게임 토너먼트를 브로드캐스팅하는 단계;및 호스팅 서비스에 의해서, 아나운서로부터의 오디오를 멀티캐스팅된 비디오 스트림 상으로 인터랙티브하게 오버레이하는 단계;를 갖추어 이루어진 것을 특징으로 하는 방법.
331 paragraphs in 1 section, as filed
HOSTING AND BROADCASTING VIRTUAL EVENTS USING STREAMING INTERACTIVE VIDEO
This application is a continuation-in-part (CIP) of 10/315,460, filed December 10, 2002, entitled "Method and Apparatus for Wireless Video Gaming," by the assignee of this CIP application. was pumped
SUMMARY OF THE INVENTION The present invention generally relates to the field of data processing systems that enhance user performance for accessing and manipulating audio and video media.
Recorded audio and video media have been an aspect of society since the days of Thomas Edison. At the beginning of the 20th century, recorded audio media (cylinders and records) and moving media (nickelodeons and movies) became widespread, but both technologies were still in their infancy. In the late 1920s on the mass market, film combined with audio, followed by color video with audio. Radio broadcasting has progressively developed into a form of advertising support largely in the broadcast mass market audio media. When television broadcasting standards were established in the mid-1940s, television combined with radio as a form of broadcast mass-market media that brought already recorded or live movies home.
By the mid-20th century, the majority of American homes had phonograph record players to reproduce recorded audio media, radios to receive live audio, and live broadcast audio/video (A/V) media to play back. I had a television set. Very often these three "media players" (record players, radio and television) are combined into one cabinet that shares a common speaker that becomes the "media center" in the home. Although media choices were limited to consumers, the media "ecosystem" was fairly stable. Most consumers knew how to use the "media player" and were able to enjoy the full range of their abilities. At the same time, media publishers (mainly video and television studios and record companies) can sell their media without going through "second sales" or widespread piracy, such as the resale of used media. It could be distributed to theaters and homes. A typical publisher does not receive any revenue from the second sale, and that alone reduces the revenue that would be gained from buyers of used media for new sales. Although second-hand records may have been sold in the mid-20th century, those sales have had little impact on record publishers. Because unlike a movie or video program, which is viewed only once or in small numbers by adults, a music track will be listened to hundreds or thousands of times. So, music media is much less "perishable" (eg has lasting value to adult consumers) than movie/video media. Once a record is purchased, if the consumer likes the music, the consumer will be able to keep it for a long time.
From the mid-20th century to the present, the media environment has undergone a series of radical changes in the interests and losses of consumers and publishers. With the widespread prevalence of audio records, especially cassette tapes with high-quality stereo sound, there is certainly a high degree of consumer convenience there. But it also marks the beginning of widespread practice (piracy) into consumer media. To be sure, many consumers used cassette tapes to tape their own records purely for convenience, but more and more consumers (e.g., students in dorms who have easy access to each other's record collections) will make infringing copies. . Also, consumers will tape-record music played over the radio rather than buying a record or tape from the publisher.
The advent of consumer VCRs has brought more consumer convenience, and since current VCRs are installed to record TV shows that can be watched late, it has also led to the creation of a video rental business, where TV programming as well as movies are " It is now accessible on an "on demand" basis. The rapid development of the mass market of home media devices since the mid-1980s has led to an unprecedented level of convenience and choice for consumers, and has led to the rapid expansion of the media publishing market.
Today, consumers are mostly tied to a particular type of media or a particular publisher, facing a plethora of media choices as well as a plethora of media devices. An avid consumer of media owns multiple devices connected to televisions and computers in various rooms of the home, resulting in a "rat's nest of personal computers (PCs) and/or one or more televisions and cables, as well as a group of wireless controls." )" results in (In the context of this application, "Personal Computer" or "PC" refers to any kind of computer suitable for home or office, which includes desktop, Macintosh<sup>&#174;</sup> or other non-Windows-based computers, Windows-based devices, Unix-like devices, laptops, etc.) These devices include video game consoles, VCRs, DVD players, audio surround-sound processors/ amplifiers (Audio sorround-sound processor/amplifier), satellite set-top boxes, cable TV set-top boxes, and the like. And for the avid consumer, there are a number of similarly functioning devices due to compatibility issues. For example, consumers can use HD-DVD and Blu-ray DVD players, or Microsoft XBOX<sup>&#174;</sup> , Sony Playstation<sup>&#174;</sup> You will own a video game system. Moreover, due to the incompatibility of game console versions between any games, consumers are<sup>&#174;</sup>You will own both later versions such as . Often, consumers are bewildered about which video input and which remote operation to use. Even after the disc is placed in the correct player (eg DVD, HD-DVD, Blu-ray, Xbox or Playstation), its video and audio inputs are selected for that device, and the correct remote control is found, consumers still face technical challenges. will be faced with For example, in the case of a wide screen DVD, the user needs to first decide and set the correct aspect ratio for the TV or monitor screen (eg 4:3, Full, Zoom, Wide Zoom, Cinema Wide, etc.) becomes Similarly, the user needs to decide first and then set it to the correct audio surround sound system format (eg AC-3, Dolby Digital, DTS, etc.). Each time, consumers are unaware that they are not enjoying the media content to the full performance of their television or audio system (eg, watching a movie distorted with the wrong aspect ratio, or audio in stereo rather than surround sound). listening to).
Gradually, Internet-based media devices are added to a stack of devices. Sonos<sup>&#174;</sup> Audio devices, such as Digital Music Systems, stream audio directly from the Internet. Similarly, the Slingbox<sup>TM</sup> Devices such as entertainment players can stream and record video over a home network or watch remotely from a PC over the Internet. And Internet Protocol Television (IPTV) services often provide cable TV-like services over a digital subscriber line (DSL) or other home Internet connection. Also Moxi<sup>&#174;</sup> There is a recent effort to consolidate multiple media functions into a single device, such as Media Center and PCs running Windows XP Media Center Edition. While each of these devices provides an element of convenience for the function it performs, each lacks ubiquitous and simple access to most media. Moreover, because of frequent expensive processing and/or the need for local storage, such devices cost hundreds of dollars to produce. Additionally, these modern consumer electronic devices typically consume a lot of power even when not in operation, which means waste of energy resources and expensive for time passing. For example, a device will continue to work if the consumer ignores it to turn it off or switch to another video input. And, as no device is a complete solution, it must be integrated with a stack of other devices within the home, leaving users with a sea of remote controls and cluttered wires.
Moreover, when many new Internet-based devices operate properly, they typically offer media in a generic form than are available in other forms. For example, devices that stream video over the Internet are simply video, games, or other non-interactive "extras" that often accompany DVDs, such as the "making of" a director's commentary. material) is streamed. This is often due to the fact that interactive materials are manufactured in a specific format intended for a specific device that handles interactivity locally. For example, DVD, HD-DVD, and Blu-ray discs each have their own specific interactive format. Any home media device or local computer that has been developed to support all popular formats requires a level of sophistication and flexibility that is complex and quite expensive for the consumer to operate.
In addition to that problem, if a new format is introduced in the future, the local device will not have the hardware capabilities to support the new format, which means consumers will have to purchase an upgraded local media device. For example, if higher-resolution video or stereoscopic video (eg, one video stream for each eye) has been introduced recently and the local device does not have the computational capabilities to decode the video, or in a new format (e.g. if stereophony is achieved via 120fps video synchronized with shuttered glasses delivered at 60fps to each eye, this option would not be available in the absence of upgraded hardware) You will not have the hardware to output video.
The issue of obsolescence and complexity of media devices is a serious problem when it becomes sophisticated interactive media, especially in video games.
Currently, video game applications are largely divided into four major non-portable hardware platforms: Sony PlayStation® 1,2 and 3 (PS1, PS2, and PS3); Microsoft Xbox®; and Nintendo Gamecube® and Wii<sup>TM</sup>; and PC-based games. Each of these platforms is different, so recorded games that run on one platform usually don't run on the other. Also, there is a problem of compatibility from one generation of devices to the next. Although the majority of software game developers create software designed independently for a particular platform, a proprietary layer of software (commonly referred to as a "game development engine") is used to run a particular game on a particular platform. called) becomes necessary to adapt the game for use on a particular platform. Each platform is sold to the consumer as a "console" (eg, a standalone box attached to a TV or monitor/speaker) or it is the PC itself. Typically, video games are sold on optical media such as Blu-ray, DVD, DVD-ROM or CD-ROM, which includes video games embodied as sophisticated real-time software applications.
The specific requirements for achieving a platform compatible with video game software will be unavoidable due to the real-time nature and high computational requirements of advanced video games. For example, from one generation of video games to the next (e.g. Xbox to Xbox 360 or Playstation 2) as there is the general compatibility of productivity applications (e.g. Microsoft Word) from one PC to another with faster processing units and cores. You can expect a fully compatible game from the "PS2") to the Playstation 3 ("PS3"). However, this is not the case with video games. Because video game manufacturers typically look for the best possible performance at a given price point when a video game generation is released, and the system's drastic structural changes are frequent, many games recorded for previous generation systems are not available for the latest generation. The system will not work. For example, the XBox is based on the x86-family processor, while the XBox 360 is based on the PowerPC family.
Techniques to emulate older architectures can be used, but a given video game cannot be run as a real-time application to achieve exactly the same behavior in emulation. This is a loss for consumers, video game console manufacturers and video game software publishers. For consumers, it means keeping both old and new generations of video game consoles connected to TVs to be able to play all the games. That means later adoption of new consoles and costs associated with emulation for console manufacturers. And for publishers, different versions of a new game must be released in order to reach all potential consumers--versions for each brand of video game (e.g. Xbox, Playstation) as well as versions of existing brands (e.g. PS2 and PS3). It means releasing a version for each version. For example, Electronic Art's "Madden NFL 08" was developed for the Xbox, XBox360, PS2, PS3, Gamecube, Wii, and PC and other platforms.
Portable devices, such as cellular ("cell") phones and portable media players, are also challenging game developers. Increasingly, such devices can connect to wireless data networks and download video games. However, there is a wide variety of cell phones and media devices on the market with different display resolutions and computing capabilities. Also, as such devices are typically constrained by power consumption, price, and weight, they typically incorporate advanced Graphics Processing Units ("GPUs") and devices, which are devices made by NVIDIA in Santa Clara, CA. It lacks the same graphics acceleration hardware. As a result, game software developers typically develop existing game titles for many different types of portable devices simultaneously. A user will find that an existing game title is not available on his cell phone or portable media player.
In the case of home game consoles, hardware platform manufacturers typically charge software game developers royalties to release games on their platform. Cell phone wireless carriers typically charge royalties to game publishers for downloading games to cell phones . In the case of PC games, there are no royalties paid for published games, but game developers typically pay high customer service costs to support a wide range of PC configurations and high costs due to possible installation issues. face a problem In addition, the barriers to piracy of gaming software are lower because PCs are typically easily reprogrammed by tech-savvy users, and games are more easily looted and distributed (eg, via the Internet). Therefore, for software game developers, there is a penalty and cost to release on game consoles, cell phones and PCs.
For console and PC software game publishers, the cost doesn't end there. To distribute games through retail channels, publishers charge a wholesale price below the selling price so that the retailer can make a profit. Publishers also typically have to pay for producing and distributing the physical media containing the game. Publishers are also entitled to "request by retailers to cover possible contingencies, such as if the game goes unsold, the price of the game drops, or if the retailer refunds part or all of the wholesale price and/or receives the game in exchange from the buyer." A "price protection fee" will be charged. In addition, retailers also typically charge a fee to publishers to help sell games in advertising flyers. Moreover, retailers gradually repurchase games from users who have finished playing, and then sell them as used games, typically sharing none of the game publisher's and used game revenues. An added cost burden for game publishers is that games are often pirated and distributed over the Internet, where users download and copy them for free.
Internet broadband speeds have increased and broadband connectivity has become more widespread in the United States and around the world, particularly in homes and Internet "cafes" where Internet-connected PCs are rented out, and games are increasingly being downloaded for PC or console. distributed through Also, broadband connections are increasingly used to play multiplayer and massively multi-user online games (both referred to as "MMOG" in this disclosure). These changes mitigate the costs and issues associated with retail distribution. Downloading online games poses some disadvantages to game publishers in that they typically have low distribution costs and little or no cost from unsold media. However, downloaded games are still prone to copyright infringement, and because of their size (often many gigabytes in size) they can take a very long time to download. In addition, multiple games can fill up small disk drives, such as those sold with portable computers or with video game consoles. However, expansion games or MMOGs require an online connection with the game to play, and piracy issues can be mitigated by requiring users to usually have a valid user account. Unlike linear media (such as video and music), which can be replicated by a microphone that records audio from a speaker or a camera that captures video of the display screen, each video game experience is unique and is simply a video game experience. /Cannot be duplicated using audio recording. Therefore, even in areas where there is a copyright law, it is not strongly sanctioned, and illegal copying is rampant. MMOGs can be protected from piracy and thus businesses can be supported. For example, Vivendi SA's "World of Warcraft" MMOG has been successfully deployed worldwide without piracy. And many online or MMOG games, such as Linden Lab's "Second Life" MMOG, generate revenue for the game's operator through an economic model built into the game where capital can be bought, sold, and created using online tools. create Thus, a mechanism in addition to a typical game software purchase or subscription may be paid for use of an online game.
On the other hand, piracy can often be mitigated due to the nature of online or MMOGs, and online game operators still face remaining challenges. Many games require substantially local (eg, home) processing resources for online or MMOG to work properly. If the user has a low performance local computer (eg without a GPU such as a low performance laptop), the game cannot be played. Moreover, depending on the age of game consoles, they are far behind the latest technology and will no longer be able to handle advanced games. Although the user's local PC can handle the computational requirements of the game, there is often the complexity of installation. There may be driver incompatibilities (eg, if a new game is downloaded, a previously installed game that relies on the older version of the graphics driver will install the new version of the graphics driver, which will not work). The console will run out of local disk space as many games are downloaded. Complex games develop over time from the game developer as bugs are discovered and fixed, or if changes are made to the game (eg, if the game developer finds the level of the game too difficult or too easy to play). Receives the typically downloaded patch. These patches require a new download. However, sometimes not all users complete the download of all patches. On the other hand, downloaded patches introduce other issues such as compatibility and consumption of disk space.
Also, while playing the game, a lot of data download is required to provide graphics or motion information for the local PC or console. For example, if a user enters an MMOG room and encounters a scene or character consisting of graphic data and actions that are not available on the user's local machine, the data of the place or character must be downloaded. This causes substantial delays during gameplay if the internet connection is not fast enough. And, if the place and character you are facing requires more than the storage space or computational power of your local PC or console, this may result in a situation in which the user cannot progress in the game or has to continue with low-quality graphics. Thus, online or MMOG games often limit their storage and/or computational complexity needs. Moreover, it often limits the transfer of large amounts of data during gaming. Online or MMOG games may also narrow the market of users who can play the game.
In addition, tech-savvy users will progressively reverse-engineer local copies of the game and transform the game to trick them. Cheats would be as easy as making a button press faster than a human could (eg, shooting a gun very repeatedly). In games that support in-game asset transactions, deception can reach a level of sophistication that results in fraudulent transactions involving assets of real economic value. When online or MMOG economic models are based on such capital transactions, this can have substantial detrimental consequences for game operators.
The cost of developing new games is increasing as PCs and consoles increasingly produce sophisticated games (eg, more realistic graphics, such as real-time ray-tracing, and more realistic behavior, such as real-time physical simulations). . In the early days of the video game business, video game development was very similar to the process for application software development; As opposed to tangible elements or "assets", they went into the development of software. Today, many sophisticated video game developments strive to more closely resemble the development of rich special effects motion pictures than software development. For example, many video games provide simulations of 3-D worlds, gradually creating photorealistic (eg, computer graphics that are photorealistic, such as lifelike images taken by photography) characters, props, and environments. One of the most challenging aspects of photorealism game development is creating a computer-generated human face that is indistinguishable from a lively working human face. Contour developed by Mova in San Francisco, California<sup>TM</sup> Face capture technology, such as Reality Capture, tracks and captures the face of an operator in motion with high resolution and precise geometry. This will allow a 3D face rendered on a PC or game console to be visually indistinguishable from the captured live facial movement. Accurately "photoreal" capturing and rendering of human faces is useful in several respects. First, recognizable celebrities or athletes are often used in video games (often employed at high cost), and flaws can be visible to the user, creating a distracting or unpleasant observation experience. Often, a high degree of detail requires rendering of high resolution textures and multiple polygons with polygons and/or textures that change from frame to frame, potentially as a face movie. required to achieve realism.
When a high polygon-count scene with detailed textures changes rapidly, a PC or game console supporting the game may have sufficient polygon and texture data to store enough polygon and texture data for the required number of animation frames being created in a game segment. You won't have RAM. Furthermore, the single optical or single disk drivers typically used in PCs or game consoles are usually slower than RAM, and typically the GPU may not be able to maintain the maximum data rate that can be tolerated for rendering polygons and textures. . Current games typically load most polygons and textures into RAM, meaning that a given scene is largely limited in complexity and persistence by the performance of RAM. In the case of facial animations, for example, this means that PCs or game consoles are limited to either low-resolution, non-photorealistic faces, or photo-realistic faces that can only be animated in a limited number of frames, before stopping the game, looking at polygons and textures. Load for many frames.
Watching a progress bar moving slowly across the screen as the PC or console displays a message similar to "Loading..." is accepted as an inherent drawback by today's users of complex video games. From disk (where "disk" is referred to as non-volatile optical or magnetic media, as well as non-disk media such as semiconductor "flash" memory, unless otherwise limited), the next scene is loaded from. The delay may take several seconds or minutes. This is a waste of time and can be quite frustrating for game players. As discussed earlier, much or all of the delay will be due to load times for polygons, textures, or other data from disk, but can also be attributed to the processor and/or GPU consumed while preparing the data for the scene on the PC or console. It will be part of the load time. For example, a soccer video game allows the player to select a number of players, teams, playgrounds and weather conditions, and the like. Thus, depending on the particular combination selected, different polygons, textures and other data (collectively "objects") are required for the scene (eg, different teams may have different colors and patterns on their uniforms). have). It would be possible to store objects on disk used to store games, precompute all or many objects in advance, and enumerate many or all of the various permutations. However, if the permutation is too large, the amount of storage required for all objects will be too large to stick to disk (or not too practical to download). Thus, existing PC or console systems are typically constrained by both complexity and duration of a given scene, resulting in long load times for complex scenes.
Another significant limitation of prior art video game systems and application software systems is that they increasingly use large databases, eg of 3D objects, such as polygons and textures, that need to be loaded into a PC or game console for processing. As mentioned above, such databases took a long time to load when stored locally on disk. However, the load time is much more severe if the database is stored in a remote location and can be accessed via the Internet. In such a situation, it may take minutes, hours, or even days to download many databases. Even worse, such databases are often very expensive (eg, detailed 3D models of tall-masted sailing ships used in games, movies, or historical documentaries) and are intended for sale to local end-users. . However, there is a risk of piracy as the database is downloaded to a local user. In many cases, users simply use a database for evaluation reasons to see if it suits the user's needs (eg, if a 3D costume for a game character has a pleasing appearance or appearance when the user performs a particular movement). want to download Long load times can deter users from evaluating 3D databases before making a purchase decision.
A similar issue arises in MMOGs such as games, especially games that allow users to gradually use custom characters. For a PC or game console to display a character, not only the behavior for the character (eg, if the character has a shield, whether that shield is strong enough to avoid spears or not), but also a 3D geometry database (polygons, textures, etc) need to be accessed. Typically when an MMOG is first played by a user, many databases for characters are already available as an initial copy of the game, which is downloaded to the game's optical disc or disc and used locally. However, as the game progresses, if the user encounters a database of characters or objects, their database is not available locally (eg, if other users have created custom characters) before the characters or objects are displayed. , its database must be downloaded. This actually causes a delay in the game.
Given the sophistication and complexity of video games, another challenge for video game developers and publishers with prior art video game consoles is that developing a video game will often take two to three years at a cost of tens of millions of dollars. Given that new video game console platforms are introduced at a rate of roughly every five years, game developers work on these game developments years in advance of the launch of new game consoles so that they have video games available at the same time as new platforms are released. It is necessary to start Some consoles from competitors are sometimes released simultaneously (eg, within a year or two), and the rest is the console's popularity, which will make the console the largest sales of video game software. For example, in the recent console cycle, Microsoft Xbox 360, Sony Playstation 3, and Nintendo Wii were introduced in the same general period. But in the years before its introduction, game developers have to "place their bets" as to whether the console platform will be more successful than the others, and they commit to matching their development resources. Video production companies must also allocate their limited production resources based on an assessment of whether the film will be a significant success prior to its release. Given the increasing level of investment required for video games, game production will gradually increase, like video production, and game production companies will typically allocate their production resources based on their assessment of the future success of a particular video game. do. However, unlike video production companies, these guesses are not based solely on the success of the production itself; Quite simply, the game is predicted based on the success of the game console on which it is intended to operate. At the same time, releasing a game on multiple consoles mitigates the risk, but this added effort increases costs and often delays the actual release of the game.
The application software and user environment on the PC are more computationally intensive, dynamic and interactive, and not only more visually expressive to the user, but also more useful and usable. For example, the new Windows Vista<sup>TM</sup> Operating system and Macintosh® Successive versions of the operating system include visual animation effects. Maya from Autodesk, Inc.<sup>TM</sup>Advanced graphics tools, such as , provide sophisticated 3D rendering and animation performance that places limits on the very latest technology CPUs and GPUs. However, the computational demands of these new tools create many practical issues for users and software developers of such products.
Since the visual display of the operating system (OS) is no longer sold--but still has to run on a wide range of computers, including previous generations, that can be upgraded to a new OS, the OS graphics requirements is limited to a large extent by the minimum common requirements of the target computers, which typically include computers without GPUs. This severely limits the graphics performance of the OS. Moreover, portable computers (eg, laptops) battery-powered limit visual display performance because the high computational performance on the CPU or GPU typically results in high power consumption and short battery life. Portable computers include software that automatically lowers processor activity to reduce power consumption when the process is not being used. On some computer models the user will manually lower the processor's operation. For example, Sony's VGN-SZ280P laptop is "Stamina" (for lower performance, longer battery life) on one side and "Speed" (for higher performance, shorter battery life) on the other. Includes a switch labeled . Running the OS on a portable computer should be able to function freely even when the computer is running at some of its peak performance. As such, OS graphics performance is often well below the available computational capabilities of the state-of-the-art.
High-end computationally intense applications such as Maya are often sold with the expectation that they will be used on high-performance PCs. This establishes at least common denominator requirements of generally very high performance, more cost and less portability. As a result, such applications have a much more limited target audience than general purpose OSs (or general purpose production applications, such as Microsoft Office), and are typically sold on a much smaller scale than general purpose OS software or general purpose application software. Potential consumers are further limited because it is difficult for prospective users to try such computationally intensive applications in advance. For example, consider that a student wants to learn how to use Maya, or a potential buyer who already has knowledge of such applications wants to try using Maya before investing in a purchase (which means that a high-end It will also involve purchasing a computer). On the other hand, although students or potential buyers can download it, they can get a physical media copy of the demo version of Maya, and if they can use Maya to the full potential of their computer (eg handling complex 3D scenes). If the running computer lacks performance, then they will not be able to fully evaluate the product information. This substantially limits the consumer for such high-end applications. It will also lead to a higher selling price because development costs are split into purchases in much smaller quantities than general purpose applications.
In addition, expensive applications may create further incentives for individuals and businesses to pirate application software. As a result, high-end application software suffers from widespread piracy despite considerable efforts by publishers of such software to mitigate such piracy by various techniques. Still, when using pirated high-end applications, users cannot eliminate the need to invest in an expensive, state-of-the-art PC to run the pirated copy. So, while they can get the use of a software application at a fraction of the actual retail price, users of pirated software still need to get or buy an expensive PC to take full advantage of the application.
The same is true for users of high-performance pirated video games. Although pirates can get the game for a fraction of the actual price, they are still asked to buy the expensive computing hardware needed to properly play the game (eg, a PC with an enhanced GPU or a high-end video game console such as the Xbox 360). do. When video games become a hobby for consumers in general, the extra cost for high-end video game systems can be prohibitive. This situation is exacerbated in countries where the average annual income of workers is currently generally quite low compared to the United States (eg China). As a result, a very small percentage of the total population owns a high-end video game system or high-end PC. In such countries, "Internet cafes", in which users pay for the use of computers connected to the Internet, are quite common. Often, these internet cafes have old model or low-end PCs that lack high performance features such as GPUs on which players can play computationally intensive video games. This is a major factor in the success of games running on low-end PCs, such as the highly successful "World of Warcraft" by Vivendi, and is commonly played in Internet cafes in China. On the other hand, computationally intensive games such as "Second life" are not quite easy to play on a PC installed from a Chinese Internet cafe. These games are virtually unacceptable to users who only have access to low-performance PCs in Internet cafes.
In addition, a barrier exists for users who are considering purchasing a video game, and the user will be more likely to try the demo version of the game by downloading the demo from home via the Internet. A video game demo is a full-fledged version of the game, often with some features disabled or limited gameplay size. This will involve a lengthy process (perhaps hours) of downloading gigabytes of data before the game is run or installed on a PC or console. In the case of a PC, understanding whether a special driver (eg DirectX, or OpenGL) is required for the game, downloading the correct version, installing it, and then determining if the PC can play the game will include The final step will involve determining if the PC has enough processing (CPU and GPU) performance, enough RAM, and a compatible OS (eg some games running on Windows XP, but not Vista). Thus, after a long process of attempting to run a video game demo, the user will be able to understand that the video game demo may not play well and given the user's PC preferences. Worse, if a user downloads a new driver to run a demo, this driver version may not be compatible with other games or applications that the user regularly uses on their PC, so installing the demo is a prerequisite for playing the game beforehand. It will either enable or disable the application. In addition to these barriers that embarrass users, they create barriers for video game software publishers and video game developers to sell their games.
Another issue that results from economic inefficiency relates to the fact that a given PC or game console is usually designed to accommodate what level of performance demands for applications and/or games. For example, some PCs have more or less RAM, slower or faster CPUs, and slower or faster GPUs if they don't have GPUs at all. Some games or applications take advantage of the sufficient computing power of a given PC or console, while not many games and applications. If a user's game or application selection falls short of the best performing performance of the local PC or console, the user will be wasting money on the PC or console for features that are not available. In the case of consoles, console manufacturers will pay more than necessary to subsidize console costs.
Another problem that exists in the marketing and enjoyment of video games involves allowing users to view other playable games before purchasing the game. Several prior art approaches exist for recording video games for later replay. For example, US Pat. No. 5,558,339 teaches recording game state information, including game controller actions, during "game play" on a video game client computer (owned by the same or a different user). This state information may later be used to replay some or all of the game action on a video game client computer (eg, PC or console). A significant drawback of this approach is that, in order for a user to view a recorded match, the user must have a video game client computer capable of playing the game and own a video game application running on the computer, so that gameplay is The state of the game is the same as when it was replayed. Beyond that, video game applications can be written in such a way that there are no possible performance differences between the recorded game and the played game.
For example, game graphics are typically computed frame by frame. In many games, the game logic sometimes follows, depending on whether a scene is particularly complex, or if there is some other delay, such as a slow-down execution (eg, other processing on the PC will be executed to take CPU cycles from the game application). It may take shorter or longer than one frame time to compute the graphics to display the frame. In such games, "threshold" frames that are computed in much less than one frame time (so-called CPU clock cycles) may eventually occur. When the same scene is computed again using exactly the same game state information, one frame time (e.g. if the internal CPU bus is slightly out of phase with the external DRAM bus) and that leads to a delay in CPU cycle time, Nevertheless, it can take less than a few milliseconds of CPU time in game processing, no significant delay in other processing that takes a few milliseconds of CPU time. Thus, when the game is played back, a frame is calculated and obtained with two frame times rather than a single frame time. Some behavior is based on how often the game computes new frames (eg, when the game samples input from the game controller). While the game is being played, this difference in time base for different actions does not affect the gameplay, but it can result in producing different results in the playback game. For example, if the trajectory of a basketball is computed at a steady 60 fps rate, but the game controller input is sampled based on the rate of the computed frame, when the game is recorded the computed frame rate will be 53 fps, but when the game is replayed Since it's 52 fps, this can make a difference between whether the basketball is blocked from entering the basket or not, resulting in different results. Thus, using game state to record a video game requires very careful game software design to ensure replay, using the same game state information that produces exactly the same results.
Another prior art approach for recording video games is simply to record the video output of a PC or video game system (eg, for a video capture board in a PC, or for a VCR or DVD recorder). The video may then be rewound and replayed, or alternatively, it may be a recorded video uploaded to the Internet after it has been compressed. A disadvantage of this approach is that when a 3D game sequence is played back, the user is limited to viewing the sequence from one point of view from which the sequence was recorded. In other words, the user cannot change the viewpoint of the scene.
Moreover, when a compressed video of a recorded game sequence played on a home PC or game console is made available to other users via the Internet, even though the video is compressed in real time, it uploads the compressed video to the Internet in real time. it would be impossible to The reason is that many homes around the world are connected to the Internet with high asymmetric broadband connections (eg, DSL and cable modems typically have much greater downstream bandwidth than upstream bandwidth). Compressed high-resolution video sequences often have higher bandwidth than the upstream bandwidth capability of the network, making real-time uploading impossible. Thus, after a game sequence has been played, there will be a significant delay (perhaps minutes or hours) before other users on the Internet can watch the game. Although this delay is tolerable in some situations (eg, viewing the results of a game player that happened at a previous time), it either eliminates the ability to view the game in real time (eg, a basketball tournament played by a champion player). , removing the "instant replay" capability when the game is playing live.
Another prior art approach allows a viewer equipped with a television receiver to watch a video game live, but only under the control of the television producers. Some television channels offer video game viewing channels in the United States and other countries, where television viewers can watch certain video game users (eg, play top tier players in tournaments) on the video game channel. This is achieved by having the video output of the video game system (PC and/or console) reflected in the video distribution and processing mechanism for the television channel. This is different from television channels broadcasting live basketball games in real time with several cameras reflecting different angles around the basketball court. Television channels can then make use of video/audio processing and effects mechanisms to handle output from various video game systems. For example, a television channel may overlay text on a video from a video game indicating the state of another player (such as may overlay text during a live basketball game), and the television channel may overlay text on a video game taking place during the game. You can overdub the audio of the narrator you can discuss. Additionally, video game output can be combined with camera recorded video of the game's real players of the game (eg, showing their emotional response to the game).
The problem with this approach is that in order to have the excitement of a live broadcast, such live video feeds must utilize the video distribution and processing machinery of the television channel in real time. However, as discussed above, this becomes impossible when a video game system is operating at home, especially if part of the broadcast includes live video from a camera that captures actual video of the game player. Also, in tournament situations, as mentioned above, there is a concern that gamers at home may modify and cheat the game. For this reason, such video game broadcasts on television channels are often combined video game systems at a common location (eg, a television studio or playground) with potential live cameras and television production equipment capable of accepting video reflected from multiple video game systems. and arranged with the player.
Although such leading video game TV channels can provide television viewers with highly exciting presentations, eg, experiences similar to live sporting events, with video game players appearing as "athletes", their forms of motion in the video game world and the real world. In all of their forms of action, these video systems are limited to situations where the player is physically very close to one another. And, after the television channels are broadcast, each broadcast channel can be viewed as only one video stream, which is selected by the television producer. Due to these restrictions and the high cost of airtime, production equipment and producers, you can only watch top-ranked player matches in top-tier tournaments.
Also, existing television channels that broadcast full screen images of video games to television viewers show only one video game at a time. This severely limits the choices of television viewers. For example, a television viewer may not be interested in watching a game at a set time. Other viewers may be only interested in seeing the gameplay of a particular player that is not characterized by the television channel at any given time. In other cases, the viewer may be interested in how a professional player can handle a particular level in the game. Still other viewers may wish to control the view from which the video game is viewed, different from that chosen by the production team. In short, television viewers will have many preferences for watching video games that are not supplied by a particular broadcast of a television network, even though several different television channels are available. For all of the above reasons, prior art video game television channels have significant limitations in presenting video games to television viewers.
Another disadvantage of prior art video game systems and application software systems is that they are complex and often suffer from errors, crashes, and/or unintended and undesirable behavior (collectively, "bugs"). Although games and applications typically go through a debugging and tuning process (often referred to as "Software Quality Assurance" or SQA) prior to release, almost invariably, bugs in the field occur when a game or application is released to a wider audience. do. Unfortunately, it is extremely difficult for software developers to find and identify many bugs after product release. It is also very difficult for software developers to find bugs. Even if they learn about a bug, much of the information available to them to recognize what is causing the bug may be limited. For example, a user will call the game developer's customer service and leave a message stating that when playing a game, the screen will start flashing and turn to a solid blue color and the PC will freeze. This provides the SQA team with very little useful information for finding bugs. Some games or applications that connect online can sometimes provide more information in some cases. For example, a "watchdog" process may sometimes be used to monitor a game or application for "crashes". When it crashes and uploads information to the SQA team over the internet, the monitoring processor can collect statistics about the state of the game or application processor (eg, the state of the stack, memory usage, how far the game or application has progressed, etc.). have. However, in a complex game or application, such information can take a very long time to read to accurately determine whether to use it by the user at the time of the crash. Even so, it would be possible to determine which sequence of events led to the collision.
Another issue with PCs and game consoles concerns service issues that are very inconvenient for consumers. Also, service issues typically affect their damaged PC or console.<b></b>It is required to be shipped in a specific box for safe shipment, after which a repair fee will be incurred if the PC or console is under warranty. Publishers of game or application software may also lose sales (or<b></b>may be affected by the use of online services).
1 shows the Sony Play Station ® 3, a prior art video gaming system such as Microsoft Xbox 360® Nintendo Wii, Windows-based personal computer or Apple Macintosh is shown. Each of these systems includes a central processing unit (CPU) for executing program code, typically a graphics processing unit (GPU) for performing advanced graphics operations, external devices and multiple types of inputs for communicating with the user. /Includes output (I/O). For simplicity, these configurations are shown combined together with a single unit 100 . The prior art video game system of FIG. 1 also includes an optical media driver 104 (eg, a DVD-ROM drive); a hard drive 103 for storing video game program code and data; a network connection 105 for playing multiplayer gameplay and downloading games, patches, demos, or other media; a random access memory (RAM) 101 for storing program codes currently executed by the CPU/GPU 100; a game controller 106 for receiving input commands from a user while playing a game; and a display device 102 (eg, SDTV/HDTV or computer monitor).
The prior art system shown in Figure 1 suffers from several limitations. First, optical drive 104 and hard drive 103 exhibit much slower access speeds compared to RAM 101 . When operating directly through the RAM 101, the CPU/GPU 100 is actually<b></b>It processes far more polygons per second than is possible when getting data and program code directly from drive 103 or optical drive 104 . The reason is that RAM 101 generally has a higher bandwidth and does not generally suffer from the long seek delays of the disk mechanism. However, only a limited amount of RAM is provided in these prior art systems (eg, 256 to 512 Mbytes). Thus, a "loading..." sequence is often required from the contiguous RAM 101 that is periodically filled with data for the next scene of the video game.
Some systems attempt to overlap the loading of program code concurrently with gameplay, but this can only be done when there is a known sequence of events (e.g., when a car goes down a road, the structure for accessing buildings on the side of the road is can be loaded while the car is running). For complex and/or rapid screen changes, this type of overlapping usually does not work. For example, if the user is at war and the RAM 101 is completely filled with data representing objects in the scene at that moment, if the user moves the scenes to the left to see an object that is not currently loaded in the RAM 101 , discontinuities in operation may occur because there is no time to load new objects from hard drive 103 or optical media 104 into RAM 101 .
Another problem with the system of FIG. 1 arises due to limited storage capacity of the hard drive 103 and optical media 104 . Although disk storage devices can be manufactured with relatively large storage space (eg, 50 gigabytes or more), they still do not provide enough storage space for certain scenarios encountered within current video games. For example, as noted above, a soccer video game may allow a user to select from 12 teams, players and stadiums throughout the world. Numerous document maps and environment maps are needed for each team, each player and each stadium to characterize the 3D appearance within that world (eg, each team has its own athletic shirt, each has its own document map).
One technique used to address this latter problem is for the game to pre-compute the documents and environment maps as soon as they are selected by the user. This may involve several computationally intensive processes, including stretching images, 3D mapping, shading, organizing data structures, and the like. As a result, there may be a delay for the user while the video game performs these operations. In principle, one way to reduce this delay is to perform all of these operations when the game is first developed, including all changes to the team, player roster, and stadium. The released version of the game then sends the already-processed data selected for the stadium selection, player roster, and team that is downloaded to the hard drive 103 via the Internet when the user selects to one or more servers or optical media 104 on the Internet. It will include all of this pre-processed data stored. As a practical matter, however, this data preloaded of all possible permutations within gameplay can easily be several terabytes, which far exceeds the capacity of today's optical media devices. Moreover, the data for a given team, player roster, and stadium selection can easily be more than a few hundred megabytes of data. With a so-called 10Mbps home network connection, it may take longer to download this data over the network connection 105 than to compute the data locally.
Thus, the conventional game structure shown in FIG. 1 introduces significant delays between key scene transitions in complex games to the user.
Another problem with the prior art, as shown in Figure 1, is that over the years, video games tend to get better and more demanding of CPU/GPU processing power. Thus, assuming an unlimited amount of RAM, video game hardware requirements will exceed the highest level of processing power available within such a system. As a result, users are required to upgrade their gaming hardware every few years to keep pace (or play new games at a lower quality level). An important conclusion among the trends for more advanced video games is that video game playing machines for home use are typically economically inefficient because their cost is usually determined by the requirements of the best performing game they can support. . For example, the Xbox 360 can be used to play games such as "Gears of War" that require hundreds of megabytes of RAM and a high performance CPU and GPU, or the Xbox 360 can have a very low performing CPU and only a few kilobytes. It can also be used to play the 1970's game Pac Man, which requires a lot of RAM. In fact, the Xbox 360 has enough computing power to host simultaneous Pac-Man games at once.
Video game machines are typically turned off for most of the time of the week. According to a June 2006 Nielsen Entertainment study of active gamers 13 and older, the average active gamer spends 14 hours per week playing console video games, or 12% of the total time per week. This means that the average video game console is out of operation 88% of the time, an inefficient use of expensive resources. This is especially important as video game consoles are often subsidized by suppliers to bring down consumer prices (with the expectation that subsidies can be reclaimed by royalties from future video game software purchases).
Video game consoles also incur costs associated with almost any consumer electronic device. For example, the electronics and mechanisms of the systems need to be enclosed and housing. The manufacturer needs to provide a service guarantee. Retailers who sell systems need to collect profits from any sale of the system and/or the sale of video game software. All these factors add to the cost of a video game console, which may be subsidized by the manufacturer, passed on to the consumer, or both.
Moreover, piracy is a major problem in the video game industry. The security mechanisms used on virtually all major video gaming systems have "cracked" over the years, leading to unauthorized piracy of video games. For example, the Xbox 360 security system was cracked in July 2006, and users can now download pirated copies online. Downloadable (eg, games for PC or Mac) games are partially vulnerable to piracy. In some parts of the world where piracy is weakly policyd, there is essentially no viable market for standalone video game software as users can buy pirated copies as easily as legitimate ones for a very small fraction of the cost. Also, with the cost of game consoles accounting for a high percentage of revenue in many parts of the world, even if breaches are controlled, few people can afford to buy a tech gaming system.
Moreover, the used game market reduces revenue for the video game industry. When users get bored with the game, they can sell the game to a store where they can resell the game to other users. This unauthorized but common practice significantly reduces the revenues of game publishers. Similarly, a 50% drop in on-demand sales usually occurs when there is a platform change every few years. This is because users stop buying games for older platforms when a newer version platform is about to be released (eg, when playstation 3 is about to be released, people stop buying playstation 2 games). Taken together, declining sales and increasing development costs associated with new platforms can have a very significant adverse impact on the suitability of game developers.
Also, new game consoles are very expensive. The Xbox 360, Nintendo Wii, and Sony PlayStation 3 all sell for hundreds of dollars. High-power PC gaming systems can cost more than $8,000. This represents a significant investment for users, especially considering the fact that many systems are sold to children and the hardware will become obsolete after a few years.
One approach to the anticipated problem is online gaming where the gaming program code and data are hosted on a server and transmitted to the ordering client machine as compressed video and audio streamed over a digital broadband network. Some companies, such as Finland's G-Cluster (now a subsidiary of Japan's SOFTBANK Broadmedia), currently offer these online services. Similar gaming services are provided by DSL and cable television providers and are available on local networks, such as networks within hotels. A major weakness of these systems is the issue of latency, ie, the amount of time it takes for a signal to travel to and from the gamer server, which is typically located within the operator's "head-end". Fast action video games (also known as "twitch" video games) require very low-latency between the time the user performs an action with a game controller and the time the display screen updates the result of the user's action. do. Low-latency is necessary for the user to have the perception that the game responds "instantly". Users may be satisfied with different latency intervals depending on the user skill level and game type. For example, a latency of 100 ms is pretty good for slow casual games (such as backgammon) or slow-action role playing games, but for fast action games, latencies of more than 70 or 80 ms will not allow the user to It can cause it to perform worse in the game, so it's unacceptable. For example, in games where fast reaction times are required, there is a steep slope in accuracy as latency increases from 50 to 100 ms.
When a game or application server is installed in close proximity, a controlled network environment, or network path to a user, is predictable and/or capable of withstanding bandwidth peaks, which reduces latency both in terms of maximum latency and consistency in latency. It is much easier to control (eg, so the user observes steady motion from streaming digital video over the network). This level of control can be achieved within a commercial office LAN environment between a cable TV network headend and a cable TV subscriber's home, or from a DSL central office to a DSL subscriber's home, or from a server or user. It is also possible to obtain specially-rated point-to-point private connections between businesses with guaranteed bandwidth and latency. However, in a game or application system that hosts games in a server center connected to the general Internet and then streams compressed video to users over a broadband connection, latency is caused by many factors, and is severe in the deployment of conventional systems. cause limitations.
In a typical broadband connection home, a user may have a DSL or cable modem for broadband service. These broadband services typically have a round-trip latency of 25ms (and more than double) between the user's home and the general Internet. Moreover, there are round-trip latencies incurred from routing data through the Internet to the server center. Latency over the Internet varies based on the delays that occur as data is given and it is routed. In addition to routing delays, round-trip latency also arises from the speed of light transmitted over the optical fibers that interconnect most of the Internet. For example, for each 1000 miles, about 22 ms is generated in round-trip latency due to the speed of light through fiber and other overhead.
Additional latency may occur due to the data rate of data streamed over the Internet. For example, if a user has a DSL service that is sold as a "6Mbps DSL service", then in practice the user will probably have less than 5Mbps of the maximum downstream, and will probably have several problems such as congestion during peak load times on the Digital Subscriber Line Access Multiplexer (DSLAM). You will see the connection periodically lowering due to factors. A similar issue exists elsewhere in the cable modem system network, or if there is congestion on a local shared coaxial cable looped through an adjacent "6Mbps cable modem service" much less than that. It may reduce the transmission rate of cable modems using connections sold as . If data packets at a stable rate of 4Mbps are streamed in one direction in User Datagram Protocol (UDP) format from the server center over this connection, then when all goes well, the data packets do not incur additional latency. will pass, but if there is congestion (or other obstacles), only 3.5 Mbps can be used to stream data to the user, so in a typical situation, packets will drop, resulting in lost data, or if they It will queue at the point of congestion until it can be transmitted, incurring additional latency in . Different points of congestion have different queuing capacity to catch delayed packets, so in some cases packets that cannot make through congestion will drop immediately. In other cases, several megabits of data will be queued up and eventually transmitted. However, in almost all cases, queues at congestion points have capacity limits, and if these limits are exceeded, the queues may overflow and packets will drop. Therefore, it is necessary to avoid exceeding the transfer rate capacity from the game or application server to the user in order to avoid the occurrence of additional latency (or worse packet loss).
Latency is also generated by the time required to compress the video within the server and decompress the video within the client device. Latency is further incurred while the video game in progress on the server is calculating the next frame to be displayed. Currently available video compression algorithms suffer from either high bitrates or high latency. For example, motion JPEG is an intraframe-only lossy compression algorithm characterized by low-latency. Each frame of video is compressed independently of other frames of video. When a client device receives a frame of compressed motion JPEG video, it will immediately stretch the frame and be able to display it, resulting in very low-latency. However, since each frame is individually compressed, the algorithm is unable to exploit the similarity between successive frames, and as a result, intraframe-only video compression algorithms suffer from very high bit rates. For example, 640x480 motion JPEG video at 60fps (frames per second) may require more than 40Mbps of data. This high bitrate for such a low resolution video window can be prohibitively expensive for many broadcast applications (and certainly for most consumer Internet-based applications). Moreover, since each frame is compressed independently, artifacts within a frame that might result from lossy compression appear to appear in different places within successive frames. This results in what appears to the viewer as visual artifacts that move when the video is stretched.
As they are used in prior art configurations, other compression algorithms such as MPEG2, H.264 or VC9 from Microsoft Corporation can achieve high compression ratios but at the cost of high latency. These algorithms perform intraframe-only compression of frames. These frames are known as key frames (typically referred to as "I" frames). Thus, these algorithms typically compare an I frame with previous and subsequent frames. Rather than compressing previous frames and successive frames independently, the algorithm is better to determine what changes within the image from frame I to previous and successive frames, and thus "B" for changes preceding the I frame. Stores these changes, called "P" frames, for changes following a frame, and an I frame. This results in a much lower bit rate than intraframe-only compression. However, it typically comes at the cost of higher latency. I frames are typically much larger (often ten times larger) than B or P frames, and consequently take proportionally longer to transmit at a given rate. For example, if I frames are 10 times the size of B and P frames, there are 29 B frames + 30 P frames = 59 interframes for every single I frame, or each "Group of Frames; There are a total of 60 frames per "GOP)". Thus, at 60 fps, there is one 60-frame GOP per second.
Imagine a transport channel having a maximum data rate of 2Mbps. To achieve the highest quality data stream within the channel, the compression algorithm will produce a 2 Mbps data stream, given the above ratio, which equals 2 megabits (Mb)/(59+10) = 30,394 bits/intraframe and 303,935 bits/I frame. When a compressed video stream is received by the decompression algorithm, in order for the video to run reliably, each frame needs to be displayed and decompressed at regular intervals (eg, 60 fps). To achieve this result, when a frame is subject to transmission latency, all frames need to be delayed at least by that latency, and thus the worst-case frame latency will define the latency for every video frame. . Since I frames are the largest, they cause the longest transmission latency, and the entire I frame will be received before the I frame is stretched and displayed (or interframe dependent on the I frame). Given a channel rate of 2Mbps, it would take 303,935/2Mb = 145ms to transmit an I frame.
An interframe video compression system as described above using a large percentage of the bandwidth of the transport channel will be subject to long latency due to the large size of the I frame, which is proportional to the average size of the frame. Or, to solve it another way, while prior art interframe compression algorithms achieve lower average data rates per frame than intraframe-only compression algorithms (eg 2 Mbps vs. 40 Mbps), they still do not because of large I frames. It suffers from a high peak per frame rate (eg, 303,935 * 60 = 18.2 Mbps). Keep in mind, with the above analysis, assume that both P and B frames are much smaller than I frames. While this is generally true, it is not true for scene changes, a lot of movement, or frames with high image complexity unrelated to previous frames. In this situation, P or B frames can be as large as I frames (if a P or B frame is larger than an I frame, a sophisticated compression algorithm will typically "force" the I frame and convert the P or B frame into an I frame. will be replaced with ). Thus, I frame-sized data rate peaks can occur at any instant within the digital video stream. Thus, with compressed video, when the average video rate approaches the rate capacity of the transport channels (if often, given the high rate demand for video), the high peak rate from an I frame or a large P or B frame is high It causes frame latency.
Naturally, the above discussion only characterizes the compression algorithm latency produced by large B, P or I frames within a GOP. If B frames are used, the latency is much higher. The reason is that before the B frame can be displayed, all of the B frames after the B frame and the I frame are received. Thus, within a group of picture (GOP) of a sequence such as BBBBBIPPPPPBBBBBIPPPPP with 5 B frames before each I frame, the first B frames are the subsequent B frames and video until the I frame is received. It cannot be displayed by a decompressor. So if the video is streamed at 60fps (i.e. 16.67ms/frame), 5 B frames and I frame takes 16.67 * 6 = 100ms to receive, and no matter how fast the channel bandwidth, it only has 5 B frames . 30 Compressed video sequences with B frames are very common. And, in a low channel bandwidth such as 2Mbps, the latency effect caused by the size of the I frame is greatly added to the latency effect because it waits for the B frames to arrive. Thus, on a 2Mbps channel, with a huge number of B frames, it is very easy to exceed 500ms of latency using conventional video compression techniques. If B frames are not used (at the cost of a low compression ratio for a given quality level), then, as described above, no B frame latency occurs, only the latency caused by peak frame sizes still occurs.
The problem is exacerbated by an important characteristic of many video games. Video compression algorithms using the GOP structure described above have been highly optimized for the use of live action or picture material intended for passive viewing. Typically, if the video or movie material is (a) typically inconvenient to watch and (b) viewed, simply if the camera or scene moves too abruptly, then usually the viewer may not be able to closely follow the action when the camera moves abruptly. Because there is, the camera (whether a real camera, or a virtual camera in the case of computer-generated animation) and scenes are relatively stable (eg, when a child blows out a candle on a birthday cake and suddenly returns to the cake and blows it again) If the camera drops when you turn it off, viewers will typically focus on the child and the cake and ignore the brief pause when the camera moves abruptly). In the case of a video interview, or video teleconferencing, the camera may be in a fixed position and not moving at all, resulting in very small data peaks. However, 3D high action video games are characterized by constant motion (eg, when the entire frame is in fast motion during a match, thinking 3D racing, or when the virtual camera is always jerking, who shoots the first person) think of). These video games result in a sequence of frames with large and frequent peaks that the user needs to clearly show what is happening during these sudden movements. As in this case, compression artifacts are not very good in 3D high action video games. Thus, the video output of many video games, by their nature, produces a compressed video stream with very high and frequent peaks.
There are limitations to server-hosted video games that stream video to the Internet, provided that users of fast-action video games have high latency, and little tolerance for all of the above given causes of latency. Moreover, users of applications that require a high degree of interactive interaction suffer similar limitations if the applications are hosted on the general Internet and stream video. In order for the route and distance from the client device to the server to be controlled to minimize latency and so that peaks can be accommodated with no latency, these services are provided by the hosting server within a headend (in the case of cable broadband) or in a central office. ) (for DSL), or requires a network configuration configured directly within the LAN (or on a specially-graded private connection) in a commercial setting. LANs with adequate bandwidth (typically with speeds of 100 Mbps-1 Gbps) and leased wires can typically support peak bandwidth requirements (eg, 18 Mbps peak bandwidth is a small fraction of 100 Mbps LAN capacity).
Also, peak bandwidth requirements can be accommodated by the home broadband infrastructure if special capacity is created. For example, on a cable TV system, digital video traffic may have a dedicated bandwidth that can handle peaks as large as I frames. And, on a DSL system, a higher speed DSL modem may be provisioned allowing higher peaks, or a specially-rated connection may be provisioned to handle higher data rates. However, the traditional cable modem and DSL infrastructure attached to the general Internet has much less tolerance for peak bandwidth requirements for compressed video. Thus, an online service that hosts a video game or application in server centers that are remote from the client device and streams compressed video output across the Internet over traditional home broadband connections, especially games that require very low-latency. and applications (eg, first person shooters and other multi-user, interactive action games, or applications that require fast reaction speed) suffer from significant latency and peak bandwidth limitations.
1 shows a prior art video gaming program architecture; 2A and 2B are diagrams illustrating a high-level system architecture according to an embodiment; 3 is a diagram showing the actual, speed, and required transmission rate for communication between a client and a server; 4A is a diagram illustrating a hosting service and an adopted client according to one embodiment; 4B illustrates typical latency associated with communication between a client and a hosting service; 4C is a diagram illustrating a client device according to an embodiment; 4D is a diagram illustrating a client device according to another embodiment; Fig. 4e shows an example of a block diagram of a client service in Fig. 4c; Fig. 4f shows an example of a block diagram of the client service in Fig. 4d; 5 shows an example of a form of video compression that may be employed in accordance with an embodiment; 6A shows an example of a form of video compression that may be employed in another embodiment; Figure 6b shows a peak in the data rate associated with a low complexity transmission, low motion video sequence; Figure 6c shows a peak in the data rate associated with a high complexity transmission, high motion video sequence; 7A, 7B illustrate examples of video compression techniques employed in one embodiment; 8 shows a further example of a video compression technique employed in one embodiment; 9A-9C illustrate examples of techniques employed in one embodiment for relaxed data rate peaks; Figures 10a and 10b illustrate one embodiment for efficiently packing an image tile in a packet. 11A-11D illustrate an embodiment employing a forward error correction technique; 12 illustrates one embodiment using a multi-core processing unit for compression; 13A, 13B illustrate geographic positioning and communication between and with a hosting service in accordance with various embodiments; 14 is a diagram illustrating typical latency associated with communication between a client and a hosting service; 15 is a diagram illustrating an example of a hosting service server center architecture; 16 shows an example of a screen shot of an embodiment of a user interface that includes multiple live video windows; Fig. 17 is an illustration of the user interface of Fig. 16 upon selection of a particular video window; Fig. 18 is a diagram showing the user interface of Fig. 17 according to the enlargement of a specific video window of full screen size; 19 shows an example of collaborative user video data overlaid on a screen of a multiplayer game; 20 is a diagram illustrating an example of a user page for a game player on a hosting service; 21 is a diagram showing an example of a 3D interactive advertisement; 22 shows an example sequence of steps for producing a photoreal image with a textured surface from a surface capture of a live performance; 23 shows an example of a user interface page allowing selection of linear media content; 24 shows a graph showing the time elapsed before a web page goes live for connection speed.
BRIEF DESCRIPTION OF THE DRAWINGS The present invention may be more fully understood from the accompanying drawings and the following detailed description. However, the specific embodiments shown do not limit the spirit of the disclosed invention, but are for understanding and explanation only.
In the following detailed description, specific details are set such as a device type, a system configuration, and a communication method in order to provide an understanding of the present disclosure. However, one of ordinary skill in the art will recognize that these specific details may not be required in order to practice the described embodiments.
2A and 2B show that video games and software applications are hosted by a hosting service 210 over the Internet 206 (or other public or private network) under a subscription service and user premises 211 ("user premises"). is connected by the client device 205 at the location where the user is located, including outside, if using a mobile device. The client device 205 may be a general-purpose computer such as Microsoft Windows or a Linux-based PC or Macintosh computer of Apple Inc. connected to the Internet by wire or wirelessly, or may be connected to an internal or external display device 222 or video and client devices such as set-top boxes (wired or wirelessly connected to the Internet) that output audio to the monitor or TV set 222 , and possibly mobile devices that are connected to the Internet by wire or wirelessly.
Some of these devices may have their own user input devices (eg, keyboards, buttons, touchscreens, trackpads, or inertial sensing bars, video capture cameras and/or motion tracking cameras, etc.), or they may be wired or wireless It uses the connected internal input device 221 (eg, keyboard, mouse, game controller, inertial sensing rod, video capture camera and/or motion tracking camera, etc.). As described in more detail below, the hosting service 210 includes servers of various levels of performance, including high-power CPU/GPU processing capabilities. During the use of an application or play of a game on the hosting server 210 , the home or office client device 205 receives keyboard and/or controller input from the user, which in turn receives the game or application software ( For example, if the user presses a button that instructs a character on the screen to move to the right, then the game program creates a sequence of video images showing the character moving to the right). It sends the controller input to a hosting service 210 that generates a frame and in response executes the gaming program code. This sequence of video images is compressed using a low-latency video compressor, and the hosting service 210 transmits the low-latency video stream over the Internet 206 . A home or office client device decodes the compressed video stream and renders the uncompressed video image to a monitor or TV. As a result, the computing and graphics hardware requirements of the client device 205 are significantly omitted. The client 205 only needs to have the processing power to send keyboard/controller input to the Internet 206 , and to decompress and decrypt the compressed video stream received from the Internet 206 , which In fact, a personal computer is capable of running software on its CPU (e.g. Intel's Core Duo CPU running at approximately 2 GHz is capable of compressing 720p HDTV encoded using Windows Media VC9 and compressors such as H.264). can be solved). And, in the case of some client devices, a dedicated chip can also perform video decompression for such standards in real time at a lower cost and less power consumption than general purpose PCs, as required for modern PCs. In order to significantly perform the functions of transmitting controller input and decompressing video, the home client device 205 may include an optical drive such as the prior art video game system shown in FIG. 1, a hard drive, or some special graphics processing unit ( GPU) is not required.
As gaming and application software become more complex and more photorealistic, they will require high performance CPUs, GPUs, more RAM, larger and faster disk drives, and computing power in hosting services 210 will be continuously upgraded, The end user will not be required to update the home or office client platform 205 as its processing needs will remain fixed for display resolution and frame rate with existing video decompression algorithms. Thus, hardware limitations and compatibility issues no longer exist in the system illustrated in Figures 2a, 2b today.
Moreover, since the games and application software run only on the servers of the hosting service 210, the user's home or office (here "office" would include any non-resident setting if any qualified, including, for example, a school room). There is never a copy of the game or application software (either in the form of optical media, or as downloaded software). This significantly mitigates the likelihood that the game or application is pirated, as well as the probability of a valuable database that can be used by the pirated game or application. Indeed, if specialized servers are required to play a game or application software that is not very effective for use in the home office, it will work in the home or office even though piracy of the game or application software has been obtained. won't be able to
In one embodiment, hosting service 210 is a game or application software developer (generally a software development company, game or movie studio, or game or application) that develops video games and designs games that can be run on hosting service 210 . It is referred to as a software publisher; 220) provides software development tools. Such tools allow developers to take advantage of features of hosting services that are not normally available on standard PCs or game consoles (e.g., complex geometries ("geometry)" polygons otherwise defined here as 3D datasets. , very fast access to a very large database of textures, rigging, lighting, motion and other elements and parameters).
Other business models are possible under this architecture. Under one model, the hosting service 210 collects a subscription fee from the end user and pays a royalty to the developer 220 , as shown in FIG. 2A . In an alternative implementation shown in FIG. 2B , the developer 220 collects a subscription fee directly from the user and pays the hosting service 210 for hosting the game or application content. This underlying theory is not limited to a specific business model for providing online gaming or application hosting.
<b>Compressed video features (</b><b>Compressed</b><b></b><b>Video</b><b></b><b>Characteristics</b><b>)</b>
As discussed above, one prominent problem with the online provision of a service of video game service or application software is latency. A latency of 70 to 80 ms (from the point of the input device actuated by the user to the response point displayed on the display device) is an upper limit for games and applications that require fast response times. However, this is very difficult to achieve with respect to the architecture shown in Figures 2a and 2b due to many practical physical constraints.
As shown in Figure 3, when a user subscribes to an Internet service, the connection is typically rated at a nominal maximum transfer rate 301 to the user's home or office. Depending on the provider's policy and the capabilities of the transmitting equipment, the maximum transfer rate is strictly enforced faster or slower, but typically the actual available transfer rate is slower for one of many other reasons. For example, there is too much network traffic on the DSL central office or local cable modem loop, there is noise on the cable causing dropped packets, or the provider will set a maximum bit per user per month. Currently, maximum downstream data rates for cable and DSL services typically range from several hundred kilobits per second (Kbps) to 30 Mbps. Cellular service is typically limited to hundreds of Kbps of downstream data. However, the speed of broadband services and the number of users subscribing to broadband services will increase dramatically over time. Currently, some analysts estimate that 33% of US broadband subscribers have a downstream speed of 2Mbps or higher. For example, some analysts predict that by 2010, more than 85% of US broadband subscribers will have a transmission rate of 2Mbps or higher.
As shown in Fig. 3, the actual available maximum transfer rate 302 will fluctuate over time. Therefore, at low-latency, the online game or application software associated with it is difficult to predict the actual available data rate in a particular video stream. If the rate 303 is actually increased than the maximum available data rate 302, and for any number of complex scenes, a given number of frames-per-second (fps) at a given resolution (eg 640 × 480 at 60 fps). ), if required to maintain a given level of quality, will result in distorted/lost images and lost data on the user's video screen. Other services temporarily buffer (eg, queue up) additional packets - and serve packets to clients at available transfer rates, increasing latency, an unacceptable result for many video games and applications. causes Finally, some Internet service providers have malicious attacks, such as denial of service attacks (users of well-known technology used by hackers to block network connections), which increase the transmission rate and disconnect users from the Internet for a certain period of time. you will see Accordingly, the embodiments described herein take steps to ensure that the required bitrate for the video game does not exceed the maximum available bitrate.
<b>Hosting service architecture (</b><b>Hosting</b><b></b><b>Service</b><b> A</b><b>rchitecture</b><b>)</b>
4A illustrates an architecture of a hosting service 210 according to one embodiment. Hosting service 210 may be located in a single server center or distributed across multiple server centers (to provide low-latency connections to users who have a lower-latency path to some server centers than others, one or more servers). to provide redundancy in case of center failure and to provide a balanced load among multiple users). The hosting service 210 will eventually include hundreds of thousands or millions of servers 402 serving a very large user base. The hosting service control system 401 controls the hosting service 201 as a whole, and oversees routers, servers, video compression systems, billing and accounting systems, and the like. In one embodiment, the hosting service control system 401 runs on a distributed processing Linux-based system coupled with a RAID array used to store a database for system statistics, server information, and user information. In the foregoing detailed description, the various operations performed by the hosting service 210 are initiated and controlled by the hosting service control system 401 unless viewed as a result of other specific systems.
The hosting service 210 includes a number of servers 402 such as those currently available from Intel, IBM, Hewlett-Packard, and others. Alternatively, the server 402 may be assembled with a customary configuration of components, or may eventually be integrated such that the entire server is implemented as a single chip. Although this diagram shows a small number of servers 402 for illustrative purposes, in an actual deployment there will be as few as one server 402, or as many as millions of servers or more. Servers 402 are configured in the same manner (as an example of any of the configuration parameters, have the same CPU type and performance; with or without GPU, have the same GPU type and performance if there is a GPU; have the same number of CPUs and GPUs; and have the same number of CPUs and GPUs; have a quantity of type/speed of RAM; and have the same RAM configuration), or different subsets of servers 402 may be configured in the same configuration (eg, 25% of servers will be configured in a particular way). (50% are configured differently, and 25% are already configured the other way), or all servers 402 will be different.
In one embodiment, the server 402 is diskless, eg, it uses local mass storage itself (which is optical or magnetic storage or semiconductor-based storage such as flash memory or other similar means of performing a similar function). Rather than having, each server accesses shared mass storage through a fast backplane or network connection. In one embodiment, this fast connection is a Storage Area Network (403) connected to a set of Redundant Arrays of Independent Disks (RAID) 405 with connections between devices running using Gigabit Ethernet. As is known in the prior art, SAN 403 can be combined with many RAID arrays 405, resulting in extremely high bandwidth - approaching or potentially exceeding the bandwidth available in RAM currently used in game consoles and PCs. will be used And while RAID arrays based on rotating media, such as magnetic media, often have significant seek time access latencies, RAID arrays based on semiconductor storage can be implemented with lower access latencies. In other configurations, some or all of the servers 402 provide some or all of their own mass storage locally. For example, server 402 may store frequently accessed information such as its operating system and its operating system and copy of a video game or application in its operating system and low-latency local flash-based storage, but based on geometry or geometry or on a less frequent basis. We will use the SAN to access the RAID array 405 based on game state information.
Moreover, in one embodiment, hosting service 210 employs low-latency video compression logic 404 detailed below. The video compression logic 404 may be implemented in software, hardware, or some combination thereof (in some embodiments described below). Video compression logic 404 includes logic for audio compression as well as visual material.
During operation, while playing a video game or using an application on the user premises 211 via a keyboard, mouse, game controller or other input device 421 , the control signal logic 413 in the client 415 allows the user transmits control signals 406a and 406b (typically in the form of UDP packets) to the hosting service 210 indicative of a button press (and other form of user input) actuated by Control signals from the existing user are sent to the appropriate server (or servers, if multiple servers respond to the user's input device). As shown in FIG. 4A , the control signal 406a will be sent to the server 402 via the SAN. Alternatively or additionally, the control signal 406b may be sent directly to the server 402 via a hosting service network (eg, an Ethernet-based local area network). Regardless of how they are transmitted, the server or servers execute game or application software responsive to control signals 406a, 406b. Although not shown in FIG. 4A , various networking elements, such as firewalls and/or gateways, may be located at the edge of the hosting service 210 and/or on the user premises 211 between the Internet 410 and the home or office client 415 . ) will handle incoming and outgoing traffic from the edge. The graphics and audio output of the running game or application software - eg, a new continuous video image - is provided to low-latency video compression logic 404 which compresses a sequence of video images according to low-latency video compression techniques as described herein. and transmits a compressed video stream, typically compressed or uncompressed audio, to the client 415 over the Internet 410 (or over the optimized high speed network service over the general Internet, as described below). ) to return Low-latency video decompression logic 412 at the client 415 then decompresses the video and audio streams and renders the decompressed video stream, typically playing the decompressed audio system on the display device 422 . . Alternatively, the audio may be reproduced on a speaker that is not separate or separate from the display device 422 , as it should be noted, although the input device 421 and the display device 422 are shown independently in FIGS. 2A and 2B . , they may be integrated into client devices such as portable computers or mobile devices.
A home or office client 415 (described previously as home or office client 205 in FIGS. 2A and 2B ) would be a very low cost, low power device with very limited computing or graphics performance, much more limited or local You won't have mass storage. In contrast, each server 402 coupled with a SAN 403 and multiple RAID 405 can be an exceptionally high performance computing system, and indeed if multiple servers are used cooperatively in a parallel processing configuration,<b></b>There are virtually no limits to the amount of computing and graphics processing power that can be purchased to withstand. And, because of the low-latency video compression 404 and the low-latency video compression 412 , perceptually to the user, the computing power of the server 402 will be provided to the user. When the user presses a button on the input device 421, the image on the display 422 is updated perceptually in response to the button press without significant delay, as if a game or application software was running locally. Thus, a very low performance computer or home or office client 415 with only a cheap chip executes the low-latency video compression and control signal logic 413, and the user is efficient at a remote location with arbitrary computing power as available locally. is provided as It presents users with the most advanced, processor-intensive (typically new) video games and highest performance applications.
4C shows a very basic and inexpensive home or office client device 465 . This device is an embodiment of a home or office client 415 from FIGS. 4A and 4B . It is approximately 2 inches long. It has an Ethernet jack 462 that interfaces with an Ethernet cable with Power over Ethernet (PoE), from which the connection between power and the Internet is obtained. It can operate NAT within a network that supports Network Address Translation (NAT). In the office environment, many new Ethernet switches are equipped with PoE and bring PoE directly to the Ethernet jack in the office. In this situation, all that is required is an Ethernet cable from the wall jack to the client 465. If the available Ethernet connection does not carry power (e.g., to a DSL or cable modem at home, but not PoE), an inexpensive available wall" brick that allows output Ethernet with PoE and Ethernet cables without power ( bricks)" (eg power supplies).
The client 465 includes control signal logic 413 ( FIG. 4A ) coupled with a Bluetooth wireless interface, which interfaces with a Bluetooth input device 479 such as a keyboard, mouse, game controller and/or microphone and/or headset. do. Additionally, one embodiment of the client 465 may alternatively provide a display device 468 capable of supporting 120 fps video and audio with a pair of shuttered glasses (typically via infrared) to shutter one eye. ) can be combined to output video at 120 fps. The effect perceived by the user is a stereoscopic 3D image "jumps out" of the display screen. A display device 468 that supports such an operation is a Samsung HL-T5076S. With the video stream for each eye separated, in one embodiment two independent video streams are compressed by the hosting service 210 , the frames are fitted in time, and the frames are decompressed into two independent decompression within the client 465 . elongated by treatment.
The client 465 also includes low-latency video decompression logic 412, which decompresses the incoming video and audio and provides the video and audio to the TV via HDMI (High-Definition Multimedia Interface) (SDTV). Standard Definition Television) or HDTV (High Definition Television) is output through a connector 463 that is plugged in, or a monitor 468 supporting HDMI. If the user's monitor 468 does not support HDMI, HDMI to DVI (Digital Visual Interface) can be used, but audio will be lost. Under the HDMI standard, display capabilities 464 (eg, supported resolutions, frame rates) communicate with display device 468 , which is then communicated to hosting service 210 via internet connection 462 and display device Compressed video can be streamed in a format suitable for
FIG. 4D shows the home or office client device 475 identically except that it further has an external interface with the home or office client device 465 shown in FIG. 4C . In addition, the client 475 may obtain power from an external power supply adapter (not shown) that plugs into the wall as well as PoE to obtain power. When client 457 uses a USB input, video camera 477 provides compressed video to client 475 , which is uploaded by client 475 for hosting service 210 in the use described below. . Consisting of camera 477 is a low-latency compressor using the compression technique described below.
In addition to an Ethernet connector for Internet connectivity, the client 475 also has an 802.11g wireless interface for the Internet. Both interfaces can use NAT within a network that supports NAT.
Additionally, in addition to having an HDMI connector for video and audio output, the client 475 also has a dual link DVI-I connector that includes an analog output (and a standard adapter cable will provide a VGA output). It also has analog outputs for composite video and S-video.
For audio, client 475 has left/right analog stereo RCA jacks and a TOSLINK output for digital audio output.
In addition to the Bluetooth wireless interface for the input device 479, a USB jack is also required to interface with the input device.
4E illustrates one embodiment of the internal architecture of client 465 . All or some of the devices shown in the diagrams may be implemented in custom designs or off-the-shelf, custom ASICs or several discrete devices, Field Programmable Logic Arrays.
Ethernet with PoE 497 is attached to Ethernet interface 481 . Power 499 is derived from Ethernet 497 with PoE and connected to the rest of the devices at client 465 . Bus 480 is a common bus for communication between devices.
Control CPU 483 (any small CPU such as the MIPS R4000 CPU series at 100 MHz with embedded RAM is usually suitable) running a small client control application from Flash 476 runs the protocol stack for the network and also It communicates with the hosting service 210 and configures all devices at the client 465 . It also interfaces with the input device 469 and, if necessary, sends packets back to the hosting service 210 with the user controller data protected by forward error correction. Control CPU 483 also monitors packet traffic (eg, if packets are lost or delayed, they also timestamp their arrival). This information is sent back to the hosting service 210 to constantly monitor the network connection and adjust what is sent accordingly. The flash memory 476 initially loads a control program for the control CPU 483 and a serial number unique to a specific client 465 unit. This serial number is allowed to uniquely identify the client 465 unit.
The Bluetooth interface 484 communicates with the input device 469 wirelessly through an antenna inside the client 465 .
Video decompressor 486 is a low-latency video decompressor configured to perform the video decompression described herein. Many video decompression devices exist either as intellectual property (IP) or off-the-shelf designs that can be integrated into FPGAs or custom ASICs. One company proposing IP for H.264 decoders is Ocean Logic of Manly, New South Wales, Australia. The advantage of using IP is that the compression techniques used here do not conform to compression standards. Some standard stretchers are sufficient here to be configured flexibly to accommodate compression techniques, while others are not. However, with IP, you have all the flexibility to redesign your stretcher as required.
The output of the video expander is coupled to a video output subsystem 487 , which is coupled to video for video output of the HDMI interface 490 .
The audio decompression subsystem 488 uses an available standard audio decompressor, it may be implemented as IP, or the audio decompressor may be implemented within a control processor 483 that may, for example, implement a Vorbis audio marketer.
The device performing audio decompression is coupled with an audio output subsystem 489 that is coupled with audio for audio output of the HDMI interface 490 .
4F illustrates one embodiment of the internal architecture of the client 475 . The architecture is the same as that of the client 465 except for an optional external DC power from the wall-plugged power supply adapter and an added interface, replacing the incoming power from the Ethernet PoE 497 if so used. The common purpose with client 465 will not be repeated below, but additional purposes are described below.
The CPU 483 configures the additional device while communicating with the additional device.
The WiFi subsystem 482 provides a wireless Internet that is alternatively connected to the Ethernet 497 through the antenna. WiFi subsystem 485 is available from a wide range of manufacturers including Atheros Communication of Santa Clara, CA.
USB subsystem 485 provides an alternative to Bluetooth communication for wired USB input device 479 . The USB subsystem is fairly standard and quite available in FPGAs and ASICs, as well as often designed as off-the-shelf devices that perform other functions such as video decompression.
Video output subsystem 487 produces a wider range of video output within client 465 . In addition to the video output provided by HDMI 490 , DVI-I 491 , S-video 492 , and composite video 493 are provided. Also, when the DVI-I 491 interface is used for digital video, the display capability 464 is passed back to the control CPU 483 at the display device, thereby announcing the hosting service 210 of the display device 478 capability. can All interfaces provided by video output subsystem 487 are fairly standard interfaces and are available in quite a number of forms.
The audio output subsystem 489 outputs digital audio through the digital interface 494 and outputs audio in analog form through the stereo analog interface 495 .
<b>round trip </b><b>latency</b><b> analysis(</b><b>Round</b><b>-</b><b>Trip</b><b></b><b>Latency</b><b></b><b>Analysis</b><b>)</b>
Of course, as can be seen from the benefit of the preceding paragraph, the round-trip latency between the user's actions using the input device 421 and showing the result of the action on the display device 420 is no greater than 70-80 ms. This latency is at the user's premises (211).<b></b>All factors in the path from the input device 421 to the hosting service 210 and back to the user premises 211 for the display device 422 should be considered. Figure 4b shows the various configurations and networks through which signals must travel, the above elements and networks being a timeline listing exemplary latencies that can be expected in practical implementation. It should be noted that Figure 4b is simplified so that only critical path routing is shown. The routing of other data used for other features of the system is described below. Double arrows (eg, arrow 453 ) indicate round-trip latency, one arrow (eg, arrow 457 ) indicates one-way latency, and "~" indicates approximate measurements. This would indicate that there will be real-world situations where the listed latencies are not achievable, but in many cases in the United States, when using DSL and cable modems to connect to user premises 211, these latencies are described in the next section. can be achieved under the circumstances. Also, while cellular wireless connections to the Internet will certainly work in the system shown, most current US cellular data systems (such as EVDO) incur very high latencies and will not be able to achieve the latencies shown in Figure 4b. will be. However, this underlying theory will be practiced in future cellular technologies that have the capability to implement these levels of latency.
Once initiated from the input device 421 at the user premises 211 , the user operates the input device 421 , and the user control signal is sent to the client 415 (which may be a basic device such as a set-top box, or it may be a PC or may be software or hardware running on another device such as a mobile device), is packetized, and the packet is given a destination address in order to reach the hosting service 210 . The packet also contains information indicating the user when the control signal comes in. The control signal packet(s) are then sent to the WAN interface 442 via the firewall/router/Network Address Translation (NAT) device 443 . The WAN interface 442 is an interface device provided to the user premises 211 by the user's Internet Service Provider (ISP). The WAN interface 442 may be a cable or DSL modem, a WIMAX transceiver, a fiber transceiver, a cellular data interface, an Internet Protocol interface on a powerline, or many other interfaces for the Internet. Moreover, the firewall/router/NAT device 443 (and potential WAN interface 442 ) will be integrated with the client 415 . An example of this would be a mobile phone, comprising software for executing the functions of a home or office client 415 as well as a means to connect and transmit wirelessly to the Internet via some standard (eg 802.11g).
The WAN interface 442 is then configured for the user's Internet Service Provider (ISP), which is a facility that provides an interface between the general Internet or private network and the WAN transport connected to the user's premises 211, here "interconnection". It transmits a control signal to what is called a "point of presence". The nature of the interconnection location will vary depending on the nature of the Internet service provided. In the case of DSL, it will typically be the central office of the telephone company where the DSLAM is located. In the case of a cable modem, it will typically be the cable Multi-System Office head end. In the case of a cellular system, this will typically be the control room associated with the cellular tower. However, whatever the nature of the interconnection location, the control signal packet(s) will be transmitted to the general Internet 410 . The control signal packet(s) will then be transmitted to the WAN interface 410 for the hosting service 210, most easily via the fiber transceiver interface. WAN 441 will then send a control signal packet to routing logic 409 (which may be implemented in many different ways, including Ethernet switches and routing servers), which evaluates the user's address and for a given user. Send the control signal(s) to the correct server 402 .
The server 402 then takes the control signal as an input for the game or application software running on the server 402 and uses the control signal to process the next frame of the game or application. When the next frame is generated, the video and audio will be output from the server 402 to the video compressor 404 . Video and audio will be output from the server 402 to the compressor 404 through various means. First, the compressor 404 will be built into the server 402 , so the compression will be performed locally within the server 402 . Alternatively, the video and/or audio may be output in packetized form over a network connection such as an Ethernet connection over a shared network such as a SAN 403 or a private network between the server 402 and the video compressor 404 . . Alternatively, the video may be output from the server 402 through a video output connector, such as a DVI or VGA connector, and then captured by the video compressor 404 . The audio is also output from the server 402 as digital audio (eg, via a TOSLINK or S/PDIF connector) or analog audio that is encoded and digitized by audio compression logic within the video compressor 404 .
A video compressor 404 captures the video frame and the audio generated during the frame time from the server 402, which will then compress the video and audio using the techniques described below. Once the video and audio are compressed, it is packetized to an address for back to the user's client 415 , which is sent to the WAN interface 441 , and then the video and audio packets are sent over the general Internet 410 and , sends video and audio packets to the user's ISP interconnection location 441 , then sends video and audio packets to the WAN interface 442 at the user's premises, and sends the video and audio packets to the firewall/router/NAT device 443 . and audio packets, and then video and audio packets to the client 415 .
The client 415 decompresses the video and audio, then displays the video on the display device 422 (or the client's integrated display device), and displays the video on the display device 422 or a separate amplifier/speaker or integrated amplifier/speaker to the client. transmit audio to
For the user to perceive the entire process described without perceptual lag, the round-trip delay needs to be 70 or 80 ms or less. Any latency delays in the described round-trip path are under the control of the hosting service 210 and/or not the user and others. Nevertheless, based on testing and analysis of a large number of real-world scenarios, the following is a rough estimate.
One-way transmission times for sending control signals 451 are typically less than 1 ms, and round-trip transmissions through user premises 452 are typically achieved using firewall/router/NAT switches over readily available consumer-grade Ethernet in approximately 1 ms. can be User ISPs vary widely in their round-trip delay 453, but with DSL and cable modem providers, we can typically see between 10 and 25 ms. Round-trip latencies on the regular Internet can vary greatly depending on how much traffic is being transmitted and what failures in the transmission, but typically the regular Internet provides a fairly optimal route (and issues these issues). is described below), latency is usually determined by the speed of light through the fiber, given the distance to its destination. As explained further below, we establish that the maximum distance we expect from the hosted service 210 away from the user's premises 211 is approximately 1000 miles. At 1000 miles (2000 miles round trip), the practical time to transmit a signal over the Internet is approximately 22 ms. The WAN interface 441 for the hosting service 210 is typically a commercial grade fiber high-speed interface with negligible latency. Thus, typical Internet latency 454 is typically between 1 and 10 ms. One-way routing 455 latency through hosting service 210 can be achieved in 1 ms or less. Server 402 will typically compute a new frame for a game or application in less than one frame time (which is 16.7 ms at 60 fps) and 16 ms is the maximum one-way latency 456 that is reasonable to use. In an optimized hardware implementation of the video compression and audio compression algorithms described herein, compression 457 can be completed in less than 1 ms. In the less optimized version, the compression may take as much as 6ms (of course longer compared to the optimized version, but such execution will affect the round-trip overall latency and shorten it to keep 70-80ms (e.g. typical The distance allowed over the Internet may be reduced) and other latencies will be required). The round-trip latency of the Internet 454 , the user ISP 453 , and the user premises routing 452 have already been considered, so what remains is the video decompression 458 latency, whether the video decompression 458 is implemented on dedicated hardware, or If run in software on the client device 415 (a PC or mobile device), it can be highly dependent on the size of the display and the performance of the expanding CPU. Typically, stretch 458 takes between 1 and 8 ms.
Thus, by adding together the worst-case latencies seen in the implementation, we can determine the worst-case round-trip latency that can be expected to be experienced by a user of the system shown in Figure 4a. They are: 1+1+25+22+1+16+6+8 = 80ms. And, in practice (with the warnings described below), this is using a prototype version of the system shown in Figure 4a, using a standardized Windows PC as a client device in the US, home DSL and cable modem connections. The round-trip latency is shown by doing this. Of course, a better-than-worst case scenario could result in much shorter latencies, but they are not required to develop widely used commercial services.
In order to achieve the latencies listed in Figure 4b on the general Internet, a video compressor 404 and a video decompressor 412 in the client 415 from Figure 4a are required to generate a packet stream of very special characteristics, resulting in The packet sequence generated through the entire path from the hosting service 210 to the display device 422 does not cause delay or excessive packet loss, and in particular, the user through the WAN interface 442 and the firewall/router/NAT 443 Consistently drop the limit of bandwidth available to users on your Internet connection. Moreover, the video compressor generates a sufficiently robust packet stream so that it can withstand the inevitable packet loss and packet reordering that normally occurs in Internet and network transmissions.
<b>low-</b><b>latency</b><b> video compression (</b><b>Low</b><b>-</b><b>Latency</b><b></b><b>Video</b><b></b><b>Compression</b><b>)</b>
In order to achieve the aforementioned goals, one embodiment leads to a new approach for video compression that reduces the peak bandwidth requirements and latency for transmitted video. Prior to a detailed description of this embodiment, an analysis of current video compression techniques will be provided with respect to FIGS. 5, 6A, and 6B. Of course, if users are provided with sufficient bandwidth to handle the data rates required by these techniques, these techniques will be adopted consistent with the underlying theory. Note that audio compression is not described here other than to mention that video compression and audio compression are synchronized and run concurrently. Prior art audio compression techniques exist to satisfy the requirements for this system.
5 shows a specific prior art technique for video compression in each individual video frame 501 - 503 being compressed by compressor logic 520 using a specific compression algorithm to produce a series of compressed frames 511 - 513 . are doing One embodiment of this technique is a "motion JPEG" in which each frame is compressed according to the Joint Picture Expert Group (JPEG) compression algorithm based on a discrete cosine transform (DCT). A variety of other types of compression algorithms will be employed, while still complying with this underlying theory (eg, wavelet-based compression algorithms such as JPEG-2000).
One problem with this form of compression is that it reduces the bit rate of each frame, but it does not exploit the similarity between successive frames to reduce the bit rate of the entire video stream. For example, as shown in Fig. 5, assuming that 604 x 480 x 24 bits/pixel = 604 * 480 * 24/8/1024 = 900 kilobytes/frame (KB/frame), for an image of a given quality, a motion JPEG compresses at a rate of only 10, resulting in a data stream of 90 KB/frame. At 60 frames/sec, this would require a channel bandwidth of 90 KB*8 bits*60 frames/sec = 42.2 Mbps, which would be far too high bandwidth for almost all home internet connections in the US today, and too much for many office internet connections. It will be high bandwidth. In fact, if a constant data stream is required at such a high bandwidth, even in an office LAN environment, it will be provided to only one user, and it will consume most of the 100Mbps Ethernet LAN bandwidth, and the Ethernet switch that supports the LAN. make it much more burdensome will it . Thus, compression for moving pictures is inefficient (as described below) compared to other compression techniques. Moreover, single-frame compression algorithms such as JPEG and JPEG-2000 may not be able to recognize as still images (e.g., artifacts in dense foliage in a scene will not appear as artifacts, as the eye does not know exactly how dense foliage can be represented). It uses lossy compression algorithms that produce unknown compression artifacts. However, once the scene is in motion, the artifacts can stand out because the eye finds artifacts that change from frame to frame, despite the fact that artifacts in one area of the scene cannot be recognized as still images. This is a result of the perception of "background noise" in the sequence of frames, similar to what appears in the "snow" noise seen during insignificant analog TV reception. Of course, this form of compression will still be used in some of the embodiments described herein, but generally speaking, to avoid background noise in the scene, a high data rate (eg, a low compression ratio) is required for a given perceptual quality.
H.264 or other forms of compression, such as Windows Media VC9, MPEG2 and MPEG4, are all effective at compressing video streams because they take advantage of the similarity between successive frames. All of these techniques follow the same general technique for compressing video. Thus, although the H.264 standard is described, the same general theory applies to other compression algorithms. A large number of H.264 compressors and decompressors are available, including the x264 open source software library for compressing H.264 and the FFmpeg open source software library for decompressing H.264.
6A and 6B show that a series of uncompressed video frames 501 - 503, 559 - 561 is a series of "I frames" 611,671 by compression logic 620; "P Frames" (612,613); and "B frames" (670). The vertical axis in Figure 6a generally represents the resulting size of each encoded frame (although the frames are not drawn to scale). As described, video coding using I-frames, B-frames, and P-frames will be readily understood by those skilled in the art, where I-frames 611 are fully uncompressed frames 501 (as described previously). DCT-based compression (similar to compressed JPEG images). P frames 612 and 613 are generally significantly smaller than I frames 611 . Because they use data from the preceding I frame or P frame; That is, they contain data indicating changes between preceding I frames or P frames. B-frame 670 is similar to that of P-frame, except that the B-frame uses the frame for the next reference frame as well as potential frames in the preceding reference frame.
In the following discussion, the requested frame rate is 60 frames/sec, each I frame is approximately 160 Kb, the average P frame and B frame is 16 Kb and a new I frame is generated every second. With this set of parameters, the average data rate will be: 160Kb + 16Kb * 59 = 1.1Mbps. This transfer rate falls within the maximum transfer rates of many current broadband Internet connections for home and office use. This technique also tends to avoid the background noise problem from only intraframe encoding because P frames and B frames track the difference between frames, so compression artifacts reduce the background noise problem described above or frame frames. It tends to disappear and not appear in two frames.
One problem with the aforementioned types of compression is that even if the average data rate is generally low (eg 1.1 Mbps), a single I frame will take several frame times to transmit. For example, a 2.2 Mbps network connection using the prior art (e.g., a DSL or cable modem with a 2.2 Mbps peak of the maximum available data rate 302 from Figure 3A) would typically run at 1.1 Mbps with 160 Kbps I frames, each 60 frames. Streaming the video would be appropriate. This can be achieved by having a stretcher that queues up the video one second before stretching it. In 1 second, 1.1 Mb of data will be transmitted, which will easily accommodate up to the maximum available data rate of 2.2 Mbps, and the available data rate will drop periodically to over 50%. Unfortunately, this prior art approach will result in 1 second latency for the video due to the 1 second video buffer at the receiver. Such a delay is suitable for many preceding applications (eg, playback of linear video), but it is a latency that is too long for fast action video games that cannot tolerate more than a latency of 70-80ms.
If an attempt was made to reduce the 1 second video buffer, it would still not result in a reasonable omission in latency for fast action video games. In some cases, as described above, the use of B frames will require reception of not only the I frame, but all B frames preceding the I frame. If we assume 59 non-I frames roughly divided between P frames and B frames, there will be at least 29 B frames and one I frame received before any B frames are displayed. Thus, despite the bandwidth of the available channel, it would require a delay of 29+1=30 frames of each 1/60 second period, or a latency of 500 ms. Obviously it's too long.
Thus, another approach would drop B frames and use only I frames and P frames. (One consequence of this would be that the bitrate would be increased for a given quality level, but for consistency in this example, each I frame is 160Kb , and the average P frame is 16Kb in size, so the data rate still assumes that it is still 1.1Mbps to remove One problem that remains with this approach is that, as is typical for most homes and many offices, I frames on low bandwidth channels are too large for average P frames, and the transmission of I frames adds substantial latency. This is shown in Figure 6b. The video stream bitrate 624 is below the maximum available bitrate 621 except for I frames, where the peak bitrate required for the I frame 623 is the maximum available bitrate 622 (and the rated even the maximum data rate 621) is exceeded. The data rate required by the P frame is less than the maximum available data rate. Although the maximum available data rate peak at 2.2 Mbps remains steady at its 2.2 Mbps peak rate, it will take 160 Kb/2.2 Mb = 71 ms to transmit an I frame, and if the maximum available data rate 622 reaches 50% ( 1.1Mbps), it will take 142ms to transmit the I frame. So, the latency of sending I frames will drop somewhat between 71 and 142ms. This latency added to the latency identified in FIG. 6b is added up to 70 ms in the worst case, which is 141 until the image appears on the display device 422 at the time the user operates the input device 421 It will result in a total round-trip latency of 222ms, which is too high. And if the maximum available data rate drops below 2.2 Mbps, the latency will increase further.
There is also "jamming," which is a serious consequence for ISPs with peak rates (623) that generally far exceed the available rates (622). In other ISPs the mechanism will behave differently, but the following behavior is quite common between DSL and cable modem ISPs when receiving packets at a rate much higher than the available rate 622: (a) queuing them It delays (inducing latency), (b) drops some or all packets, and (c) disables the connection for a period of time (mostly because of malicious attacks such as "denial of service" the ISPs associated with it). Therefore, it is not a viable option to transmit the packet stream at the maximum data rate with the characteristics shown in FIG. 6B. The peak 623 will be queued up at the hosting service 210 and transmit at a rate below the maximum available rate leading to the unacceptable latency described in the preceding paragraph.
Moreover, the video stream rate sequence 624 shown in Figure 6b is a well "tame" video stream rate sequence, the expected result of compressing the video into a video sequence that has very little motion and does not change very much. It will be a sequence of data rates of some sort (eg, the camera is in a fixed position and little movement, it will be common in video conferences).
The video stream rate sequence 634 shown in FIG. 6C is a typical sequence found in a motion picture or video game with much more motion created in a video game or some application software. Note the addition of I frame peaks 633, there are also P frame peaks (such as 635 and 636) that in many cases exceed the maximum available data rates and are quite large. Although these P-frame peaks are not quite as large as the I-frame peaks, they are still too large to be carried by the channel at their maximum data rate, and, like I-frame peaks, the P-frame peaks must be transmitted slowly (due to increased latency).
On a high bandwidth channel (e.g., a 100Mbps LAN, or a high bandwidth 100Mbps private connection) the network will be able to withstand large peaks like I frame peak 633 or P frame peak 636, and in theory low-latency would be maintained. could be However, such networks are frequently shared among many users (eg, in an office environment), and such "peaky" data, especially if network traffic is sent over a private shared connection (eg, from a data center to an office) will affect the performance of the LAN. First of all, it should be borne in mind that these examples are typically low resolution video streams 640x480 pixels at 60fps. 1920x1080 of an HDTV stream at 60fps is easily handled by modern computers and displays, and 2560x1440 resolution displays at 60fps are increasingly available (eg, Apple, Inc's 30-inch displays). A high motion video sequence of 1920x1080 at 60fps would require 4.5Mbps using H.264 compression for a reasonable quality level. If we assume an I frame peak at 10x nominal data rate, it will result in a 45 Mbps peak, but it is not only a small but still significant P frame peak. If several users receive a video stream on the same 100Mbps network (eg, a private network connection between an office and a data center), how can you tune the peaks from some users' video streams, overwhelm the network's bandwidth, and add users to the network. It's easy to know if you're potentially overpowering the bandwidth of the backplate of the switch you're supporting. Even in Gigabit Ethernet networks, if enough users are coordinated enough at once, it can overcome the network or network switch. And, once 2560x1440 resolution video becomes more common, the average video stream rate will be 9.5 Mbps, possibly resulting in a peak rate of 95 Mbps. Needless to say, 100Mbps connections between data centers and offices (which are exceptionally fast connections today) will be completely flooded with peak traffic from a single user. Thus, although LAN and private network connections can better tolerate peaky streaming video, streaming video with high peaks is undesirable and will require special planning and facilities by the IT department of the office.
Of course, for standard linear video applications this issue is not an issue. Because the bitrate is "smoothed" at the point of transmission and data for each frame below the maximum available bitrate 622, a buffer at the client stores the sequence of I,P,B frames before being stretched. Thus, the bitrate on the network remains close to the average bitrate of the video stream. Unfortunately, this introduces latency even though B frames are not used, which is unacceptable for low-latency applications such as video games and applications that require fast response times. A prior art solution to mitigate video streams with high peaks is to use an encoding technique called "Constant Bit Rate (CBR)". Although the term CBR is thought to mean that all frames are compressed to have the same bit rate (eg size), usually what it refers to is a certain number of frames (in our case, 1 frame). A compression paradigm in which the maximum bitrate traversed is allowed. For example in the case of Figure 6c, if the CBR constraint is applied to encoding a limited bit rate, then a rated maximum bit rate 621 of e.g. 70%, then the compression algorithm will limit the compression of each frame so that the rated maximum The bit rate 621 will be compressed into fewer bits. The result of this is that frames that typically require more bits to maintain a given quality level are "starved" of bits and the image quality of the frames requires more bits than 70% of the maximum data rate (621). It will be worse than any other frame that doesn't. This approach can produce acceptable results for some form of compressed video where (a) small movements or scene changes are expected and (b) the user can tolerate periodic quality reduction. A good example of an application that goes well with CBR is video teleconferencing, since there are few peaks, and if the quality is obviously reduced (eg, if captured by a camera, while being filmed without being bit bit enough for high quality image compression, it may result in reduced image quality), which is acceptable to most users. Unfortunately, CBR is not well suited for many other applications that have high complexity scenes or a lot of movement and/or where a reasonable constant level of quality is required.
The low-latency compression logic 404 employed in one embodiment uses several different techniques for streaming low-latency compressed video and conveying the scope of the problem. First, the low-latency compression logic 404 generates only I frames and P frames, thereby alleviating the need to wait several frame times to decode each B frame. Moreover, in one embodiment, as shown in FIG. 7A , the low-latency compression logic 404 may process each uncompressed frame 701 - 760 into a series of "tiles" and individually I frames or P frames. Encode each tile as A group of compressed I frames and P frames are referred to herein as "R frames" 711 - 770 . In the specific example shown in Figure 7a, each uncompressed frame is subdivided into a 4x4 matrix of 16 tiles. However, these fundamental principles are not limited to any particular subdivision scheme.
In one embodiment, the low-latency compression logic 404 divides a video frame into multiple tiles, and divides the video frame into I frames (eg, the tiles are 1/16th the size of the entire image).<sup>th</sup> is an individual video frame of , and the compression used for these "mini" frames is I-frame compression), which encodes (e.g. compresses) one tile from each frame, and P frames (e.g., each " Mini" 1/16<sup>th</sup> The compression used for the frame is P frame compression) and encode the rest of the tiles. Tiles compressed as I frames and P frames will be generally referred to as "I tiles" and "P tiles". With each successive video frame, the tile encoded as an I tile changes. Thus, at a given frame time, only one of the tiles in a video frame is an I tile, and the rest of the tiles are P tiles. For example, in FIG. 7A , tile 0 of the uncompressed frame 701 is the I tile I<sub>0</sub>and the remaining 1 - 15 tiles are P tiles P to create an R frame (711).<sub>1</sub> in P<sub>15</sub>is encoded as In the next uncompressed video frame 702, tile 1 of the uncompressed frame 701 is the I tile I<sub>1</sub> and the remaining tile 0 and tile 2 to 15 are P file P to produce an R frame 712 .<sub>0</sub><sub>,</sub> P<sub>2</sub> in P<sub>15</sub>is encoded as Thus, for a tile, the I tile and the P tile are continuously interleaved in time over successive frames. The process is that the R tile 770 is an I tile (eg, I<sub>15</sub>) until the last tile is created in the encoded matrix. The process then resumes generating another R frame, such as frame 711 (eg, encoding the I tile for tile 0). Although not shown in FIG. 7A , in one embodiment, the first R frame of a video sequence of R frames contains only I tiles (eg, subsequent P frames have reference frame data therefrom to associate an action). do. Alternatively, in one embodiment, the startup sequence generally uses the same I tile pattern, but does not include P tiles for tiles that have not yet been encoded as I tiles. In other words, some tiles are not encoded with any data until the first I tile arrives, thereby avoiding the startup peak in video stream rate 934 in FIG. 9A , which is described in detail below. Moreover, other sizes and shapes may be used for the tiles while adhering to this basic theory, as described below.
Video stretching logic 412 operating on the client 415 stretches each tile as if it were a separate video sequence of I and P frames, and then renders each tile for a frame buffer that drives the display device 422 . . For example, R frame 711 to frame 770 I<sub>0</sub> and P<sub>0</sub>is used to reconstruct tile 1 and so on. As mentioned above, stretching of I and P frames is well known in the art, and stretching of I and P tiles can be accomplished by having multiple instances of video stretching running on the client 415 . . Although the multi-operation process appears to increase the computational burden on the client 415, it is not really the case because the tile itself is proportionately smaller for many additional processings, and thus the number of displayed pixels is limited to one processing. Equivalent to using normal full size I and P frames, if any.
This R frame technique significantly mitigates the typical bandwidth peaks associated with the I frames shown in Figures 6b and 6c, since any given frame is typically composed mostly of P frames that are smaller than I frames. For example, assuming that a typical I frame is 160 Kb, the I tile of each frame shown in Figure 7a would be approximately 1/16th or 10 Kb of this size. Similarly, assuming that a typical P frame is 16 Kb, the P frame for each of the tiles shown in Figure 7a would be approximately 1 Kb. The final result of an R frame is approximately 10Kb + 15*1Kb = 25Kb. So, each 60-frame sequence is 25Kb*60 = 1.5Mbps. So at 60 frames/sec, this will require channel performance capable of sustaining a bandwidth of 1.5 Mbps, but with a much lower peak due to the I tiles being distributed throughout the 60-frame interval.
Referring to the previous example with the same assumed data rate for the I frame and the P frame, the average data rate is 1.1 Mbps. This is because, in the previous example, a new I frame was introduced only once every 60 frame times, whereas in this example, the 16 tiles making up an I frame cycle complete in 16 frame times, and this equivalent of an I frame is derived every 16 frame times, resulting in a marginally higher average data rate. Inducing more frequent I frames does not increase the data rate linearly. This is due to the fact that P frames (or P tiles) mainly encode the difference between the preceding frame and the next. So, if the preceding frame is significantly different from the next frame, the P frame will be very large. However, since P frames are derived largely from preceding frames rather than actual frames, the encoded frame result will contain more errors (eg, visual artifacts) than I frames with a reasonable number of bits. And, when one P frame follows another P frame, what can occur is an accumulation of errors that is worse when there are long sequence P frames. Now, a sophisticated video compressor will find that after P frames in a sequence the quality of the image deteriorates, and if necessary, it will allocate more bits to subsequent P frames to improve the quality, but if it is the most efficient This is a course motion, replacing the P frame with the I frame. So, when a long sequence of P frames is used (e.g. 59 frames, as in the previous example), especially when the scene has a lot of complexity and/or motion, typically more bits are removed from the I frame as more bits are removed. It becomes necessary for P frames.
Alternatively, in order to view a P frame from the opposite viewpoint, a P frame that closely follows an I frame tends to require fewer bits than a P frame that is further removed from the I frame. So, in the example shown in FIG. 7A , the P frame is not added more than 15 frames removed from the I frame, and as in the previous example, the P frame may be 59 frames removed from the I frame preceding it. Thus, as I frames are more frequent, P frames are smaller. Of course, the exact relative size will vary based on the characteristics of the video stream, but in the example of Figure 7a if an I tile is 10Kb, typically a 0.75Kb sized P tile would result in 10Kb + 15*0.75Kb = 21.25Kb, or At 60 frames per second, the data rate will be 21.25Kb*60 = 1.3 Mbps, or at 1.1 Mbps it is about 16% higher than the stream with I frames followed by 59 P frames. Again, the relevant results between these two approaches to video compression will vary depending on the video sequence, but typically, we believe that using R-frames is better than using sequences of I/P frames for a given level of quality. We find empirically that it requires about 20% more bits for . But, of course, R frames have significantly less latency than I/P frame sequences and dramatically reduce peaks that make usable video sequences.
R frames can be constructed in a variety of different ways depending on the nature of the video sequence, the reliability of the channel, and the available data rate. In an alternative embodiment, the other large number of tiles is 16 or more in a 4x4 configuration. For example, 2 tiles will be used in 2x1 or 1x2, 4 tiles will be used in 2x2, 4x1, or 1x4 configurations, or 6 tiles will be used in 3x2, 2x3, 6x1 , or 1x6 configurations, or 8 tiles will be used in 4x2 (as shown in Figure 7b), 2x4, 8x1, or 1x8 configurations. Tiles don't have to be square, and video frames don't have to be square or even rectangular. Tiles may be divided into any suitable form used for the video stream and application.
In another embodiment, the cycling of I and P tiles is not fixed to multiple tiles. For example in a 4x2 configuration 8 tiles, a 16 cycle sequence may be used as shown in FIG. 7B . The sequentially occurring uncompressed frames 721 , 722 , 723 are divided into 8 tiles, each 0-7, and each tile is individually compressed. In the R frame 731, only 0 tiles are compressed as I tiles, and the remaining tiles are compressed as P tiles. During the next R frame 732 all 8 tiles are compressed as P tiles, then during the next R frame 733 1 tile is compressed as I tiles and all other tiles are compressed as P tiles. And, the sequence continues for 16 frames, with only I frames generated from all other frames, so the last I tile is 15<sup>th</sup> During the frame time (not shown in Figure 7b), 7 tiles are created for 16 tiles.<sup>th</sup> During the frame time, the R frame 780 is compressed using all P tiles. Then, the sequence starts again with 0 tiles compressed as I tiles and the other tiles compressed as P tiles. As in the previous embodiment, the entire video sequence of the first frame will typically be all I tiles to provide a reference for the P tiles at that point. Cycling of I tiles and P tiles does not require multiple tiles. For example, with 8 tiles, each frame with I tiles can follow up to 2 frames with all P tiles until another I tile is used. In already another embodiment, some tiles will be sequenced with I tiles more often than others. For example, some areas of the screen are known to have more motion than required from frequent I tiles, while others will be sequenced with more frequent I tiles. It is more static, requiring less from I tiles (eg, showing scores for a game). Moreover, although each frame is shown as a single I tile in Figs. 7A, 7B, a plurality of I tiles will be encoded in a single frame (depending on the bandwidth of the transmission channel). Conversely, some frames or frame sequences may be transmitted without I tiles (eg, only P tiles).
A good reason for the study approach in the previous paragraph is that it does not seem to result in larger peaks without having I tiles distributed over every single frame, and the behavior of the system is not straightforward. Each tile is compressed separately from the other tiles, and tiles can be less efficient the smaller the encoding of each tile is, since the compressor of a given tile cannot use similar image features and similar behavior from other tiles. So, in general, splitting the screen into 16 tiles will be less efficient encoding than splitting the screen into 8 tiles. However, if the screen is tiled 8, it causes the data of the entire I frame to be derived in every 8 frames instead of every 16 frames, which results in a very high data rate overall. So, by deriving the entire I frame every 16 frames instead of every 8 frames, the overall data rate is reduced. Also, by using 8 larger tiles instead of 16 smaller tiles, the overall data rate will be reduced, which will also mitigate to some extent data peaks caused by large tiles.
In another embodiment, the low-latency video compression logic 404 in FIGS. 7A and 7B is preconfigured by settings based on known characteristics of the video sequence to be compressed, or automatically based on an analysis of the image quality in each tile in progress. to control the allocation of bits for various tiles in the R frame. For example, in a racing video game, the player's car front (which is usually motionless in the scene) occupies a large area in the lower half of the screen, while the upper half of the screen is filled with entirely approaching roads, buildings, and landscapes. , which is almost always in motion. If the compression logic 404 allocates the same number of bits to each tile, then in the uncompressed frame 721 in FIG. In the uncompressed frame 721, the top half of the screen will be compressed to a higher quality than the tiles (tiles 0 - 3). If this particular game, or particular scene of the game, is known to have such characteristics, the operator of the hosting service 210 may use compression logic ( 404) is constructed. Alternatively, the compression logic 404 may measure the compression quality of the tile after the frame is compressed (using one or more of a number of compression quality metrics, such as Peak-To-Noise Ratio (PSNR)), if it If determined over a certain window of time, some tiles will consistently produce better quality results, and gradually more bits will be added to the tiles producing lower quality results until the various tiles reach a similar level of quality. assign to In an alternative embodiment, the compressor logic 404 assigns bits to a particular tile or group of tiles to achieve high quality. For example, having a higher quality at the center of the screen than at the edges will give a better overall perceptual appearance.
In one embodiment, in order to improve the resolution of certain regions of the video stream, the video compression logic 404 is generally configured to have more complex scenes and/or motions than regions of the video stream with less complex scenes and/or motions. Smaller tiles are used to encode regions of the video stream. For example, as shown in FIG. 8 , smaller tiles are adopted around a moving character 805 in one region of an R frame 811 (potentially of the same tile size and followed by a series of R frames (not shown). not)). When the character 805 then moves to a new area of the image, as shown, smaller tiles are used around this new area within another R frame 812 . As noted above, a variety of other sizes and shapes may be adopted as "tiles" while still complying with this fundamental theory.
While the periodic I/P tiles described above essentially reduce the peaks in the bit rate of a video stream, they are particularly rapidly changing or very complex video images, such as those generated by motion pictures, video games, and some application software. In the case of , it does not remove the peak as a whole. For example, during a sudden scene change, a complex frame will be followed by another complex frame that is completely different. Although some I tiles may only be up to a few frame times before the transition, they cannot help in this situation because the material of the new frame is not related to the pre-I tiles. In such situations (and in other situations where many, if not all, images change), video compressor 404 determines how many if not all of the P tiles coded more efficiently, such as I tiles, and what results. It will determine if it is a very large peak in the data rate for that frame.
As discussed earlier, the case for most consumer-grade Internet connections (and many office connections) is simple, and it simply "jams" according to the rated maximum transfer rate 621, exceeding the maximum available transfer rate shown as 622 in FIG. 6C. (jam)" data is not possible. The rated maximum transfer rate 621 (eg, "6Mbps DSL") is essentially a marketing number for users considering the purchase of an Internet connection, but in general it does not guarantee any level of performance. For the purposes of this application, that is pointless, our only concern is the maximum available bitrate 622 at the time the video is streamed over the connection. Consequently, in Figures 9a and 9c, when we describe the solution for the peaking problem, the rated maximum data rate is omitted from the graph, and only the available maximum data rate 922 is shown. The video stream rate must not exceed the maximum available rate 922 .
To handle this, the first thing the video compressor 404 must do is determine the peak rate 941, which is the rate the channel can consistently handle. This speed is determined by many techniques. One such technique would be to gradually send an increasing high transfer rate test stream from the hosting service 210 to the client 415 in Figures 4a and 4b, where the client provides feedback regarding some level of packet loss and latency to the hosting service. something to do. When packet loss and/or latency starts to show a sharp increase, it is an indication that the maximum available data rate 922 is being reached. Thereafter, the hosting service 210 may gradually reduce the transfer rate of the test stream until the client 415 receives an acceptable level of packet loss and latency at an acceptable level for a reasonable period of time. Over time, the peak transfer rate 941 will fluctuate (eg, if other users in the home are using excessive Internet connections), and the client 415 determines whether there is packet loss or increased latency, the maximum available transfer rate 922 . ), it is necessary to constantly monitor it to know whether it falls below the pre-established peak data rate 941 and, if so, whether it is the peak data rate 941 . Similarly, if time passes and the client 415 finds that packet loss and latency remain at an optimal level, it can request the video compressor to slowly increase the data rate to see if the maximum available data rate has been increased ( (e.g., if other users in the home stop using excessive Internet connections), again waiting for packet loss and/or high latency to indicate that the available data rate (922) has been exceeded, and again the lower level reaches the peak data rate (941). ), but one is probably higher than that level before testing the increased bitrate. So, by using this technique (and other techniques like this) the peak data rate 941 can be found and adjusted periodically as needed. The peak rate 941 will be set to the maximum rate that can be used by the video compressor 404 to stream the video to the user. The logic for determining the peak transfer rate may be executed at the user premises 211 and/or the hosting service 210 . The client 415 at the user premises 211 performs an operation to determine the peak transfer rate and sends this information back to the hosting service 210 ; In the hosting service 210 , the server 402 in the hosting service performs an operation for determining a peak transmission rate based on statistics (eg, packet loss, latency, maximum transmission rate, etc.) received from the client 415 .
9A is an example of a video stream rate 934 having substantially complex scenes and/or operations generated using the cyclic I/P tile compression techniques shown in FIGS. 7A, 7B, and 8 and previously described. is showing The video compressor 404 is configured to output compressed video at an average rate below the peak rate 941 , and most of the time the video stream data stays below the peak rate 941 . A comparison of the bitrate 934 with the video stream rate 634 shown in FIG. 6c made using I/P/B or I/P frames shows that cyclic I/P tile compression produces a smoother bitrate. Still, frame 2X peak 952 (which approaches 2X peak data rate 942) and frame 4X peak 954 (which approaches 4X peak data rate 944), whose rate exceeds peak rate 941 and this is unacceptable. In practice, with high video from fast-changing video games, peaks above peak rate (941) occur less than 2% of frames, peaks above 2X peak rate (942) rarely occur, and peak at 3x peak rate (943) Peaks exceeding . However, when they occur (eg, during scene transitions), the data rate required by them is necessary to produce good quality video images.
One way to solve this problem is to install a video compressor 404 whose maximum rate output is the peak rate 941 . Unfortunately, the resulting video output quality during peak frames is poor because the compression algorithm is "poor" for bits. As a result of the appearance of compression artifacts when there are sudden transitions or fast movements, then, the user notices artifacts that always arise suddenly when there are sudden changes or fast motions, and they can become quite annoying.
Although the human visual system is quite sensitive to visual artifacts that appear during sudden changes or rapid movements, it is not very sensitive to detecting a decrease in frame rate in such situations. In fact, when such a sudden change occurs, it happens ahead of time as the human visual system tracks the change, and it does not notice if the frame rate clearly drops from 60 fps to 30 fps and then instantly returns to 60 fps. And, in the case of a very dramatic transition, such as a sudden scene change, the human visual system doesn't notice if the frame rate drops to 20 fps or 15 fps and then instantly returns to 60 fps. While frame rate changes occur only infrequently, to a human observer, it appears that the video is running continuously at 60 fps.
This feature of the human visual system is exploited by the technique shown in FIG. 9B . Server 402 (in FIGS. 4A and 4B ) outputs an uncompressed video output stream at a steady frame rate (60 fps in one embodiment). The timeline is 1/60th of each frame (961 - 970)<sup>th</sup> output per second. Starting with frame 961, each uncompressed video frame is a low-latency video compressor that compresses a frame in less than one frame time to produce a compressed frame 1 981 for the first frame. (404) is output. The data produced for compressed frame 1 981 may be larger or smaller depending on many factors, as described above. If the data is small enough, one frame time (1/60th<sup>th</sup>seconds) or to the client 415 at a time less than the peak transmission rate 941, followed by a transmission time (Xmit time) 991 (the length of the arrow indicates the duration of the transmission time). At the next frame time, server 402 produces uncompressed frame 2 962, which is compressed into compressed frame 2 982, which is sent to client 415 during transmission time 992, which It is less than the frame time of the peak data rate (941).
Next, at the next frame time, the server 402 generates an uncompressed frame 3 963 . When it is compressed by the video compressor 404, the resulting compressed frame 3 983 is more data than can be transmitted at peak rate 941 in one frame time. So, it is transmitted during transmission time (2X peak) 993, which occupies part of all frame time and next frame time. Now, during the next frame time, the server 402 generates another uncompressed frame 4 964 and outputs it to the video compressor 404 but the data is ignored, illustrated as 974 . This is because video compressor 404 is configured to ignore appended uncompressed video frames that arrive while still transmitting previously compressed frames. Of course the video expander of the client 415 fails to receive frame 4, but it simply continues to display frame 3 on the display device 422 for 2 frame times (eg, the frame rate goes from 60 fps to 30 fps briefly). decreased).
For the next frame 5, the server 402 outputs an uncompressed frame 5 (965) that is compressed into a compressed frame 5 (985) and transmitted within one frame during a transmission time 995. The video decompressor of the client 415 decompresses frame 5 and displays it on the display device 422 . Next, server 402 outputs uncompressed frame 6 966 compressed by video compressor 404 into compressed frame 6 986, but at this time the resulting data is very large. A compressed frame will be transmitted during a transmission time (4X peak) 996 at a peak data rate 941, but it takes almost 4 frame times to transmit the frame. During the next 3 frame times, the video compressor 404 ignores 3 frames from the server 402 and the expander at the client 415 keeps frame 6 on the display device 422 for 4 frames of time (eg, , obviously reducing the frame rate from 60fps to 15fps). Then finally, the server 402 outputs frame 10 970 , the video compressor 404 compresses it into a compressed frame 10 987 , which is transmitted during transmission time 997 , and the client 415 ) expands frame 10 and displays it to the display device 422 and once again the video starts again at 60 fps.
Although the video compressor 404 drops video frames from the video stream generated by the server 402 , no matter what the incoming audio format is, it does not drop audio data, the video frames drop and pass it back to the client 415 . As transmitted, it continues to compress the audio data, which continues to decompress the audio data and provides the audio to any device used by the user to reproduce the audio. Thus, the audio continues uninterrupted even during the period when the frame is dropped. Compared to compressed video, compressed audio usually consumes a small percentage of the bandwidth and consequently does not have a major impact on the overall data rate. Although it is not shown in any rate diagram, there is always a limited rate capability for compressed audio streams within peak rate 941.
The example shown in FIG. 9B is chosen to show how the frame rate drops during data rate peaks, but as previously described, when the cyclic I/P tile technique is used, such as data rate peaks, video games, movies, and some applications Continuously dropping frames are rare, even in highly complex scenes/high motion sequences such as those occurring in software. As a result, reduced frame rates are infrequent and obviously the human visual system does not detect them.
If the frame rate reduction mechanism is described to be applied to the video stream rate described in Fig. 9A, the resulting video stream rate is shown in Fig. 9C. In this example, the 2x peak 952 is reduced to a flattened 2x peak 953 , and the 4x peak 955 is reduced to a flattened 4x peak 955 , so that the total video stream data 934 is reduced to a peak data rate 941 . ) or below it.
Thus, using the techniques described above, high motion video streams can be transmitted over common Internet and consumer-grade Internet connections with low-latency. Moreover, in office environments, LANs (e.g. 100Mbs Ethernet or 802.11g wireless) or private networks (e.g., 100Mbps connections between data centers and offices), high motion video streams can be transmitted without peaks, so that a large number of users (e.g., Transmitting 1920×1080 at 60fps at 4.5Mbps) can use shared private data connections without overwhelming networks or network switch backplates, overlapping peaks.
<b>Adjust the baud rate (</b><b>Data</b><b></b><b>Rate</b><b></b><b>Adjustment</b><b>)</b>
In one embodiment, the hosting service 210 evaluates the maximum available bitrate 622 and the latency of the channel to determine an appropriate bitrate for the initial video stream and dynamically adjusts the bitrate in response. To adjust the bit rate, the hosting service 210 will modify, for example, the image resolution and/or the number of frames/second of the video stream sent to the client 415 . Additionally, the hosting service may adjust the quality level of the compressed video. Changing the resolution of a video stream, for example, from 1280x720 resolution to 640x360, the video stretching logic 412 on the client 415 can scale the image to a certain percentage to keep the same image size on the display screen. have.
In one embodiment, in some situations where the channel is completely missed, the hosting service 210 stops the game. In the case of multiplayer games, the hosting service reports to the other user whether the user has left the game and/or has stopped playing for the other user.
<b>Dropped or delayed packets (</b><b>Dropped</b><b></b><b>or</b><b></b><b>Delayed</b><b></b><b>Packets</b><b>)</b>
In one embodiment, if packets received in the sequence of instructions arrived too late for decompression due to packet loss between the video compressor 404 and the client 415 in FIG. 4A or 4B or due to latency requirements of uncompressed frames, If data is lost due to , the video stretching logic 412 can mitigate the visual artifact. In streaming I/P frame execution, if there are lost/delayed packets, the entire screen is affected, potentially causing a complete screen freeze for a period of time or other screen-wide visual artifacts. For example, if a lost/delayed packet causes the loss of an I frame, then the expander will run out of criteria for all P frames that follow until a new I frame is received. If a P frame is lost, then it will affect the P frame for the entire screen that follows it. Depending on how long it will take before the I frame appears, this will have a longer or shorter visual impact. With interleaved I/P tiles as shown in Figures 7a and 7b, lost/delayed packets have much less impact on the entire screen, since it will only affect the tiles included in the affected packet. If the data of each tile is sent in a separate packet, then if the packet is lost, it will affect only one tile. Of course, the duration of the visual artifact will depend on whether an I tile packet is lost, and if a P tile is lost, how many frames it will take for an I tile to appear. However, given that different tiles on the screen are updated very frequently in I frames (potentially every frame), even when one tile on the screen is affected, it may not be the other tile. Moreover, if an event immediately causes the loss of a few packets (e.g. a spike on power next to a DSL line obviously disrupting data flow), then some tiles will have more impact than others. , but because some tiles will be quickly updated with new I tiles, they will only have a simple effect. Also, with streaming I/P frame execution, the I frame is the most critical frame, but the I frame is very large, so if there is an event that results in dropped/delayed packets there, perhaps the I frame rather than the smaller I tile will be affected. (eg, if any part of an I frame is lost, it cannot be stretched at all). For all these reasons, using I/P tiles results in much less visual artifacts than when packets are dropped/delayed into I/P frames.
One embodiment attempts to reduce the effects of lost packets by intelligently packaging the compressed tiles in a transmission control protocol (TCP) packet or a user datagram protocol (UDP) packet. For example, in one embodiment, tiles are aligned with packet boundaries whenever possible. Figure 10a shows how a tile would be wrapped in a series of packets 1001 - 1005 without implementing this feature. In particular, in Fig. 10A, tiles cross packet boundaries and are inefficiently packed so that a loss of a single packet results in a loss of multiple frames. For example, if packet 1003 or 1004 is lost, 3 tiles are lost, resulting in visual artifacts.
In contrast, FIG. 10B illustrates tile packing logic 1010 for intelligently including tiles within a packet to reduce the effects of packet loss. First, the tile packing logic 1010 aligns tiles with packet boundaries. Accordingly, tiles T1, T3, T4, T7 and T2 are generally aligned at the boundary of packets 1001 - 1005. Also, the tile packing logic does not cross packet boundaries and attempts to pin tiles within the packet in the most efficient possible way. Based on the size of each tile, tiles T1 and T6 are combined into one packet 1001; Tiles T3 and T5 are combined into one packet 1002; Tiles T4 and T8 are combined into one packet 1003; tile T8 is added to packet 1004; And tile T2 is added to packet 1005 . Thus, under this structure, a single packet loss will not result in a loss of more than 2 tiles (rather than 3 tiles as shown in Fig. 10a).
An additional advantage to the embodiment shown in Figure 10b is that tiles are transmitted in a different order than they are displayed in the image. This way, if an adjacent packet is transmitted and lost in the same event interfering, it will affect areas of the screen that are not close to each other, creating barely noticeable artifacts in the display.
One embodiment employs a forward error correction (FEC) technique to protect certain portions of the video stream from channel errors. As is known in the art, FEC techniques such as Reed-Solomon and Viterbi generate and append error correction data information for data transmitted over a communication channel. If an error occurs in the underlying data (eg an I frame), then the FEC will correct the error.
FEC codes allow to increase the rate of transmission; So ideally, they are used only where they are most needed. If the data is to be sent so that it does not cause very perceptible visual artifacts, it would be desirable not to use FEC codes to protect the data. For example, a P tile is 1/60 of the screen if it is lost.<sup>th</sup>It immediately precedes the I tile, which creates only visual artifacts (eg, will not update in the tile on the screen) during the seconds of . Such visual artifacts are rarely perceived by the human eye. As the P tile is further behind the I tile, the loss of the P tile becomes progressively more significant. For example, if the tile cycle pattern is followed by an I tile up to 15 P tile before the I tile becomes available again, then immediately if the P tile following the I tile is lost, it will last for 15 frame times (at 60fps it is 250ms will result in tiles showing the wrong image. The human eye will be able to easily detect a disruption in the stream for 250ms. So, the further back the P tile is from the new I tile (eg, the closer the P tile follows the I tile), the more recognizable the artifact. As discussed above, in general, the closer a P tile follows an I tile, the smaller the data for the P tile. Thus, the P tiles following the I tiles are not more critical to protect them from being lost, and they become smaller in size. And in general, the smaller the data that needs protection, the smaller the FEC code is required to protect it.
So, because of the importance of the I tile in the video stream of one embodiment, only the I tile is provided with the FEC code, as shown in Fig. 11a. Accordingly, FEC 1101 includes error correction codes for I tile 1100 , and FEC 1104 includes error correction codes for I tile 1103 . In this embodiment, no FEC is created for the P tile.
In one embodiment shown in FIG. 11B , also an FEC code is generated for the P tile that is prone to causing artifacts of vision if lost. In this embodiment, FEC 1105 provides an error correction code for the first 3 tiles, but not for the P tiles that follow. In another embodiment, the FEC code is generated for the P tile with the smallest data size (which tends to self-select the P tile that occurs most immediately after the I tile, and is the most critical for protection) ).
In another embodiment, rather than sending some tile and FEC code, send the tile twice at separate times in different packets. If one packet is lost/delayed, another packet is used.
11C, FEC codes 1111 and 1113 generate audio packets 1110 and 1112 that are transmitted from the hosting service substantially concurrently with the video. It is particularly important to maintain the integrity of the audio in the video stream because distorted audio (eg, clicking or hissing) will result in a particularly unpleasant user experience. The FEC code helps ensure that the audio content is rendered at the client computer 415 without distortion.
In another embodiment, rather than sending the audio data and the FEC code, the audio data is sent twice each time in a different packet. If one packet is lost/delayed, another packet is used.
Additionally, in one embodiment shown in FIG. 11D , FEC codes 1121 and 1123 are typically (eg, button presses) user input instructions 1120 and 1120 sent upstream from client 415 to hosting service 210 . 1122) is used. This is important because in a video game or application, missing button presses or mouse movements can result in an unpleasant experience for the user.
In another embodiment, rather than transmitting the user input command data and the FEC code, the user input command data is transmitted twice, each time in a different packet. If one packet is lost/delayed, another packet is used.
In one embodiment, the hosting service 210 evaluates the quality of the communication channel with the client 415 to determine whether to use the FEC code, and if using the FEC code, where the FEC is applied is video, audio, and user Evaluates which part of the command. Estimating the "quality" of a channel will include functions such as measuring packet loss, latency, etc. as described above. If the channel is not particularly reliable, then the hosting service 210 will apply FEC to all I tiles, P tiles, audio and user commands. In contrast, if the channel is reliable, then the hosting service 210 will only apply FEC to audio and user commands, no FEC to audio or video, or no FEC at all. Various other application modifications of FEC can still be adopted while following this fundamental theory. In one embodiment, the hosting service 210 continuously monitors the status of the channel and changes in the FEC policy corresponding thereto.
In another embodiment, referring to Figures 4a and 4b, when any packet is lost/delayed, the FEC cannot correct the tile data loss, resulting in loss of tile data or if perhaps due to certain bad packet loss, the client 415 ) evaluates how many frames have left before a new I tile is received and compares the round trip latency from the client 415 to the hosting service 210 . If the round-trip latency is less than a number of frames before a new I tile is due to arrive, then the client 415 sends a message to the hosting service 210 to request a new I tile. These messages are routed to the video compressor 404 and generate I tiles, rather than P tiles, while the tile's data is lost. The system shown in Figures 4a and 4b is typically designed to provide a round-trip latency of less than 80 ms, which is 80 ms (at 60 fps a frame is of duration 16.67 ms, so the 80 ms latency at full frame time is corrected to within 83.33 ms). will result in a tile, which is 5 frame time - perceptible confusion, but much less perceptible than 250 ms confusion for example 15 frames). Compressor 404 generates I tiles outside of its normally cyclic order, and if the I tiles result in a frame's bandwidth exceeding the available bandwidth, the compressor 404 will delay the cycles of other tiles, so Other tiles will receive P tiles during the frame time (even if one tile normally receives I tiles during that frame), and starting with the next frame will normally continue cycling, and will normally receive I tiles in the preceding frame. The tile to be played will receive the I tile. Although this operation simply delays the state of R frame cycling, it will not be perceived as normal time.
<b>Video and Audio Compressor/</b><b>stretcher</b><b> execution(</b><b>Video</b><b></b><b>and</b><b></b><b>Audio</b><b> Compressor/Decompressor </b><b>Implementation</b><b>)</b>
12 illustrates a specific embodiment in which a multi-core and/or multi-processor 1200 is used to compress 8 tiles in parallel. In one embodiment, a dual processor, quad core Xeon CPU computer system is used, operating at 2.66 GHz or higher, with each core running the open source x264 H.264 compressor as an independent processor. However, a variety of other hardware/software configurations may be used while conforming to this fundamental theory. For example, each CPU core can be replaced by an H.264 compressor running on an FPGA. In the example shown in FIG. 12 , cores 1201 to 1208 are used to process I tiles and P tiles simultaneously as 8 independent threads. As is well known in the art, current multi-core and multi-processor computer systems are inherently capable of multi-threading when integrated into multi-threading operating systems such as Microsoft Windows XP Professional Edition (64-bit or 32-bit editions) and Linux. can
In the embodiment shown in Figure 12, each of the eight cores is responsible for only one tile, which generally operates independently of the other cores, each running individual instantiation of x264. PCI Express x1 based on DVI capture card such as Sendero Video Imaging IP Development Board from Microtronix of Oosterhout, Netherlands used to capture uncompressed video at 640×480, 800×600, or 1280×720 resolution, FPGA on card uses Direct Memory Access (DMA) to transfer captured video over the DVI bus to system RAM. The tiles are arranged in a 4x2 array 1205 (although they are shown as square tiles, in this embodiment they are 160x240 resolution). Each instance of x264 is configured to compress one of 8 160×240 tiles, and to compress an I tile followed by the 7 P tiles shown in Fig. 12, after the initial I tile compression, each core enters one cycle, and each frame is synchronized with the others after out of state.
At each frame time, using the techniques described above, the resulting compressed tiles are combined into a packet stream and then the compressed tiles are sent to the destination client 415 .
Although not shown in Fig. 12, if the data rate of the combined 8 tiles exceeds a certain peak rate 941, then all 8 x264 processes continue with as many frames as necessary until data for the combined 8 tiles is transmitted. Pause for time.
In one embodiment, the client 415 runs as software on a PC running 8 instances of FFmpeg. The receiving process receives 8 tiles, each tile being sent to an FFmpeg instance, which stretches the tile and renders it for proper tile position on the display device 422 .
The client 415 receives the input of the keyboard, mouse, and game controller from the input device driver of the PC and transmits it to the server 402 . The server 402 then applies the received input device data and applies it to a PC running Windows using an Intel 2.16GHz Core Duo CPU, a game or application running on the server 402 . The server 402 then generates a new frame and outputs it through the DVI output, and outputs it through the DVI output of the motherboard-based graphics system, or the NVIDIA 8800GTX PCI card.
At the same time, the server 402 outputs audio produced by the game or application via digital audio output (eg S/PDIF), which is combined with the digital audio input of a dual quad-core Xeon-based PC running video compression. do. The Vorbis open source audio compressor compresses video concurrently using any core available to the process thread. In one embodiment, the core performs audio compression first to finish compressing its tiles. The compressed audio is then sent along with the compressed video and is decompressed at the client 415 using the Vorbis audio decompressor.
<b>Hosting Service Server Center Deployment (</b><b>Hosting</b><b></b><b>Service</b><b></b><b>Server</b><b></b><b>Center</b><b></b><b>Distribution</b><b>)</b>
Light through glass, such as an optical fiber, travels close to the speed of light in vacuum so the exact propagation speed for light in the optical fiber can be determined. However, in practice, allowing time for routing delays, transmission inefficiencies, and other things, we observe that optimal latency on the Internet reflects transmission rates close to 50% of the speed of light. So, the optimal 1000 mile round trip latency is around 22 ms, and the optimal 3000 mile round trip latency is about 64 ms, so a single server on the US coast is too far off the other coast (which can be much more than 3000 miles away). You won't be able to provide it with the latency required for the client. However, as shown in FIG. 13A , if the hosting service 210 server center 1300 is located in the center of the United States (eg, Kansas, Nebraska, etc.), the distance to any location on the continental United States may be 1500 miles or less. and round-trip Internet latency is less than 32ms. Referring to Figure 4b, typically we observe latencies close to 10-15 ms with DSL and cable modem systems, although the worst latency allowed for ISP 453 users is 25 ms. 4B also assumes that the maximum distance from the user premises 211 to the hosting center 210 is 1000 miles. Thus, a typical ISP user would have a round-trip latency of 15 ms and a maximum internet distance of 1500 miles would have a round-trip latency of 32 ms, so the total round-trip latency at one point where the user actuates the input device 421 and sees the response on the display device 422 is 1+1+15+32+1+16+6+8 = 80ms. So an 80ms response time can typically be achieved over an Internet distance of 1500 miles. This allows any user premises in the continental United States to access a centrally located single server center with short enough user ISP latency.
In another embodiment shown in FIG. 13B , Hosting Services 210 Server Centers HS1 - HS6 are strategically located throughout the United States (or any large Hosting Services Server Center located in close proximity to a very popular center). other geopolitical regions). In one embodiment, the server centers HS1-HS6 exchange information via the Internet or a private network or a combined network 1301 thereof. With multiple server centers, the service provides low-latency to users with high user ISP latency 453 .
Although distance on the Internet is certainly a contributing factor to round-trip latency across the Internet, sometimes other factors initiate activity that is largely unrelated to latency. Sometimes the packet stream is routed over the Internet to a remote location and back again, resulting in long loop latency. Sometimes there is routing equipment on a path that does not work properly, which causes delays in transmission. Sometimes there are overloaded traffic paths, which introduces delays. Also, sometimes failures are caused by preventing the user's ISP from sending to a given destination. Thus, the general Internet usually has a fairly reliable and optimal route from one point to another and has latency determined by the distance (especially long distance links that result in routing outside of the user's local area). Connectivity is provided, whereby reliability and latency are never guaranteed and often from the user's premises to a given destination cannot be achieved on the ordinary Internet.
In one embodiment, a user client 415 is initially connected with a hosting service 210 to play a video game and use an application, and the client is initially connected to each of the hosting service server centers HS1 - HS6 available at startup. communicate with (eg, using the techniques discussed above). If the latency is low enough for a particular connection, that connection is used. In one embodiment, the client communicates with all or a subset, and the lowest latency connection among the hosting service server centers is selected. The client will select the service center with the lowest-latency connection and the service center will be identified as one with the lowest-latency connection and provides this information (eg, in the form of an Internet address) to the client.
If a particular hosting service server center is overloaded and/or the user's game or application can tolerate the latency to the lightly loaded hosting service server center, the client 415 is directed to another hosting service server center. In such a situation, the game or application the user is running on will be suspended on the server 402 at the user's overloaded server center, and the game or application status information is transmitted to the server 402 at the server center of another hosting service. will be The game or application will be resumed. In one embodiment, the hosting service 210 may naturally require a stopping point (eg, between levels in a game, or after a user initiates a "save" command in the application) to be reached in order for the game or application to transfer. will wait until In already other embodiments, the hosting service 210 will wait for user activity to cease for a specified period of time (eg, one minute) and then begin transmitting at that time.
As noted above, in one embodiment, hosting service 201 subscribes to Internet bypass service 440 of FIG. 14 in order to provide guaranteed latency to its clients. As used herein, an Internet bypass service is a service that provides a private network path from one point of the Internet to another with guaranteed characteristics (eg, latency, transfer rate, etc.). For example, if the hosting service 210 receives a large amount of traffic from users using AT&T's DSL service provided in San Francisco, rather than routing to AT&T's San Francisco-based central office, the hosting service ( 210 may lease a high-performance personal data connection from a service provider (perhaps AT&T itself or another provider) between a San Francisco-based central office and one or more server centers for hosting service 210 . Then, if all the hosting service server centers (HS1 - HS6) are routed over the general internet to users in San Francisco using AT&T DSL, it would result in too high latency, so a private data connection could be used instead. . Although private data connections are generally more expensive than routes over the general Internet, the overall cost impact will be low, and users will experience a more consistent service as the connections between users and hosting services 210 remain for a long time in such a small proportion. .
Server centers often have two layers of backup power in the event of a power outage. Typically the first tier backs up power from the battery (or alternatively a readily available energy source, such as a flywheel connected to a generator and continues to operate), which provides instantaneous power when mains power is lost and server center operation. keep If the power outage is simple, mains power returns quickly (eg, within 1 minute), then battery is required to keep the server center operating. However, if a power outage occurs for an extended period of time, typically a generator (eg diesel-powered) can run for as long as the fuel it has and is delivered for the battery. Such generators are very expensive because they are capable of producing as much power as a server center would normally get from mains power.
In one embodiment, each of the hosting services HS1 - HS5 shares user data with the other, so if one server center goes out, it can stop running games and applications, and from each server 402 to the other server center. will send game or application state data to the servers 402 of If such usage occurs occasionally, it will be permitted to transfer the user to a hosting service server center where the user cannot provide optimal latency (e.g., the user simply has to endure high latency for the duration of a power outage), which transmits It will allow many wide range choices for users. For example, differences in time zones across the United States suggest that users on the East Coast will go to bed at 11:30 PM, while users on the West Coast will start peaking in video game usage at 8:30 PM. . If at the same time the West Coast hosting service server centers are out of power, there will not be enough West Coast servers 402 in the other hosting service server centers to handle all users. In such a situation, some users may be transferred to a hosting service server center on the East Coast that has a server 402 available, and the result for the user will be only high latency. If a user is transferred from a server center that has lost power, the server center can be started with an interrupted turn of its servers and appliances, and all appliances will be shut down before the battery (or other immediate power backup) runs out. In this way, generator costs for the server center can be avoided.
In one embodiment, during heavy loading times of the hosting service 210 (due to user loading peaks or one or more server centers are down), users can access other servers based on the latency requirements of the games or applications they use. sent to the center. So users using games or applications that require low-latency will benefit from low-latency server connections available when there is a limited supply.
<b>Hosting service features (</b><b>Hosting</b><b></b><b>Service</b><b></b><b>Features</b><b>)</b>
15 shows an embodiment of the components of a server center for hosting service 210 that is used to illustrate the following features. Like the hosting service 210 shown in FIG. 2A , the components of this server center are controlled and coordinated by the hosting service 201 control system 401 , unless otherwise limited.
Inbound Internet traffic 1501 from user client 415 goes to inbound routing 1502 . Typically, inbound Internet traffic 1501 will enter the server center over a high-speed fiber optic connection for the Internet, but adequate bandwidth, reliability, and low-latency network connection means will suffice. Inbound routing 1502 is a system switch of a network (the network may be implemented over an Ethernet network, a Fiber Channel network, or some other means of transport) and an appropriate application/game ("app/game") server A routing server that routes each packet to (1521 - 1525) and supports the switch to take the packets it arrives. In one embodiment, certain packets forwarded to a particular application/game server represent a subset of data received from the client and/or by other components (eg, networking components such as gateways and routers) within the data center. will be converted/changed. In some cases, for example, if a game or application is running on multiple servers at once in parallel, the packets will be routed to one or more of the servers 1521 - 1525 at a time. The RAID arrays 1511 - 1512 are connected to the inbound routing network 1502 , and the application/game servers 1521 - 1525 can read or write the RAID arrays 1511 - 1512 . Furthermore, RAID array 1515 (which will be implemented as a multi-RAID array) is also coupled with inbound routing 1502 and data from RADI array 1515 can be read from application/game servers 1521 - 1525. Inbound routing 1502 includes a switch in a tree structure with inbound Internet traffic 1501 at the root; With a mesh structure that interconnects all the various devices; or as interconnected with a set of subnets, with traffic concentrated among devices that exchange and separate from traffic concentrated among other devices; It may be implemented in a wide range of prior art network architectures. One form of network configuration is a SAN, and although typically used for storage devices, it can also be used for general high-speed data conversion between devices. Also, application/game servers 1521 - 1525 will each have inbound routing 1502 and multiple network connections. For example, servers 1521-1525 will have subnets and network connections attached to RAID arrays 1511-1512 and subnets and other network connections attached to other devices.
As already described with respect to server 402 in the embodiment shown in Fig. 4A, application/game servers 1521-1525 will all be configured the same, some differently, all differently. In one embodiment, when using a hosting service, each user typically uses at least one application/game server 1521 - 1525 . For simplicity of explanation, it is assumed that a given user uses the application/game server 1521, but multiple servers may be used by one user, and multiple users may share a single application/game server 1521 - 1525. can The user's control input, sent from client 415 as described above, is received as inbound internet traffic 1501 and is routed via inbound routing 1502 to application/game server 1521 . The application/game server 1521 uses the user's control input as a control input of a game or application operated in the server, and calculates video and audio of the next frame related thereto. The application/game server 1521 then outputs the uncompressed video/audio 1529 to the shared video compression 1530 . The application/game server will output uncompressed video through some means, including one or more Gigabit Ethernet connections, but in some embodiments the video is output over a DVI connection and audio and other compressed and communication channel state information is It is output through a universal serial bus (USB) connection.
Shared video compression 1530 compresses uncompressed video and audio from application/game servers 1521 - 1525. Compression will probably be performed entirely as hardware or as hardware running software. There will be a dedicated compressor for each application/game server 1521-1525, but if the compressor is fast enough, a given compressor can be used to compress video/audio from more than one application/game server 1521-1525. . For example, the video frame time at 60 fps is 16.67 ms. If the compressor can compress one frame in less than 1ms, the compressor will get input from one server from as many as 16 application/game servers 1521-1525 and then from the other, so that each video/ It can be used to compress video/audio, with a compressor that stores the state of the audio compression process and switches context as cycled between video/audio streams from the server. This results in substantial cost savings in the compression hardware. In one embodiment, since different servers complete frames at different times, the compressor resource is a shared pool 1530 with shared storage means (eg RAM, flash) for storing the state of each compressor process. ), when the server 1521 - 1525 frames are complete and ready to be compressed, the control means determines whether compression resources are available at that time, and the server's compression of uncompressed frames of video/audio for compression. Provides a compressed resource with the status of processing.
Part of the state of the compression processing of each server can be used as a reference for the P tile, including the stretched frame buffer data of the previous frame, the resolution of the video output; compression quality; tiling structure; allocation of bits per tile; Compression quality, audio format (e.g. stereo, surround sound, Dolby<sup>&#174;</sup> , AC-3), note that it contains information about the compression itself. However, the state of the compression process also includes communication channel state information regarding the peak data rate 941, whether a prior frame (as shown in Fig. 9b) is currently being output (and as a result the current frame may be ignored). , and whether there are channel characteristics to consider for compression, such as potentially excessive packet loss, influence the decision for compression (eg, with respect to the frequency of the I-tile, etc.). As the peak rate 941 or other channel characteristics change over time, as determined by the application/game server 1521-1525 supporting each user monitoring data sent from the client 415, the application/game server ( 1521-1525 send related information to shared hardware compression 1530 .
In addition, shared hardware compression 1530 uses means such as those described above, applies FEC codes if appropriate, copies certain data, or expands to a high quality and with practicable stability and quality to the client 415 and to the client 415 . In order to adequately guarantee the performance of the video/audio data stream received by the controller, the compressed video/audio is packetized through other steps.
Some applications, such as those described below, require video/audio output from a given application/game server 1521 - 1525 that can be used at multiple resolutions (or other multiple formats) at the same time. If the application/game servers 1521 - 1525 may be simultaneously compressed in different formats, different resolutions, and/or different packet/error correction schemes. In some cases, certain compression resources may be shared among multiple compression processes that compress the same video/audio (eg, in many compression algorithms, there is a step in which the image is scaled to multiple sizes before compression is applied by something. (If different size images are required to be output, this step can be used to provide several compression processing at a time). In other cases, separate compression resources will be required for each format. In some cases, all the various resolutions and formats of the compressed video/audio 1539 requested for a given application/game server 1521 - 1525 will be output to the outbound routing 1540 at once. The output of compressed video/audio 1539 in one embodiment is in UDP format, and thus a unidirectional stream of packets.
The outbound routing network 1540 directs each compressed video/audio stream through an outbound Internet traffic 1599 interface (which will typically connect to the Internet and a fiber interface) to the intended user or other destination; and a set of routing servers and switches that return to delay buffer 1515, and/or return to inbound routing 1502, and/or output for video distribution over a private network (not shown) Outbound routing 1540 (as described below) will output a given video/audio stream to multiple destinations at once. In one embodiment this is performed using Internet Protocol (IP) multitasking and a given UDP stream intended to be streamed to multiple destinations at a time is broadcast, and the broadcast is performed by a routing server in outbound routing 1540 and It is repeated by the switch. A plurality of destinations of the broadcast will be to clients 415 of a plurality of users via the Internet, and/or a plurality of application/game servers 1521 - 1525 via inbound routing 1502 and/or one or more delay buffers 1515. will become Thus, the outputs of a given server 1521 - 1522 are compressed into single or multiple formats, and each compressed stream is directed to a single or multiple destinations.
Moreover, in another embodiment, if multiple application/game servers 1521 - 1525 are used concurrently by one user (eg, in a parallel processing configuration to generate 3D output of a complex scene), each server will create part of the image, the video output of multiple servers 1521 - 1525 can be combined by shared hardware compression 1530 into a combined frame, as if from a single application/game server 1521 - 1525 As such, it is treated as described above in the forward point.
In one implementation, all copies of video generated by application/game servers 1521 - 1525 are written to delay buffer 1515 for at least a few minutes (15 minutes in one embodiment). This allows each user to "rewind" the video (in the case of games) from each session to review prior actions or uses. Thus, in one embodiment, each compressed video/audio output 1539 stream routed to the user client 415 is also multicast to the delay buffer 1515 . When the video/audio is stored in the delay buffer 1515, the directory of the delay buffer 1515 provides the location of the delay buffer 1515 where the delayed video/audio can be found and the source of the delayed video/audio to the application/game server ( 1521 - 1525) as a cross-reference between the network addresses.
<b>live, instantly-visible, instantly-</b><b>to play</b><b> game that can</b><b>Live</b><b>, </b><b>Instantly</b><b>-Viewable, Instantly-playable </b><b>Games</b><b>)</b>
Application/game servers 1521 - 1525 are not only used to run a given application or video game for a user, but they also provide a hosting service 210 that supports navigation through a hosting service 210 and other features. It is used to create user interface applications for A screen shot of this user interface application, "game finder" is shown in FIG. 16 . This particular user interface screen allows the user to watch 15 games being played live (or delayed) by other users. Each "thumbnail" video window, such as 1600, is a live video window in motion showing video from one user's game. The field of view shown in the thumbnail is the same as the field of view the user is seeing, but it will be a delayed field of view (eg, if the user is playing a combat game, the user does not want other users to know where they hid until a certain amount of time, 10 You'll choose to delay some view of gameplay for a minute or so). Also, the field of view will be the field of view of the camera, different from the field of view of any user. Through menu selection (these illustrations are not shown), the user will simultaneously select a game for viewing based on various criteria. With a small sampling of typical choices, the user can play one kind of all games (all played by different players), only the top-tier players, players at a given level in the game, or lower-tier players (e.g., if players will learn the basics), mate ("buddies") (or competitor) players, games with a majority of viewers, etc. (as shown in FIG. 16 ) randomly.
In general, each user will decide whether his or her game or application can be watched by others, and if so, only the deferred ones.
The application/game servers 1521 - 1525 generating the user interface screens shown in FIG. 16 send messages to the application/game servers 1521 - 1525 while each user is requesting whose game it is 15 video/game. Requires audio feeds. The message is sent through the inbound routing 1502 or other network. The message will include the size and format of the video/audio requested, and will identify the user viewing the user interface screen. A given user will choose to select a "privacy" mode and no other user will be allowed to watch the video/audio of his game (either in his view or in any other view), or as described in the preceding paragraph, The user will choose to allow viewing video/audio from her game, but the video/audio viewing is delayed. When the user application/game server 1521 - 1525 receives and accepts a request to allow video/audio to be watched, it will grant the requesting server, it will also accept the requested format or screen size (format and screen size). It will inform the shared hardware compression 1530 of the need to generate the compressed video stream in addition to the one already created (assuming the size is different from the one already created), and also indicates the destination for the compressed video (eg, the requesting server). . If the requested video/audio is only delayed, the requesting application/game server 1521 - 1525 will be notified, and it will be notified of the location of the video/audio in the directory in the delay buffer 1515, the location of the delayed video/audio It will request the delayed video/audio from the delay buffer 1515 by finding the network address of the source application/game server 1521 - 1525 . Once all these requests are generated and handled, up to 15 thumbnail-sized video streams are routed for the application/game servers 1521 - 1525 generating user interface streams from outbound routing 1540 to inbound routing 1502. , will be expanded and displayed by the server. If the delayed video/audio stream is at a screen size that is too large, the application/game server 1521 - 1525 will expand the stream and reduce the size of the video stream to the thumbnail size. In one embodiment, requests for audio/video are directed to a central "management" similar to the hosting service control system of FIG. 4A (not shown in FIG. 15) that redirects the requests to the appropriate application/game servers 1521 - 1525. )" is sent (and managed) to the service. Moreover, in one embodiment, no request is required because the thumbnails are "pushed" to the client of users that allow it.
Audio mixed from all 15 games at the same time will create a dissonant sound. Either the user will select all the sounds together in this way (perhaps getting a sense of the "din" generated by every action shown), or the user will choose to listen to audio from one game at a time. Selection of a single game may be accomplished by moving the yellow selection box 1601 for a given game (yellow box movement is accomplished by using the arrow keys on a keyboard, by moving the mouse, by moving the joystick, or on another device such as a mobile phone). can be accomplished by pressing the direction button). Once a single game is selected, only audio is obtained from that game play. Also shown is game information 1602 . For example, in the case of this game, the publisher logo ("EA"), the game logo, "Need for Speed Carbon" and the orange vertical bar indicate items related to the number of people watching or playing the game at a particular moment. Moreover, there are 145 players actively playing 80 different instances of Need for Speed Game (eg it can be played by individual player games or multiplayer games), and there are 680 observers (one of the users). "stats" are provided to indicate These statistics (and other statistics) are collected by the hosting service control system 401 and collected by the hosting service control system 401 to maintain a log of 210 hosting service operations, moderately paying users, and paying publishers for providing content. 1512) is stored. Some statistics are recorded due to actions by the service control system 401 , and some are reported to the service control system 401 by the personal application/game servers 1521 - 1525 . For example, the application/game servers 1521 - 1525 running this game finder application send a message to the hosting service control system 401 when a game is being viewed (and when they start watching), so that how many games are being viewed. Will update if it's in progress. Some statistics can be used in user interface applications such as the game finder application.
If the user clicks the enable button on their input device, they will see a thumbnail video that zooms up in a yellow box while remaining at full screen size. This effect is illustrated by the process in FIG. 17 . Note that the video window 1700 increases in size. To implement this effect, the application/game servers 1521 - 1525 run the selected game to have a copy of the video stream for the full screen size (at the resolution of the user's display device 422). 1521 - 1525) to request a game routed for it. The application/game servers 1521 - 1525 on which the game is running share that the thumbnail-sized copies of the game are no longer needed (unless other application/game servers 1521 - 1525 require such thumbnails). It notifies the hardware compressor 1530, which in turn causes it to send a full screen size copy of the video to the application/game server 1521 - 1525 which enlarges the video. The user playing the game may or may not have the same resolution display device 422 as the user zooms into the game. Moreover, other observers of the game may or may not have a display device at the same resolution as the user zooms in on the game (and may have other means of audio reproduction, eg stereo or surround sound). Thus, the shared hardware compression 1530 determines whether a suitably compressed video/audio stream has already been created to meet the needs of the user requesting the video/audio stream, and if present, it is determined by the application/game server that augments the video. Inbound routing and an application/game server 1521 - 1525 that informs outbound routing 1540 to route a copy of the stream to 1521 - 1525, and compresses and expands the video to another video copy that is otherwise suitable for the user. Notifies outbound routing to send the stream back to 1502. A server that receives a full screen version of the video now selected will expand it and gradually scale it to full size.
18 shows what the screen will look like after the game has been enlarged to full screen and the game is viewed at the full resolution of the user's display device 422 as indicated by the image indicated by arrow 1800 . The application/game server (1521-1525) running the game finder application sends a message to other application/game servers (1521-1525) providing thumbnails that are no longer needed, and the other games are hosted by which the game is no longer being viewed. A message is also sent to the service control server 401 . At that point, the display will create an overlay 1801 on top of the screen providing only information and menu controls to the user. As the game progresses, the audience grows to 2,503 spectators. With so many spectators, there are many spectators with display devices 422 having the same or nearly the same resolution (each application/game server 1521 - 1525 can scale the video to fit the adjustment).
Since the illustrated game is a multiplayer game, the user will decide at what point to join the game. Hosting service 210 may or may not allow users to participate in games for a variety of reasons. For example, a user will have to pay to play a game and it is not a choice, the user will not have a sufficient ranking to participate in a particular game, or the user's internet connection may not have low-latency enough for the user to play. (e.g. there is no latency constraint for the game you watch, and the game you play far away (actually, on another continent) can be viewed without latency concerns, but in order to play the game, the latency is (a) to enjoy the game. It must be low enough for the user and (b) be in the same position as other players with low-latency connections). If the user is allowed to play, the application/game server 1521 - 1525, which provides the game finder user interface for the user, allows the hosting service control server 401 to load the game from the RAID array 1511 - 1512, It will ask to start (eg, browse and start) an appropriately configured application/game server 1521 - 1525 to play the game. and the hosting service control server 401 is shared to switch from compressing video/audio from the application/game server hosting the game finder application to compressing the video/audio from the application/game server hosting the current game. will instruct hardware compression 1530 . The vertical sync of the game finder application/game service and the new application/game server hosting game is not synchronized, and as a result, it is easy to cause a time difference between the two syncs. Because the shared video compression hardware 1530 will start compressing the video at the application/game server 1521 - 1525 completing the video frame, the first frame from the new server will complete sooner than the full frame time of the original server. , which will first be before the transmission of the compressed frame is complete (e.g. consider transmission time 992 in Figure 9b: if uncompressed frame 3 963 completes half of the frame time, it will be transmitted time 992 will be affected. In such a situation the shared video compression hardware 1530 will discard the first frame from the new server (eg, as frame 4 964 is ignored 974 ), and the client 415 from the old server at the rest of the frame time. It will hold the last frame, and the shared video compression hardware 1530 will start compressing the next frame time video from the new application/game server hosting the game. Apparently, to the user, the transfer from one application/game server to another will appear uniform. The hosting service control server 401 will then inform the application/game game servers 1521 - 1525 hosting the game finder to switch to the idle state, until it is needed again.
The user can then play the game. And, the exception is that the game will be played immediately perceptually (since it will be loaded into the application/game game server 1521-1525 in a RAID array 1511-1512 at gigabit per second rate), the game is an ideal driver , which has a registry configuration (in the case of Windows), and will be loaded into a server that is exactly suitable for the game, with an operating system configured correctly for the game that has no other applications running on the server that compete with the behavior of the game.
Also, as the user progresses through the game, each segment of the game will be loaded from the RAID array 1511 - 1512 to the server at the speed of gigabits per second (eg, 1 gigabyte loads in 8 seconds), and the RAID array 1511 - 1512) with vast storage capacity (which is a shared resource among many users, it is very large in size and still cost effective), geometric setups or other game segment setups can be precomputed and RAID arrays (1511 - 1512) and loaded very quickly. Moreover, since the hardware configuration and computational performance of each application/game server 1521 - 1525 are known, pixel and vertex shaders can be computed in advance.
Thus, the game would start almost immediately, it would work in an ideal environment, and the next segment would load almost immediately.
However, beyond these benefits, the user will be able to see others playing the game (either through the game finder described above, or otherwise), and tips from watching if the game is interesting and, if so, what else. learn And, users will be able to demo the game immediately without waiting for large downloads and/or installs, users will be able to play the game immediately, perhaps for a smaller fee, or longer-term play by rating criteria. will be able And users will be able to play games on their Window PCs, Macintoshes, television sets at home, and even when traveling, or on mobile phones with low-latency wireless connections that are sufficiently low. And, all of this can be accomplished without physically owning a copy of the game.
As previously mentioned, the user may not allow others to view his gameplay, may allow viewing of his game after a delay, may allow viewing of his game by a selected user, or You can make his game viewable by all users. Nevertheless, in one embodiment for 15 minutes in the delay buffer 1515, the video/audio will be stored and the user will be able to "play it back" and view his old game play, pause, slow down Play it back, play it fast forward, and so on, just like he can while watching TV with a digital video recorder (DVR). Although, in this example, the user is playing a game, and if the user is using the application, the same "DVR" capabilities are available. This can help review prior work and can be helpful for other applications as described below. Moreover, if the game is designed to have the ability to be replayed based on the available game state information, the camera field of view can be changed, then this "3D DVR" performance will also be supported, but it does not matter if the game supporting it design will be required. The "DVR" capability using the delay buffer 1515 will work for video generated when the game or application is used, and of course limitedly with any game or application, but in the case of a game with 3D DVR capability it will work for the user can control "fly through" in 3D of pre-played segments, the resulting video is stored by delay buffer 1515, and has the game state of the game segment record. Thus, certain "fly-throughs" will be recorded as compressed video, but since game state will also be recorded, other fly-throughs will be possible at a later date in the same segment of the game.
As described below, users on the hosting service 210 will have their respective user pages, where they can present information and other data about themselves. Among the things the user can inform is the video segments from the gameplay in which they are saved. For example, if a user overcomes a particular difficult challenge in a game, the user can "rewind" to just before the point of their great achievement in the game, and then make it visible to other users on the user's user page. Instructs hosting service 210 to store video segments of some duration (eg, 30 seconds). To do this, the problem with the application/game servers 1521 - 1525 is that the user will use the RAID array 1511 - 1512 to play the video stored in the delay buffer 1515 and then the video segment on the user's user page. put in the index.
If the game has the capability of a 3D DVR, then, as described above, the game state information requested for the 3D DVR can be recorded by the user and made available on the user's user page.
At events where the game is designed to have a "spectator" (eg, the user can travel through the 3D world and observe the action without participating in it), adding an active player, the Game Finder application allows the user to It will allow you to participate in the game not only as a player, but also as an audience. From a performance standpoint, there is no difference in the hosting system 210 if the user is an audience instead of an active player. The game will be loaded into the application/game server 1521 - 1525 and the user will adjust the game (eg, controlling the virtual camera to view the world). The user's gaming experience will be the only difference.
<b>Multi-user collaboration (</b><b>Multiple</b><b></b><b>User</b><b></b><b>Collaboration</b><b>)</b>
Another feature of the hosting service 210 is that multiple users can collaborate while watching live video, even though they use very different devices to watch. This is useful when playing games and when using applications.
Many PCs and mobile phones are equipped with video cameras and are capable of real-time video compression, especially when the images are small. Also, a small camera can be used to attach to a television, and it is not difficult to compress the video in real time using software or many hardware compression devices to compress the video. Also, many PCs and all mobile phones have a microphone, and a headset can be used as a microphone.
These cameras and/or microphones, combined with local video/audio compression capabilities (especially employing the low-latency video compression techniques described herein), allow users to video and/or Audio may be transmitted along with input device control data. When such a technique is employed, the following performance values shown in FIG. 19 can be achieved: a user can have his video and audio 1900 appearing on the screen within another user's game or application. An example of this is a multiplayer game in which teams compete in a car race. A user's video/audio is selectively viewable/audible only to their team. And, since there is effectively no latency there, using the techniques described above would allow users to make conversations or actions between each other in real time without perceptible delay.
Video/audio integration may be achieved by having compressed video and/or audio arriving in inbound internet traffic 1501 from the user's camera/microphone. Inbound routing 1502 sends video and/or audio to application/game game servers 1521 - 1525 where viewing/listing of video and/or audio is allowed. The user of each application/game game server 1521 - 1525 selected to use video and/or audio compresses it and incorporates it as desired for presentation within the game or application, as illustrated by 1900 .
Although the example of Figure 19 shows how such collaboration is used in a game, such collaboration can be an enormously powerful tool for an application. Considering the situation in which a large building is being designed for New York City by an architect in Chicago for a real estate developer in New York, the decision was to travel and coincidentally involve financiers in Miami Airport, satisfying investors and real estate developers. In order to do so, decisions need to be made about the design elements of the building as to how it will fit into the nearby building. A construction company has a high resolution monitor with a camera attached to a PC in Chicago, a real estate developer has a laptop with a camera in New York, and an investor has a mobile phone with a camera in Miami. Construction companies can use the hosting service 210 to host powerful architectural design applications with very realistic 3D rendering capabilities, and can use many databases of buildings in New York City as well as databases of buildings under design. One architectural design application will be executed, but many of the application/game servers 1521 - 1525 require a lot of computer power. Each of the 3 users at a separate location will be connected to the hosting service 210 , each of which will be able to simultaneously view the video output of the architectural design application, but it will depend on the given device and network connectivity characteristics each user has (e.g., a commercial A 2560×1440 60fps display over an internet connection 20Mbps, a real estate developer in New York can view a 1280×720 60fps image on his laptop over a 6Mbps DSL connection, and an investor on her mobile phone at 320× over a 250Kbps cellular data connection. 180 60fps images) will be scaled appropriately by shared hardware compression 1520. Each party will hear the other party (conference calls will be handled by any of the many available conference phone software packages on the application/game server(s) 1521 - 1525), and the button action of the user input device. Accordingly, users will be able to create a video representing themselves using their local camera. As the meeting progresses, the architects make the building appear to be rotating, with the same video and highly photorealistic 3D rendering that all groups can view, adapted to the resolution of each group's display device, and move it to the side of an area with other buildings. You can make it look like it's flying. It doesn't matter that none of the local devices used by any organization have the capability to handle 3D animation with such realism, let alone store or download the vast database required to render New York City's surrounding buildings. No problem.
From each user's point of view, despite the distance, despite being an individual local device, they will have an experience that will simply be seamless with an incredible degree of realism. And, when a group wants their faces to better convey their emotional state, they can convey it. Moreover, if a real estate developer or investor wants to control a building program or want to use their own input device (either a keyboard, mouse, keypad, or touch screen), they can control it and without perceptible latency. will respond (assuming their network connection doesn't have unreasonable latency). For example, in the case of a mobile phone, if the mobile phone connects to the WiFi network at the airport, it will have very low-latency. But if you use the cellular data networks available today in the US it will probably suffer a noticeable lag. Still, the purpose of most gatherings is for investors to watch architects control building fly-bys or to talk via video conferencing, even though cellular latency is acceptable.
Finally, at the end of the collaborative conference call, real estate developers and investors will make their comments and shut down the hosting service, and the building firm will "rewind" the meeting video stored in delay buffer 1515. rewind)" and the surface representations and/or motions and opinions applied to the 3D model of the building created during the meeting. If there is a particular segment they want to store, the segment of video/audio may be moved from the delay buffer 1515 to the RAID arrays 1511 - 1512 for archival storage and later playback.
Also, from a cost standpoint, if architects need to use computing power and New York City's huge database for a 15-minute conference call, they'll want to have an expensive copy of the large database and own a very capable workstation. Rather, you only need to pay for the time you use the resource.
<b>Video-Rich Community Services (</b><b>Video</b><b>-</b><b>rich</b><b></b><b>Community</b><b></b><b>Service</b><b>)</b>
The hosting service 210 enables an unprecedented opportunity to install the Internet's video-rich community service. 20 illustrates an exemplary user page for a game player in hosting service 210 . Like the game finder application, the user page is an application executed in one of the application/game servers 1521 - 1525. All thumbnails and video windows on this page show a continuously moving video (if the segments are short and they repeat).
By using a video camera or uploading a video, a user (which user's name is "KILLHAZARD") can post his own video 2000 for other users to view. The video is stored in RAID arrays 1511 - 1512 . Also, when another user comes to KILLHAZARD's user page, if at that time KILLHAZARD uses a hosting service 210, whatever he is doing is a live video 2001 (user viewing his user page to watch him) Assuming he allows it). This depends on whether KILLHAZARD is active and, if so, the application/game server 1521 - 1525 he uses, the application/game server 1521 - 1525 hosting the user page application requested from the service control system 401. ) will be achieved by Then, the same method is used by the game finder application, and the video stream and format compressed to the appropriate resolution is sent to the application/game server 1521 - 1525 running the user pager application and it will be displayed. If the user selects a window of KILLHAZARD's live gameplay, then clicking appropriately with their input device, the window will be enlarged (again, using the same method as the Game Finder application, depending on the nature of the viewing user's Internet connection). Where appropriate, at the resolution of the viewing user's display device 422, the live video will fill the screen).
The main advantage over the prior art approach is that it allows the user to view a user page playing a live game that the user does not own, which can be done very well without having to have a local computer or game console where the game is being played. It provides a great opportunity for the user to see the game being played "in action" shown on the user page, and it provides an opportunity for the viewing user to learn about the game they want to try or do better. .
Camera-recorded or uploaded video clips from KILLHAZARD's buddies (2002) are also displayed on the user page, with text under each video clip indicating whether the friend is playing an online game (eg, six_shot is "Eragon") playing games and MrSnuggles99 is offline, etc.). By clicking on a menu item (not shown) a friend's video clip is recorded and switched from viewing the uploaded video to a live video of a friend playing the game on the hosting service 210 operating at their game moment. Thus, it will be a game finder for grouping friends. If a friend's game is selected and it is clicked, it will be enlarged to full screen, and the user will be able to see the game being played in full screen live.
Again, a user watching a friend's game does not need to own a copy of the game or local computing/game console resources to play the game. Watching the game is effectively instantaneous.
As previously described, when a user plays a game on the hosting service 210 , the user can "rewind" the game and find the video segment he wants to save, and put the video segment on his user page. can be saved These are called "Brag Clips". The video segment 2003 is all the Bragg clips 2003 saved by KILLHAZARD from the previous games he played. Number 2004 shows how many times a Bragg clip has been viewed, when a Bragg clip is viewed, the user has the opportunity to rate them, and a number of orange keyhole-shaped icons 2005 indicate how high the rating show that it represents Bragg clip 2003 continuously loops according to the video left on the page as the user views the user page. When the user selects and clicks on one of the bragg clips 2003, it will play, pause, rewind, fast-forward, step through, etc. Zoom in to show the Bragg clip 2003 according to the DVR control that allows the clip.
The Bragg Clip 2003 playback is played by the application/game server 1521 - 1525 which loads the compressed video segments stored in the RAID array 1511 - 1512 when the user records the Bragg clip, decompresses it and plays it again. is implemented
Bragg clips 2003 may also be "3D DVR" video segments (eg, game state sequences in games that allow the user to change camera viewpoints and can be replayed) in games that support such functionality. The game state information in this case is stored plus a compressed video record of a specific "fly-through" that the user makes when a game segment is recorded. When the user page is being viewed, all thumbnails and video windows will loop continuously, and the 3D DVR Bragg Clip 2003 is a Bragg clip recorded as compressed video when the user records a "fly-through" of a game segment. (2003) will be looped continuously. However, when the user selects a 3D DVR Bragg clip 2003 and clicks on it, we add a DVR control that allows the compressed video Bragg clip to be played, and the user clicks a button giving them 3D DVR functionality for the game segment. You can do it. They will be able to control the camera "fly through" during their game segment, and if they wish (and the user with the user page allows it), they will be able to "fly through" the optional Bragg clip in compressed video form. be able to record and then make available to other viewers of the user's page (immediately or after the owner of the user's page has had the opportunity to review the brag clip).
This 3D DVR Bragg Clip 2003 performance may enable a game to replay game state information recorded in other application/game servers 1521 - 1525. Since the game can be activated almost immediately (as previously described), it is not difficult to activate that game limited to the game state recorded by the Bragg clip segment, then write the compressed video to the delay buffer 1515 . Allows the user to "fly through" to the camera while on the go. Once the user has finished running, the "fly through" game is deactivated.
From the user's point of view, an activated "fly through" with a 3D DVR Bragg clip 2003 eliminates the effort of controlling the DVR control of the linear Bragg clip 2003. They won't know how to play the game or anything about it. They are just virtual camera operators looking at the 3D world during game segments recorded by others.
Users will also be able to multiple-record their own audio as Bragg clips that are recorded or uploaded from the microphone. In this way, Bragg clips can be used to create custom animations using characters and motions from the game. This animation technique is commonly known as "machinima."
As users progress through the game, they will achieve different skill levels. The game being played reports achievements to the service control system 401, and this skill level is posted on the user page.
<b>interactive</b><b> Animated advertisements (</b><b>Interactive</b><b></b><b>Animated</b><b></b><b>Advertisements</b><b>)</b>
Online advertising has shifted from text to still images, to video, now to interactive segments, and typically thin clients like Adobe Flash have done using animations. The reason thin clients use animation is that users usually have a little patience with the privileges of a service or product sold to them. Also, thin clients continue to operate as very low performance PCs, and advertisers can have a high degree of confidence that interactive advertisements can work properly. Unfortunately, animated thin clients such as Adobe Flash are limited in the degree of interactivity and duration of the experience (to mitigate download times).
Fig. 21 shows an interactive advertisement where the user can select the exterior and interior colors of the car while rotating around the car in the showroom, and real-time ray tracing shows how the car looks. After the user has then selected an avatar to drive the car, the user can take the car for a drive to a race track or to an exotic locale such as Monaco. Users can opt for a larger engine, or better tires, and see how the changed configuration affects the car's ability to accelerate and stick to the road.
Of course, the advertisement is an effective sophisticated 3D video game. However, in order for these advertisements to be playable on a PC or video game console it will probably require a 100MB download, in the case of a PC it may require the installation of special drivers and will not work if the computing power of the appropriate CPU or GPU is insufficient. It may not be. Accordingly, such advertisements are unrealistic in prior art configurations.
In the hosting service 210, these advertisements start almost instantly, run perfectly, and it doesn't matter what the user's client 415 performance is. So they are faster than thin client interactive advertising, the experience is richer, and very reliable.
<b>Streaming geometry during real-time animation (</b><b>Streaming</b><b></b><b>Geometry</b><b></b><b>During</b><b> Real-time </b><b>Animation</b><b>)</b>
RAID arrays 1511 - 1512 and inbound routing 1502 are used to reliably deliver geometry instantly in applications or during gameplay during real-time animations (e.g., fly-through to complex databases). and designing video games and applications that rely on inbound routing 1502 to provide transfer rates that are too low and have too low latency.
With prior art systems, such as the video game system shown in Figure 1, huge storage devices are available, especially in real homes, where the geometry required during game play is so slow except in situations where the required geometry is somewhat predictable. cannot be streamed. For example, in a game of driving on a particular road, the geometry for a building coming into view can be reasonably well predicted and a huge storage device can pre-find the location where the upcoming geometry is located.
However, in complex scenes with unpredictable changes (e.g. in battle scenes with complex characters around) if the RAM in the PC or video game system is completely filled with geometry for objects in the current field of view, the user suddenly loses their character. Rotate their characters to see what's behind them, and if the geometry isn't preloaded into RAM, there will be a delay in displaying it.
The RAID arrays 1511 - 1512 in the hosting service 210 are capable of streaming data in excess of Gigabit Ethernet speeds and have an SNA network, which achieves 10 Gigabit per second rates over 10 Gigabit Ethernet or other network technologies. it is possible 10 gigabits per second will load at least 1 gigabyte per second of data. At 60 fps frame time (16.67 ms), approximately 170 megabits (21 MB) of data can be loaded. Of course, even in a RAID configuration, spinning media will still have latencies greater than one frame time, but flash-based RAID storage will eventually be as large as spinning media RAID arrays and will not experience such high latencies. In one embodiment, massive RAM write-through caching will provide very low-latency access.
Thus, mass storage with sufficiently high network speeds, and with sufficiently low enough latency, geometry can be streamed to the application/game game servers 1521 - 1525 as fast as the CPU and/or GPU can process 3D data. . So, in the example given above, if the user suddenly wants them to rotate the character and look behind them, the geometry for all the characters behind them can be loaded before the character's rotation is complete, so to the user it comes alive like live action. It will appear that he or she is in the world of porn realism.
As discussed earlier, one of the final limitations of photorealistic computer animation is the human face, and since the sensitivity of the human eye is not perfect, a slight error from the photoreal face can lead to negative reactions from the observer. 22 is Contour<sup>TM</sup> Subject of Reality Capture Technology (co-pending) application: "Method and Apparatus for Capturing Actions of Executors" 10/942,609, filed September 15, 2004; "Methods and Apparatus for Capturing Representation of Executors" ", 10/942,413, filed September 15, 2004; "Methods and Apparatus for Improving Marker Identification in Motion Capture Systems" 11/066,954, filed February 25, 2005; "Using Shutter Synchronization to Capture Motion "Method and Apparatus for Performing", 1/077,628, filed March 10, 2005; "Method and Apparatus for Capturing Motion Using Arbitrary Patterns on Capture Surfaces" 11/255,854, filed October 20, 2005; Methods and Systems for Performing Motion Capture Using Phosphor Applied Techniques", 11/449,131, filed Jun. 7, 2006; "System and Method for Performing Motion Capture Using Fluorescent Lamp Glare Effect," 11/449,043, filed Jun. 7, 2006; "Methods and Systems for Three-Dimensional Capture of Still Motion Animated Characters", 11/449,127, filed June 7, 2006, each of which is now assigned to the assignee of a CIP application) how to achieve very smooth live performance It is plotted in terms of a high polygon-count that results in a capture surface, then tracked to the surface (eg, polygon motion accurately follows the motion of the face). Finally, the video of the live performance is mapped to the tracked surface to create a textured surface and a photoreal result is generated.
Although current GPU technology can render polygons and surface textures and lights in real time on multiple tracked surfaces, if polygons and textures change every frame time (which will produce the most photorealistic results), it will quickly It will consume the available RAM of modern PCs and video game consoles.
Using the streaming geometry technology described above, it actually feeds geometry continuously to the application/game game servers 1521 - 1525 so they can continuously animate photoreal faces and create video games with faces. It is almost indistinguishable from a live, working face.
<b>interactive</b><b> Integration of linear content with features (</b><b>Integration</b><b></b><b>of</b><b></b><b>Linear</b><b></b><b>Content</b><b> with </b><b>Interactive</b><b></b><b>Features</b><b>)</b>
Movies, TV shows and audio material ("Linear Content" is widely available to home and office users in many forms) Linear Content can be acquired from physical media such as CD, DVD, HD-DVD and Blu-ray media. have. It can also be recorded with DVRs in satellite and cable TV broadcasts. And it is available as pay-per-view (PPV) content via video-on-demand (VOD) on satellite and cable TV and cable TV.
Increasingly, linear content is available over the Internet as both download and streaming content. There is no single place today to experience all the features associated with linear media. For example, DVDs and other video optical media typically have interactive features not available elsewhere, such as a director's commentary "making of" featurettes. Online music sites usually have cover art and song information not available on CDs, but not all CDs are available online. Web sites related to television programming often feature other than, blogs, sometimes holding comments from actors or staff.
Moreover, many videos or sporting events will have video games released frequently (in the case of video), often with linear media (in the case of sports), or closely linked to real-world events (eg, player trades, etc.) .
Hosting service 210 is well suited for the delivery of linear content that is linked together with individual forms of related content. Of course, the video it provides will no longer require the highly interactive video games it delivers, and the hosting service 210 can deliver linear content to a wide range of devices or mobile devices in the home or office. 23 illustrates an exemplary user interface page for hosting service 210 showing a selection of linear content.
However, unlike most linear content delivery systems, the hosting service 210 also provides associated interactive components (eg, Adobe Flash animations on websites (as described below), interactive overlays on HD-DVDs, and DVDs on DVDs). menus and functions). Accordingly, the client device 415 introduces no further restrictions as to what its functionality may be used.
In addition, the hosting system 210 may link together dynamic video game content and linear content in real time. For example, if a user is watching a Quidditch match from a Harry Potter movie, she can click a button and the movie will stop and immediately she will be taken to the Quidditch segment of the Harry Potter video game. After playing the Quidditch match, click another button, and the movie will resume immediately.
A video captured photographically with photoreal graphics and production technology is indistinguishable from a live action character, and, as described here, when a user creates a move from a Quidditch game to a live action movie to a Quidditch game to a video game on a hosting service. , the two scenes are virtually indistinguishable. This provides a new creative option for directors of linear and interactive content (eg video games) to be unable to tell the difference between the two worlds.
Utilizing the hosting service architecture shown in FIG. 14 , control of the virtual camera in the 3D movie can be provided to the viewer. For example, in a scene taking place inside a train, it would be possible to look around the train during story progression and allow the user to control the virtual camera. It is available ("assets") of all three-dimensional objects on the train with a reasonable level of computer power capable of rendering the scene in real-time as in the original movie.
Even in non-computer generated entertainment, there are very interesting interactive features that can be provided. For example, the 2005 movie "Pride and Prejudice" has many scenes of a gorgeous ancient British mansion. In certain mansion scenes, the user can pause the video and then control the camera for a tour of the mansion or surrounding area. To do this, Apple's prior art camera with a fish-eye lens that tracks its location could be carried through the mansion. QuickTime VR launches. The various frames are converted and the image is undistorted, stored in a RAID array (1511-1512) with the movie, and played back when the user makes a selection for a virtual tour.
Sporting events, such as basketball games, live sporting events will be streamed through the hosting service 201 for users to watch as they would on regular TV. After a user watches a particular match, the game's video game (a basketball player that looks as photoreal as a real player after all) starts with the player starting from the same location, and the user (perhaps each controlled by a single player) plays that game. You can run this again to see if you can do better than the player.
The hosting service 210 described herein is well-suited to support the future world because it allows it to withstand computing power and mass storage resources that would be unrealistic to be installed in a home or most office setting, and also in a home setting, As opposed to always having older generations of PCs and video games it keeps computing resources up to date, the latest computing hardware is available, and in the hosting service 210 all this computing complexity is hidden from users, even if they Although from the user's point of view you are using a very sophisticated system, from the user's point of view it is as simple as changing a channel on your television. Moreover, the user may have access to all computing power, and experience with the computing power may be drawn from any client 415 .
<b>Multiplayer games (</b><b>Multiplayer</b><b></b><b>Games</b><b>)</b>
When the game is a multiplayer game, the application/game game servers 1521 - 1525 may be connected to the network connected to the Internet and the game machine not operating in the hosting service 210 through the inbound routing 1502 network. When playing multiplayer games with a computer on the general Internet, the application/game game servers 1521 - 1525 have the advantage of being able to access the Internet very quickly (compared to the case of playing the game with a server at home), but more May be limited by the performance of other computers playing on slow connections, and potentially by the fact that game servers on the Internet designed to assign at least a common denominator, which will usually be home computers with slow consumer Internet connections.
However, a different world may be achieved if the multiplayer game takes place entirely on the hosting service 210 server center. Each application/games game server (1521 - 1525) hosting games for users can be any server hosting central control for multiplayer games with very high speed, very low-latency and vast and very fast storage arrays. Instead, it may be connected to other application/game game servers 1521 - 1525. For example, if Gigabit Ethernet is used as the inbound routing 1502 network, application/game game servers 1521 - 1525 communicate with each server and potentially play multiplayer games at gigabytes per second with 1ms latency or less. It will communicate with some server hosting a central control for Moreover, RAID arrays 1511 - 1512 can respond very quickly and transfer data at gigabit per second rates. For example, if a user customizes a character with appearance and accoutrements, the character will have unique behavior and many geometries, and with prior art systems limited to game clients running home PCs or game consoles, if When the character comes into the sight of another user, that user waits for a long and slow download to complete, loading all geometry and behavior data into their computer. Within hosting service 210, the same download may go over Gigabit Ethernet provided by RAID arrays 1511 - 1512 at the rate of gigabits per second. If a home user has an 8Mbps internet connection (which is very fast by today's standards), Gigabit Ethernet is 100x faster. So what would take an extra minute over a fast internet connection would take less than a second over Gigabit Ethernet.
<b>Top player groupings and tournaments (</b><b>Top</b><b></b><b>Player</b><b></b><b>Groupings</b><b></b><b>and</b><b></b><b>Tournaments</b><b>)</b>
Hosting service 210 is well suited for tournaments. Because if a game doesn't run on the local client, there is no chance to cheat users. Also, because of the ability of output routing 1540 to multicast UDP streams, hosting service 210 can broadcast major tournaments to thousands of people in an audience at a time.
In fact, some video streams are so popular that when a large number of users try to receive the same stream (eg, watching a major tournament), it will reach many client devices 415 by a content delivery network (CDN) such as Akamai or Limelight for mass distribution. ) to send the video stream more efficiently.
A similar level of efficiency can be achieved when CDNs are used to display game finder pages that group top players.
In major tournaments, live broadcast celebrity announcers can be used to provide live broadcasts from certain matches. Although a large number of users watch a major tournament, a relatively small number will play the tournament. The celebrity announcer's audio can be routed to an application/game server 1521 - 1525 that hosts users playing in the tournament and hosts an observer mode copy of any game in the tournament, and the audio can be recorded multiple times in addition to the game audio. . The celebrity announcer's video can be overlaid on the game, and possibly even from an observer's point of view.
<b>Acceleration of web page loading (</b><b>Acceleration</b><b></b><b>of</b><b></b><b>Web</b><b></b><b>Page</b><b></b><b>Loading</b><b>)</b>
The World Wide Web's primary transport protocol, Hypertext Transfer Protocol (HTTP), was understood and defined in an era where only businesses had high-speed Internet connections and consumers who were online used dial-up modems or ISDNs. At the same time, the "gold standard" of fast connections is the T1 line, which transmits 1.5Mbps data simultaneously (eg, with the same amount of data in both directions).
Today, the situation is completely different. In a world of many advancements, the average home connection speed via a DSL or cable modem is significantly higher downstream than a T1 line. In fact, in some parts of the world, fiber to the curb delivers transfer rates as fast as 50 to 100 Mbps, supposedly.
Unfortunately, HTTP wasn't designed to take advantage of this dramatic speedup as an effective advantage (it wasn't even realized). A website is a collection of files on a remote server. Simply put, HTTP requests the first file, waits for the first file to be downloaded, then requests the second file and waits for the file to be downloaded. In fact, HTTP allows more than one "open connection", e.g. more than one file requested at a time, but because of the agreed-upon standard (and protecting the web server from overload). ), only a few open connections are allowed. Moreover, because the path of the web page is structured, the browser often does not recognize multiple simultaneous pages available for immediate download (eg, after parsing the page's phrase, the image that needs to be downloaded). new files will appear). Thus, the files on the website are essentially loaded one at a time because of the request-response protocol used in HTTP, and there is a 100ms latency with respect to each file being loaded (the typical approach of a web server in the US).
With a relatively low speed connection, the download time of the file itself takes precedence over the time waiting for the web page, so it doesn't cause much trouble. However, problems start to arise, especially as the speed of connections to complex web pages increases.
In the example shown in Fig. 24, a typical commercial website is shown (this particular website belongs to a major athletic shoe brand). This website has 54 files. Those files include HTML, CSS, JPEG, PHP, JavaScript and Flash files, and video content. A total of 1.5 Mbytes must be loaded before the page is activated (i.e. the user can click on it and it is usable). There are many reasons to have a large number of files. First, it is a complex and sophisticated webpage, and for other reasons, it dynamically collects information based on the information of the user accessing the page (i.e., what country, what language the user is in, what language the user is in, what the user has purchased before. etc.) and depending on these factors, other files are downloaded. Still, this is a typical commercial web page.
Figure 24 shows the total amount of time that elapses before a web page is opened as the connection speed increases. With a 1.5Mbps connection speed (2401), using a traditional server with a traditional web browser, it takes 13.5 seconds for a webpage to open. With a 12Mbps connection speed of 2402, load times are reduced to 6.5 seconds, or about twice as fast. However, with a 96Mbps connection speed (2403), the load time is reduced to only 5.5 seconds. The reason is that at such a high download speed, the time to download the file itself is negligible, but the latency per file still remains approximately 100ms each, which results in a latency of 54 files*100ms = 5.4 seconds. So, no matter how fast you have a home connection, this website will always take at least 5.4 seconds to activate. Another factor is server-side queuing; All HTTP requests are added after queuing, so on a busy server this will have a significant impact because of every little item you can get from the web server, HTTP requests need to wait for its order.
One way to solve this issue is to either discard or redefine HTTP. Or, let the website owner consolidate several files into a single file (eg, in Adobe Flash format). But the real problem is that not only companies, but many others make significant investments in building websites. Moreover, if some homes are connected at 12 - 100Mbps speed, the majority of homes are connected at low speed, and HTTP works fine even at low speeds.
One alternative is to host the web browser on the application/game server 1521 - 1525, host the web server's files on the RAID array 1511 - 1512 (or potentially in RAM or the application hosting the web browser) /from local storage to game server (1521 - 1525), with very fast interconnect via inbound routing (1502) (or local storage), using HTTP rather than using HTTP to show latency of 100ms per file There will be minimal latency per file. Then, instead of having the user at her home access the webpage via HTTP, the user can access the webpage via the client 415 . Then, with a 1.5Mbps connection (because this website doesn't require much bandwidth for its video), the webpage could be active in less than 1 second (2400) per line. There will be essentially no latency before the web browser running on the application/game servers 1521 - 1525 displays the page, and there will be no appreciable latency until the client 415 displays the video output from the web browser. As the user's mouse moves and/or types on the webpage, the user's input information will be sent to the web browser running on the application/game servers 1521 - 1525, and the web browser will respond accordingly.
A disadvantage of this approach is that if the compressor is continuously sending video data, bandwidth is used, even if the webpage is static. This can be addressed by configuring the compressor to send data only when the webpage changes, and if so, to send data to the part of the page that changes. On the other hand, there are websites with some flash banners that are constantly changing, those websites tend to upset, and usually the webpages are static unless there is some cause for them to move (video clips). For such a website, less data will be transferred using the hosting service 210 than a traditional web server because only the images that are actually shown are transferred, there is no code that the thin client can execute, and the rollover image There is no such huge object as never seen.
Thus, using the hosting service 210 to host legacy webpages, webpage consumption time can be reduced where the webpage is opened, such as changing channels on a television: the webpage is effectively instantly activated do.
<b>game and </b><b>of the application</b><b> Debugging possible (</b><b>Facilitating</b><b></b><b>Debugging</b><b></b><b>of</b><b></b><b>Games</b><b> and </b><b>Applications</b><b>)</b>
As mentioned earlier, video games and applications with real-time graphics are very complex applications and typically they are brought to market with bugs. Even if software developers get feedback from users about bugs, it means returning the machine after a breakdown, which makes it difficult to pinpoint what is causing the game or real-time application to malfunction or behave improperly.
When the game or application operates in the hosting service 210 , the video/audio output of the game or application is continuously recorded in the delay buffer 1515 . Moreover, the monitoring process operates each application/game server 1521 - 1525 to periodically report to the hosting service control system 401 whether the application/game server 1521 - 1525 is operating stably. If the monitoring process fails to report, then the server control system 401 will attempt to communicate with the application/game servers 1521 - 1525, and if successful collect machine state not available. Any information available with the recorded video/audio by delay buffer 1515 will be sent to the software developer.
Thus, the game or application software developer gets a notification of the failure from the hosting service 210 and a record for each frame that caused the failure. This information is invaluable in uncovering and fixing bugs.
In addition, when the application/game servers 1521 - 1525 fail, the server restarts at the most recent restart point, and provides a message of apology for technical difficulties to the user.
<b>resource</b><b> Sharing and Savings (</b><b>Resource</b><b></b><b>Sharing</b><b></b><b>and</b><b></b><b>Cost</b><b></b><b>Savings</b><b>)</b>
The system shown in Figures 4a and 4b provides a number of advantages to users and game and application developers. For example, typically home and office client systems (eg, PCs or game consoles) are used only a small percentage of the time per week. "Active Gamer Benchmark Study" distributed October 5, 2006 by Nielsen Entertainment (http://prnewswire.com/cgi-bin/stories.pl?ACCT=104&STORY= According to /www/story/10-05-2006/0004446115&EDATE=), active gamers spend an average of 14 hours a week playing video game consoles, and using handheld computers about 17 hours a week. The data also noted that for all gaming activities (including console, hand-held and PC gaming), active gamers spend an average of 13 hours per week. Considering that the time spent playing console video games accounts for a high percentage, a week (24 hours * 7 days) is 168 hours. This means that in the homes of active gamers, video game consoles are only used 17/168 = 10% of the time of the week. Or 90% of the time the video game console is not in use. Given the high cost of video game consoles, the fact that manufacturers subsidize such devices is an inefficient use of expensive resources. In business, PCs are also used only a small amount of time a week, and especially non-portable desktop PCs often require top-notch applications like Autodesk Maya. Although some business is conducted all hours and holidays, some PCs (eg, portable PCs brought home to work in the evening) are used all hours and holidays, and most business activity is conducted Monday through Friday during the working hours. Desktop PC utilization tends to follow these operating hours, as they tend to be concentrated between 9 am and 5 pm, less use on holidays and break times (such as lunchtime), and mostly PC use. If we assume that the PC is continuously used 5 days a week from 9:00 to 5:00, the PC is used as much as 40/168 = 24% of the time of the week. A high-performance desktop PC is a very expensive investment for the office, resulting in very low levels of utilization. Schools that teach using desktops use computers for fewer hours a week, and although teaching hours vary, most teaching takes place during the day, Monday through Friday. Thus, in general, PCs and video game consoles are only utilized for a small amount of time per week.
Notably, during the non-holiday Monday-Friday daytime, many people use their computers at the office or at school, and these people generally don't play video games during these hours. And when you're playing video games, it's usually at a different time zone: evenings, weekends, and holidays.
Looking at the form of the hosting service shown in Fig. 4a, the utilization pattern is the result of very efficient utilization of resources, as described in the two paragraphs above. The number of users that can be served by the hosting service 210 at any given time will be limited, especially if the user requires real-time response of a complex application, such as a sophisticated 3D video game. However, unlike a PC used for business or a video console at home that is not in use most of the time, the server 402 may be reused by other users at different times. For example, a high-performance server 402 with a high-performance dual CPU and dual GPU and a large amount of RAM can be utilized in offices and schools between 9:00 am and 5:00 pm on weekdays, but is used to play sophisticated video games on weekends and holidays. Used by gamers. Similarly, low-performance applications can be utilized during work hours by a Celeron CPU, a low-performance server 402 without a GPU or with a low-performance GPU, limited RAM, and a low-performance game can be utilized during off-hours by a low-performance server 402 with limited RAM. can utilize
Moreover, with the deployment of hosting services described here, resources are efficiently shared among thousands if not millions of users. In general, online services have only a small percentage of total users who use the service at any given time. If you consider the Nielsen video game usage statistics published earlier, it's easy to see why. If an active gamer plays console games 17 hours a week, we will assume that the peak hours of gaming are off-hours in the evenings (5 AM - 12 PM, 7 hours*5 days = 35 hours/week) and weekends. Assuming that (8 am to 12 am, 16 hours * 2 = 32 hours / week), there is a peak time of 65 (35 + 32) hours in 17 hours of game time. The exact peak user loading of the system is difficult to estimate for a number of reasons: some users play during off-peak hours and it will certainly be daytime when users form groups, and peak times influence the type of game (e.g., Children tend to play games in the early evening). However, the average time gamers play games is much less than during the day when gamers play games, and the hosting service 210 is used by only a small number of users at any given time. For the purposes of this analysis, we will assume a peak load of 12.5%. Thus, only 12.5% of computing, compression, and bandwidth resources are used at any given time, resulting in only 12.5% of hardware costs supporting a given user to play a given level of performance game due to the reuse of resources.
Moreover, if some games and applications require more computing power than others, resources will be dynamically allocated based on the game the user is playing and the application they are running. So, a user who selects a low-performance game or application will be assigned a low-performance (less expensive) server 402, and a user who selects a high-performance game or application will allocate a high-performance server 402 (more expensive). In practice, a given game or application has both a low performance and a high performance portion of the game or application, and users can choose between portions of the game or application to keep the user running on the lowest price server 402 according to the needs of the game or application. One server 402 and another server 402 can be exchanged. A RAID array 405, which is much faster than a single disk, can be used even on a low-performance server 402, which has the advantage of speeding up the disk transfer speed. Thus, depending on both the game being played or the application being used, the average cost per server 402 is much lower than the price of the most expensive server 402 of a high-performance game or application, but even a low-performance server 402 can use RAID. A disk performance advantage will be obtained from the array 405 .
Moreover, the server 402 of the hosting service 210 is nothing more than a diskless PC motherboard or peripheral interface other than a network interface, and will only be down-integrated for a single chip with a fast network interface with the SAN 403 . Also, the RAID array 405 will be shared by more users than disks and the disk cost per active user will be less expensive than a single disk drive. All of these devices will be racked in an environmental control server room environment. If the server 402 fails, it can be easily repaired or replaced at the hosting service 210 . In contrast, in a home or office, a PC or game console is a sturdy, stand-alone device that must withstand reasonable mechanical losses from being struck or dropped, must have a housing (cover), must have at least one disk drive, and must have at least one disk drive and adverse environmental conditions. They are sold by retailers who earn a distribution margin because they have to withstand high pressure (for example, they must be placed in an overheated AV cabinet along with other devices), require a service guarantee, and must be packaged and shipped. In addition, even if a low-performance game or device is running most of the time (or a section of a game or application), the PC or game console can be used at some point in the future to provide the optimal performance of the computer-intensively expected game or application. Environment settings must be made. And, if a PC or console fails, repairing it is an expensive and time-consuming process (which adversely affects manufacturers, users, and software developers).
Thus, the system illustrated in FIG. 4A provides an opportunity for users at home, office, or school to compare with the resources of a local computer to experience a given level of computer performance and provides computing power through the architecture illustrated in FIG. 4A. would be much cheaper.
<b>Eliminate the need to upgrade (</b><b>Eliminating</b><b></b><b>The</b><b></b><b>Need</b><b></b><b>to</b><b></b><b>Upgrade</b><b>)</b>
Moreover, users no longer have to worry about upgrading their PC or console to handle new games or new high-performance applications. Regardless of what type of game or application the server 402 requires, the hosting service 210 allows the user to use any game or application almost immediately (ie, a RAID array 405 or server 402). fast loading of local storage), properly up-to-date and bug-fixed (i.e. software developers can choose to set up server 402 in an ideal environment for a given game or application, with optimal drivers After some time with the configuration of the server 402, the developer immediately provides updates and fixes bugs to all copies of the game or application on the hosting service 210). In fact, after the user starts using the hosting service 210, the user finds that they have started providing a better experience for their games and applications (eg, through updates or bug fixes) and it is probably because the user didn't even exist a year ago. A user discovers a new game or application made available for use in the service 210 that utilizes computing technology (eg, high performance GPU) after one year. Therefore, it is impossible for users to purchase the technology before playing the game or running the application after a year. Because the computing resources that play games and operate the device are invisible to the user (i.e., to the user's perception, the user simply selects a game or application that works almost immediately as much as changing the television channel), so the user's hardware is upgraded It will be an "upgrade" without recognizing it.
<b>Eliminate the need for backups (</b><b>Eliminating</b><b></b><b>the</b><b></b><b>Need</b><b></b><b>for</b><b></b><b>Backups</b><b>)</b>
Another major problem users face in business, school and at home is backups. Information stored on a local PC or video game (eg, a user's game achievement and ranking in the case of a console) may be lost due to a disc error or accidental deletion. Manuals or automatic backups are provided for PCs that can be used for many applications, and game consoles can be uploaded to an online server for backup, but local backups are typically copied to another local disk that is stored somewhere safe and organized ( or other storage devices where data is not lost), backups via online services are often limited by slow upstream speeds, typically over low-cost Internet connections. The data stored in the RAID array 405 along with the hosting service 210 of Figure 4a can be configured using the well known RAID configuration prior art, which means that no data will be lost if the server center fails. 's technician will notify you of the failed disk and replace it with the disk that has automatically executed the update, and the RAID array will withstand the failure once more. In addition, disks backed up on a regular basis on a server center or secondary storage where all of the disk drives are nearby and there is a fast local network between them via the SAN 403 and can be relocated and stored in a server center or remote location. It is not difficult to deploy everything in the system to the server center. From the point of view of the user viewing the hosting service 210 , the data is always simply protected and never have to think about backing up.
<b>access the demo (</b><b>Access</b><b></b><b>to</b><b></b><b>Demos</b><b>)</b>
Users often want to try a game or device before purchasing. As described above, there is prior art to try out games and applications (the word demo means try a demonstration version, also called demo (not a noun)), but demo games and applications. Each of them suffers from limitations and/or inconveniences. When using the hosting service 210, it is easy and convenient for users to try a demonstration. In fact, every user selects and tries a demo through the user interface (as described below). The demo will load almost immediately on the appropriate server 402 and run like any other game or application. The demo will work regardless of whether the demo requires a high performance server 402 or a low performance server 402 , or whether the user is using a home or office client 415 . Game or application demo software makers will be able to control precisely which demos are allowed to use for how long and, of course, demos will contain user interface elements that provide users with the opportunity to gain access to a full version of their game or application. there should be
Some users will use the demo over and over again (especially game demos that they enjoy playing over and over again) because demos are easy to offer below cost or for free. The hosting service 210 may apply a variety of techniques to limit the use of the demo for a given user. The simplest approach is to set a username for each user and limit the number of times a given username is allowed to run the demo.
However, the user will set up multiple user IDs, especially if it is free. One technique for dealing with this problem is to limit the number of times a given client 415 is allowed to run a demo. If the client is a standalone device, the device will have a serial number and the hosting service 210 may limit the number of times a demo can be handled by a client with that serial number. If the client 415 is running as software on a PC or other device, a serial number may be assigned by the hosting service 210 and stored on the PC will be used to limit demo use, but if the user reprograms the PC Serial number can be erased or changed, another countermeasure is that the hosting service 210 records the PC network expansion card, MAC address (and/or other mechanical special identifier such as hard drive serial number, etc.) and uses the demo. is to limit However, assuming that the MAC address of the network expansion card is changed, this is not a method without fear of failure. Another approach to limiting the number of uses of a demo is to allow it to run with a given IP address. Although IP addresses are periodically reassigned by cable modems and DSL providers, this does not happen very often in practice and once an IP within a residential area for DSL or cable modem access is determined, the small number of demo uses is typically a given assumption. can be set to Also, there may be many machines in a home under a NAT router sharing the same IP address, but in a typical residential environment there will be a limited number of such devices. If the IP address is for one company, then the majority of the demos can be set up for one company. But finally, a combination of all of the previously mentioned approaches is the best way to limit the number of demos on your PC. Although there is no firm, technical, and failure-free way to limit the repeated use of the demo by experienced users, it creates many obstacles and the majority of PC users' abuse of the demo system provides enough deterrents not to make an effort. You can create, and rather use demos as users tend to try new games and applications.
<b>school, </b><b>business</b><b> and benefits from other organizations (</b><b>Benefits</b><b></b><b>to</b><b></b><b>Schools</b><b>, Businesses </b><b>and</b><b></b><b>Other</b><b></b><b>Institutions</b><b>)</b>
Significant advantages accumulate for businesses, schools, and other organizations that utilize the system illustrated in FIG. 4A . Offices and schools install, maintain, and upgrade PCs, especially when it comes to PCs running high-performance applications like Maya. As mentioned earlier, PCs are generally used only for some time a week, and the price of a PC for a given level of performance is much higher in an office or school environment than in a server center environment.
For large businesses or schools (eg, large universities), this may be practical for IT departments such as entities that build server centers and maintain computers with remote access via LAN level connections. A number of solutions exist for remote computer access over a LAN or private high-bandwidth connection between offices. For example, through a virtual network computing device such as Microsoft Windows Terminal Server or RealVNC's VNC or Sun Microsystems' thin client, users can gain remote access to a PC or server with a certain range of quality of graphical response time or user experience. be able to In addition, such self-managed server centers are typically dedicated to one business or school, etc., and take advantage of possible overlapping usage when disparate applications (such as entertainment and office equipment) are using the same computer resources at different times of the week. can't Thus, many businesses and schools lack the scale, resources, or skills to build their own server centers with LAN-speed network connections for each user. In fact, a large percentage of schools and businesses have the same Internet connection (eg DSL, cable modem) as at home.
However, these organizations may still have a very high need for high-performance computing on a regular or periodic basis. For example, a small architectural firm will have relatively few architects, have mostly modern computing needs when working on designs, but will periodically require very high-performance 3D computing (eg, design new architectural designs for customers). when creating a three-dimensional fly-through of ). This system is well suited for such an organization as in FIG. 4a. The organization does not require any more than the same kind of network connection (eg DSL, cable modem) provided in the home, and is typically very inexpensive. They may utilize an inexpensive PC, such as the client 415 , or they may utilize a dedicated, inexpensive device that provides the PC all at once and executes only the control signal logic 413 and low-latency video decompression 412 . These features are attractive for schools that have a problem with PC theft or damage to sensitive components within the PC.
This arrangement solves many problems for such organizations (and many of these advantages can also be shared by home users doing general purpose computing). Operating costs (which ultimately have to pass through some form to users in order to have a viable business) can be very low, where (a) computing resources are shared with applications that have other peak hours of use during the rest of the week; (b) organizations can only gain access to high-performance computing resources when they are needed (and incur costs), and (c) organizations don't have to provide resources for backup in order to maintain high-performance computing resources.
<b>Elimination of piracy (</b><b>Elimination</b><b></b><b>of</b><b></b><b>Piracy</b><b>)</b>
In addition, games, applications, interactive movies, etc. can no longer be pirated as they are today. Because the game is run in a service center and users are not provided with access to the underlying program code, there is no piracy. Even if the user copied the source code, the user would not be able to run the code on a standard game console or home computer. This opens the market in places in the world like China where standard video gaming is not available. Resale of used games is also not possible.
For game developers, as is the case today, there is little market discontinuity. Hosting service 210 can be updated over time as gaming needs change gradually, in contrast to the current situation where a whole new generation of technology is being upgraded by users and developers and game developers relying on timely delivery of hardware platforms. There will be.
<b>streaming </b><b>interactive</b><b> video(</b><b>Streaming</b><b></b><b>Interactive</b><b></b><b>Video</b><b>)</b>
The above description provides a broad range of applications enabled by the new fundamental concept of a generic Internet-based, low-latency streaming interactive video (which implicitly includes audio along with video, as used herein). Prior art systems that provide streaming video over the Internet can only be enabled with applications that can run with high latency interactivity. For example, basic playback controls (eg, stop, rewind, fast forward) for linear video operate with moderately high latency, and it is possible to select among linear video feeds. And, as noted above, the nature of some video games allows them to be played with high latency. However, the high latency (or low compression ratio) of prior art approaches for streaming video severely limits the potential applications of streaming video or narrows their deployment for professional network environments, and even in such environments, the prior art It actually induces a burden on the network. The technology described herein opens the door for a wide range of applications that may have low-latency streaming interactive video over the Internet, particularly through consumer-grade Internet connections.
In fact, a client device as small as client 465 in FIG. 4C, while being small enough to provide very fast networking between powerful servers and an enhanced user experience with any amount of fast storage, and any amount of computing power, Computing of the criteria may be enabled. Moreover, since bandwidth requirements have not grown as much as the computing power of system growth (eg, bandwidth requirements are tied only to display resolution, quality and frame rate), once broadband Internet connectivity is ubiquitous (eg, widespread Through wireless coverage), reliable, high enough bandwidth to meet the needs of any user's display device 422 , and a thick client (PC running Windows, Linux, OSK, etc.) or The question of whether it is a mobile phone) or a thin client (such as Adobe Flash or Java) is necessary for typical consumer and business applications.
The advent of streaming interactive video forces us to rethink our assumptions about the structure of computing architectures. An example of this is the hosting service 210 server center embodiment shown in FIG. 15 . The video path for the delay buffer and/or group video 1550 is selectable where the multicast streaming interactive video of the application/game servers 1521 - 1525 is live via path 1552 or via path 1551 . A feedback loop that returns to the application/game servers 1521 - 1525 after a delay. This may enable a wide range of practical applications (eg, such as those shown in Figures 16,17 and 20) that are not or may not be feasible through prior art servers or local computing architectures. However, as a more general architectural feature, what feedback loop 1550 provides is a circulation of streaming interactive video levels, as the video can loop back indefinitely as the application requires. This can make the potential of a wide range of applications never available before.
Another key architectural feature is that the video stream is a one-way UDP stream. This can effectively enable streaming interactive video of any degree of multicasting (on the contrary, bidirectional streams, i.e., bidirectional streams such as TCP/IP streams, are increasingly becoming back-and-front to networks as the number of users increases). You will be able to create more traffic logjams). Multicasting is an important performance within server centers because it responds to the growing need of Internet users (and indeed the world's population) to communicate one-to-many or many-to-many. Because it allows the system to do In other words, as in FIG. 16 , the examples discussed here streaming interactive video recursion and multicasting are the tip of the big iceberg.
In one embodiment, the various functional modules illustrated herein and the associated steps include hardware logic by an application specific integrated circuit (ASIC) or any combination of programmed computer components and user-defined hardware components to perform the steps. This can be done by special hardware components.
In one embodiment, the module may be implemented with a programmable digital signal processor ("DSP"), such as Texas Instruments' TMS320x architecture (eg, TMS320C6000, TMS320C5000, ..., etc.). Various DSPs could be used while still following this fundamental theory.
An embodiment may include several steps as described above. The steps may be implemented so that a general-purpose or special-purpose processor may perform certain steps by means of machine-executable instructions. Various components that are not related to these basic principles, such as computer memories, hard drives, and input devices, which are related to these basic principles, are omitted from the drawings in order to avoid obscuring their proper aspects.
Problems of the disclosed subject matter may also be provided as machine-readable media for storing machine-executable instructions. Machine-readable media include flash memory, optical disks, CD-ROMs, DVD-ROM RAMs, EPROMs, EEPROMs, magnetic or optical cards, propagation media, and other suitable storage for electronic instructions. machine-readable media in the form of, and the like. For example, the present invention relates to a requesting computer (eg, a server) from a remote computer (eg, a server) by way of a data signal contained in a carrier wave or other propagating media via a communication link (eg, modem, network connection). For example, it may be downloaded as a computer program that can be transmitted to a client).
It may also be provided as a computer program product comprising a machine readable medium having stored thereon instructions, wherein the element in question of the disclosed subject matter can be used to program a computer to perform an operator of a sequence (eg, a processor or its electronic devices outside). Alternatively, the operation may be performed in a combination of hardware and software. Machine-readable media may include floppy diskettes, optical disks, magneto-optical disks, ROM, RAM, EPROM, EEPROM, magnetic or optical cards, radio wave media or other forms of machine-readable media suitable for storing electrical instructions. media may include, but are not limited to. For example, a problem element of the disclosed subject matter may be downloaded as a computer program product, which program may be downloaded via a communication link to a remote computer or for requesting processing according to the manner of a data signal contained in other propagated media or carrier waves via a communication link. may be transmitted to an electronic device (eg, a modem or network connection).
Additionally, although problems of the disclosed subject matter have been described in conjunction with specific embodiments, many modifications and permutations are well set forth within the scope of the present disclosure. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
41 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9055066B2 | Cited by | United States of America | Applicant |
703 members in 17 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 11999475 | United States of America | – | |
| 99947507 | United States of America | A | |
| 99947507 | United States of America | A | |
| 2007999475 | – | – | – |
| US20070999475 | – | – | – |
Members703
| Document | Office | Kind | |
|---|---|---|---|
| US2004111755A1 | United States of America | A1 | |
| US2009118017A1 | United States of America | A1 | |
| US2009118018A1 | United States of America | A1 | |
| US2009118019A1 | United States of America | A1 | |
| US2009119729A1 | United States of America | A1 | |
| US2009119730A1 | United States of America | A1 | |
| US2009119731A1 | United States of America | A1 | |
| US2009119736A1 | United States of America | A1 | |
| US2009119737A1 | United States of America | A1 | |
| US2009119738A1 | United States of America | A1 | |
| US2009124387A1 | United States of America | A1 | |
| US2009125961A1 | United States of America | A1 | |
| US2009125967A1 | United States of America | A1 | |
| US2009125968A1 | United States of America | A1 | |
| AU2008333797A1 | Australia | A1 | |
| AU2008333798A1 | Australia | A1 | |
| AU2008333799A1 | Australia | A1 | |
| AU2008333800A1 | Australia | A1 | |
| AU2008333801A1 | Australia | A1 | |
| AU2008333802A1 | Australia | A1 | |
| AU2008333803A1 | Australia | A1 | |
| AU2008333804A1 | Australia | A1 | |
| AU2008333821A1 | Australia | A1 | |
| AU2008333880A1 | Australia | A1 | |
| AU2008333881A1 | Australia | A1 | |
| CA2707576A1 | Canada | A1 | |
| CA2707578A1 | Canada | A1 | |
| CA2707579A1 | Canada | A1 | |
| CA2707583A1 | Canada | A1 | |
| CA2707605A1 | Canada | A1 | |
| CA2707606A1 | Canada | A1 | |
| CA2707607A1 | Canada | A1 | |
| CA2707608A1 | Canada | A1 | |
| CA2707609A1 | Canada | A1 | |
| CA2707610A1 | Canada | A1 | |
| CA2707674A1 | Canada | A1 | |
| CA2761151A1 | Canada | A1 | |
| WO2009073792A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073793A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073795A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073796A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073797A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073798A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073799A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073800A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073801A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073802A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009073819A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2008335471A1 | Australia | A1 | |
| CA2707696A1 | Canada | A1 | |
| WO2009076172A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2009076177A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2009196516A1 | United States of America | A1 | |
| US2009213927A1 | United States of America | A1 | |
| US2009213935A1 | United States of America | A1 | |
| US2009215531A1 | United States of America | A1 | |
| US2009215540A1 | United States of America | A1 | |
| US2009220001A1 | United States of America | A1 | |
| US2009220002A1 | United States of America | A1 | |
| US2009225220A1 | United States of America | A1 | |
| US2009225828A1 | United States of America | A1 | |
| US2009225863A1 | United States of America | A1 | |
| US2009228946A1 | United States of America | A1 | |
| WO2009076172A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AU2010202242A1 | Australia | A1 | |
| US2010166056A1 | United States of America | A1 | |
| US2010166058A1 | United States of America | A1 | |
| US2010166062A1 | United States of America | A1 | |
| US2010166063A1 | United States of America | A1 | |
| US2010166064A1 | United States of America | A1 | |
| US2010166065A1 | United States of America | A1 | |
| US2010166066A1 | United States of America | A1 | |
| US2010166068A1 | United States of America | A1 | |
| US2010167809A1 | United States of America | A1 | |
| US2010167816A1 | United States of America | A1 | |
| EP2218224A1 | European Patent Office (EPO) | A1 | |
| EP2225006A1 | European Patent Office (EPO) | A1 | |
| KR20100098668A | Republic of Korea | A | |
| EP2227728A1 | European Patent Office (EPO) | A1 | |
| EP2227745A1 | European Patent Office (EPO) | A1 | |
| EP2227747A1 | European Patent Office (EPO) | A1 | |
| EP2227748A1 | European Patent Office (EPO) | A1 | |
| EP2227752A1 | European Patent Office (EPO) | A1 | |
| EP2227901A1 | European Patent Office (EPO) | A1 | |
| EP2227903A1 | European Patent Office (EPO) | A1 | |
| EP2227904A1 | European Patent Office (EPO) | A1 | |
| EP2227905A2 | European Patent Office (EPO) | A2 | |
| KR20100101608A | Republic of Korea | A | |
| KR20100101637A | Republic of Korea | A | |
| EP2229224A1 | European Patent Office (EPO) | A1 | |
| EP2229775A1 | European Patent Office (EPO) | A1 | |
| KR20100102625AThis record | Republic of Korea | A | |
| CA2756299A1 | Canada | A1 | |
| CA2756309A1 | Canada | A1 | |
| CA2756328A1 | Canada | A1 | |
| CA2756331A1 | Canada | A1 | |
| CA2756338A1 | Canada | A1 | |
| CA2756458A1 | Canada | A1 | |
| CA2756681A1 | Canada | A1 | |
| CA2756686A1 | Canada | A1 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Application deemed withdrawn, e.g. because no request for examination was filed or no examination fee was paidWithdrawnWITN | WITN | |
| Notification of change of applicantN231 | N231 |
Numbers
- Publication
- 1020100102625
- Publication, DOCDB
- 20100102625
- Publication, EPODOC
- KR20100102625
- Application
- 1020107014296
- Application, DOCDB
- 20107014296
- Application, EPODOC
- KR20107014296
Titles4
- Korean
- 스트리밍 인터랙티브 비디오를 사용하는 가상 이벤트를 호스팅 및 브로드캐스팅하는 방법
- English
- HOSTING AND BROADCASTING VIRTUAL EVENTS USING STREAMING INTERACTIVE VIDEO
- Unlabeled
- 스트리밍 인터랙티브 비디오를 사용하는 가상 이벤트를 호스팅 및 브로드캐스팅하는 방법{HOSTING AND BROADCASTING VIRTUAL EVENTS USING STREAMING INTERACTIVE VIDEO}
- Unlabeled
- HOSTING AND BROADCASTING VIRTUAL EVENTS USING STREAMING INTERACTIVE VIDEO
Classification
- CPC, 46
- H04N21/43074
- A63F13/358
- A63F13/12
- A63F2300/402
- A63F2300/407
- A63F2300/552
- A63F2300/572
- A63F2300/577
- A63F2300/69
- H04N7/106
- H04N21/2143
- H04N21/233
- H04N21/2343
- H04N21/2381
- H04N21/439
- H04N21/4781
- H04N21/6125
- H04N21/6377
- H04N21/6405
- H04N21/658
- H04N21/6587
- H04W28/16
- H04W84/12
- H04L65/403
- H04N19/172
- H04N19/169
- H04N19/61
- H04N19/107
- H04N19/132
- H04N19/146
- H04N19/436
- H04N19/188
- H04L65/765
- H04L65/611
- H04L67/131
- A63F13/30
- A63F13/355
- A63F13/335
- H04N21/4307
- A63F13/52
- A63F9/24
- A63F2009/245
- A63F2009/247
- A63F2009/2488
- H04L67/10
- H04N21/262
- IPC, 2
- G06F17 00
- G06Q50 00